When Structure Helps Or Hurts Search¶
A physical search method does more than choose its next experiment. It also chooses how candidate interventions are represented: which differences count as nearby, which patterns are treated as related, and which observations are expected to transfer.
CCT proposes that this search grammar should be treated as a measurable design variable. A useful representation can lower discovery burden by directing scarce probes toward better programs. A mismatched representation can create negative transfer by making the wrong programs look related.
The Prospective Test¶
The benchmark compared a fixed 19-feature timing-and-coordination representation with raw and simpler alternatives across 36 new cases in three standard nonlinear model families:
- FitzHugh–Nagumo;
- Landau–Lifshitz–Gilbert;
- a zero-shot controlled Kuramoto ring.
Every workflow received the same balanced 64-policy menu, the same four starting panels, equal tuning and observation budgets, and between 4 and 16 total evaluations. The search decisions were frozen before the fresh response surfaces were opened.
What The Benchmark Found¶
The structural representation produced a clear family-specific advantage against the prespecified weighted-Hamming comparator in Landau–Lifshitz–Gilbert search. Its normalized regret-area difference was -0.03774, with a 95% interval of [-0.06819, -0.01261], and all 12 family cases favored the structural workflow.
The same map did not transfer uniformly. The FitzHugh–Nagumo estimate was near zero and uncertain. The Kuramoto point estimate was adverse and its policy ordering was unstable across independent seed halves. The family-balanced contrast was +0.00636, with a 95% interval of [-0.02190, +0.03600], so the benchmark did not establish one aggregate advantage across all three systems.
The experiment also showed that representation changes the evidence acquired. The structural and comparator workflows made 1,697 different next-probe choices, and a disjoint candidate-menu stress changed the aggregate point-estimate direction. Candidate-menu construction is therefore part of the search object, not neutral plumbing.
What This Adds To CCT¶
The earlier limited-probe pilot established the first direct case in which CCT's structural organization improved prospective experiment selection under a matched comparison. This larger benchmark asks the next question: where is that search grammar trustworthy?
The result extends CCT's regime-discovery program in an important direction:
Search grammars have regimes of validity. Their usefulness can be measured prospectively rather than assumed.
That turns representation design into a scientific problem. The relevant target is no longer simply a richer description of interventions. It is a representation whose useful range can be predicted before the search spends its scarce probes.
The Next Object¶
The immediate next step is a simpler prospective validity diagnostic. It should decide, from information available before or early in a campaign, when to use structural, raw, hybrid, or family-specific search. It should then be tested on a genuinely fresh model family or controlled physical exposure, with candidate-menu construction included in the frozen design.
The completed manuscript and replay package provide the detailed methods, comparisons, sensitivity analyses, and frozen-result lineage for this benchmark. A later selected replication release can expose that full result surface alongside the existing public simulation chains.
See What CCT Has Built and Opened · See Research Methods and Infrastructure · See the Empirical Outlook