Results & anti-cheating

Every number below is verbatim from the runs — 3 seeds each. The discipline: a judgment needs both an effect-size threshold and a random-level check, never a p-value alone.

Main result: exact decomposition of unseen composites

Trained on 4 primitives only, never on any composite. Tested on 12 unseen depth-2 composites.

COMPOSE_learned random level 0.083 same-pipeline permutation null 0.152

Anti-cheating arm: does the search do the work, or the learned primitives?

Same composition search, primitives replaced by a randomly-initialised library.

paired difference
+0.797 ± 0.066
95% CI lower bound +0.668
seeds where learned > random
3 / 3
every seed, not on average
COMPOSE_random level
0.097
indistinguishable from guessing

Anti-forgetting: the library as a shield

Forgetting after new tasks — lower is better.

with operator library fine-tune, no library

Robustness: model nonlinearity gain vs. true gain

True nonlinearity gain is 3.0. Sweeping the model's assumed gain away from the truth.

show the numbers