Every number below is verbatim from the runs — 3 seeds each. The discipline: a judgment needs both an effect-size threshold and a random-level check, never a p-value alone.
Main result: exact decomposition of unseen composites
Trained on 4 primitives only, never on any composite. Tested on 12 unseen depth-2 composites.
Anti-cheating arm: does the search do the work, or the learned primitives?
Same composition search, primitives replaced by a randomly-initialised library.
Anti-forgetting: the library as a shield
Forgetting after new tasks — lower is better.
Robustness: model nonlinearity gain vs. true gain
True nonlinearity gain is 3.0. Sweeping the model's assumed gain away from the truth.