Four ways to write experience into weights

Weight-level self-improvement on Qwen3-4B-Instruct, LoRA r=16 on q/v.

All four consume the same executor-verified experience and are budget-matched — same training examples, same gradient steps. Only the mechanism differs.

Forgetting is pass@1 on a held-out capability probe (base 0.847). Gain is pass@1 on the held-out surface (base 0.100). KL displacement is measured against the base model on the same probe.