A fixed notebook for most layers, a full archive for a few
0
total cache memory vs context length (48-layer model, illustrative)
Illustrative: 48 layers = the 8-layer pattern above × 6. Attention layer KV = 4 KiB/token (8 KV heads × 128 dim × K,V × bf16, Llama-3-8B-like). Linear layer state = 1 MiB, fixed. Ratios are the point: 1 attention layer in 4 → ¼ the slope; 1 in 8 → ⅛.
recall check: a needle written at token 2,000 (qualitative, illustrative)
pure linear — only the board
hybrid — plus an attention layer