From 2K to a Million Tokens: How Context Windows Actually Grew
A sequel to the four posts on how a transformer remembers — KV cache · Linear attention · Kernelization · Recurrent memory. You don’t need them, but a few sections below lean on them.
GPT-3 shipped with 2,048 tokens of context in 2020. Five years later, 1M is a normal number on a spec sheet, and a couple of models claim 4M or 10M. That’s a 500× jump in five years, and it’s tempting to read it as “GPUs got bigger.” They did, but that’s maybe a fifth of the story.
Most write-ups of this list the methods — RoPE scaling, FlashAttention, ring attention, GQA, sparse attention, hybrids — and stop there. A list tells you what people did. It doesn’t tell you why each one works, or why you need all of them at once. That’s what this post is for.
Here’s the frame I’ll use. A context window is held in place by five things: positions (the model only trusts distances it trained on), compute (attention is $L^2$), memory (activations in training, the KV cache at inference), data (you need text where something far back matters), and evaluation (does the model actually use the window, or just accept it). Every method below removes exactly one of those. Get rid of one, and the next one becomes the wall.
Every panel is live — press play, drag things, switch the language and the demos follow.
「transformer 怎么记东西」四篇的续集 —— KV 缓存 · 线性注意力 · 核化 · 递归记忆。不看也行,但下面有几节会借它们的光。
2020 年的 GPT-3,上下文是 2,048 个 token。五年之后,1M 在规格表上已经是个很普通的数字,还有几个模型号称 4M、10M。五年涨了 500 倍,很容易把它理解成「GPU 变大了」。确实变大了,但那大概只占这个故事的五分之一。
大多数讲这件事的文章是把方法列一遍——RoPE 缩放、FlashAttention、ring attention、GQA、稀疏注意力、混合架构——然后就结束了。列表告诉你大家做了什么,没告诉你每一招为什么管用,也没告诉你为什么必须几招一起上。这篇就是讲这个的。
我用的框架是这样:一个上下文窗口被五样东西卡住——位置(模型只信它训练时见过的距离)、算力(注意力是 $L^2$)、显存(训练时是激活,推理时是 KV cache)、数据(得有「很远之前的东西很重要」的文本)、以及评测(模型是真的用上了这个窗口,还是只是收得下)。下面每一招都恰好拆掉其中一个。拆掉一个,下一个就变成了那堵墙。
每块面板都是活的——按播放、拖滑块,切换语言,演示也会跟着切。
0. The window is just what the model practiced on
Start with the most boring fact, because everything else hangs off it. Pretraining doesn’t feed documents one at a time. The pipeline tokenizes everything, glues documents together with an end-of-document token, and chops the stream into rows of exactly $L$ tokens. That’s called packing.
So the model sees position 0, position 1, …, position $L-1$, billions of times each. It sees position $L$ exactly zero times. There’s no rule in the architecture that says “stop at $L$” — the weights would happily run on a longer input. The window is simply the set of positions the model ever got a gradient on. And anything longer than $L$ got cut in half, so the model also never learned to connect the two halves.
0. 窗口不过是模型练过的那段
先从最无聊的一个事实讲起,因为后面所有东西都挂在它上面。预训练不是一篇一篇地喂文档。数据流水线先把所有东西切成 token,用文档结束符把文档首尾粘起来,再把这条长流切成每行正好 $L$ 个 token 的样本。这叫 packing(打包)。
所以模型见过位置 0、位置 1……位置 $L-1$,每个都见过几十亿次。位置 $L$,一次都没见过。架构里并没有哪条规则写着「到 $L$ 就停」——你喂更长的输入,权重照样能算。所谓窗口,只是模型拿到过梯度的那些位置的集合。而比 $L$ 长的文档会被拦腰切断,所以模型也从来没学过怎么把两半连起来。
Two things to take from that panel. First, the window is a training fact, not an architecture fact — which is good news, because it means you can move it. Second, the triangle. Inside each row, token $i$ scores itself against every earlier token, $L(L+1)/2$ pairs. Double $L$ and you quadruple the work. That’s why nobody just pretrains at 1M from step one.
这块面板要带走两件事。第一,窗口是一个训练层面的事实,不是架构层面的——这是好消息,说明它能挪。第二,那个三角形。每一行里,第 $i$ 个 token 要和它前面所有 token 打分,一共 $L(L+1)/2$ 对。$L$ 翻倍,活儿翻四倍。这就是为什么没人从第一步就在 1M 上预训练。
1. Positions: a set of dials, and the slow ones run out
Almost every current LLM encodes position with RoPE. Split each query and key vector into pairs of dimensions. Pair $i$ of the token at position $m$ gets rotated by the angle $m\,\theta_i$, where
\[\theta_i = \text{base}^{-2i/d}\]When a query at position $m$ meets a key at position $n$, the rotations partially cancel and only $(m-n)\,\theta_i$ survives in the dot product. So attention sees relative distance. Nice.
The intuition I like: each pair is a dial. With $d=128$ and base 10,000, the fastest dial turns 1 radian (57°) per token — it wraps around every ~6 tokens and is great at telling neighbors apart. The slowest dial turns about $10^{-4}$ radians per token; one full turn takes ~54,400 tokens.
Now train at $L = 4096$. The fast dials go round and round and see every angle thousands of times. The slow dials only ever sweep a sliver of their circle. Then you hand the model a 16K input, and the slow dials point at angles the network has literally never seen. The attention scores that depend on them go weird, and perplexity explodes. This happens before memory is even a concern. Drag the position past the training line and watch the hands go red.
1. 位置:一组表盘,慢的那几个先用完
现在几乎所有 LLM 都用 RoPE 编码位置。把每个 query、key 向量按维度两两分组,位置 $m$ 上那个 token 的第 $i$ 组转过的角度是 $m\,\theta_i$,其中
\[\theta_i = \text{base}^{-2i/d}\]位置 $m$ 的 query 碰上位置 $n$ 的 key,两边的旋转一抵消,点积里只剩 $(m-n)\,\theta_i$。所以注意力看到的是相对距离。很漂亮。
我喜欢的理解方式是:每一组是一个表盘。$d=128$、base 10,000 时,最快的表盘每个 token 转 1 弧度(57°),大约 6 个 token 就转一圈,特别擅长区分相邻的 token。最慢的表盘每个 token 只转大约 $10^{-4}$ 弧度,转一整圈要 54,400 个 token 左右。
现在在 $L = 4096$ 上训练。快表盘一圈一圈地转,每个角度都见过成千上万次。慢表盘从头到尾只扫过自己圆周上的一小段。然后你喂给模型一个 16K 的输入,慢表盘就会指向网络从来没见过的角度。依赖这些表盘的注意力分数开始乱跳,困惑度直接爆掉。这一切发生在显存成为问题之前。把位置拖过训练线,看指针变红。
Every fix is one question: which dials do you slow down?
Once you see it as dials, all the position tricks collapse into the same question.
Position Interpolation (Chen et al., 2023) is the blunt answer: slow every dial down by the extension factor $s$, i.e. use position $m/s$ instead of $m$. Now no dial ever leaves its trained range. The cost: the fast dials also slow down by $s$, so two neighboring tokens now look $s$ times closer than they used to. Local word order gets blurry. A short fine-tune (~1,000 steps) mostly repairs it, and this is how early LLaMA got to 32K.
NTK-aware scaling (bloc97, 2023) noticed the fast dials were never the problem — they already wrap many times inside $L$. So instead of dividing positions, raise the base: $\text{base}’ = \text{base}\cdot s^{d/(d-2)}$. The fast dials barely change, the slowest one gets slowed by about $s$, and everything in between is smoothly in between.
Bigger base (ABF) is the same idea done during training rather than at inference: crank the base way up (10,000 → 500,000 in Llama 2 Long and Llama 3; 1,000,000 in Code Llama and Qwen3), then keep training on long data so the model gets used to its new slow dials.
YaRN (Peng et al., 2023) makes the cut explicit. For each pair, ask how many full turns it makes inside the training window, $r_i = L/\lambda_i$. If it wraps more than ~32 times, leave it alone — it’s seen everything. If it doesn’t even finish one turn, interpolate it fully. In between, ramp linearly. Plus one more piece I’ll get to in a second (temperature). It needs ~0.1% of pretraining data to fine-tune, and DeepSeek-V3 and the Qwen family both use it.
LongRoPE (Microsoft, 2024) says: why guess the curve? Search for a separate rescale factor per dimension with an evolutionary search, leave the first few positions uninterpolated, and extend in stages. That’s how Phi-3 got its 128K variants, and the paper reports going past 2M.
In the panel below the bars are dials, fast on the left, slow on the right. Flip between methods and watch which ones get stretched.
每一种修法,其实都在回答一个问题:哪些表盘该调慢?
一旦把它看成表盘,所有位置方面的招数就会收敛到同一个问题上。
Position Interpolation(位置插值,Chen 等,2023) 是最粗暴的答案:把所有表盘按扩展倍数 $s$ 一起调慢,也就是用 $m/s$ 代替 $m$。这样任何表盘都出不了训练范围。代价是快表盘也慢了 $s$ 倍,相邻两个 token 看起来比以前近了 $s$ 倍,局部词序就糊了。短暂微调(大约 1,000 步)基本能修回来,早期 LLaMA 就是这么上的 32K。
NTK-aware 缩放(bloc97,2023) 注意到快表盘从来不是问题——它们在 $L$ 之内早就转了很多圈。所以不去除位置,而是把 base 调大:$\text{base}’ = \text{base}\cdot s^{d/(d-2)}$。快表盘几乎不动,最慢的那个大约慢 $s$ 倍,中间的平滑过渡。
调大 base(ABF) 是同一个想法,只不过放在训练里做而不是推理时做:把 base 拉得很高(Llama 2 Long 和 Llama 3 从 10,000 拉到 500,000;Code Llama 和 Qwen3 拉到 1,000,000),然后在长数据上继续训练,让模型习惯新的慢表盘。
YaRN(Peng 等,2023) 把这条分界线明确画了出来。对每一组,问它在训练窗口里转了几整圈,$r_i = L/\lambda_i$。转了 32 圈以上的,别动——它什么角度都见过。连一圈都没转完的,完全插值。中间的,线性过渡。另外还有一个零件(温度),等下讲。它只需要大约 0.1% 的预训练数据来微调,DeepSeek-V3 和 Qwen 系列都在用。
LongRoPE(微软,2024) 的态度是:曲线干嘛要猜?用进化搜索给每个维度单独找一个缩放系数,开头几个位置不插值,然后分阶段扩展。Phi-3 的 128K 版本就是这么来的,论文里报告能做到 2M 以上。
下面这块面板里,每根柱子是一个表盘,左边快、右边慢。在几种方法之间切换,看哪些被拉伸了。
2. Even with perfect positions, the softmax gets diluted
Say you’ve fixed positions completely. There’s a second, quieter problem: more keys means more competitors.
One query, one key that actually matters, and $n-1$ distractors with roughly random logits. The softmax gives the right key a weight of about
\[\frac{e^{z^*}}{e^{z^*} + (n-1)\cdot \mathbb{E}[e^{z}]}\]which goes like $1/n$. At 4K the right key wins. At 1M it’s still the single highest score, but it’s holding a sliver of the total weight, and the output is mostly an average of noise. The entropy of the attention distribution climbs like $\log n$.
The fix is embarrassingly small: sharpen the logits as the context grows. Multiply them by something that rises with $\log n$. YaRN’s version is to scale $q$ and $k$ each by $0.1\ln s + 1$. Llama 4 applies a temperature scaling at inference for long inputs. Same idea everywhere: keep the peak a peak.
2. 就算位置完美,softmax 也会被稀释
假设位置问题已经彻底解决。还有第二个更安静的问题:key 越多,竞争者越多。
一个 query,一个真正相关的 key,再加 $n-1$ 个 logit 差不多随机的干扰项。softmax 分给正确那个 key 的权重大约是
\[\frac{e^{z^*}}{e^{z^*} + (n-1)\cdot \mathbb{E}[e^{z}]}\]它大致按 $1/n$ 往下掉。4K 的时候正确的 key 稳赢。到 1M,它的分数依然是全场最高,但只分到总权重的一小丝,输出基本是一堆噪声的平均。注意力分布的熵按 $\log n$ 往上涨。
修法小得令人尴尬:上下文越长,logit 越要锐化。乘一个随 $\log n$ 增长的系数就行。YaRN 的版本是把 $q$ 和 $k$ 各乘 $0.1\ln s + 1$。Llama 4 在推理长输入时做一个温度缩放。到处都是同一个想法:让尖峰保持是尖峰。
3. Or don’t stretch at all: remap the distances
All of section 1 changes the dials. There’s a sneakier option that doesn’t touch them: make sure the model never sees a distance it didn’t train on.
Dual Chunk Attention (An et al., 2024) cuts the sequence into chunks smaller than the training length. Inside a chunk, relative positions are the normal ones, so nearby tokens keep their exact order. Across chunks, the query gets a fixed large position index, so the distance gets capped near the edge of the trained range — the model sees “far, as far as I know” rather than “1.3 million tokens away”. It loses the exact distance between far-apart tokens, which it never knew how to use anyway. A third rule handles the neighboring chunk so locality doesn’t break at the boundary. Every distance the model computes lands inside its trained range. No fine-tuning needed.
This is not a toy. Qwen2.5-1M trains up to 256K and then uses DCA at inference to go 4× further, to 1M.
Llama 4 took the most radical version: in its iRoPE design some layers have no positional encoding at all. Those layers can only match by content, so there’s no dial to run out. The remaining RoPE layers attend locally, where distances stay small. That’s part of how Scout claims 10M.
3. 或者干脆不拉伸:把距离重新映射
第 1 节的所有方法都在改表盘。还有一个更狡猾的选项,表盘一点不动:保证模型永远看不到它没训练过的距离。
Dual Chunk Attention(双块注意力,An 等,2024) 把序列切成比训练长度短的块。块内,相对位置照常算,近处的 token 保持精确顺序。跨块时,query 换成一个固定的大位置编号,距离就被封顶在训练范围的边上——模型看到的是「就我所知最远的那么远」,而不是「130 万个 token 之外」。它丢掉的是远处两个 token 之间的精确距离,而这个它本来也从没学会怎么用。第三条规则专门处理相邻的块,免得局部性在块边界上断掉。模型算出的每一个距离都落在训练范围内。不用微调。
这不是玩具。Qwen2.5-1M 训练到 256K,推理时用 DCA 再往外推 4 倍,到 1M。
Llama 4 走得最极端:它的 iRoPE 设计里,有些层根本没有位置编码。这些层只能靠内容匹配,自然也就没有会用完的表盘。剩下的 RoPE 层只看局部,距离始终很小。Scout 号称 10M,这是原因之一。
4. The bill, part one: never write down the $L \times L$ matrix
Positions handled. Now the arithmetic. One head at 128K tokens has a score matrix of $131{,}072^2$ entries — 32 GiB in bf16. Per head, per layer. You can’t store that, and you don’t need to.
FlashAttention (Dao et al., 2022) computes the exact same result in tiles. Load a block of queries into the GPU’s tiny, fast on-chip SRAM, stream key/value tiles past it, and never write the full score matrix to the big, slow HBM. The catch is that softmax needs the max and the sum over the whole row, and you only have one tile at a time. The trick is the online softmax: keep a running max $m$ and running sum $\ell$, and when a new tile brings a bigger max, rescale what you’ve accumulated so far:
\[\begin{aligned} m' &= \max(m,\ \max_j s_j) \\ \ell' &= e^{m-m'}\ell + \textstyle\sum_j e^{s_j - m'} \\ O' &= e^{m-m'}O + \textstyle\sum_j e^{s_j-m'} v_j \end{aligned}\]Divide by $\ell$ at the end and you get the exact softmax. Watch the numbers below — the running output gets corrected every time a bigger score shows up, and the final answer matches the brute-force one to floating-point precision.
Memory goes from $L^2$ to $L$. Compute is still $L^2$ — FlashAttention makes long context possible, not cheap. Keep that in mind; the rest of the post is about the compute.
4. 账单之一:永远别把 $L \times L$ 矩阵写下来
位置搞定了,现在算账。一个头在 128K 时的分数矩阵有 $131{,}072^2$ 个元素——bf16 下 32 GiB。这还只是一个头、一层。你存不下,而且也不需要存。
FlashAttention(Dao 等,2022) 分块计算出完全相同的结果。把一块 query 装进 GPU 那块又小又快的片上 SRAM,让 key/value 一块一块流过去,完整的分数矩阵从来不写到又大又慢的 HBM 里。麻烦在于 softmax 需要整行的最大值和总和,而你手里每次只有一块。诀窍是 online softmax:维护一个运行中的最大值 $m$ 和运行中的和 $\ell$,新来的块带来更大的最大值时,把已经累积的部分重新缩放一下:
\[\begin{aligned} m' &= \max(m,\ \max_j s_j) \\ \ell' &= e^{m-m'}\ell + \textstyle\sum_j e^{s_j - m'} \\ O' &= e^{m-m'}O + \textstyle\sum_j e^{s_j-m'} v_j \end{aligned}\]最后除以 $\ell$,就是精确的 softmax。看下面的数字——每当出现更大的分数,累积的输出就被修正一次,最终结果和暴力算出来的一致到浮点精度。
显存从 $L^2$ 变成 $L$。算力还是 $L^2$——FlashAttention 让长上下文变得可能,没让它变得便宜。记住这一点,后面整篇都在跟算力较劲。
5. The bill, part two: one sequence, many GPUs
FlashAttention gets the score matrix out of memory, but activations still grow with $L$, and at some length one GPU just can’t hold the sequence. So split it.
Ring attention (Liu, Zaharia, Abbeel, 2023) puts the GPUs in a ring. Each one keeps its own chunk of queries and holds one chunk of keys/values. Every step, each GPU computes its queries against the KV it currently has, then passes that KV to its neighbor — and the send overlaps with the compute, so the network is mostly free. After $N$ steps, every query chunk has met every KV chunk. No GPU ever held the whole thing. Max length scales with the number of GPUs.
There’s a wrinkle, and it’s a fun one. With a causal mask, the GPU holding the first chunk has almost nothing to do (its queries can’t look at anything later), and the GPU holding the last chunk does all the work. Everyone waits for the slowest one. The fix is like dealing cards from both ends of the deck: split into $2N$ chunks and give GPU $i$ chunks $i$ and $2N-1-i$. Now every GPU gets one cheap chunk and one expensive chunk. Llama 3 does exactly this split.
The other family, DeepSpeed-Ulysses, does an all-to-all right before attention: swap from “each GPU has a slice of the sequence for all heads” to “each GPU has the whole sequence for a few heads”, run attention locally, swap back. Simple, but you can’t split wider than your number of heads.
5. 账单之二:一条序列,很多张卡
FlashAttention 把分数矩阵请出了显存,但激活还是随 $L$ 增长,长到一定程度,一张卡就是装不下这条序列。那就切开。
Ring attention(Liu、Zaharia、Abbeel,2023) 把 GPU 排成一个环。每张卡守着自己那一段 query,手里拿着一段 key/value。每一步,每张卡用自己的 query 去算手里那段 KV,然后把 KV 传给下一张卡——传输和计算重叠,网络几乎是白送的。$N$ 步之后,每一段 query 都见过了每一段 KV。没有任何一张卡拿过整条序列。最大长度随卡数线性增长。
这里有个挺好玩的小坑。加上因果掩码之后,拿着第一段的卡几乎没事干(它的 query 不能看后面的任何东西),拿着最后一段的卡干了全部的活。所有人都在等最慢的那个。修法就像从一副牌的两头发牌:切成 $2N$ 段,第 $i$ 张卡拿第 $i$ 段和第 $2N-1-i$ 段。这样每张卡都是一段轻的加一段重的。Llama 3 用的就是这种切法。
另一派是 DeepSpeed-Ulysses,在注意力之前做一次 all-to-all:从「每张卡有一段序列、所有头」换成「每张卡有整条序列、几个头」,本地算完注意力,再换回来。简单,但切分的份数不能超过头数。
6. Store less per token: GQA and MLA
Now inference. As the KV cache post showed, every token leaves a key and a value behind in every layer, and that pile grows linearly. Do the arithmetic for Llama 3 70B: 80 layers × 8 KV heads × 128 dims × 2 (K and V) × 2 bytes = 320 KiB per token. At 128K that’s 40 GiB for one sequence. At 1M ($2^{20}$ tokens), 320 GiB. The weights don’t run out first. The cache does.
So the architecture question becomes: how little can each token leave behind?
GQA (and its extreme, MQA) lets several query heads share one key/value head. Llama 3 70B has 64 query heads but only 8 KV heads, so the cache is 8× smaller than full multi-head attention would be. Quality barely notices, because the heads were storing a lot of redundant stuff anyway.
MLA (DeepSeek-V2/V3) goes further: instead of storing keys and values at all, store one compressed latent vector per token — 512 numbers — and expand it back into per-head keys and values when you need them. Better: the up-projection can be folded into the query projection, so attention runs directly in the latent space. One snag: RoPE rotates by position, and a position-dependent rotation can’t be folded into a fixed matrix. So MLA carries a separate tiny 64-dim rotary key alongside the latent. 576 numbers per token per layer, ~69 KiB per token for DeepSeek-V3. Compare to 320 KiB above.
6. 每个 token 少留点东西:GQA 和 MLA
现在讲推理。KV 缓存那篇讲过,每个 token 在每一层都会留下一个 key 和一个 value,这堆东西线性增长。给 Llama 3 70B 算一笔账:80 层 × 8 个 KV 头 × 128 维 × 2(K 和 V)× 2 字节 = 每个 token 320 KiB。128K 时就是 40 GiB,这只是一条序列。到 1M($2^{20}$ 个 token),320 GiB。先撑不住的不是权重,是缓存。
所以架构上的问题变成了:每个 token 最少能留下多少东西?
GQA(以及它的极端版本 MQA)让好几个 query 头共用一个 key/value 头。Llama 3 70B 有 64 个 query 头,但只有 8 个 KV 头,缓存比完整多头注意力小 8 倍。质量几乎感觉不到,因为那些头本来存了大量冗余。
MLA(DeepSeek-V2/V3)更进一步:干脆不存 key 和 value,每个 token 只存一个压缩过的潜向量——512 个数——需要时再展开回每个头的 key 和 value。更妙的是,这个上投影可以并进 query 的投影里,注意力直接在潜空间里算。有一个小麻烦:RoPE 是按位置旋转的,一个随位置变化的旋转没法并进一个固定矩阵。所以 MLA 在潜向量旁边额外带一个 64 维的小旋转 key。每个 token 每层 576 个数,DeepSeek-V3 下大约 69 KiB 一个 token。对比上面的 320 KiB。
7. Most layers only need to look nearby
GQA and MLA shrink each token’s footprint, but every layer still keeps every token. Here’s the observation that breaks that: most of what a layer needs is close by. Syntax, the current sentence, the last few lines of code.
So give most layers a sliding window: they only attend to the last $W$ tokens, and their cache stops growing at $W$. Stack them and information still travels — layer 1 sees $W$ back, layer 2 sees $2W$ back through layer 1, and so on. Mistral 7B used a 4K window everywhere.
But “information can travel $kW$ tokens” is not the same as “the model can exactly recall a phone number from 500K tokens ago”. Squeezing it through a chain of windows blurs it. So the modern recipe interleaves: Gemma 2 alternates local and global layers 1:1; Gemma 3 goes 5 local (window 1024) to 1 global, which makes the cache roughly a sixth of what it would be; gpt-oss alternates dense and 128-token banded layers. The global layers do the long-distance recall, the local ones do everything else for almost nothing.
The sink toggle in the panel is a side story worth knowing. StreamingLLM found that models dump a lot of attention onto the first few tokens — not because they matter, but because softmax has to put its weight somewhere. Evict those and a sliding-window model falls apart. Keep them and it’s fine. gpt-oss bakes a learned sink into each head for exactly this reason.
7. 大部分层只需要看附近
GQA 和 MLA 缩小了每个 token 的占用,但每一层还是保留着所有 token。打破这一点的观察是:一层需要的大部分东西都在附近。语法、当前这句话、前几行代码。
所以让大多数层用滑动窗口:只看最近的 $W$ 个 token,缓存长到 $W$ 就不再长了。叠起来信息照样能传——第 1 层往回看 $W$,第 2 层通过第 1 层往回看 $2W$,以此类推。Mistral 7B 就是每层都用 4K 窗口。
但「信息能传 $kW$ 个 token」不等于「模型能精确回忆起 50 万个 token 之前的一个电话号码」。挤过一串窗口传过来,早就糊了。所以现在的做法是交错:Gemma 2 局部层和全局层 1:1 交替;Gemma 3 是 5 层局部(窗口 1024)配 1 层全局,缓存大概只剩原来的六分之一;gpt-oss 是稠密层和 128-token 带状层交替。全局层负责远距离回忆,局部层几乎不花钱地把其余的事都干了。
面板里那个 sink 开关是个值得知道的支线故事。StreamingLLM 发现模型会把大量注意力倒在最开头几个 token 上——不是因为它们重要,而是因为 softmax 必须把权重放在某个地方。把它们踢出窗口,滑动窗口模型就崩了;留着,就没事。gpt-oss 正是为此在每个头里内置了一个可学习的 sink。
8. Read only the blocks that matter
Sliding windows decide where to look by a fixed rule: nearby. What if the model decided for itself?
That’s learned sparse attention, and the shape of it is always the same: a cheap scout and an expensive reader. The scout looks at everything, but with a tiny budget. The reader does full attention, but only where the scout points.
NSA (DeepSeek, 2025) runs three branches per query. A compressed branch squashes each block of keys into one summary token and attends to those — a blurry overview of the whole context. Those coarse scores pick the top few blocks, and a selected branch reads those blocks at full resolution. A sliding window branch always covers the recent tokens. Learned gates mix the three.
DSA in DeepSeek-V3.2 is leaner. A “lightning indexer” — a handful of tiny heads in FP8 — scores the query against every previous token, and the main attention reads only the top 2,048. The indexer is still technically $O(L)$ per query, but it’s so small that the expensive part, the real attention, is now a constant per query no matter how long the context is. Moonshot’s MoBA is in the same family: gate over block summaries, read the top-$k$ blocks.
Notice in the panel that the scout can miss. With a small budget, a relevant block sometimes doesn’t make the cut. That’s the real trade-off of sparse attention, and why these methods are trained end-to-end rather than bolted on.
8. 只读真正要紧的块
滑动窗口按一条固定规则决定看哪里:看附近。如果让模型自己决定呢?
这就是可学习的稀疏注意力,它的结构永远是同一个样子:一个便宜的侦察兵,加一个昂贵的阅读者。侦察兵什么都看,但预算极小。阅读者做完整的注意力,但只读侦察兵指给它的地方。
NSA(DeepSeek,2025) 每个 query 跑三个分支。压缩分支把每一块 key 压成一个摘要 token,对这些摘要做注意力——整个上下文的一个模糊概览。用这些粗粒度分数挑出最高的几块,选择分支以全分辨率去读这些块。滑动窗口分支永远覆盖最近的 token。三路结果用学出来的门控混合。
DeepSeek-V3.2 里的 DSA 更精简。一个「闪电索引器」——几个极小的头,跑在 FP8 上——给 query 和前面每一个 token 打分,主注意力只读分数最高的 2,048 个。索引器严格来说每个 query 仍是 $O(L)$,但它小到可以忽略,真正昂贵的那部分注意力,每个 query 的代价变成了常数,上下文多长都一样。月之暗面的 MoBA 也是这一家:对块摘要做门控,读 top-$k$ 块。
注意面板里侦察兵是会漏的。预算小的时候,相关的块有时候挑不进来。这是稀疏注意力真正的取舍,也是为什么这些方法要端到端训练,而不是事后硬接上去。
9. A fixed-size notebook for most layers
The last architectural move is the most aggressive: for most layers, don’t keep a cache at all. Keep a fixed-size state instead — the $d\times d$ board from the linear attention and recurrent memory posts. Each token writes into it, the board gets edited and decayed, and it never grows. A linear layer costs the same per token at 1M as at 1K.
The catch, which those posts spent a lot of time on: a fixed board has finite capacity. Push enough facts through it and old ones get blurred or overwritten. Exact recall of one specific line from 800K tokens ago is precisely what a fixed state is bad at.
So nobody ships pure linear. They ship hybrids: Jamba puts one attention layer per eight; MiniMax-01 puts a softmax layer after every seven lightning-attention layers, trains to 1M and runs inference to 4M; Qwen3-Next and Kimi Linear go 3:1. Memory grows with the slope of the attention fraction — a 7:1 hybrid’s cache grows at an eighth of the rate. The linear layers do the cheap bulk work. The few attention layers keep the archive.
9. 大多数层只配一个固定大小的笔记本
最后一个架构上的动作最激进:大多数层干脆不留缓存,改留一个固定大小的状态——就是线性注意力和递归记忆那两篇里那块 $d\times d$ 的板子。每个 token 往里写,板子被改写、被衰减,但永远不变大。线性层在 1M 时每个 token 的代价和 1K 时一样。
麻烦在于——那两篇花了很多篇幅讲这个——固定大小的板子容量有限。塞进去的事实够多,旧的就会被糊掉或覆盖。精确回忆起 80 万 token 之前的某一行,恰恰是固定状态最不擅长的事。
所以没人上纯线性,上的都是混合:Jamba 每八层放一层注意力;MiniMax-01 每七层 lightning attention 之后放一层 softmax,训练到 1M,推理到 4M;Qwen3-Next 和 Kimi Linear 是 3:1。显存的增长斜率就等于注意力层的占比——7:1 的混合,缓存涨速是原来的八分之一。线性层干便宜的大头活,少数几层注意力守着档案馆。
10. The recipe: train short, extend at the end
Now put the training run together. Attention cost per token grows with $L$, so you want to spend as few tokens as possible at long length. The standard recipe: do almost all of pretraining short, then stretch the window in a brief final phase.
Llama 3 405B pretrained at 8K, then grew to 128K in six stages over ~800B tokens — about 5% of its 15.6T. It only moved to the next stage once short-context scores had recovered and needle tests passed at the new length. DeepSeek-V3 went 4K → 32K → 128K with YaRN, two phases of 1,000 steps each. Qwen2.5-1M went 4K → 32K → 65K → 131K → 262K while raising the RoPE base to 10M, then used DCA at inference for the last 4× to 1M.
Two things the recipe quietly depends on. Data: most web documents are short, so long documents (books, whole code repos, long transcripts) get upsampled, and labs synthesize tasks where the answer depends on something far away. The surprising finding (Fu et al., 2024) is that a few billion tokens of well-mixed long data is enough — the skill is mostly latent and just needs to be unlocked. Short skills: if you only feed long data, short-context ability decays, so you keep mixing in high-quality short data.
Hit play and watch the cost meter. The long phase is a sliver of the tokens and a big chunk of the attention bill.
10. 配方:短着训,最后再拉长
现在把整个训练过程拼起来。注意力每个 token 的代价随 $L$ 增长,所以你希望在长序列上花尽量少的 token。标准配方:预训练几乎全程用短序列,最后用一个短暂的阶段把窗口拉长。
Llama 3 405B 在 8K 上预训练,然后分六个阶段、用大约 800B token 长到 128K——占它 15.6T 总量的 5% 左右。每进入下一阶段之前,都要确认短上下文的分数恢复了、新长度下的大海捞针测试过了。DeepSeek-V3 是 4K → 32K → 128K,用 YaRN,两个阶段各 1,000 步。Qwen2.5-1M 是 4K → 32K → 65K → 131K → 262K,同时把 RoPE base 拉到一千万,最后 4 倍到 1M 靠推理时的 DCA。
这个配方悄悄依赖两件事。数据:网上的文档大多很短,所以要上采样长文档(书、整个代码仓库、长对话记录),还要合成「答案取决于很远之前的内容」的任务。一个挺意外的发现(Fu 等,2024)是:几十亿 token 配比合理的长数据就够了——这个能力大部分是潜在的,只需要被解锁。短能力:只喂长数据,短上下文能力会退化,所以要一直掺着高质量的短数据。
按播放,盯着那个代价表。长序列阶段只占一小撮 token,却占了很大一块注意力账单。
11. Accepting 1M tokens is not the same as using them
Here’s the uncomfortable part. Everything so far makes a model accept a long input without breaking. None of it guarantees the model uses it.
The classic test is needle in a haystack: hide one sentence somewhere in a long document, ask for it back. Almost every model now aces it, and that’s the problem — finding a string that literally matches your question is the easiest thing attention does. RULER (NVIDIA, 2024) added multi-needle, multi-hop variable tracking, and aggregation tasks, and found that many models fell apart well before their advertised window. Lost in the Middle (Liu et al., 2023) showed accuracy is U-shaped over position: great at the start, great at the end, sagging in the middle. NoLiMa (2025) removed the word overlap between the question and the needle, and scores dropped hard with length.
So there are two numbers for every model, the advertised one and the effective one, and the gap between them is where the actual research is right now. (The heatmaps below are an illustrative pattern, not real scores — the shape is what matters.)
11. 收得下 1M,不等于用得上 1M
接下来是让人不太舒服的部分。到目前为止讲的所有东西,都是让模型收下一段长输入而不崩。没有一样能保证模型用上了它。
经典测试是大海捞针:在一篇长文档的某处藏一句话,让模型找回来。现在几乎所有模型都能满分,而这恰恰是问题——找到一个和问题字面一致的字符串,是注意力最擅长的事。RULER(NVIDIA,2024)加了多针、多跳的变量追踪、聚合类任务,发现很多模型在远没到宣称窗口的时候就散架了。Lost in the Middle(Liu 等,2023)发现准确率随位置呈 U 形:开头好,结尾好,中间塌。NoLiMa(2025)去掉了问题和针之间的字面重合,分数随长度掉得很厉害。
所以每个模型都有两个数字:宣传的那个,和有效的那个。两者之间的差距,正是现在真正在做研究的地方。(下面的热力图是示意的模式,不是真实分数——重要的是形状。)
12. Putting it together: how real models got to a million
Now the punchline. No single trick gets you to 1M. Each one removes one constraint, and the next one becomes binding. Look at how the actual million-token models stacked them:
- Qwen2.5-1M: progressive training to 256K with a huge RoPE base (positions), GQA (memory), sparse attention to speed up prefill (compute), and DCA plus YaRN’s attention scaling at inference for the last 4× (positions and dilution — without retraining).
- MiniMax-01: seven lightning-attention layers per softmax layer (memory and compute), trained to 1M, inference to 4M.
- Llama 4 Scout: iRoPE — NoPE global layers plus chunked local RoPE layers — and inference-time temperature scaling (positions + dilution). Trained at 256K, advertises 10M.
- DeepSeek-V3.2: MLA (memory) plus DSA sparse attention (compute). Stays at 128K, but makes each token there much cheaper.
- Gemini 1.5 shipped 1M (and 10M in research) in early 2024 and was the first to make it routine. The architecture isn’t public, so I won’t guess.
Build your own below. Slide the length, flip the components, and watch which gauge goes red first. It’s almost always the one you haven’t fixed yet.
12. 拼起来:真实的模型是怎么到一百万的
最后的结论。没有任何一招能单独把你带到 1M。 每一招拆掉一个约束,下一个约束就开始卡你。看看真正的百万级模型是怎么叠的:
- Qwen2.5-1M:渐进式训练到 256K,配上巨大的 RoPE base(位置);GQA(显存);用稀疏注意力加速预填充(算力);最后 4 倍靠推理时的 DCA 加上 YaRN 的注意力缩放(位置和稀释一起——不用重训)。
- MiniMax-01:每七层 lightning attention 配一层 softmax(显存和算力一起解决),训练到 1M,推理到 4M。
- Llama 4 Scout:iRoPE——没有位置编码的全局层,加分块的局部 RoPE 层——再加推理时的温度缩放(位置 + 稀释)。在 256K 上训练,宣称 10M。
- DeepSeek-V3.2:MLA(显存)加 DSA 稀疏注意力(算力)。窗口停在 128K,但让这 128K 里每个 token 便宜了很多。
- Gemini 1.5 在 2024 年初就上了 1M(研究版 10M),是第一个让百万上下文变成日常的。架构没公开,我就不瞎猜了。
下面自己搭一个。拖动长度,开关各个组件,看哪个仪表先变红。几乎总是你还没修的那一个。
So what actually happened between 2K and 1M
Squint at the whole thing and it’s five fixes for five walls.
Positions stopped being a wall once people saw RoPE as dials and asked which ones were actually out of range — interpolate the slow ones, leave the fast ones, sharpen the softmax, or remap distances so the model never sees a new one. Memory stopped being a wall in training when FlashAttention refused to write down the $L\times L$ matrix and ring attention refused to put the sequence on one GPU, and at inference when GQA, MLA, local layers and hybrids each cut what a token leaves behind. Compute is still quadratic for any layer that does full attention; the escape is to have fewer such layers and make them read less. Data turned out to be cheaper than expected — a few billion well-mixed long tokens at the end of the run. And evaluation is the one that’s still open: we got very good at making models accept a million tokens, and we’re still working on making them use all of it.
A context window is the set of distances a model practiced on, times the budget to pay for them. Everything else is engineering.
References
- Packing & recipe — Llama Team, The Llama 3 Herd of Models (arXiv:2407.21783) · DeepSeek-AI, DeepSeek-V3 Technical Report (arXiv:2412.19437) · Qwen Team, Qwen2.5-1M Technical Report (arXiv:2501.15383) · Qwen3 Technical Report (arXiv:2505.09388)
- RoPE — Su et al., RoFormer (arXiv:2104.09864)
- Position Interpolation — Chen et al. (arXiv:2306.15595) · ABF — Xiong et al., Effective Long-Context Scaling of Foundation Models (arXiv:2309.16039) · YaRN — Peng et al. (arXiv:2309.00071) · LongRoPE — Ding et al. (arXiv:2402.13753)
- Dual Chunk Attention — An et al., Training-Free Long-Context Scaling of Large Language Models (arXiv:2402.17463)
- FlashAttention — Dao et al. (arXiv:2205.14135) · Ring Attention — Liu, Zaharia, Abbeel (arXiv:2310.01889) · Striped Attention — Brandon et al. (arXiv:2311.09431) · DeepSpeed-Ulysses — Jacobs et al. (arXiv:2309.14509)
- MQA — Shazeer (arXiv:1911.02150) · GQA — Ainslie et al. (arXiv:2305.13245) · MLA — DeepSeek-AI, DeepSeek-V2 (arXiv:2405.04434)
- Sliding window & sinks — Jiang et al., Mistral 7B (arXiv:2310.06825) · Gemma Team, Gemma 3 (arXiv:2503.19786) · Xiao et al., Efficient Streaming Language Models with Attention Sinks (arXiv:2309.17453)
- Sparse attention — Yuan et al., Native Sparse Attention (arXiv:2502.11089) · Lu et al., MoBA (arXiv:2502.13189) · DeepSeek-AI, DeepSeek-V3.2-Exp (DSA, 2025)
- Hybrids — Lieber et al., Jamba (arXiv:2403.19887) · MiniMax, MiniMax-01 (arXiv:2501.08313) · Moonshot AI, Kimi Linear (arXiv:2510.26692)
- Data & evaluation — Fu et al., Data Engineering for Scaling Language Models to 128K Context (arXiv:2402.10171) · Gao et al., How to Train Long-Context Language Models (Effectively) (arXiv:2410.02660) · Hsieh et al., RULER (arXiv:2404.06654) · Liu et al., Lost in the Middle (arXiv:2307.03172) · Modarressi et al., NoLiMa (arXiv:2502.05167)
Every panel above is live, interactive and bilingual — drag the knobs, switch the language, and the demos follow.
从 2K 到 1M,到底发生了什么
眯着眼看全貌,其实就是五堵墙、五种修法。
位置不再是墙,是因为大家把 RoPE 看成了表盘,开始问到底哪几个超出了范围——慢的插值,快的不动,softmax 锐化一下,或者干脆把距离重新映射,让模型永远看不到新的距离。显存在训练侧不再是墙,是因为 FlashAttention 拒绝把 $L\times L$ 矩阵写下来、ring attention 拒绝把整条序列塞进一张卡;在推理侧,是因为 GQA、MLA、局部层、混合架构各自砍掉了一个 token 要留下的东西。算力对于任何做全注意力的层依然是平方的;出路是让这样的层少一点,再让它们读得少一点。数据比预想的便宜——训练末尾几十亿 token 配比合理的长数据就够了。而评测是唯一还没解决的:我们已经很会让模型收下一百万个 token 了,但让它用上全部,还在路上。
上下文窗口,就是模型练过的那些距离,再乘上为它们买单的预算。剩下的都是工程。
参考
- 打包与训练配方 — Llama Team, The Llama 3 Herd of Models(arXiv:2407.21783)· DeepSeek-AI, DeepSeek-V3 Technical Report(arXiv:2412.19437)· Qwen Team, Qwen2.5-1M Technical Report(arXiv:2501.15383)· Qwen3 Technical Report(arXiv:2505.09388)
- RoPE — Su 等, RoFormer(arXiv:2104.09864)
- 位置插值 — Chen 等(arXiv:2306.15595)· ABF — Xiong 等, Effective Long-Context Scaling of Foundation Models(arXiv:2309.16039)· YaRN — Peng 等(arXiv:2309.00071)· LongRoPE — Ding 等(arXiv:2402.13753)
- 双块注意力 — An 等, Training-Free Long-Context Scaling of Large Language Models(arXiv:2402.17463)
- FlashAttention — Dao 等(arXiv:2205.14135)· Ring Attention — Liu, Zaharia, Abbeel(arXiv:2310.01889)· Striped Attention — Brandon 等(arXiv:2311.09431)· DeepSpeed-Ulysses — Jacobs 等(arXiv:2309.14509)
- MQA — Shazeer(arXiv:1911.02150)· GQA — Ainslie 等(arXiv:2305.13245)· MLA — DeepSeek-AI, DeepSeek-V2(arXiv:2405.04434)
- 滑动窗口与 sink — Jiang 等, Mistral 7B(arXiv:2310.06825)· Gemma Team, Gemma 3(arXiv:2503.19786)· Xiao 等, Efficient Streaming Language Models with Attention Sinks(arXiv:2309.17453)
- 稀疏注意力 — Yuan 等, Native Sparse Attention(arXiv:2502.11089)· Lu 等, MoBA(arXiv:2502.13189)· DeepSeek-AI, DeepSeek-V3.2-Exp(DSA,2025)
- 混合架构 — Lieber 等, Jamba(arXiv:2403.19887)· MiniMax, MiniMax-01(arXiv:2501.08313)· Moonshot AI, Kimi Linear(arXiv:2510.26692)
- 数据与评测 — Fu 等, Data Engineering for Scaling Language Models to 128K Context(arXiv:2402.10171)· Gao 等, How to Train Long-Context Language Models (Effectively)(arXiv:2410.02660)· Hsieh 等, RULER(arXiv:2404.06654)· Liu 等, Lost in the Middle(arXiv:2307.03172)· Modarressi 等, NoLiMa(arXiv:2502.05167)
上面每一块面板都是实时、可交互、双语的——拖动旋钮、切换语言,演示会跟着变。