Run two unrelated documents through a large language model at the same time, and you assume the model keeps them separate — two hidden states, two outputs, no interference. A paper generating serious discussion right now says that assumption is wrong in a surprising way: mix the inputs and the model produces a precise statistical mixture of the outputs. Transformers, it turns out, are linear in a place nobody expected: the mapping from input combinations to output distributions.
The paper, “Your Transformer Can Hold Two Thoughts at Once”, introduces what the authors call the Superposition Linearity Hypothesis, and it has implications that go well beyond curiosity — from how model internals are actually organized to a decoding trick that generates two coherent texts from a single forward pass. Here is what the finding says, why it matters, and where it breaks down.
The Finding: Mixed Inputs, Mixed Outputs
The core experiment is deceptively simple. Take a prompt A and measure the model’s next-token distribution. Take an unrelated prompt B and do the same. Now create a blended input — a linear combination of the two token embedding sequences — and run it through the model. The naive expectation is garbage: neural networks are full of nonlinearities (attention softmaxes, MLP activations, layer norms), so mixing inputs should produce noise, not structure.
Instead, the output distribution for the blended input lands close to the average of the two individual output distributions. The authors quantify this with a divergence measure between the model’s actual prediction and the ideal mixture, and the divergence stays small across models and prompt pairs. The output distribution behaves as if the model had processed both thoughts independently and combined the results at the very end — a property they call superposition linearity.
The word “superposition” is doing deliberate work here. Mechanistic interpretability research has argued for years that networks represent more features than they have dimensions by storing them in overlapping, non-orthogonal directions. This paper adds an architectural twist: the linear output behavior appears to be an intrinsic property of the Transformer architecture itself, not something training had to teach. In fact, the authors observe the opposite of emergence — the linearity tends to diminish as pretraining progresses. Training spends part of its budget making the model less linear, presumably because nonlinearity buys task performance.
Why This Is Strange
It helps to be precise about what is and is not linear here. The internal computation is anything but: softmax attention reweights each token’s context nonlinearly, MLP blocks apply pointwise nonlinear activations, and normalization rescales activations in a input-dependent way. Yet the end-to-end map from “input embedding sequence” to “next-token distribution” approximately commutes with linear combination, at least for inputs drawn from distinct text streams.
That distinction matters for how you think about what LLMs compute. If the final hidden state can carry two interleaved “thoughts” whose decode-side contributions add up, then the representation is closer to a sum of quasi-independent features than to a single holistic chunk of meaning. It also explains a familiar empirical quirk: prompts that interfere with each other — where two contexts blur together — degrade outputs roughly as you would predict from averaging their individual distributions, rather than failing in unpredictable ways.
Restoring Linearity With a Cheap Fine-Tune
The second result is the practical hook. If pretraining erodes linearity, can it be restored? The authors show that lightweight fine-tuning substantially reduces the divergence between the model’s blended-input prediction and the ideal mixture — the model becomes measurably more linear at a small training cost, without collapsing its task performance.
For interpretability researchers this is significant. A large family of analysis techniques — activation patching, linear probing, feature attribution — quietly assumes that interventions combine in roughly predictable ways. A model that has been nudged back toward linearity is a cleaner substrate for those techniques. Think of it as calibrating an instrument before taking measurements: the tuned model is not smarter, but it is more legible.
Guided Decoding: Two Texts From One Pass
The showcase application is a decoding procedure that exploits the linearity directly. Run a forward pass on a superposition of two prompts, and the output distribution is approximately the mixture of the two continuations’ distributions. The guided decoding step then disentangles that mixture — separating the combined distribution back into two coherent text streams, each generated from what is effectively a shared computation.
A sketch of what this means operationally:
- Encode prompt A and prompt B as embedding sequences.
- Form a linear combination of the two (the paper’s blending scheme).
- Run a single forward pass — one sequence, not two.
- At decode time, use the guided procedure to split the resulting mixture distribution and track which continuation each token belongs to.
Whether this becomes a practical batching strategy is an open question. KV caches, positional embeddings, and attention masks all assume one coherent sequence, so a production-grade implementation would need to handle the engineering details the research code does not. But as a proof of capability it is striking: the arithmetic of the output distribution is linear enough that you can invert it.
Where It Breaks Down
The hypothesis has boundaries, and the authors are careful about them. The property holds for inputs from distinct text streams; related or adversarially chosen inputs can interact nonlinearly. It holds approximately, not exactly — divergence is reduced, not eliminated. And it concerns next-token distributions, not full-generation semantics: two superposed thoughts coexist in the distribution, but a greedy decode still collapses to one output stream unless you use the disentangling procedure.
There is also a scaling caveat. If linearity diminishes with pretraining progress, then the strongest current models are the least linear, all else equal — which is exactly why the fine-tuning result matters. The property is recoverable, but do not assume a frontier model exhibits it to the same degree as a smaller checkpoint out of the box.
Why You Should Care
Three takeaways worth carrying forward. First, for interpretability: superposition linearity gives the field a new, cheap sanity check for how much of a model’s behavior is compositional, and a training lever to increase that compositional quality when running intervention experiments. Second, for efficiency research: the single-pass, two-continuation result is an early data point for a class of techniques that amortize computation across related requests — the same instinct behind shared-prefix batching, taken further. Third, for everyone building on these models: prompt interference, that vague sense that a crowded context makes every answer a bit worse, now has a cleaner theoretical description. Mixing contexts really does mix the model’s predictions — predictably, and measurably.
The paper is on arXiv if you want the full divergence metrics and fine-tuning setup. It is a good week for it: this is the kind of result that makes the black box feel slightly less like a black box, one linear combination at a time.