Blend two texts into one input and a 3B model keeps each text’s top pick in its top ten more than a third of the time
Average two unrelated texts into one input, and a language model still ranks each text's next word near the top. This paper tests three language models, Pythia, Llama and Qwen, each with about three billion parameters. A language model is the kind of AI that writes a chatbot's replies, and parameters are the adjustable numbers it learns. The model reads text as tokens, whole words or pieces of words. Each token enters as a long list of numbers, called an embedding. The authors take two texts of the same length and average their embeddings, token by token. The authors first run each text alone and note the model's top pick. Fed the blend, the model still puts that pick in its top ten thirty-six to forty percent of the time. For scale, the authors ran Pythia on text A alone and looked for text B's real next token. It made the top ten less than three times in a hundred.
Average the two texts position by position, and the unchanged model reads the blend in one pass
Write text A as a row of tokens, and text B as a second row of the same length. Take the first token of each, and average their two lists of numbers. Do the same for the second pair, then the third, all the way along. The model reads that one averaged row in a single pass. The model itself is untouched, with no retraining. The authors then compare its answer with a target. The target is the average of the two answers the model gives when it reads each text alone.
The model is partly straight inside, so a blended input gives a partly blended answer
A straight-line machine would handle a blend perfectly. Feed it half of input A plus half of input B, and it gives back half of answer A plus half of answer B. The paper's name for this is linear superposition. A language model bends its signal at every layer, the stages each token passes through in turn. So the real question is how much straight line is left. The authors measure how far the blend's answer lands from the ideal half-and-half answer. They compare that miss with the gap between the answers to two random texts. For Pythia 2.8B on thirty-two-token texts, the miss is forty-two to seventy-one percent of that gap, across three ways of measuring the distance. The blend lands nearer than chance and far from exact, so the model is partly straight. The model averages its raw scores before it turns them into probabilities, the odds it gives each token. Averaging the scores sets each token's odds by the square root of its two odds multiplied together. A token that only text A wants gets pulled down by text B's low odds. The tokens that survive are the ones both texts would accept.
In Qwen, most top-ten hits land on predictable positions, and content words rank far lower
The authors split every position in the texts into two groups, for Qwen. About sixty-five percent are predictable spots. There, each text's own top pick sits at a median rank of six in the blend. Half the time it ranks sixth or better. The other thirty-five percent carry content, and there the median rank is two hundred and eighty-four. So in Qwen, most of the top-ten hits come from the predictable spots. The paper does not say how it drew the line between the two groups. In an appendix, the authors sample tokens that rank second to tenth. Only three of those twenty-five are content words, and most of the rest are punctuation and small words like and, in, or to.
Tuning for the blend doubles top-five hits for Pythia and Llama, and makes Pythia worse at one text
Next, the authors train the model for the blend. They call it self-distillation. They fine-tune the model so its answer to a blend matches the average of its own two separate answers. On Figure 1, the dashed curves are the tuned models. For Pythia and Llama, the share of top picks in the top five more than doubles. Pythia goes from about twenty-seven percent to sixty-two, and Llama from twenty-six to seventy-two. Qwen slips slightly, from thirty percent to twenty-seven. The tuning for Pythia 2.8B took a hundred and fourteen hours on two Nvidia A100 chips. The tuning also makes the model worse at reading one text. On a last-word test, the model reads a passage and guesses its final word. Tuned Pythia 2.8B, reading a single text, drops from fifty-four percent right to thirty-six.
Pythia is straightest in its first few hundred training steps and bends by step 10,000
Figure 2 follows the inside of the network through training. The authors track six Pythia models, from seventy million to 2.8 billion parameters, across a hundred and forty-three thousand training steps. At each checkpoint they measure linearity error. Linearity error is how far the blend's inner signals land from the average of the two texts' separate inner signals. Lower means straighter. Every size scores lowest in the first few hundred steps. The error climbs steeply until about step ten thousand, then levels off. This figure measures the network's insides, and the paper leaves the top-ten score untracked across training.
The best decoder gets 43% right on the blend, and its small helper alone gets 54%
The best decoder in the paper is Joint Contrastive decoding. It runs a small helper model, a one-billion-parameter Llama, on each text separately, and uses the helper to pull the two answers apart. On the last-word test, the blended three-billion Llama goes from eighteen percent right to forty-three. The helper on its own, reading one text, gets fifty-four percent.
The blend runs twice as fast as one text at a time, and as fast as a batch of two
The paper's conclusion promises twice the throughput, meaning tokens produced per second. On the three-billion Llama, running the two texts one after the other gives forty-one point three tokens per second. The blend in its fastest mode gives eighty-one point five. An ordinary batch of two, with both texts running side by side through the ordinary model, gives eighty-one point four. That fastest mode uses two output heads, a separate final scoring layer for each text. In the paper's accuracy table, the two-head decoder on Llama scores ten point five percent on the last-word test, below the untuned blend at eighteen. The helper method runs at fifty-two point four, slower than the plain batch. The conclusion also promises half the memory for each text, and the paper never measures that. Every model tested has between seventy million and eight billion parameters. The texts run to five hundred and twelve tokens at most, mostly in English. The authors released no code. The authors name the next tests themselves: longer texts, other languages, and models that read images as well as text.
Skim Figures 1 and 2 and Table 2 for how straight a language model is inside
Skim it. Figure 1, Figure 2 and Table 2 hold the finding. You can pass on the decoding and speed sections. Their best decoder trails its own small helper, and their fastest mode ties a plain batch of two. The paper is a version one preprint on arXiv, the open archive for research papers, posted before peer review.




















