Average two texts into one input, and a 3B model still ranks both next words

1 hour ago

@first-passSubscribe

Average two unrelated texts into one input, and a language model still ranks each text's next word near the top. This paper tests three language models, Pythia, Llama and Qwen, each with about three billion parameters. A language model is the kind of AI that writes a chatbot's replies, and parameters are the adjustable numbers it learns. The model reads text as tokens, whole words or pieces of words. Each token enters as a long list of numbers, called an embedding. The authors take two texts of the same length and average their embeddings, token by token. The authors first run each text alone and note the model's top pick. Fed the blend, the model still puts that pick in its top ten thirty-six to forty percent of the time. For scale, the authors ran Pythia on text A alone and looked for text B's real next token. It made the top ten less than three times in a hundred.

Ask

Ask about this presentation

Answers are generated from this presentation.

Chapters

  1. 0:00Blend two texts into one input and a 3B model keeps each text’s top pick in its top ten more than a third of the time
  2. 0:49Average the two texts position by position, and the unchanged model reads the blend in one pass
  3. 1:14The model is partly straight inside, so a blended input gives a partly blended answer
  4. 2:16In Qwen, most top-ten hits land on predictable positions, and content words rank far lower
  5. 2:53Tuning for the blend doubles top-five hits for Pythia and Llama, and makes Pythia worse at one text
  6. 3:38Pythia is straightest in its first few hundred training steps and bends by step 10,000
  7. 4:13The best decoder gets 43% right on the blend, and its small helper alone gets 54%
  8. 4:34The blend runs twice as fast as one text at a time, and as fast as a batch of two
  9. 5:36Skim Figures 1 and 2 and Table 2 for how straight a language model is inside
Show transcript

Blend two texts into one input and a 3B model keeps each text’s top pick in its top ten more than a third of the time

Blend two texts into one input and a 3B model keeps each text’s top pick in its top ten more than a third of the time
Out of 50,000 or more tokens · three models of about 3 billion parameters
blue Pythia · green Llama · red Qwen · solid = as released
up the side · share of positions where each text’s own top pick ranks this high or better
along the bottom · rank in the blend’s list
dashed line · top ten
36 to 40% · base models
Figure 1 · two texts averaged into one input · Pythia 2.8B, Llama 3.2 3B, Qwen 2.5 3B
Tikhonov, Korznikov, Mikhalchuk, Dragunov, Rahmatullaev, Druzhinina, Razzhigaev, Oseledets, Tutubalina · “Your Transformer Can Hold Two Thoughts at Once” · arXiv 2609.29845 v1 · Sep 2026 · preprint

Average two unrelated texts into one input, and a language model still ranks each text's next word near the top. This paper tests three language models, Pythia, Llama and Qwen, each with about three billion parameters. A language model is the kind of AI that writes a chatbot's replies, and parameters are the adjustable numbers it learns. The model reads text as tokens, whole words or pieces of words. Each token enters as a long list of numbers, called an embedding. The authors take two texts of the same length and average their embeddings, token by token. The authors first run each text alone and note the model's top pick. Fed the blend, the model still puts that pick in its top ten thirty-six to forty percent of the time. For scale, the authors ran Pythia on text A alone and looked for text B's real next token. It made the top ten less than three times in a hundred.

Average the two texts position by position, and the unchanged model reads the blend in one pass

Average the two texts position by position, and the unchanged model reads the blend in one pass
The target is the average of the model’s two separate answers
text A
text B
average
the model
unchanged · one pass
its answer
target
answer to A alone + answer to B alone, halved
compare
our diagram · from Eq. 1
Tikhonov, Korznikov, Mikhalchuk, Dragunov, Rahmatullaev, Druzhinina, Razzhigaev, Oseledets, Tutubalina · “Your Transformer Can Hold Two Thoughts at Once” · arXiv 2609.29845 v1 · Sep 2026 · preprint

Write text A as a row of tokens, and text B as a second row of the same length. Take the first token of each, and average their two lists of numbers. Do the same for the second pair, then the third, all the way along. The model reads that one averaged row in a single pass. The model itself is untouched, with no retraining. The authors then compare its answer with a target. The target is the average of the two answers the model gives when it reads each text alone.

The model is partly straight inside, so a blended input gives a partly blended answer

The model is partly straight inside, so a blended input gives a partly blended answer
The blend misses by 42 to 71% of the gap between two random texts · Pythia 2.8B, three distance measures
input
answer
text A
text B
halfway in, halfway out
a straight-line machine
a language model
red · the miss, from the ideal point to the blend’s answer
grey · the gap between answers to two random texts
the miss · 42 to 71% of the gap · three distance measures · Pythia 2.8B, 32-token texts, Table 1
text A’s odds
both
A only
B only
neither
text B’s odds
both
A only
B only
neither
the blend
both
A only
B only
neither
illustrative odds over four tokens · the blend takes the square root of A’s odds times B’s odds
our diagram · from Table 1 and Eq. 6
Tikhonov, Korznikov, Mikhalchuk, Dragunov, Rahmatullaev, Druzhinina, Razzhigaev, Oseledets, Tutubalina · “Your Transformer Can Hold Two Thoughts at Once” · arXiv 2609.29845 v1 · Sep 2026 · preprint

A straight-line machine would handle a blend perfectly. Feed it half of input A plus half of input B, and it gives back half of answer A plus half of answer B. The paper's name for this is linear superposition. A language model bends its signal at every layer, the stages each token passes through in turn. So the real question is how much straight line is left. The authors measure how far the blend's answer lands from the ideal half-and-half answer. They compare that miss with the gap between the answers to two random texts. For Pythia 2.8B on thirty-two-token texts, the miss is forty-two to seventy-one percent of that gap, across three ways of measuring the distance. The blend lands nearer than chance and far from exact, so the model is partly straight. The model averages its raw scores before it turns them into probabilities, the odds it gives each token. Averaging the scores sets each token's odds by the square root of its two odds multiplied together. A token that only text A wants gets pulled down by text B's low odds. The tokens that survive are the ones both texts would accept.

In Qwen, most top-ten hits land on predictable positions, and content words rank far lower

In Qwen, most top-ten hits land on predictable positions, and content words rank far lower
Qwen 2.5 3B · median rank 6 on predictable positions, 284 on content positions
positions in the text
share of positions
median rank of each text’s own top pick
predictable positions
about 65%
6
content positions
about 35%
284
sampled tokens ranked 2nd to 10th · “.” “,” “and” “in” “to” · 3 of 25 content words · App. H
median rank 6 · half the time the top pick ranks sixth or better
Table 2 · Qwen 2.5 3B · the paper does not define the split · samples from App. H
Tikhonov, Korznikov, Mikhalchuk, Dragunov, Rahmatullaev, Druzhinina, Razzhigaev, Oseledets, Tutubalina · “Your Transformer Can Hold Two Thoughts at Once” · arXiv 2609.29845 v1 · Sep 2026 · preprint

The authors split every position in the texts into two groups, for Qwen. About sixty-five percent are predictable spots. There, each text's own top pick sits at a median rank of six in the blend. Half the time it ranks sixth or better. The other thirty-five percent carry content, and there the median rank is two hundred and eighty-four. So in Qwen, most of the top-ten hits come from the predictable spots. The paper does not say how it drew the line between the two groups. In an appendix, the authors sample tokens that rank second to tenth. Only three of those twenty-five are content words, and most of the rest are punctuation and small words like and, in, or to.

Tuning for the blend doubles top-five hits for Pythia and Llama, and makes Pythia worse at one text

Tuning for the blend doubles top-five hits for Pythia and Llama, and makes Pythia worse at one text
Qwen stays flat · last-word test on a single text falls from 54% to 36%
dashed = tuned for the blend
dashed line · top five
Pythia · about 27% to 62%
Llama · about 26% to 72%
Qwen · about 30% to 27%
cost · 114 hours on two A100 chips · last-word test, one text: 0.544 to 0.357 (Pythia 2.8B, LAMBADA)
Figure 1 · Sec. 4 · App. A
Tikhonov, Korznikov, Mikhalchuk, Dragunov, Rahmatullaev, Druzhinina, Razzhigaev, Oseledets, Tutubalina · “Your Transformer Can Hold Two Thoughts at Once” · arXiv 2609.29845 v1 · Sep 2026 · preprint

Next, the authors train the model for the blend. They call it self-distillation. They fine-tune the model so its answer to a blend matches the average of its own two separate answers. On Figure 1, the dashed curves are the tuned models. For Pythia and Llama, the share of top picks in the top five more than doubles. Pythia goes from about twenty-seven percent to sixty-two, and Llama from twenty-six to seventy-two. Qwen slips slightly, from thirty percent to twenty-seven. The tuning for Pythia 2.8B took a hundred and fourteen hours on two Nvidia A100 chips. The tuning also makes the model worse at reading one text. On a last-word test, the model reads a passage and guesses its final word. Tuned Pythia 2.8B, reading a single text, drops from fifty-four percent right to thirty-six.

Pythia is straightest in its first few hundred training steps and bends by step 10,000

Pythia is straightest in its first few hundred training steps and bends by step 10,000
Linearity error across 143,000 steps · six sizes from 70 million to 2.8 billion parameters
six Pythia models · 70 million to 2.8 billion parameters
linearity error · lower = straighter
lowest · first few hundred steps
climbs steeply
levels off · step 10,000 to 143,000
Figure 2 · the network’s insides, not the top-ten score
Tikhonov, Korznikov, Mikhalchuk, Dragunov, Rahmatullaev, Druzhinina, Razzhigaev, Oseledets, Tutubalina · “Your Transformer Can Hold Two Thoughts at Once” · arXiv 2609.29845 v1 · Sep 2026 · preprint

Figure 2 follows the inside of the network through training. The authors track six Pythia models, from seventy million to 2.8 billion parameters, across a hundred and forty-three thousand training steps. At each checkpoint they measure linearity error. Linearity error is how far the blend's inner signals land from the average of the two texts' separate inner signals. Lower means straighter. Every size scores lowest in the first few hundred steps. The error climbs steeply until about step ten thousand, then levels off. This figure measures the network's insides, and the paper leaves the top-ten score untracked across training.

The best decoder gets 43% right on the blend, and its small helper alone gets 54%

The best decoder gets 43% right on the blend, and its small helper alone gets 54%
Llama 3.2 3B · the helper reads each text separately at every step
last-word test (LAMBADA) · share right
blend, untuned
0.182
blend with Joint Contrastive decoding
0.430
the 1B helper alone, one text
0.540
Table 5 · Llama 3.2 3B with a Llama 3.2 1B helper · last-word test (LAMBADA), share right
Tikhonov, Korznikov, Mikhalchuk, Dragunov, Rahmatullaev, Druzhinina, Razzhigaev, Oseledets, Tutubalina · “Your Transformer Can Hold Two Thoughts at Once” · arXiv 2609.29845 v1 · Sep 2026 · preprint

The best decoder in the paper is Joint Contrastive decoding. It runs a small helper model, a one-billion-parameter Llama, on each text separately, and uses the helper to pull the two answers apart. On the last-word test, the blended three-billion Llama goes from eighteen percent right to forty-three. The helper on its own, reading one text, gets fifty-four percent.

The blend runs twice as fast as one text at a time, and as fast as a batch of two

The blend runs twice as fast as one text at a time, and as fast as a batch of two
81.5 against 81.4 tokens per second on Llama 3.2 3B
one text after the other
41.3
blend, two-head mode
81.5
ordinary batch of two
81.4
blend with the helper
52.4
Table 9 · Llama 3.2 3B + 1B · tokens per second, mean of 100 runs
two-head decoder · last-word test 0.105 · untuned blend 0.182 (Table 6, Llama)
“halving the KV-cache footprint” · never measured (Sec. 6, App. G.5)
tested · 70 million to 8 billion parameters · texts up to 512 tokens · mostly English
released · no code
Table 9, Table 6, App. G.5, Sec. 8
Tikhonov, Korznikov, Mikhalchuk, Dragunov, Rahmatullaev, Druzhinina, Razzhigaev, Oseledets, Tutubalina · “Your Transformer Can Hold Two Thoughts at Once” · arXiv 2609.29845 v1 · Sep 2026 · preprint

The paper's conclusion promises twice the throughput, meaning tokens produced per second. On the three-billion Llama, running the two texts one after the other gives forty-one point three tokens per second. The blend in its fastest mode gives eighty-one point five. An ordinary batch of two, with both texts running side by side through the ordinary model, gives eighty-one point four. That fastest mode uses two output heads, a separate final scoring layer for each text. In the paper's accuracy table, the two-head decoder on Llama scores ten point five percent on the last-word test, below the untuned blend at eighteen. The helper method runs at fifty-two point four, slower than the plain batch. The conclusion also promises half the memory for each text, and the paper never measures that. Every model tested has between seventy million and eight billion parameters. The texts run to five hundred and twelve tokens at most, mostly in English. The authors released no code. The authors name the next tests themselves: longer texts, other languages, and models that read images as well as text.

Skim Figures 1 and 2 and Table 2 for how straight a language model is inside

Skim Figures 1 and 2 and Table 2 for how straight a language model is inside
The decoding and speed sections trail a small model and a plain batch of two
Skim it
Figure 1 · Figure 2 · Table 2 · the finding
best decoder 43% against its helper’s 54% · fastest mode ties a batch of two
version 1 preprint · no code released
Pavel Tikhonov, Anton Korznikov, Matvey Mikhalchuk, Nikita Dragunov, Temurbek Rahmatullaev, Polina Druzhinina, Anton Razzhigaev, Ivan Oseledets, Elena Tutubalina
“Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs”
arXiv 2609.29845 v1 · 24 September 2026 · preprint, not peer reviewed
Tikhonov, Korznikov, Mikhalchuk, Dragunov, Rahmatullaev, Druzhinina, Razzhigaev, Oseledets, Tutubalina · “Your Transformer Can Hold Two Thoughts at Once” · arXiv 2609.29845 v1 · Sep 2026 · preprint

Skim it. Figure 1, Figure 2 and Table 2 hold the finding. You can pass on the decoding and speed sections. Their best decoder trails its own small helper, and their fastest mode ties a plain batch of two. The paper is a version one preprint on arXiv, the open archive for research papers, posted before peer review.