Two claims in one context. Only one of them answers.
A sentence planted 55,000 tokens back says the code is 7431. A short record beside the question says 5824. This study replays what was actually measured: which one the model answered with, and what had to change before the newer record could be read at all.
One 64,000-token context. The original sentence is planted 15% of the way in, roughly 55,000 tokens before the question. A short record is injected beside the question itself. Choose what the model can still read, and what that record says.
02 / THE RECORDED ANSWER
7431
INERT
DIVERGENCE FROM THE UNMODIFIED RUN0.0046 BITSinside the 0.05-bit qualifying tolerance
The contradicting record is readable and is not read. The divergence sits inside the spread of the supply arms — indistinguishable from promoting nothing.
ARM / OVERRIDE K=0
03 / EVERY ARM THAT WAS RUN
SOURCE LIVE · no record · +074310.0000
SOURCE LIVE · different value · +074310.0046
SOURCE LIVE · different value · +174310.0007
SOURCE LIVE · different value · +274310.0003
SOURCE LIVE · different value · +374310.0004
SOURCE RETIRED · no record · +0corvin6.2700
SOURCE RETIRED · same value · +074310.0020
SOURCE RETIRED · same value · +174310.0017
SOURCE RETIRED · same value · +274310.0052
SOURCE RETIRED · same value · +374310.0056
SOURCE RETIRED · different value · +0582414.7914
0.00010.05 · QUALIFYING TOLERANCE20BITS · LOGARITHMIC
04 / CHOOSE WHICH READS STOP
The model has 48 attention layers. Every sixth one reads the whole context; the rest see only the last 1024 tokens. The source is far outside that window, so these eight are the only path to it. Switch one off and the record beside the question may take over.
LAYER 0SLIDING · GLOBAL · SLIDINGLAYER 47
1 READ RETIRED · LAYER 29
5824
THE PROMOTED RECORD TAKES THE ANSWER
DIVERGENCE6.3298 BITS
One blocked read. The span stays fully readable in the other seven global layers, including every later one, and the promoted record takes the answer.
ARM / ONLY_29
THE 15 SUBSETS THAT WERE RUN
WHAT YOU ARE READING
Nothing on this page runs a model. Both instruments look up arms that were already measured on google/gemma-3-12b-it under MLX · bf16 · greedy decode, in a 65,536-token context, and show you the recorded answer and its divergence from the unmodified run. Combinations that were never run return no answer, because there is nothing to return.
WHY THE GAPS MATTER
The eight global attention layers admit 256 possible retirement subsets. Fifteen were run: the empty set, each layer alone, four count-matched triples, one four-layer set and the full retirement. That is enough to show that one particular layer flips the answer by itself and that three blocked reads can behave in opposite directions — and not nearly enough to describe the other 241 subsets. The instrument stays silent about them deliberately.
ONE QUESTION, ONE SPAN, ONE MODEL
Every arm here asks the same question about the same planted sentence, at one depth, on one model.
Layer 29 is sufficient to flip the answer. It is not necessary: two other subsets flip without it.
A recorded divergence is a measurement of these arms, not a property of transformers.
An interactive study may expose recorded evidence. It must never manufacture an answer where the experiment has none.
one question, one span, one model — every arm here asks the same question about the same planted sentence, at one depth, on one model.; layer 29 is sufficient to flip the answer. it is not necessary: two other subsets flip without it.; a recorded divergence is a measurement of these arms, not a property of transformers.. An interactive study may expose recorded evidence. It must never manufacture an answer where the experiment has none.
CONTINUE THE INVESTIGATION
The notebook follows the same arc from the first promotion result to the question it leaves open: what makes a source overridable without an attention intervention at all.