CHRIS HAY

IDEAS · SYSTEMS · OBJECTS / LONDON · 2026

Which source wins?

A newer record does not win by being right.

The Apollo store query routes to window 170 and returns John Coyle with 23 bowls.
FROM THE FILM / 19:30

The Apollo document read. The system routes to a stored window, loads it, reads the passage and answers. Every step assumes one thing: that the place it went to is the place entitled to answer.

FILM & SOURCE RECORD ↗
ABOUT THIS NOTE +

When two things in one context disagree, which one answers? A recorded long-context arc measures it: a compact record can replace a source that has been retired, is completely inert against one that can still be read, and a single blocked attention layer is enough to change which of them the model answers with. The source itself is never altered, so none of this licenses forgetting it.

N-AUTHORITYOPENRECORDED 2026-09-06DRAFT · V0.1REFERENCE DRAFT
OLD SOURCE7431
↓ APPROX. 55,000 TOKENS
NEW RECORD5824

Who answers?

THE ARGUMENT IN FIVE LINES READ +

A newer record does not win by being right.

THE ARGUMENT, IN FIVE LINES

  1. Two claims disagree inside one context.

  2. The old source is still readable.

  3. The new record does nothing at all.

  4. Retire one particular attention read.

  5. The answer changes.

Everything below is how that was measured, what it cost to believe it, and where it stops being true. Read the five lines and the instrument for the shape of it; read the rest for the evidence.

01

Bring the fact to the question.

A compact record can supply a missing fact — and the model can calculate with it. Where that record arrives matters.

THE WORKING CONTEXTSCHEMATIC / NOT TO SCALE
SOURCE7431Meridian code
≈ 49,000 tokens
PROMOTED RECORD7431 / 5824correct / corrupted

What is the Meridian code? Then: add one. Reverse it.

RECORD / CONTROLCOPY+1REVERSE
Correct · 7431743174321347
Corrupted · 5824582458254285
Padding only72497897892
Original span excluded in every row. A corrupted record steers all three operations; padding alone fails all three.

Move the record close to the question, and computation returns.

But prose cannot reliably say “already computed”: a record stating the reversed code gets reversed again. A finished result needs a typed property.

FIELD NOTES / 01 MEASUREMENTS & QUALIFICATIONS +

01 / TWO CLAIMS IN ONE CONTEXT

Any memory system that summarises, supersedes or evicts eventually holds two things that disagree: an original passage and a shorter record standing in for it. The engineering question sounds like bookkeeping — which one is authoritative? The measured answer is not a preference the model expresses. It is a property of what the model can still read at the moment it is asked.

THE OPERATION BEING TESTED

Build a 64,000-token context from a verbose source. Then, beside the question roughly 55,000 tokens later, inject a compact record and cut the original sentence off from attention. This is not the same as having the record present from the first token, which is what earlier runs measured. It is the operation an eviction system actually performs: replacing something already in memory.

EVIDENCE

PROMOTION / THREE OPERATIONS

In the founding run, a 13-token record injected about 49,000 tokens downstream of the source keeps the model correct on all three tested operations — copy 7431, plus one 7432, digit reversal 1347 — with the original span excluded. Padding alone fails on every row (7249, 789, 7892) and a corrupted record steers every row (5824, 5825, 4285). Asked for the code plus one with 5824 injected, the model answers 5825: a value that appears nowhere in the context.

SUPPORTED

WHAT HAD BEEN MEASURED WAS THE DISTANCE

An earlier run concluded that compact records serve reads but not computation. Identical record, model, framing and question; only its position changes. Plus one goes from 12.98 bits, reciting 7431, when the record is ingested 49,000 tokens back, to 0.1472 and the correct 7432 when it arrives at the working boundary. Reversal goes from 15.23 to 0.0610. Moving a fact close to where the model is working returned a capability that was unavailable while it sat far away.

AND PROSE CANNOT SAY “ALREADY DONE”

Injecting “The Meridian code reversed is 1347” and then asking for the reversed code, the model reverses 1347 again and answers 7431. The corrupted twin behaves the same way, and an earlier key–value phrasing failed identically. The failure is systematic, survives rewording, and depends on the operation. A record store cannot use wording to mark a result as finished; that has to be a typed property of the record, not a sentence about it.

FROM THE REGISTER / INTO YOUR HANDS

Retire one read.
Watch who answers.

Eight layers see the whole context.
Switch one off. Read what was measured.

OPEN THE STUDY

A RECORDED ARM / LAYER 29 RETIRED

the sentence 55,000 tokens back7431
the record beside the question5824
THE MODEL
ANSWERED
5824

ONE BLOCKED READ / SEVEN STILL READING / 6.3298 BITS

Set the original sentence live or retired, choose what the promoted record says, and read the answer that was actually recorded. Then retire individual attention layers and watch which reads have to stop before the newer record can be read at all. Combinations that were never run return nothing.

EXPERIMENT & PROVENANCE ↗
02

Present does not mean heard.

A correct record fills a gap. A contradicting record beside the question leaves the live source’s answer unchanged, even with three companions.

SOURCE RETIRED

7431 ×

ONE CORRECT RECORD → ANSWER7431

One record is enough.

Supplying the original value into a gap.
SOURCE LIVE / ASSERTING 7431

5824 inert

1 RECORD7431
2 RECORDS7431
3 RECORDS7431
4 RECORDS7431
With the source retired, a contradicting 5824 record also steers the answer to 5824 in a separate run. This cross-run comparison is detailed in the field notes.

The new record isn’t losing.
It isn’t participating.

SUPPLY1

record · qualifies

DISPOSITION OVERRIDE≈4

records · qualifies without competing text

LIVE SOURCE OVERRIDE4×

records tested · does not qualify

FIELD NOTES / 02 MEASUREMENTS & QUALIFICATIONS +

02 / SUPPLYING IS NOT OVERRULING

The obvious next question is what promotion costs. If contradicting a readable passage simply needs more records than filling a gap, a planner can size its closures accordingly. So: same question, same corpus, same injection budget of 72 tokens. Restore a retired fact, or contradict one that is still there, sweeping one to four records in each.

EVIDENCE

SUPPLY / ONE RECORD IS ENOUGH

With the source retired, a single record restores the answer at 0.0020 bits and stays qualifying across zero to three companions (0.0020, 0.0017, 0.0052, 0.0056). The control holds: retiring the source with nothing promoted collapses the answer at 6.2700 bits, which is what establishes that the source was necessary.

SUPPORTED

NOT SUPPORTED

OVERRIDE / NOT AT ANY TESTED SIZE

With the source live and asserting 7431, a contradicting record plus zero to three companions leaves the answer at 7431 every time, at 0.0046, 0.0007, 0.0003 and 0.0004 bits. Adding companions does not help. The scope is one question, one model, at most three companions and a single-token answer.

INERT, NOT OUTVOTED

This is not the model weighing two claims and preferring the older one. Every override divergence sits inside the spread of the supply arms — indistinguishable from promoting nothing at all. The promoted record is not losing an argument; it is not participating in one. The decisive comparison is across two runs on a byte-identical string: with the source retired, the same twenty-odd characters that were inert here steered the answer to 5824 at 14.7914 bits. Same record, same corpus, same injection position. Only the readability of the source differs.

THREE TIERS, NOT TWO

  • Supplying a fact into a gap qualified at one record.
  • Overriding the model’s own disposition, with no competing text readable, took about four records and forty contentful tokens.
  • Overriding a live, readable source did not happen at any size tested.

Cost is not the only axis. Some operations are not expensive; they are unavailable.

EXPERIMENT & PROVENANCE ↗
03

Which read matters?

Cutting layer 29 alone flips the answer. Other combinations do too. The position of the cuts matters, and most combinations are still unmeasured.

ONE CONTEXT / TWO CLAIMSRECORDED ARMS
ORIGINAL SOURCE7431Earlier in the document
PROMOTED RECORD5824Beside the question

↓ APPROX. 55,000 TOKENS FROM SOURCE TO QUESTION

THE EIGHT GLOBAL READS OF THE ORIGINAL SOURCE

5open
11open
17open
23open
29open
35open
41open
47open
THE MODEL ANSWERS7431

Both claims are present. The original still supplies the answer.

All eight reads open: 7431. Retire layer 29 alone: 5824, with seven reads still open. Schematic · Gemma 3 12B · one synthetic 65,536-token context. Layer attribution ↗

Count is not
the mechanism.

RETIRED READS511172329354147ANSWER
None········7431
29····×···5824
5 · 23 · 47×··×···×7431
17 · 29 · 41··×·×·×·5824
35 · 41 · 47·····×××5824
5 · 11 · 17 · 23××××····5824
× Retired read · Open readA selection of six measured arms.
15 / 256SUBSETS MEASURED
241UNKNOWN
TRY ALL EIGHT SWITCHES IN THE STUDY ↗
FIELD NOTES / 03 MEASUREMENTS & QUALIFICATIONS +

03 / SO WHICH READS HAVE TO STOP?

If a live source cannot be overruled, then retirement is not a saving to be made later — it is the mechanism that makes promotion work at all. That turns a vague instruction into a measurable one. This model has 48 attention layers, and every sixth reads the whole context; the rest see only the last 1,024 tokens. The planted sentence sits far outside that window, so those eight global layers are the only path to it. Which of them have to stop reading?

EVIDENCE

ONE BLOCKED READ

Retiring the source from global layer 29 alone flips the answer to the promoted 5824 at 6.3298 bits — with the span still fully readable in the other seven global layers, including every later one. No other single layer comes close: 0.0009, 0.0008, 0.0143, 0.0019, 0.2312, 0.0097 and 0.0086 all leave the answer at 7431. Reproduction is solid: the nothing-retired arm returns 0.0046 on a sixth independent prefill, matching two earlier runs to four decimals.

SUPPORTED

COUNT IS NOT THE MECHANISM

Three blocked reads scattered across the stack ([17, 29, 41]) flip the answer at 7.5827 bits. Three blocked reads spread evenly ([5, 23, 47]) do not move it at all, at 0.0013. Same number of blocked reads, opposite outcome. An earlier sweep had only ever blocked nested groups — the first k layers or the last k — which cannot separate position from count, and its “roughly half the reads” reading does not survive.

Try the eight switches

SUFFICIENT IS NOT NECESSARY

Layer 29 flips the answer by itself, and it is still not the mechanism. The three latest layers ([35, 41, 47]) flip at 3.7205 bits without it, and the four earliest ([5, 11, 17, 23]) flip at 5.1126 bits containing neither 29 nor any late layer. At least two further routes reach the same outcome, so a single-gate story does not cover them. Fifteen of the 256 possible subsets were run. The rest are unknown, and the study says so rather than interpolating.

EXPERIMENT & PROVENANCE ↗
04

The answer changed. The source stayed.

The intervention changes later state. It leaves the stored source untouched. A persistent decision gives us no licence to free the original memory.

SOURCE CACHE
Kidenticalmax difference 0.0
Videnticalmax difference 0.0

All eight global layers.

LATER STATE / TRANSPLANT
K→ no flip
V→ flip

Value-only transplant is sufficient · ≈49 KB.

After the layer 29 intervention, divergence appears at layers 35, 41 and 47. The source’s own stored rows are unchanged.

Authority
≠ deletion.

Decision changed. Source bytes did not.

FIELD NOTES / 04 MEASUREMENTS & QUALIFICATIONS +

04 / AND THE SOURCE WAS NEVER TOUCHED

Blocking a source briefly during one question changes what the model answers to a later question, after the block is switched off. That could mean the source was permanently demoted — which would be a licence to delete it. Two runs, differing only by a mask at layer 29 across four tokens, were compared row by row.

EVIDENCE

THE STORED SOURCE IS BIT-IDENTICAL

Over the source span, the two caches agree exactly at all eight global layers: the largest difference in either the key or the value stream is 0.0. The source sits 55,600 tokens from the query, far outside the sliding window, so the global layers are the only path to it and the coverage is complete for this claim. Divergence appears only above layer 29 — at 35, 41 and 47 — exactly as masking a layer’s output but not its write predicts.

SUPPORTED

EVIDENCE

THE DECISION TRAVELS IN THE VALUE STREAM

Transplanting the masked run’s rows into the unmasked cache, at fixed positions with no tokens added or removed, flips the later question to the promoted value. The boundary rows and the model-turn rows are each independently sufficient and neither is necessary. A value-only transplant flips it; a key-only transplant does not. The sufficient object is about 49 KB.

SUPPORTED

AUTHORITY IS NOT DELETION

  • The decision persists into the next question. The old evidence is not dead — the demoted source still answers correctly whenever nothing competes with it.
  • Attention-time masking structurally cannot write to a source’s own stored rows, so no amount of scoping turns this into a cache saving.
  • Substituting either carrier region alone leaves the effect intact, so a runtime cannot revoke the decision by rewriting one site.

A persistent decision and a freed byte are different claims.

EXPERIMENT & PROVENANCE ↗
05

A good walk can take the wrong path.

Move one unrelated sentence. The same question now follows a different chain. The reference itself is unstable, so the composition test cannot be scored.

LAYOUT A / SENTENCE AT 66%
FACILITYMeridian start
CODE7431 correct
REGIONNorth correct
SUPERVISOROstran ×wrong
LAYOUT B / SENTENCE AT 10%
FACILITYMeridian start
CODE7431 correct
REGIONSouth ×wrong
SUPERVISORCorvin ×wrong
Expected: Meridian → 7431 → North → Ilex. Only one unrelated 11-token sentence moves between layouts.

The walk was correct.
The address was not.

South → Corvin correctly traverses the decoy chain. An unstable reference makes the composition result unavailable.

FIELD NOTES / 05 MEASUREMENTS & QUALIFICATIONS +

05 / AND WHEN THE ANSWER IS A KEY

Everything above concerns one fact. A planner needs more: if resolving A qualifies, and resolving B from A’s result qualifies, does walking A then B qualify too? A four-link chain tested it — facility to code, code to region, region to supervisor, supervisor’s clearance against a threshold — with a parallel decoy chain running alongside so no step is answerable by elimination.

ONE OBJECT · TWO INTERPRETATIONS

One 64,000-token context · two layouts of the same corpus

Unrelated sentence at 66%Unrelated sentence at 10%
  • Hop 1 · 7431 — correct
  • Hop 2 · North — correct
  • Hop 3 · Ostran — wrong
  • Two of three hops survive

One 64,000-token context · two layouts of the same corpus — as unrelated sentence at 66%: hop 1 · 7431 — correct, hop 2 · north — correct, hop 3 · ostran — wrong, two of three hops survive. As unrelated sentence at 10%: hop 1 · 7431 — correct, hop 2 · south — wrong, hop 3 · corvin — wrong, a coherent walk of the decoy chain.

THE PRIMARY RESULT IS UNAVAILABLE, NOT REFUTED

The reference walk — full access, nothing retired, nothing promoted — is not stable. The true and decoy links sat at identical positions in both layouts; moving a single unrelated 11-token sentence flipped hop two from North to South and hop three from Ostran to Corvin. Arms scored against a reference that flips are measuring a coin toss, so the composition question could not be answered here. A short-context screen had admitted the first three hops; that admission did not transfer to 64,000 tokens.

THE WALK WAS CORRECT. THE ADDRESS WAS NOT.

South to Corvin is a correct traversal — of the wrong chain. The relational operator works; it starts from the wrong place. At long range the model loses the key, not the ability to follow an edge. That points somewhere specific: resolve the path deterministically against an index outside the model, materialise a compact typed closure, promote it once, and let the model perform only the local operation. The run’s own depth-one result is consistent with that division — stepwise promote-then-retire was indistinguishable from replaying the resolved closure in one go, so the stepwise machinery bought nothing.

LARQL / the work

EXPERIMENT & PROVENANCE ↗
06

A late record takes a different route.

On the tested secondary destination, the document’s native binding succeeds at all three seeds. Delivering the same sentence late reproduces it at none.

NATIVE DOCUMENT3 of 3 seeds

reproduce the native route

IDENTICAL SENTENCE, DELIVERED LATE0 of 3 seeds

reproduce the native route

Secondary destination only. The primary destination failed its corpus gate and is disqualified.

Lateness changes
the route.

INSTRUMENT RULE 01

Filler is not neutral.

INSTRUMENT RULE 02

Certify the reference before scoring disagreement.

FIELD NOTES / 06 MEASUREMENTS & QUALIFICATIONS +

06 / SAYING IT LATE IS NOT SAYING IT NATIVELY

A later run asked whether a promoted record reproduces the document’s own behaviour, measured against what the document itself achieves. On the cell where a native ceiling existed, the document’s own binding supported a value-conditioned route at three seeds out of three; the identical sentence delivered late reproduced it at none, landing on a competing destination in five of six cells. The effect tracks lateness rather than provenance — a late in-document correction reaches the same competing destination — so this is not a lossy injection channel. It is a second route being activated alongside the native one.

TWO INSTRUMENT RULES EARNED THE HARD WAY

  • Filler is not automatically neutral: an arm padded with repeated “Noted.” was wrong where contentful but irrelevant assertions at the same budget were right.
  • A reference arm must be certified before disagreement with it counts as failure — controls returning the true answer were flagged as leaking because the baseline was itself wrong.
  • The write-corpus gate failed at two of three seeds, so the primary destination is disqualified and only the secondary contrast is read.

Most of an experiment’s cost is spent earning the right to believe its comparison.

EXPERIMENT & PROVENANCE ↗
OPEN

What makes a source overridable with no intervention at all?

A later run found a second fact, in the same document, where a promoted record overrides a live source unaided — so the regime measured across this whole arc does not exist for that fact. The corpus contains a natural experiment: some decoy values are planted in a type-correct role and one is not. Whatever separates them is the primitive worth building, because that path needs no attention intervention at all.

Retirement turned out to be the mechanism rather than the optimisation, and addressing rather than traversal turned out to be the bottleneck. Both point at the same engineering question the software has been circling: which operations should a model be asked to perform, and which belong to an index outside it.

SOURCES & PROVENANCE

  • Film / 370,000 tokens loaded in Context in 2.8MB, on a MacBook. · 19:30

    Chris Hay, 24 March 2026. The Apollo document passage supplies this note’s opening question, not its measurements.

  • Late promotion into already-built state

    Experiment EXP-20260804-222042-00683. Gemma 3 12B, MLX, 65,536-token context; a 13-token record injected roughly 49,000 tokens downstream of the source, with the span excluded. Three operations, five injection conditions, negative and causal controls on every row. Private register summarised for this draft; raw records are not republished.

  • Supply versus override

    Experiment EXP-20260805-113941-00687. One question, one model, 72-token injection budget, zero to three companion records, single-token answer. The 14.7914-bit figure is a cross-run comparison on a byte-identical record, not an arm of this experiment.

  • Partial retirement frontier

    Experiment EXP-20260805-125720-00688. Nested prefix and suffix sweeps over the eight global layers. Its own conclusion records that a cumulative design identifies a sufficient set, not the responsible layer, and that its “late-layer” label was wrong.

  • Layer attribution

    Experiment EXP-20260805-133312-00689. Eight single-layer arms and four count-matched triples on a byte-identical corpus. Fifteen of 256 possible subsets were run. Layer 29 is sufficient and not necessary; the residue is unresolved.

  • Commit localisation

    Experiment EXP-20260805-170039-00695. Divergence map, eviction and transplant phases. The eviction phase is uninformative and forms no part of the claims above; the transplant result carries them.

  • Composition along a relational walk

    Experiment EXP-20260805-100011-00684. Two 64K layouts differing only in the position of one unrelated 11-token span. The composition contrast is unavailable rather than refuted; the single-edge and operand-binding results replicated across both layouts.

  • Write versus document, against a native ceiling

    Experiment EXP-20260805-232158-00701. Three seeds, three-rung ladder. The primary destination is disqualified by its own corpus gate; the channel comparison on the secondary destination does not depend on that gate.

  • Film / We Don’t Need KV Cache Anymore? · 13:20

    Chris Hay, 10 March 2026. Background on retained state and rebuilding the attention cache; its figures describe its own setup.

AUTHOR / CHRIS HAY · VERSION / 0.1

REFERENCE THIS DRAFT

An unpublished working record. These references identify the draft and omit a publication date. They become version-specific publication citations when the record is released.

Chris Hay. Which source wins? [Unpublished draft, version 0.1]. https://chrishayuk.com/notebook/which-source-wins
DOWNLOAD