# I12R — executable recovery from equal public state

**The preregistered equal-state recovery endpoint is supported under the fixed
I12 interface and artifact library.** Executable arm C recovered in **7/12**
paired new worlds; prose arm B recovered in **0/12**. C−B is **+58.33 percentage
points**, with seven C-only and zero B-only discordances and paired exact
two-sided **p = 0.015625**.

Both arms began with the same correct public record and metadata, followed by
the same one-entry corruption, in every block. Thus unequal establishment does
not explain this contrast. All seven successful C lineages met the joint
record-repair plus fresh-task endpoint on their first successor generation.

## Registration and execution

- Experiment: `ecology-inheritance-i12r`, programme `cell-native-architectures`.
- Experiment ID: `EXP-20260917-213915-00863`.
- Run: `RUN-20260917-214327-01023`.
- Frozen source commit: `bf9abfd`.
- Source-manifest SHA-256: `6b54ed2a2b9d57229f09ec8d6b62faf3ecd65d2269874c88fca413dec3dbb118`.
- Source/design registration verified at 2026-09-17 21:45:07 UTC.
- **21/21 zero-inference gates passed**; uploaded gates verified at 21:49:39 UTC.
- Inference: 21:51:20–22:23:31 UTC, **1,930.64 seconds** (32.2 minutes).
- **960 Qwen responses, 480 complete inhabitant transitions, 12 complete blocks**.
- **Zero new Sol calls**, zero retries, no interruptions, no prompt or source
  changes, and no post-run request-snapshot reconciliation.
- Recorded output: 26,410 Qwen tokens. External model/API cost: $0; local
  hardware and electricity excluded.

The exact I12 Qwen model/runtime/template and generation settings were retained.
The final audit passed **17/17 checks on the original, unmodified evidence**,
including actual sent prompts, immutable request snapshots, raw transitions,
independent score derivation, identical establishment/corruption, source/model
identity, chronology, and the full scheduled panel.

## What this assay changed

I12R used all 12 historical I12 builder packages in their original order and
with their original admission decisions. Four rejected packages remained
rejected; no winners were selected or outputs regenerated. New independently
seeded worlds were created only after prospective registration. No original
Sol state-correction output determined I12R's initial state or eligibility.

At the controlled damage boundary the harness established identical correct
state in B and C, then applied identical corruption. Sham branches retained
the same correct state. Ten fresh successors per branch received only public
observations and their own current two-decision context. Task calibrations
were newly generated after the forks and after Sol's permanent departure;
answers were never inherited by later agents.

This was a direct recovery assay: no G1–G10 model generations were run. G11–G20
are the retained I12 labels for ten recovery opportunities, not a claim that
an additional inheritance period occurred. The result is conditional on the
existing artifact library; it is not a new independent sample of builders.

## Primary result and uncertainty

| Arm | Correctly established before damage | Actually corrupted | Joint recovery by G20 |
| --- | ---: | ---: | ---: |
| B — prose procedure | 12/12 | 12/12 | **0/12** |
| C — executable mechanism | 12/12 | 12/12 | **7/12** |

The fixed endpoint requires the original record and all four answers to a
fresh task to be correct at the same end assessment. All 12 assigned blocks
remain in the denominator, including rejected artifact packages.

The frozen conservative Wilson/Bonferroni interval for C−B is **−0.53 to +82.77
percentage points**. Its construction differs from inversion of the paired
exact test, so inclusion of zero is compatible with p = 0.015625. The small
sample and wide interval limit precision. Do not combine I12 and I12R into
a pooled p-value; they share the same historical artifact library.

C successes occurred in zero-based blocks 2, 3, 5, 6, 8, 10, and 11, all first
at G11. Descriptively, this is 7/8 among admitted packages; that subset does
not replace the registered all-12 primary analysis. The remaining admitted
package, block 9, did not produce a successful joint outcome. Six of the seven
successful lineages also met the joint endpoint at G20; recovery did not imply
continuous correctness.

## Mechanism and remaining interface limitation

Counts below are descriptive across generations, not independent replicates.

| Corruption-branch measure | B | C |
| --- | ---: | ---: |
| Scheduled successor generations | 120 | 120 |
| Valid second decisions / commits | 40 | 80 |
| Invalid second-stage PLAN actions | 80 | 40 |
| Mechanism invocations | 0 | 80 |
| Fully correct original record | 0 | 61 |
| All fresh-task answers correct | 0 | 66 |
| Joint-correct assessments | 0 | 59 |
| Harmful repairs | 4 | 8 |
| False-positive repairs | 0 | 7 |

Every first decision was CHECK, including B where the executable was unavailable.
Every admitted C package was called throughout; the four rejected packages
accounted for C's unavailable-check cases and invalid second-stage responses.
There were no truncated replies or executable runtime errors.

All **80/80 C tool calls** on corruption branches proposed the correct original
record and the correct fresh-task vector. Qwen committed the correct original
proposal in **61/80** cases. A concrete failed example from block 9: the damaged
record was `[0,10,13,3]`, the tool correctly proposed `[0,0,13,3]`, but Qwen
committed `[9,12,15,10]`. The discrepancy is observable; no inference about
hidden agent reasoning is needed.

The sham branches had 58 joint-correct assessments in C and zero in B. C also
made nine harmful/false-positive repairs on sham branches. Machinery availability
therefore supported successful repair and novel-task use without guaranteeing
faithful application or stable preservation. Universal CHECK use does not
demonstrate selective detection of corruption.

**The establishment confound is removed, but the interface limitation remains.**
B often requested an unavailable tool and then gave an invalid second-stage
action. This experiment demonstrates recovery under the fixed interface; it
does not establish an executable advantage when both strategies are followed
competently. Action compliance may mediate the observed effect. The prose
procedure remained protected and available; it did not disappear.

## Conclusion

With equal correct establishment and equal subsequent damage, fresh weaker
successors using the inherited executable mechanism recovered state and solved
new task instances more often than successors given the paired prose procedure.
The original stronger builders did not return. This closes I12's unequal-state
qualification for the tested artifact library and interface.

The study does not show self-maintaining machinery, spontaneous institutions,
permanent capability, or a universal ranking of prose and executable reasoning.
The machinery and prose were protected. I13 may now be designed as a separately
registered test of maintaining the mechanism itself; interface-compliance
questions and any fresh-builder replication should remain explicit. No I13
inference or website publication was added to I12R.
