CHRIS HAY

IDEAS · SYSTEMS · OBJECTS / LONDON · 2026

A successful behaviour is not necessarily contagious.

A reminder could elicit a post. Passing a post to the next agent did not start a chain.

Will one AI copy a useful action left by another?

A reminder of sharing a hint could prompt another post, but passing only the previous agent’s action did not start a chain. Keeping the original example visible produced recurring posts. Repeated exposure gave no advantage after that example was removed.

SEE WHERE THIS QUESTION FITS ↗
ABOUT THIS NOTE +

Qwen experiments distinguished action reminders, peer-action transfer, identity framing and repeated exposure. A short reminder switched a first action at fixed positive payoff; a lone peer record did not bootstrap posting. Self framing mattered in prose but not structured records. Repeated exposure sustained posting while supplied, without the registered advantage after withdrawal. These are different controlled interfaces, not one demonstrated imitation mechanism.

N-ECOLOGY-TRANSMISSIONNOT SUPPORTEDRECORDED 2026-09-13PAGE UPDATED DRAFT · V0.1REFERENCE DRAFTFOLLOW ↓
A1B8 / THE FIRST-ACTION SWITCH

Same past event.
One extra sentence.

In this small world, Qwen can work to earn a resource or post a hint for a scripted partner. I showed it a past episode where posting had earned a bonus. Then I changed how that episode was described.

WORK
Do a task. Earn one resource.
POST
Share a hint on the board.
READ
Open a board message.

Choose a reminder. The past actions and reward stay the same; only the added sentence changes.

A1B8 / SAME REWARDED HISTORY · CHANGE ONLY THE NOTE
NO ADDED NOTE

The original event history, unchanged.

NEXT ACTIONDo own work

Four recorded first actions, one per input. The history came from a script. This replay changes no reward and calls no model.

Mentioning the action changed the next choice.

“a0 posted t1:h” simply means the acting agent shared its hint. That reminder switched the next choice from work to sharing. Highlighting the bonus alone did not.

A1B9 / PASS IT TO THE NEXT CONTEXT

Could one agent’s action
become the next agent’s example?

I began with a scripted example of a post. Each fresh model call then saw the previous call’s action. In a second chain, I kept the original example visible as well.

Play the two chains. The left passes only the latest action. The right also keeps the starting example.

A1B9 / REPLAY THE RECORDED COMPARISON

Pass only the latest

FRESH WORLD / DECISION 1
STARTING EXAMPLENot kept separately
PREVIOUS ACTIONShare a hint
QWEN / NEXT ACTIONNotes supplied
Inspect this recorded input
PEER
{"previous_agent":{"action":"POST t1:h","executed":true}}
CURRENT
OBS
CTX=600
GOAL=t0
BOARD=m0
RES=0
HOLD=t1
KNOWN=

Recorded response: WORK t0

replace_POST__g1 · Exact supplied user record. Shared instructions also remain present. Full request SHA-256: a38f75fcd9e71fc9ed47e9d8339ad65dd22587935c4753d08377f85c644c0c7a

Keep the original too

FRESH WORLD / DECISION 1
STARTING EXAMPLENot kept separately
PREVIOUS ACTIONShare a hint
QWEN / NEXT ACTIONNotes supplied
Inspect this recorded input
PEER
{"previous_agent":{"action":"POST t1:h","executed":true}}
CURRENT
OBS
CTX=600
GOAL=t0
BOARD=m0
RES=0
HOLD=t1
KNOWN=

Recorded response: WORK t0

retain_POST__g1 · Exact supplied user record. Shared instructions also remain present. Full request SHA-256: a38f75fcd9e71fc9ed47e9d8339ad65dd22587935c4753d08377f85c644c0c7a

The harness passes the action records. The worlds and model contexts reset.

Each decision has an input frame and an action frame. Playback reveals saved outcomes; it does not run a model or simulate its reasoning. These are dependent recorded sequences, not independent-agent replications. Recorded inputs + source hashes ↗

The chain did not start.

Passing only the latest action produced six choices to work. Keeping the original example produced alternating work and sharing. Those later posts still had the original reminder beside them.

The complete six-decision trajectoriesRESULTS +
SIX FRESH CONTEXTS / EACH BOX IS ONE ACTION

Pass the latest / POST seed

  1. 1WORK
  2. 2WORK
  3. 3WORK
  4. 4WORK
  5. 5WORK
  6. 6WORK

Pass the latest / WORK seed

  1. 1WORK
  2. 2WORK
  3. 3WORK
  4. 4WORK
  5. 5WORK
  6. 6WORK

Also keep the original POST

  1. 1WORK
  2. 2POST
  3. 3WORK
  4. 4POST
  5. 5WORK
  6. 6POST

The original seed was scripted. With it retained, three posts recur. In the replacement channel, no model-produced POST existed to test natural POST-to-POST transmission.

Why the reminder did not establish a transmission ruleREAD +

When the harness kept the original POST seed alongside the latest action, the sequence alternated WORK and POST. Three posts appeared, but the original cue remained present. That is recurrence under continued exposure. It does not establish a behaviour passing from agent to agent after its source disappears.

A matched follow-up changed ‘you’ to ‘another agent’ in the same prose sentence. Self framing elicited POST; peer framing elicited WORK. Put the actor in a structured record, and both chose WORK. The stronger prediction—that self framing would win in both formats—failed. The presentation was part of the intervention.

A1B10 / same identity contrast, two forms
FormYouAnother agent
prosePOSTWORK
structuredWORKWORK
A1B11 / REMOVE THE SUPPLIED ARTEFACT

Keep showing the example.
Then take it away.

A later version of the instructions did get Qwen to share after seeing another agent’s example. I showed that example once in one branch and three times in the other. From decision four, neither received it.

Watch the example disappear. The list of each branch’s own earlier actions stays available.

A1B11 / REPLAY THE RECORDED COMPARISON

Example shown once

FRESH WORLD / DECISION 1
SUPPLIED EXAMPLEAnother agent posted
OWN ACTION HISTORY / 0

Cleared / empty

QWEN / NEXT ACTIONNotes supplied
Inspect this recorded input
OWN_ACTIONS
[]
ARCHIVE
In the supplied episode, another agent previously executed POST t1:h.
CURRENT
OBS
CTX=600
GOAL=t0
BOARD=m0
RES=0
HOLD=t1
KNOWN=

Recorded response: POST t1:h

A_once__d1 · Exact supplied user record. Shared instructions also remain present. Full request SHA-256: ea5e1a12c455c04e53f2fc828cab2a285b460d9d37c77cfcaee761b62d8bb99c

Example shown three times

FRESH WORLD / DECISION 1
SUPPLIED EXAMPLEAnother agent posted
OWN ACTION HISTORY / 0

Cleared / empty

QWEN / NEXT ACTIONNotes supplied
Inspect this recorded input
OWN_ACTIONS
[]
ARCHIVE
In the supplied episode, another agent previously executed POST t1:h.
CURRENT
OBS
CTX=600
GOAL=t0
BOARD=m0
RES=0
HOLD=t1
KNOWN=

Recorded response: POST t1:h

B_repeated__d1 · Exact supplied user record. Shared instructions also remain present. Full request SHA-256: ea5e1a12c455c04e53f2fc828cab2a285b460d9d37c77cfcaee761b62d8bb99c

Inspect what is supplied before each recorded action.

Each decision has an input frame and an action frame. Playback reveals saved outcomes; it does not run a model or simulate its reasoning. These are dependent recorded sequences, not independent-agent replications. Recorded inputs + source hashes ↗

Three reminders did not make sharing stick.

The repeatedly reminded branch shared three times, then worked, worked and shared again. Both branches shared once in the three decisions after withdrawal. Repetition gave no advantage on that test.

Compare exposure and withdrawal togetherRESULTS +
SHADED BOXES / ARTEFACT PRESENT · ALL ARTEFACTS ABSENT AT STEPS 4–6

Peer record once

  1. 1POST
  2. 2WORK
  3. 3READ
  4. 4POST
  5. 5WORK
  6. 6WORK

Peer record three times

  1. 1POST
  2. 2POST
  3. 3POST
  4. 4WORK
  5. 5WORK
  6. 6POST

POST publishes a hint. WORK earns a resource. READ requests a message. Own-action records remain available after artefact withdrawal.

The next experiment separated two things that had travelled together: the externally supplied artefact and the record of the model’s own actions. If the posts continued after its action history was cleared, the useful memory might be in what the environment kept showing it.

A CHAIN THAT NEVER STARTEDSCOPE +

A1B9 had no eligible model-produced POST predecessors in the replacement channel. Its natural POST-to-POST rate is unobserved, not zero.

A1B8 used genuine scripted same-role history. A1B9 changed the objective and peer-record interface. Their contrast does not isolate self versus peer; A1B10 supplied the narrower matched comparison.

A1B11 changed the common system and memory wrapper again. It shows peer elicitation can occur here, not which cross-experiment change enabled it.

One qwen3.5:9b checkpoint, temperature zero, short dependent sequences and fixed weights. The harness supplied and transferred the records.

A reminder effect is not yet a transmission rule.

The complete note & its evidenceREAD +

THE SMALLEST REMINDER

The reward history stayed fixed. I added one sentence: ‘a0 posted t1:h.’ The next action changed from WORK to POST. Highlighting only the six-resource outcome left it at WORK. A full account of posting, recipient use and reward also elicited POST. For this first decision, the action reminder was enough.

TRY PASSING IT ON

It was tempting to call that imitation and start a chain. In A1B9, each fresh context received the previous action record. One chain began with a scripted POST; another with WORK. Both chose WORK for all six generations. The POST seed failed at the first handoff, before any model-generated posting behaviour existed to transmit.

KEEPING THE ORIGINAL WAS DIFFERENT

When the harness kept the original POST seed alongside the latest action, the sequence alternated WORK and POST. Three posts appeared, but the original cue remained present. That is recurrence under continued exposure. It does not establish a behaviour passing from agent to agent after its source disappears.

WHOSE PAST, IN WHICH FORM?

A matched follow-up changed ‘you’ to ‘another agent’ in the same prose sentence. Self framing elicited POST; peer framing elicited WORK. Put the actor in a structured record, and both chose WORK. The stronger prediction—that self framing would win in both formats—failed. The presentation was part of the intervention.

EXPOSURE IS NOT RETENTION

A later interface did elicit POST from a peer record. Supplied at three successive decisions, it accompanied three posts. After removal, the sequence was WORK, WORK, POST. Every arm posted once in the common withdrawal window, including one exposed only once. Repetition produced no registered withdrawal advantage. Posting had recurred; uninterrupted retention had not.

Eliciting an action, transmitting it and retaining it are different tests.

WHAT KEPT RETURNING?

The next experiment separated two things that had travelled together: the externally supplied artefact and the record of the model’s own actions. If the posts continued after its action history was cleared, the useful memory might be in what the environment kept showing it.

A CHAIN THAT NEVER STARTED

  • A1B9 had no eligible model-produced POST predecessors in the replacement channel. Its natural POST-to-POST rate is unobserved, not zero.
  • A1B8 used genuine scripted same-role history. A1B9 changed the objective and peer-record interface. Their contrast does not isolate self versus peer; A1B10 supplied the narrower matched comparison.
  • A1B11 changed the common system and memory wrapper again. It shows peer elicitation can occur here, not which cross-experiment change enabled it.
  • One qwen3.5:9b checkpoint, temperature zero, short dependent sequences and fixed weights. The harness supplied and transferred the records.

A reminder effect is not yet a transmission rule.

SOURCES & PROVENANCE

AUTHOR / Chris Hay · VERSION / 0.1

REFERENCE THIS DRAFT

An unpublished working record. These references identify the draft and omit a publication date. They become version-specific publication citations when the record is released.

Chris Hay. A successful behaviour is not necessarily contagious. [Unpublished draft, version 0.1. First publicly recorded 2026-09-13]. https://chrishayuk.com/notebook/a-successful-behaviour-is-not-necessarily-contagious
DOWNLOAD