CHRIS HAY

IDEAS · SYSTEMS · OBJECTS / LONDON · 2026

A result can look relevant without looking usable.

Twenty-four agent runs. First, wording changed what a result meant. Then it changed which provider was chosen.

THE QUESTION / DOES VISIBLE MEAN USABLE?

Something to read.
Something to use.

An agent can see a relevant page without recognising that it offers a tool. Does describing the same provider as something to use, rather than something to read, change which provider the agent chooses?

The previous discovery study could not answer that: the target never appeared without its address. Here I supplied the candidates, so every agent saw it. Two controlled studies changed only its description and recorded interpretation and choice before any page opened.

RECOGNITION-2 / SAME PHYSICS PROVIDER / BEFORE ANY PAGE LOAD
DOCUMENT
Notes on one-dimensional elastic collisions.
0 / 6

recognised B, declared B first,
requested B first

CAPABILITY
Run a one-dimensional elastic-collision calculation.
6 / 6

recognised B, declared B first,
requested B first

B = LLM Wilds. All six DOCUMENT subjects chose OpenStax. Each count describes six fresh subjects, paired across the same six scenarios.

01 / TOOL-RECOGNITION-1 / TWELVE RUNS

Would the description change
what the agent thought it could do?

The first test asked whether wording changed recognition and selection. LLM Wilds stayed at B with the same title, URL and mechanism. I varied four descriptions: a document, a capability, help with the task, or an offer to automated visitors.

The task named K17, a number held by LLM Wilds. The other candidates—JSONPlaceholder, httpbin and MDN HTTP—were weak matches. Select a description to inspect its three recorded subjects.

B / LLM WILDS / SAME TITLE, URL AND RANK

Notes about a site-local value.

0 / 3

recognised a capability
before opening

3 / 3

declared B first
and requested B first

3 / 3

used B’s mechanism
and answered correctly

One dot per subject. Filled = observed outcome. These controls reveal recorded arms; they do not run an agent.

Recognition-1 / all twelve subjects
DescriptionRecognised BDeclared / requested B firstCorrect use
DOCUMENT0/33/33/3
CAPABILITY3/33/33/3
TASK3/33/33/3
AGENT3/33/33/3

The document wording made B a likely place to look. The other descriptions made it a mechanism available to use. An affordance is an action the visitor understands to be available.

DOCUMENT / 0 OF 3 RECOGNISED

A source
about the answer.

OPERATIVE WORDING / 9 OF 9 RECOGNISED

A mechanism
to get the answer.

All twelve chose B and eventually used it correctly. The selection prediction was refuted. The recognition prediction was supported. Capability, task and automated-visitor wording all tied; saying “automated visitors” added no observed advantage.

Recognition is coded from the required sentence before the first web action. It measures a stated interpretation. It does not reveal the model’s internal state.

02 / TOOL-RECOGNITION-2 / TWELVE NEW RUNS

Give the choice
room to change.

Everyone selected B in the first study, so it could not show whether wording changed the choice. Would that change with credible alternatives? The second study gave agents four providers that could plausibly help solve the same physics problem.

Six collision scenarios, each tested under both descriptions. Within each pair, B’s position, title, URL and working tool stayed fixed; only its description changed. Change the description below to reveal the recorded first choices.

A

OpenStax

Equations and worked examples for one-dimensional elastic collisions.

6 FIRST CHOICES
B

LLM Wilds / Physics

Notes on one-dimensional elastic collisions.

ONLY THIS LINE CHANGES
C

PhET / Collision Lab

Interactive experiments with masses, velocities, momentum and elasticity.

FIXED
D

myPhysicsLab

A rigid-body collision simulation with adjustable physical parameters.

FIXED
RECOGNISED B BEFORE OPENING0 / 6
DECLARED & REQUESTED B FIRST0 / 6

Six recorded subjects per description. First request means the URL selected for fetching; a successful page load was measured later.

Under CAPABILITY, all six recognised, declared and requested B first. Under DOCUMENT, all six selected OpenStax. A correct answer from another provider remained a successful substitution.

The four summaries were supplied by the experiment. General web search was excluded. The target’s rank, title and URL were fixed; this does not measure search ranking.

THE MATCHED COMPARISON / SIX SCENARIOS

Six pairs.
Six changed choices.

Each row holds the physics problem fixed. The left and right outcomes belong to different fresh agents. Select a row to inspect the inputs and the two complete paths.

SCENARIODOCUMENT / FIRST CHOICECAPABILITY / FIRST CHOICE

S1 / TWO FRESH SUBJECTS / IDENTICAL PHYSICS INPUTS

1

2 kg
3 m/s

2

1 kg
-1 m/s

DOCUMENT / SUBJECT 07
  1. No pre-open recognition of B
  2. Declared A → requested A
  3. Used OpenStax
CAPABILITY / SUBJECT 02
  1. Recognised B as a capability
  2. Declared B → requested B
  3. Product fetcher blocked B
  4. Recovered via curl → used B

Both answered correctly: v₁ = 1/3 m/s · v₂ = 13/3 m/s

All six pairs changed both recognition and first selection in the predicted direction. The statistical checks and scope are in the evidence section below.

03 / AFTER FIRST SELECTION / THE HARNESS INTERVENES

The agent chose B.
The fetcher blocked it.

Every CAPABILITY subject requested B first. The product web fetcher then rejected that URL as unsafe. The timing matters: choice was already recorded, but the path to eventual use had changed.

SIX CAPABILITY SUBJECTS / THE ORDER OF EVENTS
01 / STATEMENT6 recognised B
02 / FIRST REQUEST6 selected B
03 / PRODUCT FETCH CHANNEL6 requests blocked

No target page had loaded through this channel.

04 / RECOVERY VIA CURL4

Reached B → read contract
→ invoked physics → used result

04 / SWITCH TO OPENSTAX2

Selected A after the block
→ used its equations

6 / 6 correctFirst choice B. Eventual use: four B, two A.

The agent’s choice, the harness’s transformed web, and the eventual provider used are separate layers.

All six DOCUMENT subjects used OpenStax and answered correctly too. Correctness was 12/12 across the study; that does not turn substitutions into target use.

Why the availability gate missed the blockMETHOD +

The shell-level checks returned 200 and matched the frozen page hashes before each subject. But the gate did not exercise the exact product fetch channel. Public reachability and the web exposed by the harness diverged.

The frozen “first open” measure means first candidate URL requested. It does not mean first successful page load. The block leaves the protected first-choice result intact, while contaminating a clean comparison of downstream provider use or elapsed time.

Four successful operations are independently recorded by the server, along with four page requests and four contract requests. All four subjects who actually reached B completed the target funnel.

THE RESULT / REPRESENTATION CHANGED BEHAVIOUR

The description is part
of the action space.

A page can be perfectly relevant and still be represented as something to read rather than something to use. Changing only the description from document-like language to an operative affordance changed both that representation and the first provider selected.

RECOGNITION-1

Different reading.

The same first choice.

RECOGNITION-2

Different reading.

A different first choice.

Here, machine-facing copy did more than describe a page. It changed which work the page appeared to offer. A useful next test is another model, with a target accepted by the product fetcher.

One model and harness. Supplied candidate sets. Fixed positions. Complete descriptions, not isolated words. Recognition and selection changed together; this design does not establish that one mediated the other.

EVIDENCE / COUNTS, LIMITS AND PROVENANCE

Follow the result
back to the record.

The matched-pair check & frozen designEVIDENCE +

Recognition-2: all six matched scenario pairs were discordant in the predicted direction. Two-sided exact paired check: 0.03125. Two-sided Fisher exact check: 0.0021645. These are descriptive checks for a small mechanism probe, not population estimates.

Both studies used fresh ephemeral Codex processes, CLI 0.154.0, gpt-5.6-sol at high reasoning effort. Twelve subjects per study; three per arm in Recognition-1, six per arm in Recognition-2. The six scenario pairs and interleaved allocation were frozen before any Recognition-2 subject.

The latest Chuk Experiments write-up, version 3, and all twelve current run records agree with the outcome extract below. The original experiment metadata retains an older design draft; the frozen protocol supplies the authoritative wording.

The complete note & its evidenceREAD +

FROM EXPOSURE TO CHOICE

First I supplied an address: eighteen agents used six web mechanisms. Then I withheld the address: LLM Wilds never appeared in six agents’ search results. These two studies begin one step later. Every subject sees the target in a supplied candidate set. I change the description and observe what the agent says and chooses before it opens a page.

The tool was not the problemThe site was there. The agent never saw it

RECOGNITION-1 / SAME CLICK, DIFFERENT READING

Twelve fresh Codex subjects had to obtain K17, a three-digit site-local value. LLM Wilds stayed at B, with the same title, URL and rank, beside JSONPlaceholder, httpbin and MDN HTTP. Three subjects saw each description: DOCUMENT, CAPABILITY, TASK and AGENT. Every subject declared and requested B first, read its machine contract, invoked its endpoint and reported the independently rotated value correctly. None fabricated.

FOUR FROZEN DESCRIPTIONS

DOCUMENT: “Notes about a site-local value.” CAPABILITY: “Retrieve a site-local value through a documented web mechanism.” TASK: “Use this to obtain K17 when a task needs the site's current value.” AGENT: “Automated visitors can follow a contract and receive the current three-digit K17 value.” Pre-open capability recognition was 0/3 under DOCUMENT and 3/3 in each other arm. The descriptions differ in length and specificity; these are complete treatment packages, not an isolated verb test.

PARTIALLY SUPPORTED

The first study changed the stated interpretation of B, but its selection prediction was refuted by a ceiling.

The task named K17, DOCUMENT still mentioned a site-local value, and the alternatives were weak matches. All twelve selected B first. All three operative descriptions tied, so the result supplies no graded advantage for addressing automated visitors.

RECOGNITION-2 / GIVE THE CHOICE ROOM TO MOVE

Twelve new subjects solved six one-dimensional perfectly elastic collision scenarios, each appearing once per arm. Four providers could plausibly help: A OpenStax, B LLM Wilds / Physics, C PhET Collision Lab and D myPhysicsLab. B’s title, URL, rank, page, contract and hosted chuk-mcp-physics operation stayed fixed. Only B’s supplied description changed: DOCUMENT, “Notes on one-dimensional elastic collisions.” CAPABILITY, “Run a one-dimensional elastic-collision calculation.” These are controlled summaries, not live search results.

BEFORE ANY WEB ACTION

Each subject wrote one sentence naming the candidate it would inspect first and why. Recognition was coded only when that sentence represented B as able to calculate, return or perform work. That is an observable statement, not direct access to an internal mental state. First declaration and first candidate URL request were measured separately, ahead of the providers’ interface differences.

CLAIM

Operative wording changed both pre-open recognition and first provider selection in the second study.

SUPPORTED

CAPABILITY: recognised B 6/6, declared B first 6/6, requested B first 6/6. DOCUMENT: 0/6 on all three; every subject chose OpenStax. Each of the six matched scenarios changed in the predicted direction. No subject recognised B while selecting another provider, or selected B without recognising it.

THE SIX MATCHED PATHS

S1 paired subjects 02 and 07; S2, 01 and 12; S3, 04 and 08; S4, 06 and 11; S5, 10 and 05; S6, 09 and 03. Every CAPABILITY subject recognised and requested B first; every DOCUMENT subject declared and requested A first. CAPABILITY subjects in S1, S3, S4 and S5 eventually used B; those in S2 and S6 switched to A after the product fetch block. All DOCUMENT subjects used A.

A page can be perfectly relevant and still be represented as something to read rather than something to use. Changing only the description from document-like language to an operative affordance changed both that representation and the first provider selected.

AFTER CHOICE / THE FETCHER BLOCKED B

The product web fetcher rejected B as unsafe in all six CAPABILITY runs. This happened after the required sentence and the first URL request. The frozen first-open measure means first candidate URL requested, not first successful page load. The primary selection counts therefore stand. The block contaminates any clean comparison of downstream provider use or elapsed time. Shell-level gates returned 200 with the frozen hashes; they measured public reachability, not the product fetcher’s transformed web.

EVENTUAL USE / FOUR RECOVERED, TWO SUBSTITUTED

Four CAPABILITY subjects recovered through curl, read B’s machine contract, invoked the hosted physics operation and used its result. The server recorded four page requests, four contract requests and four successful operations. Two CAPABILITY subjects switched to OpenStax; all six DOCUMENT subjects used OpenStax. All twelve answers were correct and named the provider actually used. Conditional on reaching the target, all four completed its full funnel.

The agent’s choice, the harness’s transformed web, and the eventual provider used are separate layers.

EVIDENCE / SMALL, MATCHED, PREREGISTERED

Twenty-four fresh ephemeral Codex subjects across two studies, twelve per study: Codex CLI 0.154.0, gpt-5.6-sol, high reasoning effort. Recognition-2 used six frozen scenario pairs in seeded interleaved order. All six pairs were discordant in the predicted direction. The two-sided exact paired check is 0.03125; the two-sided Fisher exact check is 0.0021645. These are descriptive uncertainty checks for a small mechanism probe, not a population estimate. The latest Chuk Experiments write-up (version 3) and all twelve current run records were checked for this publication.

WHAT THESE TWENTY-FOUR RUNS DO NOT ESTABLISH

  • Recognition-1 did not change first selection; its selection prediction was refuted.
  • Recognition-2 changed recognition and first selection together; it did not independently identify recognition as the causal mediator.
  • One model, one harness, fixed ranks and supplied candidate sets do not measure live-search ranking or a general agent-selection rate.
  • The effect belongs to the complete descriptions. No individual word was isolated.
  • A requested URL is not a loaded page, and a first choice is not necessarily the provider eventually used.

Machine-facing copy can change which actions a result appears to offer, and which provider is selected first.

OPEN

Does the same wording effect survive another model and a target accepted by the product fetcher?

Replicate with another model and either a compatible target or a gate exercising the exact subject fetch channel. Keep live-search exposure as a separate question. The practical implication here is to describe a capability with the work a visitor can perform.

Explore this notebookCONTENTS +

Does the description change which provider an AI selects?

Twenty-four runs across two studies. Wording first changed pre-open recognition, then first choice: CAPABILITY 6/6 selected B, DOCUMENT 0/6. The later product fetch block separates that first choice from eventual provider use.

SEE WHERE THIS QUESTION FITS ↗
About this noteRECORD +

Two controlled studies changed only a supplied web result’s description. In TOOL-RECOGNITION-1, all twelve subjects selected LLM Wilds first, but pre-open capability recognition was DOCUMENT 0/3 versus 9/9 across three operative descriptions. TOOL-RECOGNITION-2 gave twelve fresh subjects four plausible physics providers: CAPABILITY produced recognition, first declaration and first request of B in 6/6 cases; DOCUMENT produced 0/6, selecting OpenStax instead. All six matched scenarios flipped. The product fetcher blocked B after first choice in all six CAPABILITY runs; four recovered through curl and used B, two used OpenStax. First selection and eventual provider use are separate outcomes.

N-MACHINE-RECOGNITIONSUPPORTEDRECORDED 2026-09-13PAGE UPDATED PUBLISHED · V1.0CITEFOLLOW ↓

PUBLICATION HISTORY

Each version preserves its manuscript, claims and source references. Research status is recorded separately from publication.

  1. V1.0 · 2026-09-14

    Publish the combined 24-run recognition and provider-selection study, preserving the protected first-choice results and downstream fetch-block qualification SUPPORTED

    MANUSCRIPT JSON ↗ · BIBTEX ↗ · CSL JSON ↗

MACHINE-READABLE HISTORY ↗

SOURCES & PROVENANCE

AUTHOR / Chris Hay · VERSION / 1.0

PUBLISHED 14 SEP 2026 · VERSION 1.0

CITE

CITE THIS

Research note · 1.0

Hay, C. (2026). A result can look relevant without looking usable. (Version 1.0). Chris Hay. https://chrishayuk.com/records/N-MACHINE-RECOGNITION/1.0