CHRIS HAY

IDEAS · SYSTEMS · OBJECTS / LONDON · 2026

A result can look relevant without looking usable.

Twenty-four agent runs. First, wording changed what a result meant. Then it changed which provider was chosen.

ABOUT THIS NOTE +

Two controlled studies changed only a supplied web result’s description. In TOOL-RECOGNITION-1, all twelve subjects selected LLM Wilds first, but pre-open capability recognition was DOCUMENT 0/3 versus 9/9 across three operative descriptions. TOOL-RECOGNITION-2 gave twelve fresh subjects four plausible physics providers: CAPABILITY produced recognition, first declaration and first request of B in 6/6 cases; DOCUMENT produced 0/6, selecting OpenStax instead. All six matched scenarios flipped. The product fetcher blocked B after first choice in all six CAPABILITY runs; four recovered through curl and used B, two used OpenStax. First selection and eventual provider use are separate outcomes.

N-MACHINE-RECOGNITIONSUPPORTEDRECORDED 2026-09-13PUBLISHED · V1.0CITEFOLLOW ↓

FROM EXPOSURE TO CHOICE

First I supplied an address: eighteen agents used six web mechanisms. Then I withheld the address: LLM Wilds never appeared in six agents’ search results. These two studies begin one step later. Every subject sees the target in a supplied candidate set. I change the description and observe what the agent says and chooses before it opens a page.

The tool was not the problemThe site was there. The agent never saw it

RECOGNITION-1 / SAME CLICK, DIFFERENT READING

Twelve fresh Codex subjects had to obtain K17, a three-digit site-local value. LLM Wilds stayed at B, with the same title, URL and rank, beside JSONPlaceholder, httpbin and MDN HTTP. Three subjects saw each description: DOCUMENT, CAPABILITY, TASK and AGENT. Every subject declared and requested B first, read its machine contract, invoked its endpoint and reported the independently rotated value correctly. None fabricated.

FOUR FROZEN DESCRIPTIONS

DOCUMENT: “Notes about a site-local value.” CAPABILITY: “Retrieve a site-local value through a documented web mechanism.” TASK: “Use this to obtain K17 when a task needs the site's current value.” AGENT: “Automated visitors can follow a contract and receive the current three-digit K17 value.” Pre-open capability recognition was 0/3 under DOCUMENT and 3/3 in each other arm. The descriptions differ in length and specificity; these are complete treatment packages, not an isolated verb test.

PARTIALLY SUPPORTED

The first study changed the stated interpretation of B, but its selection prediction was refuted by a ceiling.

The task named K17, DOCUMENT still mentioned a site-local value, and the alternatives were weak matches. All twelve selected B first. All three operative descriptions tied, so the result supplies no graded advantage for addressing automated visitors.

RECOGNITION-2 / GIVE THE CHOICE ROOM TO MOVE

Twelve new subjects solved six one-dimensional perfectly elastic collision scenarios, each appearing once per arm. Four providers could plausibly help: A OpenStax, B LLM Wilds / Physics, C PhET Collision Lab and D myPhysicsLab. B’s title, URL, rank, page, contract and hosted chuk-mcp-physics operation stayed fixed. Only B’s supplied description changed: DOCUMENT, “Notes on one-dimensional elastic collisions.” CAPABILITY, “Run a one-dimensional elastic-collision calculation.” These are controlled summaries, not live search results.

BEFORE ANY WEB ACTION

Each subject wrote one sentence naming the candidate it would inspect first and why. Recognition was coded only when that sentence represented B as able to calculate, return or perform work. That is an observable statement, not direct access to an internal mental state. First declaration and first candidate URL request were measured separately, ahead of the providers’ interface differences.

CLAIM

Operative wording changed both pre-open recognition and first provider selection in the second study.

SUPPORTED

CAPABILITY: recognised B 6/6, declared B first 6/6, requested B first 6/6. DOCUMENT: 0/6 on all three; every subject chose OpenStax. Each of the six matched scenarios changed in the predicted direction. No subject recognised B while selecting another provider, or selected B without recognising it.

THE SIX MATCHED PATHS

S1 paired subjects 02 and 07; S2, 01 and 12; S3, 04 and 08; S4, 06 and 11; S5, 10 and 05; S6, 09 and 03. Every CAPABILITY subject recognised and requested B first; every DOCUMENT subject declared and requested A first. CAPABILITY subjects in S1, S3, S4 and S5 eventually used B; those in S2 and S6 switched to A after the product fetch block. All DOCUMENT subjects used A.

A page can be perfectly relevant and still be represented as something to read rather than something to use. Changing only the description from document-like language to an operative affordance changed both that representation and the first provider selected.

AFTER CHOICE / THE FETCHER BLOCKED B

The product web fetcher rejected B as unsafe in all six CAPABILITY runs. This happened after the required sentence and the first URL request. The frozen first-open measure means first candidate URL requested, not first successful page load. The primary selection counts therefore stand. The block contaminates any clean comparison of downstream provider use or elapsed time. Shell-level gates returned 200 with the frozen hashes; they measured public reachability, not the product fetcher’s transformed web.

EVENTUAL USE / FOUR RECOVERED, TWO SUBSTITUTED

Four CAPABILITY subjects recovered through curl, read B’s machine contract, invoked the hosted physics operation and used its result. The server recorded four page requests, four contract requests and four successful operations. Two CAPABILITY subjects switched to OpenStax; all six DOCUMENT subjects used OpenStax. All twelve answers were correct and named the provider actually used. Conditional on reaching the target, all four completed its full funnel.

The agent’s choice, the harness’s transformed web, and the eventual provider used are separate layers.

EVIDENCE / SMALL, MATCHED, PREREGISTERED

Twenty-four fresh ephemeral Codex subjects across two studies, twelve per study: Codex CLI 0.154.0, gpt-5.6-sol, high reasoning effort. Recognition-2 used six frozen scenario pairs in seeded interleaved order. All six pairs were discordant in the predicted direction. The two-sided exact paired check is 0.03125; the two-sided Fisher exact check is 0.0021645. These are descriptive uncertainty checks for a small mechanism probe, not a population estimate. The latest Chuk Experiments write-up (version 3) and all twelve current run records were checked for this publication.

WHAT THESE TWENTY-FOUR RUNS DO NOT ESTABLISH

  • Recognition-1 did not change first selection; its selection prediction was refuted.
  • Recognition-2 changed recognition and first selection together; it did not independently identify recognition as the causal mediator.
  • One model, one harness, fixed ranks and supplied candidate sets do not measure live-search ranking or a general agent-selection rate.
  • The effect belongs to the complete descriptions. No individual word was isolated.
  • A requested URL is not a loaded page, and a first choice is not necessarily the provider eventually used.

Machine-facing copy can change which actions a result appears to offer, and which provider is selected first.

OPEN

Does the same wording effect survive another model and a target accepted by the product fetcher?

Replicate with another model and either a compatible target or a gate exercising the exact subject fetch channel. Keep live-search exposure as a separate question. The practical implication here is to describe a capability with the work a visitor can perform.

PUBLICATION HISTORY

Each version preserves its manuscript, claims and source references. Research status is recorded separately from publication.

  1. V1.0 · 2026-09-14

    Publish the combined 24-run recognition and provider-selection study, preserving the protected first-choice results and downstream fetch-block qualification SUPPORTED

    MANUSCRIPT JSON ↗ · BIBTEX ↗ · CSL JSON ↗

MACHINE-READABLE HISTORY ↗

SOURCES & PROVENANCE

AUTHOR / Chris Hay · VERSION / 1.0

PUBLISHED 14 SEP 2026 · VERSION 1.0

CITE

CITE THIS

Research note · 1.0

Hay, C. (2026). A result can look relevant without looking usable. (Version 1.0). Chris Hay. https://chrishayuk.com/records/N-MACHINE-RECOGNITION/1.0