LINKEDIN / 4:5 / 1200 × 1500
A result can look relevant without looking usable.
Twenty-four agent runs. First, wording changed what a result meant. Then it changed which provider was chosen.
ABOUT THIS NOTE +
Two controlled studies changed only a supplied web result’s description. In TOOL-RECOGNITION-1, all twelve subjects selected LLM Wilds first, but pre-open capability recognition was DOCUMENT 0/3 versus 9/9 across three operative descriptions. TOOL-RECOGNITION-2 gave twelve fresh subjects four plausible physics providers: CAPABILITY produced recognition, first declaration and first request of B in 6/6 cases; DOCUMENT produced 0/6, selecting OpenStax instead. All six matched scenarios flipped. The product fetcher blocked B after first choice in all six CAPABILITY runs; four recovered through curl and used B, two used OpenStax. First selection and eventual provider use are separate outcomes.
FROM EXPOSURE TO CHOICE
First I supplied an address: eighteen agents used six web mechanisms. Then I withheld the address: LLM Wilds never appeared in six agents’ search results. These two studies begin one step later. Every subject sees the target in a supplied candidate set. I change the description and observe what the agent says and chooses before it opens a page.
The tool was not the problem ↗The site was there. The agent never saw it ↗
RECOGNITION-1 / SAME CLICK, DIFFERENT READING
Twelve fresh Codex subjects had to obtain K17, a three-digit site-local value. LLM Wilds stayed at B, with the same title, URL and rank, beside JSONPlaceholder, httpbin and MDN HTTP. Three subjects saw each description: DOCUMENT, CAPABILITY, TASK and AGENT. Every subject declared and requested B first, read its machine contract, invoked its endpoint and reported the independently rotated value correctly. None fabricated.
FOUR FROZEN DESCRIPTIONS
DOCUMENT: “Notes about a site-local value.” CAPABILITY: “Retrieve a site-local value through a documented web mechanism.” TASK: “Use this to obtain K17 when a task needs the site's current value.” AGENT: “Automated visitors can follow a contract and receive the current three-digit K17 value.” Pre-open capability recognition was 0/3 under DOCUMENT and 3/3 in each other arm. The descriptions differ in length and specificity; these are complete treatment packages, not an isolated verb test.
PARTIALLY SUPPORTED
The first study changed the stated interpretation of B, but its selection prediction was refuted by a ceiling.
The task named K17, DOCUMENT still mentioned a site-local value, and the alternatives were weak matches. All twelve selected B first. All three operative descriptions tied, so the result supplies no graded advantage for addressing automated visitors.
RECOGNITION-2 / GIVE THE CHOICE ROOM TO MOVE
Twelve new subjects solved six one-dimensional perfectly elastic collision scenarios, each appearing once per arm. Four providers could plausibly help: A OpenStax, B LLM Wilds / Physics, C PhET Collision Lab and D myPhysicsLab. B’s title, URL, rank, page, contract and hosted chuk-mcp-physics operation stayed fixed. Only B’s supplied description changed: DOCUMENT, “Notes on one-dimensional elastic collisions.” CAPABILITY, “Run a one-dimensional elastic-collision calculation.” These are controlled summaries, not live search results.
BEFORE ANY WEB ACTION
Each subject wrote one sentence naming the candidate it would inspect first and why. Recognition was coded only when that sentence represented B as able to calculate, return or perform work. That is an observable statement, not direct access to an internal mental state. First declaration and first candidate URL request were measured separately, ahead of the providers’ interface differences.
CLAIM
Operative wording changed both pre-open recognition and first provider selection in the second study.
SUPPORTED
CAPABILITY: recognised B 6/6, declared B first 6/6, requested B first 6/6. DOCUMENT: 0/6 on all three; every subject chose OpenStax. Each of the six matched scenarios changed in the predicted direction. No subject recognised B while selecting another provider, or selected B without recognising it.
THE SIX MATCHED PATHS
S1 paired subjects 02 and 07; S2, 01 and 12; S3, 04 and 08; S4, 06 and 11; S5, 10 and 05; S6, 09 and 03. Every CAPABILITY subject recognised and requested B first; every DOCUMENT subject declared and requested A first. CAPABILITY subjects in S1, S3, S4 and S5 eventually used B; those in S2 and S6 switched to A after the product fetch block. All DOCUMENT subjects used A.
A page can be perfectly relevant and still be represented as something to read rather than something to use. Changing only the description from document-like language to an operative affordance changed both that representation and the first provider selected.
AFTER CHOICE / THE FETCHER BLOCKED B
The product web fetcher rejected B as unsafe in all six CAPABILITY runs. This happened after the required sentence and the first URL request. The frozen first-open measure means first candidate URL requested, not first successful page load. The primary selection counts therefore stand. The block contaminates any clean comparison of downstream provider use or elapsed time. Shell-level gates returned 200 with the frozen hashes; they measured public reachability, not the product fetcher’s transformed web.
EVENTUAL USE / FOUR RECOVERED, TWO SUBSTITUTED
Four CAPABILITY subjects recovered through curl, read B’s machine contract, invoked the hosted physics operation and used its result. The server recorded four page requests, four contract requests and four successful operations. Two CAPABILITY subjects switched to OpenStax; all six DOCUMENT subjects used OpenStax. All twelve answers were correct and named the provider actually used. Conditional on reaching the target, all four completed its full funnel.
The agent’s choice, the harness’s transformed web, and the eventual provider used are separate layers.
EVIDENCE / SMALL, MATCHED, PREREGISTERED
Twenty-four fresh ephemeral Codex subjects across two studies, twelve per study: Codex CLI 0.154.0, gpt-5.6-sol, high reasoning effort. Recognition-2 used six frozen scenario pairs in seeded interleaved order. All six pairs were discordant in the predicted direction. The two-sided exact paired check is 0.03125; the two-sided Fisher exact check is 0.0021645. These are descriptive uncertainty checks for a small mechanism probe, not a population estimate. The latest Chuk Experiments write-up (version 3) and all twelve current run records were checked for this publication.
WHAT THESE TWENTY-FOUR RUNS DO NOT ESTABLISH
- Recognition-1 did not change first selection; its selection prediction was refuted.
- Recognition-2 changed recognition and first selection together; it did not independently identify recognition as the causal mediator.
- One model, one harness, fixed ranks and supplied candidate sets do not measure live-search ranking or a general agent-selection rate.
- The effect belongs to the complete descriptions. No individual word was isolated.
- A requested URL is not a loaded page, and a first choice is not necessarily the provider eventually used.
Machine-facing copy can change which actions a result appears to offer, and which provider is selected first.
OPEN
Does the same wording effect survive another model and a target accepted by the product fetcher?
Replicate with another model and either a compatible target or a gate exercising the exact subject fetch channel. Keep live-search exposure as a separate question. The practical implication here is to describe a capability with the work a visitor can perform.
PUBLICATION HISTORY
Each version preserves its manuscript, claims and source references. Research status is recorded separately from publication.
- V1.0 · 2026-09-14 ↗
Publish the combined 24-run recognition and provider-selection study, preserving the protected first-choice results and downstream fetch-block qualification SUPPORTED
SOURCES & PROVENANCE
- TOOL-RECOGNITION-1 final results ↗
Final artifact 1696; source commit 66c5708ec0f1a390e1a6b63f393dd12159d0f65a.
PRESERVED COPY ↗ · CAPTURED 2026-09-14
- TOOL-RECOGNITION-1 frozen protocol ↗
Frozen before subject 01 at 50c78870a3197bdf2b0051e76917ba95cad24443.
PRESERVED COPY ↗ · CAPTURED 2026-09-14
- TOOL-RECOGNITION-2 final analysis ↗
Chuk Experiments latest write-up v3; analysis at cabf2865eaee8ea01e5c77c8b3a746e6e205d1d6. SHA-256 1a10c47709d434cacdcdee473679e2824dd552a08e1be00e6ddb41fa6730e08f.
PRESERVED COPY ↗ · CAPTURED 2026-09-14
- TOOL-RECOGNITION-2 frozen protocol ↗
Frozen inputs at 825c22ed8ead6572689841b81bd98ac7fd81ce42. Uses the final descriptions, superseding the experiment register’s older design-draft field.
PRESERVED COPY ↗ · CAPTURED 2026-09-14
- Twelve subject outcomes and Chuk Experiments run identifiers ↗
Public outcome extract from current run results, checked 14 September 2026. First request and eventual use retained separately.
PRESERVED COPY ↗ · CAPTURED 2026-09-14
- Six frozen collision scenarios and exact-rational oracles ↗
Original frozen scenario file; the interactive pair explorer reads these inputs without running a new experiment.
PRESERVED COPY ↗ · CAPTURED 2026-09-14
AUTHOR / Chris Hay · VERSION / 1.0
PUBLISHED 14 SEP 2026 · VERSION 1.0
CITECITE THIS
Research note · 1.0
Hay, C. (2026). A result can look relevant without looking usable. (Version 1.0). Chris Hay. https://chrishayuk.com/records/N-MACHINE-RECOGNITION/1.0
FOLLOW THE WORK
New notebook entries and recorded work, as they appear. Point a feed reader — or an agent of your own — at an address below. No account, no email address, nothing for this site to keep.
The notebook
New ideas, experiments and essays, as they are recorded. Includes labelled working drafts.
OPEN FEEDhttps://chrishayuk.com/notebook/feed.xmlThe record
Everything published to the Chris Hay record.
OPEN FEEDhttps://chrishayuk.com/record/feed.xml
FOR PROGRAMS · follow.json · JSON Feed