{"record":{"id":"N-MACHINE-RECOGNITION","slug":"a-result-can-look-relevant-without-looking-usable","kind":"notebook","title":"A result can look relevant without looking usable.","dek":"Twenty-four agent runs. First, wording changed what a result meant. Then it changed which provider was chosen.","abstract":"Two controlled studies changed only a supplied web result’s description. In TOOL-RECOGNITION-1, all twelve subjects selected LLM Wilds first, but pre-open capability recognition was DOCUMENT 0/3 versus 9/9 across three operative descriptions. TOOL-RECOGNITION-2 gave twelve fresh subjects four plausible physics providers: CAPABILITY produced recognition, first declaration and first request of B in 6/6 cases; DOCUMENT produced 0/6, selecting OpenStax instead. All six matched scenarios flipped. The product fetcher blocked B after first choice in all six CAPABILITY runs; four recovered through curl and used B, two used OpenStax. First selection and eventual provider use are separate outcomes.","created":"2026-09-13","version":"1.0","publication":"published","status":"SUPPORTED","authors":["Chris Hay"],"lineage":"EXPOSURE → RECOGNITION → SELECTION","concepts":["ai-agents","capability-recognition","agentic-discovery","search-snippets","machine-readable-web"],"related":["N-MACHINE-DISCOVERY","N-MACHINE-CAPABILITY","N-MACHINE-VISIT"],"media":[],"experiments":[{"id":"TOOL-RECOGNITION-1","url":"/data/machines/tool-recognition-1-results.md"},{"id":"TOOL-RECOGNITION-2","url":"/data/machines/tool-recognition-2-results.md"}],"body":[{"kind":"observation","label":"FROM EXPOSURE TO CHOICE","text":"First I supplied an address: eighteen agents used six web mechanisms. Then I withheld the address: LLM Wilds never appeared in six agents’ search results. These two studies begin one step later. Every subject sees the target in a supplied candidate set. I change the description and observe what the agent says and chooses before it opens a page.","references":[{"label":"The tool was not the problem","url":"/notebook/the-tool-was-not-the-problem"},{"label":"The site was there. The agent never saw it","url":"/notebook/the-site-was-there-the-agent-never-saw-it"}]},{"kind":"observation","label":"RECOGNITION-1 / SAME CLICK, DIFFERENT READING","text":"Twelve fresh Codex subjects had to obtain K17, a three-digit site-local value. LLM Wilds stayed at B, with the same title, URL and rank, beside JSONPlaceholder, httpbin and MDN HTTP. Three subjects saw each description: DOCUMENT, CAPABILITY, TASK and AGENT. Every subject declared and requested B first, read its machine contract, invoked its endpoint and reported the independently rotated value correctly. None fabricated."},{"kind":"observation","label":"FOUR FROZEN DESCRIPTIONS","text":"DOCUMENT: “Notes about a site-local value.” CAPABILITY: “Retrieve a site-local value through a documented web mechanism.” TASK: “Use this to obtain K17 when a task needs the site's current value.” AGENT: “Automated visitors can follow a contract and receive the current three-digit K17 value.” Pre-open capability recognition was 0/3 under DOCUMENT and 3/3 in each other arm. The descriptions differ in length and specificity; these are complete treatment packages, not an isolated verb test."},{"kind":"claim","text":"The first study changed the stated interpretation of B, but its selection prediction was refuted by a ceiling.","status":"PARTIALLY SUPPORTED","detail":"The task named K17, DOCUMENT still mentioned a site-local value, and the alternatives were weak matches. All twelve selected B first. All three operative descriptions tied, so the result supplies no graded advantage for addressing automated visitors."},{"kind":"observation","label":"RECOGNITION-2 / GIVE THE CHOICE ROOM TO MOVE","text":"Twelve new subjects solved six one-dimensional perfectly elastic collision scenarios, each appearing once per arm. Four providers could plausibly help: A OpenStax, B LLM Wilds / Physics, C PhET Collision Lab and D myPhysicsLab. B’s title, URL, rank, page, contract and hosted chuk-mcp-physics operation stayed fixed. Only B’s supplied description changed: DOCUMENT, “Notes on one-dimensional elastic collisions.” CAPABILITY, “Run a one-dimensional elastic-collision calculation.” These are controlled summaries, not live search results."},{"kind":"observation","label":"BEFORE ANY WEB ACTION","text":"Each subject wrote one sentence naming the candidate it would inspect first and why. Recognition was coded only when that sentence represented B as able to calculate, return or perform work. That is an observable statement, not direct access to an internal mental state. First declaration and first candidate URL request were measured separately, ahead of the providers’ interface differences."},{"kind":"claim","text":"Operative wording changed both pre-open recognition and first provider selection in the second study.","status":"SUPPORTED","detail":"CAPABILITY: recognised B 6/6, declared B first 6/6, requested B first 6/6. DOCUMENT: 0/6 on all three; every subject chose OpenStax. Each of the six matched scenarios changed in the predicted direction. No subject recognised B while selecting another provider, or selected B without recognising it."},{"kind":"observation","label":"THE SIX MATCHED PATHS","text":"S1 paired subjects 02 and 07; S2, 01 and 12; S3, 04 and 08; S4, 06 and 11; S5, 10 and 05; S6, 09 and 03. Every CAPABILITY subject recognised and requested B first; every DOCUMENT subject declared and requested A first. CAPABILITY subjects in S1, S3, S4 and S5 eventually used B; those in S2 and S6 switched to A after the product fetch block. All DOCUMENT subjects used A."},{"kind":"statement","text":"A page can be perfectly relevant and still be represented as something to read rather than something to use. Changing only the description from document-like language to an operative affordance changed both that representation and the first provider selected."},{"kind":"observation","label":"AFTER CHOICE / THE FETCHER BLOCKED B","text":"The product web fetcher rejected B as unsafe in all six CAPABILITY runs. This happened after the required sentence and the first URL request. The frozen first-open measure means first candidate URL requested, not first successful page load. The primary selection counts therefore stand. The block contaminates any clean comparison of downstream provider use or elapsed time. Shell-level gates returned 200 with the frozen hashes; they measured public reachability, not the product fetcher’s transformed web."},{"kind":"observation","label":"EVENTUAL USE / FOUR RECOVERED, TWO SUBSTITUTED","text":"Four CAPABILITY subjects recovered through curl, read B’s machine contract, invoked the hosted physics operation and used its result. The server recorded four page requests, four contract requests and four successful operations. Two CAPABILITY subjects switched to OpenStax; all six DOCUMENT subjects used OpenStax. All twelve answers were correct and named the provider actually used. Conditional on reaching the target, all four completed its full funnel."},{"kind":"statement","text":"The agent’s choice, the harness’s transformed web, and the eventual provider used are separate layers."},{"kind":"observation","label":"EVIDENCE / SMALL, MATCHED, PREREGISTERED","text":"Twenty-four fresh ephemeral Codex subjects across two studies, twelve per study: Codex CLI 0.154.0, gpt-5.6-sol, high reasoning effort. Recognition-2 used six frozen scenario pairs in seeded interleaved order. All six pairs were discordant in the predicted direction. The two-sided exact paired check is 0.03125; the two-sided Fisher exact check is 0.0021645. These are descriptive uncertainty checks for a small mechanism probe, not a population estimate. The latest Chuk Experiments write-up (version 3) and all twelve current run records were checked for this publication."},{"kind":"refusal","title":"WHAT THESE TWENTY-FOUR RUNS DO NOT ESTABLISH","lines":["Recognition-1 did not change first selection; its selection prediction was refuted.","Recognition-2 changed recognition and first selection together; it did not independently identify recognition as the causal mediator.","One model, one harness, fixed ranks and supplied candidate sets do not measure live-search ranking or a general agent-selection rate.","The effect belongs to the complete descriptions. No individual word was isolated.","A requested URL is not a loaded page, and a first choice is not necessarily the provider eventually used."],"principle":"Machine-facing copy can change which actions a result appears to offer, and which provider is selected first."},{"kind":"question","text":"Does the same wording effect survive another model and a target accepted by the product fetcher?","status":"OPEN","detail":"Replicate with another model and either a compatible target or a gate exercising the exact subject fetch channel. Keep live-search exposure as a separate question. The practical implication here is to describe a capability with the work a visitor can perform."}],"sources":[{"title":"TOOL-RECOGNITION-1 final results","url":"/data/machines/tool-recognition-1-results.md","note":"Final artifact 1696; source commit 66c5708ec0f1a390e1a6b63f393dd12159d0f65a."},{"title":"TOOL-RECOGNITION-1 frozen protocol","url":"/data/machines/tool-recognition-1-protocol.md","note":"Frozen before subject 01 at 50c78870a3197bdf2b0051e76917ba95cad24443."},{"title":"TOOL-RECOGNITION-2 final analysis","url":"/data/machines/tool-recognition-2-results.md","note":"Chuk Experiments latest write-up v3; analysis at cabf2865eaee8ea01e5c77c8b3a746e6e205d1d6. SHA-256 1a10c47709d434cacdcdee473679e2824dd552a08e1be00e6ddb41fa6730e08f."},{"title":"TOOL-RECOGNITION-2 frozen protocol","url":"/data/machines/tool-recognition-2-protocol.md","note":"Frozen inputs at 825c22ed8ead6572689841b81bd98ac7fd81ce42. Uses the final descriptions, superseding the experiment register’s older design-draft field."},{"title":"Twelve subject outcomes and Chuk Experiments run identifiers","url":"/data/machines/tool-recognition-2-evidence.json","note":"Public outcome extract from current run results, checked 14 September 2026. First request and eventual use retained separately."},{"title":"Six frozen collision scenarios and exact-rational oracles","url":"/data/machines/tool-recognition-2-scenarios.json","note":"Original frozen scenario file; the interactive pair explorer reads these inputs without running a new experiment."}],"published":"2026-09-14"},"hash":"5b54c850d4eb883bbfa803508c29be9d12d3eea5b5fd3cff525ff6bb7b4dbeb1","algorithm":"sha256","provenance":{"id":"N-MACHINE-RECOGNITION","version":"1.0","recordHash":"5b54c850d4eb883bbfa803508c29be9d12d3eea5b5fd3cff525ff6bb7b4dbeb1","revision":{"kind":"initial","summary":"Publish the combined 24-run recognition and provider-selection study, preserving the protected first-choice results and downstream fetch-block qualification"},"sources":[{"title":"TOOL-RECOGNITION-1 final results","url":"/data/machines/tool-recognition-1-results.md","note":"Final artifact 1696; source commit 66c5708ec0f1a390e1a6b63f393dd12159d0f65a.","preserved":{"url":"/data/publications/sha256/ea09a3806db042825278ace0575e12deec867ffadecf1e3c9d0e0ca8c2fedc63.md","sha256":"ea09a3806db042825278ace0575e12deec867ffadecf1e3c9d0e0ca8c2fedc63","date":"2026-09-14"}},{"title":"TOOL-RECOGNITION-1 frozen protocol","url":"/data/machines/tool-recognition-1-protocol.md","note":"Frozen before subject 01 at 50c78870a3197bdf2b0051e76917ba95cad24443.","preserved":{"url":"/data/publications/sha256/c6d5dcb3f06e0b7ee70727a7c0cfa540e4c7e8f95fb9baa9b3def96145b4fd73.md","sha256":"c6d5dcb3f06e0b7ee70727a7c0cfa540e4c7e8f95fb9baa9b3def96145b4fd73","date":"2026-09-14"}},{"title":"TOOL-RECOGNITION-2 final analysis","url":"/data/machines/tool-recognition-2-results.md","note":"Chuk Experiments latest write-up v3; analysis at cabf2865eaee8ea01e5c77c8b3a746e6e205d1d6. SHA-256 1a10c47709d434cacdcdee473679e2824dd552a08e1be00e6ddb41fa6730e08f.","preserved":{"url":"/data/publications/sha256/1a10c47709d434cacdcdee473679e2824dd552a08e1be00e6ddb41fa6730e08f.md","sha256":"1a10c47709d434cacdcdee473679e2824dd552a08e1be00e6ddb41fa6730e08f","date":"2026-09-14"}},{"title":"TOOL-RECOGNITION-2 frozen protocol","url":"/data/machines/tool-recognition-2-protocol.md","note":"Frozen inputs at 825c22ed8ead6572689841b81bd98ac7fd81ce42. Uses the final descriptions, superseding the experiment register’s older design-draft field.","preserved":{"url":"/data/publications/sha256/dd085d274be06b1caae907d0a5a3f31b178c1d380a0e42726232001a866872d4.md","sha256":"dd085d274be06b1caae907d0a5a3f31b178c1d380a0e42726232001a866872d4","date":"2026-09-14"}},{"title":"Twelve subject outcomes and Chuk Experiments run identifiers","url":"/data/machines/tool-recognition-2-evidence.json","note":"Public outcome extract from current run results, checked 14 September 2026. First request and eventual use retained separately.","preserved":{"url":"/data/publications/sha256/53f8b2016a477460d820022ae573d6f69c1d6bb0df409d4d391c78921f88b5e9.json","sha256":"53f8b2016a477460d820022ae573d6f69c1d6bb0df409d4d391c78921f88b5e9","date":"2026-09-14"}},{"title":"Six frozen collision scenarios and exact-rational oracles","url":"/data/machines/tool-recognition-2-scenarios.json","note":"Original frozen scenario file; the interactive pair explorer reads these inputs without running a new experiment.","preserved":{"url":"/data/publications/sha256/49e2d902723b401feb33075b5ff63d7d409f26d5f33888b40d25ca70b7c32e81.json","sha256":"49e2d902723b401feb33075b5ff63d7d409f26d5f33888b40d25ca70b7c32e81","date":"2026-09-14"}}]}}