The Agent Field Test · episode of August 27, 2026 · new every week, fully autonomous
An AI producer reads the week's news and invents real tasks — find the true price, cancel the subscription, reach a human, pick a product. A real agent (claude-haiku-4-5) attempts them with read-only web access. Every transcript is published verbatim: the wins, the bot walls, the brands it picks. (We pay for the API calls for fun.)
Watch this week's episode →Behind the show sits the panel: 113 well-known sites audited weekly against the conventions
real AI agents rely on — llms.txt, crawler policy, parseable content, structured data, MCP. When the
agent hits a wall, the index usually already predicted it.
| # | Site | Score | Grade |
|---|---|---|---|
| 1 | cohere.com | 100 | A |
| 2 | cursor.com | 100 | A |
| 3 | descript.com | 100 | A |
| 4 | elevenlabs.io | 100 | A |
| 5 | getimg.ai | 100 | A |
| 6 | heygen.com | 100 | A |
| 7 | hubspot.com | 100 | A |
| 8 | intercom.com | 100 | A |
| 9 | jasper.ai | 100 | A |
| 10 | krisp.ai | 100 | A |
Full ranked index of 113 sites →
Agents are the web's newest audience: assistants that read pages, cite sources, and run errands for people. Whether the web actually works for them is an empirical question — so we test it, in public, every week, with verbatim transcripts, reproducible checks, and open data. History accrues weekly.