Ora Measured 99% of the Web — and Found Your Site Is the Real Bottleneck
胡新宇
Published on 2026-08-24
Ora's benchmark found 99% of websites fail AI agent signup. Vercel's eve framework beat Claude Code in head-to-head traces. The bottleneck has moved from model capability to the target environment — and that changes what AI teams should ship next.
Ora Measured 99% of the Web — and Found Your Site Is the Real Bottleneck
On August 21, 2026, Vercel published a case study about a benchmarking company called Ora. Ora sends agents onto live customer websites with a simple instruction: sign up for a product, integrate with it, and pay for it. They trace every step, count cost and latency, and run the same gauntlet against Claude Code, ChatGPT, Gemini, Hermes, OpenClaw, and Vercel's own eve framework.
The line that should keep product teams up at night is at the bottom of the post: "By Ora's own measure, 99% of the web still can't handle an agent that shows up to sign up, integrate, and pay."
The easy read is that the agents aren't smart enough. The actual finding is more uncomfortable: the agents are fine. The world they land in isn't.

The harness is half the agent — and most of that half is already absorbed
For two years the AI engineering playbook has been: build a better harness. That was the right bet. Harness-Bench ran the same model over the same 106 tasks across different harnesses and got scores ranging from 52.4 to 76.2 — a 23.8-point spread with zero change to the model (Latent Space, "The Evolution of the Agent Harness", 2026-08-22). OpenAI did a similar thing on ARC-AGI-3: adding only retained reasoning and compaction, GPT-5.6 Sol's score tripled from 13.3% to 38.3%. Half the agent is the harness — that was true through 2024 and into early 2025.
Then something flipped. As Latent Space puts it, the model and the harness curves stopped running side by side and started braiding: models were trained inside the harness environment, started absorbing harness capabilities into their weights, and the harness shed the scaffold. Thariq Shihipar from Anthropic recently said the team deleted 80% of Claude Code's system prompt — capability unchanged, code cut. GPT-5.1-Codex-Max launched as "the first model natively trained to operate across multiple context windows through compaction." The loop is now