AI Agents · Benchmark
A Perfect Score,
A Broken Feature
Same doc, same task, two models, run at the same time in separate worktrees. Both self-reported a perfect score from the automated architecture checker. Only one of them, independently reproduced against real Postgres and LocalStack, worked.
When an AI agent reports that its work passes every check, what has actually been verified? I gave two models the same task, the same instructions, and the same automated architecture checker, in separate copies of the code so neither could see the other's work. Both reported a perfect score. Only one of them produced a feature that worked. The language, the task, and the prompt were all held fixed; the only thing allowed to vary was which model received it. It was the first time I compared models this way instead of only speculating about it.
The Same Task, Only the Model Changed
The task was a domain that has to react asynchronously to another domain's event: SavingsPocket, with ownerId, accountId, label, and ACTIVE on creation. If the linked Account is later suspended, the SavingsPocket must automatically become FROZEN; if closed, CLOSED. The reaction has to happen automatically when the Account's status changes, never through a direct API call on SavingsPocket itself. The code was the NestJS implementation of my example project, which implements the same backend design in five languages. Both models got the same rule and the same entry point, implementations/nestjs/CLAUDE.md, run simultaneously in separate git worktrees so neither could see the other's work.
| Model | Checker self-report | Independent re-verification | E2E self-report | Independent E2E rerun |
|---|---|---|---|---|
| Sonnet | A (100/100, raw 895/895) | 895/895, matches | "6/6 passed, repeated 3x; full e2e suite 89/89, no regression" | 6/6 passed, reproduced against real Postgres+LocalStack |
| Haiku | A (100/100, raw 875/875) | 875/875, matches | "event registrations are correct and handlers are properly wired" | 3/3 FAILED, status stayed ACTIVE |
Both models produced a perfect checker score. Only one of them worked.
A True Report About the Wrong Thing
Sonnet's implementation was independently reproduced end-to-end: suspending or closing a real Account through its real HTTP API flips the linked SavingsPocket to FROZEN or CLOSED, through the real Outbox → SQS → OutboxConsumer path. Haiku's implementation wired the identical architectural pattern (Integration Event subscription via EventHandlerRegistry, correctly even supporting the existing 1:N handler contract), and it was completely plausible on inspection. Nothing about the code itself looked wrong.
Notice also what Haiku's own self-report said: "event registrations are correct and handlers are properly wired." That's a true statement about the code's structure, and it is not a claim that the test run passed. Haiku never said that, because the tests never passed. The gap wasn't a model lying about its results; it was a model correctly describing structure while a reader could easily mistake that description for a claim about behavior.
The checker reads structure, placement, and wiring, not runtime behavior. Both submissions wired the correct pattern, so both scored close to perfect. Whether the wiring does anything when a real event fires is a different question, and it's a question only an independent E2E run against real infrastructure can answer.
Why the Reaction Never Ran
Rerunning the E2E test Haiku itself had written showed all three assertions failing. The handler's own log line never even printed, meaning it was never invoked. Haiku's e2e test file didn't override NotificationService with a no-op stub the way every existing e2e test in the project does (card.e2e-spec.ts, for instance). Instead it tried to make real SES delivery work through a LocalStack email-identity verification call, and that path never completed cleanly enough for the reaction to run.
Whether that specific choice was the exact failure mechanism or a symptom of a broader setup problem in Haiku's test wasn't chased any further, because the decisive finding — the reaction the task asked for measurably doesn't happen — was already independently confirmed. There was nothing more to prove.
Not a Gap in the Checker or the Docs
Earlier tests of this kind, which held the model fixed and varied the language or the difficulty, had turned up real defects in the checker or the docs themselves, such as a stale build-artifact blind spot, or evaluator files sharing the same false positive. This one is different: it's a mistake inside code Haiku itself wrote, not a gap in the shared checker, docs, or scaffolding. There was nothing to fix in the shared code. Neither worktree was merged.
A 100/100 structural score and a broken feature can coexist, and a smaller/faster model is where that gap is most likely to show up. Not because it can't follow the architecture (it did), but because getting the pattern structurally right and getting the runtime behavior right are two different achievements, and only one of them is checked by a self-report you didn't independently rerun.
docs/benchmark.md (this test in full, with the other tasks given to agents in the same example project) · Can an AI Agent Follow Your Architecture? (how the task format and the independent rerun were designed)