AI Agents · Benchmark
The Bug That Needed
Two Subscribers to Exist
When every candidate scores 100% on the first try, the test has said nothing about where any of them would fail. That is a ceiling effect: a test with no room to fail teaches nothing about where the edges are. Raising the difficulty one design decision at a time is what finds them. Four levels of deliberately harder tasks later, the last one found a bug none of the five implementations had ever been in a position to have: two Bounded Contexts subscribing to the same event, for the first time in the code.
The experiment gives AI coding agents the same task and scores the result with an architecture checker, a script that statically checks whether code follows the documented architecture rules. It runs on my example project, which implements the same backend design (DDD, CQRS, Outbox) in five languages side by side.
The setup was simple. The same synthetic domain, Voucher (issue to ACTIVE, redeem as a plain transition with no event, expire as an event since other parts of the system react to it), was built independently across all five languages at once, each agent given nothing but its own implementations/<lang>/CLAUDE.md as an entry point. No doc paths, no hints about the skeleton-code generator. All five hit a perfect checker score, and all five independently converged on the identical judgment: publish the event on expire() only, matching the same "does anything react?" pattern the docs already establish elsewhere. Strong evidence the docs communicate consistently across languages. Along the way, the run also surfaced three tooling regressions nobody had noticed: two of the scripts that generate a skeleton domain from a name still emitting a shape a naming rule added the same day now forbade, and the checker's versions in two languages drifting out of parity on which directories their file walkers were supposed to skip.
A Perfect Score Everywhere Is Not Reassuring
Five languages, one easy task, five first-try wins. On its own that result explains nothing about where any implementation would fail. A test that always passes has no discriminative power, and Voucher was, deliberately, an easy first task. The question was what to build next, and running the same easy shape again wasn't going to answer it. What the task needed wasn't more repeats. It needed to get harder, on purpose, in directions specifically chosen to exercise code paths nothing before this had ever exercised.
A Ladder, Not a Repeat
Level 2, Booking/Cancellation (two Aggregates inside one Bounded Context plus a Domain Service, mirroring Payment/Refund's existing RefundEligibilityService), was the first sign the ladder discriminated. All five still reached 100%, and all five independently made the same subtle judgment call the spec allowed room to get wrong (a rejected booking is never persisted, unlike Refund's persisted REJECTED state). But NestJS scored 96/100 on its first pass, a genuine defect this time, a raw string thrown where the convention requires a typed enum, then corrected it on its own.
Level 3, Membership, which needs a synchronous Adapter reading another BC's Account status, again converged on the right pattern in all five: the synchronous read, not the asynchronous Integration Event a level-4 task would have needed instead. Every language correctly translated Account's status enum into a plain boolean rather than leaking the enum itself across the boundary. Java's agent went a step further on its own, avoiding a Spring bean-name collision with Card's existing AccountAdapterImpl by noticing and reading an existing code comment that named the conflict before writing anything.
Level 4, Built to Contrast With Level 3
StandingOrder was designed specifically as level 3's mirror image. Create one against an Account, and it becomes ACTIVE; if that Account is later suspended, the StandingOrder must become PAUSED automatically, and CANCELLED if the Account is closed. The reaction has to happen the moment the Account's status changes, never through a direct call on StandingOrder itself. The correct pattern this time is the opposite of level 3's: subscribing to an asynchronous Integration Event, not a synchronous lookup. All five made that distinction, and this time verification was strengthened to match the stakes — each agent had to prove it with an end-to-end test that calls the suspend/close API and polls until the reaction completes, not a unit test asserting the handler function alone.
All five passed, independent re-verification matching every self-report. The interesting part wasn't the score.
Card was already subscribing to the same two Account events StandingOrder now needed. The moment a second subscriber existed for an eventType, it exposed that the root domain-events.md's stated principle (one event, multiple handler subscribers, 1:N) had never been load-bearing code in two of the five languages. It had been true in the docs since before this experiment existed, and false in the code the entire time, because nothing had ever tried it.
Java-springboot's handler map was built with Collectors.toMap(eventType, identity()), a shape where registering a second handler bean for an eventType already in use throws IllegalStateException: Duplicate key at boot: not at runtime under load, but the instant the application tries to start. FastAPI's build_event_handlers() returned dict[str, EventHandlerFn], one callable per key. No crash at all, just the second registration overwriting the first without a word, so only the newer subscriber would ever run. Go and NestJS had never had the problem: Go's main.go hand-assembles a plain map where adding a second call under the same key is unremarkable, and NestJS's registry was already list-shaped from the start. Each language's agent fixed its own case without coordinating with the others. Java moved to Collectors.groupingBy, producing a proper Map<String, List<OutboxEventHandler>>; FastAPI moved to dict[str, list[EventHandlerFn]] and updated its consumer, its scaffolding generator, and the doc all together.
Kotlin's fix was the odd one out: not wrong, but a different shape of workaround. Rather than restructuring its registry to be list-valued, it added the second handler call directly inside the existing per-eventType lambda:
"AccountSuspendedEvent" to { eventId, payload ->
accountSuspendedEventHandler.handle(objectMapper.readValue(payload, AccountSuspendedEvent::class.java), eventId)
standingOrderPauseHandler.handle(objectMapper.readValue(payload, AccountSuspendedEvent::class.java), eventId)
}Functionally correct for two subscribers, hardcoded rather than structural. A third subscriber to the same event will need a hand edit to this lambda rather than a new registration, unlike Java and FastAPI's now-generalized shape. Flagged, not fixed; the checker still passes, because nothing in it requires the more scalable form.
What the Task Tested
The lesson is the same one a lint rule that has only ever seen two inputs teaches, taken one step further: some bugs only exist once a specific combination of circumstances shows up in the code, and no amount of reading, no amount of repeating an easy task, and no static rule can produce that combination on its own. Only running the scenario can. Level 1's perfect scores measured whether five languages agree on an easy judgment call. Level 4 measured something a perfect score can hide entirely: whether a codebase survives the first time a real-world shape of usage (two things caring about the same event) happens to it. Three of five languages hadn't, without anyone knowing, until a task was deliberately built to make it happen.
docs/benchmark.md (the full run, every level, every self-report vs. independent re-verification table, in backend-service-playbook, my example project that implements the same backend design in five languages side by side) · OutboxEventDispatcher.java (the fix, list-valued handler map) · Can an AI Agent Follow Your Architecture? (how the tasks are designed and how the scores are re-checked by hand)