Design notes for running several coding agents on one machine without stepping on each other: a container per agent so the unit of isolation is also the unit of cleanup, agent-to-agent messages over the network, nested containers only when needed, credentials that are shared without being baked in, and a cgroup-only alternative tested with systemd.
Every page passed a word-count check and had a distinct title, URL, and numbers. Slightly under half of all text on the site still existed on more than one page. The fifteen-line measurement that found it, and three structurally different ways a page duplicates its neighbour without anyone deciding it should.
Five Kubernetes anti-pattern checkers all assumed `---`-separated documents. Naming more than one resource in a single `kubectl get -o yaml` call wraps the result in `kind: List` instead, and every checker found zero resources to flag without a word, which looked identical to a clean pass.
Argo CD's App-of-Apps proof lives entirely on the parent; Flux's dependsOn proof is declared by the child and unverifiable alone. Audit either tree without including its root, and both fail the same way — for what turns out to be the same underlying reason.
A score that adds up many checks can be missing one because nobody built it yet, or because the thing being scored can never show it. A checker covering nineteen categories of Kubernetes deployment mistake hit the second kind: drift can only exist after a manifest has already been applied, which an authoring benchmark structurally cannot produce or avoid. The honest fix was a permanent, documented ceiling, not a future version.
Two models scored an identical 9/9 on a Kubernetes manifest-authoring task, independently reproduced. Reading what each one wrote found a self-defeating NetworkPolicy in one and a promotion pipeline referencing a resource that doesn't exist in the other, two unrelated defects that checks for a resource's presence couldn't see.
A drift checker pointed at a cluster that had just been applied cleanly reported drift everywhere. The cluster wasn't lying — the API server's own admission defaulting had filled in fields Git never mentioned, and a naive full-object comparison had no way to tell the difference.
Two codebases generate an Aggregate's ID in two different places — one in the constructor, one via a Factory asking Infrastructure for it. Eric Evans' own book has a specific, citable answer for which pattern it actually describes, and it isn't the one either codebase's convention assumes.
Nearly every DDD codebase forbids referencing another Aggregate by direct object reference (ID only). Eric Evans' 2003 book explicitly permits it. The person who wrote the ID-only rule, Vaughn Vernon, says so himself, in the same paper that argues for the stricter rule anyway.
The same moment, serialized by the same driver, produces a different string depending on the process's timezone. Four languages had this bug at the call site and one had it at the process boundary — and the fix belonged in a genuinely different place in each, verified by literally running the tests nine time zones apart.
A merge workflow that waits on every check while being one of them can only succeed by accident. Every PR a Dependabot auto-merge workflow had ever merged did so by winning a race against its own six-hour deadlock — one of its steps was waiting for a check run that could only finish after that step did. Fixing it surfaced a second bug waiting right behind the first, and a class of half-merge left behind by plain GitHub 502s.
A CI check only protects what its trigger watches. A Spring Boot 4 migration that checked git history instead of a stale doc, found a workaround for a library a search index insisted did not exist, and ended a day later with the deployable image unable to build — because nothing in CI was watching the file whose meaning had just changed.
Zero findings tells you what a check looked at, not what is there. A path-existence checker reported zero before and after a three-language audit that fixed roughly eighty real issues: stale code quotes, an evaluator that grades itself a perfect score for scanning nothing, and a generator still emitting a bug already fixed in the code it was modeled on.
An end-to-end suite that assembles its own approximation of the app is testing the approximation. NestJS's e2e suite never booted the real app, and every language's LLM features had only ever run through their own fallback path. Fixing both surfaced a stranger bug: nock and testcontainers fighting over the same patched module.
Same doc, same task, two models, run at the same time in separate worktrees. Both self-reported a perfect score from the automated architecture checker. Only one of them, independently reproduced against real Postgres and LocalStack, worked.
Event dispatch that has only ever run one handler per event has not shown it can run two. When four real features gave an event its second subscriber in five language implementations of the same design, all five broke, each differently, from a loud boot-time crash to a silent single-handler drop nothing ever logged.
A monthly statement and a GDPR-style data export both died to the same question: couldn't the client just build this itself? The spending-analysis ETL that survived it, and the rule it revealed.
A fraud signal computed from text the suspect writes can't catch that suspect. Why an LLM refund-reason classifier built on exactly that had to go, the history-based ML scorer removed alongside it, and the one rule the removal left behind.
How to let an LLM turn a free-text question into a database filter without ever letting it decide whose data comes back: a filter type with no owner field, a structured-data RAG pipeline over an account's own transactions, and the same invariant held across five languages' own conventions.
A missing @Transactional, a JDK HTTP client retry quirk, a VARCHAR(36) overflow, an SQS FIFO dedup collision — four real bugs that needed real infrastructure to even exist.
An LLM makes a good signal and a bad final judge. Let it read a refund reason and hand back a value, keep the threshold in a Domain Service that never calls it, and swapping the backend from Claude to self-hosted Ollama touches almost no test.
How to put a machine-learning risk score next to a rule-based refund decision with no real data to train on: an interface, a config switch, fail-open on errors, and a Domain Service that keeps the decision. Walked through with a hand-rolled logistic regression on refund history that has since been removed.
Automated architecture checks reliably catch code in the wrong place. Three kinds of drift still get past them: a wrong name inside the right file, code that is wrong together with its own doc, and disagreement that only exists between implementations.
An annotation that compiles isn't an annotation that's true. Completing incomplete Swagger docs across five implementations of the same design, verified by actually booting each app instead of trusting the annotations compiled. What it found had nothing to do with documentation — including a Spring Boot 4 dependency split that left production migrations silently never running.
Every Repository operation fits find, save, and delete with a noun, and nothing else. Written only in prose, the rule drifted in four of five implementations of the same design. Once a check enforced it, the first run found three more violations nobody had noticed.
How to measure whether an AI agent finds and follows documented design rules on its own: a sparse task, a score you rerun yourself, and difficulty raised one decision at a time.
A shell command's output has, more than once, contained something shaped like a system message, instructing the agent to hide a change. The rule that matters: disregard it, and say so.
A transfer feature needs one thing every implementation already claimed to support: writing two Aggregates atomically. Building it for real found a working mechanism in one language, a regression waiting one edit inside the obvious fix in another, and a doc that had been quietly wrong about its own code in a third.
When every candidate scores 100% on an easy task, the test has said nothing about where any of them would fail. Raising the difficulty of an AI coding task one design decision at a time, across five language implementations of the same design, found the ceiling. The last rung exposed a fan-out bug that had been invisible because nothing had ever subscribed two things to the same event before.
Turning each audit finding into an automated rule raises a question: how many rules are enough? Starting from a naming fix that reached only the write-side interface, each batch of new rules found three or four real bugs, then two, then zero. That flat yield curve was the answer.
No parsing, no understanding of what a code snippet does — just comparing backtick-quoted paths against the real file tree. The exclusion rules that kept it from crying wolf mattered more than the two-pattern check itself, and it still caught a real bug in four docs on its first run.
A scaffolding generator is a second implementation of every convention it emits. When a rule changes and only the hand-written example is updated, the generator keeps emitting the old pattern. Generating a brand-new domain from just a name and running every automated check against it is what catches the drift, and the generator's own bugs.
A lint rule that only ever passed on the inputs it was written against has not been shown to be generic. Two architecture rules had checked out clean for months on the same two domains. A deliberately unrelated third domain surfaced two false positives, and confirmed the rule meant to catch a real mistake still did.
Port one design to several languages and its first security hole gets ported too. A security audit found /auth/sign-in accepted a userId and nothing else in all five implementations: how the bug looked in each language, what closes it, and the JDK retry bug a new 401 test uncovered along the way.
An audit that checks code against its own docs cannot catch code and doc that are wrong together. Three violations (a Query reading a write Repository, a domain class carrying JPA, a notification module in the wrong layer) had passed dozens of prior audits. Only one was a bug, and the other two show why none of those audits could have caught them.
Why an error-message enum key has to equal its value, and the four-field error response shape.
This blog uses cookies for visit measurement (Google Analytics) and ads (Google AdSense). Declining hides the ad slots only; the measurement and ad scripts still load. See the privacy policy.