Tooling · Architecture
Compliance as Code:
What an Architecture Checker Catches, and What It Keeps Missing
Turning an architecture doc into automated checks reliably catches code that ends up in the wrong place. Three kinds of drift kept getting past it anyway: a wrong name inside a correctly placed file, code that is wrong together with its own doc, and a disagreement that only exists between two implementations.
Design docs rot the same way comments do: a little at a time, until the day someone reads the doc, reads the code, and notices they've been describing two different systems for months. The fix that held up for me wasn't writing better docs. It was building an architecture checker for my example project, which implements the same backend design in five languages side by side. The checker is a script that statically checks whether code follows the documented rules and scores the result, so catching drift no longer depends on a reviewer noticing it.
Placement rules (which folder, which layer, which import direction) did their job from the start. The drift the checker missed came in three kinds, and each one had to be found by hand at least once before a rule could cover it.
What a Rule Is Allowed to Assume
The example business code (an Account domain, later joined by Card and Payment) is an illustrative sample, not a fixture the checker is allowed to depend on. A rule must never take "the Account domain behaves like this" as its premise; it has to check the architectural pattern in a way that would hold for any domain with that shape.
The checker evaluates compliance with architectural rules (layer placement, dependency direction, naming, transaction boundaries, the Outbox pattern), never whether the business logic inside those structures happens to be correct. Order cancellation rules, payment approval conditions, inventory reservation policy: none of that is in scope. That kind of content can appear in a doc's example or in the runnable examples/ code, but it can never become a required premise of a core rule. Blur that line, and a new rule couples itself to one specific business domain without anyone noticing, degrading the checker from a framework-agnostic architecture guide into an Account-service-only linter.
Many Small Assertions Over One Golden Implementation
The checker prefers partial scoring built from many small, independent assertions over pinning one single "this file must look exactly like this" reference implementation. Each rule checks individually whether a specific violation is present and the results sum into a score. The premise is that multiple valid implementations of the same principle can coexist, and a structurally sound choice shouldn't lose points just because it differs cosmetically from the example in examples/.
Blind Spot One: Names Inside the Right File
The first rules checked the things you'd expect: is there a domain folder, does the Interface layer avoid importing Infrastructure directly, does a Repository interface live where the docs say it should. That catches an entire class of drift, and it also has a precise blind spot. It says nothing about whether a method inside the right file is named correctly, or whether a class in the right folder depends on the right thing.
An audit by hand surfaced this gap. A Repository method-naming convention (list lookups always named find<Noun>s, saves always save<Noun>) was violated in four of the five implementations, in different ways each time: a bare save with no noun in two of them, a "dedicated findOne plus separate findAll" pair reintroduced in an older domain in four of them. Every one of those files passed every structural rule that existed at the time, because none of those rules had ever looked at a method name.
Each time the same class of drift turned up, the diagnosis was the same: no tool checked for this specific thing. Structural checks verify placement. They say nothing about naming conventions, dependency direction inside an already-correct folder, or exact method-name patterns. Each of those needed its own dedicated rule, written only after a manual audit found the gap by hand at least once.
Blind Spot Two: Code and Doc Wrong Together
A "does the code match its own docs" audit can pass cleanly while both the code and the doc are wrong together. One implementation's own CQRS documentation had captured a Query Handler directly using a write-capable Repository as the correct example. The code matched the doc perfectly, and both were violating the root principle that a Query side should never see a write-capable interface at all. An audit that only checks doc-vs-code agreement is structurally incapable of catching this, because agreement is what it's checking for, and here the agreement itself was the bug.
What surfaced it was comparing the local doc against the root principle instead. Once that comparison was made a standing rule instead of a one-time manual pass, it caught the same class of violation independently in other files the first time it ran.
Blind Spot Three: Disagreement Between Implementations
Audits that go one language at a time can never see a structural disagreement that only exists between languages. A notification-sending concern lived inside the domain module in two of the five implementations and in a separate shared top-level module in the other three. Neither choice was obviously wrong on its own, each passed its own language's checker cleanly, and the inconsistency was invisible to any single-language review by construction. It only became visible once the same concept was compared side by side across all five at once, which is a different kind of audit from "review this one codebase against its own docs."
Turning a Finding Into a Permanent Rule
The pattern that held up: find a violation by hand once, fix it, then write a rule that would have caught it, and run that new rule immediately, before assuming the code is now clean. The Repository-naming rule is the clearest example. Once it existed as a mechanical check rather than prose, running it against a domain nobody had thought to re-check (the authentication domain, in three of the five languages) immediately turned up three more violations that had never been in scope for any prior audit, because every earlier audit had only ever looked at the two business domains everyone kept thinking about. A dependency-direction rule, an ID-format rule, an error-response-schema rule, a soft-delete-filter rule: roughly thirty rules accumulated this way across five languages, each one born from an actual bug like this one, not a hypothetical one.
The yield drops over time, and that's expected. Early on, each new rule category found three or four violations; later, most new rules found zero, because the low-hanging cross-language drift was already closed. A rule that costs an afternoon to write and catches nothing today is still standing guard against next month's regression.
What This Buys You That a Doc Alone Never Could
A CI pipeline that runs the checker on every change means a pull request that violates layer placement, naming convention, or dependency direction fails the build before a human ever has to notice it in review, the same way a linter catches a syntax issue before a reviewer has to point it out by hand. My self-review checklist is deliberately written to double as a spec for the checker: every new checklist item gets asked, "can this also be verified mechanically," before it's accepted as prose-only.
docs/harness.md (the checker's design principles in full, in my example project that implements the same backend design in five languages) · docs/checklist.md (the self-review checklist most of these rules were built from) · The Doc Said "Done": When to Stop Adding Checks (the naming finding followed rule by rule, and how a shrinking yield showed when to stop)