Skip to main content
Architecting Software in 2026 · 2 of 6 partsEngineering11 min read

Architecture as Code Has Three Jobs: Describe, Govern, Remember

In Bullpen's first week, the planned architecture model still named a stock provider I had turned down, and a footer change and a test followed it. Nothing failed. I use that incident and two others to argue that architecture as code has three jobs: describe the system, govern it, and remember why it looks the way it does. Then I follow the first real change after the fixes, product analytics, through all three: a decision record before the model could name it, built and planned architecture that disagreed, and a rule that composes the two instead of editing either.

← All Posts
Listen to this post
0:00 / 0:00
2/4

On September 26, Bullpen's planned architecture model and my decision log disagreed. Its Part 6 moment still named Alpaca and Massive as stock providers, beside a Stripe node. The log, outside the repository, said I had picked another stock source a day earlier, with Alpaca as a backup; payments were never in the plan. A change to the footer that day followed the model: it wrote "U.S. stocks by Alpaca, from Part 6", and added a test that passed because it encoded the same wrong provider. The decision said the opposite. Nothing failed: nothing checked the model against the decisions.

Architecture as code has three jobs, and that incident shows all of them: describe the system, govern it, and remember why it looks the way it does. In week one, Bullpen did the first job well, the second in part, and the third barely at all. Whoever changes the code next trusts whichever committed file they read. So every record needs a check that keeps it true, and every rule needs a place where it fails.

Three incidents in Bullpen's first week

Bullpen is the paper-trading league I am building in public for this series, and Part 1 explains why it starts as one deployable. In its first six days, three things went wrong, and in each one the rule existed only as prose, or not at all.

DateWhat happenedWhere the rule lived
Sep 25–27Pushed history was rewritten once, closing PR #10; two fixes went straight to main; three pull requests were squash-merged against the repository's history rulesProse: handoff notes and my status notes
Sep 26The planned model still named a stock provider I had turned down, and the footer and a test followed itMy decision log, outside the repository
Sep 25–26Four decisions landed as ADRs, and none of them reached the timeline that ties decisions to the architectureNowhere: nothing compared the two

None of these did lasting damage; PR #12 replaced #10, and nothing was lost. The middle row is the worst, because the decision was not even written in the repository. It lived only in my decision log and in my head.

The three jobs of architecture as code

Many architecture-as-code setups I have seen do one of the three jobs and call it done. In one sentence, the three fit together like this: the model describes the architecture, the decision records hold its intent, and CI checks that the two stay consistent.

  • Describe means a model of the system that a machine can read and validate. In Bullpen, that is a FINOS CALM (Common Architecture Language Model) document per part of the series, plus a pattern every model must satisfy. The landing page draws its architecture explorer from those files.
  • Govern means rules the build enforces: structure rules written in an architecture definition language (ADL), a call-budget check, and Turborepo Boundaries, which reports a violation when one app imports another under a tag rule that denies it. A fitness function is just a test, and I treat it that way.
  • Remember means records of why: numbered architecture decision records (ADRs), and a CALM timeline that ties each moment of the architecture to the decisions that produced it.

None of this is new in isolation; I tried the describe-plus-govern pair before with a Structurizr workspace. What Bullpen adds is all three jobs on one system that changes every day.

Describe: a model that gets believed

CI validates Bullpen's Part 1 model and its timeline with calm validate --strict on every pull request, and main is green. That is the trap. Validation checks shape, not truth. A model can satisfy every schema rule and still name a provider nobody chose.

This is the fragment the footer change followed, from planned/part-06.architecture.json as it stood on September 26:

json
{
  "unique-id": "alpaca-provider",
  "node-type": "external-system",
  "name": "Alpaca"
}

The planned moments are my plan for Parts 2 to 6. To whoever reads one next, person or agent, a plan in a committed JSON file reads exactly like a fact. The repository held the model and the decision lived outside it, so the change sided with the repository.

This is architecture drift in its quietest form: the code was fine, and the model had drifted from the decisions. The fix is a drift check that runs on every pull request. Every external system in the current moment, and in every planned moment, must be decided in an ADR or listed as in use or proposed in docs/data-sources.md. A name that appears only as rejected does not count, because counting it would let the week-one mistake through. Every directory the ADL defines must also map to a node in the current model, and back.

The footer test and the drift check look alike, but they are different kinds of check. In a live session on October 1, Neal Ford gave a one-question litmus test: do I need any domain knowledge to write this code? The footer test needed to know which provider Bullpen uses, so it was a functional test, only as right as whoever wrote it. The drift check knows nothing about market data. It compares the model with the records, which makes it a fitness function.

To see whether it would have caught the week-one mistake, I put the Alpaca node back into the planned model on the Part 2 branch and ran the checks. Two of them failed. This is the first one, with its lines wrapped to fit:

text
✗ model currency: every external system in a CALM moment is decided
  in an ADR Decision or listed as in use or proposed in
  docs/data-sources.md
  where:   architecture/calm/planned/part-06.architecture.json
  why:     Node "alpaca-provider" is the external system "Alpaca",
           which ADR-0008 Alternatives records only as turned down
           (ADR-0009).
  fix:     Rename the node to the provider that was chosen, or write
           a new ADR that reverses the rejection before the model
           names it.

The second failure came from the check that keeps the landing page's explorer in step with the model: it found a node the explorer does not draw. The footer change in September would have failed CI on both before it could merge.

Govern: a rule has to fail in the build

A useful rule fails at the moment of the mistake, and its message says what to do instead.

The structural rules live in one ADL file, written in the style Ford and Mark Richards use in Architecture as Code. This block says which parts of Bullpen may never depend on which:

text
# Disallowed dependencies (direct and transitive)
ASSERT(Landing HAS NO DEPENDENCY ON Trading App)
ASSERT(Trading App HAS NO DEPENDENCY ON Landing)
ASSERT(UI HAS NO DEPENDENCY ON Landing, Trading App)
ASSERT(Contracts HAS NO DEPENDENCY ON Landing, Trading App)

The comment settles a question the language leaves open: a dependency laundered through a library counts. Ford has a name for the import this block forbids: cheating on dependencies. In a monorepo, the code you want is right there in the next folder, whoever is typing, and nothing but a rule stops the import.

These are boundaries with no network call, the cheapest kind Part 1 describes. The ADL keeps them honest until one of the five reasons earns a service. Every assert in the file maps to a check that can fail it, and a rule no check enforces fails the build too. Every check fails in one format: the rule exactly as the ADL states it, where it broke, why the rule exists with its ADR number, and the smallest fix.

The history rules left prose too. A CI job, commit-policy, checks the author and trailers of every commit in a pull request, and recognises the right identity by a fingerprint, so my address appears in no file. Branch protection on main requires a pull request and both checks, allows merge commits only, and has the admin bypass off. That last part catches me, since the squash merges were mine.

Remember: records drift too

The remembering job has its own drift. In week one, Bullpen's timeline linked the Part 1 moment to ADR-0001 through ADR-0003, while four more decisions landed that same week. The timeline now links every ADR from the moment of its part, and the drift check fails when one is missing.

The plans raise a harder question: which moment in time does a planned file represent? It cannot be both the plan as it stands today and the prediction I made in week one. So Bullpen keeps two records with two promises:

  • Current validity lives in planned/. Those files are the plan as it stands, and the drift check holds them to the decisions in force today. When I recorded the stock decision on September 26, Alpaca, Massive and Stripe came out of Parts 3 to 6.
  • Historical fidelity lives in git. The predictions exactly as I wrote them on September 25 stay at the tag predictions-week-1, and that is what Part 6 compares with what shipped.

Neither one is rewritten to match what shipped. A week later, that rule got its first real test.

When what runs and what is planned disagree

On October 1 I added product analytics to Bullpen, Google Analytics and Microsoft Clarity, loaded only after a visitor consents. Both joined the Part 1 model as external systems. Then I scrubbed the explorer to Part 2, and both disappeared, along with their four connections. The page now said I had removed analytics after Part 1. Nothing had removed them: the plans for Parts 2 to 6 were written before analytics existed, and the explorer drew each planned part from its plan alone.

The obvious fix was to add both systems to the five plan files. That would have put one built fact in six places, each able to drift on its own, and it would have treated a plan's silence as something to correct. A plan that says nothing about a running system is not predicting its removal.

So ADR-0011 composes the two records when the explorer draws them. Whatever is built and still running carries forward into every later planned part, labelled "Built in Part 1", and nothing is written into the plan files. A plan still wins for anything it names, so a planned replacement stays one: the Part 1 price snapshot service is not drawn beside the two services Part 3 plans in its place.

The other two jobs held the change in place. ADR-0010 recorded the decision before the model named either system, and on the Part 2 branch, renaming one of the new nodes to a system no ADR mentions fails model currency, as it should. The explorer consistency check accepts a carried card only if it exists in a built moment and names the moment that first built it, so a carried card cannot be invented.

Run the checks yourself

The repository is public, and each part of the series lives on its own branch. This part's branch is post-02-architecture-as-code. With Node 22 and pnpm 12, a fresh clone runs the same architecture checks CI does:

bash
git clone https://github.com/tiarebalbi/bullpen.git
cd bullpen
git checkout post-02-architecture-as-code
pnpm install
pnpm check:arch

On a clean clone it ends with check:arch passed: 9 checks held, 19 rules enforced. To watch one fail, add any external-system node to architecture/calm/planned/part-06.architecture.json and run it again.

Where architecture as code still breaks

These checks narrow the gap. They do not close it.

  • A check only fails what I thought to encode. Intent cannot be checked. A claim the code can falsify belongs in a check; a reason it cannot falsify belongs in an ADR.
  • The checks themselves can be wrong. Bullpen's checks were generated from the ADL, which Ford calls interpolated fitness functions, and he was blunt that the architect must review every one. Each ships with a fixture that must fail, and none should guard the build before a human has read it.
  • No second reviewer. Branch protection requires green checks but no approval, because there is no one else to approve. In a team, a required review is the one layer that can question intent.
  • A check can be routed around. I am the admin who turned the bypass off, and I can turn it back on. Boundaries is also still marked experimental, and a rule built on it inherits that status.
  • Every rule is upkeep. The drift check makes the plan something I maintain, and each new check is one more thing every change must satisfy before it ships.

A checklist for the next rule I write down

Before a rule goes into a doc, a ticket or a README, I now ask five questions:

  • Where does it fail: CI, branch protection, or nowhere?
  • Does the failure name the rule, the place, the reason and the fix?
  • Is there a fixture that proves the check can fail?
  • Which record does it cite, and what keeps that record current?
  • If it cannot fail anywhere, have I filed it as an ADR and stopped calling it a rule?

The goal was never to keep architecture files in sync for their own sake. It is to stop a committed picture of the system, like that Part 6 node naming Alpaca, from quietly overruling the decisions it is meant to follow.

Read next

Still here? You might enjoy this.

Nothing close enough — try a different angle?

Architecting Software in 2026 · 2 of 6

Part 3 is in progress.

New parts land on Mondays, 9am Pacific — leave an address and I'll send each one the day it ships. Nothing else.

Was this helpful?

Leave a rating or a quick note — it helps me improve.

Related Posts

Engineering

When to Use Microservices in 2026: Scaling Alone Is No Longer the Reason

On the day the Dow posted a then-record gain, Robinhood's customers could not trade from the open to the close. Splitting out the hot part so it scales alone was the old answer; a serverless platform now does that per function. I open this series with Bullpen, a paper-trading league on live market data, and argue when to use microservices in 2026: only for a reason I can name, with a check, a measurement or a bill behind it.

AI

Fitness Functions Are the Control Plane for Agentic Coding

The last post asked where the control point went. Here is the first answer I am willing to defend: it did not go away — part of it compiled. Architectural fitness functions, pointed at coding agents, become the control plane that lets a developer stay in charge without becoming the bottleneck: judgment compiled once into deterministic gates that enforce at machine speed, with failure messages written as prompt engineering for the retry loop. Then the complication that shapes the whole post: the moment an agent optimizes against the compilation, the compiled control becomes an object of attack — and fitness function design inherits an arms race, with a Kotlin/ArchUnit constitution to make it concrete.