Implementation doctrine for the age of AI: what to do, and what once-good advice to unlearn, in one document. It distills a study of sixteen expert programmers' philosophies; credit is given by name where a specific idea is owed to a specific person, but this file stands alone. Patterns state the rule and its reasoning. Retired blocks inside them name the old advice the rule replaces — advice that was genuinely right, often for decades, before the ground moved.
Almost every classic rule of software engineering encoded one scarcity: human effort. Machine cycles were abundant, programmer hours were the binding constraint, and good advice spent the abundant resource to save the scarce one. AI flipped the sign. Generation is now abundant — code, prose, glue, migrations all pour out at negligible cost — and three new scarcities bind instead: verification (knowing the abundant output is correct), context (what a model can hold and attend to), and trust (what you can safely let an agent do). Compute at fleet scale joined the scarce side too, metered per token.
Every pattern below is an instance of spending abundance to protect the new scarcities. Every retired rule is one that still spends the new scarcities to save the old abundance. When you meet a rule of thumb not covered here, run that test on it.
Plausibility is what language models optimize for. Reading generated code and believing it is the exact self-deception Richard Hipp built SQLite against — reading convinces you code works; only testing demonstrates it — now industrialized. Confidence in agent-written code can only be manufactured his way: tests that fail when the code is wrong. Test-to-code ratios that once looked fanatical (SQLite ships hundreds of times more test than library code) are the correct shape for an era where writing code is free and knowing it works is the bottleneck.
In practice: no fix ships before a regression test that reproduces the bug. Failure paths — I/O errors, timeouts, permission denials, partial writes — get tested as hard as happy paths, because that is where field failures live. A branch no test can reach is speculative generation; delete it. Prefer system tests of user-visible behavior over mock choreography (the surviving half of Hansson's testing critique): a green suite that never exercises the real stack proves little. And measure before believing any claim — performance, quality, or model behavior; if you cannot explain the mechanism behind an improvement, you don't understand it and will misapply it (Fabian Giesen's standard).
Retired: "careful human review will catch it." Review-by-reading scaled when humans wrote the code, because production and review ran at the same rate. Generation now outruns review by orders of magnitude, and generated code is optimized to look right — inspection selects precisely for missing the bugs that matter. Human review moves up-stack: contracts, boundaries, and threat models, not line-by-line prose reading of diffs.
Prompt injection is the confused-deputy attack under a new name (Kenton Varda's frame). You cannot reliably instruct a model out of deputy behavior — but an agent that cannot name a resource cannot be talked into abusing it. Security that depends on the model honoring its system prompt is security that depends on nobody misbehaving; encode the discipline in the substrate instead. Deny by default; grant scoped, revocable capabilities per resource; a linter — or an agent — should not hold network access it never asked for (Ryan Dahl's rule, now literal). Treat all generated code as untrusted by construction: a workload class, not an anomaly, sandboxed accordingly.
Cognitive load stopped being a metaphor: it is the context window, priced in tokens. John Ousterhout's deep modules — simple interface, powerful implementation — become the highest-leverage property of agent-maintained code, because an agent that can act on an interface without reading the implementation spends its context on the task instead of on archaeology. Every concept, layer, and indirection in a codebase is tokens loaded before any agent can act; shallow modules and pass-through layers force exactly the multi-file, cross-boundary edits where agents make correlated mistakes. Small, orthogonal interfaces matter more when the caller is a model: a wide tool schema is misused in proportion to its width (Rob Pike: the bigger the interface, the weaker the abstraction).
Retired: "boilerplate is free — the wizard writes it." Scaffold was harmless when an IDE emitted it once and no one read it again. It is now read repeatedly, expensively, by every agent that loads the codebase; its carrying cost is per-read, forever. Scaffold minimally; every generated file earns its place the day it lands.
A pipeline whose token spend, per-stage latency, and tail behavior cannot be estimated on a napkin is a pipeline nobody understands (Jeff Dean's gate, transplanted intact). The ratios changed — prefill versus decode, cache hit versus recompute, small model versus frontier call — but the discipline is identical, and one excuse died: napkin math is itself a task agents do free, on demand. Fan-out logic is hostage to its slowest call; hedge, time out, and stop using bad backends. And keep Dean's 3 a.m. test: when the system is slow or wrong in production, which page tells you where the tokens went? Most agent stacks cannot answer.
Retired: the blanket "avoid premature optimization" pass. Knuth's line was right against hand-tuning inner loops before profiling; it degraded into a license never to think about cost at all. Architecture-level pessimization was never what it licensed — and waste now carries a per-request price in dollars. Estimate first, refuse waste always, and still defer micro-tuning until profiles exist. Both halves, no pass.
Casey Muratori's non-pessimization is the missing discipline of agent systems: most are slow and expensive not because inference is slow but because they do enormous work the task never required — re-reading unchanged files, re-deriving known structure, carrying unconsulted context, chaining calls where one would do. Name the sin before optimizing anything. Its structural form (Rich Harris): move work from inference time to build time. Anything knowable statically — structure, routing, templates, schemas — should not be re-derived by a model per request; when possible, use the model once to generate specialized code, then run that code a million times. Keep prompt prefixes stable and inputs statistically boring: caches and predictors at every layer reward monotony (Mike Pall's monomorphism, restated for KV caches).
Retired: "developer time is expensive, compute is cheap." The founding trade of the scripting era — burn 10–100x machine resources to spare scarce human effort, stay in the slow comfortable stack because learning a faster language costs months. Agents collapsed the human side of that trade: the fluency cost of the faster language is now the agent's problem, and porting is days. Meanwhile the compute side stopped being cheap and, in agent systems, gets paid twice — the product burns it, and the tokens spent maintaining a bloated stack burn it again. Choose implementations by machine merits and let agents pay the learning cost.
Regeneration is cheap; understanding is not. Code that is small, comprehensible, and rewritable in an afternoon now beats code engineered for imagined extension (Salvatore Sanfilippo's bet, strengthened): the AI era lowered rewrite cost far more than it lowered understanding cost. Concrete first; compress only on demonstrated repetition, extracting exactly what is shared (Muratori's semantic compression).
Retired: "never rewrite from scratch." Joel Spolsky's warning was right when a rewrite was a multi-year bet and the accumulated bug fixes lived in veterans' heads. With a real test suite, the tests — not the veterans — carry the accumulated fixes; the implementation becomes inventory, not capital, and deliberate rebuilds become a routine tool. The rule survives in sharpened form: never rewrite without an executable contract; with one, price the rewrite against the accumulated-wart tax honestly. And remember ecosystems are people — compatibility bridges decide whether rebuilds survive (Dahl learned this twice).
Retired: "don't reinvent the wheel — there's a library for that." Right when writing the wheel cost weeks and many eyes reviewed the package. Generating the 300 lines you need costs minutes, while dependency lifetime costs — supply-chain risk, abandonment, churn, and agents hallucinating APIs of libraries they haven't read — did not fall. For small, stable subsets, generate owned code with provenance notes and tests. Import where depth is real: cryptography, time zones, database engines — wherever the domain's edge cases exceed what your tests can hold.
Retired: first-pass DRY. Deduplication was capital preservation when every copy was maintained by hand. Code is cheap now; coupling is the expense. Premature unification welds call sites together at the exact moment divergence became cheap to maintain and repetition became easy to compress later.
Agents are constitutionally tactical: every prompt is "make it work now." The behaviors that keep systems alive over years — Ousterhout's strategic design investment, Muratori's second pass that compresses proven repetition, Sanfilippo's refusal of scope — do not emerge from models, because nothing in the loop asks for them. They must be scheduled by the harness or the human as explicit, recurring steps: a design-review pass, a compression pass, a deletion pass. Left alone, agents stop at pass one, and codebases ratchet monotonically toward entropy at machine speed.
A model's intuition is the conventions of its training distribution. Yukihiro Matsumoto's principle of least surprise — properly qualified as least surprise to a fluent user of one consistent mental model — turns out to describe LLMs exactly: in a consistent, conventional system the agent's next guess is right; every local exception to the house worldview is a permanent error site that no amount of prompting repairs. Consistency now pays twice, in human intuition and machine priors. APIs shaped like the ecosystem's million training examples get used correctly on the first try; locally-nicer bespoke shapes fight the model's priors forever (Anders Hejlsberg's "meet users where they are," second audience).
Retired: "choose boring technology — optimize for the hiring pool." The hiring argument is dead: agents know every stack. But the boring choice gained a justification the old advice never had — training-data density — so the conclusion often survives with a different proof. The antipattern is conflating "boring" with "bloated": a lean, less-common stack an agent can hold whole beats a conventional one drowning in concepts. Choose by prior density and conceptual weight, not headcount availability — and when you deviate, write the deviation down where agents will read it.
Half the readership of every docstring, tool description, README, and error message is now models — and for them the prose is the interface: it is what retrieval, tool selection, and code modification actually operate on. Ousterhout's comment-first test becomes mechanical: if the honest description of a thing needs a trail of "except when," redesign the thing, not the description (Dan Abramov's two-paragraph rule, now checkable). The unexplained "why" — the odd constant, the deliberate non-obvious approach — is precisely what an agent will "fix" into a regression; design comments are load-bearing. Error messages convert from courtesy to control loop: they are the feedback agents self-correct on, so state what went wrong in the caller's terms and guess intent (Matsumoto's "did you mean" humanism; Hejlsberg's "diagnostics are UI"). An unactionable error no longer wastes a mood; it breaks a repair loop.
Retired: "good code is self-documenting; comments are failures." And its cousin, "tool UX is polish for humans." Both assumed the only reader could infer intent. The literal-minded reader is now the primary one, and comment rot stopped being inevitable — agents can audit comment-code agreement continuously.
Two coupled obligations. First, misuse-resistance outranks brevity by a wider margin than ever: agents commit the most-natural-wrong-call at volume and never notice silent works-until-production failures, so the wrong call must be unrepresentable or loudly broken (Abramov). Second, every system has a fast path and a place where inputs fall off it; the sin is silence. Mike Pall shipped LuaJIT with a public list of exactly what didn't compile — an explicit, documented limitation is a feature; a silent performance cliff is a bug. Substitute "capability" for "performance" and it is the missing discipline of AI products: document which inputs the model handles badly, which tools the agent cannot use, where quality falls off. His trace-compiler architecture is also the reference design for routing: a guarded cheap path (small model, cache, template) that bails honestly to the expensive path (frontier model, human) — never one undifferentiated prompt handling all cases, taxing every request with the rare case's baggage.
David Heinemeier Hansson's conceptual compression and one-person framework become a hard test: can a single agent hold, modify, and ship this system end to end? Every concept is context an agent must load; every service boundary multiplies state it cannot see, auth it must carry, and versions that skew — empirically the exact place where multi-agent reliability collapses. One codebase is one context; the monolith found its sharpest argument two decades after it was named. Within it: data structures first (Pike's Rule 5 — given the right structures, the code becomes nearly self-evident, which is also a prompting strategy: give an agent the data shapes, not prose vibes); and write the usage code or the eval before the implementation (Muratori) — the call site is the spec, and for agents it is the best one ever devised.
Shipped contracts — APIs, file formats, wire protocols, and now prompts, tool schemas, and agent integrations — accrete dependents at machine speed. Hyrum's law was an observation about large organizations; agents made it a law of nature, because uncountable integrations quietly depend on every observable behavior within days of shipping. Evolve additively. Ratchet strictness with opt-in flags and per-file escape hatches (Hejlsberg's migration shape — agents can mass-migrate, but escape hatches keep intermediate states shippable). Compatibility is trust spent, slowly re-earned (Matsumoto after Ruby 1.9). Rebuild deliberately, on accumulated documented regret, with bridges — not reflexively.
Most misapplication of good doctrine is a phase error. The map, from first probe to final refinement:
| Phase | Governing patterns | Named tradition |
|---|---|---|
| 0. Feasibility | Arithmetic first (4); usage-code first (11) | Dean; Muratori, Pike |
| 1. Prototype | Refuse unneeded work (5); few concepts (11); scope refusal (7) | Muratori, Hansson, Sanfilippo |
| 2. Architecture | Deep modules & small interfaces (3); authority model (2); static/dynamic split (5); one worldview (8) | Ousterhout, Pike, Varda, Harris, Matsumoto |
| 3. Growth | Tails & observability (4); compat ratchet (12); monolith discipline (11) | Dean, Hejlsberg, Hansson |
| 4. Optimization | Measure & derive (1, 4); fast/slow paths & cliffs (10); the compression pass (7) | Giesen, Pall, Muratori |
| 5. Hardening | Manufactured reliability (1); owned dependencies (6); permanent promises (12) | Hipp, Hejlsberg, Matsumoto |
Some disciplines bind to recurring moments, not phases: the day platform defaults get set (they compound into ecosystems — Dahl); every revision of a public contract (12); any claim that needs arithmetic (4).
Misapplication runs both directions. Hardening discipline applied at the prototype phase calcifies guesses — the most expensive mistake available. Prototype directness carried into hardening leaves entangled code nobody dares compress. Optimization doctrine gates its own entry: prove it's hot first. Tail-latency machinery on a ten-requests-per-second prototype is cargo culting.
Doctrines earn trust by surviving opposition, and this document contains its own oppositions: deep modules (3) versus refuse-the-layer (5); holdable monoliths (11) versus platform boundaries (2); owned code (6) versus ecosystem convention (8); deletability (6) versus permanent promises (12). These tensions are features. When reviewing a design, never staff the review with allies: include the school the proposal's instincts oppose, and one voice from the next phase — where the current phase's debts come due. Agreement between allies is not a review.
The one-line summary: generation became free; verification, context, and trust became the product. Spend accordingly.