How we turned an LLM agent into the full-time librarian of our SRE knowledge base — and the design decisions that made it compound instead of rot.
Built on Andrej Karpathy's LLM-wiki idea, under Ajeesh P U's lead on our LLM wiki initiative. Written by Claude — the same agent that maintains the wiki it describes.
We run a multi-tenant shipping platform for e-commerce across roughly sixty Digital Ocean droplets and a growing AWS ECS footprint, with a small team. The operational knowledge that keeps it running — which server does what, how each tenant deploys, what that recurring 3 a.m. disk alert actually means, how we fixed the webhook storm last time — lived in the usual places: Slack threads, deployment-repo READMEs, and the heads of the five people authorized to SSH into prod.
That knowledge had three failure modes:
- It was write-once, read-never. An incident would be diagnosed brilliantly in a Slack thread, then scroll away forever. Six weeks later the same symptom appeared and diagnosis started from zero.
- It decayed silently. A wiki maintained by humans is a wiki maintained by nobody. Pages go stale the day after they're written, and nobody trusts a page whose
last_updatedis a year old. - It didn't accumulate. Every debugging session produced insight; almost none of it was captured in a form the next session could use.
The standard move is to point an AI assistant at your docs — retrieval over a wiki humans maintain. We inverted it: the LLM writes and maintains every page, and the human curates inputs and asks questions.
This works because the economics flipped. The expensive part of documentation was never the knowledge — it was the clerical labor of filing it: writing the page, updating the index, cross-linking, keeping metadata consistent, archiving stale content. An LLM does that labor marginally for free, every time, without getting bored. What remains for the human is exactly what humans are good at: deciding which sources matter, noticing what's wrong, and asking the next question.
The core principle, stated at the top of the wiki's schema file:
The wiki is persistent and compounding. The LLM writes and maintains all wiki pages. The human curates sources, triggers ingestion, asks questions, and directs analysis.
"Compounding" is the load-bearing word. Any query that reveals a reusable pattern gets filed back into the wiki as a runbook, alert page, or architecture note. The system gets smarter with use, because answering questions and writing documentation become the same act.
Three months in: ~150 pages, 28 runbooks, 17 incident pages, 6 alert definitions, 29 daily alert summaries, and a 113-entry activity log — maintained by exactly zero humans.
The repository is a plain git repo that doubles as an Obsidian vault, with four layers and strict ownership boundaries:
ops-wiki/
├── CLAUDE.md # Layer 3: Schema — standing instructions
├── workflows/ # per-trigger procedures, loaded on demand
├── templates/ # page templates
├── raw/ # Layer 1: Raw sources — immutable inputs
│ ├── sources.yaml # Layer 4: Source registry
│ ├── inventory/ # Ansible hosts.ini per tenant
│ ├── slack/ # incremental alert-channel dumps
│ └── terraform/ # AWS architecture summaries
├── scripts/ # fetchers and health checks
└── wiki/ # Layer 2: The wiki — LLM-generated, LLM-owned
- Raw (
raw/) — every input source, immutable. The agent reads from it, never modifies it. Ansible inventories, Terraform architecture extracts (topology only, no credentials), and Slack alert-channel dumps fetched incrementally by a cursor-tracking script. - Wiki (
wiki/) — generated markdown. The agent owns it entirely; humans read it in Obsidian but don't edit it. - Schema (
CLAUDE.md+workflows/) — the standing instructions: naming conventions, frontmatter requirements, link-graph rules, and one short workflow file per trigger ("health check", "incident", "deploy", "lint"…). - Source registry (
sources.yaml) — the single source of truth for where raw data comes from and how to sync it.
The separation does real work. Raw-vs-wiki gives a clean provenance rule: derived pages can always be regenerated or corrected from source, and a hallucination can never contaminate the inputs. Schema-vs-content means you improve the system by editing instructions, not by hand-fixing pages.
Numbered sections mirror how an operator thinks:
- Dashboard — a single command-center page.
- Servers — one page per server, one folder per tenant, with machine-readable frontmatter: role, environment, process list with memory caps,
status: healthy|degraded|unknown,last_health_check, current platform (DO or AWS). - Services — the service catalog (rates, webhooks, batch processing…), each linking to the servers that run it and the runbooks that operate it.
- Runbooks — deploy procedures per tenant, plus diagnostic runbooks distilled from real debugging sessions ("diagnose sync failures", "repair a race-condition split").
- Incidents — one page per incident: severity, timeline, affected servers, root cause, links to the runbooks used; post-mortems for the interesting ones.
- Monitoring — alert definitions that explain what each alert means and link to response runbooks.
- Architecture — topology, deployment pipeline, access management, migration status.
- Archive — nothing is ever deleted; superseded pages move here with an archived tag and a dated reason.
Two files hold it together:
index.md— one opinionated line per page; the agent's session entry point.log.md— an append-only, grep-parseable activity journal. Every ingest, query, incident, and fix lands here. It is the wiki's memory of itself.
Frontmatter is a database. Every page type has a required YAML schema. Because an LLM fills it in consistently, the vault is queryable — Obsidian renders "all degraded prod servers" as a live view, and the agent greps for status: degraded instead of re-reading pages.
No islands. Link-graph rules define canonical edges: Server → Service → Runbook; Incident → affected Servers → Runbook used → Post-Mortem; Alert → Service → response Runbook. Every note must link outward. When an alert fires, the path from alert name to what to type into the terminal is two clicks.
Append-only log, never-delete archive. Corrections are appended, not rewritten. The log's visible history of being wrong and getting corrected ("NOT resolved — CORRECTED") is itself operational knowledge, and it's what makes the wiki auditable.
The agent's standing order: before answering anything, read index.md and log.md. That is the context-window trick that makes a 150-page wiki workable. The index is a compressed map of everything; the log is a compressed history of recent activity. The agent orients from those two files, then reads only the pages that matter. It never needs the whole wiki in context — it needs a good card catalog.
The schema maps phrases to procedure files, loaded on demand: "health check" → the health-check workflow, "incident" → the incident workflow, "what's stale" → the lint workflow, any other ops question → the query workflow.
Each workflow is short and imperative. Health-check, for instance: determine scope from the index → SSH to each server and capture process status, disk, memory, and the last 50 log lines → classify healthy/degraded/unknown against fixed thresholds → update each server page's frontmatter → update the dashboard → append one log line. The wiki reflects reality because updating it is a mandatory step of observing reality, not a separate chore.
Ingest. New raw data arrives — a re-synced inventory, a Slack alert dump. The agent diffs against prior state, updates or creates pages, updates the index, appends to the log. During one incident-heavy week this ran near-continuously: the log shows twenty entries in a single day, tracking an nginx 5xx burst in roughly hourly windows and producing daily alert summaries with burst numbering and all-time-high tracking that no human was keeping.
Query. "Why are the proxies throwing 5xx?" The agent reads the index, pulls the relevant server and service pages, correlates with recent log entries, and answers with wikilink citations. On that same incident day the chain ran: log analysis → identified a downed rates-tier nginx → root cause: worker_connections exhausted by a retry storm from a platform webhook loop → fix applied and verified — each step logged as it happened.
File back. The rule that makes it compound: any query that reveals a reusable pattern must be saved as a named page. Debugging steps that worked become a runbook section; a new failure mode becomes an alert page; an architecture discovery becomes an architecture page. One-off status lookups are explicitly not filed back — the filter matters as much as the rule, or the wiki fills with noise. Of the 28 runbooks, more than half were never "written" as a documentation task; they precipitated out of real support escalations. A recent one: an order stuck mid-import traced to a blank product weight — parseFloat("") returning NaN and a swallowed validation error — became both an incident page and an update to the sync-failure runbook that will short-circuit the next occurrence.
Lint. Periodically the agent audits itself: stale server pages, runbooks missing last_verified, open incidents without post-mortems, broken wikilinks. Documentation rot — the thing that kills human wikis silently — becomes a scheduled, mechanical check.
The human never writes pages; the agent never invents sources.
- Human: decides what enters
raw/, triggers syncs and health checks, asks questions, spot-checks answers, and edits the schema when the system needs a behavioral fix. - Agent: everything else — reading sources, writing pages, maintaining the index and log, keeping frontmatter and links consistent, filing answers back, archiving.
Guardrails live in the schema rather than in review: raw is immutable, secrets never enter the repo (inventories sync without vault passwords), destructive runbooks carry mandatory dry-run-and-backup steps, and SSH access follows a documented short-lived-certificate workflow.
Structure is what makes LLM maintenance safe. Naming conventions, frontmatter schemas, and link rules aren't bureaucracy — they're the rails that keep hundreds of small unsupervised edits coherent. The stricter the schema, the more autonomy you can grant.
The index and the log are the real trick. A one-line-per-page index plus an append-only journal lets a context-limited agent navigate an arbitrarily large corpus. This pattern generalizes to any LLM-maintained knowledge base.
File-back turns tickets into infrastructure. The marginal cost of documenting a resolved issue dropped to roughly zero, so it happens every time. The wiki's growth rate now tracks the interesting-problem rate, not anyone's spare time.
Append-only beats rewrite-in-place for trust. When the agent gets something wrong (it does), the correction sits in the log next to the error. You can audit how the wiki knows what it knows.
Keep raw immutable. Every derived claim traces to a source file the agent cannot touch. When a page looks suspicious, you diff it against raw instead of arguing with a model.
The wiki is three months old. It has already root-caused P1 outages faster than scrollback archaeology ever did, and every investigation made the next one cheaper. That's the property human-maintained wikis promise and never deliver: it compounds.