A pattern for building personal knowledge bases using LLMs.
This is an idea file, it is designed to be copy pasted to your own LLM Agent (e.g. OpenAI Codex, Claude Code, OpenCode / Pi, or etc.). Its goal is to communicate the high level idea, but your agent will build out the specifics in collaboration with you.
Most people's experience with LLMs and documents looks like RAG: you upload a collection of files, the LLM retrieves relevant chunks at query time, and generates an answer. This works, but the LLM is rediscovering knowledge from scratch on every question. There's no accumulation. Ask a subtle question that requires synthesizing five documents, and the LLM has to find and piece together the relevant fragments every time. Nothing is built up. NotebookLM, ChatGPT file uploads, and most RAG systems work this way.
The idea here is different. Instead of just retrieving from raw documents at query time, the LLM incrementally builds and maintains a persistent wiki — a structured, interlinked collection of markdown files that sits between you and the raw sources. When you add a new source, the LLM doesn't just index it for later retrieval. It reads it, extracts the key information, and integrates it into the existing wiki — updating entity pages, revising topic summaries, noting where new data contradicts old claims, strengthening or challenging the evolving synthesis. The knowledge is compiled once and then kept current, not re-derived on every query.
This is the key difference: the wiki is a persistent, compounding artifact. The cross-references are already there. The contradictions have already been flagged. The synthesis already reflects everything you've read. The wiki keeps getting richer with every source you add and every question you ask.
You never (or rarely) write the wiki yourself — the LLM writes and maintains all of it. You're in charge of sourcing, exploration, and asking the right questions. The LLM does all the grunt work — the summarizing, cross-referencing, filing, and bookkeeping that makes a knowledge base actually useful over time. In practice, I have the LLM agent open on one side and Obsidian open on the other. The LLM makes edits based on our conversation, and I browse the results in real time — following links, checking the graph view, reading the updated pages. Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase.
This can apply to a lot of different contexts. A few examples:
- Personal: tracking your own goals, health, psychology, self-improvement — filing journal entries, articles, podcast notes, and building up a structured picture of yourself over time.
- Research: going deep on a topic over weeks or months — reading papers, articles, reports, and incrementally building a comprehensive wiki with an evolving thesis.
- Reading a book: filing each chapter as you go, building out pages for characters, themes, plot threads, and how they connect. By the end you have a rich companion wiki. Think of fan wikis like Tolkien Gateway — thousands of interlinked pages covering characters, places, events, languages, built by a community of volunteers over years. You could build something like that personally as you read, with the LLM doing all the cross-referencing and maintenance.
- Business/team: an internal wiki maintained by LLMs, fed by Slack threads, meeting transcripts, project documents, customer calls. Possibly with humans in the loop reviewing updates. The wiki stays current because the LLM does the maintenance that no one on the team wants to do.
- Competitive analysis, due diligence, trip planning, course notes, hobby deep-dives — anything where you're accumulating knowledge over time and want it organized rather than scattered.
There are three layers:
Raw sources — your curated collection of source documents. Articles, papers, images, data files. These are immutable — the LLM reads from them but never modifies them. This is your source of truth.
The wiki — a directory of LLM-generated markdown files. Summaries, entity pages, concept pages, comparisons, an overview, a synthesis. The LLM owns this layer entirely. It creates pages, updates them when new sources arrive, maintains cross-references, and keeps everything consistent. You read it; the LLM writes it.
The schema — a document (e.g. CLAUDE.md for Claude Code or AGENTS.md for Codex) that tells the LLM how the wiki is structured, what the conventions are, and what workflows to follow when ingesting sources, answering questions, or maintaining the wiki. This is the key configuration file — it's what makes the LLM a disciplined wiki maintainer rather than a generic chatbot. You and the LLM co-evolve this over time as you figure out what works for your domain.
Ingest. You drop a new source into the raw collection and tell the LLM to process it. An example flow: the LLM reads the source, discusses key takeaways with you, writes a summary page in the wiki, updates the index, updates relevant entity and concept pages across the wiki, and appends an entry to the log. A single source might touch 10-15 wiki pages. Personally I prefer to ingest sources one at a time and stay involved — I read the summaries, check the updates, and guide the LLM on what to emphasize. But you could also batch-ingest many sources at once with less supervision. It's up to you to develop the workflow that fits your style and document it in the schema for future sessions.
Query. You ask questions against the wiki. The LLM searches for relevant pages, reads them, and synthesizes an answer with citations. Answers can take different forms depending on the question — a markdown page, a comparison table, a slide deck (Marp), a chart (matplotlib), a canvas. The important insight: good answers can be filed back into the wiki as new pages. A comparison you asked for, an analysis, a connection you discovered — these are valuable and shouldn't disappear into chat history. This way your explorations compound in the knowledge base just like ingested sources do.
Lint. Periodically, ask the LLM to health-check the wiki. Look for: contradictions between pages, stale claims that newer sources have superseded, orphan pages with no inbound links, important concepts mentioned but lacking their own page, missing cross-references, data gaps that could be filled with a web search. The LLM is good at suggesting new questions to investigate and new sources to look for. This keeps the wiki healthy as it grows.
Two special files help the LLM (and you) navigate the wiki as it grows. They serve different purposes:
index.md is content-oriented. It's a catalog of everything in the wiki — each page listed with a link, a one-line summary, and optionally metadata like date or source count. Organized by category (entities, concepts, sources, etc.). The LLM updates it on every ingest. When answering a query, the LLM reads the index first to find relevant pages, then drills into them. This works surprisingly well at moderate scale (~100 sources, ~hundreds of pages) and avoids the need for embedding-based RAG infrastructure.
log.md is chronological. It's an append-only record of what happened and when — ingests, queries, lint passes. A useful tip: if each entry starts with a consistent prefix (e.g. ## [2026-04-02] ingest | Article Title), the log becomes parseable with simple unix tools — grep "^## \[" log.md | tail -5 gives you the last 5 entries. The log gives you a timeline of the wiki's evolution and helps the LLM understand what's been done recently.
At some point you may want to build small tools that help the LLM operate on the wiki more efficiently. A search engine over the wiki pages is the most obvious one — at small scale the index file is enough, but as the wiki grows you want proper search. qmd is a good option: it's a local search engine for markdown files with hybrid BM25/vector search and LLM re-ranking, all on-device. It has both a CLI (so the LLM can shell out to it) and an MCP server (so the LLM can use it as a native tool). You could also build something simpler yourself — the LLM can help you vibe-code a naive search script as the need arises.
- Obsidian Web Clipper is a browser extension that converts web articles to markdown. Very useful for quickly getting sources into your raw collection.
- Download images locally. In Obsidian Settings → Files and links, set "Attachment folder path" to a fixed directory (e.g.
raw/assets/). Then in Settings → Hotkeys, search for "Download" to find "Download attachments for current file" and bind it to a hotkey (e.g. Ctrl+Shift+D). After clipping an article, hit the hotkey and all images get downloaded to local disk. This is optional but useful — it lets the LLM view and reference images directly instead of relying on URLs that may break. Note that LLMs can't natively read markdown with inline images in one pass — the workaround is to have the LLM read the text first, then view some or all of the referenced images separately to gain additional context. It's a bit clunky but works well enough. - Obsidian's graph view is the best way to see the shape of your wiki — what's connected to what, which pages are hubs, which are orphans.
- Marp is a markdown-based slide deck format. Obsidian has a plugin for it. Useful for generating presentations directly from wiki content.
- Dataview is an Obsidian plugin that runs queries over page frontmatter. If your LLM adds YAML frontmatter to wiki pages (tags, dates, source counts), Dataview can generate dynamic tables and lists.
- The wiki is just a git repo of markdown files. You get version history, branching, and collaboration for free.
The tedious part of maintaining a knowledge base is not the reading or the thinking — it's the bookkeeping. Updating cross-references, keeping summaries current, noting when new data contradicts old claims, maintaining consistency across dozens of pages. Humans abandon wikis because the maintenance burden grows faster than the value. LLMs don't get bored, don't forget to update a cross-reference, and can touch 15 files in one pass. The wiki stays maintained because the cost of maintenance is near zero.
The human's job is to curate sources, direct the analysis, ask good questions, and think about what it all means. The LLM's job is everything else.
The idea is related in spirit to Vannevar Bush's Memex (1945) — a personal, curated knowledge store with associative trails between documents. Bush's vision was closer to this than to what the web became: private, actively curated, with the connections between documents as valuable as the documents themselves. The part he couldn't solve was who does the maintenance. The LLM handles that.
This document is intentionally abstract. It describes the idea, not a specific implementation. The exact directory structure, the schema conventions, the page formats, the tooling — all of that will depend on your domain, your preferences, and your LLM of choice. Everything mentioned above is optional and modular — pick what's useful, ignore what isn't. For example: your sources might be text-only, so you don't need image handling at all. Your wiki might be small enough that the index file is all you need, no search engine required. You might not care about slide decks and just want markdown pages. You might want a completely different set of output formats. The right way to use this is to share it with your LLM agent and work together to instantiate a version that fits your needs. The document's only job is to communicate the pattern. Your LLM can figure out the rest.






@madgodinc is right that conflict resolution belongs at write time, and the
reason is structural rather than stylistic.
A lint pass has to find the contradiction before it can resolve it, and
finding it is the expensive half: two pages that disagree are discoverable only
by reading both and noticing. That is quadratic in pages, it runs on every
sweep, and an LLM adjudicates each comparison. At write time the comparison is
already local & you hold the arriving assertion and the one it lands on. The
cost collapses to a lookup.
The price is that you must declare in advance which claims are mutually
exclusive. We do that with predicate groups: sets of relations that cannot
simultaneously hold between the same ordered pair. On company filings,
{generates_cash_flow, consumes_cash}and{grounds, certifies}are two ofeight. A later document asserting one against a pair that already holds the
other closes the earlier edge. No sweep, no adjudication, no second LLM call.
Two design rules make this work on real corpora.
Order by document, never by extracted dates. The intuitive design asks the
model when each fact was true and sorts on that. It does not survive contact
with source material: models state a period only when the text happens to say
so, which is rare. Document order is always available - publication order for a
book series, fiscal year for filings, ingest order for everything else - and it
is robust to backfilling. Index 2019 filings after 2024 ones and resolution is
still correct, because the declared order governs, not the write order.
@gowtham0992's use of git history to spot stale memories is the same instinct;
the repository already knows which assertion came later, and that ordering is
more reliable than anything a model will tell you about when something was
true.
Close, never delete. A superseded edge is marked and timestamped, and stays
retrievable on request. This is the gist's own "compounding artifact" argument
taken to its conclusion: a wiki that silently overwrites yesterday's claim is
worth less than one that can show you the succession.
It fires on ordinary prose, not just structured filings. On the d'Artagnan
trilogy - three novels, ~645k chars, indexed in publication order - thirteen
relationships were closed by a later book, including the
ally_ofedge betweend'Artagnan and Aramis once a later volume recasts them as opponents. The
extractor supplied no dates for any of those pairs.
The pattern needs a second clock. "What is true" and "what did this system
believe in March" are different questions with different answers, because a
source published today can restate a figure from 2019. Both need to be
recoverable, and the second is the one that matters when someone asks why a
decision was made. Model them as separate axes - valid time for when a fact
held in the world, belief time for when the system came to hold it - and
reproducing a past answer is a query parameter rather than an archaeology
expedition through old snapshots. Markdown makes this genuinely hard; it is the
strongest argument for structure I know.
What this is worth, on benchmarks built around exactly this problem.
LongMemEval is long-horizon chat memory, and two of its six categories are the
gist's problem directly: knowledge-update asks whether a fact can be
distinguished from the fact that replaced it, and temporal-reasoning asks for
ordering and intervals. Full benchmark, nothing sampled, against Zep/Graphiti's
published figures:
A full-context baseline that simply puts the entire history in the prompt
scores 60.2% - worth keeping in mind whenever someone suggests a long context
window makes this architecture unnecessary.
Read that as informative about the tier rather than a controlled head-to-head:
their figures are as published, judged by a single gpt-4o where ours is a
three-model panel, and two years separate the model generations in a direction
I cannot sign. One instance of 500 is excluded (a session both extraction
prompts refused) so the denominator is 499.
The mechanism behind the temporal categories is isolable. Carrying each
relation's validity period through to synthesis (not merely extracting and
storing it) is worth +38.4 points on temporal-reasoning and +16.3 on
knowledge-update, ablated paired on one fixed graph per instance so that
re-indexing noise cannot account for it. Storing time is not the hard part.
Rendering it into the prompt is.
ECT-QA is the adversarial version: earnings-call transcripts, six companies,
sixteen consecutive quarters each, where every metric is restated every quarter
so only the date distinguishes sixteen competing values. Under that benchmark's
own element-wise protocol it scores 0.807 Correct against 0.599 for TG-RAG and
0.406/0.405 for LightRAG and GraphRAG as published - on their corpus slice with
a judge we cannot run, so treat a few points of the margin as approximate.
The reason these are worth citing in this thread is not the ranking. It is that
the two categories which separate systems most are the two that turn on time,
and the gap between handling temporal evolution and ignoring it is measurable
at benchmark scale rather than being an architectural preference.
@ShootJackal's "model output as untrusted input" is the right frame, and it
has a stricter form on the write path than on the read path.
Verifying an answer against source files catches what the model said. The
earlier problem is what the extractor is allowed to put into the store in the
first place. Every extracted record passes deterministic gates before it can
become structure, and the gates are stated as rejections rather than repairs: a
record that cannot be made into a well-formed assertion is dropped, not guessed
at. No placeholder structure, no vague predicates, no pronominal or conjunctive
entity names, no bare quantities as entities, and no inverting a predicate to
express a denial.
The placeholder rule is the one I would argue hardest for. A synthesised
stand-in written so the pipeline can continue is indistinguishable from genuine
extracted structure once stored, so a transient provider outage mid-run
permanently poisons the store; and nothing downstream can tell. Failing the
chunk loudly is the only option that keeps the artifact trustworthy. A wiki that
compounds is a wiki whose early errors compound too.
Two operations in the gist get easier once the structure is a graph, and
both for the same reason: they become computed rather than adjudicated.
The index. @kriss-b's observation that the Statement of Applicability
naturally becomes the index is, I think, the general case; every wiki grows
one, and maintaining it by hand is where the maintenance burden the gist warns
about actually lands. A clustered graph generates it instead. We cluster
entities into communities, summarise each, and then build a topic tree above
them by recursive supergraph clustering: one node per cluster, edge weights
summed across the cut, re-cluster, repeat per level.
The recursion is the part that matters, and it is worth being precise about
why. A hierarchy built by sweeping a resolution parameter gives you levels but
guarantees nothing about containment - children can sit outside their parents.
Any drill-down built on such a tree is lying: you click a topic and get
material that is not in it, or miss material that is. Recursing over the
supergraph guarantees nesting by construction, because level n+1 is built
from level n's own nodes. Parent summaries are synthesised from child
summaries rather than from raw relations, which caps the added LLM cost at
roughly the cluster count per level rather than re-reading the corpus.
The practical effect is a wiki index at several zoom levels that nobody
maintains, and that cannot drift from the pages beneath it, because it is
derived from them.
Lint's other half. The gist's lint pass looks for contradictions and gaps.
Contradictions are handled at write time, above. Gaps are the harder question,
because a wiki cannot normally tell you what it does not cover - absence has no
page. Recording which entities and communities each query touched turns that
into arithmetic: which regions of the corpus have received least retrieval
attention, and which entities have never been retrieved at all. Not an LLM
judging completeness, just a count over what was actually read. Storing a hash
of the query rather than its text keeps the telemetry from becoming a second
sensitive dataset.
The pattern across all three (conflicts, index, gaps) is the same, and
@wy-cats found the same edge from a different direction in discovering that
"pending ingest" works better as computed state than as an explicit flag. A
flag has to be maintained and can be wrong; a computed value cannot drift from
what it describes. An LLM sweep is the expensive way to rediscover something
the write path, a clustering, or a counter already knew.
One thing worth instrumenting, because it is invisible to every metric this
architecture naturally produces.
Between the retriever and the model sits an assembler that fits retrieved
material into a token budget. Any assembler that fills greedily and stops at
the first oversized passage will silently contribute nothing from that channel
and a wiki page or a filing section is routinely larger than a per-channel
budget. The model then answers "no information available" while holding the
answer, unread, in material that was retrieved successfully.
Retrieval metrics cannot see this. They are computed over what was found, not
over what survived into the prompt; recall@k reads as perfect throughout. On
LongMemEval the gap between clipping oversized passages and dropping them is
eight points of end-to-end accuracy - larger than most differences between
retrieval strategies, and attributable entirely to the step after retrieval.
Log how many retrieved characters actually reach the model. One number, and it
makes the whole failure class visible.
On markdown versus structure: markdown is readable, diffable and hand-editable,
and a graph gives all three up. Which side wins depends on whether your wiki is
read mostly by people or mostly by systems.
Both layers are open source and run entirely on PostgreSQL with pgvector for
similarity, ordinary tables for the graph, one backup, one transaction spanning
your graph and your application rows. No additional services.
post-graph-rag=> the retrieval layer described here:https://github.com/crajah/post-graph-rag
post-graph=> the graph substrate alone, no LLM dependency:https://github.com/crajah/post-graph
Entity resolution across documents and structural negation so "not a
subsidiary of" is never retrieved as "subsidiary of" are the other two write-
path problems worth comparing notes on.