Skip to content

Instantly share code, notes, and snippets.

@mpalpha
Last active August 27, 2026 10:18
Show Gist options
  • Select an option

  • Save mpalpha/de4bb77f8a62d909c603ca0ef7028b4b to your computer and use it in GitHub Desktop.

Select an option

Save mpalpha/de4bb77f8a62d909c603ca0ef7028b4b to your computer and use it in GitHub Desktop.
Portable self-maintaining agent knowledge store — file-based memory + throttled promote/prune/re-validate loop spec
name self-maintenance-pattern
description Portable self-maintaining agent knowledge store -- file-based memory + throttled promote/prune/re-validate loop
type reference
status active
confidence high
project global
created 2026-08-26
tldr File-based memory that cleans itself -- scheduler classifies, agent decides

Self-Maintaining Agent Knowledge Store

A file-based memory that cleans itself. An Obsidian-compatible markdown vault stores everything an agent learns -- decisions, research, reusable patterns -- as plain files with frontmatter and [[wikilinks]]. At session end a throttled scheduler scans the store and classifies each entry: promote (reused >= 2), prune (never reused, idle > 90 days), re-validate (hold-out elapsed). It reports candidates only, never deletes. The agent then decides -- promoting proven lessons, flagging dead ones for human confirmation, re-checking held-out lessons against their falsification condition. Promotion is gated: no evidence, no promote. One authoritative spec governs the system.


Format (SFS conformance)

The store and this spec are pure ASCII markdown. Conformance rules:

  • Pure ASCII only -- no em-dash, arrow (U+2192), minus sign (U+2212), section sign, or multiplication / less-than / greater-than glyphs. Write --, ->, -, sec, x, <=, >= instead.
  • Fenced code blocks always tagged with a language: mermaid for diagrams, json for schemas, yaml for frontmatter, markdown for markdown examples, text for non-diagram listings.
  • Mermaid uses flowchart TD / flowchart LR and stateDiagram-v2; quote any label containing >, <, ==, or (. No <br/>.
  • Fences balanced -- every open fence has a matching close; no unclosed block.

Store

Plain files, no DB. Index-first traversal: _index.md -> [[wikilinks]] -> file.

Path Holds
_index.md root map + keyword/description index -- single retrieval source
projects/{p}/memory/ decisions + rationale
projects/{p}/research/ findings + source URLs + timestamps
patterns/ pattern registry -- reusable patterns + reuse counters
references/ authoritative specs
state/ runtime telemetry

Entry = YAML frontmatter + body + [[wikilinks]]. Frontmatter carries provenance and the promotion-weight label; the registry body carries reuse counters as bold markdown (not frontmatter keys):

# frontmatter -- memory / decision files
name: my-pattern
description: one-line summary for recall  # semantic match beats keyword match -- required
type: reference             # file category: reference | research | memory | decision | ...
status: active              # lifecycle: active | held_out | archived
label: rule                 # promotion weight: fact | inferred | rule | validated-override
falsification_condition: "the helper no longer exists in the codebase"
holdout_until: "2026-09-02"  # ISO date, QUOTED -- scanner regex requires the quotes
relations: {}
sources: []
provenance: []              # receipts (URL / file:line) supporting this memory
# registry entry -- one `### <pattern-name>` block per entry; counters live in the block body
### my-pattern
**success count:** 2
**last used:** 2026-08-26

Content rules: in -- decisions, deadlines, work facts. out -- opinions, commentary on individuals, org speculation, personal info, credentials/tokens/PII. Self-audit after each write.


Store setup

Establish the store as a flat-file, index-first knowledge base (Obsidian/Zettelkasten convention): a root map-of-content file links every entry via [[wikilinks]], so traversal is _index.md -> link -> file -- no database.

  • One directory per concern -- memory/ (decisions), research/ (findings), patterns/ (registry), references/ (specs), state/ (telemetry). Each holds one kind of record.
  • One authoritative spec -- a single reference file defines the frontmatter schema (fields, label values, holdout_until format). Every other entry conforms; no ad-hoc field drift.
  • Frontmatter on every entry -- name, description (one-line recall summary), type (category), status (lifecycle), label (promotion weight), falsification_condition, holdout_until, provenance. Reuse counters (**success count:**, **last used:**) live in the registry body, not frontmatter.
  • Index is the keyword source -- one-line description + search keywords live only in _index.md. Memory files carry no per-file keyword tags; retrieval matches the index, not per-file tags, so keywords cannot drift.
  • Plain files over a database -- git-diffable, no corruption mode, portable across hosts.
  • Write discipline -- atomic write (tmp file + rename), single-writer lock (mkdir) for concurrency, append-only where possible. Never partial writes.

Roles

flowchart LR
    S["Scheduler (session-end / idle)"] -->|scan| V[(Store)]
    S -->|candidate lists| AG[Agent]
    AG -->|mutate| V
Loading
Role Runs Does
Scheduler session-end / idle scan, classify, report. never mutates.
Agent next turn read report, apply rules, mutate.

Scheduler

flowchart TD
    A[scan store] --> B{"success count >= 2"}
    B -->|yes| C[promote candidates]
    B -->|no| D{"count == 0 and idle > 90d"}
    D -->|yes| E[prune candidates]
    D -->|no| F{holdout_until passed}
    F -->|yes| G[held-out candidates]
    F -->|no| H[skip]
    C --> R[report candidates]
    E --> R
    G --> R
Loading

Guards: 1h fixed cooldown; single-writer lock (mkdir, in the state dir); re-entry guard (in-progress + live pid); stale-run recovery (10 min); atomic write (tmp + rename); force env var.

The scheduler classifies only -- the full promotion gate (re-validation + evidence receipt) is the agent's decision, below.


Agent decision loop

flowchart TD
    R[report] --> P[promote]
    R --> PR[prune]
    R --> HO[held-out]
    P -->|gate passes| M1[promote or merge]
    PR -->|human yes| M2[delete]
    HO -->|re-check pass| M3[promote]
    HO -->|re-check fail| M4[archive]
Loading
  • Promote -- gate: success count >= 2 + same task shape + stable prompt + no existing skill + held-out window passed + re-validated + evidence receipt present. else merge into existing. Held-out/re-validate applies to newly-held-out memory files; registry entries are already-validated lessons, so the count gate alone promotes them.
  • Prune -- ask human; delete on explicit yes only.
  • Held-out -- re-check falsification_condition; pass -> promote, fail -> archive.

Lifecycle state machine

stateDiagram-v2
    [*] --> active
    active --> held_out : new lesson
    held_out --> active : re-validate pass
    held_out --> archived : re-validate fail
    archived --> active : new evidence
Loading

Labels (promotion weight): fact 1x strong, inferred hold out, rule 2x, validated-override never auto-demote.


Integration

State file -- JSON in the store's state/ dir (self-maintain.json)

{
  "state": "idle",
  "lastCompleted": 0,
  "lastTriggered": 0,
  "cooldownMs": 3600000,
  "pid": null,
  "perTask": {}
}

On run: set state: in-progress + lastTriggered + pid before scanning; on completion set state: idle + lastCompleted + perTask.selfMaintain + pid: null.

Trigger -- two integration paths (pick by host capability; same throttle read on both):

  1. Native lifecycle hook (push) -- preferred. Register the scheduler on the host's session-end lifecycle event. The host invokes it automatically at close-out; no polling. Gate-obeying: runs inside the session's own hook surface, so the host's own gates still fire.

  2. Pseudo-cron emulation (poll) -- fallback when the host exposes no idle/session-end hook. Emulate a periodic trigger with whatever scheduler the host offers: a durable cron entry (survives restart, fires only while the host is idle) or an external OS/CI timer, on a fixed periodic schedule. Poll-based: no hook to push the event, so the timer checks "due?" on each tick.

Either path runs the same scheduler script under the same cooldown, so it fires at most once per interval regardless of trigger count.

Throttle -- fixed 1h cooldown (3600000 ms), not adaptive. The interval does not self-adjust to observed activity; it is a constant. Overrides: SELF_MAINTAIN_COOLDOWN_MS (interval) and SELF_MAINTAIN_FORCE=1 (skip cooldown).

# path 2 schedule -- periodic (e.g. daily); the cron expression is the host's choice
17 8 * * *

Scheduler pseudo

promote = [e for e in entries if e.success_count >= 2]
prune   = [e for e in entries if e.success_count == 0 and age(e.last_used) > 90 days]
heldout = [e for e in entries if e.holdout_until and now >= e.holdout_until]

success_count / last_used above are the parsed values extracted from the registry body's **success count:** / **last used:** lines. Two disjoint sources, not one entries collection: the promote/prune scan reads only the registry; the held-out scan walks only the memory sub-tree. Promote/prune come from the registry; held-out comes from memory files.

Agent decision loop pseudo

for e in report.promote:  verify receipt + held-out + criteria -> promote or merge
for e in report.prune:    ask human -> delete on explicit yes only
for e in report.heldout:  re-check falsification_condition -> promote or archive

Five steps: store on disk -> frontmatter each entry -> scan-and-classify script -> trigger (lifecycle hook or pseudo-cron) -> agent acts on report under its write gate. Portable -- plain files, filter-based scheduler, agent the only mutator.

Configuration -- the label enum, thresholds (cooldown, prune age, stale window, promote count), env-var names, and output prefix are cross-referenced in several places (scheduler diagram, pseudo, throttle, verify table). Changing a value in one place requires sweeping every occurrence; a stale value left elsewhere is a mismatch. After any change, verify every occurrence agrees (grep the value) before treating the install as complete.


Usage (basic workflow)

End-to-end:

  1. Session ends -- host fires the session-end hook (or the pseudo-cron timer ticks).
  2. Scheduler runs once per cooldown (idempotent; no-op inside 1h). Scans the registry + project files, prints a candidate report to stdout. Never mutates the store.
  3. Next session, the agent reads the report and applies the decision rules: promote/merge, prune (delete only on explicit human yes), re-validate held-out.
  4. Agent mutations run under the host's write gate.

Preconditions (workflow assumptions an implementer must satisfy):

  • Lifecycle-hook support -- the host must fire a session-end event (or expose a cron/poll trigger) and the scheduler must be registered on it.
  • Script runtime -- filesystem access + an epoch clock (Date.now).
  • State directory writable -- the script creates it, but the store must exist.
  • Report is stdout-only, no report file -- the host must surface the hook's stdout into the next session's context (transcript), or the report is lost. There is no automatic hand-off to the next agent; acting on the report is the agent's job, not the scheduler's.
  • Scheduler is advisory -- the agent is the only mutator; prune deletes only after human confirmation.

Verify install & setup

Systematic test -- run in order; each row states how + expected result.

# Check How Pass when
1 store present list dirs _index.md, projects/, patterns/, references/, state/ exist
2 frontmatter valid parse every .md all carry name, description, type, label; zero parse errors
3 scheduler runs run the scheduler script with SELF_MAINTAIN_FORCE=1 exit 0; prints SELF-MAINTAIN: prune candidates N | promote eligible N | held-out due N
4 classify correct seed fixtures promote (>=2), prune (0 + >90d), held-out (past) land in exactly one list each
5 throttle run twice within cooldown 2nd run silent (skips); forced run executes
6 re-entry guard state in-progress + live pid + lastTriggered < 10m old run skips
7 stale recovery in-progress + lastTriggered 10m old resets to idle, then runs
8 atomic write crash mid-write state file remains valid JSON
9 trigger registered check session-end hook / durable cron lastCompleted timestamp advances on next fire
10 decision loop feed report to agent promote needs receipt + held-out, prune asks human, held-out re-checks condition
11 round-trip teach -> schedule -> scan -> decide entry flips active <-> held_out <-> archived correctly
12 format (SFS) grep non-ASCII + parse fences zero non-ASCII; all fences tagged + balanced

Fixtures (step 4): seed three entries -- a ### foo block with **success count:** 3 and a ### bar block with **success count:** 0 + **last used:** 100 days ago (both in the registry), plus a memory file under projects/ carrying holdout_until: "2026-08-01" (a real past ISO date -- the scanner's new Date() parse skips invalid strings like "yesterday") -- then assert each appears in exactly one list and nothing else.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment