| name | self-maintenance-pattern |
|---|---|
| description | Portable self-maintaining agent knowledge store -- file-based memory + throttled promote/prune/re-validate loop |
| type | reference |
| status | active |
| confidence | high |
| project | global |
| created | 2026-08-26 |
| tldr | File-based memory that cleans itself -- scheduler classifies, agent decides |
A file-based memory that cleans itself. An Obsidian-compatible markdown vault stores
everything an agent learns -- decisions, research, reusable patterns -- as plain files with
frontmatter and [[wikilinks]]. At session end a throttled scheduler scans the store and
classifies each entry: promote (reused >= 2), prune (never reused, idle > 90 days),
re-validate (hold-out elapsed). It reports candidates only, never deletes. The agent
then decides -- promoting proven lessons, flagging dead ones for human confirmation,
re-checking held-out lessons against their falsification condition. Promotion is gated:
no evidence, no promote. One authoritative spec governs the system.
The store and this spec are pure ASCII markdown. Conformance rules:
- Pure ASCII only -- no em-dash, arrow (U+2192), minus sign (U+2212), section sign, or
multiplication / less-than / greater-than glyphs. Write
--,->,-,sec,x,<=,>=instead. - Fenced code blocks always tagged with a language:
mermaidfor diagrams,jsonfor schemas,yamlfor frontmatter,markdownfor markdown examples,textfor non-diagram listings. - Mermaid uses
flowchart TD/flowchart LRandstateDiagram-v2; quote any label containing>,<,==, or(. No<br/>. - Fences balanced -- every open fence has a matching close; no unclosed block.
Plain files, no DB. Index-first traversal: _index.md -> [[wikilinks]] -> file.
| Path | Holds |
|---|---|
_index.md |
root map + keyword/description index -- single retrieval source |
projects/{p}/memory/ |
decisions + rationale |
projects/{p}/research/ |
findings + source URLs + timestamps |
patterns/ |
pattern registry -- reusable patterns + reuse counters |
references/ |
authoritative specs |
state/ |
runtime telemetry |
Entry = YAML frontmatter + body + [[wikilinks]]. Frontmatter carries provenance and the
promotion-weight label; the registry body carries reuse counters as bold markdown (not
frontmatter keys):
# frontmatter -- memory / decision files
name: my-pattern
description: one-line summary for recall # semantic match beats keyword match -- required
type: reference # file category: reference | research | memory | decision | ...
status: active # lifecycle: active | held_out | archived
label: rule # promotion weight: fact | inferred | rule | validated-override
falsification_condition: "the helper no longer exists in the codebase"
holdout_until: "2026-09-02" # ISO date, QUOTED -- scanner regex requires the quotes
relations: {}
sources: []
provenance: [] # receipts (URL / file:line) supporting this memory# registry entry -- one `### <pattern-name>` block per entry; counters live in the block body
### my-pattern
**success count:** 2
**last used:** 2026-08-26Content rules: in -- decisions, deadlines, work facts. out -- opinions, commentary on individuals, org speculation, personal info, credentials/tokens/PII. Self-audit after each write.
Establish the store as a flat-file, index-first knowledge base (Obsidian/Zettelkasten
convention): a root map-of-content file links every entry via [[wikilinks]], so traversal
is _index.md -> link -> file -- no database.
- One directory per concern --
memory/(decisions),research/(findings),patterns/(registry),references/(specs),state/(telemetry). Each holds one kind of record. - One authoritative spec -- a single reference file defines the frontmatter schema (fields,
labelvalues,holdout_untilformat). Every other entry conforms; no ad-hoc field drift. - Frontmatter on every entry --
name,description(one-line recall summary),type(category),status(lifecycle),label(promotion weight),falsification_condition,holdout_until,provenance. Reuse counters (**success count:**,**last used:**) live in the registry body, not frontmatter. - Index is the keyword source -- one-line
description+ search keywords live only in_index.md. Memory files carry no per-file keyword tags; retrieval matches the index, not per-file tags, so keywords cannot drift. - Plain files over a database -- git-diffable, no corruption mode, portable across hosts.
- Write discipline -- atomic write (tmp file + rename), single-writer lock (mkdir) for concurrency, append-only where possible. Never partial writes.
flowchart LR
S["Scheduler (session-end / idle)"] -->|scan| V[(Store)]
S -->|candidate lists| AG[Agent]
AG -->|mutate| V
| Role | Runs | Does |
|---|---|---|
| Scheduler | session-end / idle | scan, classify, report. never mutates. |
| Agent | next turn | read report, apply rules, mutate. |
flowchart TD
A[scan store] --> B{"success count >= 2"}
B -->|yes| C[promote candidates]
B -->|no| D{"count == 0 and idle > 90d"}
D -->|yes| E[prune candidates]
D -->|no| F{holdout_until passed}
F -->|yes| G[held-out candidates]
F -->|no| H[skip]
C --> R[report candidates]
E --> R
G --> R
Guards: 1h fixed cooldown; single-writer lock (mkdir, in the state dir); re-entry guard (in-progress + live pid); stale-run recovery (10 min); atomic write (tmp + rename); force env var.
The scheduler classifies only -- the full promotion gate (re-validation + evidence receipt) is the agent's decision, below.
flowchart TD
R[report] --> P[promote]
R --> PR[prune]
R --> HO[held-out]
P -->|gate passes| M1[promote or merge]
PR -->|human yes| M2[delete]
HO -->|re-check pass| M3[promote]
HO -->|re-check fail| M4[archive]
- Promote -- gate:
success count >= 2+ same task shape + stable prompt + no existing skill + held-out window passed + re-validated + evidence receipt present. else merge into existing. Held-out/re-validate applies to newly-held-out memory files; registry entries are already-validated lessons, so the count gate alone promotes them. - Prune -- ask human; delete on explicit yes only.
- Held-out -- re-check
falsification_condition; pass -> promote, fail -> archive.
stateDiagram-v2
[*] --> active
active --> held_out : new lesson
held_out --> active : re-validate pass
held_out --> archived : re-validate fail
archived --> active : new evidence
Labels (promotion weight): fact 1x strong, inferred hold out, rule 2x,
validated-override never auto-demote.
State file -- JSON in the store's state/ dir (self-maintain.json)
{
"state": "idle",
"lastCompleted": 0,
"lastTriggered": 0,
"cooldownMs": 3600000,
"pid": null,
"perTask": {}
}On run: set state: in-progress + lastTriggered + pid before scanning; on completion
set state: idle + lastCompleted + perTask.selfMaintain + pid: null.
Trigger -- two integration paths (pick by host capability; same throttle read on both):
-
Native lifecycle hook (push) -- preferred. Register the scheduler on the host's session-end lifecycle event. The host invokes it automatically at close-out; no polling. Gate-obeying: runs inside the session's own hook surface, so the host's own gates still fire.
-
Pseudo-cron emulation (poll) -- fallback when the host exposes no idle/session-end hook. Emulate a periodic trigger with whatever scheduler the host offers: a durable cron entry (survives restart, fires only while the host is idle) or an external OS/CI timer, on a fixed periodic schedule. Poll-based: no hook to push the event, so the timer checks "due?" on each tick.
Either path runs the same scheduler script under the same cooldown, so it fires at most once per interval regardless of trigger count.
Throttle -- fixed 1h cooldown (3600000 ms), not adaptive. The interval does not
self-adjust to observed activity; it is a constant. Overrides: SELF_MAINTAIN_COOLDOWN_MS
(interval) and SELF_MAINTAIN_FORCE=1 (skip cooldown).
# path 2 schedule -- periodic (e.g. daily); the cron expression is the host's choice
17 8 * * *
Scheduler pseudo
promote = [e for e in entries if e.success_count >= 2]
prune = [e for e in entries if e.success_count == 0 and age(e.last_used) > 90 days]
heldout = [e for e in entries if e.holdout_until and now >= e.holdout_until]
success_count / last_used above are the parsed values extracted from the registry
body's **success count:** / **last used:** lines. Two disjoint sources, not one
entries collection: the promote/prune scan reads only the registry; the held-out scan
walks only the memory sub-tree. Promote/prune come from the registry; held-out comes from
memory files.
Agent decision loop pseudo
for e in report.promote: verify receipt + held-out + criteria -> promote or merge
for e in report.prune: ask human -> delete on explicit yes only
for e in report.heldout: re-check falsification_condition -> promote or archive
Five steps: store on disk -> frontmatter each entry -> scan-and-classify script -> trigger (lifecycle hook or pseudo-cron) -> agent acts on report under its write gate. Portable -- plain files, filter-based scheduler, agent the only mutator.
Configuration -- the label enum, thresholds (cooldown, prune age, stale window, promote count), env-var names, and output prefix are cross-referenced in several places (scheduler diagram, pseudo, throttle, verify table). Changing a value in one place requires sweeping every occurrence; a stale value left elsewhere is a mismatch. After any change, verify every occurrence agrees (grep the value) before treating the install as complete.
End-to-end:
- Session ends -- host fires the session-end hook (or the pseudo-cron timer ticks).
- Scheduler runs once per cooldown (idempotent; no-op inside 1h). Scans the registry + project files, prints a candidate report to stdout. Never mutates the store.
- Next session, the agent reads the report and applies the decision rules: promote/merge, prune (delete only on explicit human yes), re-validate held-out.
- Agent mutations run under the host's write gate.
Preconditions (workflow assumptions an implementer must satisfy):
- Lifecycle-hook support -- the host must fire a session-end event (or expose a cron/poll trigger) and the scheduler must be registered on it.
- Script runtime -- filesystem access + an epoch clock (
Date.now). - State directory writable -- the script creates it, but the store must exist.
- Report is stdout-only, no report file -- the host must surface the hook's stdout into the next session's context (transcript), or the report is lost. There is no automatic hand-off to the next agent; acting on the report is the agent's job, not the scheduler's.
- Scheduler is advisory -- the agent is the only mutator; prune deletes only after human confirmation.
Systematic test -- run in order; each row states how + expected result.
| # | Check | How | Pass when |
|---|---|---|---|
| 1 | store present | list dirs | _index.md, projects/, patterns/, references/, state/ exist |
| 2 | frontmatter valid | parse every .md |
all carry name, description, type, label; zero parse errors |
| 3 | scheduler runs | run the scheduler script with SELF_MAINTAIN_FORCE=1 |
exit 0; prints SELF-MAINTAIN: prune candidates N | promote eligible N | held-out due N |
| 4 | classify correct | seed fixtures | promote (>=2), prune (0 + >90d), held-out (past) land in exactly one list each |
| 5 | throttle | run twice within cooldown | 2nd run silent (skips); forced run executes |
| 6 | re-entry guard | state in-progress + live pid + lastTriggered < 10m old |
run skips |
| 7 | stale recovery | in-progress + lastTriggered 10m old |
resets to idle, then runs |
| 8 | atomic write | crash mid-write | state file remains valid JSON |
| 9 | trigger registered | check session-end hook / durable cron | lastCompleted timestamp advances on next fire |
| 10 | decision loop | feed report to agent | promote needs receipt + held-out, prune asks human, held-out re-checks condition |
| 11 | round-trip | teach -> schedule -> scan -> decide | entry flips active <-> held_out <-> archived correctly |
| 12 | format (SFS) | grep non-ASCII + parse fences | zero non-ASCII; all fences tagged + balanced |
Fixtures (step 4): seed three entries -- a ### foo block with **success count:** 3 and a
### bar block with **success count:** 0 + **last used:** 100 days ago (both in the
registry), plus a memory file under projects/ carrying
holdout_until: "2026-08-01" (a real past ISO date -- the scanner's new Date() parse
skips invalid strings like "yesterday") -- then assert each appears in exactly one list and
nothing else.