You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Version 1.3 (2026-08-21). A COMPACT, markdown-compatible package: the normative
specs + the adaptation playbook that directs an agent to RESEARCH, ANALYZE,
and ADAPT the target machine's current workflow so it functions with the
features of the spec. Pure ASCII, PII-free, Mermaid diagrams render in
CommonMark markdown. No user-identifying paths or credentials anywhere in this package.
What this package is
Portability means: research -> analyze -> adapt the destination's CURRENT
workflow so it functions with the spec's features (the capability contracts
C-1..C-9 and their observable invariants), using whatever tooling ALREADY
exists and creating a working alternative where needed. Nothing is installed
unless no suitable realization exists, and then the report says what to install.
This package is the SPEC + the ADAPTATION INSTRUCTIONS - compact, markdown-only,
markdown-compatible. It is NOT a code bundle.
Research mandate (always)
The agent performing analysis or adaptation ALWAYS uses proper, systematic,
OPTIMAL, correct, DATA-BACKED research at every step:
LOCAL: observe the destination directly (probe, files, config, logs,
installed toolset behavior) - never assume.
ONLINE: data-backed sources only (official docs, peer-reviewed/arXiv papers,
benchmarks with measured effect sizes), scoped + cited.
Evidence saved (vault Research/ or the deploy evidence store) and cited.
Never anecdote, guess, or a suboptimal shortcut. (host-deployment-spec-v1.1
section 6 Systematic analysis & research discipline; principle 13.)
Contents (operative set, versioned from 1.0)
File
Role
Status
workflow-spec-v1.1.md
The one operative workflow contract (standalone distribution spec derived from the v1.8-host realization of canonical v1.7; parity-marked SPECIFIED-NOT-IMPLEMENTED; Mermaid flows; includes the Decision-making process section sec 5)
OPERATIVE (v1.1)
host-deployment-spec-v1.1.md
The one deployment & adaptation playbook (analysis-first -> plan -> approval -> adapt; capability contracts C-1..C-9; toolset-neutral; 3 Mermaid flows)
OPERATIVE (v1.1)
workflow-plan-spec-v1.0.md
The one plan-artifact contract + PL-001..PL-010 conformance vectors (the plan gate's decision rule)
CANONICAL (v1.0)
spec-format-specification.md
The one format spec (SFS v1.1) - format/validation rules every spec MUST follow
ACTIVE
README.md
This manifest
-
Versions start at 1.0 per document; a revision bumps the stamp (e.g.
host-deployment 1.0 -> 1.1). Superseded versions are retained in full in the
source vault, never re-shipped here.
Read order
workflow-spec-v1.1.md - the one operative workflow contract (WHAT + its
host realization, with honest SPECIFIED-NOT-IMPLEMENTED markers; read
sec 5 for how the workflow directs decisions).
workflow-plan-spec-v1.0.md - the plan artifact + plan-gate rules.
host-deployment-spec-v1.1.md - HOW to research, analyze, and adapt a target
machine's current workflow to the spec.
spec-format-specification.md - the format contract for any spec document.
Adaptation flow (host-deployment-spec-v1.1)
ANALYSIS (free, data-backed research): assess the destination's current
workflow and toolset (probe or manual DP checks) -> Analysis Report.
PLAN: create the deploy plan via the workflow's own plan gate.
APPROVAL: the user approves before anything mutates the target.
ADAPT: realize the capability contracts (C-1..C-9) by adapting what exists
or creating a working alternative (researched, data-backed), verified
behaviorally on the live target; install nothing else.
VERIFY: behavioral conformance vectors on the live target.
BLOCKED case: MISSING REPORT says what needs to be installed (primitives +
example toolsets, never a forced brand).
Feasibility is decided by the capability contracts (C-1..C-9 in
host-deployment-spec-v1.1.md), not by toolset names.
Reference implementation (pointed to, NOT bundled)
The opencode reference realization of the capability contracts (engine cores,
portability probe, conformance + tests, opencode bridge plugins, tooling) lives
in the source workspace's .opencode/ directory. It is an EXAMPLE of an
adapted realization for the agent to run where compatible or adapt from. It is
intentionally NOT bundled here, keeping this package compact and markdown-compatible.
Portability invariants (enforced)
WHAT-first: contracts are observable invariants; toolsets and the reference
realization are examples only.
Research-gated: every analysis/adaptation is data-backed (local + online),
scoped, cited, saved - never guessed.
Toolset-neutral: the verdict never depends on opencode presence.
PII-free: no usernames, credentials, or user-embedding paths (checked).
ASCII: every document is pure ASCII (SFS).
Mermaid: diagrams are Mermaid (renders in CommonMark markdown).
Markdown-compatible: valid CommonMark, no BOM/CR.
Provenance note
Research/... and .opencode/... references in the spec documents are
source-vault / reference-implementation provenance (authoring evidence),
intentionally NOT shipped in this package; they resolve in the source vault,
not here. In-package references are package-relative and resolve.
CANONICAL (current). Supersedes v1.0 (2026-08-20). v1.1 removes the
opencode-centric install model. The workflow must run on the destination's
EXISTING, adaptable AI toolset; NOTHING is installed unless NO suitable
tool exists - and then the report says what needs to be installed. The goal
is that the governance workflow's observable invariants hold on the target
machine, realized through whatever primitives the already-present toolset
provides. The deploy itself follows the workflow's discipline: ANALYSIS first
(free probe + report to the user), then PLANNING via the workflow's own plan
gate, then IMPLEMENTATION only after user approval. Grounding:
Research/v1.8-host-parity-plan.md, Research/client-side-feasibility.md.
Parent contract: workflow-spec-v1.1.md (the governance this deploys; the
Decision-making process section is sec 5).
Format: Research/spec-format-specification.md (SFS).
1. Design Principles
#
Principle
Source
1
Feasibility is STATE-DECIDABLE - a deterministic probe decides, never an LLM's judgment.
Fail-closed reporting: no probe output = do not implement; blocked = MISSING REPORT, never a partial install.
workflow-spec-v1.1.md section 7 (fail-closed)
3
Self-contained playbook: a FRESH agent with no prior knowledge of this system must be able to run it.
SFS section 1 (self-contained)
4
Contracts are WHAT (observable invariants); mechanisms are HOW (per-toolset primitives) and interchangeable. The deploy payload is the portable engine, not a toolset-specific stack.
session analysis 2026-08-20; verified: all *-core.ts import only node: builtins
5
USE what the destination already has. If an existing tool can express the required contracts with its own primitives, deploy onto it; install NOTHING.
user directive 2026-08-20
6
Install only when NO suitable alternative exists - and then the MISSING REPORT says what needs to be installed (primitives + example toolsets), never a forced brand.
user directive 2026-08-20
7
Verify by execution after install: the behavioral conformance vectors pass on the target, never a self-report.
Verify-before-asserting (vault Rules)
8
The source vault is never mutated; the source is copied, never moved.
G14 migration safety
9
Complete bring-up: the portable engine AND the vault content are both deliverables.
user decision 2026-08-20
10
The MISSING REPORT uses the mandated escalation format: WHAT IS NEEDED -> WHY -> WHAT IT BLOCKS -> HOW TO OBTAIN IT.
workflow-spec-v1.1.md section 11
11
ANALYSIS FIRST, then plan, then implement-if-approved: investigation (probe + report) is FREE; the deploy plan follows the workflow's OWN plan gate; acting tools that mutate the target commit only after user approval.
Toolsets are EXAMPLES, never requirements: the normative target is the portable capability contracts (WHAT); specific installed/configured elements appear only as illustrations.
user directive 2026-08-20
13
The agent ALWAYS uses proper, systematic, OPTIMAL, correct, data-backed research at every analysis/research step: LOCAL (observed from the destination) and/or ONLINE (data-backed sources, scoped + cited), evidence saved - never anecdote, guess, or suboptimal shortcut.
user directive 2026-08-20; Rules/Self-maintenance.md (research-gated)
14
The reference realization is an EXAMPLE, not a required payload: the portable engine, probe, and conformance suite are the opencode reference implementation. The deploy realizes the capability contracts (C-1..C-6) with whatever the destination can ACTUALLY run. The installing agent MUST adapt the reference realization or CREATE A WORKING ALTERNATIVE that satisfies the intended requirements when the reference is not runnable/compatible; only NO suitable realization blocks.
user directive 2026-08-20; consistent with principle 12 (toolsets are examples)
Sources: inline citations above.
2. Architecture - the deploy loop
flowchart TB
A1["1. ANALYSIS - portability-probe.ts (reference tool; where it cannot run, perform the DP checks manually per sec 23), DP-001..N, NON-mutating, detects installed toolsets (DP-016) + workspace scan (DP-019); investigation is FREE"]
A1 --> A2["2. ANALYSIS REPORT - verdict, vectors, detected toolset, recommended path (research-backed)"]
A2 --> P["3. PLAN - deploy plan via the workflow's own plan gate (valid todowrite, forward contract)"]
P --> AP["4. APPROVAL - user approves the plan; no mutation before approval"]
AP --> B["5. BRIDGE (partial) - thin bridge for the DETECTED toolset; never install a new tool"]
B --> D["6. DEPLOY - realize capability contracts (reference engine only if runnable) + seed vault + wire bridge + apply"]
D --> V["7. VERIFY - behavioral conformance vectors on the live target + idle pass"]
V --> R["8. REPORT - completion report (verdict, vectors, adaptations, receipts)"]
A2 -. blocked .-> M["MISSING REPORT - what needs installing; pause until user obtains + re-probes"]
AP -. denied .-> STOP["Stop and report"]
D -. failure .-> RB["Rollback + notify + remediation plan (N-5)"]
V -. failure .-> RB2["Notify + remediation plan (N-6); no completion report"]
Loading
The probe is the trust root: the agent asserts nothing about the target that the
probe did not observe (principle 1). Nothing is installed on the target except
the portable engine + a bridge for a toolset that was ALREADY there (principles
5-6). The workflow's own discipline applies to the deploy itself: analysis is
free, planning is gated, implementation commits only after approval (principle 11).
Every phase transition persists its plan/research/implementation state to
<target>/.governance/ before the next phase runs (deploy tracking &
persistence, sec 8).
5. Prose->spec translation pipeline - NOT APPLICABLE
This spec IS the deployment contract. Deploy steps are executed as plan
artifacts (todowrites) created in the Planning phase, not translated through the
prose->spec pipeline (which is reference-only per v1.8-host sec 5).
Analysis phase is FREE: probing and read-only inspection of the target are
investigation (enforcement is selective - commitment actions gated, investigation
free; v1.8-host principle 3). No plan and no approval are required to probe and
report.
Approval gate (commitment): acting tools that MUTATE the target are DENIED
until ALL of: (a) a valid probe JSON exists, (b) a validated deploy plan exists
(plan gate per workflow-plan-spec-v1.0.md, PL-001..010), AND (c) the user
approved the plan. Fail-closed on any missing precondition. Re-arms per turn.
Self-commitment pre-engine (chicken-and-egg): the plan gate and approval
gate are the workflow's OWN gates. On a destination where the engine is not yet
live, the deploying agent SELF-COMMITS to them (prose discipline) until the
engine is deployed and enforcing; afterwards they become deterministic. This is
stated, not hidden - a fresh destination cannot be governed by machinery that
is being installed.
Existing-tool gate: the deploy MUST target a toolset that was ALREADY
present on the destination (probe DP-016). If the detected toolset is not
opencode, a bridge is written for it; a new tool is installed ONLY when the
MISSING REPORT path fires (no suitable existing tool). No toolset is ever
installed to satisfy a bridge that could have been written.
Realization directive: when the reference realization cannot run on the
destination (DP-002/DP-003 fail, or runtime incompatibility), the installing
agent MUST adapt it or CREATE A WORKING ALTERNATIVE that satisfies the intended
requirements - the capability contracts C-1..C-9, sec 11. A working
alternative is one whose contracts pass behavioral verification on the LIVE
target (minimal primitives present + negative controls); it is NEVER a blind
copy of the reference code. The adaptation decision is RESEARCHED with the
systematic, optimal, correct, data-backed discipline (LOCAL observation of the
destination + ONLINE data-backed sources, scoped + cited), recorded, and
reported - never guessed and never silently downgraded.
Deploy gate: steps are ordered (engine -> npm install -> seed -> bridge ->
migrate) and each step's receipt AND state update must be persisted to disk
before the next step runs (deploy tracking & persistence, sec 8). No step runs
against the source vault (copies only).
MISSING REPORT gate (fail-closed): if the analysis verdict == blocked, the
ONLY deliverable is the MISSING REPORT (schema in sec 21) delivered at the
feasibility review. The report names the missing capabilities and what needs to
be INSTALLED to obtain them (with example toolsets, never a forced brand). The
deploy pauses until the user obtains what is missing and re-probes.
Verify-before-report: the completion report is not written until the
behavioral conformance vectors pass on the LIVE destination toolset and the
idle observation pass ran. A report without fresh execution evidence is denied
(completion evidence contract, v1.8-host sec 6).
User notification + remediation proposal (issue contract): EVERY issue
notification (triggers N-1..N-7, sec 21) is delivered WITH a proposed
remediation plan. On an issue, the agent FIRST researches it systematically
(investigation is free), THEN drafts a remediation plan following the workflow's
OWN plan gate (workflow-plan-spec-v1.0.md: valid todowrite, forward contract,
plan-spec PL-001..010), and presents ISSUE + PLAN together for user approval.
NO remediation action runs before approval. On a pre-engine destination this
discipline is self-committed (prose) until the engine is live - the approval is
still the real gate.
Systematic analysis & research discipline (EVERY step, ALWAYS): the agent
performs proper, systematic, OPTIMAL, correct, data-backed research at EVERY
step of this playbook - Analysis phase, capability mapping for toolsets outside
the detection set, adaptation decisions (adapt or create a working alternative),
remediation planning (N-1..N-7), verification interpretation - following one
discipline:
LOCAL research where the question touches the destination: the probe report,
the destination's files/config/logs/state, and the installed toolset's actual
behavior - observed, never assumed.
ONLINE research where external knowledge is needed: proper, thought-out web
research from DATA-BACKED sources only (official toolset documentation,
peer-reviewed/arXiv papers, benchmarks with measured effect sizes), every
claim scoped and cited, per the vault's research-gated discipline
(Rules/Self-maintenance.md; SFS section 7 Research-source conventions).
ONLINE is MANDATORY, not optional, whenever external knowledge is needed or
local evidence is insufficient - never answer an external-knowledge question
from local inference alone. Use >= 2 independent sources, check authority +
recency (superseding docs beat the most-relevant one), and grade retrieved
evidence (relevant/authoritative/fresh) before use.
The mode is chosen per step: local for destination facts, online for external
knowledge, BOTH where a recommendation depends on both (e.g. a remediation
plan).
EVERY research step emits a RESEARCH RECEIPT (schema in
sec 21) appended to the evidence chain BEFORE the next
step runs. A claim in any report or plan is grounded in its research
receipt - traceable local observations + linked, scoped online sources -
never an unreceipted assertion.
Findings are SAVED as evidence (vault Research/ or the deploy evidence
store) and cited in the outputs, so every conclusion is traceable to observed
local facts and/or sourced online data - never anecdote or guesswork.
General gates (B1, referenced - never duplicated): the general gates
G1/G2/G3 of workflow-spec-v1.1.md section 7 apply to the deploy's own
artifact mutations - context-gain before, verify-by-execution after,
research-mode recorded. This playbook references them, never re-defines them.
The SOURCE vault is never mutated: cp -R (robocopy on Windows) semantics,
never move. G14 already proves no-lesson-loss on regeneration.
The TARGET vault is seeded with: Rules/, Lessons/, Research/,
Templates/, Index.md. Missing source content is a DP-013 failure (see
sec 23) reported in the MISSING REPORT, not silently skipped.
The target's AGENTS.md is produced by the idle regen (vault-core
regenAgents), NOT copied from the source.
Provenance: the target vault records its origin (frontmatter or a
Research/DEPLOYMENT.md note with source commit/reference + deploy date).
The vault logic is part of the portable engine (runs on ANY toolset as a
standalone process, idle/cron triggered).
Sources: none - declarative scope statement.
8. Action receipts contract (deploy receipts)
The probe report IS the feasibility receipt (JSON, sec 21 schema, written
to <target>/.governance/portability-report.json and stdout).
Each deploy step emits a receipt (tool + args hash + result hash); until the
engine is live the agent's own tool receipts on the target serve as the step
receipts.
The completion/MISSING report is a receipted-evidence record: its claims must
match the probe JSON and verification outputs (tri-state verdict logic).
Deploy tracking & persistence (plan, research, implementation): the install
plan, the research, and the implementation are all persisted to disk at each
relevant step, so the deploy is resumable, auditable, and survives session
boundaries. The deploy state lives under <target>/.governance/:
PLAN: the validated deploy plan (todowrite snapshot) is written to
<target>/.governance/deploy-plan.json at the Planning phase, before
approval; the approved plan is the committed artifact (not only in-context).
RESEARCH: every research step (systematic research discipline,
sec 6) saves its findings - local observations + online sources,
scoped and cited - to the evidence store (<target>/.governance/evidence/<step>.json
or the vault Research/) AT THAT STEP, not batched at the end. Each is a
RESEARCH RECEIPT (schema in sec 21), hash-chained with the
deploy receipts (prev_hash linked), so the research is independently
verifiable - a claim's grounding can be re-proven from the chain.
IMPLEMENTATION: each deploy step's receipt (tool, args hash, result hash,
ts) is appended to <target>/.governance/deploy-receipts.jsonl when the step
completes.
STATE:<target>/.governance/deploy-state.json is rewritten at every
phase transition (unprobed -> analyzed -> planned -> approved -> bridging / deploying -> verifying -> done | reported, sec 21) with: phase,
plan-snapshot ref, latest receipt seq, research-evidence refs, verdict, and a
resume hint. On a re-run, the agent reads this file and continues from the
recorded phase (the plan gate still applies).
Nothing that informed a decision lives only in the agent's context: every
input to a report is on disk before the report is delivered.
Sources: none - declarative scope statement.
9. Monitoring contract
Post-deploy: run the behavioral conformance vectors against the LIVE
destination toolset (WS/PL vectors; negative controls included).
PARITY (v1.1): the ONLY shipped driver for these vectors is the opencode
plugin harness (conformance.ts -> the opencode plugin test files). A
non-opencode bridge has NO shipped verification driver; until one is written,
a non-opencode deploy is verified by the deterministic engine tests plus
manual behavior checks, and the conformance claim is restricted to the
opencode path (never silently extended to a bridge without a driver).
Run verify-e2e.ts on the target (flap=0 over a scratch-copy pass) when the
target DB is available (DP-010).
Record the baseline in the target's health.json BEFORE first idle pass so
the before/after comparison is real (v1.8-host sec 9).
Sources: none - declarative scope statement.
10. Communication conventions (the reports)
Three report types, all following the escalation format
(WHAT IS NEEDED -> WHY -> WHAT IT BLOCKS -> HOW TO OBTAIN IT). Any report that
carries an ISSUE (triggers N-1..N-7, sec 21) MUST pair the issue with a
proposed remediation plan (based on LOCAL observation + ONLINE data-backed
research, spec-conformant, awaiting approval): the user is never notified of a
problem without the solution path to approve. The plan cites its sources
(probe/local evidence refs + linked online refs, sec 6).
ANALYSIS REPORT (first deliverable, free): delivered to the user BEFORE
any planning or implementation. Contains: the probe JSON, the feasibility
verdict, the detected toolset(s) + selected tool, and the recommended path.
If blocked, it is the MISSING REPORT (below). It is investigation output,
so it carries no commitment and requires no approval to produce.
MISSING REPORT (blocked case): the ONLY deliverable when no suitable
existing toolset is present. For each failed REQUIRED vector: what is
missing, why it matters, what it blocks, HOW TO OBTAIN IT - including WHAT
NEEDS TO BE INSTALLED (required primitives + example toolsets, never a forced
brand). PLUS the proposed remediation plan (steps to obtain what is missing
and re-probe, per workflow-plan-spec-v1.0.md), presented for approval. PLUS
degraded OPTIONAL capabilities the user should know about. The deploy pauses
until the user obtains what is missing and re-probes.
Completion report: verdict, probe vectors (id/name/actual/pass), the
detected toolset + bridge used, every adaptation taken, verification receipts
(behavioral vectors + idle), and the next step (await user review).
Sources: none - declarative scope statement.
11. Adaptation profile + capability contracts
The deploy USES the destination's already-installed toolset. The NORMATIVE
target is the portable CAPABILITY CONTRACTS below (WHAT). Specific toolsets
appear ONLY as EXAMPLES of tools that typically provide a contract - a tool is
never the requirement (principle 12). The probe embeds the same capability data
in machine form, so feasibility stays deterministic (principle 1).
REQUIRED capability contracts (must be enforceable; their absence blocks):
#
Capability contract (WHAT)
Minimal primitive the tool must expose
C-1
Plan gate
a PRE-TOOL hook that can DENY an acting tool with a structured reason
C-2
Commitment gate
an environment-external PERMISSION/APPROVAL boundary for irreversible actions
C-3
Receipts
access to tool-call args + results to emit a signed, hash-chained record
C-4
Completion evidence
a SESSION-END signal to finalize a tri-state evidence verdict
C-5
Knowledge lifecycle
the portable logic (promote/regen) + AN evidence feed supplied by the bridge
C-6
Vault seeding
filesystem copy + regen (standalone)
OPTIONAL capability contracts (degradable without blocking):
#
Capability contract
Minimal primitive
Degrades to
C-7
Constraint pinning
system-prompt/context control under compaction
pinned rules unenforced under compaction
C-8
Directive priority (PIM/SOM)
input hook + output review path
SOM-only
C-9
STD/STB re-injection
per-call hook for turn/token counting
re-inject on tool boundary
Examples - NOT normative (known toolsets that TYPICALLY provide a contract):
Capability
opencode
claude-code
cursor
copilot-cli
C-1
tool.execute.before
PreToolUse
none
none
C-2
permission.ask
permissions.json
approval-mode (partial)
none
C-3
tool.execute.after
PostToolUse
none
none
C-4
event idle
Stop
none
none
C-5
opencode DB feed (shipped)
bridge-supplied
bridge-supplied
bridge-supplied
C-6
standalone
standalone
standalone
standalone
C-7
experimental.session.compacting
compaction hooks (partial)
none
none
C-8
chat.message + text.complete
UserPromptSubmit (partial)
none
none
C-9
chat.params
UserPromptSubmit counting (partial)
none
none
The rows above are ILLUSTRATIVE: the CONTRACT is the requirement. A tool that
provides the primitive by any mechanism (or via a bridge the deploy writes)
satisfies it. A toolset NOT in the detection set is not auto-evaluated by the
probe; the agent must map its primitives to C-1..C-9 manually in the Analysis
report.
PARITY (v1.1) - lifecycle evidence feed: the lifecycle LOGIC
(promote/regen, vault-core.ts) is portable, but the shipped EVIDENCE FEED
reads the opencode DB (self-maintain.ts: ~/.local/share/opencode/opencode.db).
On a non-opencode destination the bridge MUST provide an evidence source (C-5),
or the lifecycle degrades (promotion/revival run on empty evidence - recorded
as an adaptation). C-5 means the logic AND a feed, not the logic alone.
Feasibility decision (deterministic, computed from the capability data)
A toolset is SUITABLE iff it provides all six REQUIRED capability contracts
(C-1..C-6), natively or via a bridge the deploy writes.
implementable = a suitable EXISTING toolset with a SHIPPED bridge
(examples: opencode) and no OPTIONAL degradation; conformance is claimed for
this path only (the WS/PL driver is opencode-bound - see sec 9).
partial = a suitable EXISTING toolset whose bridge must be written
(examples: claude-code), and/or an OPTIONAL contract degrades. Adaptations
recorded. NO new tool is installed in this path. Conformance for the bridge
is NOT claimed until a verification driver for it is written.
blocked = no suitable EXISTING toolset, or another REQUIRED vector fails.
MISSING REPORT: WHAT NEEDS TO BE INSTALLED - e.g. "a tool whose native
primitives provide C-1 (pre-tool deny), C-2 (permission boundary), C-3
(tool-result stream), C-4 (session-end signal); opencode and Claude Code are
examples."
Model differs from the source host -> recompute STD/STB + per-model rejection
rates by calibration before safety-critical paths are trusted (v1.8-host
sec 11); DP-015 is informational.
Windows targets: session.idle never fires; the bridge must verify the
re-arm uses session.status + {type:"idle"} (DP-001 + bridge wiring).
Node < 22.6 blocks --experimental-strip-types: not adaptable (REQUIRED).
Reference realization is an example (principle 14): the reference
implementation (Node >= 22.6 + npm + the portable engine + the conformance
harness) is the opencode EXAMPLE, not a requirement. If the destination
cannot run it (DP-002/DP-003 fail), the deploy ADAPTS the reference or
CREATES A WORKING ALTERNATIVE that satisfies C-1..C-9 in any compatible
runtime, researched with the systematic, optimal, correct, data-backed
discipline, and records the adaptation. A working alternative is accepted
ONLY on behavioral verification of its contracts on the live target - the
contracts, not a code copy, are the requirement.
Decision map (D7) - deploy feasibility:
flowchart TD
S{suitable existing toolset?} -- no --> B[blocked - MISSING REPORT: what to install]
S -- yes --> BS{bridge shipped?}
BS -- yes --> IM[implementable]
BS -- no --> PA[partial - write bridge + record adaptation]
Loading
PARITY (v1.1) - declared vs observed capability: suitability is DECLARED
(examples table), not observed. The actual hook surface must be validated on
the destination: opencode via sanity-load.mts post-deploy; for other
toolsets NO capability probe is shipped (DP-018), so their suitability rests
on the examples table until a probe exists. A bridge whose capabilities cannot
be observed is recorded as an adaptation, never silently claimed.
Cross-reference (v1.1): the deploy decision stack (feasibility D7 +
notification N-1..N-7, sec 12/sec 21) is the deploy-specific slice of the
workflow's decision-making process - see workflow-spec-v1.1.mdsec 5
(the cross-cutting view: who decides, the gated decision loop). Feasibility is
decided by the deterministic probe (D7 above), never by the agent's
description; the workflow's gate sequence (plan -> G1 context-gain -> G2
verify-by-execution) applies to every deploy mutation.
Sources: none - declarative scope statement.
12. Directive-priority enforcement
The user's task directive (set up the workflow) has highest priority; this
playbook is the mechanism, not an override.
Never mutate the source vault to satisfy a target requirement; propose the
change via the review queue instead.
Hard constraints (never touch the source vault, never bypass the probe, never
install a toolset when a bridge would do) are in the gates, not prompt rules.
Sources: none - declarative scope statement.
13. Drift mitigations summary
Drift type
Mitigation
Target/OS drift (Windows event gap)
DP-001 + bridge wiring check
Toolset version drift (Node/toolset)
DP-002/DP-004/DP-016 version bounds
Toolset-primitive drift (hook renamed/dropped)
behavioral vectors verify LIVE behavior, not file presence
Deterministic feasibility gate before any mutation
Tool-agnostic coverage
Capability matrix decides feasibility per toolset, no brand bias
Time-to-first-blocked-report
Probe runs in seconds; no manual checklist
Issue-to-solution latency
Every issue notification ships with a ready-for-approval remediation plan (N-1..N-7)
Verification honesty
Behavioral vectors on the live target, not file-presence checks
Sources: none - declarative scope statement.
15. Spec self-compliance & health
Self-containment: a fresh agent must be able to run this playbook with only
this document + the probe file + the source vault. If the agent must ask a
question to proceed, this spec failed its own principle 3.
The DP vectors in sec 23 EQUAL the probe's checks, and the capability-contract
examples in sec 11 EQUAL the capability data embedded in the probe (machine form). Any divergence
is a spec-health failure (mirrors plan-spec sec 15).
Cost: the probe is one local process; the gate runs before npm install.
Spec core is collapsed: the always-loaded core is the pinned core + plan
gate (per workflow-spec-v1.1.md section 7); this spec's own contracts are
on-demand.
Sources: none - declarative scope statement.
16. Enforcement/trust root
The trust root of a deployment is the PROBE OUTPUT on the target, not the
agent's description of the target. The feasibility gate reads the probe JSON
file; it does not trust the agent's summary. (Parity with v1.8-host sec 16.)
Sources: none - declarative scope statement.
17. Optional multi-agent orchestration mode - NOT APPLICABLE
Single-agent deployment on the target (v1.8-host sec 17 scoped out).
Sources: none - declarative scope statement.
18. Dynamic cascade controller - NOT APPLICABLE
No cascade on the target (v1.8-host sec 18 scoped out).
Sources: none - declarative scope statement.
19. Glossary
Target machine: the host receiving the deploy (workspace + already-
installed AI toolset analyzed by the probe).
Source vault: the immutable reference vault this spec ships with.
Portable engine: the toolset-agnostic implementation (cores, tests, vault
logic, stores) that imports only node: builtins; the deploy payload.
Bridge: the thin per-toolset translator that maps a toolset's own
primitives (hooks/permissions/events) to the portable engine. opencode's
.opencode/plugins/ factories are the shipped opencode bridge.
Probe / portability-probe.ts: the dependency-free deterministic script
that checks DP vectors, applies the capability-contract examples, and emits the
feasibility JSON.
DP vector: a numbered, state-decidable prerequisite check (DP-001..N);
each is REQUIRED or OPTIONAL.
Capability contract: a portable, tool-neutral WHAT requirement (C-1..C-9,
sec 11); a toolset satisfies one by providing the
primitive through any mechanism.
Reference realization: the opencode EXAMPLE implementation of the
capability contracts (portable engine, probe, conformance harness). An
example, not a requirement. When it cannot run on the destination, the
installing agent MUST adapt it or create a working alternative that satisfies
the intended requirements (a working alternative passes behavioral
verification of its contracts on the live target - never a blind code copy).
Suitable toolset: an already-installed toolset that provides all six
REQUIRED capability contracts (C-1..C-6).
Feasibility verdict:blocked (no suitable existing toolset, or another
REQUIRED failure -> MISSING REPORT: what to install), partial (suitable
existing toolset but a bridge must be written and/or an OPTIONAL contract
degrades), implementable (suitable existing toolset with a shipped bridge).
Adaptation: a documented, allowed deviation triggered by a to-be-written
bridge or an OPTIONAL failure (never a REQUIRED failure, never an installed
toolset).
MISSING REPORT: the fail-closed blocked-case deliverable (produced in the
Analysis phase); uses WHAT IS NEEDED -> WHY -> WHAT IT BLOCKS -> HOW TO OBTAIN
IT, and, when no suitable tool exists, names WHAT NEEDS TO BE INSTALLED.
Analysis report: the first, free deliverable (probe JSON + verdict +
detected toolset + recommended path); investigation output, requires no plan
or approval.
Approval gate: the commitment point at which the user approves the deploy
plan; no target-mutating acting tool runs before it.
Completion report: the success-case deliverable; includes the detected
toolset + bridge used and verification receipts.
Deploy receipt: the probe JSON (feasibility) + step receipts + final
report; the deploy's evidence chain.
Deploy state file:<target>/.governance/deploy-state.json - the
persisted phase, plan-snapshot ref, latest receipt seq, research-evidence
refs, verdict, and resume hint, rewritten at every phase transition
(sec 8).
Research receipt: the hash-chained evidence record every research step
emits (schema in sec 21): question, method (local/online),
local evidence, cited/scoped online sources, findings, used-in, prev_hash. A
claim is grounded in its research receipt - never an unreceipted assertion.
Inexpressible contract: a REQUIRED contract with NO primitive in the
detected toolset; its presence means that toolset is not suitable and, if no
other suitable tool exists, escalates to blocked.
Bridge driver: the shipped verification harness for a bridge. The only
shipped driver is the opencode plugin harness (conformance.ts); a
non-opencode bridge needs its own driver before conformance is claimed.
Capability probe: an observation of the destination toolset's actual hook
surface (opencode: sanity-load.mts). Not shipped for other toolsets -
their suitability is declared from the static matrix, not observed.
Issue notification: a user notification raised on trigger N-1..N-7
(sec 21); always paired with a proposed remediation plan.
Remediation plan: the local-observation + online-data-backed,
spec-conformant plan (validated todowrite per workflow-plan-spec-v1.0.md)
proposed to the user with an issue notification; no remediation action runs
before approval.
Local research: investigation of the destination itself - probe report,
files, config, logs, installed toolset behavior; observed, never assumed.
Online research: proper web research from DATA-BACKED sources only
(official toolset docs, peer-reviewed/arXiv papers, benchmarks with measured
effect sizes), every claim scoped and cited per SFS section 7 (Research-source
conventions).
Data-backed source: a verifiable, linked, HTTP-checked, scoped reference -
never an anecdote or a guess.
Sources: none - declarative scope statement.
20. Benchmark-scoping discipline
No benchmarks are cited. Feasibility is deterministic per target; effect sizes
would be meaningless.
Sources: none - declarative scope statement.
21. Contracts, schemas & state machines
Probe report schema (JSON) - the feasibility receipt
Emitted by EVERY research step (sec 6); hash-chained
with the deploy receipts (prev_hash linked). A claim in any report/plan is
grounded in its research receipt - re-provable from the chain.
User notification triggers + remediation (N-1..N-7)
An ISSUE is any condition below. The agent notifies the user AND proposes a
remediation plan together (research first, plan per the workflow's plan gate,
then user approval - sec 6). No remediation action runs before approval.
ID
Trigger
When
Content
Work pauses?
N-1
Probe failure / no valid JSON
analysis
probe failure + remediation plan (fix probe env, re-probe)
YES - no deploy
N-2
Analysis issue (failed REQUIRED/OPTIONAL vector)
analysis report
issue + adaptations/install needs + remediation plan
YES if blocked; else proceeds to plan approval
N-3
Blocked (no suitable tool / REQUIRED fail)
feasibility review
MISSING REPORT (WHAT/WHY/BLOCKS/HOW + what to install) + remediation plan (obtain + re-probe)
YES - pauses until the user acts
N-4
Plan approval (normal flow, not an issue)
planning
the validated deploy plan
YES - no mutation before approval
N-5
Deploy step failure
deploy
rollback confirmation + remediation plan (corrected steps)
YES - rollback then pause
N-6
Verify failure
verify
failing vectors + remediation plan (fix bridge / re-verify)
YES - no completion report
N-7
Completion with residual issues/adaptations
report
completion report + follow-up plan for residual items
NO - deliverable issued
Every issue row (N-1, N-2-blocked, N-3, N-5, N-6, N-7-residual) delivers:
issue (WHAT/WHY/BLOCKS/HOW) + the researched remediation plan + the approval
request, in ONE notification.
Decision map (D8) - notification trigger:
flowchart TD
I{issue?} -- N-1 probe fail --> P1[notify + remediation plan - pause]
I -- N-2 analysis issue --> P2[notify + adaptations - pause if blocked]
I -- N-3 blocked --> P3[MISSING REPORT + plan - pause]
I -- N-4 plan ready --> P4[present plan - await approval]
I -- N-5 deploy fail --> P5[rollback + notify + plan]
I -- N-6 verify fail --> P6[notify + plan - no report]
I -- N-7 residual --> P7[report + follow-up plan]
analyzed -> blocked: verdict == blocked (MISSING REPORT delivered at
feasibility review; deploy pauses until the user obtains what is missing).
analyzed -> planned: user approves the analysis; deploy plan created per the
workflow's plan gate (workflow-plan-spec-v1.0.md).
planned -> approved: plan validated by the plan gate AND user approves it.
approved -> bridging: verdict == partial (bridge for the DETECTED suitable
toolset per sec 11; never an installed tool).
bridging -> deploying: bridge written/installed for the detected toolset.
approved -> deploying: verdict == implementable (shipped bridge, e.g. opencode).
deploying -> verifying: deploy steps receipted.
verifying -> done: behavioral vectors pass on the LIVE target + idle pass ran.
Fail-closed: probe error or malformed JSON -> blocked with a probe-failure
blocker; no mutation, no deployment without analysis + plan + approval.
Every transition rewrites <target>/.governance/deploy-state.json (sec 8).
Steps, each receipted AND state-persisted before the next:
realize-engine - realize the capability contracts for the destination's
runtime. USE the reference engine (cores + package.json/tsconfig.json) ONLY
when the destination can run it (DP-002/DP-003 pass); otherwise ADAPT the
reference or CREATE A WORKING ALTERNATIVE that satisfies C-1..C-9, verified
behaviorally on the live target, and record the adaptation. Deploy into
<target>/.governance/ (or the toolset's native extension dir when one
exists, e.g. .opencode/).
npm-install - npm install in the engine dir ONLY when the reference
engine is used (DP-002/DP-003 pass); needs DP-012 or vendored node_modules.
seed-vault - copy Rules/, Lessons/, Research/, Templates/,
Index.md to the target vault dir (DP-007/DP-013 source complete).
wire-bridge - register the bridge in the DETECTED toolset's own config
(opencode: legacy .opencode/plugins/ auto-discovery or config file;
claude-code: settings.json hooks; etc.); set the vault path.
apply-migrate - npm run migrate:apply only after dry-run passes.
verify - run the behavioral conformance vectors against the LIVE toolset
(reference harness when runnable; adapted behavioral checks otherwise), one
idle observation pass (DP-010 available), baseline in health.json BEFORE
first idle pass.
flowchart LR
E["1 deploy-engine"] --> N["2 npm-install"]
N --> S["3 seed-vault"]
S --> W["4 wire-bridge"]
W --> DR["5 dry-run-migrate"]
DR --> AP["6 apply-migrate"]
AP --> VF["7 verify"]
VF --> OK[done]
AP -. failure .-> RB["Rollback + notify (N-5)"]
VF -. failure .-> RB2["Rollback + notify (N-6)"]
Loading
Rollback: before step 1, snapshot the engine dir (.governance.bak or the
toolset-native extension dir) if it already existed; on ANY deploy/verify
failure, restore the snapshot AND the pre-install AGENTS.md (G14 legacy
backup), then report. No engine change is left behind on failure.
Sources: none - declarative scope statement.
22. Privacy lifecycle contract - PROBE SCOPE
The probe NEVER reads config files for secrets; it checks existence/
accessibility/writability only, and reports booleans or safe values, never
contents.
The probe report must not embed credentials, tokens, or private key material
(report contains only paths, versions, booleans).
The MISSING REPORT tells the user WHAT NEEDS TO BE INSTALLED and HOW to obtain
it; it never fetches or carries it.
The deploy writes no secrets to the target outside the engine's own config.
Sources: none - declarative scope statement.
23. Executable conformance standard - the DP vectors
Harness contract:portability-probe.ts is dependency-free (node:builtins
only) so it can run BEFORE npm install. It is the REFERENCE analysis tool;
where Node is unavailable on the destination, the agent performs the DP checks
manually per this table (recorded as an adapted analysis). It applies the capability contracts
(sec 11; known toolsets are examples) to the DETECTED toolsets, emits the sec-21 JSON to stdout, and
writes portability-report.json into the target. Exit 0 iff a valid report was
produced; any other exit = probe failure (fail-closed).
version >= 22.6 (REQUIRED only for the REFERENCE engine realization; fail = adapt to another runtime)
DP-003
no
npm
npm --version succeeds (REQUIRED only for the reference engine install; fail = adapt)
DP-004
YES
SUITABLE existing toolset
>= 1 already-installed tool provides all 6 REQUIRED capability contracts (C-1..C-6; known toolsets are examples). None -> blocked: report what to install
Rules/, Lessons/, Research/, Templates/, Index.md all present
DP-014
no
git available
git --version succeeds (repo context)
DP-015
no
model identity
informational - reported, used for STD/STB recalibration (sec 11)
DP-016
no
toolset inventory
detected toolsets + versions reported (input to the matrix)
DP-017
no
selected suitable toolset
the EXISTING toolset chosen for the bridge (informational; NOT a verdict driver)
DP-018
no
selected-toolset capability OBSERVABLE
opencode: sanity-load.mts validates post-deploy; claude/cursor/copilot: NO capability probe shipped - suitability rests on the static matrix (declared, not observed)
DP-019
no
workspace conflict scan
reports existing .opencode/, .governance/, .claude/, AGENTS.md at the target (overwrite risk; investigation only)
REQUIRED set = DP-004, 005, 006, 007, 008, 013.
verdict = blocked if any REQUIRED fails (DP-004 fails when NO suitable existing
toolset is present - MISSING REPORT names what to install); else partial if the
suitable toolset's bridge must be written (non-opencode) OR an OPTIONAL vector
fails (incl. DP-002/DP-003, which are reference-path-only and NEVER block - a
fail means an adapted realization); else implementable. The verdict NEVER
depends on opencode presence.
DP-018/DP-019 are diagnostic (optional). DP-018 is true only when the selected
toolset's capability surface can be OBSERVED (opencode via sanity-load.mts);
for other toolsets it is false and the deploy records it as an adaptation -
never a silent claim.
Conformance: a target is "deployable" iff the probe on that target returns
verdict != blocked AND the deploy state machine (sec 21) completes through
verify. Determinism: re-running the probe on an unchanged target must produce
the same verdict.
Sources: none - declarative scope statement.
24. Evidence base
workflow-spec-v1.1.md (the governance this deploy delivers; section 7/11/12/17 contracts referenced above).
Authoring provenance (resolves in the source vault, not shipped here):
Research/workflow-spec-v1.8-host.md (the operative host realization this derives from),
Research/v1.8-host-parity-plan.md (implementation reality the deploy must reproduce),
Research/client-side-feasibility.md (verified opencode hook surface; the shipped bridge precedent),
Research/host-deployment-spec-v1.0.md (superseded opencode-centric version; retained as history),
.opencode/portability-probe.ts (the reference probe; DP vectors + matrix),
.opencode/plugins/*.ts (the reference opencode bridge),
Live implementation .opencode/ (the reference engine; cores import only node: builtins).
Analysis-first / plan-then-approve / implement-on-approval structure: user
directive 2026-08-20; grounded in the workflow's own plan gate
(workflow-plan-spec-v1.0.md) and selective-enforcement principle
(workflow-spec-v1.1.md principle 3).
Meta-specification. Compiled 2026-08-13; updated 2026-08-21 (v1.1: added the
extension-sections rule - a spec may add numbered extension sections after
the mandatory 24). Defines the required structure,
conventions, and validation rules for every workflow spec document in this
vault (e.g., workflow-spec-v1.7.md). Other spec docs MUST follow this format
so they are self-contained, machine-checkable, and markdown-compatible.
Status: ACTIVE. Applies to the current canonical spec and all future versions.
1. Purpose
This document standardizes how spec documents are written so that:
Every spec is self-contained (all contracts inlined, no cross-file
pointers as the contract).
Every spec is machine-checkable (schemas, state machines, and conformance
vectors are explicit).
Every spec is markdown-compatible (pure ASCII, standard markdown; Mermaid
diagrams render in CommonMark markdown - Mermaid support since 2022-02-28,
github.blog changelog).
Every spec carries verifiable, scoped research sources for each claim.
Every spec complies with its own rules (collapsed core, small always-
loaded section, no circular references).
Sections 21-23 (schemas, state machines, conformance) are mandatory in any
spec that claims to be enforceable. Sections that do not apply to a spec's
scope MUST be stated as "Not applicable to this spec" rather than omitted
silently.
Extension sections (v1.1): the 24 numbered sections above are the
MANDATORY core and stay in that order. A spec MAY add numbered extension
sections AFTER section 24 (or at a documented logical position) when a new
top-level concern warrants it. Each extension section MUST follow all SFS
conventions (Sources paragraph, glossary terms defined before use, pure ASCII,
mermaid for diagrams, no bare secN) and MUST be listed in the section
inventory. An extension section is never a silent renumber of the mandatory
core - the mandatory 24 keep their numbers.
3. YAML frontmatter conventions
Every spec file begins with YAML frontmatter:
---
type: specstatus: canonical | supersededversion: "1.7"supersedes: ["1.0", ..., "1.6"] # only on canonicalsuperseded_by: "1.8"# only on supersededproject: snes
---
Rules:
status: canonical appears ONLY on the current operative version.
A superseded spec carries superseded_by pointing to its immediate successor.
A canonical spec carries supersedes listing all prior versions.
The canonical file is referenced by later specs as the operative document;
prior versions are archived history, never cited as the operative contract.
4. H1 title + status notice
# AGENT GOVERNANCE & WORKFLOW SPECIFICATION v1.7> **CANONICAL (current).** Status: `supersedes` v1.0-v1.6. [one-line change> summary]. Grounding: `Research/<source>.md`. Greenfield design; all prior> versions retained as archived history.
For superseded files:
# AGENT GOVERNANCE & WORKFLOW SPECIFICATION v1.6> **SUPERSEDED** by v1.7 (`workflow-spec-v1.7.md`). Retained as archived> revision history. Do not treat as current.
The status notice MUST be the first content after the H1 (not before it).
5. Text encoding rules (markdown compatibility)
Pure ASCII only. No em-dashes (use -), no minus signs (use -), no
arrows (use ->), no box-drawing characters, no section sign (write sec),
no <= from U+2264 (write <= ASCII), no x from U+00D7 (write x).
Mermaid is ALLOWED and PREFERRED for diagrams. CommonMark markdown renders
Mermaid since 2022-02-28 (github.blog changelog, verified 2026-08-20); use
```mermaid fences for all diagrams. ```text is reserved for
non-diagram listings only; schemas use ```json.
No BOM at file start.
LF line endings (no CR).
UTF-8 encoding, validated (a fatal-decode must pass).
Rationale: a Windows clipboard / CP1252 paste path mangles non-ASCII into
bytes CommonMark renderers flag as corrupt. ASCII-only survives every transport.
6. Markdown conventions
Headings:## N. Title numbered sequentially; H1 only for the document
title.
Tables: standard pipe tables with a header separator row. Every table
cell is single-line where possible.
Code blocks: always tagged with a language: ```mermaid for diagrams
(flowcharts, stateDiagram-v2), ```json for schemas, ```text only
for non-diagram listings, ```yaml for frontmatter examples.
Cross-references: use markdown anchor links to headings
([sec 18](#18-dynamic-cascade-controller)), NOT bare sec18 text and NOT
self-referential anchors (a section must never link to itself).
Inline code: backticks for identifiers, file names, tool names.
Bold for contract names and key terms.
Sources: every section ends with a **Sources:** paragraph of linked
citations.
7. Research-source conventions
Every principle, contract, and claim MUST carry a linked, verifiable source.
Scoped evidence discipline: benchmark numbers are annotated with their
domain/model/caveat and marked "(scoped: ...)". A spec never presents a
single-benchmark effect size as a universal law. If the source paper itself
flags a result as non-robust, the spec repeats that caveat.
Link verification: external URLs are verified (HTTP status + title match)
before inclusion; a fragile link is replaced with a stable one or cited by
title/venue without a link.
Vault-internal references use relative paths (Research/name.md).
8. Schemas & state machines conventions
JSON schemas are presented in ```json blocks with field name, type,
and required/optional semantics per field.
State machines are presented as a ```mermaidstateDiagram-v2
diagram PLUS the textual list of states, transitions (from -> to, on trigger),
and fail-closed behavior. Example:
Every state machine and schema is referenced from the contract that uses it
via anchor link, and each defines its inputs and outputs explicitly.
9. Conformance conventions
Every spec claiming enforceability MUST include an Executable Conformance
Standard section: a table of deterministic test vectors (ID | Contract |
Test | Expected), conformance levels, and a runtime-agnostic harness contract.
Vectors are PASS/FAIL; a FAIL vector is one that must be rejected with a
specified failure reason.
Collapsed core: the spec must state which of its own contracts compress
into the ~5 always-loaded rules and which are on-demand. A spec that exceeds
its own compliance ceiling fails its own principle 5.
No circular references: no section links to itself; cross-version
references point to actual files, not to the current section.
No undefined terms: every specialized term appears in the Glossary before
or at first use.
11. Validation checklist (before publishing a spec)
Run this checklist; a spec is not complete until all pass:
[ ] YAML frontmatter: type=spec, status, version, supersedes/superseded_by
[ ] H1 + status notice first
[ ] All 24 mandatory sections present (or explicitly N/A); extension sections documented
[ ] Pure ASCII: 0 non-ASCII code points
[ ] UTF-8 fatal decode passes; no BOM; no CR; no NUL
[ ] Mermaid fences used for all diagrams (no ASCII/text diagram blocks)
[ ] Code fences balanced (even count); every ```mermaid block is valid syntax
[ ] No bare 'secN' text; all cross-refs are anchor links
[ ] No section links to itself
[ ] Every section has a **Sources:** paragraph
[ ] Every benchmark claim carries a (scoped: ...) caveat
[ ] Glossary defines every specialized term used
[ ] Schemas + state machines present (sec 21)
[ ] Executable conformance standard present (sec 23)
[ ] All external links return success + title-match
[ ] Spec-health: collapsed core stated (sec 15)
12. Versioning
A new canonical version supersedes the prior; the prior's frontmatter flips
to status: superseded with superseded_by: "<new>".
Versions are sequential (v1.0 -> v1.1 -> ...); no overwriting of an archived
version's content after it is superseded.
Each version's file is workflow-spec-vX.Y.md; the file name encodes the
version.
13. Sources
Derived from the revision practice of workflow-spec-v1.0 through v1.7 and
the failures identified in Research/spec-v1.6-weaknesses.md.
Mermaid-in-markdown support: CommonMark markdown renders Mermaid since
2022-02-28 (github.blog changelog, verified 2026-08-20). Prior versions of
this spec banned Mermaid on the outdated premise that the renderer did not
support it.
Encoding: Windows clipboard/CP1252 transport (session-observed).
PLAN SPECIFICATION v1.0 (spec-based plan artifact)
CANONICAL (contract). Defines the machine-checkable Plan artifact that
the workflow requires before acting tools (bash/edit/write/task). Parent
contract: workflow-spec-v1.0.md (Plan gate, Plan definition,
enforcement-loop state machine). This spec makes the plan FORMAT spec-based
and ties the gate's validation to conformance vectors, so the documented plan
and the enforced plan are the same thing.
Format: follows spec-format-specification.md (SFS). Pure ASCII; Mermaid
is allowed per the SFS (CommonMark renders it since 2022-02-28); all 24 sections
present or explicitly N/A.
1. Design Principles
#
Principle
Source
1
Plan before act: a plan must exist before any acting tool (bash/edit/write/task) executes.
The format is enforced, not just stated: the gate validates structure, not the model's judgment.
Rules/Tracked-plan-mandatory.md (vault)
5
A plan is a forward contract: completed/cancelled steps do not satisfy it.
plan-enforce.ts validatePlan
6
The spec and the gate must agree: conformance vectors are the gate's rejection logic, verbatim.
SFS section 9
Sources: none - declarative scope statement.
2. Architecture - the plan in the workflow
flowchart LR
A[routing] --> B[planning]
B --> C["gated: plan gate - valid todowrite + Plan artifact schema"]
C --> D[verifying: action + receipt]
D --> E[completing: canary self-check]
E --> F[idle: plan-gate re-arm]
F --> A
Loading
The plan lives at the planning -> gated transition (the plan gate): the
artifact is a todowrite call with a valid step list (Plan artifact schema,
sec 21); enforcement is
plan-enforce.ts validatePlan (external to the model); the canary logs gate
state (silent-dead gate = critical).
5. Prose->spec translation (how a plan is produced)
A Plan artifact is produced directly as a structured todowrite, not via the
prose-to-spec pipeline (which is for long-lived rules). The plan must be
expressible in the Plan artifact schema (sec 21); a task that cannot be expressed as steps must be split or asked about.
Sources: source vault: prose-to-spec-translation.md (the pipeline is for long-lived rules; plan artifacts are produced directly).
6. Enforcement contracts - the Plan gate
Plan gate: acting tools (bash/edit/write/task) are DENIED until a Plan
artifact (a todowrite whose todos pass validation) exists. Re-arms each turn on
idle. Host mechanism: tool.execute.before in plan-enforce.ts. The gate
rejects a plan when validatePlan returns a non-null reason:
empty (no todos) - PL-003
no actionable step (all completed/cancelled) - PL-004
status not in pending|in_progress|completed|cancelled - PL-004
priority not in high|medium|low - PL-010
Read-evidence requirement (PL-009): beyond a valid plan, the gate requires
HOOK-OBSERVED evidence that the agent READ the plan spec
(workflow-plan-spec-v1.0.md) this session, before acting tools are allowed.
A read of the spec stamps a per-session mark (observed via
tool.execute.after); a write/edit of the spec invalidates the mark (re-read
required). Without the mark the acting tool is denied with a corrective hint
(Read-Gate / deer-flow version-gate pattern). Fail-closed: no mark = deny.
On rejection the gate escalates a review proposal and the agent must supply a
valid Plan (and, for PL-009, a read of the spec) before retrying the acting
tool.
A Plan artifact is transient per-turn state, not persisted knowledge. It does
not enter the vault. (The lessons about plan enforcement are vault knowledge,
governed by the workflow-spec-v1.0.md knowledge lifecycle, out of scope here.)
Sources: workflow-spec-v1.0.md section 7 (knowledge lifecycle, out of scope here).
8. Action receipts contract (plan-related)
The plan gate's BLOCKED / REJECTED events are canary-logged and escalated to
the review queue. Full step receipts are governed by the workflow receipts
contract, out of scope here.
Sources: none - declarative scope statement.
9. Monitoring contract
Plan-gate self-verification: canary logs gate state each load/re-arm; a
silent-dead gate is detectable (workflow enforcement contracts).
Block counter per session feeds the behavioral re-injection (plan-first
directive). Repeated blocks in a session escalate visibility.
Sources: none - declarative scope statement.
10. Communication conventions
The plan-first directive (injected on block) is prose guidance to the model;
the gate itself is deterministic. Communication never gates.
Sources: none - declarative scope statement.
11. Adaptation profile
No per-model adaptation in this spec. (STD/STB re-injection is workflow-scoped.)
Sources: none - declarative scope statement.
12. Directive-priority enforcement (plan-related)
The plan-first requirement is a must-hold directive: it is structurally
enforced by the gate (not just stated). A block + injected directive is the
PIM-style response to a plan-skipping action.
validatePlan rejects empty/placeholder plans (precision over recall)
Sources: none - declarative scope statement.
15. Spec self-compliance & health
The plan spec's own conformance vectors (sec 23) must pass for the spec to be valid.
The gate's rejection reasons ARE the vectors: any divergence between
plan-enforce.ts and this spec is a spec-health failure.
Revision trigger: if the block counter rises without a compliance change, the
plan-first directive or the format is wrong, not the agent.
Collapsed core: the always-loaded core is the plan-first directive and the
gate's rejection reasons (PL-003..PL-010); everything else in this spec is
on-demand detail.
Sources: none - declarative scope statement.
16. Enforcement/trust root
The trust root is the external gate (plan-enforce.ts), not the model. The plan
spec does not change this; it makes the gate's decision rule explicit and
machine-checkable.
Sources: none - declarative scope statement.
17. Optional multi-agent orchestration mode - NOT APPLICABLE
MAS is scoped out on this host (workflow-spec-v1.0.md); single-agent only; no sub-agent plan gates.
every todo satisfies the schema: content is a string of length >= 8, content
does not match PLACEHOLDER, status is one of pending|in_progress|completed|
cancelled, priority is one of high|medium|low, AND
at least one todo is a forward step (status not completed/cancelled).
An acting tool (bash/edit/write/task) additionally requires a per-session mark
proving the agent READ the plan spec this session. Stamp: tool.execute.after
on read of the spec path. Invalidate: tool.execute.after on edit/write of the
spec path. Fail-closed: no mark = deny with a corrective hint.
OPERATIVE (self-contained distribution contract, v1.1). This is the single
operative workflow spec in this package. It is the host realization of the
canonical workflow-spec-v1.7.md (retained in full in the source vault),
folded into a standalone document so no parent reference is required for a
new install. Where a contract is marked SPECIFIED-NOT-IMPLEMENTED, that
marking is the operative truth: the contract is design/reference, not live -
an adapting agent must realize it (or a working alternative) to satisfy it.
This package copy carries the decision maps (D1-D9) of the operative spec.
v1.1 (2026-08-21): adds the Decision-making process top-level section
(sec 5) per SFS v1.1 extension-sections rule;
sections 5-24 renumbered to 6-25.
Trust root = the external enforcement layer, NOT the planner. The planner is a high-risk, low-trust role: most attacked, most consequential, most defended.
Conformance must be executable: deterministic vectors, runtime-agnostic harness, verifiable receipts - a spec without a conformance suite is unenforceable.
BRANCH: client-side-only enforcement on this host. Contracts are realized with the opencode plugin hook surface; a contract that needs server enforcement is realized as a post-hoc, independently recomputable approximation and documented as such - never silently downgraded to prose.
BRANCH (host realization): on this host, Layer 3 is the opencode plugin
layer (verified hook surface from Research/client-side-feasibility.md):
event, chat.message, chat.params, chat.headers, permission.ask,
command.execute.before, tool.execute.before, shell.env,
tool.execute.after, experimental.chat.messages.transform,
experimental.chat.system.transform, experimental.session.compacting,
experimental.compaction.autocontinue, experimental.provider.small_model,
tool.definition, experimental.text.complete. Execution mode is single-agent
only on this host (sec 19/sec 20 scoped out). Layer 0 is the Obsidian vault;
Layer 5 signing keys are out-of-process (Phase 2 of the plan).
flowchart LR
S[Session start / resume - pin + snapshot + spec-core] --> T[Turn start - intent routing + PIM]
T --> P[Plan gate - valid todowrite + PL-009 read evidence]
P --> A[Pre-action - commitment gate + per-tool receipt]
A --> C[Periodic - re-inject omission rules at STD]
C --> F[Completion - post-hoc tri-state evidence verdict on idle]
F --> I[Idle - read-gate, knowledge persist, self-check, plan-gate re-arm]
I --> T
Loading
BRANCH (host realization): the fire points are mapped to the verified
opencode events/hooks above (verified in Research/client-side-feasibility.md).
The plan-gate re-arm uses session.status + {type:"idle"} (the event that
actually fires on the Windows build), not session.idle. session.next.step.ended
and session.next.text.ended are NOT delivered on this host, so completion
evidence is finalized on IDLE from accumulated receipts - a post-hoc evidence
verdict, because no blocking finalize hook exists client-side (see sec 8).
Ambiguous -> ask the user with recommended resolution options.
Everything else -> respond in prose, no spec machinery.
Decision map (D1) - intent routing:
flowchart TD
R[incoming request] --> Q{task type?}
Q -- must-hold / actionable --> A[ensure a spec exists; plan gate applies]
Q -- ambiguous --> B[ask the user with recommended options]
Q -- everything else --> C[respond in prose, no spec machinery]
Loading
BRANCH (host realization): the "ambiguous -> ask with options" branch is
prose/behavioral (G8) - the plugin observes it but does not gate it. This is a
documented non-enforced path, not an omission.
Extension section (SFS v1.1 rule; user-approved 2026-08-21). Makes the
workflow's decision process explicit: how decisions are classified, who
decides, the gated decision loop, and the behaviors the workflow fosters and
structurally blocks. The decision maps D1-D9 and the gates G1/G2/G3/G5/G6 are
defined in their owning sections; this section is the cross-cutting view.
5.1 Decision taxonomy (who decides what)
Every incoming request is first classified (D1, sec 4):
must-hold/actionable -> spec + plan gate; ambiguous -> ask the user with
recommended options; everything else -> prose. Then the state-decidability
boundary (sec 17) splits every rule:
state-decidable -> gate; judgment-required/ambiguity -> prose + user.
Who decides:
The model decides freely for investigation, research, reasoning, and
drafting (enforcement is selective - commitment actions gated, investigation
free, principle 3).
The workflow decides for gated transitions: plan gate (PL-001..010 +
PL-009), G1 context-gain, G2 verify-by-execution, G5 commitment gate, G3
research mode, G6 completion evidence.
The user decides for ambiguity, irreversibility, and approval
(sec 18; D1 ask branch).
5.2 The gated decision loop
Gate
Decision
Fail-closed behavior
Plan gate
may I act at all?
deny until valid todowrite + PL-009 spec read (sec 7)
G1 context-gain
may I mutate?
deny until read (LOW) or reference-graph map shown (HIGH)
G2 verify-by-execution
safe to proceed after a change?
deny until applicable check ran + output shown
G5 commitment gate
may I do something irreversible?
escalate to user's native permission prompt (ask)
G3 research mode
how do I acquire knowledge?
LOCAL for observed state; ONLINE mandatory for external knowledge; research receipt emitted
G6 completion evidence
is it actually done?
claims reconciled to receipts -> VERIFIED/INVALID/UNVERIFIABLE
5.3 The fostered sequence (behavioral convergence)
Classify the request (D1).
Plan before acting (valid todowrite + spec read).
Ask when ambiguous or irreversible (D1 ask branch; G5).
BRANCH (host realization): REFERENCE DESIGN ONLY. This pipeline is not
implemented on this host. Must-hold rules are encoded directly as spec
documents; plan artifacts are produced directly as structured todowrites per
workflow-plan-spec-v1.0.md section 6, not via this pipeline.
Every contract below states its client-side mechanism and, where it is an
approximation of a server-enforced gate, states that explicitly (principle 28).
Plan gate: acting tools denied until a plan exists; re-arms each turn.
Host mechanism: tool.execute.before todowrite-gate in plan-enforce.ts (a
todowrite is a TOOL call; command.execute.before is for slash commands and
cannot intercept it); requires the plan format per workflow-plan-spec-v1.0.md;
re-arms on eventsession.status + {type:"idle"}. Self-verification:
heartbeat/canary logs gate state so a silent-dead gate is detectable, not
silently ignored (fixes the field bug where the gate reset on session.idle
which never fires on Windows). PARITY (v1.8-host) - two host semantics:
(a) PL-009 read evidence - beyond a valid plan, the gate requires a
hook-observed read of workflow-plan-spec-v1.0.md this session
(tool.execute.after stamps the mark; edit/write of the spec invalidates it;
fail-closed: no mark = deny with a corrective hint - workflow-plan-spec-v1.0.md
section 22). (b) planCreated is process-global, not per-session - one
session's todowrite arms the gate for every session until the next idle
re-arms it; PL-009 read-marks remain per-session. The asymmetry is documented,
known behavior.
Commitment gate: for write/send/irreversible only. MUST be environment-
external (framework gates leak sibling effects during pauses - Stop-Means-Stop). Host mechanism: the plugin's permission.ask
hook sets status: "ask" so the user's native permission prompt is the
authoritative environment-external boundary (see the second Commitment gate
paragraph below); deny wins over ask. PARITY (v1.8-host): the gate is
user-mediated, NOT a deterministic deny; sabotage-validation and
precision-auditing are SPECIFIED-NOT-IMPLEMENTED, and the gate has zero test
coverage (as of 2026-08-20). Verdict schema in sec 23.
Verify-before-retry: never blindly retry; verify postcondition first +
idempotency key. PARITY (v1.8-host): PARTIAL. The plugin keeps a
(tool, args-hash) -> result store fed by tool.execute.after; tool.execute.before
checks only for store EXISTENCE and emits a retry_hint evidence marker when a
prior identical call succeeded. The postcondition PROBE (postconditionMet) is
test-only (enforcement-core.ts); post_condition_probe_hash is always null
live. The full Verified-Tool-Calls probe is SPECIFIED-NOT-IMPLEMENTED.
(Verified-Tool-Calls; G4)
Verify-between-cycles: unverified intermediate state never committed.
Invariant monitor:PARITY (v1.8-host): SPECIFIED-NOT-IMPLEMENTED. No
deviation/drift-detection layer exists on this host; no frozen admission-time
snapshot is maintained. The nearest live mechanisms are: the AGENTS.md budget
guard (25KB/200 lines, vault-core.ts), the health counters, and read-gate
miss detection at idle. (From-Admission-to-Invariants)
Completion evidence - BRANCH (post-hoc receipted evidence): "done" requires
spec acceptance + fresh execution-tested evidence + a deterministic
recomputation. Host realization: because no blocking finalize hook exists
client-side AND session.next.step.ended/session.next.text.ended are NOT
delivered on this host, the record is finalized on IDLE from accumulated
receipts as a post-hoc, independently recomputable evidence record
(evidence-of-action family) with a tri-state verdict - VERIFIED, INVALID,
UNVERIFIABLE. PARITY (v1.8-host): live VERIFIED verdicts are predominantly
"execution evidence present" (claimed_totals empty) because claim-vs-receipt
reconciliation only fires when final prose contains work-done claims; the one
live INVALID (evidence seq 93) had a parser-artifact claim key; UNVERIFIABLE
has never occurred live. Trajectory-level Near-Miss pre-conditions are
SPECIFIED-NOT-IMPLEMENTED. This is an honest post-hoc approximation, NOT a
blocking finalize gate (server-side-only; documented, not claimed). Verdict
schema in sec 23.
flowchart TD
F[final prose claims] --> H{work-done claims?}
H -- yes --> C{claims match receipts?}
C -- yes --> V[VERIFIED]
C -- no --> I[INVALID - corrected to receipted figure]
H -- no --> R{receipts exist?}
R -- yes --> V
R -- no --> U[UNVERIFIABLE]
Loading
Read-gate (idle): dormant-signature skips flagged and counted. Host
realization: plan-enforce.ts stamps every read of a Lessons/*.md file
(.plan-enforce-lesson-reads.json, shared with self-maintain); on idle,
self-maintain flags a lesson whose signature recurs but whose vault file was
never read (a SKIP - visible, not silent), deduped against its own revive
proposals.
Commitment gate (permission boundary): for irreversible/destructive actions
(task tool; write/edit targeting outside the project dir) the plugin's
permission.ask hook sets status: "ask" so the user's native permission
prompt is the authoritative environment-external boundary. Non-destructive
actions pass through (plan gate covers them); deny wins over ask. Ordering:
plan-enforce's tool.execute.before is the deny gate; enforcement's
permission.ask only escalates to ask when nothing denied.
PIM: cheap scriptable gate on every chat.message (keyword/priority-pair
match vs the pinned directive list), writing a pim_conflict evidence marker
on a flag. PARITY (v1.8-host): the small-model judge is NOT wired -
experimental.provider.small_model is never invoked by any plugin; the cheap
gate is the whole PIM. Bounds per-message cost/latency (LAS: cheap gate
first) are trivially satisfied.
SOM: review-then-repair - review each drafted response against the pinned
higher-priority directives and repair via experimental.text.complete
(or filter via experimental.chat.messages.transform). A post-hoc
correction approximation, NOT a pre-send block. Evidence taxonomy
(Global Evidence Directive): countable work-done claims are split against
session receipts - VERIFIED claims need no label (high confidence default);
INVALID claims are corrected to the receipted figure; unverified claims the
model labeled (FACT / REASONING / [INFERENCE] / UNVERIFIED / UNKNOWN) are
allowed; only unverified AND unlabeled claims get a targeted
"[UNVERIFIED: ]" marker. Never manufacture certainty.
Safe Turn Depth / Safe Token Budget: re-inject omission (prohibition) rules
every k < STD; cap session at STB for safety-critical constraints. Values are
per-model, computed by calibration. Host mechanism: chat.params fires per LLM
call; plugin counts turns/tokens and re-injects omission rules at k < STD.
PARITY (v1.8-host): the stb_over_budget marker fires on essentially every
turn after turn ~8 across sessions (82% of the live evidence log) - it is a
MARKER ONLY (no enforcement action) and over-triggers; the calibration is
flagged as suspect. (SRD)
Principled stopping:PARITY (v1.8-host): SPECIFIED-NOT-IMPLEMENTED.
Belief-based stop/re-pair is advisory only; the actual recovery discipline is
fixed: retry once -> escalate -> 2-fail stop -> failed verify = revert
(sec 12). (VRR-Stop)
Enforcement self-verification: gates sabotage-validated on a schedule;
heartbeat + canary independent of the failing subject; enforcement events
receipted; silent-dead gate = critical alert. Host mechanism: the plan-gate
canary + a conformance check (WS-018). (Monitoring; plan-enforce field bug)
Context-gain gate (G1, risk-graded): acting tools that MUTATE an
interlinked artifact set are DENIED until local evidence shows the agent
mapped the set's reference graph this cycle. Depth scales with mutation class:
LOW (single-file content edit AND isolation VERIFIED - no external refs to the
changed file) requires only the changed file's relevant region; HIGH
(rename/remove/merge a referenced file, cross-file refs, >= 2 interlinked
files) requires the full graph - which files exist, which reference which
(parent/child, links, filenames, anchors) - SHOWN. The map is computed by
grep/glob (a few tokens); full file reads happen only for the subset that will
change. Moderate uncertainty about isolation is material uncertainty and MUST
be resolved by the map before mutating; never classify LOW on assumption.
Generalizes the plan-spec PL-009 read-evidence pattern to every artifact
class. Fail-closed: no map = deny.
(Read-Gate arXiv:2608.02011;
deer-flow read-before-write)
Verify-by-execution gate (G2, risk-graded): after a state-changing step,
the next acting step is DENIED until the APPLICABLE deterministic verification
passes and its output is shown. LOW mutations run the cheap applicable check
or a read-back re-fetch of the changed state; HIGH mutations run the full
applicable verification. Selector (existing checks, not new tools): spec ->
sfs-validate; code -> tests/conformance; vault -> vault-audit; doc
set/distribution -> reference-resolution (links, filenames, parents, anchors
resolve); any artifact -> read-back re-fetch. Deterministic checks are non-LLM
and low-token. Fail-closed: any gate error = deny, never fail-open.
(Self-Healing Agentic Orchestrators;
Verified Tool Calls;
Near-Miss; vault lesson
verify-after-change)
Research-mode decision (G3): every research step records its mode
(local / online / both). "Local-only" is a deliberate labeled choice valid only
for local-answerable questions (e.g. reference-graph resolution); unperformed
local research (acting without reading/mapping) is a G1 violation. ONLINE
research is MANDATORY, not optional, whenever external knowledge is needed or
local evidence is insufficient - never answer an external-knowledge question
from local inference alone (CRAG corrective fallback). Online sources are
data-backed, scoped + cited, >= 2 independent sources with authority + recency
checked (superseding docs beat the most-relevant one); retrieved evidence is
GRADED (relevant/authoritative/fresh) before use - noisy or low-authority docs
degrade correctness.
Decision maps (D2-D4):
flowchart TD
M[mutate artifact set] --> I{existing file?}
I -- new --> OK[proceed]
I -- existing --> R{read this session?}
R -- yes --> OK
R -- no --> U{moderate uncertainty about refs?}
U -- no --> LOW[LOW: read the file region; cheap verify or read-back]
U -- yes --> MAP[map the reference graph - grep]
MAP --> X{external refs found?}
X -- no --> LOW
X -- yes --> HIGH[HIGH: full map SHOWN + full verification]
M -- rename / remove / merge / cross-file --> HIGH
LOW --> A[verify-by-execution - fail-closed]
HIGH --> A
Loading
flowchart TD
A[state change] --> C{artifact class}
C -- spec --> S[sfs-validate]
C -- code --> T[tests / conformance]
C -- vault --> VA[vault-audit]
C -- doc set / distribution --> R[reference-resolution]
C -- other --> RB[read-back re-fetch]
S --> P{pass?}
T --> P
VA --> P
R --> P
RB --> P
P -- yes --> N[next step]
P -- no --> D[deny - fail-closed]
Loading
flowchart TD
Q[research question] --> L{answerable from observed local state?}
L -- yes --> LO[LOCAL: read/map the relevant subset - record mode]
L -- no --> E{external knowledge needed?}
E -- yes --> ON[ONLINE: MANDATORY, data-backed, scoped + cited, >= 2 sources, authority + recency]
E -- both --> B[BOTH]
LO --> RR[research receipt]
ON --> RR
B --> RR
Loading
8. Knowledge lifecycle (labeled store contract)
Every entry carries: label (fact | inferred | rule | validated-override);
provenance (writer identity, evidence reference, derivation, valid-time,
transaction-time - survives conflict resolution); three trust axes
(relevance/reliability/staleness, independent); scope (cross-user AND
cross-project; a locally-valid project rule is never applied to another project
without explicit elevation).
Security:PARITY (v1.8-host): PARTIAL. Implemented: the write-path
evidence gate at lesson promotion (EGM)
and the PII policy (forbidden categories rejected at the promotion write gate;
PII flagged, never rejected; receipts store hashed args/results).
SPECIFIED-NOT-IMPLEMENTED: origin-bound writes (lessons carry no writer
identity in frontmatter), source isolation, quarantine-on-risk, artifact-level
scope control. The gate covers ONE store (lesson promotion) - AGENTS.md
rendering, receipts, evidence, and review proposals are not gated.
Evaluator locking:PARITY (v1.8-host): SPECIFIED-NOT-IMPLEMENTED. Trust
signals are not computed under a locked regime on this host.
(RewardHackingAgents)
Held-out validation for promotion: a lesson promotes on 2x evidence in
window N; correctness confirmed on windows N+k (k>0) or cross-project/cross-
domain data it never saw. PARITY (v1.8-host): promotion and held-out window
RECORDING are implemented; the demotion path (held-out FAIL -> demote ->
AGENTS.md regen + re-verify) is SPECIFIED-NOT-IMPLEMENTED
(recordPromotion is test-only; promotion_holdout_failures never advances
live). (SpecBench)
BRANCH (host realization):
Time-anchored revival (flap fix): when a lesson is archived, set
archivedAt in frontmatter. Revive ONLY when the signature recurs with
time_created > archivedAt, not across the whole 7-day window. Eliminates
the live revived 33 / archived 33 oscillation (root cause of the AGENTS.md
47KB corruption).
Budget = category-aware split: always-loaded AGENTS.md capped at 25KB /
200 lines; large prose sections (Rules project notes) move to on-demand vault
notes loaded via path-scoping, NOT archived away. Budget enforced by a guard
so the cap cannot be exceeded by future additions.
Safe migration (G14): regenAgents runs dry-run to scratch output first,
verifies budget AND that no lesson is lost (each permanent present once, each
archived present in the archive index - content-based, not count-based);
only then writes the real AGENTS.md; prior AGENTS.md kept as legacy until the
new one verifies. (aletheia migration pattern)
archivedAt persistence (v1.8-host, verified 2026-08-20):archivedAt IS
written to frontmatter at archive time; legacy lessons without it are
backfilled to "now" in-memory so pre-archive evidence cannot revive them.
Knowledge-drift monitoring (v1.8-host): SPECIFIED-NOT-IMPLEMENTED. The
only "drift" artifact is a review-proposal item type label. The nearest live
mechanism is read-gate miss detection at idle.
Every step emits a receipt with schema in sec 23: post-condition probes on
writes (falsifiable claims); receipt-derived summary (verified Actions vs LLM
Narrative vs auto-flagged Divergence, claims_verified/claims_total);
non-execution proof for denied/aborted steps; human-view observation as a
first-class detection signal; hash-chained, signed, fail-closed; behavioral
channel is the ONLY observability of process fidelity (DPI).
PARITY (v1.8-host): signatures are ed25519; chain verification
(verifyReceiptChain) is offline/test-only, not invoked by any plugin.
outcome is derived from the result title containing "error";
reversibility is false for bash/edit/write/task; parent_seq,
post_condition_probe_hash, and evidence_ref are always null live (the
schema is broader than the client realization).
Key custody: signing keys out-of-process / separate from the agent
(principle 19). In-process keys = evidence-grade but NOT non-repudiable (stated
limitation). Host: Phase 2 moves keys out-of-process; until then the stated
limitation applies.
BRANCH (host realization): the plugin writes receipts from observed events
(tool.execute.before/after + event), not from a server-authoritative trace.
Client-side replay is approximate; authoritative full-trace replay is
server-side-only and out of scope (documented, not claimed).
Multi-signal, never a single LLM monitor; measured MR/FAR per deployment
(AutoMonitor-Bench).
PARITY (v1.8-host): the live monitor is SINGLE-signal
(completion_evidence only); plan-gate blocks, privacy denials, and read-gate
skips never reach monitor.jsonl; MR/FAR is computed in tests only, never
live.
Never LLM-judge pass/fail on correctness-critical metrics.
Detection is model-dependent; use per-model rejection rates.
Category-aware: on-demand/skill constraints get stronger, not weaker,
enforcement (OCTOBENCH).
Latent sub-agent failure modes detected + attributed, never repaired by retry
(OrchestraBench - MAS-related; retained as a principle, not exercised on this
single-agent host).
BRANCH (host realization): baselines recorded BEFORE Phase 1 so success is
before/after, not asserted: claims_verified/claims_total, AGENTS.md
bytes/lines, flap count per pass, plan-gate-fires count. The spec-health
counters (sec 17) report against these baselines.
PARITY (v1.8-host): only claims_verified/claims_total advance live
(verified rate ~62.5% as of 2026-08-20); promotions,
promotion_holdout_failures, and constraint_survival never advance (their
recorders are test-only).
11. Communication conventions (Layer 4 - prose, non-enforced)
Confidence footer: High = receipt-verified, Medium = labeled inference, Low =
unverifiable. Communication only - never gates.
Evidence taxonomy (Global Evidence Directive): FACT (established), REASONING
(derived), [INFERENCE] (deduction), UNVERIFIED (claim/report not established),
UNKNOWN (insufficient evidence). Verify material uncertainty, not everything;
never manufacture certainty. User escalation format: WHAT IS NEEDED -> WHY ->
WHAT IT BLOCKS -> HOW TO OBTAIN IT.
Facts | Reasoning | [Inference] labeling.
Failure recovery: retry once -> escalate -> 2-fail stop -> failed verify =
revert (stopping is belief-based per sec 8).
flowchart TD
C[claim] --> V{receipt-verified?}
V -- yes --> FACT[FACT - high confidence]
V -- no --> R{derived from evidence?}
R -- yes --> REASON[REASONING]
R -- no --> I{deduction?}
I -- yes --> INF[INFERENCE]
I -- no --> U{established?}
U -- no --> UNK[UNKNOWN - insufficient evidence - ask]
U -- yes --> UNV[UNVERIFIED - label as such]
Instruction priority is not reliably model-enforced, so it is structurally
enforced (principle 9):
Parallel Input Monitor (PIM): a separate thread checks each new
low-priority message against higher-priority instructions; on conflict,
discard speculative output, inject a sanitized warning, re-run. Host: cheap
scriptable gate on chat.message only; PARITY (v1.8-host): the
small-model judge is NOT wired (never invoked).
Sequential Output Monitor (SOM): reviews drafted responses for
violations of higher-priority instructions - catches realization failures.
Host: review-then-repair via experimental.text.complete /
experimental.chat.messages.transform.
Critical rules first + re-injection: the ~3-5 highest-value constraints
in the strongest-recency position; re-inject near the end of each request.
Positive phrasing, not prohibitions; prune history demonstrating banned
behavior (examples beat instructions).
Priority as data: explicit privilege values (ManyIH-style), not role
labels.
Spec core is collapsed: the spec defines which contracts are the ~5
always-loaded rules and which are on-demand - it must not exceed its own
compliance ceiling.
Spec-health metric: the spec is revised when its own metrics don't move -
e.g., if the completion evidence verdicts don't reduce false-success, the
evidence mechanism is wrong, not the agent. A revision trigger exists.
PARITY (v1.8-host): the only live counters are
claims_verified/claims_total; promotions/holdout-failures/constraint-
survival are library-only and MUST NOT be reported as live signals.
Cost budget: the full stack's combined cost is budgeted; escalate through
the cheapest layer that holds.
Human-oversight budget: human time sized and capped (~15min/day review
queue). BRANCH: the review surface is a reviews/ note dir + notification
via the idle toast/log; propose-then-approve for drift and lesson promotions
(G13).
State-decidability boundary: every rule is classified - state-decidable ->
gate; judgment-required/ambiguity -> prose + user. Rules that can't be
state-decided are explicitly NOT gated.
Verify-gate health (v1.8-host): the context-gain (G1) and
verify-by-execution (G2) gates are live-counter candidates - context-gain
denials and verify-after denials per cycle are recorded; if the count rises
without artifact-quality change, the gate or its calibration is wrong, not
the agent.
BRANCH: client-side boundary is a documented scope line, not an
afterthought: completion evidence = post-hoc tri-state verdict (not
blocking); SOM = review-then-repair (not pre-send block); replay =
approximate (not authoritative). PARITY (v1.8-host): the scope line also
covers the SPECIFIED-NOT-IMPLEMENTED contracts - commitment-gate
sabotage-validation, verify-before-retry probe, invariant monitor, security
set, knowledge-drift/evaluator locking, principled stopping, and the
prose->spec pipeline. Any future claim that these are implemented on this
host is a conformance failure.
Sources: none - declarative scope statement.
17. Enforcement/trust root
The trust root is the environment-external enforcement layer (the
deterministic gates, receipts, and their out-of-process keys),
NOT the planner and NOT any model. Per principle 4, enforcement runs outside the
model's context and cannot be overruled; per principle 19, keys live
out-of-process. The model (including the planner) is never the trust root.
BRANCH (host realization): the trust root is the plugin gate layer
(tool.execute.before, permission.ask, plan gate) + the receipt store + the
out-of-process key (Phase 2). The learn session and the planner are never the
trust root.
18. Optional multi-agent orchestration mode - SCOPED OUT (BRANCH)
Not applicable to this spec (stated, not omitted): single-agent is the
default and the ONLY mode on this host. MAS orchestration (delegation gate,
critic-veto, context firewalls, latent-mode containment, orchestration
simulation, role-aware assignment) is not implemented on this single-user
deployment. The principles remain recorded (below) for a future multi-user or
regulated scope, per the plan's decision. Re-enter MAS only when the cascade
controller (sec 20) is also re-entered.
Delegation gate, centralized hub-and-spoke topology, planner as high-risk
low-trust role, critic-veto, context firewalls, latent-mode containment,
orchestration simulation, role-aware model assignment (retained from v1.7 for
future scope; see Research/multi-agent-orchestration.md).
19. Dynamic cascade controller - SCOPED OUT (BRANCH)
Not applicable to this spec (stated, not omitted): there is no cascade
controller on this host; the mode is always single-agent. The v1.7 cascade
design (start single-agent, cheap cascade gate, escalate-to-MAS, DAG topology,
route-tasks-not-calls, routing evidence) is retained only as reference for a
future scope. Bound: if multi-user/regulated scope is ever adopted, sec 19
and 18 are re-entered together, with routing fidelity held-out validated per
the v1.7 contract.
State-decidable rule: a rule whose compliance can be determined from the
current tool arguments and observable state alone (no judgment, no ambiguity).
State-decidable rules may be gates; others are prose + user.
Must-hold rule: a non-negotiable, binary, checkable directive (vs.
contextual guidance, which stays prose).
Spec: the machine-checkable encoding of a must-hold rule (the sec 7
pipeline is REFERENCE DESIGN ONLY on this host; specs are authored directly).
Plan: an explicit step-list artifact created before acting tools are
permitted (the plan gate's precondition). The plan FORMAT and validation
contract are defined in workflow-plan-spec-v1.0.md.
Held-out validation: verifying a lesson/routing decision on data/windows
it did not see during promotion (window N+k, or cross-project/cross-domain).
Safe Turn Depth (STD): the model-specific turn depth before omission
(prohibition) compliance degrades; re-inject at k < STD.
Safe Token Budget (STB): STD x mean tokens/turn; cap session at STB for
safety-critical constraints.
Receipt-grounded: a claim backed by a signed, hash-chained step receipt
(the confidence-footer "High" basis).
Divergence flag: an automated flag when the LLM narrative and the
receipt-verified actions disagree (claims_verified/claims_total).
Quarantine: demote or block a memory entry from retrieval/promotion while
flagged risky, without deleting it. (SPECIFIED-NOT-IMPLEMENTED on this host.)
Evaluator locking: computing the trust signal from pristine sources under
a locked regime, independent of the agent's reported value.
(SPECIFIED-NOT-IMPLEMENTED on this host.)
Crypto-shredding: erasing by destroying a per-subject/per-cell key so
ciphertext becomes permanently unrecoverable while the audit chain stays
intact (GDPR erasure over append-only stores).
Retraction tombstone: an additive record asserting a prior fact is
retracted as of an instant; satisfies audit/replay without destructive delete.
BRANCH - Evidence record (receipted evidence): a post-hoc, independently
recomputable statement about an action, carrying a tri-state verdict
(VERIFIED / INVALID / UNVERIFIABLE). An evidence record does not gate
anything; its verification produces a verdict about the record
(IETF draft-msebenzi-evidence-action-00).
BRANCH - Cheap gate: a deterministic, scriptable pre-check run before any
LLM-based monitor, to bound cost/latency (LAS: cheap gate first).
BRANCH - Review-then-repair: SOM's client-side realization: review a
drafted response and correct it post-hoc via text.complete; a post-hoc
correction, not a pre-send block.
BRANCH - Baseline: a pre-implementation measured value used as the
before/after reference for a metric target (G10).
BRANCH - Observation window: the acceptance criterion for deployment-
only phenomena (e.g., flap = 0 over 3 consecutive idle passes) that unit tests
cannot prove (G11).
BRANCH - Canary/heartbeat: an independent, deterministic signal that a
gate is alive; silent-dead gate = critical alert (principle 16).
BRANCH - SPECIFIED-NOT-IMPLEMENTED: a contract retained as design/
reference but with no live implementation on this host. Marking it is the
parity mechanism: the contract is neither silently downgraded to prose nor
falsely claimed as live.
Sources: inline citations above.
21. Benchmark-scoping discipline
Every measurement cited in this spec is a benchmark result, not a universal
law. Each claim carries its domain/model/caveat, as marked inline (e.g.,
"(scoped: ...)"). The spec deliberately avoids presenting single-benchmark
effect sizes as general guarantees; where a paper itself flags a result as
non-robust (e.g., Nature MI's ~45% threshold is a validated selection rule, not
a coefficient-level scaling law), the spec repeats that caveat. Conformance
testing (sec 25) measures the spec's own deployment, not benchmark
extrapolation.
The full v1.7 contract (immutable fact vs. erasable content, crypto-shredding,
eraseByScope on every store, retention TTLs, verified-erasure checks,
user-accessible memory visibility) is staged on this host:
Implemented (minimal path - current scope): classify entries at the
lesson-promotion write gate (forbidden categories: credentials, secrets, keys
rejected); never store raw credentials/PII in lessons or markers; PII-bearing
entries flagged (advisory only). PARITY (v1.8-host): the gate covers ONE
store (lesson promotion) - AGENTS.md rendering, receipts, evidence, and
review proposals are not gated. Decision-proof retention is SPECIFIED-NOT-
IMPLEMENTED: decisionProofHash is library/test-only, never called live.
This covers the actual single-user risk (accidental secret/PII persistence in
promoted lessons).
Deferred (only if multi-user/regulated): full crypto-shredding
(per-subject keys + eraseByScope fan-out across every store), retention
TTLs default-on for PII, verified-erasure checks, user-accessible memory
visibility. The contract is staged minimal-now /
full-when-scope-requires.
The append-only / "never delete" invariant and the right-to-erasure conflict
resolution method (key destruction, retraction tombstones, decision-proof
records) remain the reference design and are NOT contradicted by the deferral.
24. Executable conformance standard - BRANCH (host-scaled)
A governance spec without a conformance suite is unenforceable. This spec
defines a deterministic, runtime-agnostic conformance suite; the host's suite is
the test files under .opencode/ run against the plugin API directly.
Harness contract (runtime-agnostic): a single harness drives YAML/JSON
cases against any implementation through a thin subprocess adapter (JSON over
stdin/stdout - per Open Agent Spec conformance).
LLM calls MUST be mocked/stubbed; tests assert runtime BEHAVIOR, never LLM
output content. Cases carry requires: capability tags; unsupported features
report UNSUPPORTED, not PASS/FAIL.
Deterministic vectors (per H33):
each vector is an input -> expected_output JSON contract; byte-identical output
required; determinism verified (same vector 1,000x -> 1,000 identical);
immutable once published.
Normative vectors (PASS and FAIL) - BRANCH level table:
ID
Contract
Test
Expected
WS-001
Commitment gate
Submit action matching DENY policy
PARITY: user-mediated permission.ask escalation (deny wins over ask); NOT a deterministic deny gate; no test coverage as of 2026-08-20
WS-002
Commitment gate
Make gate unavailable, submit action
PARITY: deterministic fail-closed SPECIFIED-NOT-IMPLEMENTED; fail-open bypass behavior NOT verified (no test)
WS-003
Plan gate
Acting tool without plan
Tool denied; structured reason returned
WS-004
Receipts
Generate receipt for ALLOW action
Receipt present with required fields; signature verifies offline
Constraint may drop (violation detectable) - negative case untested
BRANCH addition - completion evidence vectors:
ID
Contract
Test
Expected
WS-021
Completion evidence
Run states total matching receipted results
VERIFIED verdict recorded
WS-022
Completion evidence
Run states total contradicting receipted results
INVALID verdict recorded
WS-023
Completion evidence
Run claims a result with no receipt
UNVERIFIABLE verdict recorded
Deployment verification tooling (v1.8-host):verify-e2e.ts replays
promote+regen passes on a scratch vault copy against the real DB and asserts
steady-state flap = 0; the observation-window criterion (G11, e.g. flap = 0
over 3 consecutive idle passes) is the acceptance test for deployment-only
phenomena that unit tests cannot prove. Neither is scheduled in CI; the
observation-window check is manual.
Core (WS-001..006): gates + receipts + fail-closed. PARITY: WS-001/002
are user-mediated commitment escalation; the deterministic deny + fail-closed
vectors are NOT implemented - the Core level claim must reflect this.
Extended (Core + WS-007, WS-010..013): store + lifecycle + directive
monitor. MUST all pass on this host. WS-008/009 report UNSUPPORTED. WS-011
reports SPECIFIED-NOT-IMPLEMENTED.
Full (Extended + WS-017..020 + WS-021..023): invariant + self-verification
An implementation SHALL be conformant at a level iff all vectors in that
level's set produce byte-identical expected outputs, deterministically, with
UNSUPPORTED (not FAIL) for scoped-out features.
Verifiable conformance (per OAP):
two tiers - self-issued (operator runs the suite, signs the result with an
out-of-process key, anchors it) verified by anyone re-running the same suite;
and independent-witness (a second party runs the same suite, signs a matching
attestation). The verifiable conformance receipt is machine-readable and
re-runnable, not a narrative audit certificate. (BRANCH: self-issued tier is
the current host scope.)