Skip to content

Instantly share code, notes, and snippets.

@mpalpha
Last active August 22, 2026 02:01
Show Gist options
  • Select an option

  • Save mpalpha/cd391ce91bdca79257089cc4a5ef0949 to your computer and use it in GitHub Desktop.

Select an option

Save mpalpha/cd391ce91bdca79257089cc4a5ef0949 to your computer and use it in GitHub Desktop.
Agent governance workflow - portable spec + deployment package (operative set, decision maps, CommonMark-compatible)

Agent Governance Workflow - Portable Spec & Adaptation Package

Version 1.3 (2026-08-21). A COMPACT, markdown-compatible package: the normative specs + the adaptation playbook that directs an agent to RESEARCH, ANALYZE, and ADAPT the target machine's current workflow so it functions with the features of the spec. Pure ASCII, PII-free, Mermaid diagrams render in CommonMark markdown. No user-identifying paths or credentials anywhere in this package.

What this package is

Portability means: research -> analyze -> adapt the destination's CURRENT workflow so it functions with the spec's features (the capability contracts C-1..C-9 and their observable invariants), using whatever tooling ALREADY exists and creating a working alternative where needed. Nothing is installed unless no suitable realization exists, and then the report says what to install.

This package is the SPEC + the ADAPTATION INSTRUCTIONS - compact, markdown-only, markdown-compatible. It is NOT a code bundle.

Research mandate (always)

The agent performing analysis or adaptation ALWAYS uses proper, systematic, OPTIMAL, correct, DATA-BACKED research at every step:

  • LOCAL: observe the destination directly (probe, files, config, logs, installed toolset behavior) - never assume.
  • ONLINE: data-backed sources only (official docs, peer-reviewed/arXiv papers, benchmarks with measured effect sizes), scoped + cited.
  • Evidence saved (vault Research/ or the deploy evidence store) and cited.
  • Never anecdote, guess, or a suboptimal shortcut. (host-deployment-spec-v1.1 section 6 Systematic analysis & research discipline; principle 13.)

Contents (operative set, versioned from 1.0)

File Role Status
workflow-spec-v1.1.md The one operative workflow contract (standalone distribution spec derived from the v1.8-host realization of canonical v1.7; parity-marked SPECIFIED-NOT-IMPLEMENTED; Mermaid flows; includes the Decision-making process section sec 5) OPERATIVE (v1.1)
host-deployment-spec-v1.1.md The one deployment & adaptation playbook (analysis-first -> plan -> approval -> adapt; capability contracts C-1..C-9; toolset-neutral; 3 Mermaid flows) OPERATIVE (v1.1)
workflow-plan-spec-v1.0.md The one plan-artifact contract + PL-001..PL-010 conformance vectors (the plan gate's decision rule) CANONICAL (v1.0)
spec-format-specification.md The one format spec (SFS v1.1) - format/validation rules every spec MUST follow ACTIVE
README.md This manifest -

Versions start at 1.0 per document; a revision bumps the stamp (e.g. host-deployment 1.0 -> 1.1). Superseded versions are retained in full in the source vault, never re-shipped here.

Read order

  1. workflow-spec-v1.1.md - the one operative workflow contract (WHAT + its host realization, with honest SPECIFIED-NOT-IMPLEMENTED markers; read sec 5 for how the workflow directs decisions).
  2. workflow-plan-spec-v1.0.md - the plan artifact + plan-gate rules.
  3. host-deployment-spec-v1.1.md - HOW to research, analyze, and adapt a target machine's current workflow to the spec.
  4. spec-format-specification.md - the format contract for any spec document.

Adaptation flow (host-deployment-spec-v1.1)

  1. ANALYSIS (free, data-backed research): assess the destination's current workflow and toolset (probe or manual DP checks) -> Analysis Report.
  2. PLAN: create the deploy plan via the workflow's own plan gate.
  3. APPROVAL: the user approves before anything mutates the target.
  4. ADAPT: realize the capability contracts (C-1..C-9) by adapting what exists or creating a working alternative (researched, data-backed), verified behaviorally on the live target; install nothing else.
  5. VERIFY: behavioral conformance vectors on the live target.
  6. BLOCKED case: MISSING REPORT says what needs to be installed (primitives + example toolsets, never a forced brand).

Feasibility is decided by the capability contracts (C-1..C-9 in host-deployment-spec-v1.1.md), not by toolset names.

Reference implementation (pointed to, NOT bundled)

The opencode reference realization of the capability contracts (engine cores, portability probe, conformance + tests, opencode bridge plugins, tooling) lives in the source workspace's .opencode/ directory. It is an EXAMPLE of an adapted realization for the agent to run where compatible or adapt from. It is intentionally NOT bundled here, keeping this package compact and markdown-compatible.

Portability invariants (enforced)

  • WHAT-first: contracts are observable invariants; toolsets and the reference realization are examples only.
  • Research-gated: every analysis/adaptation is data-backed (local + online), scoped, cited, saved - never guessed.
  • Toolset-neutral: the verdict never depends on opencode presence.
  • PII-free: no usernames, credentials, or user-embedding paths (checked).
  • ASCII: every document is pure ASCII (SFS).
  • Mermaid: diagrams are Mermaid (renders in CommonMark markdown).
  • Markdown-compatible: valid CommonMark, no BOM/CR.

Provenance note

Research/... and .opencode/... references in the spec documents are source-vault / reference-implementation provenance (authoring evidence), intentionally NOT shipped in this package; they resolve in the source vault, not here. In-package references are package-relative and resolve.

type spec
status canonical
version 1.1
supersedes
1.0
project snes

HOST DEPLOYMENT SPECIFICATION v1.1 - Tool-Neutral Workflow Deployment

CANONICAL (current). Supersedes v1.0 (2026-08-20). v1.1 removes the opencode-centric install model. The workflow must run on the destination's EXISTING, adaptable AI toolset; NOTHING is installed unless NO suitable tool exists - and then the report says what needs to be installed. The goal is that the governance workflow's observable invariants hold on the target machine, realized through whatever primitives the already-present toolset provides. The deploy itself follows the workflow's discipline: ANALYSIS first (free probe + report to the user), then PLANNING via the workflow's own plan gate, then IMPLEMENTATION only after user approval. Grounding: Research/v1.8-host-parity-plan.md, Research/client-side-feasibility.md. Parent contract: workflow-spec-v1.1.md (the governance this deploys; the Decision-making process section is sec 5). Format: Research/spec-format-specification.md (SFS).

1. Design Principles

# Principle Source
1 Feasibility is STATE-DECIDABLE - a deterministic probe decides, never an LLM's judgment. Real-Time Detection
2 Fail-closed reporting: no probe output = do not implement; blocked = MISSING REPORT, never a partial install. workflow-spec-v1.1.md section 7 (fail-closed)
3 Self-contained playbook: a FRESH agent with no prior knowledge of this system must be able to run it. SFS section 1 (self-contained)
4 Contracts are WHAT (observable invariants); mechanisms are HOW (per-toolset primitives) and interchangeable. The deploy payload is the portable engine, not a toolset-specific stack. session analysis 2026-08-20; verified: all *-core.ts import only node: builtins
5 USE what the destination already has. If an existing tool can express the required contracts with its own primitives, deploy onto it; install NOTHING. user directive 2026-08-20
6 Install only when NO suitable alternative exists - and then the MISSING REPORT says what needs to be installed (primitives + example toolsets), never a forced brand. user directive 2026-08-20
7 Verify by execution after install: the behavioral conformance vectors pass on the target, never a self-report. Verify-before-asserting (vault Rules)
8 The source vault is never mutated; the source is copied, never moved. G14 migration safety
9 Complete bring-up: the portable engine AND the vault content are both deliverables. user decision 2026-08-20
10 The MISSING REPORT uses the mandated escalation format: WHAT IS NEEDED -> WHY -> WHAT IT BLOCKS -> HOW TO OBTAIN IT. workflow-spec-v1.1.md section 11
11 ANALYSIS FIRST, then plan, then implement-if-approved: investigation (probe + report) is FREE; the deploy plan follows the workflow's OWN plan gate; acting tools that mutate the target commit only after user approval. workflow-spec-v1.1.md section 3 (plan gate); principle 20 (human oversight budget); user directive 2026-08-20
12 Toolsets are EXAMPLES, never requirements: the normative target is the portable capability contracts (WHAT); specific installed/configured elements appear only as illustrations. user directive 2026-08-20
13 The agent ALWAYS uses proper, systematic, OPTIMAL, correct, data-backed research at every analysis/research step: LOCAL (observed from the destination) and/or ONLINE (data-backed sources, scoped + cited), evidence saved - never anecdote, guess, or suboptimal shortcut. user directive 2026-08-20; Rules/Self-maintenance.md (research-gated)
14 The reference realization is an EXAMPLE, not a required payload: the portable engine, probe, and conformance suite are the opencode reference implementation. The deploy realizes the capability contracts (C-1..C-6) with whatever the destination can ACTUALLY run. The installing agent MUST adapt the reference realization or CREATE A WORKING ALTERNATIVE that satisfies the intended requirements when the reference is not runnable/compatible; only NO suitable realization blocks. user directive 2026-08-20; consistent with principle 12 (toolsets are examples)

Sources: inline citations above.

2. Architecture - the deploy loop

flowchart TB
    A1["1. ANALYSIS - portability-probe.ts (reference tool; where it cannot run, perform the DP checks manually per sec 23), DP-001..N, NON-mutating, detects installed toolsets (DP-016) + workspace scan (DP-019); investigation is FREE"]
    A1 --> A2["2. ANALYSIS REPORT - verdict, vectors, detected toolset, recommended path (research-backed)"]
    A2 --> P["3. PLAN - deploy plan via the workflow's own plan gate (valid todowrite, forward contract)"]
    P --> AP["4. APPROVAL - user approves the plan; no mutation before approval"]
    AP --> B["5. BRIDGE (partial) - thin bridge for the DETECTED toolset; never install a new tool"]
    B --> D["6. DEPLOY - realize capability contracts (reference engine only if runnable) + seed vault + wire bridge + apply"]
    D --> V["7. VERIFY - behavioral conformance vectors on the live target + idle pass"]
    V --> R["8. REPORT - completion report (verdict, vectors, adaptations, receipts)"]
    A2 -. blocked .-> M["MISSING REPORT - what needs installing; pause until user obtains + re-probes"]
    AP -. denied .-> STOP["Stop and report"]
    D -. failure .-> RB["Rollback + notify + remediation plan (N-5)"]
    V -. failure .-> RB2["Notify + remediation plan (N-6); no completion report"]
Loading

The probe is the trust root: the agent asserts nothing about the target that the probe did not observe (principle 1). Nothing is installed on the target except the portable engine + a bridge for a toolset that was ALREADY there (principles 5-6). The workflow's own discipline applies to the deploy itself: analysis is free, planning is gated, implementation commits only after approval (principle 11). Every phase transition persists its plan/research/implementation state to <target>/.governance/ before the next phase runs (deploy tracking & persistence, sec 8).

Sources: probe-over-judgment (Real-Time Detection); fail-closed (workflow-spec-v1.1.md section 7); migration safety (G14, migrate-g14.ts); escalation format (workflow-spec-v1.1.md section 11).

3. The deployment workflow loop - phase-anchored triggers

Phase boundary Fire point What happens
Task start user task matches "set up the governance workflow on this machine" Load this spec
Analysis task start Run portability-probe.ts --target <ws> --vault <src> (free, non-mutating); read-only workspace inspection; analysis conclusions follow the systematic research discipline (sec 6)
Analysis report probe JSON available Deliver the ANALYSIS REPORT to the user (verdict, vectors, detected toolset, recommended path; blocked -> what needs installing)
Feasibility review report delivered User reviews; blocked -> MISSING REPORT path (stop until the user obtains what is missing and re-probes)
Planning user approves the analysis Create the deploy plan following the WORKFLOW'S OWN plan gate (workflow-plan-spec-v1.0.md: valid todowrite, forward contract); present the plan
Approval plan validated User approves the plan; DENIED -> stop and report
Bridge verdict == partial, plan approved Write/install the thin bridge for the DETECTED toolset (sec 11); no new tool installed
Deploy plan approved Deploy engine -> npm install -> seed vault -> wire bridge -> configure -> dry-run -> apply
Verify deploy complete Behavioral conformance vectors pass on the LIVE destination toolset + one idle observation pass
Report verify complete (or any N-trigger) Report to the user; every issue report pairs with a proposed remediation plan (N-1..N-7, sec 21)

Fail-closed: if the probe cannot run or emits no valid JSON, the agent MUST NOT deploy and MUST report the probe failure as the blocker.

Sources: phase-anchored triggers (workflow-spec-v1.1.md section 3); deterministic recomputation (workflow-spec-v1.1.md section 7).

4. Intent routing policy

  • Task is "set up/install/deploy the governance workflow on this machine" -> run this playbook. Analysis (probe + report) is INVESTIGATION and is free.
  • Task asks about the workflow but does NOT request deployment -> prose, no probe.
  • Ambiguous -> ask the user with options.

Sources: nl2spec; ARTEMIS.

5. Prose->spec translation pipeline - NOT APPLICABLE

This spec IS the deployment contract. Deploy steps are executed as plan artifacts (todowrites) created in the Planning phase, not translated through the prose->spec pipeline (which is reference-only per v1.8-host sec 5).

Sources: source vault: prose-to-spec-translation.md (reference design only).

6. Enforcement contracts

Analysis phase is FREE: probing and read-only inspection of the target are investigation (enforcement is selective - commitment actions gated, investigation free; v1.8-host principle 3). No plan and no approval are required to probe and report.

Approval gate (commitment): acting tools that MUTATE the target are DENIED until ALL of: (a) a valid probe JSON exists, (b) a validated deploy plan exists (plan gate per workflow-plan-spec-v1.0.md, PL-001..010), AND (c) the user approved the plan. Fail-closed on any missing precondition. Re-arms per turn.

Self-commitment pre-engine (chicken-and-egg): the plan gate and approval gate are the workflow's OWN gates. On a destination where the engine is not yet live, the deploying agent SELF-COMMITS to them (prose discipline) until the engine is deployed and enforcing; afterwards they become deterministic. This is stated, not hidden - a fresh destination cannot be governed by machinery that is being installed.

Existing-tool gate: the deploy MUST target a toolset that was ALREADY present on the destination (probe DP-016). If the detected toolset is not opencode, a bridge is written for it; a new tool is installed ONLY when the MISSING REPORT path fires (no suitable existing tool). No toolset is ever installed to satisfy a bridge that could have been written.

Realization directive: when the reference realization cannot run on the destination (DP-002/DP-003 fail, or runtime incompatibility), the installing agent MUST adapt it or CREATE A WORKING ALTERNATIVE that satisfies the intended requirements - the capability contracts C-1..C-9, sec 11. A working alternative is one whose contracts pass behavioral verification on the LIVE target (minimal primitives present + negative controls); it is NEVER a blind copy of the reference code. The adaptation decision is RESEARCHED with the systematic, optimal, correct, data-backed discipline (LOCAL observation of the destination + ONLINE data-backed sources, scoped + cited), recorded, and reported - never guessed and never silently downgraded.

Deploy gate: steps are ordered (engine -> npm install -> seed -> bridge -> migrate) and each step's receipt AND state update must be persisted to disk before the next step runs (deploy tracking & persistence, sec 8). No step runs against the source vault (copies only).

MISSING REPORT gate (fail-closed): if the analysis verdict == blocked, the ONLY deliverable is the MISSING REPORT (schema in sec 21) delivered at the feasibility review. The report names the missing capabilities and what needs to be INSTALLED to obtain them (with example toolsets, never a forced brand). The deploy pauses until the user obtains what is missing and re-probes.

Verify-before-report: the completion report is not written until the behavioral conformance vectors pass on the LIVE destination toolset and the idle observation pass ran. A report without fresh execution evidence is denied (completion evidence contract, v1.8-host sec 6).

User notification + remediation proposal (issue contract): EVERY issue notification (triggers N-1..N-7, sec 21) is delivered WITH a proposed remediation plan. On an issue, the agent FIRST researches it systematically (investigation is free), THEN drafts a remediation plan following the workflow's OWN plan gate (workflow-plan-spec-v1.0.md: valid todowrite, forward contract, plan-spec PL-001..010), and presents ISSUE + PLAN together for user approval. NO remediation action runs before approval. On a pre-engine destination this discipline is self-committed (prose) until the engine is live - the approval is still the real gate.

Systematic analysis & research discipline (EVERY step, ALWAYS): the agent performs proper, systematic, OPTIMAL, correct, data-backed research at EVERY step of this playbook - Analysis phase, capability mapping for toolsets outside the detection set, adaptation decisions (adapt or create a working alternative), remediation planning (N-1..N-7), verification interpretation - following one discipline:

  • LOCAL research where the question touches the destination: the probe report, the destination's files/config/logs/state, and the installed toolset's actual behavior - observed, never assumed.
  • ONLINE research where external knowledge is needed: proper, thought-out web research from DATA-BACKED sources only (official toolset documentation, peer-reviewed/arXiv papers, benchmarks with measured effect sizes), every claim scoped and cited, per the vault's research-gated discipline (Rules/Self-maintenance.md; SFS section 7 Research-source conventions). ONLINE is MANDATORY, not optional, whenever external knowledge is needed or local evidence is insufficient - never answer an external-knowledge question from local inference alone. Use >= 2 independent sources, check authority + recency (superseding docs beat the most-relevant one), and grade retrieved evidence (relevant/authoritative/fresh) before use.
  • The mode is chosen per step: local for destination facts, online for external knowledge, BOTH where a recommendation depends on both (e.g. a remediation plan).
  • EVERY research step emits a RESEARCH RECEIPT (schema in sec 21) appended to the evidence chain BEFORE the next step runs. A claim in any report or plan is grounded in its research receipt - traceable local observations + linked, scoped online sources - never an unreceipted assertion.
  • Findings are SAVED as evidence (vault Research/ or the deploy evidence store) and cited in the outputs, so every conclusion is traceable to observed local facts and/or sourced online data - never anecdote or guesswork.

General gates (B1, referenced - never duplicated): the general gates G1/G2/G3 of workflow-spec-v1.1.md section 7 apply to the deploy's own artifact mutations - context-gain before, verify-by-execution after, research-mode recorded. This playbook references them, never re-defines them.

Sources: fail-closed and completion evidence (workflow-spec-v1.1.md section 7); verify-by-execution (vault Rules Verify-before-asserting).

7. Knowledge lifecycle - vault seeding

  • The SOURCE vault is never mutated: cp -R (robocopy on Windows) semantics, never move. G14 already proves no-lesson-loss on regeneration.
  • The TARGET vault is seeded with: Rules/, Lessons/, Research/, Templates/, Index.md. Missing source content is a DP-013 failure (see sec 23) reported in the MISSING REPORT, not silently skipped.
  • The target's AGENTS.md is produced by the idle regen (vault-core regenAgents), NOT copied from the source.
  • Provenance: the target vault records its origin (frontmatter or a Research/DEPLOYMENT.md note with source commit/reference + deploy date).
  • The vault logic is part of the portable engine (runs on ANY toolset as a standalone process, idle/cron triggered).

Sources: none - declarative scope statement.

8. Action receipts contract (deploy receipts)

  • The probe report IS the feasibility receipt (JSON, sec 21 schema, written to <target>/.governance/portability-report.json and stdout).
  • Each deploy step emits a receipt (tool + args hash + result hash); until the engine is live the agent's own tool receipts on the target serve as the step receipts.
  • The completion/MISSING report is a receipted-evidence record: its claims must match the probe JSON and verification outputs (tri-state verdict logic).

Deploy tracking & persistence (plan, research, implementation): the install plan, the research, and the implementation are all persisted to disk at each relevant step, so the deploy is resumable, auditable, and survives session boundaries. The deploy state lives under <target>/.governance/:

  • PLAN: the validated deploy plan (todowrite snapshot) is written to <target>/.governance/deploy-plan.json at the Planning phase, before approval; the approved plan is the committed artifact (not only in-context).
  • RESEARCH: every research step (systematic research discipline, sec 6) saves its findings - local observations + online sources, scoped and cited - to the evidence store (<target>/.governance/evidence/<step>.json or the vault Research/) AT THAT STEP, not batched at the end. Each is a RESEARCH RECEIPT (schema in sec 21), hash-chained with the deploy receipts (prev_hash linked), so the research is independently verifiable - a claim's grounding can be re-proven from the chain.
  • IMPLEMENTATION: each deploy step's receipt (tool, args hash, result hash, ts) is appended to <target>/.governance/deploy-receipts.jsonl when the step completes.
  • STATE: <target>/.governance/deploy-state.json is rewritten at every phase transition (unprobed -> analyzed -> planned -> approved -> bridging / deploying -> verifying -> done | reported, sec 21) with: phase, plan-snapshot ref, latest receipt seq, research-evidence refs, verdict, and a resume hint. On a re-run, the agent reads this file and continues from the recorded phase (the plan gate still applies).
  • Nothing that informed a decision lives only in the agent's context: every input to a report is on disk before the report is delivered.

Sources: none - declarative scope statement.

9. Monitoring contract

  • Post-deploy: run the behavioral conformance vectors against the LIVE destination toolset (WS/PL vectors; negative controls included). PARITY (v1.1): the ONLY shipped driver for these vectors is the opencode plugin harness (conformance.ts -> the opencode plugin test files). A non-opencode bridge has NO shipped verification driver; until one is written, a non-opencode deploy is verified by the deterministic engine tests plus manual behavior checks, and the conformance claim is restricted to the opencode path (never silently extended to a bridge without a driver).
  • Run verify-e2e.ts on the target (flap=0 over a scratch-copy pass) when the target DB is available (DP-010).
  • Record the baseline in the target's health.json BEFORE first idle pass so the before/after comparison is real (v1.8-host sec 9).

Sources: none - declarative scope statement.

10. Communication conventions (the reports)

Three report types, all following the escalation format (WHAT IS NEEDED -> WHY -> WHAT IT BLOCKS -> HOW TO OBTAIN IT). Any report that carries an ISSUE (triggers N-1..N-7, sec 21) MUST pair the issue with a proposed remediation plan (based on LOCAL observation + ONLINE data-backed research, spec-conformant, awaiting approval): the user is never notified of a problem without the solution path to approve. The plan cites its sources (probe/local evidence refs + linked online refs, sec 6).

  • ANALYSIS REPORT (first deliverable, free): delivered to the user BEFORE any planning or implementation. Contains: the probe JSON, the feasibility verdict, the detected toolset(s) + selected tool, and the recommended path. If blocked, it is the MISSING REPORT (below). It is investigation output, so it carries no commitment and requires no approval to produce.
  • MISSING REPORT (blocked case): the ONLY deliverable when no suitable existing toolset is present. For each failed REQUIRED vector: what is missing, why it matters, what it blocks, HOW TO OBTAIN IT - including WHAT NEEDS TO BE INSTALLED (required primitives + example toolsets, never a forced brand). PLUS the proposed remediation plan (steps to obtain what is missing and re-probe, per workflow-plan-spec-v1.0.md), presented for approval. PLUS degraded OPTIONAL capabilities the user should know about. The deploy pauses until the user obtains what is missing and re-probes.
  • Completion report: verdict, probe vectors (id/name/actual/pass), the detected toolset + bridge used, every adaptation taken, verification receipts (behavioral vectors + idle), and the next step (await user review).

Sources: none - declarative scope statement.

11. Adaptation profile + capability contracts

The deploy USES the destination's already-installed toolset. The NORMATIVE target is the portable CAPABILITY CONTRACTS below (WHAT). Specific toolsets appear ONLY as EXAMPLES of tools that typically provide a contract - a tool is never the requirement (principle 12). The probe embeds the same capability data in machine form, so feasibility stays deterministic (principle 1).

REQUIRED capability contracts (must be enforceable; their absence blocks):

# Capability contract (WHAT) Minimal primitive the tool must expose
C-1 Plan gate a PRE-TOOL hook that can DENY an acting tool with a structured reason
C-2 Commitment gate an environment-external PERMISSION/APPROVAL boundary for irreversible actions
C-3 Receipts access to tool-call args + results to emit a signed, hash-chained record
C-4 Completion evidence a SESSION-END signal to finalize a tri-state evidence verdict
C-5 Knowledge lifecycle the portable logic (promote/regen) + AN evidence feed supplied by the bridge
C-6 Vault seeding filesystem copy + regen (standalone)

OPTIONAL capability contracts (degradable without blocking):

# Capability contract Minimal primitive Degrades to
C-7 Constraint pinning system-prompt/context control under compaction pinned rules unenforced under compaction
C-8 Directive priority (PIM/SOM) input hook + output review path SOM-only
C-9 STD/STB re-injection per-call hook for turn/token counting re-inject on tool boundary

Examples - NOT normative (known toolsets that TYPICALLY provide a contract):

Capability opencode claude-code cursor copilot-cli
C-1 tool.execute.before PreToolUse none none
C-2 permission.ask permissions.json approval-mode (partial) none
C-3 tool.execute.after PostToolUse none none
C-4 event idle Stop none none
C-5 opencode DB feed (shipped) bridge-supplied bridge-supplied bridge-supplied
C-6 standalone standalone standalone standalone
C-7 experimental.session.compacting compaction hooks (partial) none none
C-8 chat.message + text.complete UserPromptSubmit (partial) none none
C-9 chat.params UserPromptSubmit counting (partial) none none

The rows above are ILLUSTRATIVE: the CONTRACT is the requirement. A tool that provides the primitive by any mechanism (or via a bridge the deploy writes) satisfies it. A toolset NOT in the detection set is not auto-evaluated by the probe; the agent must map its primitives to C-1..C-9 manually in the Analysis report.

PARITY (v1.1) - lifecycle evidence feed: the lifecycle LOGIC (promote/regen, vault-core.ts) is portable, but the shipped EVIDENCE FEED reads the opencode DB (self-maintain.ts: ~/.local/share/opencode/opencode.db). On a non-opencode destination the bridge MUST provide an evidence source (C-5), or the lifecycle degrades (promotion/revival run on empty evidence - recorded as an adaptation). C-5 means the logic AND a feed, not the logic alone.

Feasibility decision (deterministic, computed from the capability data)

  1. A toolset is SUITABLE iff it provides all six REQUIRED capability contracts (C-1..C-6), natively or via a bridge the deploy writes.
  2. implementable = a suitable EXISTING toolset with a SHIPPED bridge (examples: opencode) and no OPTIONAL degradation; conformance is claimed for this path only (the WS/PL driver is opencode-bound - see sec 9).
  3. partial = a suitable EXISTING toolset whose bridge must be written (examples: claude-code), and/or an OPTIONAL contract degrades. Adaptations recorded. NO new tool is installed in this path. Conformance for the bridge is NOT claimed until a verification driver for it is written.
  4. blocked = no suitable EXISTING toolset, or another REQUIRED vector fails. MISSING REPORT: WHAT NEEDS TO BE INSTALLED - e.g. "a tool whose native primitives provide C-1 (pre-tool deny), C-2 (permission boundary), C-3 (tool-result stream), C-4 (session-end signal); opencode and Claude Code are examples."
  5. Model differs from the source host -> recompute STD/STB + per-model rejection rates by calibration before safety-critical paths are trusted (v1.8-host sec 11); DP-015 is informational.
  6. Windows targets: session.idle never fires; the bridge must verify the re-arm uses session.status + {type:"idle"} (DP-001 + bridge wiring).
  7. Node < 22.6 blocks --experimental-strip-types: not adaptable (REQUIRED).
  8. Reference realization is an example (principle 14): the reference implementation (Node >= 22.6 + npm + the portable engine + the conformance harness) is the opencode EXAMPLE, not a requirement. If the destination cannot run it (DP-002/DP-003 fail), the deploy ADAPTS the reference or CREATES A WORKING ALTERNATIVE that satisfies C-1..C-9 in any compatible runtime, researched with the systematic, optimal, correct, data-backed discipline, and records the adaptation. A working alternative is accepted ONLY on behavioral verification of its contracts on the live target - the contracts, not a code copy, are the requirement.

Decision map (D7) - deploy feasibility:

flowchart TD
    S{suitable existing toolset?} -- no --> B[blocked - MISSING REPORT: what to install]
    S -- yes --> BS{bridge shipped?}
    BS -- yes --> IM[implementable]
    BS -- no --> PA[partial - write bridge + record adaptation]
Loading

PARITY (v1.1) - declared vs observed capability: suitability is DECLARED (examples table), not observed. The actual hook surface must be validated on the destination: opencode via sanity-load.mts post-deploy; for other toolsets NO capability probe is shipped (DP-018), so their suitability rests on the examples table until a probe exists. A bridge whose capabilities cannot be observed is recorded as an adaptation, never silently claimed.

Cross-reference (v1.1): the deploy decision stack (feasibility D7 + notification N-1..N-7, sec 12/sec 21) is the deploy-specific slice of the workflow's decision-making process - see workflow-spec-v1.1.md sec 5 (the cross-cutting view: who decides, the gated decision loop). Feasibility is decided by the deterministic probe (D7 above), never by the agent's description; the workflow's gate sequence (plan -> G1 context-gain -> G2 verify-by-execution) applies to every deploy mutation.

Sources: none - declarative scope statement.

12. Directive-priority enforcement

  • The user's task directive (set up the workflow) has highest priority; this playbook is the mechanism, not an override.
  • Never mutate the source vault to satisfy a target requirement; propose the change via the review queue instead.
  • Hard constraints (never touch the source vault, never bypass the probe, never install a toolset when a bridge would do) are in the gates, not prompt rules.

Sources: none - declarative scope statement.

13. Drift mitigations summary

Drift type Mitigation
Target/OS drift (Windows event gap) DP-001 + bridge wiring check
Toolset version drift (Node/toolset) DP-002/DP-004/DP-016 version bounds
Toolset-primitive drift (hook renamed/dropped) behavioral vectors verify LIVE behavior, not file presence
Source vault drift DP-013 structure check + copy-never-move
Model drift on target DP-015 recalibration step (sec 11)
Stale feasibility claim Probe JSON is the only evidence accepted (sec 6)

Sources: none - declarative scope statement.

14. Metric-improvement changes

Metric Change
Deploy success rate Deterministic feasibility gate before any mutation
Tool-agnostic coverage Capability matrix decides feasibility per toolset, no brand bias
Time-to-first-blocked-report Probe runs in seconds; no manual checklist
Issue-to-solution latency Every issue notification ships with a ready-for-approval remediation plan (N-1..N-7)
Verification honesty Behavioral vectors on the live target, not file-presence checks

Sources: none - declarative scope statement.

15. Spec self-compliance & health

  • Self-containment: a fresh agent must be able to run this playbook with only this document + the probe file + the source vault. If the agent must ask a question to proceed, this spec failed its own principle 3.
  • The DP vectors in sec 23 EQUAL the probe's checks, and the capability-contract examples in sec 11 EQUAL the capability data embedded in the probe (machine form). Any divergence is a spec-health failure (mirrors plan-spec sec 15).
  • Cost: the probe is one local process; the gate runs before npm install.
  • Spec core is collapsed: the always-loaded core is the pinned core + plan gate (per workflow-spec-v1.1.md section 7); this spec's own contracts are on-demand.

Sources: none - declarative scope statement.

16. Enforcement/trust root

The trust root of a deployment is the PROBE OUTPUT on the target, not the agent's description of the target. The feasibility gate reads the probe JSON file; it does not trust the agent's summary. (Parity with v1.8-host sec 16.)

Sources: none - declarative scope statement.

17. Optional multi-agent orchestration mode - NOT APPLICABLE

Single-agent deployment on the target (v1.8-host sec 17 scoped out).

Sources: none - declarative scope statement.

18. Dynamic cascade controller - NOT APPLICABLE

No cascade on the target (v1.8-host sec 18 scoped out).

Sources: none - declarative scope statement.

19. Glossary

  • Target machine: the host receiving the deploy (workspace + already- installed AI toolset analyzed by the probe).
  • Source vault: the immutable reference vault this spec ships with.
  • Portable engine: the toolset-agnostic implementation (cores, tests, vault logic, stores) that imports only node: builtins; the deploy payload.
  • Bridge: the thin per-toolset translator that maps a toolset's own primitives (hooks/permissions/events) to the portable engine. opencode's .opencode/plugins/ factories are the shipped opencode bridge.
  • Probe / portability-probe.ts: the dependency-free deterministic script that checks DP vectors, applies the capability-contract examples, and emits the feasibility JSON.
  • DP vector: a numbered, state-decidable prerequisite check (DP-001..N); each is REQUIRED or OPTIONAL.
  • Capability contract: a portable, tool-neutral WHAT requirement (C-1..C-9, sec 11); a toolset satisfies one by providing the primitive through any mechanism.
  • Reference realization: the opencode EXAMPLE implementation of the capability contracts (portable engine, probe, conformance harness). An example, not a requirement. When it cannot run on the destination, the installing agent MUST adapt it or create a working alternative that satisfies the intended requirements (a working alternative passes behavioral verification of its contracts on the live target - never a blind code copy).
  • Suitable toolset: an already-installed toolset that provides all six REQUIRED capability contracts (C-1..C-6).
  • Feasibility verdict: blocked (no suitable existing toolset, or another REQUIRED failure -> MISSING REPORT: what to install), partial (suitable existing toolset but a bridge must be written and/or an OPTIONAL contract degrades), implementable (suitable existing toolset with a shipped bridge).
  • Adaptation: a documented, allowed deviation triggered by a to-be-written bridge or an OPTIONAL failure (never a REQUIRED failure, never an installed toolset).
  • MISSING REPORT: the fail-closed blocked-case deliverable (produced in the Analysis phase); uses WHAT IS NEEDED -> WHY -> WHAT IT BLOCKS -> HOW TO OBTAIN IT, and, when no suitable tool exists, names WHAT NEEDS TO BE INSTALLED.
  • Analysis report: the first, free deliverable (probe JSON + verdict + detected toolset + recommended path); investigation output, requires no plan or approval.
  • Approval gate: the commitment point at which the user approves the deploy plan; no target-mutating acting tool runs before it.
  • Completion report: the success-case deliverable; includes the detected toolset + bridge used and verification receipts.
  • Deploy receipt: the probe JSON (feasibility) + step receipts + final report; the deploy's evidence chain.
  • Deploy state file: <target>/.governance/deploy-state.json - the persisted phase, plan-snapshot ref, latest receipt seq, research-evidence refs, verdict, and resume hint, rewritten at every phase transition (sec 8).
  • Research receipt: the hash-chained evidence record every research step emits (schema in sec 21): question, method (local/online), local evidence, cited/scoped online sources, findings, used-in, prev_hash. A claim is grounded in its research receipt - never an unreceipted assertion.
  • Inexpressible contract: a REQUIRED contract with NO primitive in the detected toolset; its presence means that toolset is not suitable and, if no other suitable tool exists, escalates to blocked.
  • Bridge driver: the shipped verification harness for a bridge. The only shipped driver is the opencode plugin harness (conformance.ts); a non-opencode bridge needs its own driver before conformance is claimed.
  • Capability probe: an observation of the destination toolset's actual hook surface (opencode: sanity-load.mts). Not shipped for other toolsets - their suitability is declared from the static matrix, not observed.
  • Issue notification: a user notification raised on trigger N-1..N-7 (sec 21); always paired with a proposed remediation plan.
  • Remediation plan: the local-observation + online-data-backed, spec-conformant plan (validated todowrite per workflow-plan-spec-v1.0.md) proposed to the user with an issue notification; no remediation action runs before approval.
  • Local research: investigation of the destination itself - probe report, files, config, logs, installed toolset behavior; observed, never assumed.
  • Online research: proper web research from DATA-BACKED sources only (official toolset docs, peer-reviewed/arXiv papers, benchmarks with measured effect sizes), every claim scoped and cited per SFS section 7 (Research-source conventions).
  • Data-backed source: a verifiable, linked, HTTP-checked, scoped reference - never an anecdote or a guess.

Sources: none - declarative scope statement.

20. Benchmark-scoping discipline

No benchmarks are cited. Feasibility is deterministic per target; effect sizes would be meaningless.

Sources: none - declarative scope statement.

21. Contracts, schemas & state machines

Probe report schema (JSON) - the feasibility receipt

{
  "spec": "host-deployment-spec-v1.1",
  "ts": "ISO-8601",
  "target": "string (workspace path probed)",
  "vault": "string (source vault path probed)",
  "os": { "platform": "win32|linux|darwin", "version": "string" },
  "toolset": {
    "inventory": [ { "name": "opencode|claude|cursor|copilot", "version": "string" } ],
    "suitable": "string|null (first suitable EXISTING toolset)",
    "suitable_bridge_shipped": true,
    "expressibility": [
      { "toolset": "string", "contract": "string (required contract)",
        "expressible": true, "primitive": "string" }
    ]
  },
  "vectors": [
    {
      "id": "DP-001",
      "name": "string",
      "required": true,
      "pass": true,
      "actual": "string (observed value)",
      "expected": "string (predicate satisfied / minimum)"
    }
  ],
  "verdict": "blocked|partial|implementable",
  "adaptations": ["string (only when partial)"],
  "missing": [
    { "what": "string", "why": "string", "blocks": "string", "how": "string" }
  ]
}

MISSING REPORT schema

{
  "type": "missing-report",
  "target": "string",
  "probe_ts": "ISO-8601",
  "blockers": [
    { "vector": "DP-00N", "what": "string", "why": "string",
      "blocks": "string", "how": "string" }
  ],
  "install_requirements": [
    { "primitive": "pre-tool hook|permission boundary|tool-result stream|session-end signal",
      "needed_for": "string (contract)", "example_toolsets": ["opencode", "claude-code"] }
  ],
  "degraded_optional": [
    { "vector": "DP-00N", "what": "string", "impact": "string", "how": "string" }
  ],
  "inexpressible_contracts": [
    { "contract": "string (workflow-spec section 7 contract name)",
      "available_primitive": "string|null (partial primitive, if any)",
      "impact": "string", "how": "string" }
  ],
  "proposed_plan": [ "string (remediation steps, research-backed, per plan-spec v1.0)" ],
  "no_partial_install": true
}

Research receipt schema (JSON)

Emitted by EVERY research step (sec 6); hash-chained with the deploy receipts (prev_hash linked). A claim in any report/plan is grounded in its research receipt - re-provable from the chain.

{
  "id": "string (uuid or sha256-truncated)",
  "kind": "research",
  "session_id": "string",
  "seq": "int (monotonic in the evidence chain)",
  "ts": "ISO-8601",
  "question": "string (what was researched)",
  "method": ["local", "online"],
  "local_evidence": ["string (paths/files/configs/logs/toolset-behavior observed)"],
  "online_sources": [
    { "url": "string", "title": "string", "scoped": "string (domain/model/caveat)" }
  ],
  "findings": "string (summary)",
  "used_in": "string|null (analysis | remediation N-x | adaptation | verify)",
  "prev_hash": "sha256",
  "signature": "ed25519|null (out-of-process key when available)"
}

User notification triggers + remediation (N-1..N-7)

An ISSUE is any condition below. The agent notifies the user AND proposes a remediation plan together (research first, plan per the workflow's plan gate, then user approval - sec 6). No remediation action runs before approval.

ID Trigger When Content Work pauses?
N-1 Probe failure / no valid JSON analysis probe failure + remediation plan (fix probe env, re-probe) YES - no deploy
N-2 Analysis issue (failed REQUIRED/OPTIONAL vector) analysis report issue + adaptations/install needs + remediation plan YES if blocked; else proceeds to plan approval
N-3 Blocked (no suitable tool / REQUIRED fail) feasibility review MISSING REPORT (WHAT/WHY/BLOCKS/HOW + what to install) + remediation plan (obtain + re-probe) YES - pauses until the user acts
N-4 Plan approval (normal flow, not an issue) planning the validated deploy plan YES - no mutation before approval
N-5 Deploy step failure deploy rollback confirmation + remediation plan (corrected steps) YES - rollback then pause
N-6 Verify failure verify failing vectors + remediation plan (fix bridge / re-verify) YES - no completion report
N-7 Completion with residual issues/adaptations report completion report + follow-up plan for residual items NO - deliverable issued

Every issue row (N-1, N-2-blocked, N-3, N-5, N-6, N-7-residual) delivers: issue (WHAT/WHY/BLOCKS/HOW) + the researched remediation plan + the approval request, in ONE notification.

Decision map (D8) - notification trigger:

flowchart TD
    I{issue?} -- N-1 probe fail --> P1[notify + remediation plan - pause]
    I -- N-2 analysis issue --> P2[notify + adaptations - pause if blocked]
    I -- N-3 blocked --> P3[MISSING REPORT + plan - pause]
    I -- N-4 plan ready --> P4[present plan - await approval]
    I -- N-5 deploy fail --> P5[rollback + notify + plan]
    I -- N-6 verify fail --> P6[notify + plan - no report]
    I -- N-7 residual --> P7[report + follow-up plan]
Loading

Feasibility state machine

States: unprobed -> analyzed -> planned -> approved -> bridging -> deploying -> verifying -> done (or blocked -> reported).

  • unprobed -> analyzed: probe emits valid JSON (free investigation).
  • analyzed -> blocked: verdict == blocked (MISSING REPORT delivered at feasibility review; deploy pauses until the user obtains what is missing).
  • analyzed -> planned: user approves the analysis; deploy plan created per the workflow's plan gate (workflow-plan-spec-v1.0.md).
  • planned -> approved: plan validated by the plan gate AND user approves it.
  • approved -> bridging: verdict == partial (bridge for the DETECTED suitable toolset per sec 11; never an installed tool).
  • bridging -> deploying: bridge written/installed for the detected toolset.
  • approved -> deploying: verdict == implementable (shipped bridge, e.g. opencode).
  • deploying -> verifying: deploy steps receipted.
  • verifying -> done: behavioral vectors pass on the LIVE target + idle pass ran. Fail-closed: probe error or malformed JSON -> blocked with a probe-failure blocker; no mutation, no deployment without analysis + plan + approval. Every transition rewrites <target>/.governance/deploy-state.json (sec 8).
stateDiagram-v2
    [*] --> unprobed
    unprobed --> analyzed: probe emits valid JSON (free)
    unprobed --> blocked: probe failure / malformed JSON
    analyzed --> blocked: verdict == blocked
    analyzed --> planned: user approves analysis
    planned --> approved: plan gate validates AND user approves
    approved --> bridging: verdict == partial
    approved --> deploying: verdict == implementable
    bridging --> deploying: bridge written for the detected toolset
    deploying --> verifying: deploy steps receipted
    verifying --> done: behavioral vectors pass + idle pass
    blocked --> reported: MISSING REPORT (only deliverable)
Loading

Install (deploy) state machine

Steps, each receipted AND state-persisted before the next:

  1. realize-engine - realize the capability contracts for the destination's runtime. USE the reference engine (cores + package.json/tsconfig.json) ONLY when the destination can run it (DP-002/DP-003 pass); otherwise ADAPT the reference or CREATE A WORKING ALTERNATIVE that satisfies C-1..C-9, verified behaviorally on the live target, and record the adaptation. Deploy into <target>/.governance/ (or the toolset's native extension dir when one exists, e.g. .opencode/).
  2. npm-install - npm install in the engine dir ONLY when the reference engine is used (DP-002/DP-003 pass); needs DP-012 or vendored node_modules.
  3. seed-vault - copy Rules/, Lessons/, Research/, Templates/, Index.md to the target vault dir (DP-007/DP-013 source complete).
  4. wire-bridge - register the bridge in the DETECTED toolset's own config (opencode: legacy .opencode/plugins/ auto-discovery or config file; claude-code: settings.json hooks; etc.); set the vault path.
  5. dry-run-migrate - npm run migrate:dry (G14: scratch regen, no-lesson- lost, budget).
  6. apply-migrate - npm run migrate:apply only after dry-run passes.
  7. verify - run the behavioral conformance vectors against the LIVE toolset (reference harness when runnable; adapted behavioral checks otherwise), one idle observation pass (DP-010 available), baseline in health.json BEFORE first idle pass.
flowchart LR
    E["1 deploy-engine"] --> N["2 npm-install"]
    N --> S["3 seed-vault"]
    S --> W["4 wire-bridge"]
    W --> DR["5 dry-run-migrate"]
    DR --> AP["6 apply-migrate"]
    AP --> VF["7 verify"]
    VF --> OK[done]
    AP -. failure .-> RB["Rollback + notify (N-5)"]
    VF -. failure .-> RB2["Rollback + notify (N-6)"]
Loading

Rollback: before step 1, snapshot the engine dir (.governance.bak or the toolset-native extension dir) if it already existed; on ANY deploy/verify failure, restore the snapshot AND the pre-install AGENTS.md (G14 legacy backup), then report. No engine change is left behind on failure.

Sources: none - declarative scope statement.

22. Privacy lifecycle contract - PROBE SCOPE

  • The probe NEVER reads config files for secrets; it checks existence/ accessibility/writability only, and reports booleans or safe values, never contents.
  • The probe report must not embed credentials, tokens, or private key material (report contains only paths, versions, booleans).
  • The MISSING REPORT tells the user WHAT NEEDS TO BE INSTALLED and HOW to obtain it; it never fetches or carries it.
  • The deploy writes no secrets to the target outside the engine's own config.

Sources: none - declarative scope statement.

23. Executable conformance standard - the DP vectors

Harness contract: portability-probe.ts is dependency-free (node:builtins only) so it can run BEFORE npm install. It is the REFERENCE analysis tool; where Node is unavailable on the destination, the agent performs the DP checks manually per this table (recorded as an adapted analysis). It applies the capability contracts (sec 11; known toolsets are examples) to the DETECTED toolsets, emits the sec-21 JSON to stdout, and writes portability-report.json into the target. Exit 0 iff a valid report was produced; any other exit = probe failure (fail-closed).

ID Required Check Expected (pass predicate)
DP-001 no OS platform win32/linux/darwin detected (Windows: idle-event wiring note)
DP-002 no Node runtime version >= 22.6 (REQUIRED only for the REFERENCE engine realization; fail = adapt to another runtime)
DP-003 no npm npm --version succeeds (REQUIRED only for the reference engine install; fail = adapt)
DP-004 YES SUITABLE existing toolset >= 1 already-installed tool provides all 6 REQUIRED capability contracts (C-1..C-6; known toolsets are examples). None -> blocked: report what to install
DP-005 YES target workspace writable fs.access W_OK on target dir
DP-006 YES engine deploy dir creatable <target>/.governance/ (or toolset-native extension dir) absent (creatable) OR present (overwrite flagged)
DP-007 YES source vault exists vault dir present at given path
DP-008 YES target vault creatable target vault dir creatable OR a vault path env (e.g. OPCODE_VAULT) set
DP-009 no toolset config writable the detected toolset's config (opencode.json / settings.json / equivalent) writable
DP-010 no toolset DB readable opencode.db (or a toolset-equivalent evidence store) read-access (self-maintain/verify-e2e)
DP-011 no out-of-process signing key dir creatable keys dir creatable under the engine dir (Phase 2 receipts)
DP-012 no registry reachable fetch https://registry.npmjs.org within 3s (else vendor node_modules)
DP-013 YES source vault structure complete Rules/, Lessons/, Research/, Templates/, Index.md all present
DP-014 no git available git --version succeeds (repo context)
DP-015 no model identity informational - reported, used for STD/STB recalibration (sec 11)
DP-016 no toolset inventory detected toolsets + versions reported (input to the matrix)
DP-017 no selected suitable toolset the EXISTING toolset chosen for the bridge (informational; NOT a verdict driver)
DP-018 no selected-toolset capability OBSERVABLE opencode: sanity-load.mts validates post-deploy; claude/cursor/copilot: NO capability probe shipped - suitability rests on the static matrix (declared, not observed)
DP-019 no workspace conflict scan reports existing .opencode/, .governance/, .claude/, AGENTS.md at the target (overwrite risk; investigation only)

REQUIRED set = DP-004, 005, 006, 007, 008, 013. verdict = blocked if any REQUIRED fails (DP-004 fails when NO suitable existing toolset is present - MISSING REPORT names what to install); else partial if the suitable toolset's bridge must be written (non-opencode) OR an OPTIONAL vector fails (incl. DP-002/DP-003, which are reference-path-only and NEVER block - a fail means an adapted realization); else implementable. The verdict NEVER depends on opencode presence. DP-018/DP-019 are diagnostic (optional). DP-018 is true only when the selected toolset's capability surface can be OBSERVED (opencode via sanity-load.mts); for other toolsets it is false and the deploy records it as an adaptation - never a silent claim.

Conformance: a target is "deployable" iff the probe on that target returns verdict != blocked AND the deploy state machine (sec 21) completes through verify. Determinism: re-running the probe on an unchanged target must produce the same verdict.

Sources: none - declarative scope statement.

24. Evidence base

  • workflow-spec-v1.1.md (the governance this deploy delivers; section 7/11/12/17 contracts referenced above).
  • Authoring provenance (resolves in the source vault, not shipped here):
    • Research/workflow-spec-v1.8-host.md (the operative host realization this derives from),
    • Research/v1.8-host-parity-plan.md (implementation reality the deploy must reproduce),
    • Research/client-side-feasibility.md (verified opencode hook surface; the shipped bridge precedent),
    • Research/host-deployment-spec-v1.0.md (superseded opencode-centric version; retained as history),
    • Research/plan-gap-check.md (G14 migration safety; G10 baselines),
    • .opencode/portability-probe.ts (the reference probe; DP vectors + matrix),
    • .opencode/plugins/*.ts (the reference opencode bridge),
    • Live implementation .opencode/ (the reference engine; cores import only node: builtins).
  • Analysis-first / plan-then-approve / implement-on-approval structure: user directive 2026-08-20; grounded in the workflow's own plan gate (workflow-plan-spec-v1.0.md) and selective-enforcement principle (workflow-spec-v1.1.md principle 3).

Sources: none - declarative scope statement.

SPEC FORMAT SPECIFICATION (SFS) v1.1

Meta-specification. Compiled 2026-08-13; updated 2026-08-21 (v1.1: added the extension-sections rule - a spec may add numbered extension sections after the mandatory 24). Defines the required structure, conventions, and validation rules for every workflow spec document in this vault (e.g., workflow-spec-v1.7.md). Other spec docs MUST follow this format so they are self-contained, machine-checkable, and markdown-compatible. Status: ACTIVE. Applies to the current canonical spec and all future versions.

1. Purpose

This document standardizes how spec documents are written so that:

  1. Every spec is self-contained (all contracts inlined, no cross-file pointers as the contract).
  2. Every spec is machine-checkable (schemas, state machines, and conformance vectors are explicit).
  3. Every spec is markdown-compatible (pure ASCII, standard markdown; Mermaid diagrams render in CommonMark markdown - Mermaid support since 2022-02-28, github.blog changelog).
  4. Every spec carries verifiable, scoped research sources for each claim.
  5. Every spec complies with its own rules (collapsed core, small always- loaded section, no circular references).

2. Mandatory structure

A spec document MUST contain, in this order:

01  YAML frontmatter          (type, status, version, supersedes, project)
02  H1 title + canonical/superseded notice
03  ## 1. Design Principles   (numbered table: # | Principle | Source)
04  ## 2. Architecture        (mermaid diagram: ```mermaid)
05  ## 3. Workflow loop        (phase table: Phase | Fire point | What happens)
06  ## 4. Intent routing policy
07  ## 5. Prose->spec pipeline
08  ## 6. Enforcement contracts   (self-contained, defined terms)
09  ## 7. Knowledge lifecycle     (labeled store contract)
10  ## 8. Action receipts contract
11  ## 9. Monitoring contract
12  ## 10. Communication conventions
13  ## 11. Adaptation profile
14  ## 12. Directive-priority enforcement
15  ## 13. Drift mitigations summary
16  ## 14. Metric-improvement changes
17  ## 15. Spec self-compliance & health
18  ## 16. Enforcement/trust root
19  ## 17. Optional multi-agent orchestration mode
20  ## 18. Dynamic cascade controller
21  ## 19. Glossary (definitions)      (MANDATORY - every specialized term defined)
22  ## 20. Benchmark-scoping discipline (MANDATORY - scoped, not universal)
23  ## 21. Contracts, schemas & state machines (MANDATORY)
24  ## 22. Privacy lifecycle contract     (if the spec touches persistence)
25  ## 23. Executable conformance standard (MANDATORY)
26  ## 24. Evidence base

Sections 21-23 (schemas, state machines, conformance) are mandatory in any spec that claims to be enforceable. Sections that do not apply to a spec's scope MUST be stated as "Not applicable to this spec" rather than omitted silently.

Extension sections (v1.1): the 24 numbered sections above are the MANDATORY core and stay in that order. A spec MAY add numbered extension sections AFTER section 24 (or at a documented logical position) when a new top-level concern warrants it. Each extension section MUST follow all SFS conventions (Sources paragraph, glossary terms defined before use, pure ASCII, mermaid for diagrams, no bare secN) and MUST be listed in the section inventory. An extension section is never a silent renumber of the mandatory core - the mandatory 24 keep their numbers.

3. YAML frontmatter conventions

Every spec file begins with YAML frontmatter:

---
type: spec
status: canonical | superseded
version: "1.7"
supersedes: ["1.0", ..., "1.6"]   # only on canonical
superseded_by: "1.8"             # only on superseded
project: snes
---

Rules:

  • status: canonical appears ONLY on the current operative version.
  • A superseded spec carries superseded_by pointing to its immediate successor.
  • A canonical spec carries supersedes listing all prior versions.
  • The canonical file is referenced by later specs as the operative document; prior versions are archived history, never cited as the operative contract.

4. H1 title + status notice

# AGENT GOVERNANCE & WORKFLOW SPECIFICATION v1.7

> **CANONICAL (current).** Status: `supersedes` v1.0-v1.6. [one-line change
> summary]. Grounding: `Research/<source>.md`. Greenfield design; all prior
> versions retained as archived history.

For superseded files:

# AGENT GOVERNANCE & WORKFLOW SPECIFICATION v1.6

> **SUPERSEDED** by v1.7 (`workflow-spec-v1.7.md`). Retained as archived
> revision history. Do not treat as current.

The status notice MUST be the first content after the H1 (not before it).

5. Text encoding rules (markdown compatibility)

  • Pure ASCII only. No em-dashes (use -), no minus signs (use -), no arrows (use ->), no box-drawing characters, no section sign (write sec), no <= from U+2264 (write <= ASCII), no x from U+00D7 (write x).
  • Mermaid is ALLOWED and PREFERRED for diagrams. CommonMark markdown renders Mermaid since 2022-02-28 (github.blog changelog, verified 2026-08-20); use ```mermaid fences for all diagrams. ```text is reserved for non-diagram listings only; schemas use ```json.
  • No BOM at file start.
  • LF line endings (no CR).
  • UTF-8 encoding, validated (a fatal-decode must pass).
  • Rationale: a Windows clipboard / CP1252 paste path mangles non-ASCII into bytes CommonMark renderers flag as corrupt. ASCII-only survives every transport.

6. Markdown conventions

  • Headings: ## N. Title numbered sequentially; H1 only for the document title.
  • Tables: standard pipe tables with a header separator row. Every table cell is single-line where possible.
  • Code blocks: always tagged with a language: ```mermaid for diagrams (flowcharts, stateDiagram-v2), ```json for schemas, ```text only for non-diagram listings, ```yaml for frontmatter examples.
  • Cross-references: use markdown anchor links to headings ([sec 18](#18-dynamic-cascade-controller)), NOT bare sec18 text and NOT self-referential anchors (a section must never link to itself).
  • Inline code: backticks for identifiers, file names, tool names.
  • Bold for contract names and key terms.
  • Sources: every section ends with a **Sources:** paragraph of linked citations.

7. Research-source conventions

  • Every principle, contract, and claim MUST carry a linked, verifiable source.
  • Scoped evidence discipline: benchmark numbers are annotated with their domain/model/caveat and marked "(scoped: ...)". A spec never presents a single-benchmark effect size as a universal law. If the source paper itself flags a result as non-robust, the spec repeats that caveat.
  • Link verification: external URLs are verified (HTTP status + title match) before inclusion; a fragile link is replaced with a stable one or cited by title/venue without a link.
  • Vault-internal references use relative paths (Research/name.md).

8. Schemas & state machines conventions

  • JSON schemas are presented in ```json blocks with field name, type, and required/optional semantics per field.
  • State machines are presented as a ```mermaid stateDiagram-v2 diagram PLUS the textual list of states, transitions (from -> to, on trigger), and fail-closed behavior. Example:
stateDiagram-v2
    [*] --> candidate
    candidate --> permanent: 2x evidence
    permanent --> archived: budget overflow
    archived --> permanent: evidence AFTER archivedAt
    candidate --> retired: older than RETIRE_DAYS
Loading
  • Every state machine and schema is referenced from the contract that uses it via anchor link, and each defines its inputs and outputs explicitly.

9. Conformance conventions

  • Every spec claiming enforceability MUST include an Executable Conformance Standard section: a table of deterministic test vectors (ID | Contract | Test | Expected), conformance levels, and a runtime-agnostic harness contract.
  • Vectors are PASS/FAIL; a FAIL vector is one that must be rejected with a specified failure reason.
  • Conformance claims require byte-identical deterministic output.

10. Self-compliance rules

  • Collapsed core: the spec must state which of its own contracts compress into the ~5 always-loaded rules and which are on-demand. A spec that exceeds its own compliance ceiling fails its own principle 5.
  • No circular references: no section links to itself; cross-version references point to actual files, not to the current section.
  • No undefined terms: every specialized term appears in the Glossary before or at first use.

11. Validation checklist (before publishing a spec)

Run this checklist; a spec is not complete until all pass:

[ ] YAML frontmatter: type=spec, status, version, supersedes/superseded_by
[ ] H1 + status notice first
[ ] All 24 mandatory sections present (or explicitly N/A); extension sections documented
[ ] Pure ASCII: 0 non-ASCII code points
[ ] UTF-8 fatal decode passes; no BOM; no CR; no NUL
[ ] Mermaid fences used for all diagrams (no ASCII/text diagram blocks)
[ ] Code fences balanced (even count); every ```mermaid block is valid syntax
[ ] No bare 'secN' text; all cross-refs are anchor links
[ ] No section links to itself
[ ] Every section has a **Sources:** paragraph
[ ] Every benchmark claim carries a (scoped: ...) caveat
[ ] Glossary defines every specialized term used
[ ] Schemas + state machines present (sec 21)
[ ] Executable conformance standard present (sec 23)
[ ] All external links return success + title-match
[ ] Spec-health: collapsed core stated (sec 15)

12. Versioning

  • A new canonical version supersedes the prior; the prior's frontmatter flips to status: superseded with superseded_by: "<new>".
  • Versions are sequential (v1.0 -> v1.1 -> ...); no overwriting of an archived version's content after it is superseded.
  • Each version's file is workflow-spec-vX.Y.md; the file name encodes the version.

13. Sources

  • Derived from the revision practice of workflow-spec-v1.0 through v1.7 and the failures identified in Research/spec-v1.6-weaknesses.md.
  • Mermaid-in-markdown support: CommonMark markdown renders Mermaid since 2022-02-28 (github.blog changelog, verified 2026-08-20). Prior versions of this spec banned Mermaid on the outdated premise that the renderer did not support it.
  • Encoding: Windows clipboard/CP1252 transport (session-observed).
type spec
status canonical
version 1.0
supersedes
project snes

PLAN SPECIFICATION v1.0 (spec-based plan artifact)

CANONICAL (contract). Defines the machine-checkable Plan artifact that the workflow requires before acting tools (bash/edit/write/task). Parent contract: workflow-spec-v1.0.md (Plan gate, Plan definition, enforcement-loop state machine). This spec makes the plan FORMAT spec-based and ties the gate's validation to conformance vectors, so the documented plan and the enforced plan are the same thing.

Format: follows spec-format-specification.md (SFS). Pure ASCII; Mermaid is allowed per the SFS (CommonMark renders it since 2022-02-28); all 24 sections present or explicitly N/A.

1. Design Principles

# Principle Source
1 Plan before act: a plan must exist before any acting tool (bash/edit/write/task) executes. workflow-spec-v1.0.md (Plan gate)
2 The plan is an artifact, not an intent: it is a machine-checkable step list, not a paragraph. workflow-spec-v1.0.md (Plan definition)
3 Enforcement outside the model: the gate decides, the model cannot overrule it. workflow-spec-v1.0.md principle 4
4 The format is enforced, not just stated: the gate validates structure, not the model's judgment. Rules/Tracked-plan-mandatory.md (vault)
5 A plan is a forward contract: completed/cancelled steps do not satisfy it. plan-enforce.ts validatePlan
6 The spec and the gate must agree: conformance vectors are the gate's rejection logic, verbatim. SFS section 9

Sources: none - declarative scope statement.

2. Architecture - the plan in the workflow

flowchart LR
    A[routing] --> B[planning]
    B --> C["gated: plan gate - valid todowrite + Plan artifact schema"]
    C --> D[verifying: action + receipt]
    D --> E[completing: canary self-check]
    E --> F[idle: plan-gate re-arm]
    F --> A
Loading

The plan lives at the planning -> gated transition (the plan gate): the artifact is a todowrite call with a valid step list (Plan artifact schema, sec 21); enforcement is plan-enforce.ts validatePlan (external to the model); the canary logs gate state (silent-dead gate = critical).

Sources: workflow-spec-v1.0.md (workflow table, state machine).

3. The workflow loop - plan placement

Phase boundary Fire point What happens
Turn start chat.message Intent routing + PIM
Plan todowrite (via tool.execute.before) Plan gate: valid Plan artifact arms; empty/placeholder rejected
Pre-action tool.execute.before Acting tool denied if no valid Plan (re-arms on idle)
Idle event session.status idle Read-gate + plan-gate re-arm

Sources: workflow-spec-v1.0.md (workflow table).

4. Intent routing policy (plan scope)

  • Must-hold / actionable -> a Plan artifact is required before acting.
  • Ambiguous -> ask the user (no plan machinery).
  • Everything else -> prose, no plan. Routing itself is out of this spec's scope; the plan requirement applies to the actionable branch only.

Sources: workflow-spec-v1.0.md (intent routing).

5. Prose->spec translation (how a plan is produced)

A Plan artifact is produced directly as a structured todowrite, not via the prose-to-spec pipeline (which is for long-lived rules). The plan must be expressible in the Plan artifact schema (sec 21); a task that cannot be expressed as steps must be split or asked about.

Sources: source vault: prose-to-spec-translation.md (the pipeline is for long-lived rules; plan artifacts are produced directly).

6. Enforcement contracts - the Plan gate

Plan gate: acting tools (bash/edit/write/task) are DENIED until a Plan artifact (a todowrite whose todos pass validation) exists. Re-arms each turn on idle. Host mechanism: tool.execute.before in plan-enforce.ts. The gate rejects a plan when validatePlan returns a non-null reason:

  • empty (no todos) - PL-003
  • no actionable step (all completed/cancelled) - PL-004
  • placeholder-only content (todo/tbd/placeholder/fill/doing/later/maybe/ something/could, or content < 8 chars) - PL-005/PL-006
  • status not in pending|in_progress|completed|cancelled - PL-004
  • priority not in high|medium|low - PL-010

Read-evidence requirement (PL-009): beyond a valid plan, the gate requires HOOK-OBSERVED evidence that the agent READ the plan spec (workflow-plan-spec-v1.0.md) this session, before acting tools are allowed. A read of the spec stamps a per-session mark (observed via tool.execute.after); a write/edit of the spec invalidates the mark (re-read required). Without the mark the acting tool is denied with a corrective hint (Read-Gate / deer-flow version-gate pattern). Fail-closed: no mark = deny.

On rejection the gate escalates a review proposal and the agent must supply a valid Plan (and, for PL-009, a read of the spec) before retrying the acting tool.

Sources: workflow-spec-v1.0.md (enforcement contracts); plan-enforce.ts validatePlan; Read-Gate arXiv:2608.02011; deer-flow read-before-write middleware (github.com/bytedance/deer-flow).

7. Knowledge lifecycle (plan-specific)

A Plan artifact is transient per-turn state, not persisted knowledge. It does not enter the vault. (The lessons about plan enforcement are vault knowledge, governed by the workflow-spec-v1.0.md knowledge lifecycle, out of scope here.)

Sources: workflow-spec-v1.0.md section 7 (knowledge lifecycle, out of scope here).

8. Action receipts contract (plan-related)

The plan gate's BLOCKED / REJECTED events are canary-logged and escalated to the review queue. Full step receipts are governed by the workflow receipts contract, out of scope here.

Sources: none - declarative scope statement.

9. Monitoring contract

  • Plan-gate self-verification: canary logs gate state each load/re-arm; a silent-dead gate is detectable (workflow enforcement contracts).
  • Block counter per session feeds the behavioral re-injection (plan-first directive). Repeated blocks in a session escalate visibility.

Sources: none - declarative scope statement.

10. Communication conventions

The plan-first directive (injected on block) is prose guidance to the model; the gate itself is deterministic. Communication never gates.

Sources: none - declarative scope statement.

11. Adaptation profile

No per-model adaptation in this spec. (STD/STB re-injection is workflow-scoped.)

Sources: none - declarative scope statement.

12. Directive-priority enforcement (plan-related)

The plan-first requirement is a must-hold directive: it is structurally enforced by the gate (not just stated). A block + injected directive is the PIM-style response to a plan-skipping action.

Sources: none - declarative scope statement.

13. Drift mitigations summary (plan-related)

Drift type Mitigation Source
Enforcement drift (dead gate) Canary + sabotage validation workflow-spec-v1.0.md enforcement contracts
Behavioral bypass (block->todowrite->retry) Per-session block counter + escalating plan-first directive plan-enforce.ts

Sources: none - declarative scope statement.

14. Metric-improvement changes (plan-related)

Metric Change
Plan compliance Block counter + re-injection reduce unplanned acting calls
Plan quality validatePlan rejects empty/placeholder plans (precision over recall)

Sources: none - declarative scope statement.

15. Spec self-compliance & health

  • The plan spec's own conformance vectors (sec 23) must pass for the spec to be valid.
  • The gate's rejection reasons ARE the vectors: any divergence between plan-enforce.ts and this spec is a spec-health failure.
  • Revision trigger: if the block counter rises without a compliance change, the plan-first directive or the format is wrong, not the agent.
  • Collapsed core: the always-loaded core is the plan-first directive and the gate's rejection reasons (PL-003..PL-010); everything else in this spec is on-demand detail.

Sources: none - declarative scope statement.

16. Enforcement/trust root

The trust root is the external gate (plan-enforce.ts), not the model. The plan spec does not change this; it makes the gate's decision rule explicit and machine-checkable.

Sources: none - declarative scope statement.

17. Optional multi-agent orchestration mode - NOT APPLICABLE

MAS is scoped out on this host (workflow-spec-v1.0.md); single-agent only; no sub-agent plan gates.

Sources: none - declarative scope statement.

18. Dynamic cascade controller - NOT APPLICABLE

No cascade controller on this host (workflow-spec-v1.0.md).

Sources: none - declarative scope statement.

19. Glossary

  • Plan: an explicit step-list artifact (todowrite) created before acting tools are permitted; the plan gate's precondition.
  • Plan artifact: a todowrite whose todos array is a valid Plan per the schema in sec 21.
  • Actionable step: a step with content >= 8 chars, not placeholder, status pending or in_progress (not completed/cancelled).
  • Forward contract: a plan must contain at least one actionable step.
  • Plan-first directive: the escalating system-prompt guidance injected on repeated blocks.

Sources: none - declarative scope statement.

20. Benchmark-scoping discipline

No benchmarks are cited in this spec. Behavior is verified by conformance vectors (deterministic), not extrapolated effect sizes.

Sources: none - declarative scope statement.

21. Contracts, schemas & state machines

Plan artifact schema (JSON) - the todowrite todos contract

{
  "type": "object",
  "properties": {
    "todos": {
      "type": "array",
      "items": {
        "type": "object",
        "required": ["content", "status", "priority"],
        "properties": {
          "content": { "type": "string", "minLength": 8 },
          "status": { "type": "string", "enum": ["pending", "in_progress", "completed", "cancelled"] },
          "priority": { "type": "string", "enum": ["high", "medium", "low"] }
        }
      },
      "minItems": 1
    }
  },
  "required": ["todos"]
}

Validation rule (equals plan-enforce.ts validatePlan)

A Plan is VALID iff:

  • todos is a non-empty array, AND
  • every todo satisfies the schema: content is a string of length >= 8, content does not match PLACEHOLDER, status is one of pending|in_progress|completed| cancelled, priority is one of high|medium|low, AND
  • at least one todo is a forward step (status not completed/cancelled).

PLACEHOLDER = regex \b(todo|tbd|placeholder|(fill|to do|doing|later|maybe|something|could) )\b|^[^a-z]{0,3}$ (case-insensitive)

Read evidence (PL-009)

An acting tool (bash/edit/write/task) additionally requires a per-session mark proving the agent READ the plan spec this session. Stamp: tool.execute.after on read of the spec path. Invalidate: tool.execute.after on edit/write of the spec path. Fail-closed: no mark = deny with a corrective hint.

Plan gate state machine

States: no_plan -> armed -> (acting tool) -> no_plan (re-arm on idle). Transitions:

  • no_plan -> armed: validatePlan returns null (valid Plan).
  • no_plan -> no_plan: acting tool attempted (DENIED, block counter++).
  • armed -> no_plan: idle re-arm.
  • any: validatePlan returns reason (REJECTED, review proposal escalated).

Fail-closed: any gate error denies the acting tool.

Sources: none - declarative scope statement.

22. Privacy lifecycle contract - NOT APPLICABLE

Plans are transient per-turn artifacts with no PII persistence; the workflow privacy lifecycle governs the vault and is out of scope here.

Sources: none - declarative scope statement.

23. Executable conformance standard

Deterministic vectors; each maps to plan-enforce.ts validatePlan.

ID Contract Test Expected
PL-001 Plan gate Acting tool (bash) without any todowrite Tool denied; reason "Must perform systematic research and properly plan before bash."
PL-002 Plan gate Valid Plan (1+ actionable step) then bash Bash allowed
PL-003 Plan content Empty todos Rejected: "todowrite had no todos"
PL-004 Plan content Todos all completed/cancelled Rejected: "no actionable"
PL-005 Plan content Placeholder content ("todo") Rejected: placeholder
PL-006 Plan content Content < 8 chars Rejected: placeholder/substantive
PL-007 Plan re-arm Idle after armed plan Gate re-arms to no_plan
PL-008 Self-check Gate disabled (sabotage) Canary flags silent-dead gate
PL-009 Read evidence Acting tool with plan but spec NOT read this session Denied; corrective hint "read the plan spec"
PL-010 Plan content Step priority not in high/medium/low Rejected: "priority must be one of"

Conformance level: all PL-001..010 must pass deterministically.

Sources: none - declarative scope statement.

24. Evidence base

  • workflow-spec-v1.0.md (parent contract: Plan gate, Plan definition, state machine).
  • spec-format-specification.md (SFS format this spec conforms to).
  • Authoring provenance (resolves in the source vault, not shipped here): Rules/Tracked-plan-mandatory.md (numbered/outcome/status format intent), .opencode/plugins/plan-enforce.ts (validatePlan - the conformance oracle, including PL-009 read evidence and PL-010 priority enum), .opencode/test/plan-gate.test.ts (conformance coverage PL-001..010), Research/plan-gate-read-evidence.md (PL-009 grounding: Read-Gate arXiv:2608.02011, deer-flow read-before-write, FAVA arXiv:2607.27267).

Sources: none - declarative scope statement.

type spec
status canonical
version 1.1
project snes

AGENT GOVERNANCE & WORKFLOW SPECIFICATION v1.1 - Operative

OPERATIVE (self-contained distribution contract, v1.1). This is the single operative workflow spec in this package. It is the host realization of the canonical workflow-spec-v1.7.md (retained in full in the source vault), folded into a standalone document so no parent reference is required for a new install. Where a contract is marked SPECIFIED-NOT-IMPLEMENTED, that marking is the operative truth: the contract is design/reference, not live - an adapting agent must realize it (or a working alternative) to satisfy it. This package copy carries the decision maps (D1-D9) of the operative spec. v1.1 (2026-08-21): adds the Decision-making process top-level section (sec 5) per SFS v1.1 extension-sections rule; sections 5-24 renumbered to 6-25.

1. Design Principles

# Principle Source
1 Three enforcement tiers: prose (advisory) -> hooks/rules (deterministic) -> gates (fail-closed). AgentPatterns
2 Enforcement anchored to phase boundaries, never per-step-everything. ECLoop; Reason-Less-Verify-More
3 Enforcement selective: commitment actions gated; investigation free. ECLoop ablation
4 Enforcement outside the model's context - the model cannot overrule a gate. hooks-vs-prompts; Stop-Means-Stop
5 Always-loaded instruction set stays small; critical rules first (primacy). IFScale; IHEval
6 Knowledge self-maintains: evidence-gated promotion, never deleted, always reachable, research-gated authoring. session research
7 Verification by execution, not assertion - negative controls. verification-gap
8 Every claim traceable - receipts, provenance, labels survive conflict resolution. receipts; Mnemoverse
9 Directive priority is NOT reliably model-enforced - structurally enforced. Control Illusion; IHEval
10 Memory is an attack surface - origin-bound, scope-enforced, quarantine-on-risk. MemSecBench; TMA-NM
11 Governance decays under context management - constraints pinned, not just stated. Governance Decay
12 Prohibition constraints rot faster than requirements - re-inject at Safe Turn Depth; cap at Safe Token Budget. SRD
13 Behavioral compliance undetectable from text alone - receipts/tool-logs observe it. Compliance Gap
14 Deterministic verification over LLM judges for correctness-critical metrics. AgentProp-Bench; Real-Time Detection
15 Metrics themselves are gameable - held-out validation; treat evaluation as adversarial. SpecBench; HackDetect
16 The enforcement layer must verify itself - gates can die silently. plan-enforce field bug; Monitoring
17 Principled stopping beats fixed-round retry. VRR-Stop
18 The spec must comply with its own rules - collapse to a small always-loaded core. IFScale (self-referential)
19 Trust root (keys) must be out-of-process. receipts; Agent Receipts
20 Human oversight is a budgeted resource. Slite; human-view observation
21 A strong single agent is the default; multi-agent is a condition-gated execution mode. Google; Nature MI
22 MAS helps on parallelizable + critic-verified work; degrades sequential; the harness is the model-invariant multiplier. Google; Team of Rivals; Harness Effect
23 Dynamic selection is a cascade, not a one-shot classifier. LLM-as-Scheduler
24 Topology from the task DAG; transparent cascade + static fallback over opaque learned router. AdaptOrch; MetaRoute
25 Privacy is a first-class lifecycle contract: immutable fact vs. erasable content, crypto-shredding for erasure, decision-proof retention. Eidentic; ClawQL; WunderOS
26 Trust root = the external enforcement layer, NOT the planner. The planner is a high-risk, low-trust role: most attacked, most consequential, most defended. PEAR; OrchestraBench; Nature MI
27 Conformance must be executable: deterministic vectors, runtime-agnostic harness, verifiable receipts - a spec without a conformance suite is unenforceable. H33; AARM; OAP
28 BRANCH: client-side-only enforcement on this host. Contracts are realized with the opencode plugin hook surface; a contract that needs server enforcement is realized as a post-hoc, independently recomputable approximation and documented as such - never silently downgraded to prose. client-side-feasibility; client-side-gap-fill

Sources: inline citations above.

2. Architecture - five layers

flowchart TB
    L0[Layer 0 SOURCE OF TRUTH - immutable, evidence-gated vault] --> L1
    L1[Layer 1 ALWAYS-LOADED - pinned collapsed core, ~5 rules, compaction-exempt] --> L2
    L2[Layer 2 ON-DEMAND STORE - labeled memory fact/inferred/rule/override + research] --> L3
    L3[Layer 3 ENFORCEMENT - external gates, constraint pinning, self-check] --> L4
    L4[Layer 4 COMMUNICATION - prose conventions, confidence footer, labels] --> L5
    L5[Layer 5 RECEIPTS - tamper-evident, signed, hash-chained step records]
Loading

BRANCH (host realization): on this host, Layer 3 is the opencode plugin layer (verified hook surface from Research/client-side-feasibility.md): event, chat.message, chat.params, chat.headers, permission.ask, command.execute.before, tool.execute.before, shell.env, tool.execute.after, experimental.chat.messages.transform, experimental.chat.system.transform, experimental.session.compacting, experimental.compaction.autocontinue, experimental.provider.small_model, tool.definition, experimental.text.complete. Execution mode is single-agent only on this host (sec 19/sec 20 scoped out). Layer 0 is the Obsidian vault; Layer 5 signing keys are out-of-process (Phase 2 of the plan).

Sources: receipts/tamper-evident chains (action-receipts; AERF; IETF CCS; Tesserae); labeled store (labeled-memory-store; OB1); enforcement = external gates (Reason-Less-Verify-More; Stop-Means-Stop; ECLoop); pinned projection (Governance Decay); evidence-gated vault (EGM; projectmem).

3. The workflow loop - phase-anchored triggers

Phase boundary Fire point What happens
Session start / resume plugin load Constraint Pinning + admission-snapshot + anchor + spec-core injection
Turn start chat.message Intent routing + directive-conflict monitor (PIM cheap gate)
Mode decision - SCOPED OUT on this host: single-agent only (sec 19/sec 20)
Plan tool.execute.before (todowrite) Plan gate (re-arms on idle)
(background) write-path, amortized Prose->spec translation - REFERENCE ONLY (not implemented; specs/plans produced directly)
Pre-action tool.execute.before Commitment gate (external) + per-tool receipt
Pre-mutation before acting tools on an interlinked artifact set Context-gain gate (G1): reference graph mapped + shown (risk-graded)
Post-mutation after a state-changing step Verify-by-execution gate (G2): applicable verification run + output shown (risk-graded, fail-closed)
Pre-retry before any retry Verify-before-retry (partial: retry store + retry_hint marker; postcondition probe test-only)
Periodic every k < STD turns Constraint re-injection (omission) + anchor
Cycle boundary before persist Verify + budget/health checks (invariant monitor not implemented)
Sub-agent completion - SCOPED OUT on this host: no MAS (sec 19/sec 20)
Completion event idle (step.ended/text.ended NEVER fire on this host) Completion evidence - post-hoc tri-state verdict from accumulated receipts
Idle event session.status/session.idle Read-gate + knowledge persist + enforcement self-check + plan-gate re-arm
flowchart LR
    S[Session start / resume - pin + snapshot + spec-core] --> T[Turn start - intent routing + PIM]
    T --> P[Plan gate - valid todowrite + PL-009 read evidence]
    P --> A[Pre-action - commitment gate + per-tool receipt]
    A --> C[Periodic - re-inject omission rules at STD]
    C --> F[Completion - post-hoc tri-state evidence verdict on idle]
    F --> I[Idle - read-gate, knowledge persist, self-check, plan-gate re-arm]
    I --> T
Loading

BRANCH (host realization): the fire points are mapped to the verified opencode events/hooks above (verified in Research/client-side-feasibility.md). The plan-gate re-arm uses session.status + {type:"idle"} (the event that actually fires on the Windows build), not session.idle. session.next.step.ended and session.next.text.ended are NOT delivered on this host, so completion evidence is finalized on IDLE from accumulated receipts - a post-hoc evidence verdict, because no blocking finalize hook exists client-side (see sec 8).

Sources: phase-anchored triggers (enforcement-trigger-points); PEV + verify-before-commit (javatask); constraint pinning (Governance Decay); Safe Turn Depth (SRD); verify-before-retry (Verified-Tool-Calls); deterministic recomputation (Real-Time Detection); Read-Gate (arXiv:2608.02011); evidence of action (IETF draft-msebenzi-evidence-action-00).

4. Intent routing policy (once per turn)

  • Must-hold / actionable -> ensure a spec exists (translate amortized, background).
  • Ambiguous -> ask the user with recommended resolution options.
  • Everything else -> respond in prose, no spec machinery.

Decision map (D1) - intent routing:

flowchart TD
    R[incoming request] --> Q{task type?}
    Q -- must-hold / actionable --> A[ensure a spec exists; plan gate applies]
    Q -- ambiguous --> B[ask the user with recommended options]
    Q -- everything else --> C[respond in prose, no spec machinery]
Loading

BRANCH (host realization): the "ambiguous -> ask with options" branch is prose/behavioral (G8) - the plugin observes it but does not gate it. This is a documented non-enforced path, not an omission.

Sources: routing once-per-turn (confidence-gated; ProactAgent); prose-vs-spec (spec-vs-prose; hooks-vs-prompts); clarify-with-options (nl2spec; ARTEMIS).

5. Decision-making process

Extension section (SFS v1.1 rule; user-approved 2026-08-21). Makes the workflow's decision process explicit: how decisions are classified, who decides, the gated decision loop, and the behaviors the workflow fosters and structurally blocks. The decision maps D1-D9 and the gates G1/G2/G3/G5/G6 are defined in their owning sections; this section is the cross-cutting view.

5.1 Decision taxonomy (who decides what)

Every incoming request is first classified (D1, sec 4): must-hold/actionable -> spec + plan gate; ambiguous -> ask the user with recommended options; everything else -> prose. Then the state-decidability boundary (sec 17) splits every rule: state-decidable -> gate; judgment-required/ambiguity -> prose + user.

Who decides:

  • The model decides freely for investigation, research, reasoning, and drafting (enforcement is selective - commitment actions gated, investigation free, principle 3).
  • The workflow decides for gated transitions: plan gate (PL-001..010 + PL-009), G1 context-gain, G2 verify-by-execution, G5 commitment gate, G3 research mode, G6 completion evidence.
  • The user decides for ambiguity, irreversibility, and approval (sec 18; D1 ask branch).

5.2 The gated decision loop

Gate Decision Fail-closed behavior
Plan gate may I act at all? deny until valid todowrite + PL-009 spec read (sec 7)
G1 context-gain may I mutate? deny until read (LOW) or reference-graph map shown (HIGH)
G2 verify-by-execution safe to proceed after a change? deny until applicable check ran + output shown
G5 commitment gate may I do something irreversible? escalate to user's native permission prompt (ask)
G3 research mode how do I acquire knowledge? LOCAL for observed state; ONLINE mandatory for external knowledge; research receipt emitted
G6 completion evidence is it actually done? claims reconciled to receipts -> VERIFIED/INVALID/UNVERIFIABLE

5.3 The fostered sequence (behavioral convergence)

  1. Classify the request (D1).
  2. Plan before acting (valid todowrite + spec read).
  3. Ask when ambiguous or irreversible (D1 ask branch; G5).
  4. Research in the correct mode (G3).
  5. Read context before mutating (G1).
  6. Verify by execution after mutating (G2).
  7. Label every claim (FACT / REASONING / [INFERENCE] / UNVERIFIED / UNKNOWN).
  8. Recover from failure in bounded steps (retry once -> verify-before-retry -> 2-fail stop+report -> failed verify = revert).
  9. Reconcile claims to receipts before claiming done (G6).

5.4 Structurally blocked behaviors

  • Acting without a plan (plan gate deny).
  • Blind retry (verify-before-retry; postcondition probe).
  • Silent gate death (canary/heartbeat - principle 16).
  • Compliance-with-form (empty/placeholder plans rejected, PL-003..006/PL-010).
  • Manufactured certainty (evidence taxonomy + SOM marking).
  • Unlabeled inference presented as fact (D6 evidence taxonomy).
  • Local-only research for external-knowledge questions (G3/D4: ONLINE mandatory).
  • The model overruling a gate (gates are outside model context - principle 4).
  • Trusting the planner/model as the trust root (principle 26; enforcement layer is the trust root - sec 18).
  • Claiming "done" without execution evidence (G6 tri-state).

5.5 Summary decision flow

flowchart TD
    R[request] --> D1{intent}
    D1 -- ambiguous --> ASK[ask user with options]
    D1 -- prose --> PSE[respond in prose]
    D1 -- actionable --> PG[plan gate: valid todowrite + PL-009 read]
    PG -- denied --> RETRY[fix plan / read spec]
    PG -- armed --> RS{research mode G3}
    RS -- local-answerable --> LO[local + receipt]
    RS -- external --> ON[online MANDATORY + receipt]
    LO --> G1{context-gain: read/map?}
    ON --> G1
    G1 -- deny --> READ[read / map ref graph]
    G1 -- ok --> ACT[act; commitment gate asks if irreversible]
    ACT --> G2{verify-by-execution}
    G2 -- deny --> VERIFY[run applicable check]
    G2 -- ok --> NEXT[next step]
    NEXT --> CMP[completion: claims vs receipts -> tri-state]
    CMP --> IDLE[idle: read-gate, persist, self-check, re-arm]
Loading

Sources: decision maps D1-D9 + gates G1/G2/G3/G5/G6 (sec 7); principles 3/4/16/26 (sec 1).

6. Prose->spec translation pipeline (for must-hold rules)

  1. Extract atomic requirements (one rule per unit, with ID)
  2. Intermediate representation
  3. Grammar-constrain the output space
  4. Verifiability filter - discard non-expressible; escalate ambiguity
  5. Translate via decomposed, interactive sub-translations
  6. Validate continuously (syntax/type/SMT) + semantic-alignment judge
  7. Verify by execution with negative controls

BRANCH (host realization): REFERENCE DESIGN ONLY. This pipeline is not implemented on this host. Must-hold rules are encoded directly as spec documents; plan artifacts are produced directly as structured todowrites per workflow-plan-spec-v1.0.md section 6, not via this pipeline.

Sources: prose-to-spec; ARTEMIS; nl2spec; Doc2Spec; arXiv:2604.18228; Expecto.

7. Enforcement contracts (host realization - client-side-only)

Every contract below states its client-side mechanism and, where it is an approximation of a server-enforced gate, states that explicitly (principle 28).

Plan gate: acting tools denied until a plan exists; re-arms each turn. Host mechanism: tool.execute.before todowrite-gate in plan-enforce.ts (a todowrite is a TOOL call; command.execute.before is for slash commands and cannot intercept it); requires the plan format per workflow-plan-spec-v1.0.md; re-arms on event session.status + {type:"idle"}. Self-verification: heartbeat/canary logs gate state so a silent-dead gate is detectable, not silently ignored (fixes the field bug where the gate reset on session.idle which never fires on Windows). PARITY (v1.8-host) - two host semantics: (a) PL-009 read evidence - beyond a valid plan, the gate requires a hook-observed read of workflow-plan-spec-v1.0.md this session (tool.execute.after stamps the mark; edit/write of the spec invalidates it; fail-closed: no mark = deny with a corrective hint - workflow-plan-spec-v1.0.md section 22). (b) planCreated is process-global, not per-session - one session's todowrite arms the gate for every session until the next idle re-arms it; PL-009 read-marks remain per-session. The asymmetry is documented, known behavior.

Commitment gate: for write/send/irreversible only. MUST be environment- external (framework gates leak sibling effects during pauses - Stop-Means-Stop). Host mechanism: the plugin's permission.ask hook sets status: "ask" so the user's native permission prompt is the authoritative environment-external boundary (see the second Commitment gate paragraph below); deny wins over ask. PARITY (v1.8-host): the gate is user-mediated, NOT a deterministic deny; sabotage-validation and precision-auditing are SPECIFIED-NOT-IMPLEMENTED, and the gate has zero test coverage (as of 2026-08-20). Verdict schema in sec 23.

Verify-before-retry: never blindly retry; verify postcondition first + idempotency key. PARITY (v1.8-host): PARTIAL. The plugin keeps a (tool, args-hash) -> result store fed by tool.execute.after; tool.execute.before checks only for store EXISTENCE and emits a retry_hint evidence marker when a prior identical call succeeded. The postcondition PROBE (postconditionMet) is test-only (enforcement-core.ts); post_condition_probe_hash is always null live. The full Verified-Tool-Calls probe is SPECIFIED-NOT-IMPLEMENTED. (Verified-Tool-Calls; G4)

Verify-between-cycles: unverified intermediate state never committed.

Invariant monitor: PARITY (v1.8-host): SPECIFIED-NOT-IMPLEMENTED. No deviation/drift-detection layer exists on this host; no frozen admission-time snapshot is maintained. The nearest live mechanisms are: the AGENTS.md budget guard (25KB/200 lines, vault-core.ts), the health counters, and read-gate miss detection at idle. (From-Admission-to-Invariants)

Completion evidence - BRANCH (post-hoc receipted evidence): "done" requires spec acceptance + fresh execution-tested evidence + a deterministic recomputation. Host realization: because no blocking finalize hook exists client-side AND session.next.step.ended/session.next.text.ended are NOT delivered on this host, the record is finalized on IDLE from accumulated receipts as a post-hoc, independently recomputable evidence record (evidence-of-action family) with a tri-state verdict - VERIFIED, INVALID, UNVERIFIABLE. PARITY (v1.8-host): live VERIFIED verdicts are predominantly "execution evidence present" (claimed_totals empty) because claim-vs-receipt reconciliation only fires when final prose contains work-done claims; the one live INVALID (evidence seq 93) had a parser-artifact claim key; UNVERIFIABLE has never occurred live. Trajectory-level Near-Miss pre-conditions are SPECIFIED-NOT-IMPLEMENTED. This is an honest post-hoc approximation, NOT a blocking finalize gate (server-side-only; documented, not claimed). Verdict schema in sec 23.

Decision map (D9) - completion evidence tri-state:

flowchart TD
    F[final prose claims] --> H{work-done claims?}
    H -- yes --> C{claims match receipts?}
    C -- yes --> V[VERIFIED]
    C -- no --> I[INVALID - corrected to receipted figure]
    H -- no --> R{receipts exist?}
    R -- yes --> V
    R -- no --> U[UNVERIFIABLE]
Loading

Read-gate (idle): dormant-signature skips flagged and counted. Host realization: plan-enforce.ts stamps every read of a Lessons/*.md file (.plan-enforce-lesson-reads.json, shared with self-maintain); on idle, self-maintain flags a lesson whose signature recurs but whose vault file was never read (a SKIP - visible, not silent), deduped against its own revive proposals.

Commitment gate (permission boundary): for irreversible/destructive actions (task tool; write/edit targeting outside the project dir) the plugin's permission.ask hook sets status: "ask" so the user's native permission prompt is the authoritative environment-external boundary. Non-destructive actions pass through (plan gate covers them); deny wins over ask. Ordering: plan-enforce's tool.execute.before is the deny gate; enforcement's permission.ask only escalates to ask when nothing denied.

Directive-priority monitor: parallel input monitor + sequential output monitor. Host realization (G1a/G1b):

  • PIM: cheap scriptable gate on every chat.message (keyword/priority-pair match vs the pinned directive list), writing a pim_conflict evidence marker on a flag. PARITY (v1.8-host): the small-model judge is NOT wired - experimental.provider.small_model is never invoked by any plugin; the cheap gate is the whole PIM. Bounds per-message cost/latency (LAS: cheap gate first) are trivially satisfied.
  • SOM: review-then-repair - review each drafted response against the pinned higher-priority directives and repair via experimental.text.complete (or filter via experimental.chat.messages.transform). A post-hoc correction approximation, NOT a pre-send block. Evidence taxonomy (Global Evidence Directive): countable work-done claims are split against session receipts - VERIFIED claims need no label (high confidence default); INVALID claims are corrected to the receipted figure; unverified claims the model labeled (FACT / REASONING / [INFERENCE] / UNVERIFIED / UNKNOWN) are allowed; only unverified AND unlabeled claims get a targeted "[UNVERIFIED: ]" marker. Never manufacture certainty.

Constraint Pinning: pinned buffer, compaction-exempt, re-injected verbatim, integrity-checked. Host mechanism: experimental.session.compacting (BEFORE compaction) + experimental.chat.messages.transform / system.transform per turn. (Governance Decay; G3 revised feasible)

Safe Turn Depth / Safe Token Budget: re-inject omission (prohibition) rules every k < STD; cap session at STB for safety-critical constraints. Values are per-model, computed by calibration. Host mechanism: chat.params fires per LLM call; plugin counts turns/tokens and re-injects omission rules at k < STD. PARITY (v1.8-host): the stb_over_budget marker fires on essentially every turn after turn ~8 across sessions (82% of the live evidence log) - it is a MARKER ONLY (no enforcement action) and over-triggers; the calibration is flagged as suspect. (SRD)

Principled stopping: PARITY (v1.8-host): SPECIFIED-NOT-IMPLEMENTED. Belief-based stop/re-pair is advisory only; the actual recovery discipline is fixed: retry once -> escalate -> 2-fail stop -> failed verify = revert (sec 12). (VRR-Stop)

Enforcement self-verification: gates sabotage-validated on a schedule; heartbeat + canary independent of the failing subject; enforcement events receipted; silent-dead gate = critical alert. Host mechanism: the plan-gate canary + a conformance check (WS-018). (Monitoring; plan-enforce field bug)

Sources: commitment gate (Stop-Means-Stop; Reason-Less-Verify-More); verify-before-retry (Verified-Tool-Calls); invariant monitor (From-Admission-to-Invariants); completion evidence (Real-Time Detection; IETF draft-msebenzi-evidence-action-00); directive monitors (Where-IH-Breaks); pinning (Governance Decay); STD/STB (SRD); stopping (VRR-Stop); host hook surface (client-side-feasibility; client-side-gap-fill).

Context-gain gate (G1, risk-graded): acting tools that MUTATE an interlinked artifact set are DENIED until local evidence shows the agent mapped the set's reference graph this cycle. Depth scales with mutation class: LOW (single-file content edit AND isolation VERIFIED - no external refs to the changed file) requires only the changed file's relevant region; HIGH (rename/remove/merge a referenced file, cross-file refs, >= 2 interlinked files) requires the full graph - which files exist, which reference which (parent/child, links, filenames, anchors) - SHOWN. The map is computed by grep/glob (a few tokens); full file reads happen only for the subset that will change. Moderate uncertainty about isolation is material uncertainty and MUST be resolved by the map before mutating; never classify LOW on assumption. Generalizes the plan-spec PL-009 read-evidence pattern to every artifact class. Fail-closed: no map = deny. (Read-Gate arXiv:2608.02011; deer-flow read-before-write)

Verify-by-execution gate (G2, risk-graded): after a state-changing step, the next acting step is DENIED until the APPLICABLE deterministic verification passes and its output is shown. LOW mutations run the cheap applicable check or a read-back re-fetch of the changed state; HIGH mutations run the full applicable verification. Selector (existing checks, not new tools): spec -> sfs-validate; code -> tests/conformance; vault -> vault-audit; doc set/distribution -> reference-resolution (links, filenames, parents, anchors resolve); any artifact -> read-back re-fetch. Deterministic checks are non-LLM and low-token. Fail-closed: any gate error = deny, never fail-open. (Self-Healing Agentic Orchestrators; Verified Tool Calls; Near-Miss; vault lesson verify-after-change)

Research-mode decision (G3): every research step records its mode (local / online / both). "Local-only" is a deliberate labeled choice valid only for local-answerable questions (e.g. reference-graph resolution); unperformed local research (acting without reading/mapping) is a G1 violation. ONLINE research is MANDATORY, not optional, whenever external knowledge is needed or local evidence is insufficient - never answer an external-knowledge question from local inference alone (CRAG corrective fallback). Online sources are data-backed, scoped + cited, >= 2 independent sources with authority + recency checked (superseding docs beat the most-relevant one); retrieved evidence is GRADED (relevant/authoritative/fresh) before use - noisy or low-authority docs degrade correctness.

Decision maps (D2-D4):

flowchart TD
    M[mutate artifact set] --> I{existing file?}
    I -- new --> OK[proceed]
    I -- existing --> R{read this session?}
    R -- yes --> OK
    R -- no --> U{moderate uncertainty about refs?}
    U -- no --> LOW[LOW: read the file region; cheap verify or read-back]
    U -- yes --> MAP[map the reference graph - grep]
    MAP --> X{external refs found?}
    X -- no --> LOW
    X -- yes --> HIGH[HIGH: full map SHOWN + full verification]
    M -- rename / remove / merge / cross-file --> HIGH
    LOW --> A[verify-by-execution - fail-closed]
    HIGH --> A
Loading
flowchart TD
    A[state change] --> C{artifact class}
    C -- spec --> S[sfs-validate]
    C -- code --> T[tests / conformance]
    C -- vault --> VA[vault-audit]
    C -- doc set / distribution --> R[reference-resolution]
    C -- other --> RB[read-back re-fetch]
    S --> P{pass?}
    T --> P
    VA --> P
    R --> P
    RB --> P
    P -- yes --> N[next step]
    P -- no --> D[deny - fail-closed]
Loading
flowchart TD
    Q[research question] --> L{answerable from observed local state?}
    L -- yes --> LO[LOCAL: read/map the relevant subset - record mode]
    L -- no --> E{external knowledge needed?}
    E -- yes --> ON[ONLINE: MANDATORY, data-backed, scoped + cited, >= 2 sources, authority + recency]
    E -- both --> B[BOTH]
    LO --> RR[research receipt]
    ON --> RR
    B --> RR
Loading

8. Knowledge lifecycle (labeled store contract)

Every entry carries: label (fact | inferred | rule | validated-override); provenance (writer identity, evidence reference, derivation, valid-time, transaction-time - survives conflict resolution); three trust axes (relevance/reliability/staleness, independent); scope (cross-user AND cross-project; a locally-valid project rule is never applied to another project without explicit elevation).

Security: PARITY (v1.8-host): PARTIAL. Implemented: the write-path evidence gate at lesson promotion (EGM) and the PII policy (forbidden categories rejected at the promotion write gate; PII flagged, never rejected; receipts store hashed args/results). SPECIFIED-NOT-IMPLEMENTED: origin-bound writes (lessons carry no writer identity in frontmatter), source isolation, quarantine-on-risk, artifact-level scope control. The gate covers ONE store (lesson promotion) - AGENTS.md rendering, receipts, evidence, and review proposals are not gated.

Evaluator locking: PARITY (v1.8-host): SPECIFIED-NOT-IMPLEMENTED. Trust signals are not computed under a locked regime on this host. (RewardHackingAgents)

Held-out validation for promotion: a lesson promotes on 2x evidence in window N; correctness confirmed on windows N+k (k>0) or cross-project/cross- domain data it never saw. PARITY (v1.8-host): promotion and held-out window RECORDING are implemented; the demotion path (held-out FAIL -> demote -> AGENTS.md regen + re-verify) is SPECIFIED-NOT-IMPLEMENTED (recordPromotion is test-only; promotion_holdout_failures never advances live). (SpecBench)

BRANCH (host realization):

  • Time-anchored revival (flap fix): when a lesson is archived, set archivedAt in frontmatter. Revive ONLY when the signature recurs with time_created > archivedAt, not across the whole 7-day window. Eliminates the live revived 33 / archived 33 oscillation (root cause of the AGENTS.md 47KB corruption).
  • Budget = category-aware split: always-loaded AGENTS.md capped at 25KB / 200 lines; large prose sections (Rules project notes) move to on-demand vault notes loaded via path-scoping, NOT archived away. Budget enforced by a guard so the cap cannot be exceeded by future additions.
  • Safe migration (G14): regenAgents runs dry-run to scratch output first, verifies budget AND that no lesson is lost (each permanent present once, each archived present in the archive index - content-based, not count-based); only then writes the real AGENTS.md; prior AGENTS.md kept as legacy until the new one verifies. (aletheia migration pattern)
  • archivedAt persistence (v1.8-host, verified 2026-08-20): archivedAt IS written to frontmatter at archive time; legacy lessons without it are backfilled to "now" in-memory so pre-archive evidence cannot revive them.
  • Knowledge-drift monitoring (v1.8-host): SPECIFIED-NOT-IMPLEMENTED. The only "drift" artifact is a review-proposal item type label. The nearest live mechanism is read-gate miss detection at idle.

Privacy lifecycle contract: see sec 24.

Lifecycle state machine in sec 23.

Sources: labeled store (labeled-memory-store; OB1); origin-bound (TMA-NM; MemSecBench); scope (arXiv:2604.01350); evaluator locking (RewardHackingAgents); held-out (SpecBench); drift (Slite; Armalo); migration safety (aletheia UPGRADING).

9. Action receipts contract (traceability layer)

Every step emits a receipt with schema in sec 23: post-condition probes on writes (falsifiable claims); receipt-derived summary (verified Actions vs LLM Narrative vs auto-flagged Divergence, claims_verified/claims_total); non-execution proof for denied/aborted steps; human-view observation as a first-class detection signal; hash-chained, signed, fail-closed; behavioral channel is the ONLY observability of process fidelity (DPI). PARITY (v1.8-host): signatures are ed25519; chain verification (verifyReceiptChain) is offline/test-only, not invoked by any plugin. outcome is derived from the result title containing "error"; reversibility is false for bash/edit/write/task; parent_seq, post_condition_probe_hash, and evidence_ref are always null live (the schema is broader than the client realization).

Key custody: signing keys out-of-process / separate from the agent (principle 19). In-process keys = evidence-grade but NOT non-repudiable (stated limitation). Host: Phase 2 moves keys out-of-process; until then the stated limitation applies.

BRANCH (host realization): the plugin writes receipts from observed events (tool.execute.before/after + event), not from a server-authoritative trace. Client-side replay is approximate; authoritative full-trace replay is server-side-only and out of scope (documented, not claimed).

Sources: action-receipts; AERF; IETF CCS; Tesserae; Xiaoguai ADR-0008; Compliance Gap.

10. Monitoring contract

  • Multi-signal, never a single LLM monitor; measured MR/FAR per deployment (AutoMonitor-Bench). PARITY (v1.8-host): the live monitor is SINGLE-signal (completion_evidence only); plan-gate blocks, privacy denials, and read-gate skips never reach monitor.jsonl; MR/FAR is computed in tests only, never live.
  • Never LLM-judge pass/fail on correctness-critical metrics.
  • Detection is model-dependent; use per-model rejection rates.
  • Category-aware: on-demand/skill constraints get stronger, not weaker, enforcement (OCTOBENCH).
  • Latent sub-agent failure modes detected + attributed, never repaired by retry (OrchestraBench - MAS-related; retained as a principle, not exercised on this single-agent host).
  • Routing-policy monitoring: routing fidelity measured separately from task quality; routing policy held-out validated (DecisionBench).

BRANCH (host realization): baselines recorded BEFORE Phase 1 so success is before/after, not asserted: claims_verified/claims_total, AGENTS.md bytes/lines, flap count per pass, plan-gate-fires count. The spec-health counters (sec 17) report against these baselines. PARITY (v1.8-host): only claims_verified/claims_total advance live (verified rate ~62.5% as of 2026-08-20); promotions, promotion_holdout_failures, and constraint_survival never advance (their recorders are test-only).

Sources: AutoMonitor-Bench; OCTOBENCH; DecisionBench; plan-gap-check G10 (baselines).

11. Communication conventions (Layer 4 - prose, non-enforced)

  • Confidence footer: High = receipt-verified, Medium = labeled inference, Low = unverifiable. Communication only - never gates.
  • Evidence taxonomy (Global Evidence Directive): FACT (established), REASONING (derived), [INFERENCE] (deduction), UNVERIFIED (claim/report not established), UNKNOWN (insufficient evidence). Verify material uncertainty, not everything; never manufacture certainty. User escalation format: WHAT IS NEEDED -> WHY -> WHAT IT BLOCKS -> HOW TO OBTAIN IT.
  • Facts | Reasoning | [Inference] labeling.
  • Failure recovery: retry once -> escalate -> 2-fail stop -> failed verify = revert (stopping is belief-based per sec 8).

Decision maps (D5-D6):

flowchart TD
    F[failure] --> R1[retry once]
    R1 --> V1{verified?}
    V1 -- yes --> OK[done]
    V1 -- no --> R2[retry - verify-before-retry]
    R2 --> V2{2 fails?}
    V2 -- yes --> ST[stop + report]
    V2 -- no --> F
    F -- failed verify --> RV[revert]
Loading
flowchart TD
    C[claim] --> V{receipt-verified?}
    V -- yes --> FACT[FACT - high confidence]
    V -- no --> R{derived from evidence?}
    R -- yes --> REASON[REASONING]
    R -- no --> I{deduction?}
    I -- yes --> INF[INFERENCE]
    I -- no --> U{established?}
    U -- no --> UNK[UNKNOWN - insufficient evidence - ask]
    U -- yes --> UNV[UNVERIFIED - label as such]
Loading

Sources: Where-IH-Breaks; Control Illusion; Verified-Tool-Calls; VRR-Stop.

12. Adaptation profile + model-switch recalibration

  • Enforcement depth scales with model reliability/context.
  • Collapsed priority core: ~3-5 enforced rules first, remainder on-demand.
  • Drift-aware per-model thresholds (Safe Turn Depth, Safe Token Budget), computed by calibration.
  • Model-switch recalibration: on any model change, thresholds and rejection rates recomputed before trust in safety-critical paths.
  • Role-aware model assignment evaluated per-domain; planner is always a strong model with memory.

BRANCH (host realization): the flash-fallback model inherit (learn session) is the only model routing on this host; recalibration applies to it.

Sources: SRD (STD/STB calibration); model-switch recalibration per sec 8.

13. Directive-priority enforcement (self-contained)

Instruction priority is not reliably model-enforced, so it is structurally enforced (principle 9):

  1. Parallel Input Monitor (PIM): a separate thread checks each new low-priority message against higher-priority instructions; on conflict, discard speculative output, inject a sanitized warning, re-run. Host: cheap scriptable gate on chat.message only; PARITY (v1.8-host): the small-model judge is NOT wired (never invoked).
  2. Sequential Output Monitor (SOM): reviews drafted responses for violations of higher-priority instructions - catches realization failures. Host: review-then-repair via experimental.text.complete / experimental.chat.messages.transform.
  3. Critical rules first + re-injection: the ~3-5 highest-value constraints in the strongest-recency position; re-inject near the end of each request.
  4. Positive phrasing, not prohibitions; prune history demonstrating banned behavior (examples beat instructions).
  5. Priority as data: explicit privilege values (ManyIH-style), not role labels.
  6. Hard constraints in gates, never prompt rules.

Sources: Control Illusion; IHEval; ManyIH; IH-Benchmark; Where-IH-Breaks; Multigrid; host mechanisms (client-side-gap-fill).

14. Drift mitigations summary (self-contained)

Drift type Mitigation Source
Governance decay (compaction) Constraint Pinning (host: experimental.session.compacting + per-turn transform) Governance Decay
Omission decay (SRD) Re-inject at STD; cap at STB SRD
Compliance gap (process) Receipts/tool-logs (behavioral) Compliance Gap
Goal/persona drift Frozen admission snapshot + anchor arXiv:2603.03456; ContextEcho
Knowledge-base drift SPECIFIED-NOT-IMPLEMENTED (no drift monitor); nearest live: read-gate miss detection + propose-then-approve review queue Slite
Manufactured certainty Evidence taxonomy pinned + SOM marks only unverified-unlabeled claims; never manufacture certainty Global Evidence Directive
Dormant-signature skip Read-gate: recurred-but-unread lesson flagged at idle Read-Gate arXiv:2608.02011
Enforcement drift Sabotage-validated gates + self-check (host: plan-gate canary) plan-enforce bug; Monitoring

Sources: inline citations above.

15. Metric-improvement changes

Metric Change Source
False-success rate Post-hoc receipted-evidence recomputation over LLM judges Real-Time Detection
Whole-build verification Verify-before-retry + idempotency Verified-Tool-Calls
Constraint survival Constraint Pinning (model-independent) Governance Decay
Poisoning defense Evaluator locking / reference metric RewardHackingAgents
Promotion precision Held-out validation SpecBench
Prohibition recall Multi-signal monitoring with measured MR/FAR AutoMonitor-Bench
Scope leakage Artifact-level scope control arXiv:2604.01350
Token efficiency Deterministic over LLM calls Real-Time Detection

Sources: inline citations above.

16. Spec self-compliance & health

  • Spec core is collapsed: the spec defines which contracts are the ~5 always-loaded rules and which are on-demand - it must not exceed its own compliance ceiling.
  • Spec-health metric: the spec is revised when its own metrics don't move - e.g., if the completion evidence verdicts don't reduce false-success, the evidence mechanism is wrong, not the agent. A revision trigger exists. PARITY (v1.8-host): the only live counters are claims_verified/claims_total; promotions/holdout-failures/constraint- survival are library-only and MUST NOT be reported as live signals.
  • Cost budget: the full stack's combined cost is budgeted; escalate through the cheapest layer that holds.
  • Human-oversight budget: human time sized and capped (~15min/day review queue). BRANCH: the review surface is a reviews/ note dir + notification via the idle toast/log; propose-then-approve for drift and lesson promotions (G13).
  • State-decidability boundary: every rule is classified - state-decidable -> gate; judgment-required/ambiguity -> prose + user. Rules that can't be state-decided are explicitly NOT gated.
  • Verify-gate health (v1.8-host): the context-gain (G1) and verify-by-execution (G2) gates are live-counter candidates - context-gain denials and verify-after denials per cycle are recorded; if the count rises without artifact-quality change, the gate or its calibration is wrong, not the agent.
  • BRANCH: client-side boundary is a documented scope line, not an afterthought: completion evidence = post-hoc tri-state verdict (not blocking); SOM = review-then-repair (not pre-send block); replay = approximate (not authoritative). PARITY (v1.8-host): the scope line also covers the SPECIFIED-NOT-IMPLEMENTED contracts - commitment-gate sabotage-validation, verify-before-retry probe, invariant monitor, security set, knowledge-drift/evaluator locking, principled stopping, and the prose->spec pipeline. Any future claim that these are implemented on this host is a conformance failure.

Sources: none - declarative scope statement.

17. Enforcement/trust root

The trust root is the environment-external enforcement layer (the deterministic gates, receipts, and their out-of-process keys), NOT the planner and NOT any model. Per principle 4, enforcement runs outside the model's context and cannot be overruled; per principle 19, keys live out-of-process. The model (including the planner) is never the trust root.

BRANCH (host realization): the trust root is the plugin gate layer (tool.execute.before, permission.ask, plan gate) + the receipt store + the out-of-process key (Phase 2). The learn session and the planner are never the trust root.

Sources: hooks-vs-prompts; Stop-Means-Stop; action-receipts.

18. Optional multi-agent orchestration mode - SCOPED OUT (BRANCH)

Not applicable to this spec (stated, not omitted): single-agent is the default and the ONLY mode on this host. MAS orchestration (delegation gate, critic-veto, context firewalls, latent-mode containment, orchestration simulation, role-aware assignment) is not implemented on this single-user deployment. The principles remain recorded (below) for a future multi-user or regulated scope, per the plan's decision. Re-enter MAS only when the cascade controller (sec 20) is also re-entered.

  • Delegation gate, centralized hub-and-spoke topology, planner as high-risk low-trust role, critic-veto, context firewalls, latent-mode containment, orchestration simulation, role-aware model assignment (retained from v1.7 for future scope; see Research/multi-agent-orchestration.md).

Sources: multi-agent-orchestration; plan gap-check G6 (scope decision).

19. Dynamic cascade controller - SCOPED OUT (BRANCH)

Not applicable to this spec (stated, not omitted): there is no cascade controller on this host; the mode is always single-agent. The v1.7 cascade design (start single-agent, cheap cascade gate, escalate-to-MAS, DAG topology, route-tasks-not-calls, routing evidence) is retained only as reference for a future scope. Bound: if multi-user/regulated scope is ever adopted, sec 19 and 18 are re-entered together, with routing fidelity held-out validated per the v1.7 contract.

Sources: dynamic-mas-selection; plan gap-check G6 (scope decision).

20. Glossary (definitions)

  • State-decidable rule: a rule whose compliance can be determined from the current tool arguments and observable state alone (no judgment, no ambiguity). State-decidable rules may be gates; others are prose + user.
  • Must-hold rule: a non-negotiable, binary, checkable directive (vs. contextual guidance, which stays prose).
  • Spec: the machine-checkable encoding of a must-hold rule (the sec 7 pipeline is REFERENCE DESIGN ONLY on this host; specs are authored directly).
  • Plan: an explicit step-list artifact created before acting tools are permitted (the plan gate's precondition). The plan FORMAT and validation contract are defined in workflow-plan-spec-v1.0.md.
  • Held-out validation: verifying a lesson/routing decision on data/windows it did not see during promotion (window N+k, or cross-project/cross-domain).
  • Safe Turn Depth (STD): the model-specific turn depth before omission (prohibition) compliance degrades; re-inject at k < STD.
  • Safe Token Budget (STB): STD x mean tokens/turn; cap session at STB for safety-critical constraints.
  • Receipt-grounded: a claim backed by a signed, hash-chained step receipt (the confidence-footer "High" basis).
  • Divergence flag: an automated flag when the LLM narrative and the receipt-verified actions disagree (claims_verified/claims_total).
  • Quarantine: demote or block a memory entry from retrieval/promotion while flagged risky, without deleting it. (SPECIFIED-NOT-IMPLEMENTED on this host.)
  • Evaluator locking: computing the trust signal from pristine sources under a locked regime, independent of the agent's reported value. (SPECIFIED-NOT-IMPLEMENTED on this host.)
  • Crypto-shredding: erasing by destroying a per-subject/per-cell key so ciphertext becomes permanently unrecoverable while the audit chain stays intact (GDPR erasure over append-only stores).
  • Retraction tombstone: an additive record asserting a prior fact is retracted as of an instant; satisfies audit/replay without destructive delete.
  • BRANCH - Evidence record (receipted evidence): a post-hoc, independently recomputable statement about an action, carrying a tri-state verdict (VERIFIED / INVALID / UNVERIFIABLE). An evidence record does not gate anything; its verification produces a verdict about the record (IETF draft-msebenzi-evidence-action-00).
  • BRANCH - Cheap gate: a deterministic, scriptable pre-check run before any LLM-based monitor, to bound cost/latency (LAS: cheap gate first).
  • BRANCH - Review-then-repair: SOM's client-side realization: review a drafted response and correct it post-hoc via text.complete; a post-hoc correction, not a pre-send block.
  • BRANCH - Baseline: a pre-implementation measured value used as the before/after reference for a metric target (G10).
  • BRANCH - Observation window: the acceptance criterion for deployment- only phenomena (e.g., flap = 0 over 3 consecutive idle passes) that unit tests cannot prove (G11).
  • BRANCH - Canary/heartbeat: an independent, deterministic signal that a gate is alive; silent-dead gate = critical alert (principle 16).
  • BRANCH - SPECIFIED-NOT-IMPLEMENTED: a contract retained as design/ reference but with no live implementation on this host. Marking it is the parity mechanism: the contract is neither silently downgraded to prose nor falsely claimed as live.

Sources: inline citations above.

21. Benchmark-scoping discipline

Every measurement cited in this spec is a benchmark result, not a universal law. Each claim carries its domain/model/caveat, as marked inline (e.g., "(scoped: ...)"). The spec deliberately avoids presenting single-benchmark effect sizes as general guarantees; where a paper itself flags a result as non-robust (e.g., Nature MI's ~45% threshold is a validated selection rule, not a coefficient-level scaling law), the spec repeats that caveat. Conformance testing (sec 25) measures the spec's own deployment, not benchmark extrapolation.

Sources: none - declarative scope statement.

22. Contracts, schemas & state machines

Receipt schema (JSON)

{
  "id": "string (uuid)",
  "session_id": "string",
  "seq": "int (monotonic per session)",
  "parent_seq": "int|null",
  "tool": "string",
  "tool_call_id": "string",
  "args_hash": "sha256 (canonical-JCS of args)",
  "result_hash": "sha256 (canonical-JCS of result)",
  "post_condition_probe_hash": "sha256|null",
  "authority": "string (gate/approval reference)",
  "policy_hash": "sha256",
  "evidence_ref": "string|null (link to store entry)",
  "ts": "ISO-8601",
  "prev_hash": "sha256 (prev receipt envelope)",
  "signature": "ed25519 (out-of-process key)",
  "outcome": "success|failure|pending|denied|aborted",
  "reversibility": "bool"
}

Completion evidence record schema (BRANCH - NEW)

{
  "id": "string (uuid)",
  "session_id": "string",
  "seq": "int",
  "ts": "ISO-8601",
  "tool_result_receipts": ["sha256 of each receipted tool result"],
  "final_text_hash": "sha256 (of final response)",
  "claimed_totals": { "key": "string", "claimed": "number", "receipted_sum": "number" },
  "verdict": "VERIFIED|INVALID|UNVERIFIABLE",
  "verdict_reason": "string",
  "prev_hash": "sha256",
  "signature": "ed25519"
}

Labeled-store entry schema (JSON)

{
  "id": "string",
  "label": "fact|inferred|rule|validated-override",
  "provenance": { "writer": "string", "evidence_ref": "string", "derivation": "string",
                  "valid_time": "ISO-8601", "transaction_time": "ISO-8601" },
  "trust": { "relevance": "float 0-1", "reliability": "float 0-1", "staleness": "bool" },
  "scope": { "project": "string", "user": "string|null" },
  "status": "candidate|permanent|archived|retired",
  "superseded_by": "string|null",
  "archived_at": "ISO-8601|null",
  "held_out_validation": { "promoted_window": "int", "validated_windows": "int[]" }
}

Gate verdict schema (JSON)

{
  "gate": "commitment|completion|plan|read|critic-veto|directive",
  "decision": "allow|deny|block|defer|escalate",
  "reason": "string (structured, actionable)",
  "spec_ref": "string|null (spec section / acceptance predicate)",
  "evidence": "string[] (what was checked)",
  "ts": "ISO-8601",
  "fail_closed": "bool (MUST be true for correctness-critical gates)"
}

Cascade decision schema (JSON) - NOT APPLICABLE (BRANCH)

The cascade schema is not applicable on this host (sec 20 scoped out); retained from v1.7 as reference only.

Enforcement loop state machine

States: routing -> planning -> gated -> verifying -> completing -> idle (and back). Transitions:

  • routing -> planning: on must-hold/actionable intent (once per turn).
  • planning -> gated: plan exists (plan gate).
  • gated -> verifying: commitment action approved.
  • verifying -> completing: verify-between-cycles passes + budget/health checks OK (invariant monitor is SPECIFIED-NOT-IMPLEMENTED).
  • completing -> idle: completion evidence verdict recorded (post-hoc recomputation + trajectory pre-conditions + receipts) - BRANCH: recorded post-hoc, not blocking.
  • any -> routing: session idle (read-gate + knowledge persist + self-check + plan-gate re-arm). Fail-closed: any gate error transitions to denied (per verdict schema).
stateDiagram-v2
    [*] --> routing
    routing --> planning: must-hold / actionable intent
    planning --> gated: plan exists (plan gate)
    gated --> verifying: commitment action approved
    verifying --> completing: verify-between-cycles + budget/health checks
    completing --> idle: completion evidence verdict (post-hoc)
    idle --> routing: session idle
    routing --> denied: gate error (fail-closed)
    planning --> denied: gate error (fail-closed)
    gated --> denied: gate error (fail-closed)
    verifying --> denied: gate error (fail-closed)
    completing --> denied: gate error (fail-closed)
    denied --> [*]
Loading

Knowledge lifecycle state machine

States: candidate -> permanent -> archived -> retired (plus rollback). Transitions:

  • candidate -> permanent: 2x evidence + write-gate + held-out validate.
  • permanent -> archived: budget overflow.
  • archived -> permanent: evidence AFTER archivedAt (time-anchored - BRANCH).
  • permanent -> demote: held-out FAIL -> regen + re-verify (rollback to archived). PARITY (v1.8-host): SPECIFIED-NOT-IMPLEMENTED - no demotion path exists in code.
  • candidate -> retired: older than RETIRE_DAYS, never promoted.
  • (privacy) any -> erased-by-key-destruction: GDPR/legal erasure via crypto-shredding (deferred on this host; see sec 24).
stateDiagram-v2
    [*] --> candidate
    candidate --> permanent: 2x evidence + write-gate + held-out validate
    permanent --> archived: budget overflow
    archived --> permanent: evidence AFTER archivedAt (time-anchored)
    candidate --> retired: older than RETIRE_DAYS, never promoted
    candidate --> erased: GDPR erasure (deferred, crypto-shredding)
    permanent --> erased: GDPR erasure (deferred)
    archived --> erased: GDPR erasure (deferred)
    retired --> [*]
    note right of permanent
      held-out FAIL -> demote is SPECIFIED-NOT-IMPLEMENTED
    end note
Loading

Sources: none - declarative scope statement.

23. Privacy lifecycle contract - BRANCH (minimal path implemented)

The full v1.7 contract (immutable fact vs. erasable content, crypto-shredding, eraseByScope on every store, retention TTLs, verified-erasure checks, user-accessible memory visibility) is staged on this host:

  • Implemented (minimal path - current scope): classify entries at the lesson-promotion write gate (forbidden categories: credentials, secrets, keys rejected); never store raw credentials/PII in lessons or markers; PII-bearing entries flagged (advisory only). PARITY (v1.8-host): the gate covers ONE store (lesson promotion) - AGENTS.md rendering, receipts, evidence, and review proposals are not gated. Decision-proof retention is SPECIFIED-NOT- IMPLEMENTED: decisionProofHash is library/test-only, never called live. This covers the actual single-user risk (accidental secret/PII persistence in promoted lessons).
  • Deferred (only if multi-user/regulated): full crypto-shredding (per-subject keys + eraseByScope fan-out across every store), retention TTLs default-on for PII, verified-erasure checks, user-accessible memory visibility. The contract is staged minimal-now / full-when-scope-requires.
  • The append-only / "never delete" invariant and the right-to-erasure conflict resolution method (key destruction, retraction tombstones, decision-proof records) remain the reference design and are NOT contradicted by the deferral.

Sources: Eidentic 15-data-governance; ClawQL 29; SAIHM; WunderOS PLRN/007; Tianpan GDPR; Zylos; ArkForge.

24. Executable conformance standard - BRANCH (host-scaled)

A governance spec without a conformance suite is unenforceable. This spec defines a deterministic, runtime-agnostic conformance suite; the host's suite is the test files under .opencode/ run against the plugin API directly.

Harness contract (runtime-agnostic): a single harness drives YAML/JSON cases against any implementation through a thin subprocess adapter (JSON over stdin/stdout - per Open Agent Spec conformance). LLM calls MUST be mocked/stubbed; tests assert runtime BEHAVIOR, never LLM output content. Cases carry requires: capability tags; unsupported features report UNSUPPORTED, not PASS/FAIL.

Deterministic vectors (per H33): each vector is an input -> expected_output JSON contract; byte-identical output required; determinism verified (same vector 1,000x -> 1,000 identical); immutable once published.

Normative vectors (PASS and FAIL) - BRANCH level table:

ID Contract Test Expected
WS-001 Commitment gate Submit action matching DENY policy PARITY: user-mediated permission.ask escalation (deny wins over ask); NOT a deterministic deny gate; no test coverage as of 2026-08-20
WS-002 Commitment gate Make gate unavailable, submit action PARITY: deterministic fail-closed SPECIFIED-NOT-IMPLEMENTED; fail-open bypass behavior NOT verified (no test)
WS-003 Plan gate Acting tool without plan Tool denied; structured reason returned
WS-004 Receipts Generate receipt for ALLOW action Receipt present with required fields; signature verifies offline
WS-005 Receipts Tamper with receipt payload Signature verification fails
WS-006 Receipts Tamper with a middle receipt in the chain Chain break detected (prev_hash mismatch)
WS-007 Store write-gate Write entry without evidence/gate-result Write rejected; actionable rejection returned
WS-008 Store erasure Call eraseByScope(subject) UNSUPPORTED on this host (deferred, sec 24)
WS-009 Store erasure Append-only integrity after erasure UNSUPPORTED on this host (deferred, sec 24)
WS-010 Lifecycle Candidate with 2x evidence + gate + held-out Promoted to permanent
WS-011 Lifecycle Permanent, held-out FAIL SPECIFIED-NOT-IMPLEMENTED: no demotion path exists; recordPromotion test-only
WS-012 Lifecycle Archived, evidence AFTER archivedAt Revived to permanent (time-anchored)
WS-013 Directive monitor Conflicting low-priority message Conflict detected (PIM cheap gate; small-model judge NOT wired); sanitized warning injected; higher-priority followed
WS-014 Cascade Trivial task UNSUPPORTED on this host (sec 20 scoped out)
WS-015 Cascade Parallelizable task UNSUPPORTED on this host (sec 20 scoped out)
WS-016 Critic-veto Sub-agent output without critic approval UNSUPPORTED on this host (sec 19 scoped out)
WS-017 Invariant monitor Drift within admissible space SPECIFIED-NOT-IMPLEMENTED: no drift-detection layer on this host
WS-018 Self-verification Gate disabled (sabotage) Heartbeat/canary flags silent-dead gate as critical
WS-019 Pinning Compaction with pinned constraint Pinned constraint survives; entailed post-compaction
WS-020 Pinning Compaction without pinning Constraint may drop (violation detectable) - negative case untested

BRANCH addition - completion evidence vectors:

ID Contract Test Expected
WS-021 Completion evidence Run states total matching receipted results VERIFIED verdict recorded
WS-022 Completion evidence Run states total contradicting receipted results INVALID verdict recorded
WS-023 Completion evidence Run claims a result with no receipt UNVERIFIABLE verdict recorded

Deployment verification tooling (v1.8-host): verify-e2e.ts replays promote+regen passes on a scratch vault copy against the real DB and asserts steady-state flap = 0; the observation-window criterion (G11, e.g. flap = 0 over 3 consecutive idle passes) is the acceptance test for deployment-only phenomena that unit tests cannot prove. Neither is scheduled in CI; the observation-window check is manual.

Conformance levels (per AARM / H33) - BRANCH:

  • Core (WS-001..006): gates + receipts + fail-closed. PARITY: WS-001/002 are user-mediated commitment escalation; the deterministic deny + fail-closed vectors are NOT implemented - the Core level claim must reflect this.
  • Extended (Core + WS-007, WS-010..013): store + lifecycle + directive monitor. MUST all pass on this host. WS-008/009 report UNSUPPORTED. WS-011 reports SPECIFIED-NOT-IMPLEMENTED.
  • Full (Extended + WS-017..020 + WS-021..023): invariant + self-verification
    • pinning + completion evidence. PARITY: WS-017 reports SPECIFIED-NOT-IMPLEMENTED; WS-020 negative case untested. WS-014..016 report UNSUPPORTED (MAS/cascade scoped out).

An implementation SHALL be conformant at a level iff all vectors in that level's set produce byte-identical expected outputs, deterministically, with UNSUPPORTED (not FAIL) for scoped-out features.

Verifiable conformance (per OAP): two tiers - self-issued (operator runs the suite, signs the result with an out-of-process key, anchors it) verified by anyone re-running the same suite; and independent-witness (a second party runs the same suite, signs a matching attestation). The verifiable conformance receipt is machine-readable and re-runnable, not a narrative audit certificate. (BRANCH: self-issued tier is the current host scope.)

Sources: H33 conformance vectors; AARM; OAP verifiable conformance; Open Agent Spec suite; Microsoft AGT conformance; OASB-2.

25. Evidence base

  • Research/workflow-spec-v1.7.md (the canonical parent this branch derives from).
  • Research/workflow-spec-v1.7-host.md (the superseded host branch this revision derives from; retained as history).
  • Research/v1.8-host-parity-plan.md (this revision's parity delta set; the gap analysis against the live implementation, 2026-08-20).
  • Research/v1.7-implementation-plan.md (gap-fill revision 2026-08-13 - the plan changes folded into this branch).
  • Research/plan-gap-check.md (G1-G15 gap audit).
  • Research/client-side-feasibility.md (verified plugin hook surface).
  • Research/client-side-gap-fill.md (evidence-of-action, SOM repair, PIM cheap gate, migration safety).
  • All prior session Research/ notes listed in the v1.7 evidence base.
  • All sources cited inline above.

Sources: none - declarative scope statement.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment