| name | plan-critic-pattern |
|---|---|
| description | Portable adversarial plan reviewer -- pre-write Verifier + post-phase completeness critic + dimension registry |
| type | reference |
| status | active |
| confidence | high |
| project | global |
| created | 2026-08-26 |
| tldr | Independent critic checks a plan before anyone sees it -- receipts, not confidence scores |
An independent reviewer that checks a plan before it reaches the user, and again after each phase before it is marked done. Two critics, one discipline: pre-write (the Verifier -- adversarial, receipt-grounded, checks the plan against the reference checklist and live sources) and post-phase (the completeness critic -- sweeps for what is missing, stale, or contradictory). Both run in a separate context window so their errors decorrelate from the drafting agent; both return findings, never edits; both grade findings so only blocking ones stop progress. A dimension registry gives the sweep a fixed checklist of quality axes plus a self-extending "meta" dimension.
This spec and the artifacts it reviews are pure ASCII markdown. Conformance rules:
- Pure ASCII only -- no em-dash, arrow (U+2192), minus sign (U+2212), section sign, or
multiplication / less-than / greater-than glyphs. Write
--,->,-,sec,x,<=,>=instead. - Fenced code blocks always tagged with a language:
mermaidfor diagrams,jsonfor schemas,yamlfor frontmatter,markdownfor markdown examples,textfor non-diagram listings. - Mermaid uses
flowchart TD/flowchart LRandstateDiagram-v2; quote any label containing>,<,==, or(. No<br/>. - Fences balanced -- every open fence has a matching close; no unclosed block.
Two independent critics over one review discipline. No database -- the reviewer is spawned fresh each time, so its state lives in the plan's own metadata, not in the reviewer.
| Component | Holds |
|---|---|
verifier.md |
pre-write reviewer definition -- constraints + decision loop |
pre-planning.md |
the reference standard the plan is checked against (every section) |
dimensions.json |
review dimensions + convergence settings (the sweep's checklist) |
completeness critic |
post-phase "what is missing" prompt (general-purpose spawn) |
plan file |
the draft under review -- carries Status + Last verified metadata |
The plan file's metadata is the critic's integration anchor:
Status: NOT IMPLEMENTED | IMPLEMENTED | SUPERSEDED
Last verified: <ISO date>Status records lifecycle; Last verified records when external-dependency claims were
last checked against live sources -- the critic reads this before trusting any dependency
claim.
flowchart LR
V["Verifier (pre-write)"] -->|findings block| O[Orchestrator]
C["Completeness critic (post-phase)"] -->|findings block| O
O -->|revise or proceed| D[Plan / work]
| Role | Runs | Does |
|---|---|---|
| Verifier | before plan reaches user | adversarial check vs checklist + live sources; return findings |
| Completeness critic | after a phase completes | sweep for missing / stale / contradictory; grade findings |
| Orchestrator | after findings arrive | resolve blocking findings, re-run critic, then proceed |
The critics never edit. The orchestrator alone decides whether to revise or proceed.
flowchart TD
S[spawn with draft + checklist + task spec] --> P{is a plan?}
P -->|yes| CK["check vs pre-planning checklist (every section)"]
P -->|no| VC[verify each claim vs receipt]
CK --> VC
VC --> SV[spot-verify highest severity]
SV --> FB[findings block]
FB --> R{verdict}
R -->|approve| PR[present to user]
R -->|request-revision| RV[resolve high findings]
RV --> S
Constraints (non-negotiable):
- Single round -- 2 rounds max; never loop past 3. More rounds add noise, not accuracy.
- Receipt-grounded -- every finding cites a receipt (URL, path + line, or a live tool-call result from this run). A claim that cannot be verified is a finding, not a silent acceptance.
- Anti-sycophancy -- assume the draft may be wrong; confirm only what a receipt supports.
- No self-scored confidence -- never emit a percentage or the bare word "verified". State per-claim status: CONFIRMED / REFUTED / UNVERIFIED.
Return shape:
VERDICT: APPROVE | REQUEST-REVISION
- claim: <the claim>
status: CONFIRMED | REFUTED | UNVERIFIED
receipt: <URL / path:line / tool output>
note: <one-line why>
If the draft is a plan, every section of the reference checklist is reported SATISFIED / N-A (with why) / MISSED -- each MISSED section becomes a finding with its section number.
Decorrelation: a separate context window makes the reviewer an independent evaluator, not a self-eval (Multiagent Debate, arXiv:2305.14325). A same-model critic adds self-preference bias (arXiv:2306.05685).
stateDiagram-v2
[*] --> drafted
drafted --> under_review : critic invoked
under_review --> approved : APPROVE
under_review --> revised : REQUEST-REVISION
revised --> under_review : re-run critic
approved --> implemented : work done
implemented --> superseded : replaced
drafted / implemented / superseded map to the plan Status field
(NOT IMPLEMENTED / IMPLEMENTED / SUPERSEDED). The critic gates the drafted -> approved
edge: a plan with a blocking finding never reaches the user.
The review sweep runs against a fixed registry of quality dimensions. Core and cross-cutting dimensions are always checked; domain dimensions fire only when the plan content matches their trigger.
| Dimension | Category | Trigger | Key question |
|---|---|---|---|
| correctness | core | always | Do plan claims match reality? |
| completeness | core | always | Are all required elements present? |
| consistency | core | always | Any internal contradictions? |
| clarity | core | always | Unambiguous and measurable? |
| feasibility | core | always | Implementable within constraints? |
| testability | core | always | Each phase has verifiable completion criteria? |
| traceability | cross-cutting | always | Every phase traces to a goal, gap, or issue? |
| maintainability | cross-cutting | always | Can a stranger read it cold and act? |
| actionability | cross-cutting | always | Tasks specific enough to execute without more research? |
| citation-quality | cross-cutting | always | Decisions cite verifiable sources? |
| security-impact | domain | auth / secrets / input handling | Plan touches auth, data access, or secrets? |
| performance-impact | domain | queries / render / data processing | N+1 risk, batching, caching addressed? |
| mobile-scope | domain | UI changes | Mobile always in scope -- verified per phase? |
| rollback-recovery | domain | deploy / migration / destructive | Rollback, backups, irreversible flags present? |
| dependency-addition | domain | new packages | Each new dependency authorized + assessed? |
| breaking-change | domain | public API / shared component | Consumers identified + migration documented? |
The registry grows: a "meta" dimension asks "what quality dimension, if it existed, would flag an issue no current dimension catches?" A newly found dimension is added to the registry and the sweep re-runs.
Spawn points -- two critics, two hooks in the workflow.
-
Pre-write (plan presentation gate) -- after a plan is written and before the user sees it, spawn the Verifier with (a) the plan, (b) the reference checklist, (c) the task spec. Resolve every high-grade finding first. A plan is never presented un-reviewed.
-
Post-phase (completion gate) -- after an implementation or audit/analysis phase completes and before it is marked done, spawn the completeness critic. Its findings become the next round of work.
Severity grading -- findings are graded; only the blocking grade stops progress. High / major findings block; minor / clean do not. The Verifier's APPROVE / REQUEST-REVISION verdict is the plan-level gate; the completeness critic's major grade is the phase-level gate.
Convergence config (lives with the dimension registry):
{
"dry_rounds_required": 2,
"max_rounds": 5,
"stagnation_threshold": 2,
"dimension_inflation_threshold": 3
}dry_rounds_required-- K dry rounds (zero new findings, zero new dimensions) to converge; LLM non-determinism means one clean pass is not enough.max_rounds-- hard cap; past this, escalate to the user rather than self-debug.stagnation_threshold-- consecutive rounds with no change halts the loop.dimension_inflation_threshold-- guard against unbounded registry growth.
Report integrity -- before accepting a critic's report, cross-check its tool-call count against its claimed scope. A report claiming many independent checks but backed by few tool calls is unreliable; discard and re-run. An independent reviewer must have its own context window -- one that received the original report as ground truth cannot decorrelate its errors.
Pre-write critic pseudo
spawn verifier with (plan, checklist, task spec)
findings = verifier.return_findings()
for round in 1..2:
if no blocking finding: break
resolve blocking findings in the plan
findings = verifier.return_findings() # fresh spawn, re-check
present plan to user only when findings.verdict == APPROVE
Post-phase critic pseudo
spawn completeness critic with phase summary
findings = critic.return_findings()
for f in findings.major:
fix f
re-check f against disk (grep / find / jq), not the plan text
mark phase done only after all major findings resolve
Configuration cross-reference -- the severity grades, convergence numbers, dimension categories, and spawn-point inputs appear in several places (diagram, pseudo, verify table). Changing one requires sweeping every occurrence; a stale value left elsewhere is a mismatch. After any change, grep the value and confirm every occurrence agrees.
End-to-end:
- Draft the plan (
Status: NOT IMPLEMENTED,Last verifiedset after checking live dependencies). - Spawn the Verifier with the plan, the reference checklist, and the task spec.
- Verifier checks every checklist section (SATISFIED / N-A / MISSED), verifies claims against live sources, spot-verifies the highest-severity claims, returns a findings block with a verdict.
- Orchestrator resolves every blocking finding, re-runs the critic if REQUEST-REVISION.
- Plan reaches the user only after APPROVE.
- After each implementation or audit phase, spawn the completeness critic; its findings become the next round of work before the phase is marked done.
Preconditions (workflow assumptions an implementer must satisfy):
- Subagent spawn -- a separate context window per critic; the critic is a spawn, not a self-check.
- Reviewer definition present -- the Verifier's constraints and loop, or an equivalent.
- Reference standard present -- the checklist the plan is checked against.
- Dimension registry present -- the sweep's dimension list + convergence config.
- Live-source access -- receipts require re-fetching claims, not trusting stored receipts (a stored receipt may be stale).
- Critics are advisory, not editors -- they return findings only; the orchestrator applies changes under its own write gate.
Systematic test -- run in order; each row states how + expected result.
| # | Check | How | Pass when |
|---|---|---|---|
| 1 | reviewer present | list agent definitions | Verifier (or equivalent) definition exists |
| 2 | reference standard present | read checklist | every section present, including the numbered sub-sections |
| 3 | dimension registry parses | parse dimensions file | dimensions + convergence meta load; zero parse errors |
| 4 | critic invoked on a plan | spawn with a draft plan | returns a findings block with a VERDICT |
| 5 | checklist coverage | inspect findings | every checklist section reported SATISFIED / N-A / MISSED |
| 6 | receipt grounding | inspect findings | every finding cites a receipt; no bare "verified" |
| 7 | anti-sycophancy | seed a wrong claim (nonexistent path) | claim returned REFUTED with a receipt, not accepted |
| 8 | no self-score | inspect findings | no percentage or bare confidence word |
| 9 | no draft edit | diff the draft | critic returned findings only; draft unchanged |
| 10 | single round | run on a clean plan | 2 rounds max; never loops past 3 |
| 11 | blocking gate | seed a high finding | plan not presented; fixed, then re-run before APPROVE |
| 12 | format (SFS) | grep non-ASCII + parse fences | zero non-ASCII; all fences tagged + balanced |
Fixtures (steps 7 and 11): seed a plan whose file list names a path that does not exist -- assert the critic returns REFUTED citing the missing path (step 7). Seed a plan missing a checklist section -- assert that section is reported MISSED with its number, and that the verdict is REQUEST-REVISION until the section is added (step 11).