Skip to content

Instantly share code, notes, and snippets.

@danong
Created August 25, 2026 04:03
Show Gist options
  • Select an option

  • Save danong/f113b2cd316bdba518659a78c61a2a12 to your computer and use it in GitHub Desktop.

Select an option

Save danong/f113b2cd316bdba518659a78c61a2a12 to your computer and use it in GitHub Desktop.
[
{
"role": "system",
"content": "You are an expert coding assistant operating inside pi, a coding agent harness. You help users by reading files, executing commands, editing code, and writing new files.\n\nAvailable tools:\n- read: Read file contents\n- bash: Execute bash commands (ls, grep, find, etc.)\n- edit: Make precise file edits with exact text replacement, including multiple disjoint edits in one call\n- write: Create or overwrite files\n- checklist: Track progress through task requirements (init/done/status)\n\nIn addition to the tools above, you may have access to other custom tools depending on the project.\n\nGuidelines:\n- Use bash for file operations like ls, rg, find\n- Use read to examine files instead of cat or sed.\n- You can inspect PI_* environment variables for current model and session details.\n- Use edit for precise changes (edits[].oldText must match exactly)\n- When changing multiple separate locations in one file, use one edit call with multiple entries in edits[] instead of multiple edit calls\n- Each edits[].oldText is matched against the original file, not after earlier edits are applied. Do not emit overlapping or nested edits. Merge nearby changes into one edit.\n- Keep edits[].oldText as small as possible while still being unique in the file. Do not pad with large unchanged regions.\n- Use write only for new files or complete rewrites.\n- Be concise in your responses\n- Show file paths clearly when working with files\n\nPi documentation (read only when the user asks about pi itself, its SDK, extensions, themes, skills, or TUI):\n- Main documentation: /usr/local/lib/node_modules/@earendil-works/pi-coding-agent/README.md\n- Additional docs: /usr/local/lib/node_modules/@earendil-works/pi-coding-agent/docs\n- Examples: /usr/local/lib/node_modules/@earendil-works/pi-coding-agent/examples (extensions, custom tools, SDK)\n- When reading pi docs or examples, resolve docs/... under Additional docs and examples/... under Examples, not the current working directory\n- When asked about: extensions (docs/extensions.md, examples/extensions/), themes (docs/themes.md), skills (docs/skills.md), prompt templates (docs/prompt-templates.md), TUI components (docs/tui.md), keybindings (docs/keybindings.md), SDK integrations (docs/sdk.md), custom providers (docs/custom-provider.md), adding models (docs/models.md), pi packages (docs/packages.md), environment variables (docs/environment-variables.md)\n- When working on pi topics, read the docs and examples, and follow .md cross-references before implementing\n- Always read pi .md files completely and follow links to related docs (e.g., tui.md for TUI API details)\n\nYou are implementing a coding task. Explore the codebase, make changes,\nand call yield() when complete.\n\nMake atomic jj commits as you complete each requirement.\nRun verification commands after your changes.\nWhen the project has tests, work test-first: write a failing test, verify it fails, then implement until it passes (red-green-refactor).\nWrite scratch/debug probes under /tmp, never in the repo — check jj file list before yielding so no debug files are tracked.\n\nEfficiency — make every tool call count:\n- Batch independent tool calls into one turn; never run sequential calls you could run together.\n- Never dump large outputs: truncate bash (head/tail/grep, head -c ~2000), never cat whole files or logs.\n- Never re-run an identical command; re-derive with a smaller query if you lost context.\n- Run the full test suite at most twice (after your changes and before yield); use targeted files otherwise.\n- Each turn must visibly advance the task — finish in the fewest turns you can.\n\nYour first edit should be your most confident change.\n\nUse checklist() to track your progress through requirements.\n\n<project_context>\n\nProject-specific instructions and guidelines:\n\n<project_instructions path=\"/root/.pi/agent/AGENTS.md\">\n# Global Rules\n\n## Behavior\n\n- Execute commands directly. Do not write instructions for the user to run.\n\n## Code quality bar\n\n- Handle errors explicitly — no empty catch blocks\n- No commented-out code in committed files\n- Always keep documentation up to date: update documentation before running jj commit\n- Test aggressively to prevent regressions\n\n## Decision recording\n\n- When choosing between valid approaches, state the tradeoffs explicitly in 1 sentence (e.g., \"chose X over Y because Z\")\n\n## Documentation\n\n- No hardcoded counts (items, tests, assertions) — they go stale. Teach readers where to find the data.\n- No exhaustive enum/list values — say \"includes but not limited to\" and point to the source file.\n- No machine-specific paths (binary paths, home dirs) in README or AGENTS files — use env vars or relative paths.\n- Docs explain the setup and teach readers how to find data, not snapshot the current state.\n\n\n</project_instructions>\n\n</project_context>\n\nCurrent working directory: /workspace/pi-task"
},
{
"role": "user",
"content": [
{
"type": "text",
"text": "## Codebase map\nEntry points: extensions/task/index.ts\nPatterns: Schema-driven tool contracts (TypeBox schemas enforce parameter shapes); Event-driven model swapping (prewalk → execute on first edit); Budget tier resolution chain (locked flag > locked param > config default > auto heuristic); jj workspace isolation for parallel workers (squash-merge into task base); Fork-and-prune adversarial review (forked session inherits context, strips commitment); Typed schema validation at RPC boundaries (yield, findings, spec parsing)\nTest layout: Fast hermetic tests live in test*.ts under extensions/task/ and packages/core-v2/test/ — run via `mise run test` (zero LLM, ≤15s); real-LLM e2e tests (test-e2e.ts, smoke-*.ts) run manually with a 15min timeout.\n\nFiles most relevant to this task (read these first):\n- extensions/task/batch.ts — Provider-agnostic batch inference lane (M2): turns a spec's typed requirements into single-turn prompt items each with an output contract, submits/polls/collects async batch jobs, validates output (code-fence stripping), extracts files with path-safety/conflict checks, and reports typed error codes (submit_failed, poll_timeout, aborted, ...). [symbols: BATCH_ERROR_CODES, BatchErrorCode, BatchErrorDetail, BatchError, BatchOutputContract, BatchPromptItem, stripCodeFences, m, validateBatchOutput, text, value, missing, BATCH_FILE_CONTRACT, BatchFile, requirementId]\n- packages/core-v2/src/daemon/parallel.ts — M2.d parallel pipeline: fan-out N worker sessions across WorkspaceDriver workspaces; combine task-base/feature-branch; verify once on integrated tree; per-attempt + aggregate receipts; conflicts -> escalated. [symbols: SessionHost, UsageSnapshot, RunParallelOptions, WorkerObservation, attemptNumberOf, m, makeReceipt, RunParallelResult, store, env, gateway, pluginHookCtx, parsed, familyBase, isTaskBase]\n- extensions/task/test-index.ts — Hermetic test suite for index.ts — covers budget schema-locking (budget param removed when tier locked), sub_specs-vs-spec precedence and normalization, budget tier resolution including dynamic tier sets (Phase 11), session goals/budget override persistence from session entries, TaskResult-to-tool-return mapping, review-report rendering with finding caps, TUI renderInPlace reuse contract, and the run-plan + progress-state render pipeline. [symbols: BudgetTierConfig, ProgressState, RunPlan, TaskToolReturn, schemaProperties, json, tierConfig, propertyJson, fakeResult, fakeWorker, testBudgetSchemaLocking, check, unlocked, locked, budgetSchema]\n- extensions/task/index.ts — pi-task extension entry point (Phase 9) — registers the `task` tool (orchestrator interface), enforces budget via `--task-budget`/`/task-budget`, persists session goals, injects codebase-map and workflow-contract into the main prompt, streams TUI progress (zero LLM tokens), and supports detached dispatch. Contains pure budget-resolution helpers, schema builders, result mappers, and progress render logic. [symbols: DetachedRunOptions, BudgetMode, BudgetTier, BudgetTierConfig, TaskConfig, isLockedBudget, normalizeBudgetMode, resolveBudgetMode, autoTierForRequirements, small, large, heuristicDefaultTier, names, notMax, countSpecRequirements]\n- extensions/task/orchestrator.ts — Orchestrates task execution, spawns workers, validates specs, and merges workspace commits. [symbols: WorkerFailureDiagnostics, WorkerResult, SwapInfo, BatchLaneConfig, SandboxConfig, TaskShape, BatchFile, BatchProvider, MergeMetrics, RunManifest, MergeFailureInfo, writeMergeFailureArtifact, artifact, FAILURE_PATH_JJ_TIMEOUT_MS, isFinalizationIncomplete]\n- packages/core-v2/test/test-plugins-lifecycle.ts — Hermetic suite for the lifecycle-collector trigger plugin that exercises the REAL plugin path: loaded by path from task.toml, trigger side wired through a real InMemoryTaskGateway subscription, and hook isolation (throwing plugins never crash the pipeline). [symbols: PLUGIN_PATH, TASK_RUNNER_PATH, rowsDouble, errors, check, dir, tomlPath, plugins, plugin, gateway, journal, secondSeen, off, journalLenBefore, freshGateway]\n- extensions/task/test-review.ts — Hermetic tests for review.ts's pure logic: buildReviewPrompt, settleReview outcome settling, and the first-event/no-progress watchdog decisions (forkedReview itself is e2e territory since it spawns a real pi process). [symbols: REVIEW, errors, check, prompt, mk, finding, single, merged, worst, empty, ok, fail, quiet, long, detail]\n- extensions/task/progress.ts — Provides pure functions for generating task-tool progress view, including dispatch plan and phase sequencing. [symbols: WorkerPhase, PlanPhase, RunPlan, BuildRunPlanOptions, buildRunPlan, phases, initialPhaseOf, WorkerProgress, ProgressState, createProgressState, workers, phase, recountDone, n, extractFileScope]\n- extensions/task/test-e2e.ts — Consolidated real-LLM end-to-end test (run manually, ~3-5 min) exercising real pi + a single cheap hardcoded model across steering, sub-specs, fix-loop, and verify paths. [symbols: MODEL, TASK_EXTENSION_PATH, SINGLE_TOOL_PROMPT, SUBSPECS_TOOL_PROMPT, SYSTEM_PROMPT, HELLO_SPEC, MULTI_SPEC, FIXLOOP_SPEC, NO_VERIFY_PROMPT, errors, check, makeRepo, dir, parseRemaining, m]\n- packages/core-v2/src/daemon/task-runner.ts — M1.4 pipeline entry: validate->route->host->yield->verify->persist receipt; deterministic prefix (R5) pure of timestamps; bundle telemetry; plugin extractions (handoff-cap + lifecycle-collector moved out of core). [symbols: USAGE_UNAVAILABLE, UsageSnapshot, totalInputTokens, computeCor, estimateGroundingTokens, emptyUsage, stats, sumUsage, total, receiptUsageFields, SpecValidationError, ParsedTaskSpec, parseTaskSpec, sections, current]\n- extensions/task/test-sandbox.ts — Sandbox builder tests: pure arg-set unit tests (strict ro-root policy, bind ordering, cwd/agent-dir shadowing) plus one guarded real-bwrap probe. [symbols: findTriple, sandboxConfig, jj, bwrapAvailable, testDisabled, check, args, testStrictPolicy, roRoot, dev, proc, tmp, testRwBinds, cwd, agentDir]\n- packages/core-v2/src/plugins/hooks.ts — Isolated per-call plugin hooks: try/catch with error sink, never crashes pipeline; transformExecutionBundle/Handoff re-validated against zod schemas; emitLifecycleEvent fan-out; registerPluginTriggers isolated. [symbols: PluginHookContext, report, sink, current, emitLifecycleEventToPlugins, registerPluginTriggers, callIsolatedSync]\n- packages/core-v2/src/grounding/prewalk.ts — Prewalk cost-break-even policy (M3, OFF by default): pure function decidePrewalkSwap computes whether swapping from strong to cheap model saves money given context size, remaining turns, and per-model pricing; attachPrewalk wires it to a session handle to fire once on first successful edit/write. [symbols: TokenPricing, PrewalkPricing, PrewalkDecisionInput, PrewalkDecision, DEFAULT_OUTPUT_TOKENS_PER_TURN, perM, decidePrewalkSwap, ctx, n, out, strong, cheap, stayCostUsd, swapPenaltyUsd, swapCostUsd]\n- extensions/task/test-checklist-relay.ts — Hermetic tests for the checklist relay reducer against an in-process fake session, asserting R4 progress tracking with zero LLM tokens and observer-only behavior. [symbols: ChecklistProgress, ChecklistRelayState, start, end, FakeSession, makeFakeSession, listeners, commands, session, i, errors, check, state, i1, i2]\n- docs/pi-task-testing-spec.md — Phase 6.5 spec: replace six LLM-driven smoke tests with a fast hermetic deterministic suite (pure-unit + real-jj/real-bash, zero LLM, <=15s) plus one manually run consolidated e2e. [symbols: interface, function]\n\nRead the files above first; explore beyond them only if they do not cover the task.\n\n## Goal\nWrite a single file confirming the task pipeline works end-to-end.\n\n## Requirements\n- R1: Create the file /tmp/pi-task-smoke-test.txt containing exactly the text \"ok\" (a trailing newline is fine).\n\n## Verification\ntest -f /tmp/pi-task-smoke-test.txt"
}
]
},
{
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "af7699f1-204a-4a02-be45-a406adee9283",
"type": "function",
"function": {
"name": "write",
"arguments": "{\"path\":\"/tmp/pi-task-smoke-test.txt\",\"content\":\"ok\\n\"}"
}
}
],
"reasoning_content": ""
},
{
"role": "tool",
"content": "Successfully wrote 3 bytes to /tmp/pi-task-smoke-test.txt",
"tool_call_id": "af7699f1-204a-4a02-be45-a406adee9283"
},
{
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "f3622891-58ff-4ae4-8fd1-acc6cb5e2691",
"type": "function",
"function": {
"name": "bash",
"arguments": "{\"command\":\"test -f /tmp/pi-task-smoke-test.txt && cat /tmp/pi-task-smoke-test.txt\"}"
}
}
],
"reasoning_content": ""
},
{
"role": "tool",
"content": "ok\n",
"tool_call_id": "f3622891-58ff-4ae4-8fd1-acc6cb5e2691"
}
]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment