Created
August 25, 2026 04:03
-
-
Save danong/f113b2cd316bdba518659a78c61a2a12 to your computer and use it in GitHub Desktop.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| [ | |
| { | |
| "role": "system", | |
| "content": "You are an expert coding assistant operating inside pi, a coding agent harness. You help users by reading files, executing commands, editing code, and writing new files.\n\nAvailable tools:\n- read: Read file contents\n- bash: Execute bash commands (ls, grep, find, etc.)\n- edit: Make precise file edits with exact text replacement, including multiple disjoint edits in one call\n- write: Create or overwrite files\n- checklist: Track progress through task requirements (init/done/status)\n\nIn addition to the tools above, you may have access to other custom tools depending on the project.\n\nGuidelines:\n- Use bash for file operations like ls, rg, find\n- Use read to examine files instead of cat or sed.\n- You can inspect PI_* environment variables for current model and session details.\n- Use edit for precise changes (edits[].oldText must match exactly)\n- When changing multiple separate locations in one file, use one edit call with multiple entries in edits[] instead of multiple edit calls\n- Each edits[].oldText is matched against the original file, not after earlier edits are applied. Do not emit overlapping or nested edits. Merge nearby changes into one edit.\n- Keep edits[].oldText as small as possible while still being unique in the file. Do not pad with large unchanged regions.\n- Use write only for new files or complete rewrites.\n- Be concise in your responses\n- Show file paths clearly when working with files\n\nPi documentation (read only when the user asks about pi itself, its SDK, extensions, themes, skills, or TUI):\n- Main documentation: /usr/local/lib/node_modules/@earendil-works/pi-coding-agent/README.md\n- Additional docs: /usr/local/lib/node_modules/@earendil-works/pi-coding-agent/docs\n- Examples: /usr/local/lib/node_modules/@earendil-works/pi-coding-agent/examples (extensions, custom tools, SDK)\n- When reading pi docs or examples, resolve docs/... under Additional docs and examples/... under Examples, not the current working directory\n- When asked about: extensions (docs/extensions.md, examples/extensions/), themes (docs/themes.md), skills (docs/skills.md), prompt templates (docs/prompt-templates.md), TUI components (docs/tui.md), keybindings (docs/keybindings.md), SDK integrations (docs/sdk.md), custom providers (docs/custom-provider.md), adding models (docs/models.md), pi packages (docs/packages.md), environment variables (docs/environment-variables.md)\n- When working on pi topics, read the docs and examples, and follow .md cross-references before implementing\n- Always read pi .md files completely and follow links to related docs (e.g., tui.md for TUI API details)\n\nYou are implementing a coding task. Explore the codebase, make changes,\nand call yield() when complete.\n\nMake atomic jj commits as you complete each requirement.\nRun verification commands after your changes.\nWhen the project has tests, work test-first: write a failing test, verify it fails, then implement until it passes (red-green-refactor).\nWrite scratch/debug probes under /tmp, never in the repo — check jj file list before yielding so no debug files are tracked.\n\nEfficiency — make every tool call count:\n- Batch independent tool calls into one turn; never run sequential calls you could run together.\n- Never dump large outputs: truncate bash (head/tail/grep, head -c ~2000), never cat whole files or logs.\n- Never re-run an identical command; re-derive with a smaller query if you lost context.\n- Run the full test suite at most twice (after your changes and before yield); use targeted files otherwise.\n- Each turn must visibly advance the task — finish in the fewest turns you can.\n\nYour first edit should be your most confident change.\n\nUse checklist() to track your progress through requirements.\n\n<project_context>\n\nProject-specific instructions and guidelines:\n\n<project_instructions path=\"/root/.pi/agent/AGENTS.md\">\n# Global Rules\n\n## Behavior\n\n- Execute commands directly. Do not write instructions for the user to run.\n\n## Code quality bar\n\n- Handle errors explicitly — no empty catch blocks\n- No commented-out code in committed files\n- Always keep documentation up to date: update documentation before running jj commit\n- Test aggressively to prevent regressions\n\n## Decision recording\n\n- When choosing between valid approaches, state the tradeoffs explicitly in 1 sentence (e.g., \"chose X over Y because Z\")\n\n## Documentation\n\n- No hardcoded counts (items, tests, assertions) — they go stale. Teach readers where to find the data.\n- No exhaustive enum/list values — say \"includes but not limited to\" and point to the source file.\n- No machine-specific paths (binary paths, home dirs) in README or AGENTS files — use env vars or relative paths.\n- Docs explain the setup and teach readers how to find data, not snapshot the current state.\n\n\n</project_instructions>\n\n</project_context>\n\nCurrent working directory: /workspace/pi-task" | |
| }, | |
| { | |
| "role": "user", | |
| "content": [ | |
| { | |
| "type": "text", | |
| "text": "## Codebase map\nEntry points: extensions/task/index.ts\nPatterns: Schema-driven tool contracts (TypeBox schemas enforce parameter shapes); Event-driven model swapping (prewalk → execute on first edit); Budget tier resolution chain (locked flag > locked param > config default > auto heuristic); jj workspace isolation for parallel workers (squash-merge into task base); Fork-and-prune adversarial review (forked session inherits context, strips commitment); Typed schema validation at RPC boundaries (yield, findings, spec parsing)\nTest layout: Fast hermetic tests live in test*.ts under extensions/task/ and packages/core-v2/test/ — run via `mise run test` (zero LLM, ≤15s); real-LLM e2e tests (test-e2e.ts, smoke-*.ts) run manually with a 15min timeout.\n\nFiles most relevant to this task (read these first):\n- extensions/task/batch.ts — Provider-agnostic batch inference lane (M2): turns a spec's typed requirements into single-turn prompt items each with an output contract, submits/polls/collects async batch jobs, validates output (code-fence stripping), extracts files with path-safety/conflict checks, and reports typed error codes (submit_failed, poll_timeout, aborted, ...). [symbols: BATCH_ERROR_CODES, BatchErrorCode, BatchErrorDetail, BatchError, BatchOutputContract, BatchPromptItem, stripCodeFences, m, validateBatchOutput, text, value, missing, BATCH_FILE_CONTRACT, BatchFile, requirementId]\n- packages/core-v2/src/daemon/parallel.ts — M2.d parallel pipeline: fan-out N worker sessions across WorkspaceDriver workspaces; combine task-base/feature-branch; verify once on integrated tree; per-attempt + aggregate receipts; conflicts -> escalated. [symbols: SessionHost, UsageSnapshot, RunParallelOptions, WorkerObservation, attemptNumberOf, m, makeReceipt, RunParallelResult, store, env, gateway, pluginHookCtx, parsed, familyBase, isTaskBase]\n- extensions/task/test-index.ts — Hermetic test suite for index.ts — covers budget schema-locking (budget param removed when tier locked), sub_specs-vs-spec precedence and normalization, budget tier resolution including dynamic tier sets (Phase 11), session goals/budget override persistence from session entries, TaskResult-to-tool-return mapping, review-report rendering with finding caps, TUI renderInPlace reuse contract, and the run-plan + progress-state render pipeline. [symbols: BudgetTierConfig, ProgressState, RunPlan, TaskToolReturn, schemaProperties, json, tierConfig, propertyJson, fakeResult, fakeWorker, testBudgetSchemaLocking, check, unlocked, locked, budgetSchema]\n- extensions/task/index.ts — pi-task extension entry point (Phase 9) — registers the `task` tool (orchestrator interface), enforces budget via `--task-budget`/`/task-budget`, persists session goals, injects codebase-map and workflow-contract into the main prompt, streams TUI progress (zero LLM tokens), and supports detached dispatch. Contains pure budget-resolution helpers, schema builders, result mappers, and progress render logic. [symbols: DetachedRunOptions, BudgetMode, BudgetTier, BudgetTierConfig, TaskConfig, isLockedBudget, normalizeBudgetMode, resolveBudgetMode, autoTierForRequirements, small, large, heuristicDefaultTier, names, notMax, countSpecRequirements]\n- extensions/task/orchestrator.ts — Orchestrates task execution, spawns workers, validates specs, and merges workspace commits. [symbols: WorkerFailureDiagnostics, WorkerResult, SwapInfo, BatchLaneConfig, SandboxConfig, TaskShape, BatchFile, BatchProvider, MergeMetrics, RunManifest, MergeFailureInfo, writeMergeFailureArtifact, artifact, FAILURE_PATH_JJ_TIMEOUT_MS, isFinalizationIncomplete]\n- packages/core-v2/test/test-plugins-lifecycle.ts — Hermetic suite for the lifecycle-collector trigger plugin that exercises the REAL plugin path: loaded by path from task.toml, trigger side wired through a real InMemoryTaskGateway subscription, and hook isolation (throwing plugins never crash the pipeline). [symbols: PLUGIN_PATH, TASK_RUNNER_PATH, rowsDouble, errors, check, dir, tomlPath, plugins, plugin, gateway, journal, secondSeen, off, journalLenBefore, freshGateway]\n- extensions/task/test-review.ts — Hermetic tests for review.ts's pure logic: buildReviewPrompt, settleReview outcome settling, and the first-event/no-progress watchdog decisions (forkedReview itself is e2e territory since it spawns a real pi process). [symbols: REVIEW, errors, check, prompt, mk, finding, single, merged, worst, empty, ok, fail, quiet, long, detail]\n- extensions/task/progress.ts — Provides pure functions for generating task-tool progress view, including dispatch plan and phase sequencing. [symbols: WorkerPhase, PlanPhase, RunPlan, BuildRunPlanOptions, buildRunPlan, phases, initialPhaseOf, WorkerProgress, ProgressState, createProgressState, workers, phase, recountDone, n, extractFileScope]\n- extensions/task/test-e2e.ts — Consolidated real-LLM end-to-end test (run manually, ~3-5 min) exercising real pi + a single cheap hardcoded model across steering, sub-specs, fix-loop, and verify paths. [symbols: MODEL, TASK_EXTENSION_PATH, SINGLE_TOOL_PROMPT, SUBSPECS_TOOL_PROMPT, SYSTEM_PROMPT, HELLO_SPEC, MULTI_SPEC, FIXLOOP_SPEC, NO_VERIFY_PROMPT, errors, check, makeRepo, dir, parseRemaining, m]\n- packages/core-v2/src/daemon/task-runner.ts — M1.4 pipeline entry: validate->route->host->yield->verify->persist receipt; deterministic prefix (R5) pure of timestamps; bundle telemetry; plugin extractions (handoff-cap + lifecycle-collector moved out of core). [symbols: USAGE_UNAVAILABLE, UsageSnapshot, totalInputTokens, computeCor, estimateGroundingTokens, emptyUsage, stats, sumUsage, total, receiptUsageFields, SpecValidationError, ParsedTaskSpec, parseTaskSpec, sections, current]\n- extensions/task/test-sandbox.ts — Sandbox builder tests: pure arg-set unit tests (strict ro-root policy, bind ordering, cwd/agent-dir shadowing) plus one guarded real-bwrap probe. [symbols: findTriple, sandboxConfig, jj, bwrapAvailable, testDisabled, check, args, testStrictPolicy, roRoot, dev, proc, tmp, testRwBinds, cwd, agentDir]\n- packages/core-v2/src/plugins/hooks.ts — Isolated per-call plugin hooks: try/catch with error sink, never crashes pipeline; transformExecutionBundle/Handoff re-validated against zod schemas; emitLifecycleEvent fan-out; registerPluginTriggers isolated. [symbols: PluginHookContext, report, sink, current, emitLifecycleEventToPlugins, registerPluginTriggers, callIsolatedSync]\n- packages/core-v2/src/grounding/prewalk.ts — Prewalk cost-break-even policy (M3, OFF by default): pure function decidePrewalkSwap computes whether swapping from strong to cheap model saves money given context size, remaining turns, and per-model pricing; attachPrewalk wires it to a session handle to fire once on first successful edit/write. [symbols: TokenPricing, PrewalkPricing, PrewalkDecisionInput, PrewalkDecision, DEFAULT_OUTPUT_TOKENS_PER_TURN, perM, decidePrewalkSwap, ctx, n, out, strong, cheap, stayCostUsd, swapPenaltyUsd, swapCostUsd]\n- extensions/task/test-checklist-relay.ts — Hermetic tests for the checklist relay reducer against an in-process fake session, asserting R4 progress tracking with zero LLM tokens and observer-only behavior. [symbols: ChecklistProgress, ChecklistRelayState, start, end, FakeSession, makeFakeSession, listeners, commands, session, i, errors, check, state, i1, i2]\n- docs/pi-task-testing-spec.md — Phase 6.5 spec: replace six LLM-driven smoke tests with a fast hermetic deterministic suite (pure-unit + real-jj/real-bash, zero LLM, <=15s) plus one manually run consolidated e2e. [symbols: interface, function]\n\nRead the files above first; explore beyond them only if they do not cover the task.\n\n## Goal\nWrite a single file confirming the task pipeline works end-to-end.\n\n## Requirements\n- R1: Create the file /tmp/pi-task-smoke-test.txt containing exactly the text \"ok\" (a trailing newline is fine).\n\n## Verification\ntest -f /tmp/pi-task-smoke-test.txt" | |
| } | |
| ] | |
| }, | |
| { | |
| "role": "assistant", | |
| "content": null, | |
| "tool_calls": [ | |
| { | |
| "id": "af7699f1-204a-4a02-be45-a406adee9283", | |
| "type": "function", | |
| "function": { | |
| "name": "write", | |
| "arguments": "{\"path\":\"/tmp/pi-task-smoke-test.txt\",\"content\":\"ok\\n\"}" | |
| } | |
| } | |
| ], | |
| "reasoning_content": "" | |
| }, | |
| { | |
| "role": "tool", | |
| "content": "Successfully wrote 3 bytes to /tmp/pi-task-smoke-test.txt", | |
| "tool_call_id": "af7699f1-204a-4a02-be45-a406adee9283" | |
| }, | |
| { | |
| "role": "assistant", | |
| "content": null, | |
| "tool_calls": [ | |
| { | |
| "id": "f3622891-58ff-4ae4-8fd1-acc6cb5e2691", | |
| "type": "function", | |
| "function": { | |
| "name": "bash", | |
| "arguments": "{\"command\":\"test -f /tmp/pi-task-smoke-test.txt && cat /tmp/pi-task-smoke-test.txt\"}" | |
| } | |
| } | |
| ], | |
| "reasoning_content": "" | |
| }, | |
| { | |
| "role": "tool", | |
| "content": "ok\n", | |
| "tool_call_id": "f3622891-58ff-4ae4-8fd1-acc6cb5e2691" | |
| } | |
| ] |
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment