Skip to content

Instantly share code, notes, and snippets.

@arubis
Created June 9, 2026 19:18
Show Gist options
  • Select an option

  • Save arubis/b59163ada3f7015d60afc1bb7333a985 to your computer and use it in GitHub Desktop.

Select an option

Save arubis/b59163ada3f7015d60afc1bb7333a985 to your computer and use it in GitHub Desktop.

Task Review Guide

Reference for reviewing Horizon evaluation tasks. Applies to all tasks (standalone and subtasks).


Acceptance Criteria

For a task to be acceptable, it must meet all criteria:

Criterion Threshold Verification
Solvable solution.sh passes grader horizon test-solution <task-dir>
Challenging Mean score ≤0.85 (teapot) or <0.50 (docker) on biggie-nebula horizon evaluations run <task-dir> --runs 8 --model biggie-nebula
Substantial ≥4 hours for senior engineer Manual assessment

Backend-Specific Thresholds

The difficulty threshold depends on the evaluation backend:

Backend Threshold How to identify
Teapot (hosted/firecracker) mean score ≤0.85 First user message is the bare task.yaml prompt — no outer wrapper, no <problem> tags, no <TASK_DONE> instruction, no Initial prompt: JSON envelope. Just the raw prompt text.
Docker (local) mean score <0.50 First user message has a structured wrapper around the task.yaml content: an Initial prompt: JSON envelope, <TASK_DONE> end-of-task instruction, <problem>/</problem> tags around the task content, and a working directory is /workdir note.

To determine which backend was used for an eval run, check the first user message in any transcript:

  • Docker wraps the prompt: Initial prompt: {'prompt': '...<TASK_DONE>...<problem>[task.yaml content]</problem>...working directory is /workdir...'}
  • Teapot passes the raw task.yaml content with no envelope or tags

Note: some tasks (e.g., mail tasks) naturally reference Nebula services like Maddy, Mattermost, or wiki URLs in their task.yaml — these are part of the task scenario, not backend injection. The presence of <problem> tags and the <TASK_DONE> instruction is what distinguishes docker from teapot.


Eval Quality

Model Requirements

Task acceptance requires evaluation with biggie-nebula. You can optionally use smalli-nebula during development for faster iteration, but final acceptance is based on biggie-nebula results.

Model Runs Pass rate target Purpose
biggie-nebula 8 ≤0.85 (teapot) or <0.50 (docker) Required for task acceptance
smalli-nebula 8–10 mean score <0.50 Optional, development iteration only

Mean Score Assessment

The threshold is the mean score across all runs, not the percentage of runs scoring 1.0. For example, scores of [1.0, 0.5, 1.0, 0.0, 1.0, 0.5, 0.5, 0.0, 1.0, 0.5] have a mean of 0.60, which passes both backend thresholds.

  • biggie-nebula: Run 8 evaluations; target depends on backend (see Backend-Specific Thresholds above)
  • 0.0 mean score is only acceptable when failures are clearly genuine difficulty, not due to task clarity, blockers, timing, or grader issues — otherwise treat 0.0 as NEEDS WORK
  • Mean score above the backend-specific threshold indicates task is too easy

Genuine vs Artificial Failures

Genuine failures (task is appropriately difficult):

  • Multiple different failure modes across runs
  • Agent attempts reasonable approaches
  • Failures occur at different stages
  • Agent runs out of turns mid-progress

Artificial failures (task or grader needs fixing):

  • All evals fail the same check
  • Agent completes work but grader fails
  • Timeouts, connection refused errors
  • Services not ready when grader checks
Failure Analysis Flow:

All fail same check? → Yes → Grader timing issue
                   → No  → Agent tried correct approach?
                              → No  → Task unclear, revise task.yaml
                              → Yes → Multiple failure modes?
                                        → Yes → Genuine (task is good)
                                        → No  → Investigate blocker

Grader Quality

Wait Time Requirements

Graders must wait for async operations to complete before checking.

Scenario Minimum Maximum
Single pod ready 30s 3 min
Deployment rollout 1 min 5 min
Full service mesh 3 min 10 min
CI/CD pipeline completion 2 min 10 min
Complex multi-service 5 min 15 min (apex max)

Wait Pattern

def wait_for_condition(check_fn, timeout=300, interval=10):
    """Standard wait pattern for graders."""
    start = time.time()
    while time.time() - start < timeout:
        if check_fn():
            return True
        time.sleep(interval)
    return False

# Example: wait for deployment
def wait_for_deployment(name, namespace, timeout=300):
    def check():
        code, _, _ = run(f"kubectl rollout status deployment/{name} -n {namespace} --timeout=10s")
        return code == 0
    return wait_for_condition(check, timeout)

Partial Grading Quality

Partial grading is preferred. When a grader uses subscores, verify:

  • Each subscore represents a working milestone, not a prerequisite (no "file exists" or "no syntax errors" checks)
  • Each subscore is objective — deterministic, measurable
  • Each subscore is not gameable — agent can't get reward without actual work
  • All weights are equal (e.g., 3 subscores = 0.33 each) — this is a platform requirement, not a tuning knob. Never recommend reweighting. If "easy" checks dilute the difficulty signal, recommend grouping related checks into a single subscore instead. Rounding to sum to exactly 1.0 is expected and not a violation (e.g., .33 + .33 + .34 = 1.0 for thirds, .25 + .25 + .25 + .25 = 1.0 for quarters). Do not flag rounding differences as unequal weights.
  • Agent cannot score > 0.3 without meaningful progress

Quick test for each subscore:

  • "If agent ONLY gets this subscore, did they make real progress?" → must be YES
  • "Can agent get this reward without working toward the actual goal?" → must be NO

If the grader uses binary scoring, consider whether partial grading would provide better signal.

Variance check: At least one subscore must show variance across 10 rollouts (the reviewer bot enforces this). If all subscores are identical across every run, the task is either trivially easy (all 1s) or impossibly hard (all 0s) — see Subscore Variance Requirement.

Grader Alignment

  • Checks align with task.yaml deliverables
  • Requirements are inferrable by senior DevOps engineer from task.yaml
  • Grader doesn't check for specific implementation details (multiple valid approaches)
  • Error messages are descriptive for debugging

Task Quality

task.yaml Assessment

  • Clear, specific objective
  • Sufficient context about environment
  • Doesn't reveal grading criteria (no gaming)
  • Doesn't overspecify solution approach
  • Includes standard reminders (no interactive commands, etc.)

Scope Assessment

  • Scope is substantial (≥4 hours)
  • Task feels cohesive (not disconnected pieces)
  • Clear start and end states
  • Agent has necessary permissions/access

Information Isolation

The agent's only source of task intent should be task.yaml. Everything else in the environment should require genuine investigation and expertise to interpret. Review these categories for leaks:

  • Filesystem: Dockerfile does not COPY solution.sh or grader.py into the image (harness mounts them externally at runtime)
  • task.yaml specificity: Prompt describes objectives, not exact grader checks. Agent should need expertise to determine how to achieve each objective, not just follow a mechanical checklist
  • Setup artifacts: Resources created by setup.sh (annotations, labels, ConfigMaps, comments) don't explain the intended fix or describe what's broken in a way that bypasses investigation
  • Git repos: Gitea repositories don't contain committed runbooks, solution files, or incident docs that hand the agent the answer
  • Environment variables: No env vars in Dockerfile or setup.sh that name the solution approach or expected end state
  • Prior run artifacts: test-solution runs don't leave behind temp files, logs, or K8s resources visible to agents in subsequent evals

Common Issues & Fixes

Problem Diagnosis Fix
All evals fail same check Grader timing Add wait with appropriate timeout
Agent does wrong thing Unclear task Clarify task.yaml deliverables
Solution passes, evals fail Race condition Increase grader wait times
Mean score above threshold Task too easy Increase scope or remove hints
Agent runs out of turns Scope too broad Consider breaking into subtasks
solution.sh/grader.py in image Dockerfile COPYs answer files Remove COPY lines; harness mounts them at runtime
Bootstrap errors Environment issue docker system prune -af && rm -rf horizon_env && ./apex-arena-install.sh
test-solution hangs locally but passes hosted Local machine has internet; air-gapped commands (e.g., docker push) hang instead of failing fast Not a task bug — see Local vs Hosted

Review Workflow

  1. Verify solvability: horizon test-solution <task-dir>
  2. Run evals: horizon evaluations run <task-dir> --runs 8 --model biggie-nebula
  3. Analyze failures: Categorize as genuine vs artificial
  4. Check grader: Review wait times, alignment with task.yaml
  5. Assess task.yaml: Clarity, scope, no gaming potential
  6. Iterate: Fix issues, re-run evals until criteria met

See Also

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment