Reference for reviewing Horizon evaluation tasks. Applies to all tasks (standalone and subtasks).
For a task to be acceptable, it must meet all criteria:
| Criterion | Threshold | Verification |
|---|---|---|
| Solvable | solution.sh passes grader | horizon test-solution <task-dir> |
| Challenging | Mean score ≤0.85 (teapot) or <0.50 (docker) on biggie-nebula |
horizon evaluations run <task-dir> --runs 8 --model biggie-nebula |
| Substantial | ≥4 hours for senior engineer | Manual assessment |
The difficulty threshold depends on the evaluation backend:
| Backend | Threshold | How to identify |
|---|---|---|
| Teapot (hosted/firecracker) | mean score ≤0.85 | First user message is the bare task.yaml prompt — no outer wrapper, no <problem> tags, no <TASK_DONE> instruction, no Initial prompt: JSON envelope. Just the raw prompt text. |
| Docker (local) | mean score <0.50 | First user message has a structured wrapper around the task.yaml content: an Initial prompt: JSON envelope, <TASK_DONE> end-of-task instruction, <problem>/</problem> tags around the task content, and a working directory is /workdir note. |
To determine which backend was used for an eval run, check the first user message in any transcript:
- Docker wraps the prompt:
Initial prompt: {'prompt': '...<TASK_DONE>...<problem>[task.yaml content]</problem>...working directory is /workdir...'} - Teapot passes the raw task.yaml content with no envelope or tags
Note: some tasks (e.g., mail tasks) naturally reference Nebula services like Maddy, Mattermost, or wiki URLs in their task.yaml — these are part of the task scenario, not backend injection. The presence of <problem> tags and the <TASK_DONE> instruction is what distinguishes docker from teapot.
Task acceptance requires evaluation with biggie-nebula. You can optionally use smalli-nebula during development for faster iteration, but final acceptance is based on biggie-nebula results.
| Model | Runs | Pass rate target | Purpose |
|---|---|---|---|
biggie-nebula |
8 | ≤0.85 (teapot) or <0.50 (docker) | Required for task acceptance |
smalli-nebula |
8–10 | mean score <0.50 | Optional, development iteration only |
The threshold is the mean score across all runs, not the percentage of runs scoring 1.0. For example, scores of [1.0, 0.5, 1.0, 0.0, 1.0, 0.5, 0.5, 0.0, 1.0, 0.5] have a mean of 0.60, which passes both backend thresholds.
biggie-nebula: Run 8 evaluations; target depends on backend (see Backend-Specific Thresholds above)- 0.0 mean score is only acceptable when failures are clearly genuine difficulty, not due to task clarity, blockers, timing, or grader issues — otherwise treat 0.0 as NEEDS WORK
- Mean score above the backend-specific threshold indicates task is too easy
Genuine failures (task is appropriately difficult):
- Multiple different failure modes across runs
- Agent attempts reasonable approaches
- Failures occur at different stages
- Agent runs out of turns mid-progress
Artificial failures (task or grader needs fixing):
- All evals fail the same check
- Agent completes work but grader fails
- Timeouts, connection refused errors
- Services not ready when grader checks
Failure Analysis Flow:
All fail same check? → Yes → Grader timing issue
→ No → Agent tried correct approach?
→ No → Task unclear, revise task.yaml
→ Yes → Multiple failure modes?
→ Yes → Genuine (task is good)
→ No → Investigate blocker
Graders must wait for async operations to complete before checking.
| Scenario | Minimum | Maximum |
|---|---|---|
| Single pod ready | 30s | 3 min |
| Deployment rollout | 1 min | 5 min |
| Full service mesh | 3 min | 10 min |
| CI/CD pipeline completion | 2 min | 10 min |
| Complex multi-service | 5 min | 15 min (apex max) |
def wait_for_condition(check_fn, timeout=300, interval=10):
"""Standard wait pattern for graders."""
start = time.time()
while time.time() - start < timeout:
if check_fn():
return True
time.sleep(interval)
return False
# Example: wait for deployment
def wait_for_deployment(name, namespace, timeout=300):
def check():
code, _, _ = run(f"kubectl rollout status deployment/{name} -n {namespace} --timeout=10s")
return code == 0
return wait_for_condition(check, timeout)Partial grading is preferred. When a grader uses subscores, verify:
- Each subscore represents a working milestone, not a prerequisite (no "file exists" or "no syntax errors" checks)
- Each subscore is objective — deterministic, measurable
- Each subscore is not gameable — agent can't get reward without actual work
- All weights are equal (e.g., 3 subscores = 0.33 each) — this is a platform requirement, not a tuning knob. Never recommend reweighting. If "easy" checks dilute the difficulty signal, recommend grouping related checks into a single subscore instead. Rounding to sum to exactly 1.0 is expected and not a violation (e.g.,
.33 + .33 + .34 = 1.0for thirds,.25 + .25 + .25 + .25 = 1.0for quarters). Do not flag rounding differences as unequal weights. - Agent cannot score > 0.3 without meaningful progress
Quick test for each subscore:
- "If agent ONLY gets this subscore, did they make real progress?" → must be YES
- "Can agent get this reward without working toward the actual goal?" → must be NO
If the grader uses binary scoring, consider whether partial grading would provide better signal.
Variance check: At least one subscore must show variance across 10 rollouts (the reviewer bot enforces this). If all subscores are identical across every run, the task is either trivially easy (all 1s) or impossibly hard (all 0s) — see Subscore Variance Requirement.
- Checks align with task.yaml deliverables
- Requirements are inferrable by senior DevOps engineer from task.yaml
- Grader doesn't check for specific implementation details (multiple valid approaches)
- Error messages are descriptive for debugging
- Clear, specific objective
- Sufficient context about environment
- Doesn't reveal grading criteria (no gaming)
- Doesn't overspecify solution approach
- Includes standard reminders (no interactive commands, etc.)
- Scope is substantial (≥4 hours)
- Task feels cohesive (not disconnected pieces)
- Clear start and end states
- Agent has necessary permissions/access
The agent's only source of task intent should be task.yaml. Everything else in the environment should require genuine investigation and expertise to interpret. Review these categories for leaks:
- Filesystem: Dockerfile does not COPY
solution.shorgrader.pyinto the image (harness mounts them externally at runtime) - task.yaml specificity: Prompt describes objectives, not exact grader checks. Agent should need expertise to determine how to achieve each objective, not just follow a mechanical checklist
- Setup artifacts: Resources created by
setup.sh(annotations, labels, ConfigMaps, comments) don't explain the intended fix or describe what's broken in a way that bypasses investigation - Git repos: Gitea repositories don't contain committed runbooks, solution files, or incident docs that hand the agent the answer
- Environment variables: No env vars in Dockerfile or setup.sh that name the solution approach or expected end state
- Prior run artifacts:
test-solutionruns don't leave behind temp files, logs, or K8s resources visible to agents in subsequent evals
| Problem | Diagnosis | Fix |
|---|---|---|
| All evals fail same check | Grader timing | Add wait with appropriate timeout |
| Agent does wrong thing | Unclear task | Clarify task.yaml deliverables |
| Solution passes, evals fail | Race condition | Increase grader wait times |
| Mean score above threshold | Task too easy | Increase scope or remove hints |
| Agent runs out of turns | Scope too broad | Consider breaking into subtasks |
| solution.sh/grader.py in image | Dockerfile COPYs answer files | Remove COPY lines; harness mounts them at runtime |
| Bootstrap errors | Environment issue | docker system prune -af && rm -rf horizon_env && ./apex-arena-install.sh |
| test-solution hangs locally but passes hosted | Local machine has internet; air-gapped commands (e.g., docker push) hang instead of failing fast |
Not a task bug — see Local vs Hosted |
- Verify solvability:
horizon test-solution <task-dir> - Run evals:
horizon evaluations run <task-dir> --runs 8 --model biggie-nebula - Analyze failures: Categorize as genuine vs artificial
- Check grader: Review wait times, alignment with task.yaml
- Assess task.yaml: Clarity, scope, no gaming potential
- Iterate: Fix issues, re-run evals until criteria met
- Task Eval Analysis Guide — Deep dive on analyzing eval results
- Subtask Review Additions — Additional criteria for subtasks
- Creating Horizon Tasks — Task creation reference