Run completed on 2026-07-02 against the ctm-vllm-small service using:
- Model:
nvidia/Gemma-4-31B-IT-NVFP4 - API path/model alias:
/small/v1,gemma4 - Max context:
262144 - GPU layout: 6 GPUs, data parallel 3, tensor parallel 2
- Benchmark: DeepSWE full run
- Job:
deepswe-gemma4-256k-full-20260701-223016
- Total completed trials: 113 / 113
- Errored trials: 11
- Scored eval trials: 112
- Solved tasks / reward: 0
- Partial score: 0.5744747444
- f2p: 0.0738640636
- p2p: 0.7902985828
- Started:
2026-07-01T22:30:18.403753 - Finished:
2026-07-02T06:28:55.670597 - Runtime: about 7h 58m 37s
- Input tokens: 877,229,315
- Cache tokens: 854,443,808
- Output tokens: 4,186,337
- End-to-end benchmark output throughput: about 145.8 output tok/s
This throughput is the full DeepSWE agent wall-clock rate, not the raw small-context vLLM serving rate.
VerifierTimeoutError: 1AgentTimeoutError: 6NonZeroAgentExitCodeError: 4
The move from 128k to 256k fixed the earlier 128k context failures, but 4 tasks still hit the 256k cap because the prompt plus the 8192-token output budget exceeded 262144.
Context-limit error tasks:
fastapi-deprecation-response-hea__C27FXDEkoota-entity-snapshot-rollback__ZiwKUPHscriggo-method-declarations__tSrhb6Xsql-formatter-bigquery-pipe-form__3gh8Lic
result.json: full benchmark result JSONdeepswe-gemma4-256k-full-20260701-223016.clean.20260701-223016.console.log: cleaned console log