GLM-5.2 completed the full Datacurve DeepSWE benchmark with 23 full solves out of 113 tasks.
- Benchmark: DeepSWE full run, 113 tasks
- Agent harness: Pier + mini-swe-agent
- Model endpoint: OpenAI-compatible
https://rig.ctmdev.us/v1, served modelGLM-5.2 - Job:
deepswe-glm52-test-full - Completed trials: 113 / 113
- Full solves / reward=1: 23 / 113
- Official reward: 0.2035398230
- Partial score: 0.8813117941
- f2p: 0.6833060027
- p2p: 0.9895487289
- Errored trials: 14
- Total duration: 22h 58m 16s
- Started:
2026-07-02T16:18:28.624239 - Finished:
2026-07-03T15:16:44.937405
- Eval key:
mini-swe-agent__GLM-5.2__tasks - Scored eval trials: 112
- Eval errors: 14
- Concurrency: 2 Pier trials (
--n-concurrent 2) - Model config:
benchmarks/deepswe/glm.pier.yaml - Task source: local DeepSWE checkout under
benchmarks/deepswe/deep-swe/tasks - DeepSWE setup ref used by runner:
3cda4081fed96103a6395de39c85e9b20275e307 - Output throughput over end-to-end benchmark wall clock: 58.42 output tokens/s
Token totals reported by Pier:
- Input tokens: 2,785,220,197
- Cache tokens: 2,766,974,976
- Output tokens: 4,830,867
Benchmark requests were served from rig.ctmdev.us.
- CPU: AMD Ryzen Threadripper PRO 9965WX, 24 cores / 48 threads
- Memory: 251 GiB RAM, 71 GiB swap
- Installed GPUs: 6x NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 97,887 MiB each
- NVIDIA driver: 595.80
- GPU power limit: 300 W per GPU
- Storage mount for run host:
/dev/nvme2n1p3, 1.9 TiB filesystem
Active GLM-5.2 serving configuration during the run:
- Checkpoint mount:
/active-models/madeby561--GLM-5.2-NVFP4-REAP-504B-term - GPUs used by vLLM: 0, 1, 2, 3
- Tensor parallel: 4
- Pipeline parallel: 1
- Decode context parallel: 4
- Max model length: 250,000
- Max sequences: 2
- Max batched tokens: 16,384
- KV cache dtype: FP8
- Quantization:
modelopt_fp4 - Attention backend:
B12X_MLA_SPARSE - MoE backend:
b12x - Speculative decoding: same-model MTP, 3 speculative tokens
- Prefix caching: enabled
- Chunked prefill: enabled
- Reward: 0.20353982300884957 (23 full solves / 113 total tasks)
- Partial: 0.88131179414746
- f2p: 0.6833060027158157 (36.60176991150443 / 51.56637168141593)
- p2p: 0.9895487288954923 (2045.1150442477876 / 2045.3097345132744)
Direct per-trial reward counts:
reward=1: 23reward=0: 89- no verifier reward due to error: 1
- NonZeroAgentExitCodeError: 11
- AgentTimeoutError: 2
- VerifierTimeoutError: 1
arcane-drift-detection-baselinesbandit-interprocedural-taint-checksclack-async-autocomplete-optionsclaude-code-by-agents-recursive-delegationdasel-html-document-formatdrizzle-orm-window-function-buildersdynamodb-toolbox-conditional-attribute-requirementshappy-dom-abort-pending-body-readsnarwhals-rolling-window-suitepsd-tools-blend-range-apiquery-persist-restored-query-statescc-bounded-memory-spillingskrub-duration-encodingsql-formatter-bigquery-pipe-formattingsqlite-utils-safe-import-checkpointstomlkit-toml-table-converterstrue-myth-iterable-collection-combinatorsvitest-duration-shardingvulture-persistent-analysis-cachewasmi-trap-coredumpsyaegi-go-embed-directivesyjs-map-conflict-detectionytt-jsonpath-query-api
oxvg-structural-selector-preservation: AgentTimeoutErrortextual-kitty-key-phases: AgentTimeoutErroradaptix-name-mapping-aliases: NonZeroAgentExitCodeErroranko-typed-variable-bindings: NonZeroAgentExitCodeErroreffect-sse-httpapi-streaming: NonZeroAgentExitCodeErrorfastapi-deprecation-response-headers: NonZeroAgentExitCodeErrorhelm-array-merge-strategies: NonZeroAgentExitCodeErrormashumaro-flattened-dataclass-fields: NonZeroAgentExitCodeErrormnamer-daemon-watch-lifecycle: NonZeroAgentExitCodeErroropa-rego-rule-profiling: NonZeroAgentExitCodeErrorpython-statemachine-state-data-scoping: NonZeroAgentExitCodeErrorreturns-validated-error-accumulation: NonZeroAgentExitCodeErrorsqlfmt-create-table-ddl-formatting: NonZeroAgentExitCodeErrorlangchain-request-coalescing: VerifierTimeoutError
The aggregate Pier result was stored locally at:
benchmarks/deepswe/jobs/deepswe-glm52-test-full/result.json
This gist intentionally publishes the benchmark summary and derived task lists, not the full per-step agent trajectories.