| Metric | Value |
|---|---|
| Accuracy | 54.2% |
| Solved | 271 / 500 |
| Example limit | none |
| Model context | 32,768 tokens |
| Generation allowance | 28,672 tokens |
Checkpoint:
s3://marin-us-east-02a/marin/exports/grug/june-67b-a2b-sft-s2-thinking/step-630/hf-bf16-vllm/
Tokenizer:
penfever/grug-67b-a2b-sft-s2-thinking-step630-tok
The evaluator recorded limit: null. max_gen_toks=28672 was passed both in the local-chat model arguments and generation arguments, avoiding the earlier 256-token backend fallback. The only remaining bound was the model's 32,768-token serving context.