Skip to content

Instantly share code, notes, and snippets.

@Helw150
Created July 18, 2026 23:04
Show Gist options
  • Select an option

  • Save Helw150/10aa5fe8533b4139512df485ecc717ec to your computer and use it in GitHub Desktop.

Select an option

Save Helw150/10aa5fe8533b4139512df485ecc717ec to your computer and use it in GitHub Desktop.
Grug step-630 MATH500 full-context results

Grug step-630 · MATH500 full-context results

Metric Value
Accuracy 54.2%
Solved 271 / 500
Example limit none
Model context 32,768 tokens
Generation allowance 28,672 tokens

Checkpoint:

s3://marin-us-east-02a/marin/exports/grug/june-67b-a2b-sft-s2-thinking/step-630/hf-bf16-vllm/

Tokenizer:

penfever/grug-67b-a2b-sft-s2-thinking-step630-tok

The evaluator recorded limit: null. max_gen_toks=28672 was passed both in the local-chat model arguments and generation arguments, avoiding the earlier 256-token backend fallback. The only remaining bound was the model's 32,768-token serving context.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment