A step-by-step guide to reproduce cold-start GRPO training on
SciKnowEval with
prime-rl. By the end you
will have RL-trained Qwen3-8B from scratch (no SFT warmstart) on
multi-domain scientific MCQ and evaluated the result.
This guide is written from an actual run (Aug 2026) on an 8×H100 node. Everything here was verified to work. Companion to the