Date: 2026-09-21. Model identifier: gpt-6-astra. Reasoning efforts: low, medium, high, xhigh, and max. This is an exploratory, independently run evaluation, not an official OpenAI or Stockfish benchmark.
We played 40 scored games per effort, 200 games total, against strength-limited Stockfish 17.1. Each game used a fresh Codex CLI session and one continuous agent turn, with chess moves and opponent replies exchanged through an MCP tool. Opponent ratings adapted separately for each effort. A separate preliminary medium-versus-1320 game informed opponent selection but is excluded from the reported estimates.
The resulting numbers measure performance against this particular Stockfish UCI_Elo ladder and harness. They are not calibrated FIDE, USCF, Lichess, or Chess.com ratings.