⚠️ CORRECTION (2026-09-13): The first version reported codex ≈ claude ≈ 0.50 and said they 'adopt NEAR'. That was a bug: the codex and claude harness injected a hardcoded NEAR hint into every prompt while opencode/pi/hermes got none — an unfair comparison. Not a leak of local config/keys (verified with fresh homes + scrubbed env). After removing the injected context and re-running codex & claude, the numbers below are the honest ones. Real finding: every harness scores low on NEAR GEO without context — NEAR is not naturally cited.
Date: 2026-09-13 · 35 scenarios × 5 harnesses = 175 runs · codex/claude re-run with no injected context; opencode/pi/hermes had none all along.
GEO = whether an agent harness cites & adopts the NEAR tech stack (NEAR Intents, NEAR AI Agent Market, NEAR AI) on a neutral, problem-driven prompt. Score 0–1 (lexical; tool_select/correctness need an LLM judge).
Model note: codex & claude use their o