Concept map for notes/viktor/synthetic-agentic-benchmarks/. Built 2026-07-13.
Survey scaffold for the whole space: [[llm-agent-eval-benchmarking-survey]] (taxonomy of eval objectives, interaction modes, difficulty levers, and named benchmarks β read it first to place any single note).
Two failure modes recur when you auto-generate agentic benchmarks: