3,200 papers indexed across cs.LG, cs.CV, cs.CL, cs.AI, cs.RO and friends.
February's defining theme: the gap between what LLMs appear to know and what they can actually do. Multiple papers converged on the same uncomfortable finding — scale isn't fixing the reasoning problem, it's just hiding it better. Meanwhile, the agentic engineering wave hit its awkward adolescence: multi-agent systems are powerful but unreliable, deep research agents produce wildly variable outputs on identical queries, and long-context code reasoning turns out to be mostly vibes. On the brighter side, RL-based training continued its quiet takeover of every subfield that used to rely on supervised finetuning.