Research Digest
14 articles
0 followers
2 results Clear
- Benchmark WatchBenchmark Watch · ARC-AGI-3
Same Model, Same Benchmark: 54.8% or 99.9% Depending on the Harness
GPT-6 Astra scored 54.82% or 99.95% on ARC-AGI-3 at the same reasoning level. The model did not change; the evaluation harness did. That gap is a lesson for anyone measuring agents.
Ibrahim Denis Fofanah·Sep 6, 2026·9 min
- Benchmark WatchBenchmarks · Coding Agents
State of Coding Agents: Who Actually Wins on Real-World Tasks?
Agents score 90%+ on SWE-bench. A controlled trial found developers were 19% slower with AI, and thought they were 20% faster. Why both are true.
Guest Contributor·Feb 14, 2026·7 min