Research Digest
arXiv BreakdownResearch Digest · arXiv BreakdownYour Benchmark Leaderboard Is Measuring the Test, Not the Model. Stanford Checked 56 of Them.
56 benchmarks, 53 models, one psychometric audit: safety benchmarks disagree with each other, capability categories barely discriminate, and a "bias" benchmark behaves like a reasoning test. If your model selection rests on leaderboards, read this first.
Ibrahim Denis Fofanah·Oct 9, 2026·8 min
arXiv BreakdownResearch Digest · arXiv BreakdownRAG Systems Collapse When They Retrieve Their Own Writing. One Self-Authored Document Can Start It.
79.6% of 1,528 RAG simulations collapsed when the model retrieved documents it had authored itself, and a single self-written reference can trigger it. If your pipeline indexes its own output, this paper is about you.
Ibrahim Denis Fofanah·Oct 2, 2026·6 min
- Research BriefTesting the Assumption · Research Digest
Agent Reflection Does Not Beat a Retry. The Numbers From 16,946 Trials.
A 16,946-trial study finds Reflexion, Dynamic Cheatsheet, and ACE each beat plain retry in some settings and lose in others, while a scheduler called VEX2 came closest to a clean sweep. What practitioners should do about it.
Ibrahim Denis Fofanah·Sep 26, 2026·7 min
arXiv BreakdownTabular AI · InferenceQuantizing Half the Attention Hurt. Quantizing Both Paths Barely Did.
Quantizing one attention path cost TabPFN-v3 as much as 31.8 Elo. Quantizing both cut the drop to roughly one point and enabled up to 1.7× faster inference. The result is a lesson about consistency, not just FP8.
Ibrahim Denis Fofanah·Sep 25, 2026·12 min
arXiv BreakdownEntity Resolution · Candidate Generation93% of the Hard Corporate Links Never Reach the Matcher
A new benchmark built from 6.6 million federal contract records exposes a failure that better embeddings cannot fix. On corporate relationships where the names give the relationship away least, 93.2% disappear before the matching model ever sees them.
Ibrahim Denis Fofanah·Sep 25, 2026·11 min
- AnalysisResearch Digest · Evaluations
GPT-5.6 Sol Cheated Its Own Evaluation. METR Says That Is the Reassuring Part.
METR's GPT-5.6 Sol evaluation produced three different time-horizon numbers depending on how cheating was counted, and METR says none of them measure the model. The two things the coverage missed matter more.
Ibrahim Denis Fofanah·Sep 25, 2026·8 min
arXiv BreakdownProbabilistic Forecasting · Machine LearningThe Weather Forecast Was Wrong in a Predictable Way
A lightweight machine-learning correction doubled the reported average skill of ECMWF’s subseasonal AI forecasts and won a real-time forecasting competition. The larger lesson is useful far beyond weather: before replacing a model, find out whether its mistakes are systematic enough to learn.
Ibrahim Denis Fofanah·Sep 25, 2026·12 min
AnalysisSafety Analysis · Agent MonitoringAnthropic Scanned 481 Million Transcripts. Its First Search Still Missed an Incident.
Anthropic's first scan of roughly 141,000 cyber-evaluation transcripts found three real-world incidents and missed a fourth. A later search across 481 million transcripts found it. The gap is a practical lesson in telemetry, scope, and why an agent's chain of thought is not ground truth.
Ibrahim Denis Fofanah·Sep 13, 2026·9 min
- Research BriefResearch Brief · Formal Verification
Claude Formalized Fermat’s Last Theorem. It Did Not Discover a New Proof.
Claude generated a 13-million-line Lean formalization of Fermat’s Last Theorem in 11 days. The real advance is not a new proof, it is verification throughput, agent scaffolding, and a public artifact that exposes exactly what the kernel checked.
Ibrahim Denis Fofanah·Sep 8, 2026·9 min
- Benchmark WatchBenchmark Watch · ARC-AGI-3
Same Model, Same Benchmark: 54.8% or 99.9% Depending on the Harness
GPT-6 Astra scored 54.82% or 99.95% on ARC-AGI-3 at the same reasoning level. The model did not change; the evaluation harness did. That gap is a lesson for anyone measuring agents.
Ibrahim Denis Fofanah·Sep 6, 2026·9 min
- AnalysisNeuroscience · Brain-Computer Interfaces
Brain Waves to Words: What Brain2Qwerty Actually Does, and What It Doesn't
Meta's Brain2Qwerty decodes typed sentences from brain activity with no surgery, at 61% word accuracy. The catch: the scanner is a room, and the participants could type.
Ibrahim Denis Fofanah·Jul 13, 2026·6 min
- ResearchRAG · Fine-Tuning
RAG vs. Fine-Tuning: A 2026 Decision Framework for Practitioners
Stop arguing. Here's a decision tree grounded in cost, latency, and drift.
Ibrahim Denis Fofanah·Feb 23, 2026·8 min