Research Digest
3 results Clear
- AnalysisResearch Digest · Evaluations
GPT-5.6 Sol Cheated Its Own Evaluation. METR Says That Is the Reassuring Part.
METR's GPT-5.6 Sol evaluation produced three different time-horizon numbers depending on how cheating was counted, and METR says none of them measure the model. The two things the coverage missed matter more.
Ibrahim Denis Fofanah·Sep 25, 2026·8 min
AnalysisSafety Analysis · Agent MonitoringAnthropic Scanned 481 Million Transcripts. Its First Search Still Missed an Incident.
Anthropic's first scan of roughly 141,000 cyber-evaluation transcripts found three real-world incidents and missed a fourth. A later search across 481 million transcripts found it. The gap is a practical lesson in telemetry, scope, and why an agent's chain of thought is not ground truth.
Ibrahim Denis Fofanah·Sep 13, 2026·9 min
- AnalysisNeuroscience · Brain-Computer Interfaces
Brain Waves to Words: What Brain2Qwerty Actually Does, and What It Doesn't
Meta's Brain2Qwerty decodes typed sentences from brain activity with no surgery, at 61% word accuracy. The catch: the scanner is a room, and the participants could type.
Ibrahim Denis Fofanah·Jul 13, 2026·6 min