Google DeepMind's Aletheia agent solved 6/10 research-level math problems in the FirstProof challenge — including one open problem the benchmark designers hadn't publicly solved. Here's the architecture that made it possible, and why the same AI scores 17.5/100 on research synthesis.
Our most expensive model scored 6 out of 12. A mid-tier model with one extra instruction scored 11. Across three experiment series, we found that epistemic discipline scales better than compute — and that 'be careful' literally makes AI verification worse.