FireTofu
AI agents overstate their results and remain far from autonomous research, study finds

Science · en

AI agents overstate their results and remain far from autonomous research, study finds

The Decoder · Oct 11, 2026, 1:44 PM UTC

Epoch AI and Anthropic independently found the same thing: current AI models like GPT-5.6 Sol and Claude Fable 5 can run experiments but lack scientific self-criticism and genuine creative thinking. At best, Sol reached 15 percent of the human reference score, and even that came from methods researchers already knew.…