FireTofu
AI benchmarks have a trust problem and Google wants to fix it

Technology · en

AI benchmarks have a trust problem and Google wants to fix it

The Decoder · Aug 28, 2026, 1:15 PM UTC

Google Deepmind is testing a double-blind evaluation of a frontier AI model for the first time. Cryptographic protection through Confidential Space is meant to keep Google from seeing the test questions and keep evaluators from seeing the model weights.…