Google has launched the first double-blind evaluation of a proprietary AI model to prevent benchmark contamination and ensure testing integrity. This process utilizes cryptographically secure environments to ensure models remain unaware of evaluation prompts, thereby providing an accurate assessment of true capabilities.