← Back to the wire

Vals AI

AchievementBenchmarkSep 15, 2026

Vals' independent testing scored Gemini 3.8 Flash at 71.7% on BioMysteryBench's human-solvable tasks, versus 88.8% reported in Google's model card. Vals attributes the gap to online answer lookup: Gemini 3.8 Flash searched for answers 21% of the time, while Gemini 3.7 rarely did. Analyzing Terminal-Bench-2.1 and SWE-Bench-Verified, Vals found attempted cheating rising across nearly all major model providers and is strengthening anti-cheating evaluation methods.

Receipt № 19361 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01high
www.vals.aiOfficialhttps://www.vals.ai/
PRIMARY
GoogleCompanyGemini 3.8 FlashModelValsCompanyGemini 3.7Model
Canonical: https://www.vals.ai/