← Back
July 2026 · Changelog

Dataset cleanup v2

A full read of every row across the three scored benchmarks turned up duplicate questions, one broken multiple-choice item, and a taxonomy that didn't match between tiers. This is the changelog for the fixes. Nothing about the tier design changed — the questions that remain are the same questions, minus the redundant and broken ones.

bench-effortless-6-2026

bench-easy-6-2026  269 → 238

bench-mid-6-2026  155 → 143

Leaderboard


Re-run any model against the cleaned data with script.py from the leaderboard Space.