Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint
Data sources: ZENODO
addClaim

The Measurement Gap

Authors: Srihari R Mysore;

The Measurement Gap

Abstract

This essay argues that the AI capability question posed by "Situational Awareness" (Aschenbrenner, 2024) has, on its own terms, mostly been answered, and has turned into a measurement question that no current benchmark can settle. In the first week of September 2026, OpenAI's GPT-6 Astra scored 62.7% and 99.9% on the ARC-AGI-3 benchmark with the same underlying weights, depending only on which evaluation harness was used. This essay treats that result as the central data point. Part I re-grades the 2024 "Situational Awareness" predictions against September 2026 evidence. Part II decomposes the Astra result, defines a "harness ratio" (H) for separating model capability from scaffold-provided capability, and fits a half-life to the ARC-AGI benchmark series (roughly 61 months, 12 months, then 6 months across three generations). Part III proposes DRIFT (Dynamic Rule Inference and Fast Transfer), a benchmark that measures adaptation to unannounced mid-episode rule changes as a rate against a human baseline, rather than a static pass/fail score. A minimum-viable prototype design is included. Part IV states nine dated, falsifiable predictions through September 2028. Written with Claude Fable 5.1 as a research and drafting collaborator; the argument, editing, and any errors are the author's own. All figures were checked against primary sources as of September 7, 2026. Companion site with rendered mathematics: https://measurementgap.com

Powered by OpenAIRE graph
Found an issue? Give us feedback