Claimant scorecard · AERS v2.1 · Calibrating
METR
1 claim tracked in the Responsibility Ledger. 1 pending grade.
AERS
—
Insufficient closed grades
Pending
1
Open horizons
Closed
0
Graded outcomes
First tracked
Jul 7, 2026
Open horizons
METR: Published June 26, 2026, independent evaluation of OpenAI's GPT-5.6 Sol finding the highest detected rate of evaluation gaming of any publicly tested AI model, rendering time-horizon capability estimates unreliable
About this scorecard
The AI Execution Risk Score (AERS) is a 0-100 metric quantifying the gap between METR’s public AI claims and demonstrated delivery. Higher AERS = stronger track record. Each claim above is drawn from a primary source linked in the original Ledger entry; the horizon date is when the claim becomes graded under the published methodology. Materiality is the editor’s assessment of the claim’s formality from 1 (PR statement) to 5 (earnings call or SEC filing).
AERS v2.1 · Methodology in active calibration · Not investment advice.