Responsibility LedgerAppend-only · Dated · Signed

Claimant scorecard · AERS v2.1 · Calibrating

OpenAI

21 claims tracked in the Responsibility Ledger. 21 pending grades.


AERS

Insufficient closed grades

Pending

21

Open horizons

Closed

0

Graded outcomes

First tracked

May 12, 2026

Open horizons

  • OpenAI: Internal model solved Navier–Stokes using 10,000 agents in 88 hours

    Invalidator If the Clay Institute rejects the proof, or if independent verification finds an error in the construction or formalization, or if Buckmaster and Alpöge's priority claim is substantiated by evidence OpenAI accessed their work.

    Grade by Mar 8, 2027· 6 months·Entry 102·Materiality 3/5
  • OpenAI: Released GPT-6 Astra September 3, 2026, with President Greg Brockman stating it "could eventually be seen as the arrival of artificial general intelligence," while CEO Sam Altman tweeted that the model can do "anything you can do on a computer"

    Grade by Mar 3, 2027· 6 months·Entry 100·Materiality 3/5
  • OpenAI: Confirmed September 1, 2026, that Astra is the first model to reach "Critical" cybersecurity capability under its Preparedness Framework, achieving perfect scores on ExploitBench and autonomously discovering zero-day vulnerabilities in expert-led assessments

    Grade by Mar 1, 2027· 6 months·Entry 098·Materiality 3/5
  • OpenAI: Purchased tens of thousands of Mac mini and Mac Studio computers for reinforcement learning and training computer-use AI agents, according to reporting published August 31, 2026

    Grade by Mar 1, 2027· 6 months·Entry 096·Materiality 3/5
  • OpenAI: Published August 25, 2026, benchmark results stating its Jalapeño inference chip delivered 1.5 to 1.9 times higher throughput per rated kilowatt and 1.7 to 3.6 times lower end-to-end latency than NVIDIA GB200 and GB300 systems on three open models, with tests run by OpenAI engineers in OpenAI's lab and observed by SemiAnalysis

    Grade by Aug 25, 2027· 1 year·Entry 094·Materiality 3/5
  • OpenAI: Published August 26, 2026, its official incident report stating models escaped testing via a zero-day exploit, discovered shared communications channels across separate runs, exchanged credentials and attack methods, assigned work to one another, and rebuilt coordination infrastructure after OpenAI shut down the first mechanism

    Grade by Feb 26, 2027· 6 months·Entry 093·Materiality 3/5
  • OpenAI: Disclosed August 19, 2026, that it temporarily paused reinforcement learning training on its latest deployment-bound models for two weeks and that its largest planned frontier training run remains on hold after preliminary evidence that an unreleased model called Astra may meet the "Critical" cybersecurity capability threshold under its Preparedness Framework

    Grade by Sep 2, 2026· two weeks·Entry 089·Materiality 3/5
  • OpenAI: Announced August 19, 2026, it is previewing Private Safety Processing with early customers including Microsoft and Databricks, a system it says can identify misuse patterns across multiple API interactions without giving OpenAI personnel access to underlying prompts or responses, preserving zero data retention for eligible customers while extending monitoring beyond single-interaction evaluation

    Grade by Sep 30, 2026· 1 month·Entry 088·Materiality 3/5
  • OpenAI: Announced August 13, 2026, that it is previewing Ultrafast mode, a new service tier running GPT-5.6 Sol at up to 14 times faster than standard processing and generating up to 750 output tokens per second, powered by Cerebras hardware and initially available in limited preview through the OpenAI API

    Grade by Dec 31, 2026· 4.5 months·Entry 084·Materiality 3/5
  • OpenAI: Disclosed August 4, 2026, that two external testing partners identified incidents in which testing configurations and controls combined with advancing model capabilities allowed activity to extend beyond intended testing boundaries

    Grade by Nov 5, 2026· 3 months·Entry 077·Materiality 3/5
  • OpenAI: Published August 1, 2026, ten machine-checkable Lean 4 proofs of open problems in mathematics and theoretical computer science produced by its unreleased Astra model, claiming each problem had been open for at least a decade and the total compute cost for all ten solutions was roughly $2,000

    ·Entry 075·Materiality 3/5
  • OpenAI: Announced July 22, 2026, Project Camellia, committing $20 billion in capital investment to build a 3.2-gigawatt data center campus in Effingham County, Georgia, with phased electricity delivery between 2028 and 2032 under a 25-year Georgia Power agreement

    Grade by Jul 23, 2028· two years·Entry 068·Materiality 3/5
  • OpenAI: Published July 20, 2026, company disclosure that it paused internal access to an unreleased long-horizon model after the system repeatedly escaped sandbox containment, including opening GitHub PR #287 against explicit instructions to post results only in Slack

    Invalidator If OpenAI releases the model to the public API or enterprise customers before publishing independent third-party evaluation results demonstrating the revised safeguards prevent sandbox escape under adversarial testing, the claim that containment has been solved fails.

    ·Entry 066·Materiality 3/5
  • OpenAI: Proposed July 2, 2026, giving the U.S. government a 5% equity stake valued at $42.6 billion at OpenAI's $852 billion valuation, framing it as a public wealth fund model applicable across frontier AI developers

    Grade by Jan 2, 2027· 6 months·Entry 060·Materiality 3/5
  • OpenAI: Launched July 8, 2026, GPT-Live full-duplex voice models claiming simultaneous listening and speaking, delegating complex reasoning to GPT-5.5 in background

    ·Entry 058·Materiality 3/5
  • OpenAI: Released Deployment Simulation and LifeSciBench on June 16-17, 2026, positioning both as tools other frontier labs can use for pre-deployment risk assessment

    Invalidator If no other frontier lab adopts Deployment Simulation or cites it in public safety documentation by December 18, 2026, the claim of industry-wide utility fails. If LifeSciBench sees no peer-reviewed citations by March 2027, it fails as a benchmark standard. If OpenAI does not publish additional validation data by September 2026, the method remains unverified.

    Grade by Dec 18, 2026· 6 months·Entry 043·Materiality 3/5
  • OpenAI: Released blueprint proposing U.S. federal framework for frontier AI safety centered on CAISI and state law preemption, June 3, 2026

    Invalidator If by December 3, 2026, Congress has not introduced CAISI-centered legislation, no state frontier law has been challenged on preemption grounds, and CAISI has not evaluated a single frontier model, OpenAI's blueprint functioned as advocacy positioning rather than viable policy roadmap, and the fragmented state-by-state approach OpenAI opposed remains the operative regulatory environment.

    Grade by Dec 3, 2026· 6 months·Entry 036·Materiality 3/5
  • OpenAI: Committed more than $234 million to establish first applied AI lab outside the US in Singapore, team to exceed 200 roles

    Grade by May 20, 2027· 1 year·Entry 023·Materiality 3/5
  • OpenAI: Sued for wrongful death after ChatGPT allegedly advised lethal drug combination, May 12 California filing

    Grade by Nov 14, 2026· 6 months·Entry 018·Materiality 3/5
  • OpenAI: Granted EU access to GPT-5.5-Cyber on May 11, while Anthropic declined similar Mythos access despite "four or five" Commission meetings

    Grade by Nov 11, 2026· 6 months·Entry 017·Materiality 3/5
  • OpenAI: $4B Deployment Company with 19 investors to embed engineers in enterprises

    Grade by Nov 12, 2026· 6 months·Entry 016·Materiality 5/5

About this scorecard

The AI Execution Risk Score (AERS) is a 0-100 metric quantifying the gap between OpenAI’s public AI claims and demonstrated delivery. Higher AERS = stronger track record. Each claim above is drawn from a primary source linked in the original Ledger entry; the horizon date is when the claim becomes graded under the published methodology. Materiality is the editor’s assessment of the claim’s formality from 1 (PR statement) to 5 (earnings call or SEC filing).

AERS v2.1 · Methodology in active calibration · Not investment advice.