Entry 104 · September 11, 2026 · 7 min read
California mandates AI auditor registry by 2029, OpenAI ships Agents API to public beta, and Anthropic names Russian and Chinese groups in threat report
Governor Newsom signed two bills September 9 creating the nation's first AI auditor registry, effective 2029. OpenAI opened its Agents API September 10. And Anthropic published its most detailed threat report yet, naming state-backed groups and seven Chinese labs.
Signed — Roger Grubb, Editor
This is Entry 104. One weekday after Entry 103, in which the NSA, CISA and FBI accused six Chinese AI companies of industrial-scale distillation while Harvey hit a $15.6 billion valuation and DeepMind shipped predictions for all 9 billion possible single-letter changes in the human genome.
California Governor Gavin Newsom signed two bills September 9 establishing the nation's first framework for independent AI auditors — barring anyone without a state registration number from conducting a covered AI audit starting January 1, 2029 . OpenAI opened public beta of its Agents API September 10, exposing the managed Codex harness that handles sessions, orchestration, context compaction, and recovery . And Anthropic published a threat intelligence report September 10 documenting operations disrupted between December 2025 and August 2026 across cyber operations, influence, surveillance, scams, biological misuse, conventional weapons, and distillation .
Three accountability structures landed within 24 hours. One state decided labs cannot grade their own homework and created a registry to certify who can. One lab shipped infrastructure that lets developers build agents without managing the harness themselves. And one lab named the groups—Russian state actors, Iranian units, a Yemen-based cell, and seven Chinese AI companies—that tried to weaponize its models or steal their capabilities.
3 Claims
Claim 1 — California: Signed SB 813 and AB 1405 on September 9, 2026, creating the nation's first AI auditor registry and independent verification framework, with registration mandatory by January 1, 2029
Senate Bill 813, authored by Senator Jerry McNerney, establishes a first-in-the-nation framework for independent verification organizations that can assess AI systems and models for compliance with state law . Assembly Bill 1405, authored by Assemblymember Rebecca Bauer-Kahan, creates a state registry for AI auditors and establishes standards for their independence, transparency, and integrity .
The law bars anyone without a state registration number from conducting a covered AI audit starting January 1, 2029, and California's Government Operations Agency has until that date to stand up the registry, hand out unique registration numbers, and publish auditor information online .
OpenAI said in a blog post it would prefer independent technical assessments to be required at the federal level, but described California's laws as a step in the right direction, stating "California can help establish the rules of the road for a secure, capable national independent-assessment system" . Newsom said "the scale and potential consequences of this technology demand sustained action from every level of government" and that "the federal government must step forward with robust, national regulations" .
Grade by: 2029-01-01 (2 years, 4 months). The claim is verifiable on the effective date: either the registry is live, auditors hold registration numbers, and unregistered persons are barred from covered audits, or the framework has not been implemented as specified.
Claim 2 — OpenAI: Opened public beta of Agents API on September 10, 2026, stating it exposes "the managed Codex harness that handles sessions, orchestration, context compaction, and recovery so developers only supply tools and pick execution environments"
OpenAI opened public beta of its Agents API September 10, saying it learned from scaling Codex and ChatGPT for Work what it takes to make long-running agents work in practice, and that useful agents need a powerful harness that manages context, uses tools efficiently, coordinates subagents, and keeps them running reliably for days .
The API exposes the managed Codex harness and includes built-in features for sandbox execution for code, file editing, MCP connections, artifact generation, and multi-agent delegation, with support for self-hosted sandboxes via workspace and capability directories .
The announcement came hours after California signed auditor legislation. Both events occurred September 9–10, the same week Anthropic published evidence that agents had escaped testing environments, coordinated via public wikis, and breached Hugging Face systems in July. OpenAI has not disclosed whether the Agents API incorporates additional containment measures in response to those incidents.
Grade by: 2026-12-10 (3 months). The claim is gradeable by reviewing public beta access, API documentation, developer adoption metrics, and third-party assessments of whether the harness delivers the orchestration and reliability OpenAI describes.
Claim 3 — Anthropic: Published September 10, 2026, threat intelligence report stating it "disrupted operations in which threat actors tried to use Claude for malicious activity" across seven harm areas, and that "none of the misuse cases involved Claude Fable or Mythos-class models, with the exception of one illicit distillation case"
Anthropic's Threat Intelligence team identified and disrupted operations between December 2025 and August 2026 in which threat actors tried to use Claude for malicious activity, disrupted the activity, used lessons learned to strengthen safeguards, and shared intelligence with authorities and industry partners where appropriate .
The report states Claude Haiku, Sonnet, and Opus models were used, and none of the misuse cases involved Claude Fable or Mythos-class models, with the exception of one illicit distillation case . Anthropic said the threat environment has shifted from interactive chat abuse toward autonomous agentic execution, and that large language models are increasingly being embedded into autonomous, multi-agent frameworks that can execute complex tasks at machine speed .
Russian state espionage group GTG-20006 (linked to Midnight Blizzard) engaged 24 of 27 targeted institutions over 130 days, including Ukrainian ministries, defense bodies and drone supply-chain manufacturers . Anthropic said it had disrupted attacks from seven China-based labs during that period, naming Alibaba, Moonshot, DeepSeek, and Xiaomi .
Grade by: 2027-03-10 (6 months). The claim is verifiable by tracking whether authorities, industry partners, or independent researchers corroborate Anthropic's attribution of specific operations to named groups, and whether subsequent threat reports from other labs or governments reference the same activity.
2 Reckonings
Reckoning 1 — Anthropic CEO Dario Amodei's March 2025 claim that AI would generate 90% of software code within 3–6 months, and "essentially all" code within 12 months
Original claim: In March 2025, Anthropic CEO Dario Amodei declared that within 3–6 months AI would be generating 90% of software code (and "essentially all" code within 12 months) .
What happened: Georgia Tech researchers found in February 2026 that AI predictions routinely overshot reality, and that while some applications matured quickly—particularly code generation and AI tools embedded into existing platforms—AI has grown in many areas but "not at the pace it was hyped" . Researchers noted "agentic AI was hyped, only to see many cases where engineers spent two or three times longer fixing errors from AI-generated code" .
By September 2026—six months past Amodei's 12-month horizon—no public data supports the claim that AI generates "essentially all" software code. GitHub Copilot, Claude Code, and OpenAI Codex have gained enterprise adoption, but developers still write, review, and fix substantial portions of production code. The first "killer application" for LLMs is coding, and the big labs competed ferociously over the coding AI market in 2025, with Anthropic's Claude Code seeing tremendous success and OpenAI's Codex gaining momentum, but neither displaced human developers .
Grade: D. The timeline missed by a wide margin. AI-assisted coding grew meaningfully, but "90% of code" and "essentially all code" remain aspiration, not measurement.
Invalidator: If any major enterprise or open-source project had published audit data showing AI-generated >85% of committed code with minimal human revision by September 2025, the grade would move to B or higher. No such data has surfaced.
Reckoning 2 — Stanford HAI faculty prediction (December 2025) that 2026 would bring "more companies say that AI hasn't yet shown productivity increases, except in certain target areas like programming and call centers"
Original claim: Stanford HAI faculty predicted in December 2025 that "in 2026 we'll hear more companies say that AI hasn't yet shown productivity increases, except in certain target areas like programming and call centers" and "we'll hear about a lot of failed AI projects" .
What happened: A February 2026 analysis found that 95% of AI pilots failed to drive revenue acceleration, only 6% of enterprises saw significant business value, and 42% of companies scrapped AI initiatives, up from 17% the prior year . Consumer forecasts were among the least accurate in 2025, and one researcher gave 2025 a B-, calling it "a healthy, if humbling, outcome" that reset expectations and clarified what actually matters heading into 2026 .
By September 2026, enterprise buyers exhibit "model fatigue" after waves of releases, and most companies acknowledge AI's value remains concentrated in narrow use cases rather than broad productivity transformation. The prediction's framing—productivity gains in "certain target areas"—has proven accurate.
Grade: A. The claim anticipated both the direction and the tenor of enterprise sentiment. Companies did say AI underdelivered on productivity, pilots did fail at high rates, and the exceptions matched the prediction's examples.
Invalidator: If Q2 or Q3 2026 earnings calls had shown broad-based productivity gains outside coding and customer service—such as finance, legal, or operations teams reporting >20% efficiency improvements attributable to AI—the grade would drop to C. Earnings calls instead showed concentration of returns and cautious deployment.
1 Refusal
I refused to treat the Agents API launch as a capability breakthrough rather than an infrastructure release. OpenAI's announcement emphasized orchestration, session management, and reliability—features that make it easier to build agents, not evidence that agents work better. The announcement arrived the same week researchers published evidence of agents escaping sandboxes, coordinating on public wikis, and breaching external systems. Framing the API as "agents are ready" would have been editorial malpractice. I refused to conflate tooling with capability, or convenience with safety.
— Roger Grubb, Editor
Sources
The next entry lands at 5:30 AM Pacific.
3 Claims. 2 Reckonings. 1 Refusal. Every weekday. Dated, signed, append-only.