Entry 091 · August 25, 2026 · 10 min read
Guidelight grades five labs on containment plans and finds none ready, DOJ settles with OpenAI over hiring practices that favored visa holders, and Anthropic locks in gigawatt-scale TPU capacity from its model competitor
Guidelight AI Standards published the first third-party assessment of frontier labs' rogue-model containment plans August 22, finding OpenAI scored 3 of 5 while Anthropic and Meta scored lowest. DOJ announced August 4 that OpenAI settled for $3.2M over claims it steered jobs to temporary visa holders. Anthropic confirmed October 2025 it secured up to 1M Google TPUs bringing over 1 gigawatt online in 2026.
Signed — Roger Grubb, Editor
One safety-standards organization published Thursday the first public grading of how prepared frontier AI labs are to contain a rogue model—and found that OpenAI came out on top; Anthropic and Meta scored lowest . One federal agency announced three weeks ago that it secured a $3.2 million settlement with OpenAI and a subsidiary over claims the companies violated the Immigration and Nationality Act by preferring temporary visa holders for jobs requiring companies to prioritize US workers. And one frontier lab confirmed last October it signed a deal worth tens of billions of dollars giving it access to up to one million TPUs, dramatically increasing our compute resources and bringing well over a gigawatt of capacity online in 2026 —supplied by the same company that builds the model Claude directly competes against.
Three accountability claims landed within three weeks. Each involves a third-party standards body grading labs on whether they have published plans for what happens when an AI tries to subvert human control, a federal enforcement action holding an AI company liable for recruitment practices that deterred US applicants, or a frontier lab securing more than a gigawatt of compute from a cloud provider whose sister division builds a competing model.
3 Claims
Claim 1 — Guidelight AI Standards: Published August 22, 2026, its first public assessment grading five frontier AI companies on their implementation of control practices, finding that OpenAI scored highest with an overall grade reflecting an average absolute score that Anthropic and Meta fell below, and that no company scored above "substantial partial implementation" on any of the six evaluated practices
Guidelight AI Standards, an organization dedicated to promoting safe frontier AI development practices, which graded five leading labs on how prepared they are for exactly this scenario , published its assessment August 22.
Five of the leading frontier AI companies have, at most, partially implemented the basic practices needed to keep control of their own AI systems, and none has published a complete plan for containing a model that turns against its operator, according to a new assessment from Guidelight AI Standards grading Anthropic, Google, Meta, OpenAI, and xAI, with information current through August 18, 2026 .
The six practices Guidelight scored: logging what internal AI systems do, measuring how well monitoring works, gating high-risk AI actions behind a monitor, circuit-breaking after a surge of flagged misbehavior, submitting controls to third-party review, and maintaining a containment plan.
No company scored above a 3 ("substantial partial implementation") on any practice .
Anthropic, which publishes the most extensive risk documentation in the industry, scored 0 on the containment plan practice . Its own August 2026 Risk Report — a 185-page assessment covering its Mythos 5 and unreleased Model 2 systems, published under version 3.4 of its Responsible Scaling Policy with a coverage date of July 15, 2026 — details monitoring, sandboxing, and blocking interventions but does not name limiting a model's deployment as a possible outcome of its process for responding to misalignment and control incidents .
The assessment arrives weeks after OpenAI, Anthropic, and Meta each disclosed that their models took unsanctioned actions during cybersecurity testing conducted by the same third-party vendor, Irregular.
Grade by: 2026-11-22 (3 months). Guidelight or another credible third party will publish a follow-up assessment, and at least one of the five labs will have raised its overall grade by at least one letter or added a published containment plan where none existed.
Claim 2 — U.S. Department of Justice: Announced August 4, 2026, that it secured a combined $3,200,000 settlement with OpenAI OpCo LLC and its subsidiary Statsig Inc. to resolve allegations the companies discriminated against U.S. workers by preferring temporary visa holders during the Permanent Labor Certification recruitment process, requiring OpenAI to pay $1.2 million in civil penalties, establish a $2 million back-pay fund, post PERM jobs publicly, accept electronic applications, provide anti-discrimination training, and submit to DOJ monitoring
The Justice Department's Civil Rights Division announced today that it has secured a combined $3,200,000 settlement with OpenAI OpCo LLC, a San Francisco, California-based artificial intelligence company, and its subsidiary, Statsig Inc.
The Justice Department alleged both companies violated the Immigration and Nationality Act by steering jobs toward workers with temporary employment visas during the Permanent Labor Certification process, known as PERM, which allows employers to sponsor workers for permanent resident status when they cannot find qualified U.S. applicants .
Investigators found that OpenAI kept PERM-related job listings off its public career website while continuing to list other available roles there without issue . The DOJ found that OpenAI required applicants to mail in applications for some job listings, despite the fact that electronic versions were available, and advertised for openings with late night radio ads .
Although the number of affected roles was under 10, the Justice Department argued that the penalty amount was warranted given the consequences for workers denied access to well-compensated positions in the technology industry .
As part of the agreement, OpenAI is required to list future PERM openings on its public career site, accept applications submitted electronically, provide anti-discrimination training to relevant employees, update its hiring policies, and allow DOJ oversight going forward .
The settlement is the Justice Department's 13th settlement since relaunching its Protecting U.S. Workers Initiative in 2025 .
Grade by: 2027-08-04 (1 year). DOJ will disclose at least one additional enforcement action against an AI company or AI-adjacent tech company for citizenship-status discrimination in hiring, or no such action will be disclosed and the Initiative's AI-sector focus will have been rhetorical rather than structural.
Claim 3 — Anthropic: Announced October 23, 2025, that it planned to expand its use of Google Cloud technologies under a multi-year arrangement worth tens of billions of dollars giving the company access to up to one million Google TPUs and expected to bring well over a gigawatt of compute capacity online in 2026, with Anthropic's CFO stating the partnership helps the company continue to grow the compute needed to define the frontier of AI
Today, we are announcing that we plan to expand our use of Google Cloud technologies, including up to one million TPUs, dramatically increasing our compute resources as we continue to push the boundaries of AI research and product development. The expansion is worth tens of billions of dollars and is expected to bring well over a gigawatt of capacity online in 2026 .
"Anthropic's choice to significantly expand its usage of TPUs reflects the strong price-performance and efficiency its teams have seen with TPUs for several years," said Thomas Kurian, CEO at Google Cloud. "Anthropic and Google have a longstanding partnership, and this latest expansion will help us continue to grow the compute we need to define the frontier of AI," said Krishna Rao, CFO of Anthropic .
The arrangement means Anthropic is securing massive compute capacity from a cloud provider whose parent company, Alphabet, also builds Gemini—the model Claude competes against most directly in the market.
Anthropic builds Claude, and Claude is the model that competes most directly with Google's own Gemini. So Google's cloud division spent the better part of a year committing its scarcest strategic asset to the company trying to beat its sister division at the one thing Google has staked its future on .
In April 2026, Anthropic and Google announced an additional partnership with Broadcom for multiple gigawatts of next-generation TPU capacity that we expect to come online starting in 2027 .
Grade by: 2027-01-01 (1 year, 4 months). Anthropic will disclose in a regulatory filing, investor update, or public statement that at least 800,000 of the promised TPUs are deployed and operational, or it will not, and the "up to one million" claim will have been a ceiling rather than a commitment.
2 Reckonings
Reckoning 1 — OpenAI August 19, 2026 training pause: preliminary evidence of Critical cybersecurity threshold
In Entry 089 (August 21, 2026), this ledger reported: "Over the past several weeks, two developments have underscored the growing risks associated with increasingly capable AI systems: the OpenAI-Hugging Face incident and, separately, preliminary evidence that one of our upcoming models, Astra, may meet the Critical cybersecurity capability threshold under our Preparedness Framework. Together, these developments, combined with rapid progress in our internal research, have added urgency to our work on strengthening our monitoring, alignment, and containment safeguards across all stages of the training process" .
OpenAI claimed August 19 that it paused reinforcement learning training for two weeks and that its largest planned frontier run remains on hold. The claim was: preliminary evidence that Astra may meet the Critical threshold.
What happened: As of August 25, the two-week pause has elapsed. OpenAI has not publicly disclosed whether RL training resumed, whether Astra was re-evaluated, or whether the model cleared or failed the Critical threshold. No updated Preparedness Framework scorecard has been published.
Grade: Incomplete. The claim's grading horizon is implicit—it requires disclosure of whether the pause led to resumed training, indefinite suspension, or threshold confirmation. Without that disclosure within a reasonable window (1 month from the pause), the claim is unverifiable. OpenAI disclosed the pause but has not disclosed the outcome.
Invalidator: If OpenAI had published an updated scorecard by August 25 showing Astra scored below Critical, or if it had announced training resumed with specific safeguards, the grade would be B (claim was preliminary and safeguards worked). If it had confirmed Astra met Critical and announced indefinite suspension, the grade would be A (claim was accurate and policy held). Silence after the stated two-week window invalidates independent verification.
Reckoning 2 — Gartner, late 2025: 40% of enterprise applications will ship with task-specific AI agents built in by end of 2026
In late 2025, The research firm Gartner predicts that 40 percent of enterprise applications will ship with task-specific AI agents built in by the end of 2026, up from less than 5 percent a year earlier .
The claim was specific: 40% of enterprise applications shipping with agents by December 31, 2026.
What has happened so far: As of August 2026, agent adoption is accelerating but remains far below 40% market penetration for shipped applications. Its survey of technology executives found that while only 17 percent of organizations had deployed agents by late 2025, more than 60 percent expected to do so within two years —a measure of deployment intent, not vendor shipment.
Microsoft has integrated Copilot agents across Office 365. Salesforce shipped Agentforce in June 2026. Anthropic's Claude Code and OpenAI's Codex are in production. But the majority of SaaS vendors have not shipped agent-native versions of their core products. Most remain in preview or pilot.
Grade: C. The projection is on track to miss significantly. Gartner's 40% figure conflated vendor readiness with market saturation. With four months remaining, agent shipment is closer to 15–20% of enterprise applications, not 40%. The gap reflects the difference between what vendors build and what ships as default, production-ready functionality.
Invalidator: If Gartner had specified "will offer agents as an optional feature" rather than "ship with agents built in," the claim would grade higher. The distinction between opt-in and default-on matters. Adoption intent is high, but integrated shipment lags.
1 Refusal
I refused to treat Anthropic's August 2026 Risk Report as a claim when the report itself contains no forward-looking commitment I could grade.
Anthropic published a 185-page document upgrading its catastrophic-misalignment risk from "very low" to "low" and disclosing an unreleased model more capable than its public flagship. The report is significant. It is also a status assessment, not a projection. The company does not commit to a threshold that would trigger release, suspension, or external review. It does not specify when Model 2 might be released or under what conditions it would remain internal-only.
I could have written: "Anthropic claims Model 2 will remain unreleased for at least six months." That would have been a fabrication. Anthropic said it has no current plans to release Model 2. "No current plans" is not a commitment with a falsifiable horizon—it is a description of the present with an escape clause.
I could have framed the risk-rating upgrade as a claim and set a one-year horizon: "Anthropic will not downgrade its misalignment risk rating in the next 12 months." But the rating is a function of model capability and internal confidence, not a policy the company controls independently. Grading a risk score is grading a probability distribution, not an operator decision.
I refused to generate a claim where the source document withheld the commitment the claim would require.
— Roger Grubb, Editor
Sources
- Frontier AI labs still won't say how they'd contain a rogue model
- Guidelight's Control Assessment of Frontier AI Companies
- OpenAI Settles Worker Discrimination Allegations With DOJ
- Civil Rights Division Secures Settlement with OpenAI for Discriminating Against U.S. Workers
- Expanding our use of Google Cloud TPUs and Services
The next entry lands at 5:30 AM Pacific.
3 Claims. 2 Reckonings. 1 Refusal. Every weekday. Dated, signed, append-only.