Future of Life Gives Frontier AI Labs Weak Safety Grades
Future of Life Institute's Summer 2026 AI Safety Index graded nine frontier AI companies and found weak safety and governance performance across the sector. Anthropic led with a C+, OpenAI and Google DeepMind received Cs, Meta received a D+, and xAI, DeepSeek and Mistral failed overall, according to FLI's report. The useful practitioner takeaway is vendor diligence: the index pushes teams to ask for concrete risk thresholds, external testing access, incident reporting, whistleblower protections, military-use policies and deployment guardrails before adopting frontier models. Axios separately reported that reviewers flagged weakened pause commitments and rising military-use concerns, so the story is a governance benchmark rather than a capability leaderboard.
The LDS value in this report is vendor diligence: frontier-model buyers need more than model cards, benchmark claims and broad safety language. FLI's index gives procurement, governance and platform teams a current checklist for asking how labs define risk thresholds, who gets external testing access, and what happens when deployment risks rise.
What happened
The Future of Life Institute published its Summer 2026 AI Safety Index for nine frontier AI companies: Anthropic, OpenAI, Google DeepMind, Meta, Z.ai, Alibaba Cloud, xAI, DeepSeek and Mistral. The index grades 37 indicators across domains including risk assessment, current harms, safety frameworks, existential safety, governance and information sharing. Anthropic ranked first with a C+ overall grade, OpenAI and Google DeepMind received C grades, Meta improved to D+, and xAI, DeepSeek and Mistral received failing overall grades.
Policy context
FLI is an advocacy organization focused on catastrophic AI risk, so the index should be read as an expert-panel governance assessment with a clear risk lens rather than as a neutral regulatory finding. That limitation matters, but the report is still useful because it turns vague safety promises into observable questions about thresholds, disclosure, red-team depth and accountability.
For practitioners
Teams evaluating frontier vendors can use the report as a prompt for procurement questions: whether deployment gates are documented, whether external evaluators can test models, whether incident reporting is public enough for enterprise risk teams, and whether military or high-risk uses are governed by enforceable policies.
What to watch
The next useful signal is whether labs respond with updated safety frameworks, third-party evaluation access or clearer pause commitments. If the grades remain low while capabilities rise, AI governance teams should treat vendor controls as a first-class integration requirement rather than a policy appendix.
Key Points
- 1FLI graded nine frontier AI companies across 37 safety and governance indicators in its Summer 2026 index.
- 2Anthropic led with a C+, while OpenAI and Google DeepMind received Cs and three labs failed overall.
- 3The report gives model buyers a concrete checklist for risk thresholds, testing access and incident-reporting diligence.
Scoring Rationale
This is a notable governance signal for teams deploying frontier models because it compares major labs on safety frameworks, disclosure and risk controls. It is not a binding regulatory action, but it is current, primary-source-backed and broadly relevant to model procurement, compliance and AI risk management.
Sources
Public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems


