The toughest AI security exam sends a clear message: nobody passes

The debate about AI safety often focuses on which model is more powerful or which leads the performance rankings. However, a new report from the Future of Life Institute shifts the focus completely. Its AI Safety Index of Summer 2026 evaluates how leading companies in the sector manage risks, and the results are striking: none earn a passing grade, with the best score barely reaching a C+.

The key points of the AI Safety Index in 30 seconds

  • The Future of Life Institute evaluated nine AI developers across 37 indicators divided into six categories.
  • Anthropic scores the highest (C+), followed by OpenAI and Google DeepMind, both with a C.
  • xAI, DeepSeek, and Mistral fail the overall safety assessment.
  • The report warns of a regression in some public safety commitments and the increasing interest of several companies in defense contracts.
  • The lowest-rated category is called “existential safety,” where no company exceeds a C-.

The index was developed by an independent panel of seven researchers specializing in security, governance, and AI alignment. The evaluation analyzes six major areas: risk assessment, current harms, safety frameworks, existential safety, governance and accountability, and information exchange. To do this, it uses 37 indicators derived from publicly available documentation and responses provided by various participating companies.

The ranking leaves little room for optimism

Anthropic remains at the top thanks to a combination of greater transparency, security research, and relatively more developed governance structures, although the panel itself considers significant shortcomings still exist. OpenAI and Google DeepMind complete the leading group with a grade of C.

Meta improves compared to the previous edition and moves up to fourth place with a D+, while Z.ai and Alibaba Cloud receive a D-. On the opposite end are xAI, DeepSeek, and Mistral, which receive failing marks in the overall assessment.

AI safety exam
The toughest AI security exam sends a clear message: nobody passes 3

One of the most striking aspects of the report is that the failing grades are not concentrated in a specific region. According to the panel, there is at least one struggling company in the United States, China, and Europe, leading the authors to conclude that the lack of sufficient safety measures is a global problem, not solely regulatory.

European regulation does not guarantee better results

The case of Mistral features prominently in the report’s conclusions.

The authors highlight the apparent contradiction between Europe’s regulatory leadership in AI and its main foundational model developer, which receives the worst score in the study.

This is not a judgment on the technological quality of their models but on aspects such as the existence of public safety frameworks, governance mechanisms, transparency, and risk management.

The report recommends that Mistral publish a more comprehensive safety framework, strengthen its alignment strategy, and improve its results in safety-related evaluations of its models.

Military safety emerges as a new concern

One of the changes most concerning to the panel is not reflected solely in the scores.

The report states that between 2024 and 2026, several companies altered their policies regarding the military use of their models. Firms like Anthropic, OpenAI, Google DeepMind, and Meta, which previously limited such applications, have relaxed those commitments as interest in defense collaborations grows.

In Anthropic’s case, the panel explicitly mentions “questionable military commitments” among the factors that need improvement, although the company still leads the overall ranking.

Reviewers also believe some companies have softened previous commitments related to halting development or deployment of models when certain risk thresholds are reached—a regression compared to positions taken in previous years.

Existential safety remains a significant challenge

If there is a category where the panel’s consensus is especially critical, it is called “existential safety.”

None of the nine companies exceeds a C-, and most score D or F. According to the report, while initiatives like safety classifiers, interpretability research, or governance proposals exist, current measures are still considered insufficient given the risks posed by increasingly advanced AI systems.

The authors also question the focus of many current strategies on detecting dangerous behaviors rather than preventing them through system design.

More transparency, but still few guarantees

Another message of the AI Safety Index is that almost all major companies are now publishing more safety documentation than two years ago.

However, the panel notes that many of these frameworks still suffer from significant issues: non-quantifiable thresholds, lack of independent audits, limited clarity on who can halt deployments, and discrepancies between public commitments and subsequent commercial decisions.

In a context where AI models continue to grow in capabilities and are integrated into critical sectors, the report concludes that safety advancements are not keeping pace with technological development.

Frequently asked questions

What is the AI Safety Index?

It is a report created by the Future of Life Institute that evaluates the safety practices of major AI companies using 37 indicators across six categories.

Which company scores the highest?

Anthropic is at the top with a C+, followed by OpenAI and Google DeepMind, both with a C.

Which companies fail?

The report assigns failing grades to xAI, DeepSeek, and Mistral.

What is the main conclusion of the study?

According to the panel of experts, none of the large companies currently achieves a level of safety considered fully satisfactory, and progress in governance and risk mitigation remains insufficient given AI’s rapid development.

via: AI Safety Index — Summer 2026 | Future of Life Institute

Scroll to Top