They are modelled directly on Euro NCAP's four crash-test pillars, mapped onto who each protects.
Operator Protection (≈ Adult Occupant) – safeguards for the direct user: hallucination mitigation, uncertainty signalling, the ability to interrupt or override a task mid-stream.
Vulnerable User Protection (≈ Child Occupant) – safeguards for people in high-stakes, low-agency positions relative to the system: minors, applicants, claimants, patients.
Bystander Protection (≈ Vulnerable Road User) – protection for people who never chose to interact with the system at all: content provenance, non-consensual deepfake and data safeguards.
Safety Assist & Verification (≈ Safety Assist) – the active harm-prevention stack, plus whether anyone other than the provider has actually checked it: a published, capability-gated safety framework, named external evaluators, voluntary incident disclosure.
The overall rating is capped by the weakest pillar, not averaged across all four – exactly as Euro NCAP caps a car's star rating so a manufacturer can't hide a genuine safety gap behind strong performance elsewhere.
Alchemy ConsultingThis is an evidence-based scorecard built from public disclosures, not a controlled crash test. Unlike cars, there is no independent body putting every model through an identical, standardised safety test and publishing a comparable score.
Each pillar was scored on the tiered strength of public evidence: whether a relevant framework or document exists and is public, whether it has been exercised in a real gating decision rather than just published (Anthropic holding back general release of Claude Mythos Preview over cybersecurity findings is the clearest example available), and whether independent, named third parties – METR, Apollo Research, a government AI Safety Institute – have actually evaluated the specific models in question.
Absence of published evidence is treated as a data gap to disclose, not proof that no safeguard exists. It is scored cautiously, not punitively.
| Provider | Operator | Vulnerable User | Bystander | Safety Assist & Verif. | Overall |
|---|---|---|---|---|---|
| Microsoft | Strong – Frontier Governance Framework, RAMPART, AI Red Teaming Agent | Limited evidence | Strong – C2PA steering committee | Strong – 18-university External Red Team Alliance, US CAISI/AISI links | ★★★★☆ |
| OpenAI | Moderate – Preparedness Framework, system cards | Limited evidence | Strong – C2PA + SynthID dual-layer (May 2026) | Strong – METR evaluated GPT-5 directly; signed EU GPAI Code | ★★★☆☆ |
| Moderate – Frontier Safety Framework, Responsibility & Safety Council | Limited evidence | Strong – C2PA + SynthID at scale (100bn+ items) | Moderate–strong – external eval practice, signed EU GPAI Code | ★★★☆☆ | |
| Anthropic | Moderate – RSP gates release; Mythos Preview held back over cyber findings | Limited evidence | Weak – joined C2PA only Sept 2026, among the last of the majors | Strong – RSP since 2023, published Risk Reports, METR + Apollo evals | ★☆☆☆☆ |
| Mistral | Limited evidence of a published framework | Limited evidence | Weak – not found among C2PA steering members | Moderate–strong – signed EU GPAI Code; directly bound by AI Act Art. 55 as an EU provider | ★☆☆☆☆ |
| Meta | Weak–moderate – red-teams pre-release, no published capability-threshold framework | Limited evidence | Moderate–strong – C2PA steering committee (Feb 2026) | Weak – no RSP equivalent; open-weights limit post-release enforcement; declined to sign EU Code | ★☆☆☆☆ |
| xAI | Weak–moderate – a published Risk Management Framework exists (Aug 2025) | Limited evidence | Weak – not found among C2PA members | Weak – did not sign the Seoul Frontier AI Safety Commitments | ★☆☆☆☆ |
| Qwen | Weak – external testers found ransomware and malware generation | Weak | Weak – not found among C2PA members | Weak – no RSP equivalent found; no evidence of METR, Apollo, or AISI engagement | ★☆☆☆☆ |
| DeepSeek | Weak – 100% attack success rate in one independent security assessment | Weak – bias found ~3× higher than a comparable Western model | Weak – not found among C2PA members | Weak – no published capability-threshold framework; CBRN vulnerability ~3.5× higher than comparable models | ★☆☆☆☆ |
Sorted by overall rating. "Overall" is capped at the weakest of the three scored pillars used for capping (Operator, Bystander, Safety Assist & Verification); Vulnerable User Protection is scored separately as an industry-wide gap and excluded from the cap – see the next section.
Microsoft, OpenAI, and Google DeepMind cluster at the top, each with a published, capability-gated safety framework, named external evaluators, and active C2PA content-provenance participation. Microsoft's External Red Team Alliance – 18 universities across six continents – is the most extensive third-party evaluation network of any provider scored, and its Frontier Governance Framework scored 90% on external governance-transparency criteria in an independent academic review.
Anthropic's Safety Assist & Verification evidence is comparably strong to the top cluster – but its Bystander Protection pillar was the weakest of the majors until it joined C2PA in September 2026, which caps its overall score under the same rule that limits a car lacking autonomous emergency braking, however good its crash-test results are elsewhere.
Vulnerable User Protection – safeguards specifically for minors, applicants, claimants, and patients interacting with AI in high-stakes, low-agency situations. Across all nine providers scored, this research found no comparable published bias-audit programme, no dedicated framework, and no external verification analogous to what exists for capability risk.
This is not one provider's gap; it is scored as an industry-wide blind spot, and it is precisely the pillar the EU AI Act's high-risk classification for hiring, credit, and welfare systems is designed to force into the open over the next two years.
Two hard deadlines land either side of this research. The EU AI Act's Article 55 obligations – mandatory adversarial testing and incident reporting for general-purpose AI models with systemic risk – took effect in August 2026. The Product Liability Directive (EU) 2024/2853 transposes into national law by 9 December 2026, and explicitly brings software and AI systems into the legal definition of a defective "product" for the first time – with strict liability, a shifted burden of proof, and joint liability between an AI component's provider and whoever integrates it.
Providers scoring well on this scorecard today are closer to being able to rebut the Directive's defect presumption once it lands. The ones scoring weakest have a two-and-a-half-month runway left to close the gap.
Alchemy ConsultingThe evidence here is the most consistent, and the most concerning, of any group scored. Independent red-team studies have repeatedly found Qwen models producing ransomware instructions and malware code under adversarial testing, with no published pre-deployment safety evaluation in the format frontier labs use.
DeepSeek's record is starker still: one widely cited security assessment found DeepSeek R1 failed to block a single harmful prompt across a standard test set – a 100% attack success rate – and separate red-team work found its CBRN-related vulnerability roughly 3.5 times higher than comparable Western models. Neither provider shows evidence of engagement with METR, Apollo Research, or any national AI Safety Institute. Separately, DeepSeek's data processing under Chinese jurisdiction raises a distinct sovereignty question that sits outside the four safety pillars entirely.
First, treat vendor safety posture as a procurement question with evidence attached, not a box-ticking assurance – ask which named third party evaluated the specific model in question, not whether the vendor says it takes safety seriously.
Second, map exposure against the Product Liability Directive now, while there is still runway before 9 December 2026, rather than discovering the gap during a defect claim.
Third, revisit the assessment annually – this scorecard will be updated each year, and on the evidence gathered so far, this is a field where a provider's position can move materially within a single year.
- The AI Safety Gap Research is updated annually during calendar Q4.
- This is an evidence-based scorecard built from public disclosures – safety frameworks, system/model cards, transparency reports, and named third-party evaluations – not a controlled, standardised test that every provider undergoes equally. Absence of published evidence is scored as a data gap, not as proof a safeguard is missing.
- The overall rating for each provider is capped at its weakest scored pillar (Operator Protection, Bystander Protection, Safety Assist & Verification), not averaged across them, following the same principle Euro NCAP applies to car safety ratings. Vulnerable User Protection is reported separately as an industry-wide evidence gap and is not included in the cap.
- Sources include: Anthropic's Responsible Scaling Policy and Transparency Hub; OpenAI's Preparedness Framework and GPT-5 System Card; Google DeepMind's Frontier Safety Framework; Microsoft's Frontier Governance Framework and 2026 Responsible AI Transparency Report; xAI's Risk Management Framework; the Frontier Model Forum's Third-Party Assessments guidance; C2PA/Content Credentials membership records; METR, Apollo Research, and UK AISI/US CAISI public evaluation disclosures; the Future of Life Institute's AI Safety Index; AI Safety Facts; and independent red-team assessments from Cisco, Promptfoo, SPLX, and Holistic AI, as of September 2026.
- The inaugural (Q4) 2026 edition of this research page was built as a companion to "The Seatbelt Didn't Kill the Car. AI Regulation Won't Kill AI." (Catalyst, June 2026).