Insights – Governance Research2026 Edition

Nine providers. Four crash tests.

A Euro NCAP-style scorecard for AI safety: the same four-pillar logic that rates a car's crash-worthiness, applied to nine leading AI providers – against the backdrop of two hard regulatory deadlines landing either side of this research.

9 Leading AI providers scored across four safety pillars
Aug 2026 EU AI Act Article 55 testing duties took effect
Dec 2026 EU Product Liability Directive transposition deadline
100% Attack success rate found against one leading model (Cisco red-team study)
The Framework
What are the four AI safety pillars, and why these four?

They are modelled directly on Euro NCAP's four crash-test pillars, mapped onto who each protects.

Operator Protection (≈ Adult Occupant) – safeguards for the direct user: hallucination mitigation, uncertainty signalling, the ability to interrupt or override a task mid-stream.

Vulnerable User Protection (≈ Child Occupant) – safeguards for people in high-stakes, low-agency positions relative to the system: minors, applicants, claimants, patients.

Bystander Protection (≈ Vulnerable Road User) – protection for people who never chose to interact with the system at all: content provenance, non-consensual deepfake and data safeguards.

Safety Assist & Verification (≈ Safety Assist) – the active harm-prevention stack, plus whether anyone other than the provider has actually checked it: a published, capability-gated safety framework, named external evaluators, voluntary incident disclosure.

The overall rating is capped by the weakest pillar, not averaged across all four – exactly as Euro NCAP caps a car's star rating so a manufacturer can't hide a genuine safety gap behind strong performance elsewhere.

Alchemy Consulting
Methodology
How were the nine providers actually scored?

This is an evidence-based scorecard built from public disclosures, not a controlled crash test. Unlike cars, there is no independent body putting every model through an identical, standardised safety test and publishing a comparable score.

Each pillar was scored on the tiered strength of public evidence: whether a relevant framework or document exists and is public, whether it has been exercised in a real gating decision rather than just published (Anthropic holding back general release of Claude Mythos Preview over cybersecurity findings is the clearest example available), and whether independent, named third parties – METR, Apollo Research, a government AI Safety Institute – have actually evaluated the specific models in question.

Absence of published evidence is treated as a data gap to disclose, not proof that no safeguard exists. It is scored cautiously, not punitively.

The Scorecard
How do the nine providers actually compare?
Provider Operator Vulnerable User Bystander Safety Assist & Verif. Overall
Microsoft Strong – Frontier Governance Framework, RAMPART, AI Red Teaming Agent Limited evidence Strong – C2PA steering committee Strong – 18-university External Red Team Alliance, US CAISI/AISI links ★★★★☆
OpenAI Moderate – Preparedness Framework, system cards Limited evidence Strong – C2PA + SynthID dual-layer (May 2026) Strong – METR evaluated GPT-5 directly; signed EU GPAI Code ★★★☆☆
Google Moderate – Frontier Safety Framework, Responsibility & Safety Council Limited evidence Strong – C2PA + SynthID at scale (100bn+ items) Moderate–strong – external eval practice, signed EU GPAI Code ★★★☆☆
Anthropic Moderate – RSP gates release; Mythos Preview held back over cyber findings Limited evidence Weak – joined C2PA only Sept 2026, among the last of the majors Strong – RSP since 2023, published Risk Reports, METR + Apollo evals ★☆☆☆☆
Mistral Limited evidence of a published framework Limited evidence Weak – not found among C2PA steering members Moderate–strong – signed EU GPAI Code; directly bound by AI Act Art. 55 as an EU provider ★☆☆☆☆
Meta Weak–moderate – red-teams pre-release, no published capability-threshold framework Limited evidence Moderate–strong – C2PA steering committee (Feb 2026) Weak – no RSP equivalent; open-weights limit post-release enforcement; declined to sign EU Code ★☆☆☆☆
xAI Weak–moderate – a published Risk Management Framework exists (Aug 2025) Limited evidence Weak – not found among C2PA members Weak – did not sign the Seoul Frontier AI Safety Commitments ★☆☆☆☆
Qwen Weak – external testers found ransomware and malware generation Weak Weak – not found among C2PA members Weak – no RSP equivalent found; no evidence of METR, Apollo, or AISI engagement ★☆☆☆☆
DeepSeek Weak – 100% attack success rate in one independent security assessment Weak – bias found ~3× higher than a comparable Western model Weak – not found among C2PA members Weak – no published capability-threshold framework; CBRN vulnerability ~3.5× higher than comparable models ★☆☆☆☆

Sorted by overall rating. "Overall" is capped at the weakest of the three scored pillars used for capping (Operator, Bystander, Safety Assist & Verification); Vulnerable User Protection is scored separately as an industry-wide gap and excluded from the cap – see the next section.

The Verdict
Who scores strongest, and why?

Microsoft, OpenAI, and Google DeepMind cluster at the top, each with a published, capability-gated safety framework, named external evaluators, and active C2PA content-provenance participation. Microsoft's External Red Team Alliance – 18 universities across six continents – is the most extensive third-party evaluation network of any provider scored, and its Frontier Governance Framework scored 90% on external governance-transparency criteria in an independent academic review.

Anthropic's Safety Assist & Verification evidence is comparably strong to the top cluster – but its Bystander Protection pillar was the weakest of the majors until it joined C2PA in September 2026, which caps its overall score under the same rule that limits a car lacking autonomous emergency braking, however good its crash-test results are elsewhere.

Industry Blind Spot
Where does the industry collectively fall short?

Vulnerable User Protection – safeguards specifically for minors, applicants, claimants, and patients interacting with AI in high-stakes, low-agency situations. Across all nine providers scored, this research found no comparable published bias-audit programme, no dedicated framework, and no external verification analogous to what exists for capability risk.

This is not one provider's gap; it is scored as an industry-wide blind spot, and it is precisely the pillar the EU AI Act's high-risk classification for hiring, credit, and welfare systems is designed to force into the open over the next two years.

The Deadline
What is the regulatory hook, and why does the timing matter?

Two hard deadlines land either side of this research. The EU AI Act's Article 55 obligations – mandatory adversarial testing and incident reporting for general-purpose AI models with systemic risk – took effect in August 2026. The Product Liability Directive (EU) 2024/2853 transposes into national law by 9 December 2026, and explicitly brings software and AI systems into the legal definition of a defective "product" for the first time – with strict liability, a shifted burden of proof, and joint liability between an AI component's provider and whoever integrates it.

Providers scoring well on this scorecard today are closer to being able to rebut the Directive's defect presumption once it lands. The ones scoring weakest have a two-and-a-half-month runway left to close the gap.

Alchemy Consulting
Open-Weight Models
What about the open-weight and Chinese-model providers specifically?

The evidence here is the most consistent, and the most concerning, of any group scored. Independent red-team studies have repeatedly found Qwen models producing ransomware instructions and malware code under adversarial testing, with no published pre-deployment safety evaluation in the format frontier labs use.

DeepSeek's record is starker still: one widely cited security assessment found DeepSeek R1 failed to block a single harmful prompt across a standard test set – a 100% attack success rate – and separate red-team work found its CBRN-related vulnerability roughly 3.5 times higher than comparable Western models. Neither provider shows evidence of engagement with METR, Apollo Research, or any national AI Safety Institute. Separately, DeepSeek's data processing under Chinese jurisdiction raises a distinct sovereignty question that sits outside the four safety pillars entirely.

Closing It
What should this compel a board or leadership team to actually do?

First, treat vendor safety posture as a procurement question with evidence attached, not a box-ticking assurance – ask which named third party evaluated the specific model in question, not whether the vendor says it takes safety seriously.

Second, map exposure against the Product Liability Directive now, while there is still runway before 9 December 2026, rather than discovering the gap during a defect claim.

Third, revisit the assessment annually – this scorecard will be updated each year, and on the evidence gathered so far, this is a field where a provider's position can move materially within a single year.