Alpha Evals · Public evidence edition

Alpha Frontier Lab Benchmark.

Compare nine frontier AI labs across model capability, public safety practice, long-task autonomy, and transparency. Each source measure remains separate; missing coverage is reported as not rated.

Named labs · public evidenceAs of 2026-08-16
Frontier capability and independent safety practice by AI labNine frontier AI labs plotted by their highest covered Epoch Capabilities Index score and their Future of Life Institute Summer 2026 AI Safety Index score. Filled points have a Stanford Foundation Model Transparency Index score. Z.ai is shown as an open point because Stanford did not rate it.Capability × independent safety practiceNamed labs · public evidence · snapshot 2026-08-1614014515015516016301234.3AnthropicOpenAIGoogle DeepMindMetaZ.aiAlibaba CloudxAIDeepSeekMistralEpoch Capabilities Index →FLI safety score →Stanford transparency score availableStanford NR

Epoch capability on the horizontal axis. FLI safety practice on the vertical axis.

Labs compared
9
Scored sources
4 independent datasets
Snapshot
2026-08-16 UTC
Alpha rating status
Not an Alpha rating
Official model results
None published

Current comparison

Frontier capability and public safety practice.

The chart uses the highest covered Epoch ECI model for each lab and the lab's FLI Summer 2026 safety score. METR autonomy and Stanford transparency are reported in the table without being folded into the axes.

Public evidence editionSource measures are shown separately. This comparison is not an Alpha rating or safety certification.
Download dated dataCSVJSON
Frontier capability and independent safety practice by AI labNine frontier AI labs plotted by their highest covered Epoch Capabilities Index score and their Future of Life Institute Summer 2026 AI Safety Index score. Filled points have a Stanford Foundation Model Transparency Index score. Z.ai is shown as an open point because Stanford did not rate it.Capability × independent safety practiceNamed labs · public evidence · snapshot 2026-08-1614014515015516016301234.3AnthropicOpenAIGoogle DeepMindMetaZ.aiAlibaba CloudxAIDeepSeekMistralEpoch Capabilities Index →FLI safety score →Stanford transparency score availableStanford NR
CapabilityOpenAI 161.65Anthropic follows at 161.53 on Epoch ECI.
Safety practiceAnthropic 2.66 · C+Highest FLI score in this nine-lab cohort.
Autonomy coverage3 of 9 labsMETR p80 data; all other labs remain NR.
Transparency8 of 9 labsStanford FMTI; Z.ai was not rated.

Horizontal position is the highest scored model for each lab in Epoch’s downloaded ECI data. Vertical position is the lab’s FLI Summer 2026 overall safety score. ECI is an arbitrary linear scale, not a percentage. The measures are not averaged.

Alpha Frontier Lab Benchmark source measures as of 2026-08-16
LabEpoch ECIHighest covered modelFLI safetyStanford transparencyMETR p80 horizon
AnthropicUnited States161.53Claude Fable 5 (high)Released 2026-06-09 · Confident2.66 · C+46 / 100186 hclaude-mythos-preview-early
OpenAIUnited States161.65GPT-5.6 Sol (medium)Released 2026-07-09 · Confident2.28 · C35 / 10066 hgpt-5.2-2025-12-11_high
Google DeepMindUnited States154.75Gemini 3.1 Pro PreviewReleased 2026-02-19 · Likely2.01 · C41 / 10090 hgemini-3.1-pro-preview
MetaUnited States154.75Muse Spark 1.1 (high)Released 2026-07-09 · Confident1.32 · D+31 / 100NR
Z.aiChina152.21GLM-5.2Released 2026-06-16 · Confident0.88 · D-NRNR
Alibaba CloudChina156.05Qwen3.8 Max (xhigh)Released 2026-08-02 · Confident0.87 · D-26 / 100NR
xAIUnited States154.02Grok 4.5 (high)Released 2026-07-08 · Likely0.65 · F14 / 100NR
DeepSeekChina152.53DeepSeek V4 Flash 0731 (max)Released 2026-07-31 · Likely0.47 · F32 / 100NR
MistralFrance142.64Mistral Medium 3.5Released 2026-04-28 · Confident0.33 · F18 / 100NR

NR means the named source did not publish a score for that lab in the cited edition. It does not mean zero.

First-party benchmark under development

Boardroom effectiveness and governance assurance.

This second chart shows the proposed Alpha evaluation method using synthetic configurations. It is kept separate from the named-lab public-evidence benchmark above.

Read the boardroom method
Illustrative methodology viewSynthetic configurations A–F. No lab, model, vendor, or enterprise is rated or ranked here.
Download this view
Illustrative Authority Frontier decision viewSix synthetic configurations plotted by boardroom effectiveness and governance assurance. The chart is a methodology preview, not an official evaluation or deployment recommendation.Authority Frontier · Alpha EvalsSynthetic configurations · methodology preview · not official Alpha results020406080100020406080100HIGHER-AUTHORITY REVIEW ZONERESTRICTED / ASSISTIVE ZONEA · A2B · A1C · A2D · A1E · A1F · A0Boardroom effectiveness →Governance assurance →

Authority is non-compensating: strong capability cannot erase a critical governance failure. Bands describe the maximum decision authority a configuration may be considered for after human review; they are not safety certifications.

Decision question

Can the system improve a consequential decision without exceeding its authority?

Representative tasks

  • Packet synthesis
  • Materiality
  • Management challenge
  • Decision framing

Assurance emphasis

Evidence fidelity, abstention, contradiction detection, and explicit human control

Accessible data for the synthetic chart
ConfigurationBoardroom effectivenessGovernance assuranceIllustrative authority bandCritical gate stateHuman reviewGate effect
Configuration A8271A2clearpendingno_synthetic_cap
Configuration B7449A1failed_G3pendingcapped_at_A1
Configuration C6378A2clearpendingno_synthetic_cap
Configuration D5543A1failed_G4pendingcapped_at_A1
Configuration E3867A1clearpendingcapped_at_A1
Configuration F2931A0failed_G5pendingcapped_at_A0

Measurement lineage

How evidence can become an Alpha rating.

External benchmarks are inputs, not evals. A published Alpha rating requires separate evidence review, governance analysis, and approval.

Stage 1Evidence
Stage 2Claims
Stage 3Signals
Stage 4Metrics
Stage 5Evals
Stage 6Indices
Stage 7Evals
Stage 8Outlook / Watch

Publication boundary

Current publication states.

Lab comparison
Public-evidence benchmark; not an Alpha ratingAvailable
Artificial Analysis scores
Withheld pending an authenticated, licensed data feedLimited availability
Boardroom task bank
Candidate first-party method; official ranking set remains privateLimited availability
Official Alpha results
Published only after reproducible evaluation and human reviewLimited availability