Alpha Evals · Public evidence edition
Alpha Frontier Lab Benchmark.
Compare nine frontier AI labs across model capability, public safety practice, long-task autonomy, and transparency. Each source measure remains separate; missing coverage is reported as not rated.
Current comparison
Frontier capability and public safety practice.
The chart uses the highest covered Epoch ECI model for each lab and the lab's FLI Summer 2026 safety score. METR autonomy and Stanford transparency are reported in the table without being folded into the axes.
Source ledger
What is used, what is missing, and why.
Public scores retain their original scale and date. Alpha does not convert missing coverage to zero or average unlike measures into a single rank.
First-party benchmark under development
Boardroom effectiveness and governance assurance.
This second chart shows the proposed Alpha evaluation method using synthetic configurations. It is kept separate from the named-lab public-evidence benchmark above.
Measurement lineage
How evidence can become an Alpha rating.
External benchmarks are inputs, not evals. A published Alpha rating requires separate evidence review, governance analysis, and approval.
Publication boundary
Current publication states.
- Lab comparison
- Public-evidence benchmark; not an Alpha ratingAvailable
- Artificial Analysis scores
- Withheld pending an authenticated, licensed data feedLimited availability
- Boardroom task bank
- Candidate first-party method; official ranking set remains privateLimited availability
- Official Alpha results
- Published only after reproducible evaluation and human reviewLimited availability

