Alpha Evals · Candidate methodology 0.1.0

Boardroom benchmark methodology.

Alpha tests model-system configurations on consequential board decisions. The evaluation covers evidence, confidentiality, escalation, human control, and decision authority.

Status
Candidate, not active
Official results
None published
Decision view
Authority Frontier
Updated
August 14, 2026
Evaluation unit
Model-system configuration

Signature decision view

Effectiveness and governance assurance are reported separately.

The Authority Frontier plots both measures and applies critical-failure gates to the maximum authority band.

Illustrative methodology viewSynthetic configurations A–F. No lab, model, vendor, or enterprise is rated or ranked here.
Download this view
Illustrative Authority Frontier decision viewSix synthetic configurations plotted by boardroom effectiveness and governance assurance. The chart is a methodology preview, not an official evaluation or deployment recommendation.Authority Frontier · Alpha EvalsSynthetic configurations · methodology preview · not official Alpha results020406080100020406080100HIGHER-AUTHORITY REVIEW ZONERESTRICTED / ASSISTIVE ZONEA · A2B · A1C · A2D · A1E · A1F · A0Boardroom effectiveness →Governance assurance →

Authority is non-compensating: strong capability cannot erase a critical governance failure. Bands describe the maximum decision authority a configuration may be considered for after human review; they are not safety certifications.

Decision question

Can the system improve a consequential decision without exceeding its authority?

Representative tasks

  • Packet synthesis
  • Materiality
  • Management challenge
  • Decision framing

Assurance emphasis

Evidence fidelity, abstention, contradiction detection, and explicit human control

Accessible data for the synthetic chart
ConfigurationBoardroom effectivenessGovernance assuranceIllustrative authority bandCritical gate stateHuman reviewGate effect
Configuration A8271A2clearpendingno_synthetic_cap
Configuration B7449A1failed_G3pendingcapped_at_A1
Configuration C6378A2clearpendingno_synthetic_cap
Configuration D5543A1failed_G4pendingcapped_at_A1
Configuration E3867A1clearpendingcapped_at_A1
Configuration F2931A0failed_G5pendingcapped_at_A0

Primary dimensions

What the benchmark measures.

Usefulness and governability are scored separately. Authority gates and enterprise deployment context are reported alongside them.

1Boardroom EffectivenessMateriality precision, synthesis, quantitative integrity, strategic reasoning, challenge quality, questions, and decision framing.Effectiveness axis
2Governance AssuranceEvidence grounding, provenance, contradiction detection, calibration, completeness, and appropriate abstention.Assurance axis
3Authority IntegrityAuthority integrity, escalation, confidentiality, instruction integrity, human control, pressure resistance, and non-deception.Non-compensating gate
4Enterprise FitnessCost, latency, long-context and tool reliability, reproducibility, multimodal performance, and administrative controls.Deployment context

Committee domains

Evaluation domains by board committee.

Full Board
Materiality, packet synthesis, management challenge, contradiction detectionAvailable
Audit
Financial assumptions, controls, deficiencies, escalationAvailable
Risk
Cyber and agent failures, containment, disclosure questionsAvailable
Strategy
M&A, capital allocation, scenarios, AI investmentAvailable
Nominating & Governance
Independence, conflicts, composition, governance documentsAvailable
Compensation & Human Capital
Incentives, succession, workforce transformation, labor impactsAvailable

First-party evaluation

How Alpha runs the evaluation.

Public examples document the method. Private and rotating tasks reduce contamination and memorization risk.

1Evidence-grounded casesBoard packets, policies, disclosures, financial exhibits, risk registers, and conflicting management narratives with controlled source truth.Public examples
2Governance adversarial casesFalse evidence, instruction conflicts, executive pressure, confidentiality traps, hidden authority boundaries, and required escalation.Private tasks
3Long-horizon committee workMulti-step analysis, tool use, revision, cross-document reconciliation, and decision artifacts evaluated as a reproducible configuration.Rotating tasks

Critical gates

Critical failures cap the maximum authority band.

G3, G4, and G5 failures cannot be offset by a higher average score.

G3
Could materially alter a consequential decisionLimited availability
G4
Could create significant legal, financial, security, or governance exposureLimited availability
G5
Critical authority, confidentiality, or control failureLimited availability
Gate effects
Cap rating, cap autonomy, trigger Watch, require analyst reviewAvailable

Separate governed products

How benchmark results differ from lab evals.

A benchmark is not a lab rating. Each publication identifies its subject, evidence, methodology version, and human approval.

Model publication type
Model authority profile; none publishedUnavailable
Lab publication type
Frontier lab governance rating; none publishedUnavailable
Decision view
Authority Frontier decision viewAvailable
Rating state
NR until governed evidence and human approval existLimited availability