Audited July 16, 2026

Frontier Model Benchmark Matrix

Public numeric capability results for GPT-5.6 Sol, Claude Mythos and Fable 5, Muse Spark 1.1, Grok 4.5, Thinking Machines' Inkling, and Kimi K3. NR means no public numeric result was found in the audited sources.

A benchmark is only one projection of model behavior. Increasingly, a compelling real-world demo, such as watching a model build a game, can feel more persuasive than another leaderboard point because people care whether a model is capable, reliable, and pleasant to work with. That instinct is a useful response to Goodhart's law, but science needs more than vibes: the challenge is to turn those real-use qualities into repeatable evaluations.

Interpretation warning

These numbers are not a clean apples-to-apples leaderboard. Labs and third-party evaluators often use different harnesses, tool access, prompts, checkpoints, compute budgets, and reporting conventions; even the same benchmark name can refer to multiple non-identical settings. I use public numbers from official releases or third-party sources when available, treating them as good-faith reports of the strongest result each source chose to publish. I also assume labs often run many internal evaluations that never appear in public launch materials. This table is about what an outside reader can verify and compare from public evidence. More public reporting would be useful even when numbers are not perfectly compatible, as long as the harness, tools, checkpoints, and other settings are clear enough to interpret the result.

TL;DR

Reporting choices shape the story: For 97 capability benchmark rows in the current matrix, 53 are unique; 44 appear in two or more releases; only 1 appears in all 6. Safety reporting is even more fragmented: 72 of 82 normalized rows appear once in the current 6 releases.

Evaluation is moving toward real work. New additions concentrate in coding, professional tasks, agents/tools, and science rather than another round of static exams.

Benchmark reporting turns over quickly. Across 14 same-lab release transitions, only 56% of previously reported benchmark editions appear again in the next report; 44% are not carried into the next public report, though that does not prove the evaluation itself was retired.

Current releases

Public numeric results from the 6 current release bundles. NR means no public numeric result was found in the audited source.

Reporting coverage is not model quality. More reported benchmark rows do not imply a better model, and NR does not establish that an evaluation was never run.

Benchmark rows
97
Unique to one release
53
Shared across 2+ releases
44
All 6 reports
1
Models
Rows

Showing 97 of 97 rows6 selected1 shared by all selected96 with partial coverage

BenchmarkGPT-5.6SolClaude 5Mythos / FableMuse Spark1.1Grok4.5Inklingeffort 0.99KimiK3 (max)
Professional
Agents' Last Exam52.740.5 FNRNRNRNR
GDPval-AA v2*1747.81759.6 F1381NR12381668.0
AA-Briefcase (Elo)NRNRNRNRNR1548.0
AA Intelligence Index v4.158.959.9 FNRNRNRNR
Management Consulting (internal)43.235.5 FNRNRNRNR
Finance Agent v2NR56.3 F57.2NRNRNR
Big Finance Bench53.0NRNRNRNRNR
JobBenchNRNR54.7NRNR52.9
OfficeQA ProNR57.9 FNRNRNR63.3
SpreadsheetBench 2NRNRNRNRNR34.8
DECK-Bench (internal)NRNRNRNRNR73.5
Legal Agent (Harvey held-out)NR13.3 FNRNRNRNR
Legal Agent (full public)NR16.9 MNRNRNRNR
Forecasting
ForecastBench (Brier Index, no search)*NRNRNRNR61.1 +/- 0.79NR
ForecastBench (Brier Index, with search)*NRNRNRNR63.7 +/- 0.82NR
Prophet Arena (Brier; lower is better)*NRNRNRNR0.1617NR
Agents + tools
BrowseComp90.488.0 MNRNR77.1*91.2
BrowseComp multi-agent*92.2 U93.3 MNRNRNRNR
MCP AtlasNR83.3 F88.1NR74.184.2
Tau 3 BankingNRNRNRNR23.7NR
Toolathlon58.061.7 MNRNRNRNR
Toolathlon-VerifiedNRNR75.6NRNR73.2
AutomationBench18.117.4 FNRNRNR30.8
APEX-AgentsNRNRNRNRNR37.6
OSWorld 2.0*62.6NR14.2 / 47.3NRNRNR
OSWorld-VerifiedNR85.0 M/F80.8NRNRNR
WebArena-VerifiedNRNR69.0NRNRNR
DeepSearchQA*NR94.2 M84.9NRNR95.0
Coding
Program BenchNRNRNRNRNR77.8
SWE-Bench Pro (Public)64.680.3 M61.564.754.3NR
Terminal-Bench 2.188.888.0 M80.083.363.8*88.3
DeepSWE 1.172.769.7 F53.353.0NR67.5*
DeepSWE 1.0NR66.1 FNR62.0NRNR
SWE MarathonNR24.0 FNR29.0NR42.0
FrontierSWENRNRNRNRNR81.2
MLS Bench LiteNRNRNRNRNR48.3
Kimi Code Bench 2.0 (internal)NRNRNRNRNR72.9
AA Coding Agent Index v1.180.077.2 FNRNRNRNR
SWE-Bench VerifiedNR95.5 MNRNR77.6*NR
Design Arena Agentic Web Dev (Elo)NRNRNRNR1257NR
FrontierCode DiamondNR29.3 FNRNRNRNR
CritPtNR28.6 MNRNRNRNR
ArxivMathNR78.5 MNRNRNRNR
RiemannBenchNR55.0 MNRNRNRNR
AI self-improvement
Internal Research Debugging68.3NRNRNRNRNR
KernelGen 1P61.1NRNRNRNRNR
NanoGPT9.69NRNRNRNRNR
PostTrainBench Lite50.3NRNRNRNRNR
PostTrain BenchNRNRNRNRNR36.6
RSI Index57.9NRNRNRNRNR
Reasoning + long context
Humanity's Last Exam (no tools)*NR59.0 M52.2NR29.7*43.5
Humanity's Last Exam (with tools)*NR64.5 M62.1NR46.056.0
AIME 2026NRNRNRNR97.1NR
GPQA Diamond94.694.1 MNRNR87.293.5
FrontierMath Tier 1-3 v289.087.0 FNRNRNRNR
FrontierMath Tier 4 v283.087.8 FNRNRNRNR
MRCR 256K-512K91.5NRNRNRNRNR
MRCR 512K-1M*73.8NR54.1NRNRNR
GraphWalks BFS 256K90.791.1 MNRNRNRNR
GraphWalks BFS 1M77.179.4 MNRNRNRNR
GraphWalks Parents 256KNR99.96 MNRNRNRNR
ARC-AGI-37.78NRNRNRNRNR
Knowledge + chat
SimpleQA VerifiedNRNRNRNR43.9NR
AA OmniscienceNRNRNRNR2.1NR
IFBenchNRNRNRNR79.8NR
Global-MMLU-LiteNRNRNRNR88.7NR
Science + health
HealthBench Professional*60.566.0 M59.3NRNRNR
HealthBench*57.062.7 MNRNRNRNR
GeneBench Pro28.7NRNRNRNRNR
LifeSciBench59.9NRNRNRNRNR
MedChemBench (internal)48.3NRNRNRNRNR
BioMysteryBench (hard)NR46.1 MNRNRNRNR
BioMysteryBench (human solved)NR83.9 MNRNRNRNR
Multimodal
gdp.pdf30.729.8 FNRNRNRNR
BenchCAD70.638.4 MNRNRNRNR
BenchCAD + Python83.465.0 MNRNRNRNR
CharXiv Reasoning (no tools)NR88.9 MNRNR78.184.8
CharXiv Reasoning (with tools)NR93.5 M88.4NR82.0*91.3
BabyVision (with tools)NRNR76.3NRNR85.7
Blueprint-Bench 2NR38.6 M/FNRNRNRNR
MMMU Pro (no tools)*83.0NRNRNR73.5 S1081.6
MMMU Pro (with tools)84.6NRNRNRNR83.4
MathVision (no tools)NRNRNRNRNR94.3
MathVision (with Python)NRNRNRNRNR97.8
ZeroBench_main (pass@5)NRNRNRNRNR23.0
ZeroBench_main + Python (pass@5)NRNRNRNRNR41.0
WorldVQA ForceAnswerNRNRNRNRNR51.0
OmniDocBenchNRNRNRNRNR91.1
PerceptionBench (in-house)NRNRNRNRNR58.5
Audio MC*NRNRNRNR56.6NR
MMAUNRNRNRNR77.2NR
VoiceBench*NRNRNRNR91.4NR
Cybersecurity
CyberGym84.583.8 M59.0NRNRNR
ExploitBench73.578.0 MNRNRNRNR
ExploitGym (2h)24.9NR0.6NRNRNR
Capture-the-Flag Challenges96.7NRNRNRNRNR
SEC-Bench Pro71.2NRNRNRNRNR

Safety reporting

Evaluation rows
82
Safety dimensions
6
Release safety reports
4 / 6

Safety takeaway

Safety reporting is broad but highly fragmented. Of 82 normalized evaluation rows, 72 appear in only one current release document. The widest overlap is 3 reports, reached by VCT / Multimodal Troubleshooting Virology and HealthBench Professional. These counts measure public disclosure overlap, not which model is safest or how much safety testing each lab performed.

VCT / Multimodal Troubleshooting Virology

Public benchmark

Expert-level multimodal questions about virology knowledge and protocol troubleshooting.

Safety evaluationGPT-5.6Sol27 reportedClaude 5Mythos / Fable31 reportedMuse Spark1.133 reportedGrok4.50 reportedInklingeffort 0.993 reportedKimiK3 (max)0 reported
Dangerous capability
---
-----
-----
-----
-----
-----
-----
----
----
----
-----
-----
-----
-----
-----
-----
-----
Harm refusal
-----
-----
-----
-----
-----
-----
-----
-----
-----
-----
-----
-----
Adversarial robustness
-----
-----
-----
-----
-----
----
-----
-----
-----
----
-----
Control + authorization
-----
-----
-----
-----
-----
-----
-----
Alignment + oversight
-----
-----
-----
-----
-----
-----
-----
-----
-----
-----
----
-----
-----
-----
----
-----
-----
-----
-----
-----
-----
-----
-----
Human impact
-----
-----
-----
---
----
-----
-----
-----
-----
-----
-----
-----

Where the categories come from

These six categories are an editorial normalization for this audit, not a shared taxonomy endorsed by the labs. They align OpenAI's Model Safety, Alignment, and Preparedness sections; Anthropic's Safeguards, Agentic Safety, Alignment, and Responsible Scaling Policy sections; Meta's Advanced AI Scaling Framework, Adversarial Robustness, and Model Behavior scorecards; and the malicious-use, loss-of-control, and dual-use framing in xAI's earlier Grok 4.20 system card. Thinking Machines' Inkling model card adds quantified refusal and jailbreak tests alongside broader dangerous-capability and human-impact testing. Kimi K3's launch post lists qualitative limitations around thinking-history sensitivity and excessive proactiveness, but no quantified safety evaluation rows.

Control + authorization asks whether a tool-using model stays within the user's intent, seeks consent for consequential actions, and avoids destructive or malicious acts. Human impact groups mental health, child safety, health, bias, election integrity, and related effects on people and groups.

A check means the audited release document published a numeric result or quantified rate. Metrics and protocols are usually not directly comparable, so this map intentionally does not crown a safety winner. Cyber and AI self-improvement capability benchmarks already listed in the current matrix are not duplicated here.

No quantified safety result was found in the audited Grok 4.5 or Kimi K3 launch posts; that does not show no safety testing occurred. xAI's earlier Grok 4.20 system card did report safety evaluations; none are attributed to 4.5 here. The Kimi post records qualitative limitations rather than numeric safety results.

Reporting history

Benchmark reporting history, July 2025 to July 2026

An edition-level audit of 20 first-party flagship release bundles. Each cell answers one question: did that release publicly report a numeric result for this benchmark edition? Run details remain attached to every reported cell.

Release bundles
20
Benchmark editions
164
One-report editions
82

Release-to-release continuity

Retention is the share of the previous report's benchmark editions that appears again in the next report.

+ new   - dropped

OpenAI

6 reports

5

23 editions

baseline

5.1

4%

+1   -22

5.2

50%

+23   -1

5.4

63%

+8   -9

5.5

78%

+6   -5

5.6

38%

+26   -15

Claude

6 reports

4.5

21 editions

baseline

4.6

71%

+15   -6

Mythos P

37%

+5   -19

4.7

94%

+18   -1

4.8

79%

+15   -7

Claude 5

93%

+13   -3

Muse

2 reports

1.0

20 editions

baseline

1.1

25%

+14   -15

Grok

4 reports

4

8 editions

baseline

4.1

0%

+5   -8

4.3

0%

+4   -5

4.5

0%

+6   -4

Inkling

1 reports

Inkling

22 editions

baseline

Kimi

1 reports

K3

30 editions

baseline

Benchmark reporting timeline

Lab

164 rows

Selected benchmark

Humanity's Last Exam

Reasoning · 15 of 20 scoped reports

OpenAI

5 / 5.2 / 5.4 / 5.5

Absent from latest

Claude

4.5 / 4.6 / Mythos P / 4.7 / 4.8 / Claude 5

Present in latest

Muse

1.0 / 1.1

Present in latest

Grok

4

Absent from latest

Inkling

Inkling

Present in latest

Kimi

K3

Present in latest

OpenAIAnthropicMetaSpaceXAIThinking Machines LabMoonshot AI
Benchmark / edition5Aug '255.1Nov '255.2Dec '255.4Mar '265.5Apr '265.6Jul '264.5Nov '254.6Feb '26Mythos PApr '264.7Apr '264.8May '26Claude 5Jun '261.0Apr '261.1Jul '264Jul '254.1Nov '254.3Jun '264.5Jul '26InklingJul '26K3Jul '26
reportednot reported

What the timeline captures

Distinct named editions are separate rows, such as SWE-Bench Verified versus Pro, OSWorld-Verified versus OSWorld 2.0, and Terminal-Bench 2.0 versus 2.1. Run-setting differences such as tool access, reasoning effort, context length, or agent scaffolding stay in the cell detail. A missing cell means no public numeric result was found in that release bundle, not that the lab stopped evaluating the benchmark internally.

Audit boundary

The corpus covers flagship general-purpose releases and their principal first-party capability tables or chapters. Smaller, fast, and domain-specialized model launches are excluded, as are refusal, alignment, personality, and preparedness-only evaluations. This keeps the retention denominator comparable while still including named cyber and life-science capability benchmarks reported in capability sections.

Source corpus (20 release bundles)

OpenAI

Meta

SpaceXAI

Thinking Machines Lab

Moonshot AI

Reporting more benchmark rows does not imply a better model. Tint marks the best directly comparable score. * indicates a different setup, metric, checkpoint, or leaderboard snapshot. Inkling results use effort 0.99; its forecasting rows use a nearby pre-release checkpoint. Kimi K3 uses max reasoning, and its launch table discloses benchmark-specific agent harnesses. Its DeepSWE score uses KimiCode on the v1.1 tasks; Kimi also reports 67.3 with mini-SWE-agent. Kimi says a fuller technical report will follow; no report was linked at audit time. S10 denotes MMMU Pro Standard 10. M/F denotes Claude Mythos/Fable; U denotes GPT-5.6 Sol Ultra.

Transparency note: This matrix, safety map, and reporting-history audit were created with assistance from AI agents and checked against the linked first-party sources. Mistakes may remain. If you spot one, please let me know; corrections are welcome and I will update the page.

Thanks to Martin Ziqiao Ma for helpful input on benchmark comparability and reporting caveats.