Audited July 16, 2026
Frontier Model Benchmark Matrix
Public numeric capability results for GPT-5.6 Sol, Claude Mythos and Fable 5, Muse Spark 1.1, Grok 4.5, Thinking Machines' Inkling, and Kimi K3. NR means no public numeric result was found in the audited sources.
A benchmark is only one projection of model behavior. Increasingly, a compelling real-world demo, such as watching a model build a game, can feel more persuasive than another leaderboard point because people care whether a model is capable, reliable, and pleasant to work with. That instinct is a useful response to Goodhart's law, but science needs more than vibes: the challenge is to turn those real-use qualities into repeatable evaluations.
Interpretation warning
These numbers are not a clean apples-to-apples leaderboard. Labs and third-party evaluators often use different harnesses, tool access, prompts, checkpoints, compute budgets, and reporting conventions; even the same benchmark name can refer to multiple non-identical settings. I use public numbers from official releases or third-party sources when available, treating them as good-faith reports of the strongest result each source chose to publish. I also assume labs often run many internal evaluations that never appear in public launch materials. This table is about what an outside reader can verify and compare from public evidence. More public reporting would be useful even when numbers are not perfectly compatible, as long as the harness, tools, checkpoints, and other settings are clear enough to interpret the result.
TL;DR
Reporting choices shape the story: For 97 capability benchmark rows in the current matrix, 53 are unique; 44 appear in two or more releases; only 1 appears in all 6. Safety reporting is even more fragmented: 72 of 82 normalized rows appear once in the current 6 releases.
Evaluation is moving toward real work. New additions concentrate in coding, professional tasks, agents/tools, and science rather than another round of static exams.
Benchmark reporting turns over quickly. Across 14 same-lab release transitions, only 56% of previously reported benchmark editions appear again in the next report; 44% are not carried into the next public report, though that does not prove the evaluation itself was retired.
Current releases
Public numeric results from the 6 current release bundles. NR means no public numeric result was found in the audited source.
Reporting coverage is not model quality. More reported benchmark rows do not imply a better model, and NR does not establish that an evaluation was never run.
- Benchmark rows
- 97
- Unique to one release
- 53
- Shared across 2+ releases
- 44
- All 6 reports
- 1
Showing 97 of 97 rows6 selected1 shared by all selected96 with partial coverage
| Benchmark | GPT-5.6Sol | Claude 5Mythos / Fable | Muse Spark1.1 | Grok4.5 | Inklingeffort 0.99 | KimiK3 (max) |
|---|---|---|---|---|---|---|
| Professional | ||||||
| Agents' Last Exam | 52.7 | 40.5 F | NR | NR | NR | NR |
| GDPval-AA v2* | 1747.8 | 1759.6 F | 1381 | NR | 1238 | 1668.0 |
| AA-Briefcase (Elo) | NR | NR | NR | NR | NR | 1548.0 |
| AA Intelligence Index v4.1 | 58.9 | 59.9 F | NR | NR | NR | NR |
| Management Consulting (internal) | 43.2 | 35.5 F | NR | NR | NR | NR |
| Finance Agent v2 | NR | 56.3 F | 57.2 | NR | NR | NR |
| Big Finance Bench | 53.0 | NR | NR | NR | NR | NR |
| JobBench | NR | NR | 54.7 | NR | NR | 52.9 |
| OfficeQA Pro | NR | 57.9 F | NR | NR | NR | 63.3 |
| SpreadsheetBench 2 | NR | NR | NR | NR | NR | 34.8 |
| DECK-Bench (internal) | NR | NR | NR | NR | NR | 73.5 |
| Legal Agent (Harvey held-out) | NR | 13.3 F | NR | NR | NR | NR |
| Legal Agent (full public) | NR | 16.9 M | NR | NR | NR | NR |
| Forecasting | ||||||
| ForecastBench (Brier Index, no search)* | NR | NR | NR | NR | 61.1 +/- 0.79 | NR |
| ForecastBench (Brier Index, with search)* | NR | NR | NR | NR | 63.7 +/- 0.82 | NR |
| Prophet Arena (Brier; lower is better)* | NR | NR | NR | NR | 0.1617 | NR |
| Agents + tools | ||||||
| BrowseComp | 90.4 | 88.0 M | NR | NR | 77.1* | 91.2 |
| BrowseComp multi-agent* | 92.2 U | 93.3 M | NR | NR | NR | NR |
| MCP Atlas | NR | 83.3 F | 88.1 | NR | 74.1 | 84.2 |
| Tau 3 Banking | NR | NR | NR | NR | 23.7 | NR |
| Toolathlon | 58.0 | 61.7 M | NR | NR | NR | NR |
| Toolathlon-Verified | NR | NR | 75.6 | NR | NR | 73.2 |
| AutomationBench | 18.1 | 17.4 F | NR | NR | NR | 30.8 |
| APEX-Agents | NR | NR | NR | NR | NR | 37.6 |
| OSWorld 2.0* | 62.6 | NR | 14.2 / 47.3 | NR | NR | NR |
| OSWorld-Verified | NR | 85.0 M/F | 80.8 | NR | NR | NR |
| WebArena-Verified | NR | NR | 69.0 | NR | NR | NR |
| DeepSearchQA* | NR | 94.2 M | 84.9 | NR | NR | 95.0 |
| Coding | ||||||
| Program Bench | NR | NR | NR | NR | NR | 77.8 |
| SWE-Bench Pro (Public) | 64.6 | 80.3 M | 61.5 | 64.7 | 54.3 | NR |
| Terminal-Bench 2.1 | 88.8 | 88.0 M | 80.0 | 83.3 | 63.8* | 88.3 |
| DeepSWE 1.1 | 72.7 | 69.7 F | 53.3 | 53.0 | NR | 67.5* |
| DeepSWE 1.0 | NR | 66.1 F | NR | 62.0 | NR | NR |
| SWE Marathon | NR | 24.0 F | NR | 29.0 | NR | 42.0 |
| FrontierSWE | NR | NR | NR | NR | NR | 81.2 |
| MLS Bench Lite | NR | NR | NR | NR | NR | 48.3 |
| Kimi Code Bench 2.0 (internal) | NR | NR | NR | NR | NR | 72.9 |
| AA Coding Agent Index v1.1 | 80.0 | 77.2 F | NR | NR | NR | NR |
| SWE-Bench Verified | NR | 95.5 M | NR | NR | 77.6* | NR |
| Design Arena Agentic Web Dev (Elo) | NR | NR | NR | NR | 1257 | NR |
| FrontierCode Diamond | NR | 29.3 F | NR | NR | NR | NR |
| CritPt | NR | 28.6 M | NR | NR | NR | NR |
| ArxivMath | NR | 78.5 M | NR | NR | NR | NR |
| RiemannBench | NR | 55.0 M | NR | NR | NR | NR |
| AI self-improvement | ||||||
| Internal Research Debugging | 68.3 | NR | NR | NR | NR | NR |
| KernelGen 1P | 61.1 | NR | NR | NR | NR | NR |
| NanoGPT | 9.69 | NR | NR | NR | NR | NR |
| PostTrainBench Lite | 50.3 | NR | NR | NR | NR | NR |
| PostTrain Bench | NR | NR | NR | NR | NR | 36.6 |
| RSI Index | 57.9 | NR | NR | NR | NR | NR |
| Reasoning + long context | ||||||
| Humanity's Last Exam (no tools)* | NR | 59.0 M | 52.2 | NR | 29.7* | 43.5 |
| Humanity's Last Exam (with tools)* | NR | 64.5 M | 62.1 | NR | 46.0 | 56.0 |
| AIME 2026 | NR | NR | NR | NR | 97.1 | NR |
| GPQA Diamond | 94.6 | 94.1 M | NR | NR | 87.2 | 93.5 |
| FrontierMath Tier 1-3 v2 | 89.0 | 87.0 F | NR | NR | NR | NR |
| FrontierMath Tier 4 v2 | 83.0 | 87.8 F | NR | NR | NR | NR |
| MRCR 256K-512K | 91.5 | NR | NR | NR | NR | NR |
| MRCR 512K-1M* | 73.8 | NR | 54.1 | NR | NR | NR |
| GraphWalks BFS 256K | 90.7 | 91.1 M | NR | NR | NR | NR |
| GraphWalks BFS 1M | 77.1 | 79.4 M | NR | NR | NR | NR |
| GraphWalks Parents 256K | NR | 99.96 M | NR | NR | NR | NR |
| ARC-AGI-3 | 7.78 | NR | NR | NR | NR | NR |
| Knowledge + chat | ||||||
| SimpleQA Verified | NR | NR | NR | NR | 43.9 | NR |
| AA Omniscience | NR | NR | NR | NR | 2.1 | NR |
| IFBench | NR | NR | NR | NR | 79.8 | NR |
| Global-MMLU-Lite | NR | NR | NR | NR | 88.7 | NR |
| Science + health | ||||||
| HealthBench Professional* | 60.5 | 66.0 M | 59.3 | NR | NR | NR |
| HealthBench* | 57.0 | 62.7 M | NR | NR | NR | NR |
| GeneBench Pro | 28.7 | NR | NR | NR | NR | NR |
| LifeSciBench | 59.9 | NR | NR | NR | NR | NR |
| MedChemBench (internal) | 48.3 | NR | NR | NR | NR | NR |
| BioMysteryBench (hard) | NR | 46.1 M | NR | NR | NR | NR |
| BioMysteryBench (human solved) | NR | 83.9 M | NR | NR | NR | NR |
| Multimodal | ||||||
| gdp.pdf | 30.7 | 29.8 F | NR | NR | NR | NR |
| BenchCAD | 70.6 | 38.4 M | NR | NR | NR | NR |
| BenchCAD + Python | 83.4 | 65.0 M | NR | NR | NR | NR |
| CharXiv Reasoning (no tools) | NR | 88.9 M | NR | NR | 78.1 | 84.8 |
| CharXiv Reasoning (with tools) | NR | 93.5 M | 88.4 | NR | 82.0* | 91.3 |
| BabyVision (with tools) | NR | NR | 76.3 | NR | NR | 85.7 |
| Blueprint-Bench 2 | NR | 38.6 M/F | NR | NR | NR | NR |
| MMMU Pro (no tools)* | 83.0 | NR | NR | NR | 73.5 S10 | 81.6 |
| MMMU Pro (with tools) | 84.6 | NR | NR | NR | NR | 83.4 |
| MathVision (no tools) | NR | NR | NR | NR | NR | 94.3 |
| MathVision (with Python) | NR | NR | NR | NR | NR | 97.8 |
| ZeroBench_main (pass@5) | NR | NR | NR | NR | NR | 23.0 |
| ZeroBench_main + Python (pass@5) | NR | NR | NR | NR | NR | 41.0 |
| WorldVQA ForceAnswer | NR | NR | NR | NR | NR | 51.0 |
| OmniDocBench | NR | NR | NR | NR | NR | 91.1 |
| PerceptionBench (in-house) | NR | NR | NR | NR | NR | 58.5 |
| Audio MC* | NR | NR | NR | NR | 56.6 | NR |
| MMAU | NR | NR | NR | NR | 77.2 | NR |
| VoiceBench* | NR | NR | NR | NR | 91.4 | NR |
| Cybersecurity | ||||||
| CyberGym | 84.5 | 83.8 M | 59.0 | NR | NR | NR |
| ExploitBench | 73.5 | 78.0 M | NR | NR | NR | NR |
| ExploitGym (2h) | 24.9 | NR | 0.6 | NR | NR | NR |
| Capture-the-Flag Challenges | 96.7 | NR | NR | NR | NR | NR |
| SEC-Bench Pro | 71.2 | NR | NR | NR | NR | NR |
Safety reporting
- Evaluation rows
- 82
- Safety dimensions
- 6
- Release safety reports
- 4 / 6
Safety takeaway
Safety reporting is broad but highly fragmented. Of 82 normalized evaluation rows, 72 appear in only one current release document. The widest overlap is 3 reports, reached by VCT / Multimodal Troubleshooting Virology and HealthBench Professional. These counts measure public disclosure overlap, not which model is safest or how much safety testing each lab performed.
VCT / Multimodal Troubleshooting Virology
Public benchmarkExpert-level multimodal questions about virology knowledge and protocol troubleshooting.
| Safety evaluation | GPT-5.6Sol27 reported | Claude 5Mythos / Fable31 reported | Muse Spark1.133 reported | Grok4.50 reported | Inklingeffort 0.993 reported | KimiK3 (max)0 reported |
|---|---|---|---|---|---|---|
| Dangerous capability | ||||||
| - | - | - | ||||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | |||
| - | - | - | - | |||
| - | - | - | - | |||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| Harm refusal | ||||||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| Adversarial robustness | ||||||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | |||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | |||
| - | - | - | - | - | ||
| Control + authorization | ||||||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| Alignment + oversight | ||||||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | |||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | |||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| Human impact | ||||||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | ||||
| - | - | - | - | |||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
| - | - | - | - | - | ||
Where the categories come from
These six categories are an editorial normalization for this audit, not a shared taxonomy endorsed by the labs. They align OpenAI's Model Safety, Alignment, and Preparedness sections; Anthropic's Safeguards, Agentic Safety, Alignment, and Responsible Scaling Policy sections; Meta's Advanced AI Scaling Framework, Adversarial Robustness, and Model Behavior scorecards; and the malicious-use, loss-of-control, and dual-use framing in xAI's earlier Grok 4.20 system card. Thinking Machines' Inkling model card adds quantified refusal and jailbreak tests alongside broader dangerous-capability and human-impact testing. Kimi K3's launch post lists qualitative limitations around thinking-history sensitivity and excessive proactiveness, but no quantified safety evaluation rows.
Control + authorization asks whether a tool-using model stays within the user's intent, seeks consent for consequential actions, and avoids destructive or malicious acts. Human impact groups mental health, child safety, health, bias, election integrity, and related effects on people and groups.
A check means the audited release document published a numeric result or quantified rate. Metrics and protocols are usually not directly comparable, so this map intentionally does not crown a safety winner. Cyber and AI self-improvement capability benchmarks already listed in the current matrix are not duplicated here.
No quantified safety result was found in the audited Grok 4.5 or Kimi K3 launch posts; that does not show no safety testing occurred. xAI's earlier Grok 4.20 system card did report safety evaluations; none are attributed to 4.5 here. The Kimi post records qualitative limitations rather than numeric safety results.
Reporting history
Benchmark reporting history, July 2025 to July 2026
An edition-level audit of 20 first-party flagship release bundles. Each cell answers one question: did that release publicly report a numeric result for this benchmark edition? Run details remain attached to every reported cell.
- Release bundles
- 20
- Benchmark editions
- 164
- One-report editions
- 82
Release-to-release continuity
Retention is the share of the previous report's benchmark editions that appears again in the next report.
+ new - dropped
OpenAI
6 reports
5
23 editions
baseline
5.1
4%
+1 -22
5.2
50%
+23 -1
5.4
63%
+8 -9
5.5
78%
+6 -5
5.6
38%
+26 -15
Claude
6 reports
4.5
21 editions
baseline
4.6
71%
+15 -6
Mythos P
37%
+5 -19
4.7
94%
+18 -1
4.8
79%
+15 -7
Claude 5
93%
+13 -3
Muse
2 reports
1.0
20 editions
baseline
1.1
25%
+14 -15
Grok
4 reports
4
8 editions
baseline
4.1
0%
+5 -8
4.3
0%
+4 -5
4.5
0%
+6 -4
Inkling
1 reports
Inkling
22 editions
baseline
Kimi
1 reports
K3
30 editions
baseline
Benchmark reporting timeline
164 rows
Selected benchmark
Humanity's Last Exam
Reasoning · 15 of 20 scoped reports
OpenAI
5 / 5.2 / 5.4 / 5.5
Absent from latest
Claude
4.5 / 4.6 / Mythos P / 4.7 / 4.8 / Claude 5
Present in latest
Muse
1.0 / 1.1
Present in latest
Grok
4
Absent from latest
Inkling
Inkling
Present in latest
Kimi
K3
Present in latest
| OpenAI | Anthropic | Meta | SpaceXAI | Thinking Machines Lab | Moonshot AI | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Benchmark / edition | 5Aug '25 | 5.1Nov '25 | 5.2Dec '25 | 5.4Mar '26 | 5.5Apr '26 | 5.6Jul '26 | 4.5Nov '25 | 4.6Feb '26 | Mythos PApr '26 | 4.7Apr '26 | 4.8May '26 | Claude 5Jun '26 | 1.0Apr '26 | 1.1Jul '26 | 4Jul '25 | 4.1Nov '25 | 4.3Jun '26 | 4.5Jul '26 | InklingJul '26 | K3Jul '26 |
What the timeline captures
Distinct named editions are separate rows, such as SWE-Bench Verified versus Pro, OSWorld-Verified versus OSWorld 2.0, and Terminal-Bench 2.0 versus 2.1. Run-setting differences such as tool access, reasoning effort, context length, or agent scaffolding stay in the cell detail. A missing cell means no public numeric result was found in that release bundle, not that the lab stopped evaluating the benchmark internally.
Audit boundary
The corpus covers flagship general-purpose releases and their principal first-party capability tables or chapters. Smaller, fast, and domain-specialized model launches are excluded, as are refusal, alignment, personality, and preparedness-only evaluations. This keeps the retention denominator comparable while still including named cyber and life-science capability benchmarks reported in capability sections.
Source corpus (20 release bundles)
OpenAI
Anthropic
- Claude Opus 4.5 Nov 2025
- Claude Opus 4.6 Feb 2026
- Claude Mythos Preview Apr 2026
- Claude Opus 4.7 Apr 2026
- Claude Opus 4.8 May 2026
- Claude Mythos / Fable 5 Jun 2026
Meta
- Muse Spark 1.0 Apr 2026
- Muse Spark 1.1 Jul 2026
Thinking Machines Lab
- Inkling Jul 2026
Moonshot AI
- Kimi K3 Jul 2026
