Frontier model roster
The top 20, by AA-Omniscience
Ranked by AA-Omniscience Index, which nets a model’s knowledge against how often it invents an answer. Everything to the right of it is the trade you make to get that: how well it holds a long context, drives a tool, reads an image, and what a single task actually costs once you count the tokens it burns getting there.
| Claude Fable 5 (with fallback) | Anthropic | 1M | 40 | 53 | 70 | — | 60 | 80.65s | 66 | $3.15 | 19.0 | 0.004 |
| Gemini 3.1 Pro Preview | 1M | 33 | 21 | 73 | 82 | 46 | 21.34s | 125 | $0.34 | 135 | 0.116 | |
| Claude Opus 5 (max) | Anthropic | 1M | 31 | 55 | 70 | 85 | 61 | 83.18s | 53 | $2.34 | 26.1 | 0.005 |
| Claude Opus 4.8 (max) | Anthropic | 1M | 27 | 47 | 68 | — | 56 | 14.19s | 53 | $2.03 | 27.6 | 0.021 |
| Grok 4.5 (high) | SpaceXAI | 500K | 26 | 46 | 68 | 80 | 54 | 7.85s | 56 | $0.44 | 123 | 0.135 |
| Claude Opus 4.7 (max) | Anthropic | 1M | 26 | 44 | 70 | 79 | 54 | 18.38s | 47 | $2.22 | 24.3 | 0.015 |
| Gemini 3.6 Flash | 1M | 24 | 39 | 70 | 83 | 50 | 15.55s | 212 | $0.56 | 89.3 | 0.100 | |
| Gemini 3.5 Flash | 1M | 23 | 37 | 71 | 84 | 50 | 19.55s | 174 | $0.69 | 72.5 | 0.065 | |
| GPT-5.6 Sol (max) | OpenAI | 1M | 22 | 54 | 74 | 83 | 59 | 137.84s | 68 | $1.86 | 31.7 | 0.004 |
| GPT-5.5 (xhigh) | OpenAI | 922K | 20 | 45 | 74 | 81 | 55 | 73.05s | 68 | $1.17 | 47.0 | 0.011 |
| Kimi K3 (max) | Kimi | 1.1M | 18 | 50 | 75 | 81 | 57 | 3.29s | 35 | $0.86 | 66.3 | 0.015 |
| Muse Spark 1.1 (xhigh) | Meta | 1.1M | 18 | 38 | 63 | — | 51 | 2.72s | 135 | $0.29 | 176 | 0.162 |
| Grok 4.3 (high) | SpaceXAI | 1M | 18 | 24 | 65 | 78 | 38 | 17.31s | 143 | $0.15 | 253 | 0.321 |
| Gemini 3 Pro Preview (high) | 1M | 16 | — | 71 | 80 | 40 | — | — | — | — | — | |
| Claude Sonnet 5 (max) | Anthropic | 1M | 15 | 47 | 71 | 77 | 53 | 191.94s | 75 | $1.72 | 30.8 | 0.003 |
| Grok 4.20 0309 v2 | SpaceXAI | 2M | 15 | — | 58 | 75 | 37 | 13.42s | 177 | — | — | — |
| Qwen3.7 Max | Alibaba | 1M | 14 | 31 | 69 | — | 46 | 2.40s | 202 | $1.28 | 35.9 | 0.047 |
| Claude Opus 4.6 (max) | Anthropic | 1M | 14 | — | 71 | 75 | 44 | 17.24s | 42 | — | — | — |
| Claude Opus 4.5 | Anthropic | 200K | 13 | — | 74 | 74 | 41 | 15.88s | 48 | — | — | — |
| Grok 4.20 0309 | SpaceXAI | 2M | 13 | — | 59 | 73 | 37 | — | — | — | — | — |
How to read it
The three cost columns are the ones that change decisions. Cost per task is the measured bill for one unit of work. Intelligence per cost says what that money buys in capability; E2E per cost says what it buys in responsiveness. A model can top the intelligence column and still lose both, which is usually the moment a cheaper model turns out to be the right one for the job in front of you.
Sort by any column. The two per-cost columns are shaded by rank, so the brighter the cell, the more capability or speed that model returns per dollar.
What these numbers are and are not
Cost is measured per completed task, not quoted per million tokens, so a verbose model is charged for its verbosity. Two models on the same headline token price can differ several-fold here.
An em dash means Artificial Analysis has not measured that model on that axis. It is absent from the comparison, not scored zero, and it sorts to the bottom of that column in both directions.
Figures are a snapshot taken on 2 August 2026, not a live feed. Every number is Artificial Analysis’ own measurement, not this site’s.