Home

Frontier model roster

The top 20, by AA-Omniscience

Ranked by AA-Omniscience Index, which nets a model’s knowledge against how often it invents an answer. Everything to the right of it is the trade you make to get that: how well it holds a long context, drives a tool, reads an image, and what a single task actually costs once you count the tokens it burns getting there.

Top 20 scored models by Omniscience · captured 2 August 2026Data: Artificial Analysis
Claude Fable 5 (with fallback)Anthropic1M4053706080.65s66$3.1519.00.004
Gemini 3.1 Pro PreviewGoogle1M332173824621.34s125$0.341350.116
Claude Opus 5 (max)Anthropic1M315570856183.18s53$2.3426.10.005
Claude Opus 4.8 (max)Anthropic1M2747685614.19s53$2.0327.60.021
Grok 4.5 (high)SpaceXAI500K26466880547.85s56$0.441230.135
Claude Opus 4.7 (max)Anthropic1M264470795418.38s47$2.2224.30.015
Gemini 3.6 FlashGoogle1M243970835015.55s212$0.5689.30.100
Gemini 3.5 FlashGoogle1M233771845019.55s174$0.6972.50.065
GPT-5.6 Sol (max)OpenAI1M2254748359137.84s68$1.8631.70.004
GPT-5.5 (xhigh)OpenAI922K204574815573.05s68$1.1747.00.011
Kimi K3 (max)Kimi1.1M18507581573.29s35$0.8666.30.015
Muse Spark 1.1 (xhigh)Meta1.1M183863512.72s135$0.291760.162
Grok 4.3 (high)SpaceXAI1M182465783817.31s143$0.152530.321
Gemini 3 Pro Preview (high)Google1M16718040
Claude Sonnet 5 (max)Anthropic1M1547717753191.94s75$1.7230.80.003
Grok 4.20 0309 v2SpaceXAI2M1558753713.42s177
Qwen3.7 MaxAlibaba1M143169462.40s202$1.2835.90.047
Claude Opus 4.6 (max)Anthropic1M1471754417.24s42
Claude Opus 4.5Anthropic200K1374744115.88s48
Grok 4.20 0309SpaceXAI2M13597337

How to read it

The three cost columns are the ones that change decisions. Cost per task is the measured bill for one unit of work. Intelligence per cost says what that money buys in capability; E2E per cost says what it buys in responsiveness. A model can top the intelligence column and still lose both, which is usually the moment a cheaper model turns out to be the right one for the job in front of you.

Sort by any column. The two per-cost columns are shaded by rank, so the brighter the cell, the more capability or speed that model returns per dollar.

What these numbers are and are not

Cost is measured per completed task, not quoted per million tokens, so a verbose model is charged for its verbosity. Two models on the same headline token price can differ several-fold here.

An em dash means Artificial Analysis has not measured that model on that axis. It is absent from the comparison, not scored zero, and it sorts to the bottom of that column in both directions.

Figures are a snapshot taken on 2 August 2026, not a live feed. Every number is Artificial Analysis’ own measurement, not this site’s.