Home

Frontier model roster

The top 20, by AA-Omniscience

Ranked by AA-Omniscience Index, which nets a model’s knowledge against how often it invents an answer. Everything to the right of it is the trade you make to get that: how well it holds a long context, drives a tool, reads an image, and what a single task actually costs once you count the tokens it burns getting there.

Top 20 scored models by Omniscience · captured 2 September 2026Data: Artificial Analysis
Claude Fable 5.1 (max with fallback)Anthropic1M43618066296.81s66$3.6917.90.001
Claude Fable 5 (with fallback)Anthropic1M43577762120.99s67$3.1419.70.002
Claude Opus 5 (max)Anthropic1M375979856371.65s56$2.3426.90.005
Gemini 3.1 Pro PreviewGoogle1M322379824822.15s103$0.331450.112
Grok 4.6 (high)SpaceXAI500K3059796142.00s53$0.9464.90.021
Gemini 3.8 Flash (high)Google1M305082865913.39s305$0.581020.115
Claude Opus 4.8 (max)Anthropic1M2949735745.05s58$2.0328.10.009
Muse Spark 1.1 (xhigh)Meta1.1M284081531.45s178$0.291830.223
Muse Spark 1.2 (xhigh)Meta1.1M2749835715.39s154$0.401430.079
Claude Opus 4.7 (max)Anthropic1M274675795515.33s48$2.2324.70.017
Gemini 3.7 Flash (high)Google1M264581855610.85s279$0.401400.198
Grok 4.5 (high)SpaceXAI500K254974805611.20s51$0.431300.111
GPT-5.6 Sol (max)OpenAI1M2258788361100.85s72$0.9564.20.010
Gemini 3.6 FlashGoogle1M224179835218.49s168$0.341530.137
GPT-5.5 (xhigh)OpenAI922K214779815658.98s79$1.1947.10.013
Gemini 3.5 FlashGoogle1M214081845213.89s189$0.6975.40.088
Kimi K3 (max)Kimi1.1M20548381604.81s38$0.8471.40.017
Grok 4.3 (high)SpaceXAI1M182468783818.46s113$0.152530.291
Claude Sonnet 5 (max)Anthropic1M1650777755178.65s70$1.7232.00.003
Gemini 3 Pro Preview (high)Google1M15738041

How to read it

The three cost columns are the ones that change decisions. Cost per task is the measured bill for one unit of work. Intelligence per cost says what that money buys in capability; E2E per cost says what it buys in responsiveness. A model can top the intelligence column and still lose both, which is usually the moment a cheaper model turns out to be the right one for the job in front of you.

Sort by any column. The two per-cost columns are shaded by rank, so the brighter the cell, the more capability or speed that model returns per dollar.

What these numbers are and are not

Cost is measured per completed task, not quoted per million tokens, so a verbose model is charged for its verbosity. Two models on the same headline token price can differ several-fold here.

An em dash means Artificial Analysis has not measured that model on that axis. It is absent from the comparison, not scored zero, and it sorts to the bottom of that column in both directions.

Figures are a snapshot taken on 2 September 2026, not a live feed. Every number is Artificial Analysis’ own measurement, not this site’s.