Snapshot of where the frontier AI models stand as of mid-July 2026. Benchmarks move fast — treat this as a point-in-time comparison, not a permanent ranking. Numbers pulled from [LM Council](https://lmcouncil.ai/benchmarks) (Epoch AI / Scale AI data, last updated 2026-07-01) and [BenchLM](https://benchlm.ai/llm-pricing) pricing tables (2026-07-17). ## The frontier lineup - **Claude Opus 4.8** (Anthropic) — released May 2026 - **GPT-5.5** (OpenAI) — released April 2026 - **Gemini 3.1 Pro** (Google) — flagship since February 2026 - **Grok 4.3** (xAI) No single model wins outright — the top handful are within a few points of each other on most composite indexes. The practical choice comes down to task fit, not raw smarts. ## Pricing (per million tokens, API) | Model | Input | Output | Context window | |---|---|---|---| | Claude Opus 4.8 | $5.00 | $25.00 | 1M | | GPT-5.5 | $5.00 | $30.00 | 1M | | Gemini 3.1 Pro | $2.00 | $12.00 | 1M–2M (reports vary) | | Grok 4.3 | $1.25 | $2.50 | 1M | Grok 4.3 is the cheapest frontier-class option by a good margin, but locks its best features behind the $300/month SuperGrok Heavy tier. Claude and GPT-5.5 sit at similar price points; Gemini undercuts both while offering the largest context window. ## Where each one wins **Coding.** Claude leads — Opus 4.7 tops SWE-bench Verified (83.5%) and Terminal-Bench 2.0 (90.2%), and the Opus line sweeps the top of Text Arena (Coding) and GSO (code optimization). GPT-5.5 is close behind on SWE-bench (80.6%). **Reasoning / general knowledge.** Gemini 3.1 Pro leads Humanity's Last Exam (46.4%) and GPQA Diamond is a near-tie at the top between GPT-5.4 Pro, Gemini 3.1 Pro, and GPT-5.5. **Math.** Split between OpenAI and Anthropic — GPT-5.5 Pro and Claude Fable 5 trade the top spot across FrontierMath tiers and OTIS Mock AIME, both near-perfect on competition-style problems. **Long-horizon / agentic tasks.** Claude dominates METR's time-horizon benchmark by a wide margin (Claude Mythos Preview: ~1045 min task length at 50% success vs. ~385 min for the next closest, Gemini 3.1 Pro). **Common-sense reasoning.** Claude Fable 5 leads SimpleBench (81.9%), ahead of Gemini 3.1 Pro and GPT-5.5 Pro. **Multimodal/visual & spatial reasoning.** Gemini's line (3 Pro Preview, 3.1 Pro Preview) leads BALROG (game-playing) and VPCT (physics/visual reasoning) — consistent with Google's multimodal-first positioning. **Value for money.** Not close — Chinese open-weight models (DeepSeek V4, GLM-5.x, MiniMax) crush the frontier labs on score-per-dollar. DeepSeek V4 Pro scores ~76-80 for $0.43/$0.87 per million tokens, an order of magnitude cheaper than Claude/GPT/Gemini at similar capability tiers. ## Caveats - These are third-party benchmark numbers (Epoch AI, Scale AI), not vendor self-reported — may diverge from marketing claims. - "Score" fields on BenchLM's pricing table are provisional composites; treat directional (A > B) more than the absolute numbers. - Fast-moving space — Anthropic, OpenAI, and Google were all shipping point releases (4.6 → 4.7 → 4.8, 5.2 → 5.5, 3 → 3.1) within the same quarter, so exact model names will likely be stale within weeks. Sources: [LM Council benchmarks](https://lmcouncil.ai/benchmarks), [BenchLM LLM pricing](https://benchlm.ai/llm-pricing)