Model Rankings

AI models ranked

Curated rankings of frontier and open-source AI models for developers. Scores are calibrated against the current frontier so older generations do not look artificially competitive.

62 models - page 1 of 6

🥇

GPT-6 Astra

âš¡ Top PickTextCodeImage
OpenAI1.05M ctx

OpenAI's most capable model for the hardest end-to-end work across complex reasoning, software engineering, research, computer use, and professional tools. Its 1.05M context window and 128K maximum output support long multistep workflows, while premium pricing makes it best suited to high-value tasks.

Frontier-relative93
Coding
100
Reasoning
100
Instruction
99
Speed
52
Cost eff.
34
Cost$10 in / $50 out per 1M
Best for
Frontier codingDeep reasoningComputer useAgentic tools
🥈

Claude Fable 5.1

🧠 Best ReasoningTextCodeImage
Anthropic1M ctx

Anthropic's most capable generally available model for demanding reasoning, long-horizon agentic coding, research, and professional work. Published evaluations place it effectively level with Astra on Terminal-Bench 4.0, while its lower cache-read price improves repeated-context agent economics.

Frontier-relative92
Coding
99
Reasoning
99
Instruction
98
Speed
48
Cost eff.
38
Cost$10 in / $50 out per 1M
Best for
Long-horizon agentsFrontier codingResearch1M context
🥉

Claude Opus 5

âš¡ Top PickTextCodeImage
Anthropic1M ctx

Anthropic's recommended starting point for complex agentic coding and enterprise work. It trails the premium Fable tier on the hardest published evaluations, but combines strong long-horizon performance with moderate latency and half the standard token price.

Frontier-relative91
Coding
95
Reasoning
96
Instruction
96
Speed
62
Cost eff.
44
Cost$5 in / $25 out per 1M
Best for
Agentic codingEnterprise workModerate latency1M context
#4

Claude Mythos 5.1

🔒 Limited AccessTextCodeImage
Anthropic1M ctx

The same underlying model as Fable 5.1 with reduced cybersecurity and biology safeguards for vetted Project Glasswing participants. Its higher specialist ceiling is reflected in coding, while limited availability lowers its general Frontier Fit rather than its capability grades.

Frontier-relative90
Coding
100
Reasoning
99
Instruction
98
Speed
48
Cost eff.
38
Cost$10 in / $50 out per 1M
Best for
CybersecurityLife sciencesFrontier codingInvite-only access
#5

GPT-5.6 Sol

âš¡ Top PickTextCodeImage
OpenAI1.05M ctx

OpenAI's GPT-5.6 flagship tier for complex professional work, frontier reasoning, long-horizon coding agents, and demanding tool workflows. It supports configurable reasoning through max effort and remains a strong alternative when Astra's additional capability does not justify its premium price.

Frontier-relative88
Coding
89
Reasoning
91
Instruction
93
Speed
58
Cost eff.
56
Cost$4 in / $20 out per 1M
Best for
Frontier codingDeep reasoningCybersecuritySubagent orchestration
#6

GPT-5.5

âš¡ Top PickTextCode
OpenAI1M ctx

OpenAI's newest flagship model for agentic coding, professional knowledge work, data analysis, computer use, and long-running tool workflows. It is positioned as a step up from GPT-5.4 with stronger system understanding, better debugging behavior, and API availability for production developers.

Frontier-relative88
Coding
89
Reasoning
94
Instruction
90
Speed
64
Cost eff.
48
Cost$5 in / $30 out per 1M
Best for
Agentic codingComputer useTool workflowsProfessional work
#7

Gemini 3.5 Flash

TextCodeImage
Google1.048576M ctx

Google's stable Gemini 3.5 production model for sustained frontier performance with strong coding, agentic loops, grounding, tool use, and multimodal inputs. It balances capability, 1M-token context, and production cost better than older Gemini 2.5 entries.

Frontier-relative87
Coding
85
Reasoning
90
Instruction
86
Speed
78
Cost eff.
68
Cost$1.5 in / $9 out per 1M
Best for
Agent loops1M contextGroundingMultimodal input
#8

Grok 4.3

TextCodeImage
xAI1M ctx

xAI's newest flagship chat model with configurable reasoning, agentic tool calling, low hallucination positioning, image input, and a 1M-token context window. Use server-side search tools when current events or live data matter.

Frontier-relative87
Coding
84
Reasoning
90
Instruction
82
Speed
82
Cost eff.
78
Cost$1.25 in / $2.5 out per 1M
Best for
Agentic toolsReasoning control1M contextSearch workflows
#9

DeepSeek V4 Pro

🔓 Open SourceTextCode
DeepSeek1M ctx

DeepSeek's V4 Pro open-weight model for agentic coding, math, STEM reasoning, and 1M-context workflows. It supports OpenAI-compatible and Anthropic-compatible APIs with thinking and non-thinking modes.

Frontier-relative87
Coding
86
Reasoning
91
Instruction
81
Speed
64
Cost eff.
96
Cost$0.435 in / $0.87 out per 1M
Best for
Open weights1M contextThinking modeLow API cost
#10

GPT-5.6 Terra

💻 Best CodingTextCodeImage
OpenAI1.05M ctx

OpenAI's balanced GPT-5.6 tier for everyday agentic coding, product engineering, analysis, tool workflows, and production assistants. Terra lowers token cost substantially relative to Sol, making it the default choice when the flagship tier's deepest reasoning is not required.

Frontier-relative86
Coding
82
Reasoning
86
Instruction
90
Speed
74
Cost eff.
68
Cost$2 in / $12 out per 1M
Best for
Everyday agentsCoding valueTool workflowsBalanced cost
#11

GPT-5.6 Luna

âš¡ FastestTextCodeImage
OpenAI1.05M ctx

OpenAI's fast, lowest-cost GPT-5.6 tier for high-volume assistance, routing, extraction, quick code edits, support automation, and latency-sensitive workflows. Luna gives teams a current-generation model for scale while preserving stronger reasoning than older budget models.

Frontier-relative86
Coding
76
Reasoning
82
Instruction
87
Speed
92
Cost eff.
96
Cost$0.2 in / $1.2 out per 1M
Best for
High volumeFast assistanceRoutingLow-cost coding
#12

Claude Fable 5

TextCodeImage
Anthropic1M ctx

Anthropic's first Fable-generation model for demanding reasoning and long-horizon agentic work. Fable 5.1 succeeds it with stronger coding, research, and workflow performance plus cheaper cache reads.

Frontier-relative86
Coding
87
Reasoning
94
Instruction
91
Speed
50
Cost eff.
34
Cost$10 in / $50 out per 1M
Best for
Long-horizon agentsComplex reasoningHigh autonomy1M context