AI models ranked
Curated rankings of frontier and open-source AI models for developers. Scores are calibrated against the current frontier so older generations do not look artificially competitive.
62 models - page 1 of 6
OpenAI's most capable model for the hardest end-to-end work across complex reasoning, software engineering, research, computer use, and professional tools. Its 1.05M context window and 128K maximum output support long multistep workflows, while premium pricing makes it best suited to high-value tasks.
Anthropic's most capable generally available model for demanding reasoning, long-horizon agentic coding, research, and professional work. Published evaluations place it effectively level with Astra on Terminal-Bench 4.0, while its lower cache-read price improves repeated-context agent economics.
Anthropic's recommended starting point for complex agentic coding and enterprise work. It trails the premium Fable tier on the hardest published evaluations, but combines strong long-horizon performance with moderate latency and half the standard token price.
The same underlying model as Fable 5.1 with reduced cybersecurity and biology safeguards for vetted Project Glasswing participants. Its higher specialist ceiling is reflected in coding, while limited availability lowers its general Frontier Fit rather than its capability grades.
OpenAI's GPT-5.6 flagship tier for complex professional work, frontier reasoning, long-horizon coding agents, and demanding tool workflows. It supports configurable reasoning through max effort and remains a strong alternative when Astra's additional capability does not justify its premium price.
OpenAI's newest flagship model for agentic coding, professional knowledge work, data analysis, computer use, and long-running tool workflows. It is positioned as a step up from GPT-5.4 with stronger system understanding, better debugging behavior, and API availability for production developers.
Google's stable Gemini 3.5 production model for sustained frontier performance with strong coding, agentic loops, grounding, tool use, and multimodal inputs. It balances capability, 1M-token context, and production cost better than older Gemini 2.5 entries.
xAI's newest flagship chat model with configurable reasoning, agentic tool calling, low hallucination positioning, image input, and a 1M-token context window. Use server-side search tools when current events or live data matter.
DeepSeek's V4 Pro open-weight model for agentic coding, math, STEM reasoning, and 1M-context workflows. It supports OpenAI-compatible and Anthropic-compatible APIs with thinking and non-thinking modes.
OpenAI's balanced GPT-5.6 tier for everyday agentic coding, product engineering, analysis, tool workflows, and production assistants. Terra lowers token cost substantially relative to Sol, making it the default choice when the flagship tier's deepest reasoning is not required.
OpenAI's fast, lowest-cost GPT-5.6 tier for high-volume assistance, routing, extraction, quick code edits, support automation, and latency-sensitive workflows. Luna gives teams a current-generation model for scale while preserving stronger reasoning than older budget models.
Anthropic's first Fable-generation model for demanding reasoning and long-horizon agentic work. Fable 5.1 succeeds it with stronger coding, research, and workflow performance plus cheaper cache reads.