AI Model Performance Comparison Tool
A weighted LLM leaderboard that lets you compare AI models side by side across real-world coding (SWE-bench), terminal and agent skills, PhD-level reasoning (GPQA Diamond), frontier intelligence (ARC-AGI-2 and Humanity’s Last Exam), web and computer use, Chatbot Arena human preference, and speed and cost. Drag the priority sliders and the LLM rankings below re-order instantly for your exact workload.
Benchmark data compiled from SWE-bench Verified, Terminal-Bench, GPQA Diamond, ARC-AGI (ARC Prize), Humanity’s Last Exam, BrowseComp, OSWorld-Verified, LMArena (Chatbot Arena), and Artificial Analysis speed and pricing measurements.
Embedded dataset · September 2026Set Your Priorities, Get Your LLM Rankings
Every AI leaderboard tells a slightly different story because every team weighs capability differently. A startup shipping a coding agent cares about SWE-bench Verified scores; a research team lives on GPQA Diamond; a consumer chatbot product should watch Chatbot Arena Elo, where millions of real users vote blind on which answer they prefer. Pick a preset or drag the sliders, and the composite score, gauge, and full model leaderboard update in real time across all 22 models — including GPT-6 Astra, launched in early September 2026.
Best match for your priorities
—
—
—
Full Weighted LLM Leaderboard
Unpublished benchmarks are excluded from a model’s weighted average rather than counted as zero — no model is punished with a fake zero for a missing score. To keep that fair, each score is then multiplied by a data-coverage confidence factor: models covering at least 85% of your selected priority weight score at full confidence — so one unpublished benchmark never decides a matchup — while sparser coverage is discounted toward a 50% floor. The coverage badge on each card shows how many of the seven pillars are measured.
Head-to-Head: Compare Any Two AI Models Side by Side
Pick any two of the 22 models for a direct AI model comparison across all seven pillars — the fastest way to answer questions like GPT-6 Astra vs Claude Fable 5.1 or Kimi K3 vs DeepSeek V4 Pro. Bars show each model’s normalized pillar score; a pillar only names a leader when both models have published data for it. When a headline pillar has no shared score, the tool drops to raw shared benchmarks underneath — so a matchup like GPT-6 Astra vs Claude Fable 5.1 still gets a real coding and human-preference comparison even though neither model has published a SWE-bench Verified figure.
—
“No published score yet” means exactly that — no verified figure exists for that model on that pillar, so the pillar names no winner and the tool looks for shared evidence underneath instead. GPT-6 Astra, for example, has no SWE-bench Verified result and no overall Arena Elo yet (its votes have only accumulated since its September 4 Arena debut), but it does report DeepSWE v1.1 and holds a WebDev Code Arena rating — so those appear as raw shared rows. A comparison tool that filled the gap with an invented number would be guessing; this one shows you the evidence that actually exists.
Who Leads Each AI Benchmark Right Now
The current #1 model on every benchmark this tool tracks, computed live from the dataset — a visual answer to “which AI model is the smartest” that shows why the honest answer is always “at what?”
Bars are drawn to each benchmark’s own scale; Arena Elo is drawn on its normalized band. When the live data feed updates, this chart updates with it.
AI Model Benchmark Comparison Table (22 Models, 8 Benchmarks)
The raw scores behind the LLM rankings above. This table answers the questions people type into Google and AI assistants — which AI model is best for coding in 2026, which LLM tops the Chatbot Arena leaderboard, and which model gives the best speed and cost balance. A dash means the score has not been published on a source we track.
| Model | Vendor | SWE-bench Verified | Terminal-Bench 2.1 | GPQA Diamond | ARC-AGI-2 | HLE | BrowseComp | OSWorld-Verified | OSWorld 2.0 | Arena Elo | API Price (in/out per 1M) | Speed |
|---|
← Swipe sideways to see all columns →
*Vendor-reported on the provider’s own harness. †Introductory pricing. Scores blend independent leaderboards (LLM Stats, BenchLM, Scale SEAL, Vals AI, Vellum, LMArena via Hugging Face) with vendor system cards; agentic benchmarks are scaffold-dependent, so the same model can score several points apart between harnesses. Arena Elo figures use a single consistent LMArena snapshot — Elo from different trackers is not directly comparable because baselines get re-calibrated. Speed figures are median output tokens per second from independent throughput measurements where published.
LLM Leaderboard vs Chatbot Arena: Two Ways to Rank AI Models
Most LLM rankings you find online come from one of two philosophies, and understanding the difference will save you from picking the wrong model for the wrong reason.
Benchmark leaderboards — like the weighted one on this page — score models against fixed, verifiable tasks: did the patch pass the test suite, did the agent finish the terminal job, did the answer match the correct chemistry solution. They measure capability objectively but only for the skills the benchmark covers.
Chatbot Arena (LMArena) takes the opposite approach: millions of real users chat with two anonymous models side by side and vote for the better answer. The resulting Elo rating captures something benchmarks miss — helpfulness, tone, formatting, and how answers feel to actual humans. As of the September 2026 LMArena snapshot, Claude Opus 5 holds the top Elo position with Claude Fable 5 and Gemini’s Flash tier close behind, while benchmark leaderboards for coding show GPT-5.6 Sol in front. Neither ranking is wrong; they measure different things.
Practical rule: building an agent or developer tool? Weight the benchmark pillars. Building a consumer-facing chat product? Give the Human Preference slider real weight — Arena Elo is the closest thing the industry has to a customer-satisfaction score. That is exactly why this tool includes both in one weighted LLM leaderboard.
The Arena also runs specialized sub-leaderboards, and those move faster than the overall board: within two days of joining, GPT-6 Astra took the WebDev Code Arena at 1,797 Elo, edging Claude Fable 5.1’s 1,762 — the first community-vote signal on the new flagship while its overall text Elo is still accumulating votes. Also note that Arena Elo can be gamed by style — models that format nicely and answer confidently gain points independent of accuracy — which is why LMArena added Style Control filtering and why serious model selection uses Arena alongside benchmarks, never instead of them.
What Each Benchmark Actually Measures — and Why It Matters
Real-World Coding: SWE-bench Verified
SWE-bench Verified drops a model into a real open-source repository with a genuine bug report and asks it to produce a patch that passes the project’s own test suite. It is the closest public proxy we have for “can this AI fix my production code.” Frontier models now resolve over 90% of these issues, which is why harder follow-ups like SWE-bench Pro exist — the same models drop 20 to 40 points on unseen, contamination-resistant codebases. If you are choosing the best LLM for coding, look at Verified for headline capability and Pro for humility.
Terminal & Agent Skills: Terminal-Bench 2.1
Terminal-Bench puts a model in a live shell and gives it multi-step jobs — compile this project, migrate that database, debug a failing pipeline — with no human help between steps. It measures the agentic skills that separate a chatbot from a genuine AI coworker: planning, recovering from errors, and chaining dozens of commands. This is the benchmark to watch if you plan to run autonomous coding agents.
Advanced Reasoning & Science: GPQA Diamond
GPQA Diamond is a set of graduate-level physics, chemistry, and biology questions written so that even skilled non-experts with full internet access score near 34%. Frontier models now cluster between 91% and 95%, so it is approaching saturation — small gaps here matter less than they look. Harder evals like Humanity’s Last Exam and competition math (AIME, where open-weights GLM 5.2 posted a striking 99.2% in 2026) still spread the field.
Frontier Intelligence: ARC-AGI-2 and ARC-AGI-3
The ARC-AGI benchmark leaderboard, run by the ARC Prize Foundation, measures something no other test does: fluid intelligence — the ability to solve abstract visual puzzles the model has never seen, with no memorized knowledge to lean on. Average individual human performance on ARC-AGI-2 sits around 66%, and the benchmark stayed brutal for AI until this generation: GPT-6 Astra now leads at 95%, with GPT-5.6 Sol at 92.5% and Claude Opus 5 at 90.4%. ARC-AGI-3 goes further into interactive agentic reasoning and remains wide open — Claude Opus 5 holds the standard-harness record at roughly 30%, and headline claims of 99.9% come from custom provider scaffolds that the ARC Prize discloses separately. If you want to know which AI model actually thinks rather than recalls, the ARC-AGI scores are the column to watch.
Frontier Knowledge: Humanity’s Last Exam
Humanity’s Last Exam (HLE) is a 2,500-question gauntlet of extreme graduate-level science and expert knowledge, deliberately built to stay unsaturated as models improve. Where GPQA Diamond now clusters everyone in the low-to-mid 90s, HLE still spreads the frontier wide apart: Claude Fable 5.1 leads at 65% and Claude Opus 5 sits at 64.7%, while several capable models land in the 40s. That spread makes HLE one of the most honest difficulty signals left in AI benchmarking — and a key input to this tool’s Frontier Intelligence pillar, which averages each model’s normalized HLE and ARC-AGI-2 scores.
Web Use & Computer Control: BrowseComp and OSWorld-Verified
BrowseComp tests long-horizon web research — hunting down obscure, verifiable facts across many sites — while OSWorld-Verified measures whether a model can operate real desktop software: clicking, typing, and completing office tasks across applications. Together they predict how well a model will handle browser automation, form-filling, and computer-use agents.
Human Preference: LMArena Chatbot Arena Elo
Chatbot Arena is the largest blind taste test in AI: users compare two anonymous model responses and vote, and an Elo rating emerges from millions of pairwise battles. It is the industry’s proxy for “which model do people actually like talking to.” Because Elo baselines get re-calibrated over time, this tool normalizes ratings from a single consistent snapshot rather than mixing numbers from different trackers.
Speed and Cost: Throughput and Blended Price
Capability without affordability is a demo, not a product. The Speed & Cost pillar averages two things: independently measured output tokens per second, and a price-efficiency score built from a blended 80/20 input-output token mix — the ratio most production workloads actually see. Cheaper models score higher on a square-root curve so a large price gap does not completely drown out capability differences, and models without published throughput are scored on price alone.
GPT-6 Astra vs Claude Fable 5.1 vs Gemini: How the Newest Flagship Changes the Rankings
OpenAI began rolling out GPT-6 Astra on September 3, 2026 — the first GPT-6 family model — and this tool already includes its launch numbers. The headline: Astra takes the #1 spot on both GPQA Diamond (96.0%) and ARC-AGI-2 (95%), posts a verified 87.3% on Terminal-Bench 2.1, and matches Claude Fable pricing exactly at $10 input / $50 output per million tokens.
The head-to-head tool above defaults to exactly this matchup, and the verified evidence splits cleanly. Astra dominates the newest computer-automation evals — 72.6% versus Claude Fable 5.1’s 41.7% on the long-horizon OSWorld 2.0 benchmark (both figures from the companies’ own system cards; note OSWorld 2.0 is a far harder successor to the OSWorld-Verified column in our table), plus leads on Terminal-Bench 4.0 and FrontierMath. Astra’s computer-operator design shows up everywhere OpenAI could measure it: 92.7% on ScreenSpot Pro visual UI grounding and 64.6% versus Fable 5.1’s 52.6% on Terminal-Bench Science. On human preference, the picture is early but real: Astra entered the Arena on September 4 and immediately took the WebDev Code Arena crown at 1,797 Elo versus Fable 5.1’s 1,762 — community votes, not vendor claims — while its overall Chatbot Arena Elo is still accumulating votes, which is why the head-to-head tool above honestly shows “no published score yet” on that pillar rather than inventing a number. Availability and cache economics cut the other way: Fable 5.1 has been generally available across the API, AWS, Google Cloud, and Azure since September 1 with cached input at $0.25 per million tokens, while Astra is in limited availability during its phased rollout and prices cached input at $1 — which makes cache-heavy agent loops measurably cheaper on Fable despite identical $10/$50 headline rates.
Coding is where the “most intelligent model” framing gets softer, and it is worth reading carefully because neither flagship has published a SWE-bench Verified score. On the benchmarks they do share, Astra leads narrowly: 74.1% to 67.4% on DeepSWE v1.1 (113 long-horizon engineering tasks) and 57.7% to 55.8% on Terminal-Bench 4.0. But the field bunches up tightly there — Claude Opus 5 sits at 73.7% and even Gemini 3.8 Flash reaches 73.8% on DeepSWE — so it is a thin lead, not a generational gap. Independent testing tilts the other way: Artificial Analysis’s Coding Agent Index, which runs each model in its native harness, puts Claude Fable 5.1 in Claude Code at 70 against Astra in Codex at 67, and its aggregate Intelligence Index has Fable 5.1 at 65.7 versus Astra’s 61.2. The honest summary is the one practitioners keep reaching: Astra wins localized terminal execution and raw token efficiency, Fable 5.1 wins long-horizon context retention and complex system maintenance.
But the GPT-6 Astra vs Claude question is not one-sided. Claude Fable 5.1 still leads Humanity’s Last Exam at 65% versus Astra’s 57.2%, Claude Opus 5 keeps the ARC-AGI-3 standard-harness record and the #1 Chatbot Arena Elo, and Astra’s SWE-bench Verified score is not yet published — OpenAI’s launch table used newer coding benchmarks instead. Several Astra figures are still vendor-reported pending independent verification, its full capability is gated behind OpenAI’s trusted-access program due to a Critical cybersecurity rating, and rollout to all paid tiers is completing over the days following launch. Set your sliders to your workload and let the weighted math decide; on a balanced profile, the top three now sit within a few points of each other.
How to Choose the Best AI Model in 2026: A Practical Framework
After thirteen-plus years of building and ranking web tools, the pattern I see with AI model selection is the same one I see with any technology purchase: people buy the leaderboard winner, then discover their actual bottleneck was something the leaderboard never measured. Use this three-step framework instead:
1. Weight by workload, not by headlines. If 80% of your token spend is code review, a 2-point SWE-bench gap matters more than a 5-point GPQA gap. That is exactly what the sliders above simulate.
2. Check the cost per completed task, not per token. A cheaper model that needs three attempts can cost more than a premium model that succeeds first try. In recent independent measurements, mid-tier frontier models completed standard tasks for roughly $1 or less while flagship tiers ran nearly double.
3. Re-test quarterly. The LLM rankings in this tool changed materially several times in the last twelve months — Chatbot Arena’s #1 spot alone changed hands repeatedly, and GPT-6 Astra rearranged two leaderboards the week it launched. Model choice in 2026 is a subscription decision, not a marriage.
One more pattern worth naming: there is no single “smartest AI in the world” anymore, and any AI model ranking that gives you one number is hiding a weighting decision from you. Astra leads fluid reasoning, Fable 5.1 leads frontier knowledge, Sol leads real-repository coding, Opus 5 leads human preference, and open-weights models like Kimi K3 lead value. This tool exists precisely to make that weighting decision yours — visible, adjustable, and honest about missing data.
Frequently Asked Questions
More Interactive Tools From DexoCalc
If you are budgeting for AI API spend alongside the rest of your finances, our mortgage calculators hub helps you model the biggest line item most households carry, and the Trump Accounts Calculator projects long-term savings growth with the same interactive, slider-driven approach used in this tool.
