Skip to main content
Technology

AI Models Ranked: The Best Bang For Your Buck Right Now

We ranked the top AI models of 2026 on benchmark score against real token cost. See which ones win on value, plus the best local picks.

AI Models | Madison Ave Magazine
Courtesy Derek Madison Media

The AI race stopped looking like a race in 2026. It looks like a market now. A dozen labs ship AI models near the top, and the gaps between them keep shrinking. Meanwhile the prices swing wildly. One model can charge a hundred times more per output token than another scoring within five points of it. So the interesting question is no longer which lab is winning. The question is what you get for a dollar.

We pulled the current numbers from independent testing houses and from every provider rate card we could verify. Then we lined them up. What follows is a snapshot of the AI models people are really buying in September 2026, plus the token math that decides your bill.

 

Why AI models are so hard to rank now

Old benchmarks broke. Nearly every frontier model now clears 88 percent on MMLU, the test that defined progress for years. So the scores at the top are noise rather than signal. Stanford’s 2026 AI Index makes the same point. Tests built to last for years get saturated in months.

So the field moved on. The rankings that matter for AI models now blend agentic work, coding, and long-horizon tasks where a model has to actually finish something. The most cited of these is the Artificial Analysis Intelligence Index, which weights agentic tasks at 34 percent of the total score. Others, like the BenchAlign board behind BenchLM, run a wider composite across hundreds of benchmarks.

Neither one is gospel. But both tell about the same story, so we use them as a starting point rather than a verdict.

 

The AI models leading the board

Here is where the outside scoring lands. Notice how tight the top is. Seven points separate the best model from the tenth, and several of those ten cost a fraction of the leader.

 

Artificial Analysis Intelligence Index

Independent composite score, best reasoning effort. Read late August 2026. Scale shown 0 to 70.

Claude Opus 5 63

Claude Fable 5 62

Grok 4.6 61

GPT-5.6 Sol 61

GLM-5.3 60

Kimi K3 60

Qwen3.8-Max 58

Claude Opus 4.8 57

GLM-5.3-Flash 57

Gemini 3.7 Flash 56

Claude Sonnet 5 55

Qwen3.8-27B (open, runs local) 52

Muse Glimmer 30B (open, runs local) 35

Gold bars are US closed models. Grey bars are Chinese or open-weight models. Source: Artificial Analysis, compiled by Capital & Compute.

 

Among closed AI models, Anthropic still holds the ceiling. Claude Opus 5 arrived on July 24 and has led the index since. Its bigger sibling, Claude Fable 5.1, landed on September 1 and posts stronger vendor numbers again, though outside testing has not caught up yet.

Look further down, though, and the picture flips. Three of the top seven models come out of Chinese labs, and two of those three ship downloadable weights.

 

What the top AI models really cost

Now put prices next to those scores for the same AI models. This is where the ranking stops making sense as a shopping list.

ModelLabScoreInput / Output per 1M tokens
Claude Fable 5.1AnthropicNot yet scored$10.00 / $50.00
Claude Opus 5Anthropic63$5.00 / $25.00
GPT-5.6 SolOpenAI61$4.00 / $20.00 (promo)
Grok 4.6xAI61$2.00 / $6.00
Kimi K3Moonshot AI60$3.00 / $15.00
GLM-5.3Z.ai60$1.40 / $4.40
Qwen3.8-MaxAlibaba58$2.00 / $6.00
GLM-5.3-FlashZ.ai57$0.15 / $0.50
Gemini 3.7 FlashGoogle56$0.75 / $3.75 (rises Jan 1)
Claude Sonnet 5Anthropic55$2.00 / $10.00
DeepSeek V4 ProDeepSeekMid-tier$1.32 / $3.96 peak
GPT-5.6 LunaOpenAIEffort-dependent$0.20 / $1.20

Read the GLM-5.3-Flash row again. It lands six points behind the leader and charges a fiftieth as much for output tokens. That one line explains most of what happened to this market in 2026.

 

Token use is where your money really goes

Here is the part almost every buyer of AI models gets wrong. You do not pay per question. You pay per token, and reasoning models emit two kinds of output. There is the answer you see, plus a much larger stream of hidden thinking tokens the model burns on the way there. Those thinking tokens bill at the output rate, and they usually dominate the invoice.

A 2026 Microsoft Research preprint put hard numbers on this. Researchers ran eight frontier models across twelve task suites and compared listed price against real finished-work cost. In 32 percent of head-to-head pairs, the cheaper-per-token model cost more to finish the job. The worst gap hit 28x.

 

A rate card tells you what a model charges. Only the invoice tells you what it costs.

 

Their cleanest example was brutal. Gemini 3 Flash listed 80 percent below GPT-5.4, yet it cost 38 percent more to complete the same twelve suites. Cheaper on paper, pricier in the account.

Why the spread is so wide

On an identical query, one model can burn 900 percent more thinking tokens than another. On agentic work, one can take ten times more turns in the environment. Worse, the cost is not even stable. Repeat one prompt on one model, and the study found swings of up to 9.7x. It depends entirely on how long the model happened to think that run.

You can read the full working in the preprint, or the plain-English breakdown at Capital & Compute.

 

How much AI models really spend to think

There is a public version of this metric. Artificial Analysis publishes the actual dollar cost for each model to complete its whole benchmark suite. Same tests, same tasks, wildly different bills.

 

Cost to run one full Intelligence Index evaluation

Real measured spend, including thinking tokens. Lower is better.

Claude Fable 5 about $6,200

Claude Opus 4.7 about $4,400

Claude Opus 5 $3,836

Claude Opus 4.8 about $3,700

GPT-5.5 (xhigh effort) about $2,900

MiniMax M2.7 about $175

Source: Artificial Analysis published cost-to-run figures, 2026.

 

Opus 4.7 costs more to run than the higher-scoring Opus 4.8. So spend does not buy score in any clean way. It buys verbosity.

The labs know it, and they are fixing it

Token thrift became a selling point for AI models this year. Look at what vendors now brag about.

ModelClaimed token behavior
Grok 4.5About 4.2x fewer output tokens than Opus 4.8 on SWE-Bench Pro
Microsoft MAI-Code-1-FlashUp to 60 percent fewer tokens on SWE-bench Verified
Kimi K2.7 CodeAbout 30 percent fewer reasoning tokens than K2.6
Gemini 3.6 FlashRebuilt for token efficiency after complaints about verbosity
GLM-5.3Tuned for stronger results at smaller output lengths
Claude Sonnet 5Newer tokenizer emits about 30 percent more tokens for the same text

That last row matters more than it looks. A tokenizer change quietly raises your bill even when the sticker price never moves.

Artificial Analysis now measures this directly. Running its full test battery, GLM-5.3 burns about 18,700 output tokens per task. Kimi K3 needs about 14,700 for a matching score of 60. So one open model spends 27 percent more than the other to land in the same place.

Habits differ too. In the Agents on Rails benchmark, Opus 5 wrote the wordiest code diffs in the field, at 1.65 times the task median. GLM-5.3 re-ran the test suite about twenty times per task while most models ran it three to six times. Neither habit improved the solve rate. Both spend your money.

 

The best value AI models per dollar

So flip the leaderboard. Divide benchmark score by blended token price and the order almost inverts. The chart below weights input to output at three to one, which mirrors how coding agents actually bill.

 

Coding points bought per dollar

Composite coding score divided by blended price per 1M tokens. Higher is better.

GPT-5.6 Luna 166.7

MiniMax M3 111.6

Qwen3.7 Plus 99.8

Nemotron 3 Ultra 36.5

Kimi K2.7 Code 35.5

GLM-5.2 32.0

Grok 4.5 25.3

GPT-5.6 Terra 17.1

Gemini 3.1 Pro 15.3

GPT-5.6 Sol 10.0

Claude Opus 4.8 7.4

Claude Fable 5 3.8

Source: Capital & Compute value leaderboard, built on Price Per Token and Artificial Analysis composites.

 

GPT-5.6 Luna buys about seventeen times more coding ability per dollar than GPT-5.6 Sol, its own flagship sibling. Fable 5 sits dead last on value while sitting near first on raw ability. Both facts are true at once, and neither one settles the argument for you.

 

What one job costs across AI models

Value ratios are abstract, so here is a modeled multi-file coding task priced across the major AI models. Same job, different bills.

ModelSticker (in/out)Modeled cost per task
GLM-5.3-Flash$0.15 / $0.50$0.083
GPT-5.6 Luna$0.20 / $1.20$0.105
DeepSeek V4 Flash$0.44 / $1.32$0.138
Gemini 3.7 Flash$0.75 / $3.75$0.364
DeepSeek V4 Pro$1.32 / $3.96$0.416
Kimi K2.7 Code$0.95 / $4.00$0.559
GLM-5.2$1.40 / $4.40$0.737
Qwen3.8-Max$2.00 / $6.00$0.877
Claude Sonnet 5$2.00 / $10.00$0.970
Grok 4.6$2.00 / $6.00$1.21
Kimi K3$3.00 / $15.00$1.46
GPT-5.6 Sol$4.00 / $20.00$1.94
Claude Opus 5$5.00 / $25.00$2.42
Claude Fable 5.1$10.00 / $50.00$3.84

Notice Fable 5.1 versus Fable 5. Anthropic left the headline price alone on September 1 but cut cache reads by 75 percent, from $1.00 down to $0.25 per million. That never shows up in a sticker comparison. Yet it lowers a heavy agent workload by up to about 45 percent, which is a bigger swing than most version bumps deliver.

 

Chinese AI models own the cheap seats

Two years ago Chinese AI models were the budget option. Now they hold the middle of the board and most of the value ranking. Stanford’s index puts Alibaba and DeepSeek inside the top tier of Arena ratings, right alongside the American labs. The open weight gap has not closed, though. As of March 2026 the best closed model still led the best open one by 3.3 percent, up from 0.5 percent in August 2024.

August made the point loudly. Z.ai shipped GLM-5.3 on August 14 at the exact price of GLM-5.2. Then, twelve days later, it shipped GLM-5.3-Flash at about a tenth of that rate. Artificial Analysis scores it three points lower, at 57, but measures its cost per task at $0.09 against $0.68 for the full model.

Alibaba published Qwen3.8-Max weights at 2.4 trillion parameters. Meituan’s LongCat-2.0, meanwhile, was trained entirely on domestic Chinese chips with no Nvidia hardware at all. Tencent and ByteDance both shipped flagship tiers of their own during the same stretch.

 

Watch the license, not the headline

Qwen3.8-Max: Weights published, but a custom license. Resellers need a separate agreement above $50M revenue.

Kimi K3: Weights published under a custom license, gated above $20M revenue.

GLM-5.3: Weights landed in late August under a bespoke license, not the MIT terms its predecessor carried.

GLM-5.3-Flash: MIT licensed, no revenue gate.

Qwen3.8-27B and Muse Glimmer: Both Apache 2.0, the cleanest terms on the board.

 

Prices move in both directions too. DeepSeek switched to peak and off-peak billing on August 16, and even the cheapest hour now costs more than double what the same model charged before. So an open weight license is not automatically a cheap license.

 

Local AI models you can run at home

The most underrated story of 2026 is which AI models now fit on a single graphics card. A 27 billion parameter model now scores 52 on the same independent index where the global leader scores 63. That model costs nothing per token because it runs on your own machine.

What your hardware can handle

Memory availableModel size that fitsGood pick
8 GB3B to 4BGemma 3 4B
16 GB7B to 14BPhi-4, Mistral Small
24 GB27B to 30B at 4-bitQwen3.8-27B
32 GB or more30B dense, small MoEMuse Glimmer 30B
Multi-GPU server320B sparse MoEGLM-5.3-Flash

The two local models worth your time

Alibaba’s Qwen3.8-27B is the current default. It is dense, Apache 2.0, and multimodal, with a 262K native context that stretches toward a million. At 4-bit it fits in about 14 to 17 GB, so a 24 GB card handles it comfortably. It scores 52 on the Intelligence Index against 38 for the model it replaced.

Meta’s Muse Glimmer 30B is the cleaner license story. Apache 2.0 with no acceptable-use addendum and no revenue threshold, about 29.6 billion parameters, quantized under 20 GB. Meta reports 76 percent on SWE-bench Verified and 83.5 percent on GPQA Diamond. Independent scoring is far lower at 35, though, so test it before you trust the launch table.

Worth noting: Meta shipped no Llama model in this window. Its open weight work has moved to the Muse family entirely.

 

Prices for AI models are collapsing

Step back, and the trend across all AI models beats any single release. Average inference pricing fell to about $1.16 per million tokens in early August, down 43 percent from $2.04 at the end of May. Frontier power now prices near 12 percent of what it cost in March 2023.

An analysis by CatalystNeuro traced the actual cheapest route to each skill level over a year. The pattern repeats with unsettling regularity.

Skill tierFirst reachedCost collapse sinceHalving time
Index 30 and upAug 202529xAbout 73 days
Index 40 and upFeb 202656xAbout 28 days
Index 50 and upMar 202635xAbout 29 days
Index 60 and upJun 20263.8xAbout 34 days

Every tier follows one arc. A premium flagship gets there first, then a distilled or open weight model drags the price down by an order of magnitude within months. The skill level that cost $1.22 per task in February costs about two cents now.

 

Locking one model name into your product has become the surest way to overpay for it.

 

Where AI model benchmarks still mislead

Treat every ranking of AI models above as directional. Several problems dent every leaderboard on the internet right now.

  • Vendor numbers are vendor numbers. Most launch tables have never been reproduced by anyone outside the lab that published them.
  • The harness changes the result. Run one task suite through a different scaffold and the scores move, which is why any public board rarely predicts your workload.
  • Some benchmarks are simply broken. An OpenAI audit in July put it that about 30 percent of the public SWE-bench Pro task split is unusable, and the company pulled its earlier advice to lean on it.
  • Lab scores drift from production. Real company rollouts show a wide gap between test scores and live results, with huge cost swings for similar accuracy.

A number that looks decisive on a chart can be close to useless on your own codebase.

 

How to choose AI models for real work

So what should you actually do with all of this? A few rules survive the noise.

  • Rank on cost to finish your own tasks, never on the pricing page. One in three comparisons points the wrong way.
  • Budget against the expensive tail rather than the average, because thinking-token variance can swing a single task by nearly ten times.
  • Check cache read rates separately. For a standing agent, cached input often dominates the bill more than the headline output rate.
  • Route by tier instead of picking one model. Send easy work to a cheap tier and reserve the flagship for tasks that really need it.
  • Read the license before you build on open weights. Downloadable is not the same as unrestricted.
  • Re-check everything quarterly. Records at each tier have been halving about every four to ten weeks.

The easy answer would be that one model wins. It does not. Claude Opus 5 and Fable 5.1 hold the ceiling for truly hard problems. GPT-5.6 Luna and GLM-5.3-Flash own the high-volume work at a fraction of the price. Meanwhile a 27B model on your own graphics card quietly handles a startling amount of the middle. Choosing well now means knowing which of those jobs you are doing.

 

The short version

Best raw score: Claude Opus 5 at 63, with Claude Fable 5.1 posting stronger vendor numbers as of September 1.

Best value overall: GPT-5.6 Luna, at about 167 coding points per dollar.

Best cheap frontier-adjacent pick: GLM-5.3-Flash, scoring 57 at $0.15 in and $0.50 out.

Best local model: Qwen3.8-27B, Apache 2.0, running in about 14 to 17 GB at 4-bit.

Biggest hidden cost: Thinking tokens, which can vary by 900 percent between models on one identical query.

 

Figures reflect published rates and outside tests as of September 2, 2026. Model pricing changes weekly, so verify before committing a budget. Additional model data compiled from the DemandSphere AI Frontier Model Tracker.

DEVARIO JOHNSON

Devario Johnson is the founder and creative lead of Madison Avenue Magazine and Derek Madison Media, where he shapes culture through editorial storytelling, original photography, and platform design. As a fashion editor, media entrepreneur, and senior technology leader, he blends style, innovation, and narrative across every venture. As a former world-class athlete, he brings the same discipline and vision to all his creative pursuits.