← Today's edition

AI & Compute PLATFORM

AI Has a Wholesale Price. It Is Much Lower Than the Sticker Price.

Gateway pricing puts frontier models 60% to 75% below first-party API rates—and capable images near two cents an output.

GPU server racks receding like a commodities warehouse, with abstract glowing token streams between anonymous market stalls and a distant price board

For most of the generative-AI boom, intelligence had a single sticker price on the model maker's site. A survey of multi-model gateways in October 2026 finds frontier language models advertised far below those rates, with image generation clustering around pennies. The question is shifting from which model is best to where inference clears in the open market.

For most of the generative-AI boom, the price of intelligence appeared straightforward.

OpenAI, Anthropic, Google or xAI published a price for a million tokens. Developers chose a model, entered a credit card and paid it.

That is no longer the whole market.

We reviewed current pricing across third-party multi-model AI gateways and found something remarkable: access to some of the industry’s newest models is being advertised for 40%, 60%, even 70% or more below their creators’ standard API prices.

One frontier OpenAI model officially priced at $2 per million input tokens and $10 per million output tokens was available through a third-party gateway for $0.60 and $3 respectively.

A Claude model officially priced at $2/$10 was available for $0.80/$4.

A current Grok model officially priced at $2/$6 could be found for $0.80/$2.40.

And Google’s newest Flash-class model, officially offered at an introductory $0.75/$3.75, was being advertised elsewhere for roughly $0.225/$1.125.

These are not comparisons against old models or promotional predecessors. OpenAI currently lists GPT-6 Sol at $2 per million input tokens and $10 per million output tokens, while GPT-6 Luna costs just $0.10/$0.50. Anthropic prices Claude Sonnet 5.5 at $2/$10 and Opus 5.5 at $4/$20. Google is offering Gemini 3.8 Flash at $0.75/$3.75 through December 31, 2026. And xAI lists Grok 4.7 at $2/$6.

The difference is large enough that the official API price is beginning to look less like a universal market price and more like a manufacturer’s suggested retail price.

And underneath it, a wholesale market for AI inference appears to be forming.

The Same Model, A Very Different Price

Our October 2026 market snapshot produced comparisons like these:

Model classFirst-party standard priceObserved gateway priceApprox. discount
Frontier OpenAI reasoning model$2 input / $10 output$0.60 / $3.0070%
OpenAI high-end model$10 / $50$2.80 / $1472%
Claude Sonnet-class$2 / $10$0.80 / $460%
Claude Opus-class$4 / $20$1.60 / $860%
Grok frontier chat$2 / $6$0.80 / $2.4060%
Gemini Flash-class$0.75 / $3.75$0.225 / $1.12570%

Gateway prices are a snapshot of rates observed during our research and can change independently of first-party prices.

A billion input tokens through a model costing $2 per million is $2,000.

At $0.60, it is $600.

If the workload also produces 200 million output tokens at $10 per million, the official standard-api bill becomes another $2,000. At $3 through the gateway, it is $600.

That is:

$4,000 versus $1,200.

Scale that workload by 100 and the difference becomes:

$400,000 versus $120,000.

For a casual chatbot, perhaps that does not matter much.

For an autonomous research system, coding agent, search engine, synthetic-data pipeline, customer-support operation or any product that can burn billions of tokens, it matters enormously.

Two invoice stacks on a trader's desk—one large marked $400K, one thin marked $120K—with server reflections and a mechanical calculator

The Obvious Explanation Only Gets Us Halfway There

The first objection is straightforward: the model companies themselves already discount large or flexible workloads.

Correct.

OpenAI’s Batch API costs 50% less than synchronous standard inference and permits jobs to complete within a 24-hour window. OpenAI also prices GPT-6 Luna’s Batch and Flex processing at 50% below Standard.

Anthropic similarly advertises 50% savings through batch processing.

Google’s paid Gemini tier includes a Batch API offering a 50% cost reduction.

So comparing a third-party gateway solely with the full first-party synchronous price exaggerates the mystery.

But it does not eliminate it.

Consider a model priced officially at:

$2 input / $10 output

A 50% batch rate would imply approximately:

$1 / $5

We observed:

$0.60 / $3

The gateway is therefore not merely 70% below standard pricing.

It remains 40% below the already-discounted first-party batch price.

For the Claude example:

Standard: $2 / $10

50%-off batch: $1 / $5

Observed gateway: $0.80 / $4

Still another 20% below batch.

For the Google example:

Standard: $0.75 / $3.75

Approximate half-price batch equivalent: $0.375 / $1.875

Observed gateway: $0.225 / $1.125

Again, roughly 40% lower.

Something more than the obvious public discount is happening.

What Exactly Are You Buying?

This is where the comparison becomes more complicated — and more interesting.

A model name is not necessarily a complete description of an inference product.

Two APIs can both advertise access to the same named model while differing in:

  • latency;
  • concurrency;
  • service priority;
  • availability guarantees;
  • routing;
  • geographic processing;
  • context limits;
  • caching;
  • tool availability;
  • retention policies;
  • rate limits;
  • upstream infrastructure;
  • model snapshot;
  • and contractual guarantees.

This matters because first-party vendors themselves already demonstrate that the same intelligence can have several prices depending on how it is served.

OpenAI sells Standard, Batch, Flex and faster processing tiers.

xAI does something similar in the other direction: Grok 4.7 Fast is the same model on faster infrastructure and costs more than the standard service.

The model may be the same.

The infrastructure contract is not.

That distinction may explain part of what the gateway market is selling.

A developer buying directly is often purchasing not just intelligence but priority, provenance, support, predictable routing and a direct contractual relationship with the model maker.

A gateway customer may instead be buying something closer to:

Give me access to this intelligence wherever it is economically available.

That looks much more like a commodity market.

AI Inference Is Starting to Resemble Cloud Compute

Cloud infrastructure already has this structure.

A company can purchase an on-demand virtual machine at one price.

Reserve capacity and it pays less.

Accept interruptible or spot capacity and it pays dramatically less.

Commit to large volumes and another price appears.

A reseller or managed provider may negotiate yet another rate.

The physical computation does not suddenly become different mathematics.

What changes is when it runs, where it runs, how guaranteed it is, and who bears the utilization risk.

AI inference appears to be moving in the same direction.

A gateway that aggregates workloads from thousands of developers has advantages an individual developer does not.

It can potentially:

  • pool demand;
  • negotiate volume contracts;
  • exploit reserved capacity;
  • distribute workloads among regions;
  • make greater use of caching;
  • route traffic among compatible infrastructure;
  • monetize otherwise-idle capacity;
  • smooth workloads through queues;
  • commit to large volumes;
  • or accept weaker guarantees in exchange for lower upstream prices.

None of these possibilities alone proves how any particular gateway achieves its pricing.

The important observation is that the market now contains enough price dispersion for these strategies to matter.

Sealed metal containers with abstract token symbols exchanged between server farms like commodities on a trading floor

Then There Are Images

The same phenomenon is appearing outside language models.

In our survey, capable image models were available around:

1.6 cents

2 cents

2.4 cents

3 cents

per generated image.

Not thumbnails from obviously obsolete models.

Current systems from major AI developers.

For comparison, xAI currently charges $0.04 for a 1K low-quality output from Grok Imagine Image 2.0, rising to $0.08 for a 2K medium-quality image. We observed the 1K version offered through a gateway at $0.02 per image.

This is particularly striking because the previous Grok Imagine model itself costs $0.02 per output through xAI. The intermediary market can therefore make the newer generation available for approximately what the manufacturer charges for the older one.

OpenAI’s image economics illustrate another complication: image generation is increasingly tokenized rather than sold as a simple flat rate. GPT Image 2’s official output cost varies considerably with resolution and quality; OpenAI currently estimates a square image at roughly $0.006 on Low, $0.053 on Medium and $0.211 on High before applicable inputs.

That makes simplistic ”$ per image” comparisons increasingly dangerous.

The right unit is becoming:

cost per accepted output at the required quality.

Cheap Is Not Cheap If You Have To Run It Three Times

Suppose Model A costs:

$0.0162 per image

and Model B costs:

$0.02.

Model A appears 19% cheaper.

But suppose Model A produces a usable result half the time while Model B succeeds on the first attempt.

Two Model A generations cost:

$0.0324

One Model B generation:

$0.02

The apparently more expensive model is now 38% cheaper per accepted result.

This applies even more strongly to language models.

A model that uses fewer tokens, requires fewer retries, invokes fewer tools or makes fewer errors can have higher nominal token prices while producing a lower cost per completed task.

Anthropic makes precisely this point about Claude Sonnet 5.5: its token prices are unchanged from Sonnet 5, but Anthropic says the newer model can cost up to 30% less per task because it needs fewer tokens to complete the same work.

The invoice for intelligence increasingly has two layers:

unit price

and

efficiency of the intelligence consuming those units.

Multiple Outputs Can Distort Image Comparisons Too

There is another trap in image pricing.

Some gateways charge for a generation task that returns several candidate images.

One service we reviewed, for example, advertised a $0.02 generation that could return several images from the same prompt.

That sounds spectacular on a per-image calculation.

But if an editor needs only one photograph for a story, five additional variations may have little economic value. Someone still has to inspect them.

A six-image task costing two cents can technically be described as a fraction of a cent per output.

For the user, however, the economically relevant number may still be:

two cents per decision.

This is why AI pricing comparisons are going to need to become more sophisticated.

The meaningful units are increasingly:

  • cost per accepted image;
  • cost per solved coding task;
  • cost per research report;
  • cost per thousand classified documents;
  • cost per successful agent run;
  • cost per minute saved.

Tokens and generations remain useful billing units.

They are becoming weaker measures of economic output.

Intelligence Is Getting Astonishingly Cheap

Perhaps the most consequential number in the entire survey is not attached to a flagship model.

It is attached to the cheap ones.

OpenAI’s GPT-6 Luna currently costs just $0.10 per million input tokens and $0.50 per million output tokens at standard rates.

At one gateway we reviewed, the corresponding price was approximately:

$0.03 input

and

$0.15 output

per million tokens.

At that rate:

one billion input tokens costs $30.

Ten billion cost $300.

One billion output tokens cost $150.

There was a time, not very long ago, when feeding billions of words through a state-of-the-art language model sounded like something only a giant technology company could afford.

Now the raw inference charge can be less than dinner for two.

The most important consequence may therefore not be that existing AI products become cheaper.

It may be that entire categories of applications become economical only because developers can now afford to be extraordinarily wasteful with intelligence.

Agents can read more documents.

Search systems can query more sources.

Software agents can attempt more solutions.

Classifiers can inspect nearly everything.

Businesses can run intelligence over data that previously was not worth processing at all.

The history of computing repeatedly shows that when the unit cost of a resource collapses, developers do not simply spend less.

They use dramatically more of it.

A vast paper archive feeding into a compact GPU server the size of a briefcase, suggesting billions of words processed for pocket change

Why Would Anyone Still Buy Direct?

Because price is not the only thing companies buy.

Direct first-party access can provide clearer assurances around:

  • contractual data handling;
  • enterprise support;
  • regional processing;
  • security controls;
  • auditability;
  • model provenance;
  • uptime;
  • feature availability;
  • and escalation when something breaks.

OpenAI, for example, explicitly charges a premium for some regional processing options. xAI offers a U.S.-regional endpoint that carries a 10% premium.

That premium is itself revealing.

Where inference occurs has a price.

How fast it occurs has a price.

When it occurs has a price.

Who guarantees it has a price.

And increasingly, the model itself is only one component of that package.

For sensitive medical data, unreleased financial information, privileged legal material, government workloads or valuable proprietary code, the cheapest gateway may be completely irrelevant if it does not offer the contractual protections required by the customer.

For public-web research, bulk summarization, synthetic data, editorial image generation or low-risk agent workloads, the calculation can be very different.

The market is segmenting.

What We Do Not Know

There are limits to what public price sheets can establish.

We cannot determine from advertised pricing alone:

  • what private volume contracts a gateway has negotiated;
  • whether some prices are temporarily subsidized;
  • whether cloud credits play a role;
  • the gateway’s gross margin;
  • the exact infrastructure path behind every request;
  • or whether every named model is contractually guaranteed to be served identically to its first-party counterpart.

Those distinctions matter.

So does one important wording choice:

These services should not automatically be described as selling the identical first-party product for less.

They are selling access to the same named models, often with a different service layer around them.

That is enough to create a market.

It is not enough to assume perfect equivalence.

But The Price Signal Is Real

Even after all of those caveats, the observed discounts are too large and too widespread to dismiss as noise.

We found the pattern across:

  • OpenAI;
  • Anthropic;
  • Google;
  • xAI;
  • ByteDance;
  • Alibaba;
  • Black Forest Labs;
  • and other model makers.

We saw it in:

  • text inference;
  • cached tokens;
  • image generation;
  • image editing;
  • upscaling;
  • and multimodal workflows.

And crucially, the first-party vendors themselves are moving in the same direction.

OpenAI says improvements in inference and caching allowed it to cut GPT-6 Sol and Luna API prices substantially relative to the previous generation. Google explicitly uses batch discounts to sell idle-tolerant compute more cheaply. Anthropic does the same.

The gateway market is not contradicting this trend.

It is accelerating it.

The Sticker Price May No Longer Be The Market Price

For years, AI comparisons revolved around benchmark scores.

Which model can code better?

Which reasons better?

Which understands images?

Which hallucinates less?

Those questions still matter.

But another axis is becoming impossible to ignore:

What does a unit of useful intelligence actually clear for in the open market?

The answer is becoming surprisingly difficult to infer from an official pricing page.

A model listed at $10 may trade elsewhere at $3.

A four-cent image may appear at two cents.

A billion tokens may cost tens of dollars.

And the gap between those numbers is giving rise to something the AI industry did not have in its early years:

price discovery.

That is what mature commodity markets do.

Different suppliers expose the same underlying resource through different contracts, priorities, guarantees and channels. Buyers decide how much certainty they need and how much they are willing to pay for it.

Cloud computing went through this transition.

Bandwidth went through it.

Storage went through it.

Compute went through it.

AI inference may be next.

The industry’s most important pricing number may soon cease to be the one printed by the company that trained the model.

It may be the price at which someone, somewhere, is willing to run it for you.

And right now, that price is falling remarkably fast.

Methodology

Pricing was reviewed on October 6, 2026. First-party rates were checked against current model-maker documentation. Third-party figures represent advertised rates observed across multi-model API gateways during our research. Gateway pricing can change rapidly and may involve service characteristics that differ from first-party APIs. Discounts refer to advertised unit prices and should not be interpreted as proof of identical latency, provenance, service guarantees or contractual terms.

Continue reading

Sources

October 6, 2026 pricing survey of first-party model documentation and advertised multi-model API gateway rates; methodology note at article end.

More in AI & Compute

View hub →