LLM API Prices

What text models cost per million tokens — collected from public sources, not hand-typed.

Updated 2026-08-27 20:27:48 UTC
01

Cheapest right now

Ranked by input + output price per million tokens, within tier — comparing across tiers is not a fair comparison, because a flash model is cheap for a reason.

02

The models people actually use

input $/M output $/M bar length = price within tier (√ scale)
03

Method

How the numbers get here, and what they deliberately leave out.

Where the numbers come from

Two machine-readable public sources, re-read on every run: the OpenRouter models API (aggregator prices) and the LiteLLM community price file (direct provider prices). No pages are scraped and no prices are typed by hand.

Direct vs aggregator

When a model appears in both, the direct provider price is shown and the OpenRouter price is listed underneath it. 69 of 402 models carry both. Re-hosts and gateways (Bedrock, Azure, Fireworks, Together…) are excluded, so one open-weight model does not appear fifteen times.

When sources disagree

If the two prices differ by more than 20%, the row is flagged differs and neither is assumed correct. 7 rows are flagged right now. Nothing is averaged.

How tiers are assigned

By name, using a keyword map — opus/ultra/pro from a frontier lab is flagship; haiku/flash/mini/nano and sub-10B models are small; the rest is mid. This is a rough heuristic, not a benchmark. Anything the map cannot place confidently goes to Unclassified rather than being forced into a tier.

Core selection vs the full catalogue

The default view is a curated core of 42 models — the families people actually reach for, chosen by hand and kept as an allowlist in core_models.py. It is an editorial selection, not a ranking: a model is absent from it because it is niche, superseded, or a variant, never because it scored badly. Nothing is discarded — all 402 priced models stay in the data and in the show all view, and the cheapest-per-tier cards recompute for whichever view you are looking at.

A page per model

All 402 models have a static page of their own under /m/ — click any row. Each one carries the same prices at reading size, the context window, where the number came from, how the model sits against its tier, and its nearest neighbours by price.

What is left out

Text chat/completion models only — no embeddings, image, audio, video or rerank models. Rows with a zero or missing price are dropped rather than guessed, which also removes free tiers. Batch, cached-input and priority SKUs are not shown; the price here is standard synchronous inference.

Known gaps

Only providers that publish machine-readable pricing, or that OpenRouter resells, can appear. A provider that lists prices solely on a marketing page is absent by design — an absent model is better than an invented number.