LLM API PricesUpdated 2026-08-27 21:58:52 UTC

Method

Every price on this site is read from a public source on a schedule and written to a static page. Nothing is typed by hand, nothing is estimated, and a model we cannot price confidently is dropped rather than guessed.

Sources

SourceHowFragilityOffers
OpenRouterhttps://openrouter.ai/modelsjson apiAPI-stable297
Provider directhttps://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.jsonjson fileAPI-stable171
OrcaRouterhttps://www.orcarouter.ai/modelsjson apiAPI-undocumented147
TokenRouterhttps://www.tokenrouter.com/modelsjson apiAPI-undocumented + derived ratio106
Together AIhttps://www.together.ai/pricinghtml scrapeHTML-scrape38
Nahcrofhttps://ai.nahcrof.com/pricingjson apiAPI-stable22
Google AIhttps://ai.google.dev/gemini-api/docs/pricinghtml scrapeHTML-scrape + dated prices16
Anthropichttps://docs.claude.com/en/docs/about-claude/pricinghtml scrapeHTML-scrape15
InferXhttps://inferx.net/pricing/endpointshtml scrapeHTML-scrape11
OpenAIhttps://platform.openai.com/docs/pricingembedded jsonembedded-JSON9
RunInfrahttps://runinfra.ai/inference-apiembedded jsonembedded-JSON7
DeepSeekhttps://api-docs.deepseek.com/quick_start/pricinghtml scrapeHTML-scrape6

12 of 12 sources answered on the last run. Each is fetched once per run with a User-Agent naming this site, and each source's robots.txt was checked before it was added. A source that returns a well-formed page with zero rows is treated as a failure, not as an empty catalogue: its last known prices are carried forward and marked stale rather than blanked.

Offers, and the price you see

A model is not one price. 158 of 476 models here are sold by more than one source, at genuinely different rates — the widest spread in the catalogue right now is qwen/qwen3.8-27b at 18.9× ($0.0225 to $0.425 per 1M input). So each model carries an offers list: one row per source that sells it, with input, output and cached-input prices and the time it was read. The number in the table is the cheapest of those offers; the model's own page shows all of them.

Excluded from "cheapest", deliberately: promotional $0.00 rates, off-peak tiers such as DeepSeek's discounted window, and any offer missing half its price. All three are still listed on the model page, labelled — they are just not allowed to win.

Metrics beyond price

Only what a source actually publishes, labelled with that source. Where a seller publishes a cached-input rate it renders in that seller's row on the model page (118 models have at least one); where nobody does, the column is absent rather than filled with dashes. Throughput in tokens per second shows on the 29 models whose seller states one — Nahcrof and RunInfra do, nobody else here does. Availability comes from OpenRouter's public API, which reports uptime per serving endpoint; the model page shows the median across those endpoints over the last 30 minutes, for 21 core models.

Latency is shown nowhere, and throughput is never inferred. OpenRouter's website displays both, but its public API returns null for latency_last_30m and throughput_last_30m on every endpoint we measured — 175 of 175. An unpublished number is not a number, so this site leaves those blank rather than estimating them from a benchmark it did not run.

When sources disagree

The first-party price and the OpenRouter price are still cross-checked on every model that carries both. If they differ by more than 20%, the row is flagged differs and neither is assumed correct. 7 rows are flagged right now. Nothing is ever averaged.

Model identity

Comparing sellers only works if the same model is recognised across sources that call it anthropic/claude-opus-5, Claude Opus 5 and claude-opus-5-20260101. Names are normalized to a canonical key, and a variant is never merged with its base — not -fast, not -turbo, not -vision, not a dated snapshot — unless an explicit, hand-checked alias says so. That rule is not caution for its own sake: two sources price deepseek-v4-pro and deepseek-v4-pro-0813 differently, so merging them would average two real offers into one that nobody sells. Unmatched stays unmatched.

Tiers

Assigned by name, not by benchmark. Flagship — frontier-lab top models — the opus / ultra / pro tier. Mid — the workhorse tier: sonnet, gpt-class, gemini, deepseek, large open models. Small — fast and cheap: haiku, flash, mini, nano, lite, and sub-10B open models. Anything the map cannot place goes to Unclassified rather than being forced into a tier it may not belong to.

Core selection

The default view on the models page is a curated core of 42 models — the families in common use, kept as a hand-edited allowlist. It is an editorial selection, not a ranking: a model is outside it because it is niche, superseded, or a variant, never because it priced badly. Nothing is discarded; all 476 priced models are one click away in show all.

Logos

The developer marks are the monochrome set from lobehub/lobe-icons (MIT). Where that set has no mark for a developer, the site draws a plain letter monogram instead — never a similar company's logo and never one redrawn from memory. All product names, logos and brands are property of their respective owners; used for identification only. Not affiliated with any provider.

What is left out

Text chat and completion models only — no embeddings, image, audio, video or rerank models. Zero-priced and half-priced rows are dropped. Batch, flex and priority SKUs are not collected; the headline figure is standard synchronous inference.

Gateways that publish no catalogue are absent rather than estimated. Merge.dev, for example, states only that it charges a 5.0% markup over the underlying provider — that is a formula, not a price list, so no Merge prices are published here. GMI Cloud's open endpoint is a nine-model marketing list behind an inferred scaling constant, and is not collected either. An absent model is better than an invented number.

← Back to the prices