Skip to content

CALCULATOR · Finance

AI API Cost Calculator

What your workload costs per month on each large language model, with prices read live from an open price list.

Show formula
Monthly cost = (input tokens × input price + output tokens × output price) ÷ 1,000,000 × requests per month

Method reviewed:Per-token list prices, with long-context tiers and cache-read prices where listed, read live from LiteLLM’s open price list · reviewed October 2026

Inputs

$0.05 per 1M input tokens · $0.40 per 1M output tokens

API calls in a month: 1,000 a day is about 30,000.

Input is everything sent — system prompt, chat history, documents and the question. Output includes reasoning tokens. 1,000 tokens is roughly 750 English words.

—

The part of each prompt that repeats exactly — a fixed system prompt or document — and is read from the provider’s prompt cache at a lower price.

Prices per 1M tokens

Enter your requests and tokens to price a workload. Until then, here is every model’s list price.

US dollars per million tokens

US dollars per million tokens
ModelInput / 1MCached / 1MOutput / 1M
gpt-5-nanoOpenAI$0.05$0.005$0.40
gpt-6-lunaOpenAI$0.10$0.01$0.50
gpt-4o-miniOpenAI$0.15$0.075$0.60
gpt-5.6-lunaOpenAI$0.20$0.02$1.20
gpt-5.4-nanoOpenAI$0.20$0.02$1.25
gpt-5-miniOpenAI$0.25$0.025$2.00
gpt-4.1-miniOpenAI$0.40$0.10$1.60
gpt-5.4-miniOpenAI$0.75$0.075$4.50
gpt-5OpenAI$1.25$0.125$10.00
gpt-5.1OpenAI$1.25$0.125$10.00
gpt-5.2OpenAI$1.75$0.175$14.00
gpt-5.3-codexOpenAI$1.75$0.175$14.00
Saved price list

Showing the list from Oct 2, 2026. Your browser checks for newer prices when the page opens.

Source: LiteLLM’s open price list, kept in step with each provider’s own pricing page. Check OpenAI’s official prices

Standard list prices in US dollars, before tax. Batch, priority and regional endpoints, cache writes, tools and images are billed differently.

How it works

  1. Pick a model. Prices load live from LiteLLM’s open price list when the page opens, and the box under the result shows when they were checked.
  2. Enter the requests per month, the average input and output tokens per request, and the share of the input your prompt cache serves.
  3. Read the monthly and yearly cost, the split between input and output, and the same workload priced on every other model.

Frequently asked questions

How many tokens is a word?

In English, about three quarters of a word per token, so 1,000 tokens is roughly 750 words. Other languages, code and numbers use more tokens per word, and each model family has its own tokenizer; the usage figures in the API’s responses give your exact counts.

Why do output tokens cost more than input tokens?

Output is generated one token at a time while the whole input is read in one pass, so providers price output at several times the input rate. Reasoning models bill their hidden reasoning as output too, which is why long reasoning raises the bill.

What is prompt caching?

When the start of a prompt repeats exactly across requests — a long system prompt or a document — providers serve it from a cache and charge a fraction of the input price. Enter the share of input that repeats to see the saving; the first request that writes the cache can cost extra, which this page leaves out.

Where do the prices come from, and how current are they?

From LiteLLM’s open price list, which its maintainers keep in step with each provider’s pricing page, usually within days of a change. The page reads it live in your browser and shows when it last checked; if the list cannot be reached it shows the saved copy and its date. Confirm on the provider’s own page before committing a budget.

Examples

A model at $1 per million input tokens and $4 per million output: 10,000 requests a month of 2,000 input and 500 output tokens
$0.004 a request: $40 a month and $480 a year, half for input and half for output.
The same workload with half of each prompt read from a cache priced at $0.10 per million
$31 a month: the cached half of the input costs $1 instead of $10.