What is the difference between input and output token pricing?

Updated October 2026 · How we answer

Short answerInput tokens are the text you send to the model; output tokens are the text it generates. Output tokens usually cost more—often 2x to 4x the input rate—because generation is more compute-intensive.

Why the rates differ

Input tokens are processed in parallel, so they are cheaper per token. Output tokens are generated one at a time, sequentially, which uses more compute per token. That is why providers charge more for output. Typical ratios range from 1.5x to 4x, but some models have equal rates or even higher output multipliers.

The exact ratio depends on the model and provider. For example, one model might charge $1 per million input tokens and $3 per million output tokens. Another might charge $0.25 input and $1.25 output. Always check the specific model's pricing.

Practical impact

If your application generates long responses, output tokens will dominate your bill. If you send large documents as context, input tokens will dominate. You can reduce cost by trimming prompts, caching repeated input, or limiting output length.

Some providers offer prompt caching, where repeated input tokens are charged at a lower rate. This can significantly cut costs for applications with long system prompts or few-shot examples.

  • Input: your prompt, system message, chat history, documents
  • Output: the model's generated response
  • Output usually costs more per token
  • Prompt caching can lower input costs for repeated text
  • Set max_tokens to control output cost

Common mistakes

  • Thinking input and output tokens cost the same; output is often 2-4x more expensive.
  • Overlooking that chat history counts as input tokens on every turn, increasing cost.
  • Assuming prompt caching is automatic; it often requires explicit setup and may have minimum token requirements.
From our shopsSwiftCase: Curated phone cases that ship in 48 hours.