What is the difference between input and output token pricing?
Why the rates differ
Input tokens are processed in parallel, so they are cheaper per token. Output tokens are generated one at a time, sequentially, which uses more compute per token. That is why providers charge more for output. Typical ratios range from 1.5x to 4x, but some models have equal rates or even higher output multipliers.
The exact ratio depends on the model and provider. For example, one model might charge $1 per million input tokens and $3 per million output tokens. Another might charge $0.25 input and $1.25 output. Always check the specific model's pricing.
Practical impact
If your application generates long responses, output tokens will dominate your bill. If you send large documents as context, input tokens will dominate. You can reduce cost by trimming prompts, caching repeated input, or limiting output length.
Some providers offer prompt caching, where repeated input tokens are charged at a lower rate. This can significantly cut costs for applications with long system prompts or few-shot examples.
- Input: your prompt, system message, chat history, documents
- Output: the model's generated response
- Output usually costs more per token
- Prompt caching can lower input costs for repeated text
- Set max_tokens to control output cost
Common mistakes
- Thinking input and output tokens cost the same; output is often 2-4x more expensive.
- Overlooking that chat history counts as input tokens on every turn, increasing cost.
- Assuming prompt caching is automatic; it often requires explicit setup and may have minimum token requirements.
