Is LLM API pricing based on tokens or requests?
Token-based is standard
For chat and completion models, token-based pricing is the norm. You pay for input and output tokens separately. The number of requests does not directly affect cost, though more requests mean more tokens overall. This model aligns cost with usage and is easy to estimate.
Token pricing is usually quoted per million tokens. For example, $0.50 per million input tokens and $1.50 per million output tokens. Your bill is the sum of input tokens times input rate plus output tokens times output rate.
When requests matter
Some APIs charge per request for specialized tasks like image generation, speech-to-text, or embeddings. Others may have a per-request fee in addition to token fees, especially for high-reliability or low-latency endpoints. Always check the pricing page for the specific model or endpoint.
If you are comparing providers, convert everything to a common unit, such as cost per 1,000 tokens or cost per typical request, to make a fair comparison.
- Text LLMs: usually per token
- Image models: often per image or per request
- Embeddings: per token or per request
- Speech: per minute or per request
- Some APIs: hybrid token + request fee
Common mistakes
- Assuming all LLM APIs charge per request; most text models charge per token.
- Ignoring per-request fees on top of token costs for some endpoints.
- Comparing providers without normalizing to a common unit like cost per 1,000 tokens.
