How is LLM API pricing calculated?
The token-based model
The dominant pricing model for text LLMs is per token. A token is a chunk of text, roughly 4 characters or 0.75 words in English. When you send a prompt, the provider counts the input tokens; when the model replies, it counts the output tokens. You pay a rate for each, usually quoted per million tokens.
For example, a model might cost $0.50 per million input tokens and $1.50 per million output tokens. If your request uses 1,000 input tokens and generates 500 output tokens, the cost is $0.0005 + $0.00075 = $0.00125. Prices vary widely by model and provider, so always check the current pricing page.
Other pricing dimensions
Some APIs charge per request on top of tokens, especially for specialized endpoints like image generation or embeddings. Others use flat monthly subscriptions with usage caps. Fine-tuning often has its own training and inference rates. Always read the provider's pricing details because the model can be mixed.
- Per-token: input and output rates
- Per-request: fixed fee per API call
- Per-image or per-minute: for multimodal or audio
- Per-fine-tuning-job: training plus inference
- Subscription tiers: monthly fee with included usage
Common mistakes
- Assuming all LLM APIs charge only per token; some add per-request or per-image fees.
- Forgetting that output tokens often cost more than input tokens, which can double your effective rate.
- Ignoring that token counts include system prompts and chat history, not just the user's message.
