How is LLM API pricing calculated?

Updated October 2026 · How we answer

Short answerMost LLM APIs charge per token, with separate rates for input (prompt) and output (completion) tokens. Some providers also charge per request, per image, or per fine-tuning job.

The token-based model

The dominant pricing model for text LLMs is per token. A token is a chunk of text, roughly 4 characters or 0.75 words in English. When you send a prompt, the provider counts the input tokens; when the model replies, it counts the output tokens. You pay a rate for each, usually quoted per million tokens.

For example, a model might cost $0.50 per million input tokens and $1.50 per million output tokens. If your request uses 1,000 input tokens and generates 500 output tokens, the cost is $0.0005 + $0.00075 = $0.00125. Prices vary widely by model and provider, so always check the current pricing page.

Other pricing dimensions

Some APIs charge per request on top of tokens, especially for specialized endpoints like image generation or embeddings. Others use flat monthly subscriptions with usage caps. Fine-tuning often has its own training and inference rates. Always read the provider's pricing details because the model can be mixed.

  • Per-token: input and output rates
  • Per-request: fixed fee per API call
  • Per-image or per-minute: for multimodal or audio
  • Per-fine-tuning-job: training plus inference
  • Subscription tiers: monthly fee with included usage

Common mistakes

  • Assuming all LLM APIs charge only per token; some add per-request or per-image fees.
  • Forgetting that output tokens often cost more than input tokens, which can double your effective rate.
  • Ignoring that token counts include system prompts and chat history, not just the user's message.
From our shopsCaseMorph: Type an idea, see a custom phone case in seconds, then print a one-of-one.