What is a token and how does it affect pricing?
Token basics
Tokens are the units of text that models read and generate. They can be whole words, parts of words, or punctuation. For English, 1 token ≈ 4 characters or 0.75 words. For other languages, token counts can be higher.
Pricing is usually quoted per 1,000 tokens (e.g., $0.01 per 1K tokens) or per 1 million tokens. Output tokens often cost more than input tokens because generation is more compute-intensive.
- Input tokens: the prompt you send.
- Output tokens: the model's response.
- Some providers charge differently for cached input tokens.
- Token counts vary by model and tokenizer.
Why it matters
Your cost is directly proportional to the number of tokens processed. Long prompts, verbose outputs, or multiple retries can quickly increase costs. Understanding tokenization helps you optimize prompts and set budgets.
For example, if a model costs $0.50 per 1M input tokens and $1.50 per 1M output tokens, a request with 1,000 input and 500 output tokens costs $0.0005 + $0.00075 = $0.00125.
Common mistakes
- Assuming a token equals a word; it's often less.
- Forgetting that output tokens usually cost more than input tokens.
- Not accounting for tokens in system prompts or few-shot examples.
