How can I estimate token usage before making API calls?
Use the right tokenizer
Each model family uses a specific tokenizer. For OpenAI models, the tiktoken library is common. For open-source models like Llama or Mistral, use the tokenizer from the Hugging Face transformers library. These tools give exact token counts for a given text.
Token counts vary by language and content. English text often averages around 4 characters per token, but code, JSON, and non-English languages can use more tokens per character. Always test with representative samples from your own data.
Estimate the full request
Count tokens in your system prompt, user message, and any few-shot examples. Then estimate output tokens based on the maximum length you allow or the typical response length you've seen. Add a buffer for safety.
If you're building a cost model, multiply input tokens by the input price and output tokens by the output price. Remember that some providers count special tokens or formatting overhead. Check the provider's documentation for exact rules.
- Use tiktoken for OpenAI models.
- Use Hugging Face tokenizers for open-source models.
- Test with your own text, not just generic examples.
- Include system prompts and examples in your count.
- Add 10–20% buffer for unexpected output length.
Common mistakes
- Assuming one token equals one word; it's often closer to 0.75 words in English.
- Forgetting that non-English text and code can use many more tokens per character.
- Ignoring output tokens in cost estimates, even though they often cost more.
