How do I avoid paying for unnecessary tokens?
Reduce input tokens
Start by cutting anything the model doesn't need: repeated instructions, verbose examples, or old conversation history. Summarize long histories instead of resending them. If you use retrieval, only include the most relevant chunks, not entire documents.
Prompt caching can help if you reuse the same long prefix across calls. Providers like OpenAI and Anthropic offer discounts on cached input tokens, sometimes 50-90% off. This is ideal for system prompts or fixed knowledge.
Control output and model choice
Set a max_tokens limit to prevent overly long responses. Use stop sequences to end generation when a pattern appears. For simple tasks like classification or extraction, use a smaller, cheaper model. Reserve large models for complex reasoning.
Batch API options often offer 50% discounts for non-urgent jobs. Also monitor your usage with provider dashboards to spot waste.
- Remove duplicate instructions and examples.
- Summarize chat history instead of resending it.
- Use prompt caching for repeated prefixes.
- Set max_tokens and stop sequences.
- Choose smaller models for easy tasks.
- Use batch APIs for non-urgent work.
Common mistakes
- Sending the entire conversation history every time without summarizing.
- Using a top-tier model for simple tasks that a cheaper model handles well.
- Ignoring caching features that could cut input costs significantly.
