Token Management
- What is a token and how does it affect pricing?
A token is a piece of text (roughly 4 characters or 0.75 words in English) that LLMs process. API pricing is typically per 1,000 or 1 million tokens, with input and output tokens priced separately. - How can I estimate token usage before making API calls?
Use a tokenizer library for your model family to count tokens in your prompt and expected output. For rough estimates, assume about 4 characters per token in English, but verify with the actual tokenizer. - Why do long prompts cost more?
LLM APIs charge per token, so a longer prompt means more input tokens and a higher bill. The model must process every token you send, and you pay for that processing even if the output is short. - Can I limit the number of tokens in a response to save money?
Yes, most LLM APIs let you set a maximum output token limit (often called max_tokens or max_output_tokens). This caps the response length and prevents runaway costs, but it does not reduce input token costs. - Is there a way to count tokens for free?
Yes, you can count tokens for free using tokenizer libraries, online tools, or the API's own token counting endpoint. These methods don't require a paid API call and work for most popular models. - How do I avoid paying for unnecessary tokens?
Trim prompts, remove redundant context, use caching, and set output limits. Also choose cheaper models for simple tasks and batch requests when possible.