Why do long prompts cost more?
How token-based pricing works
Most LLM APIs charge separately for input (prompt) tokens and output (completion) tokens. Input tokens are usually cheaper per token than output tokens, but they still add up. For example, if an API charges $0.50 per million input tokens, a 10,000-token prompt costs about half a cent before the model writes anything.
The model reads the entire prompt to generate a response, so every extra word, example, or instruction increases the input token count. Even if you only need a one-word answer, the model still processes the whole prompt.
Why long prompts add up fast
Long prompts often include system instructions, conversation history, retrieved documents, or few-shot examples. In a chat app, the entire history is resent with each new message, so costs grow as the conversation continues. A 5,000-token history sent 20 times means 100,000 input tokens just for context.
Some providers offer prompt caching or batch discounts, which can reduce the cost of repeated long prompts. But without caching, you pay full price for every token every time.
- System prompts and instructions count as input tokens.
- Conversation history is resent with each turn, multiplying cost.
- Retrieved documents or search results add many tokens.
- Few-shot examples can be hundreds or thousands of tokens.
- Output tokens are usually priced higher than input tokens.
Common mistakes
- Thinking you only pay for the output; input tokens are also billed.
- Assuming a long prompt is free if the model ignores part of it; all tokens are processed.
- Forgetting that chat history is resent and billed again on every turn.
