Can I limit the number of tokens in a response to save money?

Updated October 2026 · How we answer

Short answerYes, most LLM APIs let you set a maximum output token limit (often called max_tokens or max_output_tokens). This caps the response length and prevents runaway costs, but it does not reduce input token costs.

How to set a token limit

In the API request, you can specify a parameter like max_tokens (OpenAI, Anthropic, etc.) or max_output_tokens (Google Gemini). The model will stop generating once it reaches that limit. For example, setting max_tokens=100 means the response will be at most 100 tokens, which is roughly 75 words.

This is useful for controlling costs because output tokens are often the most expensive part of a request. However, if the model hits the limit mid-sentence, the response may be cut off, so choose a limit that fits your use case.

What a token limit does not do

A max_tokens limit only controls the output length. It does not reduce the number of input tokens you send, so a long prompt still costs the same. Also, some APIs count the limit as part of the context window, meaning a very high max_tokens can leave less room for input.

If you need to save on input costs, you must shorten the prompt itself or use caching. Combining a reasonable max_tokens with a concise prompt gives the best savings.

  • Set max_tokens to a value that matches your expected answer length.
  • Check if the limit includes reasoning tokens for models that think step-by-step.
  • Use stop sequences to end generation early when a pattern appears.
  • Monitor actual output lengths to tune the limit over time.

Common mistakes

  • Believing a max_tokens limit also reduces input token charges.
  • Setting the limit too low and getting truncated, unusable answers.
  • Ignoring that some models count internal reasoning tokens against the limit.
From our shopsSwiftCase: Curated phone cases that ship in 48 hours.