What is the best way to implement rate limiting to control costs?

Updated October 2026 · How we answer

Short answerUse a layered approach: enforce per-user and per-API-key limits at the application level, set hard spending caps with your LLM provider, and monitor usage in real time. Combine token-based limits with request-based limits for best results.

Start with provider-level controls

Most LLM providers offer built-in rate limits and spending caps. Set a monthly budget cap and a per-minute request limit in your provider dashboard. These act as a safety net if your own code fails.

Provider limits are usually coarse: they apply to your whole account, not individual users. So they stop runaway costs but don't prevent one user from consuming your entire quota. Use them as a backstop, not your primary control.

Check your provider's docs for exact limits—they vary widely. Some allow custom limits; others have fixed tiers. Typical ranges: 60–1,000 requests per minute and $50–$5,000 monthly caps, depending on your plan.

Implement application-level rate limiting

Track usage per user, API key, or IP address. A common pattern is a token bucket or sliding window counter stored in a fast data store like Redis. For each request, check if the user has enough tokens left; if not, return a 429 error.

Limit by tokens, not just requests. A single request can cost 100x more than another if it includes a long prompt or generates a long response. Estimate token usage before calling the LLM and deduct from the user's budget.

Set different tiers for different user groups. Free users might get 10,000 tokens per day; paid users get more. This aligns cost control with your business model.

  • Use a sliding window or token bucket algorithm for smooth limiting.
  • Store counters in Redis or a similar in-memory store for low latency.
  • Estimate tokens with a tokenizer library before sending the request.
  • Return clear error messages with retry-after headers.
  • Log every request and its token count for auditing.

Monitor and adjust

Set up alerts for when usage hits 50%, 80%, and 100% of your budget. Without monitoring, you won't know a limit is too loose until you get a surprise bill.

Review your limits monthly. As your app grows, user behavior changes. What worked at 100 users may fail at 10,000. Adjust limits based on actual usage patterns, not guesses.

Common mistakes

  • Relying only on provider-level limits, which don't protect against one user hogging all resources.
  • Limiting only by request count and ignoring token size, letting expensive requests slip through.
  • Setting limits too high because you fear blocking legitimate users, which defeats the purpose of cost control.
From our shopsSwiftCase: Curated phone cases that ship in 48 hours.