How does pay-as-you-go pricing work for LLM APIs?
The basics of pay-as-you-go
With pay-as-you-go, you add a payment method and are charged for each API call based on the number of tokens processed. There's no monthly fee or minimum. Pricing is typically listed per million tokens, e.g., $0.50 per million input tokens and $1.50 per million output tokens.
Providers often have different rates for different models. A small model might cost $0.10 per million tokens, while a large model could be $10 or more. You can track usage in real time through the provider's dashboard and set spending limits.
What affects your bill
Your bill depends on how many tokens you send and receive, and which model you use. Long prompts, verbose outputs, and chatty conversations increase costs. Some providers also charge for fine-tuning or embeddings separately.
Many providers offer free tiers or credits for new users, but after that, it's purely usage-based. You can often set hard limits to avoid surprise bills.
- No upfront fees; pay only for what you use.
- Input and output tokens priced differently.
- Rates vary widely by model and provider.
- Usage dashboards show real-time spend.
- Spending limits can prevent overages.
Common mistakes
- Thinking pay-as-you-go means unlimited use for a flat fee.
- Forgetting that output tokens often cost more than input tokens.
- Not setting a spending limit and getting an unexpectedly high bill.
