Which LLM API is cheapest for basic text generation?
Typical cheap options
For basic text generation—like short replies, summaries, or simple content—you don't need a top-tier model. Small, fast models from major providers are priced low, often under $0.50 per million input tokens. Open-source models hosted by third parties can be cheaper still, sometimes under $0.20 per million tokens.
However, 'cheapest' depends on your specific needs. Some providers charge less for input but more for output. Others offer free tiers with rate limits. Always check the current pricing page because these numbers change frequently.
- OpenAI GPT-4o mini: ~$0.15/M input, $0.60/M output
- Google Gemini 1.5 Flash: ~$0.075/M input, $0.30/M output
- Anthropic Claude 3 Haiku: ~$0.25/M input, $1.25/M output
- Together AI (Llama 3 8B): ~$0.20/M tokens
- Fireworks AI (Llama 3 8B): ~$0.20/M tokens
What to watch for
The absolute cheapest per-token price isn't always the lowest total cost. If a model requires more retries or produces lower quality, you may spend more overall. Also, some providers have minimum charges or subscription fees.
For basic text generation, start with a small model from a major provider. If costs are still too high, try an open-source model on a serverless platform. Test quality on your data before committing.
Common mistakes
- Assuming the cheapest per-token price always means the lowest total cost.
- Ignoring output token prices, which can be 3–5x higher than input prices.
- Not checking for free tiers or credits that can cover low-volume usage.
