How do I choose between a cheap and an expensive LLM?
Task complexity matters most
Cheap models are fine for tasks with clear rules: sentiment analysis, keyword extraction, formatting, or answering FAQs from a fixed knowledge base. Expensive models excel at open-ended reasoning, creative writing, coding, and handling ambiguous instructions.
A good rule: if a human with average skills could do the task in a few seconds without deep thought, a cheap model can probably handle it. If it requires expertise or multiple steps, consider a more expensive model.
- Cheap: classification, extraction, simple summaries, templated replies
- Expensive: complex reasoning, long-form writing, code generation, nuanced chat
- Test both on 50–100 real examples
- Calculate cost per correct answer, not just cost per token
Consider volume and latency
At high volume, even small price differences matter. If you process millions of requests, a cheap model can save thousands of dollars per month. But if quality suffers, you may lose customers or spend on manual fixes.
Latency also varies. Some cheap models are faster; some are slower. If you need real-time responses, test speed as well as cost. Sometimes a mid-tier model offers the best balance.
Common mistakes
- Choosing a model based only on price without testing quality on your specific task.
- Using an expensive model for simple tasks out of habit or fear of poor quality.
- Ignoring latency and user experience when optimizing for cost.
