What is the best way to monitor LLM API usage and costs?
Start with your provider's dashboard
Every major LLM API provider—OpenAI, Anthropic, Google, Cohere—offers a web dashboard that shows token usage and estimated costs. These dashboards typically break down usage by model, API key, and time period, and they're the fastest way to spot trends.
Check the dashboard at least weekly if you're in active development, and daily if you're in production. Most dashboards let you export CSV data, which you can feed into a spreadsheet or BI tool for custom reports.
Add a proxy or gateway for deeper insights
If you need per-user, per-feature, or per-customer cost tracking, consider a lightweight proxy like Helicone, OpenMeter, or LiteLLM. These sit between your app and the provider, logging every request and calculating cost based on token counts and model pricing.
They also let you set budget alerts, cache responses to reduce spend, and fall back to cheaper models when a budget threshold is hit. Setup usually takes an hour or two and requires only a URL change in your code.
- Helicone: open-source, easy setup, good for per-user tracking
- OpenMeter: usage-based billing and metering
- LiteLLM: proxy with budget limits and model routing
- LangSmith: tracing with cost estimates for LangChain apps
- Custom logging: log token counts and multiply by current prices
Set up alerts and reviews
Most dashboards allow you to set spending alerts via email or Slack. Configure alerts at 50%, 80%, and 100% of your expected monthly budget. Also, review your usage weekly to catch anomalies early.
For teams, assign someone to own cost monitoring. A simple weekly email with top spenders and cost per feature can prevent surprises.
Common mistakes
- Relying only on the provider's dashboard and missing per-user or per-feature cost drivers.
- Forgetting that token counts include both input and output, and output tokens often cost more.
- Assuming all models have the same price—prices vary widely, so mixing models can skew your cost estimates.
