Pricing Basics

  • How is LLM API pricing calculated?
    Most LLM APIs charge per token, with separate rates for input (prompt) and output (completion) tokens. Some providers also charge per request, per image, or per fine-tuning job.
  • What is the difference between input and output token pricing?
    Input tokens are the text you send to the model; output tokens are the text it generates. Output tokens usually cost more—often 2x to 4x the input rate—because generation is more compute-intensive.
  • Why do some LLM APIs cost more than others?
    Prices reflect model size, compute cost, provider margins, and features like context length, speed, and reliability. Larger, more capable models generally cost more per token.
  • Can I get a free tier for LLM APIs?
    Yes, many providers offer free tiers with limited usage, free credits for new users, or free open-source models you can self-host. Limits vary widely and often require a credit card or phone verification.
  • Is LLM API pricing based on tokens or requests?
    Most text LLM APIs price by tokens, not requests. However, some providers charge per request for certain endpoints, and a few use a hybrid model with both token and request fees.
  • How much does it cost to run a chatbot with an LLM API?
    Costs vary widely by model and usage. For a typical chatbot, expect anywhere from a few cents to several dollars per 1,000 conversations, depending on the model and message length.