An AI token is the basic unit of text that a large language model (LLM) reads and writes: a whole word, part of a word, a punctuation mark or a few characters, depending on the model. Models break input into tokens, process them and generate their reply token by token. Because tokens measure how much work a model does, most AI APIs price usage per token, and model limits are stated in tokens.
Not to be confused with tokenization in data security, which replaces sensitive values such as card numbers with substitute tokens.
At a glance
- A token is a chunk of text, roughly a short word or part of a longer one; the exact split depends on the model’s tokenizer.
- Both input (your prompt, documents and history) and output (the model’s reply) are counted in tokens.
- Most AI APIs bill per token, often with a different price for input and output, so costs grow with usage.
- A model’s context window, the most it can consider at once, is measured in tokens.
- Images, audio and other inputs are usually converted to tokens too, at rates that vary by provider.
What problem it solves
Models do not work on letters or whole sentences directly. Splitting text into tokens gives them a fixed vocabulary of pieces they can turn into numbers and process. For buyers, the token matters less as a technical idea and more as a unit of measure: it is how AI usage is metered, priced and limited.
Understanding tokens explains why the same task can cost very different amounts. A short question with a short answer uses few tokens. A question that pulls in long documents through retrieval-augmented generation (RAG), or an agentic AI task that loops through many steps, can use far more.
How it works
Each model has a tokenizer that splits text into tokens from its vocabulary. Common words may be one token; long or rare words, names, numbers and code may be several. As a rough guide for English, providers often say a token averages around three-quarters of a word, but this varies by model, language and content.
Usage is typically counted in two parts:
- Input tokens include the system instructions, the user’s prompt, any documents or search results added to the prompt and, in a conversation, the earlier messages that are sent again with each turn.
- Output tokens are the reply. Some models also generate internal reasoning tokens before answering, which some providers bill as output.
Providers typically price per thousand or per million tokens, with output usually costing more than input, and larger models costing more than smaller ones. Some offer discounts for cached input or batch processing. Applications built on top of models may hide tokens behind a per-user price or credits, but token use often still drives their limits.
The context window caps the total tokens a model can handle in one request. If a conversation or document exceeds it, older content is dropped or summarized.
When it matters for buyers
- When building on AI APIs. Usage-based pricing makes costs variable. Forecast from realistic token counts for your workloads and set budgets and alerts.
- When comparing models. Compare price per task, not just price per token, since tokenizers and output lengths differ.
- When rolling out agents or RAG. Long context and multi-step loops can multiply token use quickly.
- When negotiating AI contracts. Ask whether committed-spend discounts, rate limits and overage terms are stated in tokens, credits or dollars. See our artificial intelligence overview for help comparing providers.
Questions to ask vendors
- What are your prices for input, output, cached and any reasoning tokens for the models we plan to use?
- How do you count tokens for images, audio and files?
- What context window does each model support, and does pricing change for long contexts?
- What rate limits apply, and what happens when we hit them?
- What budget controls, alerts and per-team usage reporting do you provide?
- If your product is priced per user or by credits, how do those map to underlying token use and limits?
How it differs from per-user licensing
Per-user licensing charges a fixed fee for each person entitled to use software, regardless of how much they use it. Token pricing is usage-based: you pay for the volume of text the model processes, so cost depends on how much and how heavily the AI is used, not on how many people have access. Many AI products combine the two, for example a per-user subscription with token-based usage caps, or an enterprise plan with a pool of tokens or credits. Per-user pricing is easier to budget; token pricing can be cheaper for light use and more expensive for heavy or automated use.
