AI Tokens Defined
AI tokens are the small units of text, such as words, word fragments, or punctuation, that a large language model reads and writes. Models count input and output in tokens, so tokens determine how much text fits in a context window, how fast a response arrives, and how much each AI interaction costs.
Every AI answer has a unit cost, and that unit is the token. When a support team asks why an AI agent can only read so much of a long email thread, why one model responds faster than another, or why AI pricing varies by provider, the answer usually comes back to tokens.
What are AI tokens?
AI tokens are the pieces of text a large language model processes. Before a model reads anything, a tokenizer splits the text into tokens: common words often become a single token, while rare words, names, and code are split into several fragments. Punctuation and spaces count too.
As a rule of thumb for English, one token is roughly four characters, or about three quarters of a word. A 100-word customer message is around 130 tokens. Other languages often need more tokens for the same meaning, because tokenizers are usually trained mostly on English text.
Tokens come in two kinds:
- Input tokens: everything the model reads, including the customer's message, conversation history, retrieved help articles, and system instructions
- Output tokens: everything the model writes back
Why AI tokens matter
Tokens set three practical limits on any AI system:
- Capacity: a model's context window is measured in tokens. A long ticket history, a large policy document, and the instructions for the AI all compete for the same space.
- Speed: models generate output one token at a time, so longer answers take longer to arrive. This matters most on voice calls, where a pause of a second or two feels broken.
- Cost: most model providers bill per million input and output tokens, with output tokens usually priced several times higher than input tokens.
For customer service leaders, token counts are an input to cost, not the cost itself. A useful comparison is what each resolved conversation costs, which is why teams track fully loaded cost per conversation rather than raw token spend.
How AI tokens work in a support conversation
Here is what happens to tokens when a customer asks an AI agent, "Can I get a refund on an order I placed three weeks ago?"
- Tokenize the question: the message is split into about 15 tokens.
- Add context: the system adds instructions, recent conversation history, and the customer's order details. The input grows to a few thousand tokens.
- Retrieve knowledge: retrieval-augmented generation finds the refund policy and adds the relevant passages, not the whole help center, to keep the input focused.
- Generate the answer: the model writes a reply token by token, perhaps 80 to 150 tokens for a clear, complete answer.
- Check and send: many systems run extra model calls to validate the answer before it reaches the customer, and each of those calls uses tokens as well.
The design choice that matters is what goes into the input. Stuffing every document into the prompt wastes tokens and can make answers worse, because the model has more irrelevant text to sort through. Good retrieval sends fewer, better tokens.
AI tokens vs. embeddings
| Tokens | Embeddings | |
|---|---|---|
| What it is | A unit of text | A list of numbers representing meaning |
| Used for | Reading and generating language | Searching and comparing meaning |
| Measured in | Count of tokens | Vector dimensions |
Tokens are how a model reads text. Embeddings are how a system represents the meaning of that text so it can find similar content, such as the help article that best matches a question.
How Fin handles tokens
Fin customers do not buy or manage tokens. Fin runs a multi-step pipeline for each answer, including retrieval, reranking, generation, and validation, and prices on outcomes rather than on the number of tokens those steps consume. Teams can focus on whether a conversation was resolved, not how many tokens it took.
Frequently asked questions
How many tokens is a word?
In English, one word averages about 1.3 tokens. Short common words are usually one token; long or unusual words can be three or more.
Do images and audio use tokens?
Yes. Multimodal models convert images and audio into tokens too. A single image can use hundreds or thousands of tokens depending on its size and the model.
Are tokens the same across AI models?
No. Each model family uses its own tokenizer, so the same sentence can produce a different token count on different models. Compare prices per task, not per token.