AI Concepts & Fundamentals

Token (AI)

The small chunks of text an AI model actually reads, counts, and gets billed for.

Quick Answer

A token is the basic unit of text an AI model processes — typically a word, part of a word, or punctuation mark, averaging about four characters in English. Tokens set two practical limits: how much text fits in a model's context window, and how much a request costs, since usage is billed per token.

The Full Picture

Models don't read raw text the way people do; they need text converted into discrete units their neural network can map to numbers. Tokens are that translation layer — an intermediate vocabulary sized so common words get one token while rare or compound words split into pieces.

A tokenizer breaks input text into these units — "construction" might be one token, "prefabrication" might split into two or three. The model has a fixed vocabulary of possible tokens, commonly tens of thousands, and processes a sequence of them. The model's context window — how much it can hold at once — is measured in tokens, not words or pages, and both the prompt and the generated response count against it.

Cost and capability trade-offs follow directly from token counts. A 50-page specification document might run 30,000–40,000 tokens once converted; a model with a 128,000-token context window can hold that alongside a lengthy prompt and still generate a response, while a smaller-window model would need the document chunked or summarized first.

Being token-aware matters operationally: teams minimize wasted tokens — redundant instructions, unnecessarily verbose documents — both to control cost and stay under context limits, and choose models with context windows sized to their actual documents rather than over- or under-provisioning.

Real Examples

→Long document chunking: A 200-page project manual exceeds a model's context window in tokens, so the document is split into overlapping chunks before an AI system processes it.
→Cost estimate: An API call that sends a 2,000-token prompt and receives a 500-token response is billed for roughly 2,500 tokens combined, at the provider's per-token rate.
→Context window limit: A chatbot summarizing a full RFI log hits its token limit mid-document and truncates the oldest messages first unless the conversation is managed deliberately.

Common Misconceptions

People assume: A token is the same as a word.

Actually: Tokens often split words into pieces, especially uncommon or technical terms — a word like "geosynthetic" might be two or three tokens. As a rough rule of thumb, one token is about four characters or three-quarters of a word in English, not a full word.

People assume: A bigger context window means the model "understands" more.

Actually: A larger window lets more tokens fit in a single pass, but models can still lose track of details buried in the middle of a very long input — fitting text isn't the same as reasoning well over all of it.

Frequently Asked Questions

What is a token in AI, exactly?

A token is the smallest unit of text a language model reads or generates — often a word, a piece of a word, or a punctuation mark. Text is converted into a sequence of tokens before a model can process it, and the model's output is also produced one token at a time.

How many tokens is a word?

Roughly 0.75 words per token in English, though it varies. Common short words are usually a single token; longer, technical, or compound words often split into two or more.

What is a context window?

A context window is the maximum number of tokens a model can process in a single interaction — including both the input prompt and the generated response. Once that limit is reached, older content typically has to be dropped, summarized, or chunked separately.

Why do AI tools charge per token?

Processing each token consumes compute, so providers meter usage by token count rather than by request or by word, giving a direct, granular measure of how much work a given input and output actually required.

Related Terms

More AI Concepts & Fundamentals Terms

Sources

  1. NIST — AI Risk Management Framework
  2. Hugging Face — Tokenizer Summary (Transformers documentation)
MELTPLAN