Large Language Models (LLMs) don't read text the way you do, character by character or word by word. They break text into chunks called tokens — which might be a whole word, part of a word, or even a single punctuation mark, depending on how common that chunk is.
As a rough rule of thumb in English, 1 token ≈ ¾ of a word, so 100 words is roughly 130-140 tokens.
Why tokens matter to you
Two practical things depend directly on tokens:
- Cost and speed. Many AI tools charge (or measure limits) per token processed, not per word or per request.
- The context window. This is the maximum number of tokens a model can "see" at once — including your question, any files or history you've provided, and its own response. Once a conversation exceeds that limit, the oldest content gets dropped or summarized, which is why long conversations can sometimes seem to "forget" earlier details.
A useful analogy
Think of the context window like a desk, not a filing cabinet. The model can only "see" what's currently spread out on the desk (within the token limit). It has no separate long-term memory of your conversation unless a product specifically saves and re-feeds that information back in — which is a feature the surrounding application builds, not something the raw model does automatically.
What this looks like in practice
Paste a long report into a chat tool and ask for a summary. Everything you pasted, plus your instruction, plus every earlier message in that conversation, plus the summary as it is being written, all has to fit inside the context window at the same time.
If it fits, you get a summary drawn from the whole document. If it does not, the tool has to do something with the excess, and every option loses something: drop the earliest material, compress it into a shorter form, or refuse the request. Most products choose quietly. Nothing on screen tells you which part of your document stopped being visible.
That is why a summary of a long document can be confidently wrong about its ending, or repeat a point from early on while missing the conclusion entirely. The model summarized what it could see, and it had no way to know that the rest was ever there.
Two habits that follow from this
Start a new conversation when the subject changes. Everything earlier in the thread is still competing for the same space. A fresh conversation is not only tidier — it gives the question you actually care about more room.
Put the important instruction near your request, not far above it. A constraint you gave twenty messages ago may already have been dropped or compressed. Repeating it in the message that needs it costs a few tokens and saves a re-run.
One misunderstanding is worth naming directly: a larger context window is not memory. It is a bigger desk, not a filing cabinet. Close the conversation and the desk is cleared. Any product that seems to remember you between sessions is storing that information somewhere separate and feeding it back in at the start of the next one.
Go deeper: Token in the glossary · Context window in the glossary · AI hallucinations explained
Key takeaway: tokens are the model's unit of "reading," and the context window is the hard limit on how much it can consider at once.