Why ‘Word Count’ Is the Wrong Mental Model
Every time you send a prompt to an AI chatbot, the model doesn’t see your sentence the way you do. It breaks your text into tokens — chunks that are often smaller than a full word. ‘Token’ isn’t a vague technical buzzword here; it’s the literal unit the model was trained on and the literal unit you’re billed for. As a rough rule of thumb, 1 token is about 4 characters of English text, or roughly 0.75 words, meaning 100 words of input is closer to 130-140 tokens. Short, common words like ‘the’ or ‘and’ are usually a single token. Longer or unusual words — brand names, typos, technical jargon, non-English text — often get split into two, three, or more token pieces.
Why This Matters for Your Prompts
Tokenization explains some genuinely strange AI behavior. Ask a model to count the letters in an uncommon word and it sometimes gets it wrong, because it never actually ‘saw’ individual letters — it saw a token that represents a chunk of that word as one unit. Pricing, context limits, and response length are all measured in tokens, not words or characters, which is why a request with a lot of code, JSON, or a foreign language can burn through a token budget far faster than the same amount of plain English prose would suggest.
The Cost Angle Most People Miss
Nearly every paid AI API charges separately for input tokens (what you send) and output tokens (what the model generates), with output tokens typically costing 3-5x more than input tokens. This means a prompt that asks for a long, detailed answer costs meaningfully more than one that asks for a tight, bulleted response — even if your question itself was identical. If you’re building anything on top of an AI API, the cheapest optimization available to you usually isn’t a fancier model; it’s trimming unnecessary repetition from your system prompts and asking for concise output when a concise answer actually serves the task.
Practical Moves That Come From Understanding Tokens
Paste only the relevant section of a document instead of an entire file when you can, since every extra paragraph is tokens spent whether the model needs it or not. When a chatbot seems to ‘forget’ earlier parts of a long conversation, that’s usually a token-budget limit (the context window) being hit, not a design flaw — older messages get truncated to make room. And when giving instructions, dense and specific phrasing almost always beats padded, conversational phrasing: both cost tokens, but only one gives the model more to work with per token spent. Understanding this one mechanical detail — that AI reads in token chunks, not words — explains a surprising share of the quirks people run into when using these tools daily.