Skip to content

Tokens, context windows, and why your bill is what it is

The unit of billing for language models is not the request or the word. Once you understand tokens and how a context window fills up, the invoice stops being a surprise.

4 min read
Contents

A client of mine put a chat feature behind a login, got 200 users, and received a bill roughly eleven times what the estimate said. Nothing was broken. Nobody was abusing it. The feature simply cost what it cost, and the estimate had been built on the wrong unit.

The unit is the token. If you are going to ship anything on top of a language model, this is the hour of reading that pays for itself.

What a token is

A token is a chunk of text, usually a common word or a piece of one. Models do not see letters and they do not see words. They see token ids.

Some rough shapes for English:

  • hello — one token
  • unbelievable — three or so (un, believ, able)
  • azeemhassni.com — five or six, because it is unusual
  • A space before a word is normally part of that word’s token

The rule of thumb everyone quotes is one token ≈ four characters ≈ 0.75 words for English. It is close enough for budgeting. It is badly wrong for other things: JSON is token-hungry because of all the punctuation, code is worse, and languages other than English can cost two to three times more per word because the vocabulary was fitted mostly to English text. If your product serves Urdu or Arabic users, measure it — do not assume the English ratio.

The context window is a desk, not a memory

The context window is the maximum number of tokens the model can look at in one go — your system prompt, the conversation so far, whatever documents you pasted in, and the answer it is about to write, all counted together.

The mental model that gets people into trouble is “memory”. The model does not remember the conversation. It re-reads the entire thing, from the top, on every single turn.

That is the whole explanation for the bill.

A ten-turn conversation does not cost ten messages. It costs the sum of every prefix:

turn 1:  system (400) + user (50)                       =    450 tokens in
turn 2:  system (400) + turn 1 (450) + reply (300) + 50 =  1,200 tokens in
turn 3:  ...                                            =  2,300 tokens in
...
turn 10:                                                = 18,000 tokens in

The cost of a conversation grows with the square of its length. Ten turns is not ten times one turn — on the numbers above it is closer to sixty.

Input and output are not priced the same

Every provider charges separately for tokens you send and tokens the model writes, and output is typically several times more expensive per token. This flips a lot of intuitions:

  • Pasting a 3,000-token document in is cheap. Asking for a 3,000-token document back is not.
  • “Be concise” in your prompt is a cost control, not just a style note.
  • A max_tokens limit is a spending cap. Set one. Always.

The four levers that actually move the number

I have now done this optimisation on three products. These are the levers, in the order I reach for them.

1. Cache the prefix. Every major provider now supports caching the unchanging front of your prompt — the system message, the tool definitions, the style guide — at a large discount on subsequent calls. If your system prompt is 2,000 tokens and you send it on every request, this is the single biggest win available and it usually takes an afternoon.

2. Stop sending the whole conversation. Keep the last few turns verbatim, and replace the older ones with a short running summary you regenerate occasionally. The user notices nothing. The curve goes from quadratic back to roughly linear.

3. Use a smaller model for the boring calls. Classification, routing, extracting a date from a sentence, deciding whether a message needs the expensive model at all — a small fast model does these at a fraction of the price, and you will not be able to tell the difference in the output. Reserve the large model for work that actually needs judgement.

4. Retrieve less, better. Most retrieval pipelines paste in the top ten chunks because ten felt like a nice number. Try three. Measure whether the answers got worse. In my experience they usually do not, and you just cut your input tokens by 70%.

Latency has the same shape

Two numbers matter and they behave differently. Time to first token is mostly a function of how much you sent, because the model has to read all of it before it says anything. Tokens per second after that is a property of the model and the provider.

Which means: a long prompt hurts the moment the user cares about most, the wait before anything appears. Streaming the response papers over slow generation beautifully, but it cannot hide a slow start. If your feature feels sluggish, look at input size before you look at the model.

What to actually do on Monday

Log two numbers on every call: input tokens and output tokens. Most SDKs hand them back in the response and most teams throw them away.

const res = await client.messages.create({ /* … */ });
await metrics.record({
  feature: 'support-reply',
  model: res.model,
  inputTokens: res.usage.input_tokens,
  outputTokens: res.usage.output_tokens,
  userId: user.id,
});

A week of that data tells you which feature is expensive, which user is an outlier, and whether your cache is working. Without it you are guessing, and the invoice arrives monthly.

One more thing: set a hard per-user daily cap before you launch, not after. The first time someone finds a way to loop your feature, you want the ceiling to already be there.

Share