Everything else in this path — KV caches, batching, PagedAttention — exists to make one loop faster. So start with the loop.
An LLM is, from the outside, a function. You give it a list of tokens (integer ids for word pieces; English averages a bit under one word per token) and it returns a probability for every token in its vocabulary being the next one. One call of that function is a forward pass: the input runs through every layer of the network once.
That is all a single pass gives you: one next-token distribution. Not a sentence. Not a paragraph.
async function generate(prompt: number[], maxNew: number) {
const tokens = [...prompt];
for (let i = 0; i < maxNew; i++) {
const probs = await model.forward(tokens); // one forward pass
const next = sample(probs); // pick one token id
if (next === EOS) break; // model said "done"
tokens.push(next); // feed it back in
}
return tokens.slice(prompt.length);
}Read the loop as: predict one token, append it, call the model again on the longer list, and stop at a special end-of-sequence token or a length limit. That is exactly how the Hugging Face team describes inference: the model processes the input, predicts the next token, appends it, and repeats until a stopping criterion is hit 1.
The word for this is autoregressive: each output becomes part of the input for the next step.
Using the generate function above, the prompt is 1,000 tokens and the model emits 200 tokens and then EOS. How many times is model.forward called?
201. The first call reads all 1,000 prompt tokens and yields reply token 1. Calls 2–200 each yield one more reply token. Call 201 yields EOS, which isn't kept. So the count tracks the output length, not the prompt length. In practice people say "≈ one pass per output token" and don't fuss about the ±1.
Notice the 1,000-token prompt cost only one pass. The whole prompt is known up front, so the network processes all of its positions at once inside that single pass. Only tokens you don't know yet force extra passes.
## Why can't we parallelise the output?
GPUs love parallel work, and transformers are internally parallel: inside one pass, every input position is processed simultaneously. So why not compute all 200 output tokens at once?
Because token 2's input includes token 1, and token 1 doesn't exist until pass 1 finishes and sample() picks it. Sampling involves a choice — the model returns probabilities, not a fixed answer — so there's no way to know token 1 in advance and skip ahead. The DeepMind scaling book puts it bluntly: during generation, each request's forward passes happen one token at a time because there's a sequential dependency between steps 2.
The same source points at the escape hatch: since one request can't be parallelised over its own tokens, servers get their GPU utilisation by running many requests' steps together 2. That's batching, and it's coming later in the path.
How long does a reply take?
A model's forward pass takes about 20 ms at decode time on your GPU. A user sends a 1,000-token prompt and gets a 200-token answer. Ignore network time. Roughly how long until the last token, and what dominates?
Count the passes. One pass reads the prompt and produces token 1, then one pass per remaining token. That's ≈ 200 passes.
Separate the first pass. Pass 1 handles 1,000 tokens at once, so it costs more than a 20 ms decode pass — say a few hundred ms on a big prompt. Call it ~0.3 s. This is the wait before the first token appears.
Multiply the rest. 199 passes × 20 ms ≈ 4.0 s.
Total: ≈ 0.3 s + 4.0 s ≈ 4.3 s. The 200 sequential passes dominate, even though the prompt was 5× longer than the answer.
Reply length, not prompt length, is usually what makes generation slow. When you design a feature that calls an LLM, the cheapest speedup is often asking for a shorter output (e.g. JSON with terse keys instead of prose).
A model answers a 50-token question with 500 tokens. A second request sends a 500-token question and gets a 50-token answer. Same model, same GPU. Which takes longer?
It needs ~500 sequential passes. The long prompt in the other request is handled in parallel within its first pass.
Total tokens isn't the right measure. Prompt tokens are processed together in one pass; output tokens each need their own pass.
A long prompt makes the first pass heavier, but it's still one pass. 450 extra sequential passes cost far more.
Why can't a server compute the 200 tokens of a single reply in parallel on 200 GPUs?
That sequential dependency is the core of autoregressive generation.
They're highly parallel inside a pass — all positions of a known input run at once. The limit is between passes, not within them.
One pass yields one next-token distribution, never a whole reply.
Streaming text in a chat UI appears token by token. What's actually happening?
This is the misconception this lesson targets. There's no finished answer to animate until the loop ends.
Streaming exposes the loop directly — that's why the first token can show up long before the last one exists.
Write a TypeScript generate that streams: it should be an async function* that yields each token as soon as it's produced, and also counts forward passes. Use this mock model:
``
const EOS = 0;
const model = { forward: async (t: number[]) => t.length < 8 ? t.length + 1 : EOS };
``
(The mock returns the next token directly instead of probabilities, so there's no sampling step.) Then: with a 3-token prompt [1, 2, 3] and maxNew = 20, predict how many tokens get yielded and how many passes run, before you run it.
- Hint 1Keep the
forloop from the lesson; replacereturnat the end with ayieldinside the loop. - Hint 2Increment a counter every time you call
model.forward, including the call that returnsEOS. - Hint 3The mock returns
t.length + 1: with input length 3 it returns 4, and it returnsEOSonce the list reaches length 8.
It yields 5 tokens (4, 5, 6, 7, 8) and runs 6 passes: five that produce a token, plus one that returns EOS. Same ±1 as the predict block — the count follows output length, and the 3-token prompt cost nothing extra.
const EOS = 0;
const model = { forward: async (t: number[]) => t.length < 8 ? t.length + 1 : EOS };
let passes = 0;
async function* generate(prompt: number[], maxNew: number) {
const tokens = [...prompt];
for (let i = 0; i < maxNew; i++) {
passes++;
const next = await model.forward(tokens);
if (next === EOS) return;
tokens.push(next);
yield next;
}
}
(async () => {
for await (const tok of generate([1, 2, 3], 20)) console.log(tok);
console.log('passes:', passes); // 6
})();