Review from autoregressive-loop
Your MCP tool calls a model that takes about 25 ms per decode pass. It returns a 400-token answer. Ignoring the first pass and the network, roughly how long does generating the answer take?
One forward pass per output token: 400 passes x 25 ms is about 10 s. Output length sets the cost.
A forward pass produces one next-token distribution, not a reply. Streaming shows you the tokens as each pass finishes.
The GPU is parallel inside a pass, across the input positions. Across output tokens it can't be, because each pass needs the token chosen by the previous one.
After pass 1 finishes, why can't the server start pass 3 before pass 2 is done?
Each step appends the chosen token and feeds the longer list back in. Until
sample()picks token 2, there is no input for pass 3. This lesson opens up thatsample()call.Weights are read-only during inference, and many requests read them at once. The blocker is data: the next input isn't known yet.
Caching is a speedup that comes later in the path. Even without any cache, the sequential dependency between tokens would remain.
Last lesson's loop had a line you skipped past: const next = sample(probs). This lesson opens that function up.
The model's last layer doesn't output probabilities. It outputs logits: one raw, unbounded score per token in the vocabulary (tens to hundreds of thousands of numbers). Higher means "more likely next". Turning that array into exactly one token id is decoding, and it's the part of the per-token computation you control at request time. Every temperature, top_k and top_p knob in an API is a parameter to this one small function.
Here is the whole thing in plain TypeScript. Run it with node sampler.ts.
// logits: one raw score per vocabulary token, straight out of the model
function softmax(logits: number[], temperature: number): number[] {
const scaled = logits.map((l) => l / temperature); // the only line temperature touches
const max = Math.max(...scaled); // subtract max so exp() can't overflow
const exps = scaled.map((x) => Math.exp(x - max));
const sum = exps.reduce((a, b) => a + b, 0);
return exps.map((e) => e / sum); // now a probability distribution
}
function greedy(logits: number[]): number {
return logits.indexOf(Math.max(...logits)); // argmax: no randomness at all
}
function sample(probs: number[], rand: () => number): number {
let r = rand(); // uniform in [0, 1)
for (let id = 0; id < probs.length; id++) {
r -= probs[id]; // walk the cumulative distribution
if (r <= 0) return id;
}
return probs.length - 1; // float rounding safety net
}
const tokens = ["users", "data", "items", "result", "banana"];
const logits = [3.0, 2.0, 1.5, 1.0, -2.0];
for (const t of [1, 0.5, 0.2]) {
console.log(t, softmax(logits, t).map((p) => p.toFixed(3)).join(" "));
}
console.log("greedy:", tokens[greedy(logits)]);
console.log("sampled:", tokens[sample(softmax(logits, 1), Math.random)]);Three ideas are in that file.
- Greedy takes the argmax. Same logits in, same token out, every time. It is the default decoding strategy in Hugging Face Transformers, and it works well for short outputs where creativity isn't the goal, but it starts repeating itself on long outputs 4.
- Sampling treats the probabilities as a weighted die and rolls it. Any token with non-zero probability can come out, which reduces repetition and gives more diverse text 5.
- Softmax turns logits into probabilities. It exponentiates each score and divides by the total, so the results are positive and add up to 1.
Think of rand as the only source of randomness in the whole function. Swap Math.random for a seeded generator and sampling becomes reproducible on your machine. Hold that thought for later.
## Temperature: one division
Temperature divides every logit before softmax. Dividing by a number below 1 stretches the gaps between scores, so the top token pulls further ahead. Dividing by a number above 1 shrinks the gaps, and the distribution flattens. Lowering the temperature makes the distribution sharper: likely tokens get likelier and unlikely ones get less likely 1.
Output of the file above for the logits [3, 2, 1.5, 1, -2]:
T = 1: users 0.577, data 0.212, items 0.129, result 0.078, banana 0.004T = 0.5: users 0.831, data 0.112, items 0.041, ...T = 0.2: users 0.993, data 0.007, ...
As T approaches 0, all the mass piles onto the argmax, so sampling at a very low temperature becomes greedy decoding. You can't literally divide by 0, so most hosted APIs treat temperature: 0 as "just take the argmax" (greedy). In Hugging Face Transformers you don't pass 0 at all: you set do_sample=False instead.
At T = 1 the junk token banana has probability 0.004. A reply is 300 tokens long, and suppose each step has a similar long tail of junk tokens adding up to about 0.4%. Will a 300-token reply at T = 1 with plain sampling (no filtering) usually contain at least one junk pick?
Usually yes. The chance of avoiding junk at every step is roughly 0.996^300, which is about 0.30. So about 70% of replies contain at least one weird token. Each one is unlikely on its own, but with 300 rolls one of them usually comes up. That is the problem top-k and top-p solve: they remove the tail before the die is rolled.
## Top-k and top-p: cut the tail, then roll
Both filters run after temperature and before the roll. They throw out unlikely tokens and rescale the survivors so they add up to 1 again.
- Top-k keeps the
kmost likely tokens and redistributes the probability among just those 2. Its weakness is thatkis fixed. When the model is confident,k = 50still lets 49 bad options in. When the model is genuinely unsure,k = 5might cut good ones. - Top-p (nucleus sampling) keeps the smallest set of top tokens whose combined probability exceeds
p3. The set grows when the model is unsure and shrinks when it's confident. That adaptivity is why top-p is the more commonly tuned filter. Some APIs (OpenAI, for example) expose onlytop_p, while others expose both.
Temperature and top-p interact. Lower the temperature and the top tokens hold more of the mass, so fewer of them are needed to reach p. The worked example shows this with numbers.
Pick settings for inline code completion vs a tagline generator
You're building two features on the same model. A: ghost-text code completion in an editor, which finishes the current line. B: a "give me 5 tagline ideas" button. Use the logits [3, 2, 1.5, 1, -2] for [users, data, items, result, banana] to see what each setting does, then choose.
Ask what a good answer looks like. For A there is usually one right continuation. A variable called
usersis either correct or a bug, and the user wants the same suggestion every time they pause on that line. For B the user explicitly wants different answers each time they click. That points to low or zero randomness for A and real randomness for B.Feature A: greedy. Set
temperature: 0, which means argmax, so you getusersevery time. Greedy's known weakness is repetition on long outputs 4, and a one-line completion is short. If you show several alternatives, use a low temperature such as 0.2 instead:usersstays at 0.993, so suggestions rarely change.typescriptconst completionSettings = { temperature: 0, maxTokens: 64 };Feature B: try
T = 1, top_p = 0.9. The sorted probabilities atT = 1add up as 0.577, then 0.789, then 0.918, which passes 0.9. So three tokens survive (users, data, items). Rescaled: 0.577/0.918 = 0.628, 0.212/0.918 = 0.231, 0.129/0.918 = 0.141.bananaandresultcan never come out, but there is still real variety.Check the interaction. Lower B to
T = 0.5while keepingtop_p = 0.9. Now it's 0.831, then 0.943, so only two tokens survive anduserswins about 88% of the time (0.831/0.943). That's too repetitive for five ideas. Raise it toT = 1.5instead and the cumulative sum reaches 0.9 only at the fourth token (0.459, 0.694, 0.863, 0.984), soresultgets in too. The higher the temperature, the more tokens top-p lets through.Ship it. For B, start around
T = 0.9, top_p = 0.95and tune by reading real outputs: if it rambles, lower T; if the 5 ideas feel the same, raise it. Adjust one knob at a time, because each one changes what the other does.typescriptconst taglineSettings = { temperature: 0.9, topP: 0.95, n: 5 };
Ask whether the task has one right answer (greedy or low T) or many good answers (T around 0.7 to 1 plus top-p to cut the tail). Temperature controls how flat the distribution is, and top-p controls how much of the tail gets dropped. Since temperature also changes how many tokens top-p keeps, tune them together.
## Temperature 0 on a server: why it still varies
Greedy has no randomness, so temperature: 0 should give identical output every time. Run the same request against a real API a few times, though, and you'll sometimes get different answers. Setting temperature to 0 makes sampling deterministic in theory, but LLM APIs are still not deterministic in practice 6.
Greedy itself is deterministic. The logits it's given are not quite the same from run to run. The root cause is floating-point arithmetic: adding numbers in a different order can give a different result 8. Try it in Node:
const a = 0.1, b = 1e20, c = -1e20;
console.log((a + b) + c); // 0
console.log(a + (b + c)); // 0.1
// Two logits that are nearly tied, computed by summing the same terms in two orders
const terms = [0.1, 0.2, 0.3];
const left = (terms[0] + terms[1]) + terms[2]; // 0.6000000000000001
const right = terms[0] + (terms[1] + terms[2]); // 0.6
console.log(left === right, left > right); // false trueA forward pass is billions of these additions. GPU kernels choose how to split and order the sums depending on the shape of the work, and on a server that shape includes how many other users' requests are batched with yours. Your request's numbers can change with the batch size, and the batch size depends on how busy the server is. That load-dependent batch size is the main reason nearly all inference endpoints are nondeterministic 7.
Usually the differences are in the last few bits and the argmax doesn't change. But when the top two logits are almost tied (say data and items differ by 1e-7), a tiny difference can swap the winner. After that, the loop feeds a different token back in and the rest of the reply diverges.
In practice: don't write tests or caches that assume a served model gives byte-identical output for the same prompt at temperature: 0. Compare on structure or meaning instead, or store the output you got. Self-hosting doesn't fix this by itself either: even on your own hardware with an open-source engine like vLLM or SGLang, sampling still isn't deterministic 9, because those servers batch requests internally too. True determinism needs batch-invariant kernels, which is what the Thinking Machines post builds.
You're adding an "autofix this lint error" feature that returns a small patch. Which settings fit best?
There's one correct fix and the output is short, which is exactly where greedy works well. Its weakness, repeating itself on long outputs, doesn't come up here.
A flatter distribution gives you variety you don't want. Randomly picking a less likely token in a patch usually means a bug.
Unfiltered sampling keeps the whole tail, so over a patch of a few dozen tokens some junk tokens are likely to slip in.
You pay for five generations, then add your own dice roll. Low temperature still lets an unlikely token through sometimes, so the same lint error can get a different patch per click. One right answer calls for greedy.
With top_p = 0.9, you lower temperature from 1.0 to 0.5. What happens to the number of tokens that survive the top-p filter?
Lower T sharpens the distribution, so the top few tokens reach 0.9 sooner. In the worked example it dropped from 3 tokens to 2.
Top-p cuts by cumulative probability, and temperature changes those probabilities before the cut. The two settings interact.
That's what raising T does: a flatter distribution needs more tokens to reach the same 0.9.
Your snapshot test calls a hosted model at temperature 0 and compares the output to a stored string. It fails about 1 run in 20 with a slightly different reply. What's the most likely cause?
Greedy is deterministic for given logits. On a shared server the logits wobble with batch size, which depends on load, and one flipped token changes everything after it.
This is the misconception. At T=0 the pick is the argmax, with no dice roll. The variation comes from the numbers going into the argmax.
A model version change would cause much bigger and more consistent differences. Rare, small differences on a fixed model point to numerical nondeterminism.
At T=0 only the single top token is considered, so top-p has nothing to filter. It can't add variation.
Extend the sampler into sample(logits, settings, rand) where settings is { temperature: number; topK?: number; topP?: number }:
temperature === 0returns the argmax (don't divide by zero).- Otherwise: softmax at that temperature, sort by probability, apply
topKthentopP(keep the smallest set whose total exceedstopP), renormalise the survivors, and roll the die. - Use this seeded generator instead of
Math.random, so your runs are reproducible:
function mulberry32(seed: number) { return () => { seed = (seed + 0x6d2b79f5) | 0; let t = Math.imul(seed ^ (seed >>> 15), 1 | seed); t = (t + Math.imul(t ^ (t >>> 7), 61 | t)) ^ t; return ((t ^ (t >>> 14)) >>> 0) / 4294967296; }; }
Use the tokens ["=>", "{", "(", "async", "//"] with logits [2.5, 2.2, 1.0, 0.3, -0.5]. Before running it, predict how many tokens survive topP: 0.9 at temperature: 1 and at temperature: 0.7. Then sample 1000 times at { temperature: 1, topP: 0.9 } with seed 42 and check that the counts roughly match your rescaled probabilities.
- Hint 1Write a
candidates(logits, settings)helper that returns{ id, p }[]after filtering and renormalising. Thensampleis just the cumulative walk from the lesson over that list. - Hint 2For top-p, walk down the sorted list adding probabilities, and stop right after the token that takes the total past
p. Keep that token. - Hint 3At T=1 the probabilities are about 0.471, 0.349, 0.105, 0.052, 0.023. At T=0.7 they're about 0.548, 0.357, 0.064, ...
At T=1, the running total is 0.471, then 0.819, then 0.924, so 3 tokens survive (=>, {, (), rescaled to 0.509, 0.377, 0.114. At T=0.7 it's 0.548, then 0.905, so only 2 tokens survive (=> 0.606, { 0.394). With seed 42, 1000 samples come out around 493 / 393 / 114, which is close to the rescaled probabilities. async and // never appear. At T=0 you always get =>.
function mulberry32(seed: number) {
return () => {
seed = (seed + 0x6d2b79f5) | 0;
let t = Math.imul(seed ^ (seed >>> 15), 1 | seed);
t = (t + Math.imul(t ^ (t >>> 7), 61 | t)) ^ t;
return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
};
}
type Settings = { temperature: number; topK?: number; topP?: number };
function softmax(logits: number[], temperature: number): number[] {
const scaled = logits.map((l) => l / temperature);
const max = Math.max(...scaled);
const exps = scaled.map((x) => Math.exp(x - max));
const sum = exps.reduce((a, b) => a + b, 0);
return exps.map((e) => e / sum);
}
function candidates(logits: number[], s: Settings): { id: number; p: number }[] {
if (s.temperature === 0) {
const best = logits.indexOf(Math.max(...logits));
return [{ id: best, p: 1 }];
}
let c = softmax(logits, s.temperature)
.map((p, id) => ({ id, p }))
.sort((a, b) => b.p - a.p);
if (s.topK !== undefined) c = c.slice(0, s.topK);
if (s.topP !== undefined) {
const kept: { id: number; p: number }[] = [];
let cum = 0;
for (const x of c) {
kept.push(x);
cum += x.p;
if (cum > s.topP) break;
}
c = kept;
}
const total = c.reduce((a, x) => a + x.p, 0);
return c.map((x) => ({ id: x.id, p: x.p / total }));
}
function sample(logits: number[], s: Settings, rand: () => number): number {
const c = candidates(logits, s);
let r = rand();
for (const x of c) {
r -= x.p;
if (r <= 0) return x.id;
}
return c[c.length - 1].id;
}
const tokens = ["=>", "{", "(", "async", "//"];
const logits = [2.5, 2.2, 1.0, 0.3, -0.5];
for (const s of [
{ temperature: 1, topP: 0.9 },
{ temperature: 0.7, topP: 0.9 },
{ temperature: 0 },
] as Settings[]) {
const c = candidates(logits, s);
console.log(JSON.stringify(s), c.map((x) => `${tokens[x.id]} ${x.p.toFixed(3)}`).join(", "));
}
const rand = mulberry32(42);
const counts: Record<string, number> = {};
for (let i = 0; i < 1000; i++) {
const t = tokens[sample(logits, { temperature: 1, topP: 0.9 }, rand)];
counts[t] = (counts[t] ?? 0) + 1;
}
console.log(counts);