Skill map · llm-inference

LLM Inference

GoalBuild something with LLM inference/servingbuild mode

1 of 24 mastered1 in progress

  • Mastered 1
  • In progress 1
  • Ready 1
  • Locked 21

Up nextSampling and decoding strategies · Attention through the lens of inference

Stage 1
Stage 2
ReadyAttention through the lens of inferenceWhat queries, keys and values are, and which of them a new token actually needs from the past.
Stage 3
LockedRun a model on your MacServe an open model locally with Ollama or llama.cpp, call its OpenAI-compatible API from TypeScript, stream tokens and read the timing fields.Needs Sampling and decoding strategies
LockedThe KV cacheStoring per-layer keys and values so each decode step only computes the newest token.Needs Attention through the lens of inference
Stage 4
LockedKV cache size mathComputing KV cache bytes from layers, heads, head dim, precision, context and batch.Needs The KV cache
LockedPrefill vs decodeThe two phases of a request: parallel prompt processing, then sequential token generation.Needs The KV cache
Stage 5
LockedMQA and GQASharing key/value heads across query heads to shrink the KV cache.Needs KV cache size math
LockedTTFT, TPOT and throughputThe metrics a serving system is judged by, and how they trade off.Needs Prefill vs decode
LockedMemory-bound vs compute-boundGPU bandwidth, arithmetic intensity and the roofline: why decode is bandwidth-limited.Needs Prefill vs decode
Stage 6
LockedSpeculative decodingDrafting several tokens cheaply and verifying them in one target-model pass.Needs Memory-bound vs compute-bound, Sampling and decoding strategies
LockedFlashAttentionExact attention that's faster by avoiding round-trips to slow GPU memory.Needs Memory-bound vs compute-bound, Attention through the lens of inference
LockedBenchmark an endpoint correctlyMeasure TTFT, inter-token latency and throughput from a client under controlled concurrency, with warm-up and percentiles instead of one-off averages.Needs TTFT, TPOT and throughput, Run a model on your Mac
LockedStreaming tokens to the clientDelivering tokens over SSE as they decode, and what that means for perceived latency.Needs TTFT, TPOT and throughput
LockedQuantizationLower-precision weights and KV cache: why it speeds decode and frees memory.Needs Memory-bound vs compute-bound
LockedTensor and pipeline parallelismSplitting a model across GPUs when it doesn't fit or is too slow on one.Needs Memory-bound vs compute-bound
LockedStatic vs continuous batchingWhy batching is nearly free when memory-bound, and how continuous batching fills slots as requests finish.Needs Memory-bound vs compute-bound, TTFT, TPOT and throughput
Stage 7
LockedPagedAttentionManaging the KV cache in fixed-size blocks to stop fragmentation and fit more requests.Needs KV cache size math, Static vs continuous batching
LockedPrefill/decode disaggregationRunning prefill and decode on separate GPU pools to meet TTFT and TPOT targets.Needs Static vs continuous batching
LockedCost and capacity planningTurn throughput, KV memory and latency targets into GPUs needed, cost per million tokens, and a self-host vs API break-even.Needs Benchmark an endpoint correctly, Static vs continuous batching +1
LockedBuild a tiny scheduler simulatorSimulate static vs continuous batching with a per-step token budget and watch TTFT and throughput trade off.Needs Static vs continuous batching
Stage 8
LockedPrefix cachingReusing KV blocks across requests that share a system prompt or conversation history.Needs PagedAttention
LockedBuild a paged KV allocatorImplement fixed-size KV blocks, per-request block tables, shared prefix blocks with reference counts, and preemption when memory runs out.Needs PagedAttention
Stage 9
LockedRunning a serving enginePutting it together with vLLM or SGLang: config knobs, OpenAI-compatible API, and benchmarking.Needs Benchmark an endpoint correctly, Build a paged KV allocator +5

Made with byagent

Published with byagent