You know gradient descent. This is everything between that and the LoRA run you did in 40 seconds on your Mac — built up one honest step at a time, nothing skipped.
Picture a giant mixing console, like a sound engineer's desk with thousands of sliders. You feed in a row of numbers on the left. Every slider on the right is just some weighted blend of all the sliders on the left — turn some inputs up, some down, add them together. That blend is the entire "layer". The slider positions are the knobs, and there can be millions of them.
A neural network is just several of these mixing consoles chained one after another, the output of one feeding the input of the next.
Chain a few dozen consoles together and the whole pipeline can approximate absurdly complicated things — "is this a cat" or "what word comes next". Training is the loop you already know, just at scale: run an example through, see how wrong the answer was, then figure out which sliders to nudge and in which direction to make it less wrong.
Working out "which direction" for a slider buried in the third console, when the wrongness is only measured at the very end, is the one genuinely new piece. It's solved by working backwards: a coach reviewing game footage in reverse, starting from the missed goal, and tracing back through each player's touch to say exactly how much each one contributed and which way they should have moved. That backward pass is called backpropagation. Once every slider has its "which way, how much", that's a gradient — and you already know what to do with a gradient.
Plain mixing consoles have a blind spot: a word's meaning depends on other words far away in the sentence, but a slider only ever sees its own fixed input slot — it has no way to reach across and check a different word. You do this reaching-across constantly without noticing: reading a mystery novel and hitting "he did it", your brain instantly flips back through everyone introduced so far and weighs which one "he" most plausibly points at. Attention is a layer built to do exactly that flip-back-and-check, for every word, automatically.
The trophy didn’t fit in the suitcase because it was too big.
| other word | raw score | attention weight |
|---|---|---|
| trophy | 4.8 | |
| suitcase | 3.1 | |
| big | 2.6 | |
| because | 0.4 |
Swap “big” for “small” and every one of these weights recomputes — now suitcase wins. Same word “it”, same position, different answer, because the answer was never stored anywhere — it’s recomputed from context every time.
Every word does two things at once: it puts out a search request saying what it's looking for, and it wears a tag advertising what it's about. "It" puts out a search request that basically reads "I'm a thing that can be too big or too small, find me a noun." Every other word's tag gets checked against that request — "trophy" wears a tag like "physical object, has a size" and matches well; "because" wears a tag like "connector word" and matches badly. The match strength is the score in the table above. Once you know the scores, you also pull each word's actual content — its value, roughly "what I mean" — and blend it in proportion to how well it matched.
In the jargon: the search request is called a Query, the tag is a Key, the content is a Value. Same idea as typing into a search box (your Query) and it matching against every webpage's index terms (their Keys), then returning the actual page content (the Value) for the best matches.
Stack a run of attention layers and plain layers together, and that stack is what people call a transformer.
Qwen3-4B has roughly four billion sliders. Fine-tuning the ordinary way means figuring out, storing, and nudging a correction for every single one — far more memory than an M-series Mac has.
Think of the frozen model as a huge, expensively-printed textbook. You're not allowed to reprint it — too big, too expensive, and you'd need to redo all four billion words in it. But you're allowed to clip a thin insert of sticky notes to it. When reading, you read the original page and your sticky note together, and mentally add the correction on top. That's the whole trick: freeze the textbook (W), never touch it again, and only ever write and rewrite the sticky notes (A and B below).
output = W·input + B·(A·input)
Instead of writing a full, detailed sticky note for all 4096 outputs individually, A forces the correction to first get summarised down into just 8 core adjustments — that “8” is the rank in Low-Rank Adaptation, and it's like writing a short summary note instead of retyping the whole page. B then expands those 8 core adjustments back out to touch all 4096 outputs. Squeezing through that narrow 8-wide bottleneck is the entire reason it's cheap: you're writing roughly 4096×8 + 8×4096 numbers of sticky note instead of retyping 4096×4096 — about two orders of magnitude less to write, for this one layer.
Only the sticky notes, A and B, ever get corrected during training. The textbook, W, sits there frozen — you read from it every single time, but you never once pick up a pen and change it.
Matrix multiplication distributes over addition. That means:
The left side and the right side land on the exact same number. So the sticky-note correction behaves exactly as if it had been written straight into the textbook, even though you never touched the original pages — you only ever read the page and the note side by side and added them in your head.
Fuse takes that literally: it computes what the sticky-note correction would look like written out in full (the same shape as a whole page, 4096×4096), and writes a brand new page that already has the correction baked in. After fusing, there's no separate sticky note left to flip to — you get back one ordinary textbook, at the original textbook's normal reading speed, with the correction now permanent and invisible.
None of the above changes for audio. Orpheus-TTS is the same decoder-only shape as Qwen3-4B, trained with the same LoRA trick, on the same kind of dataset-→forward-pass-→loss-→backprop-→gradient-descent loop. The one genuinely new piece is what the tokens are: for text they’re word-pieces, for Orpheus they’re integers produced by a neural audio codec (SNAC) that turns a sound wave into a sequence, the same way a tokenizer turns a sentence into one. You’re not implementing that codec — you’re fine-tuning on top of it, exactly the way this fine-tune sat on top of Qwen3 without you writing any of Qwen3’s code.