How the fine-tune you ran actually works

Frozen & Trained

You know gradient descent. This is everything between that and the LoRA run you did in 40 seconds on your Mac — built up one honest step at a time, nothing skipped.

frozen — never updated trained — gradient descent touches this
Step 1

A layer is one matrix multiply

Picture a giant mixing console, like a sound engineer's desk with thousands of sliders. You feed in a row of numbers on the left. Every slider on the right is just some weighted blend of all the sliders on the left — turn some inputs up, some down, add them together. That blend is the entire "layer". The slider positions are the knobs, and there can be millions of them.

A neural network is just several of these mixing consoles chained one after another, the output of one feeding the input of the next.

One mixing console, unrolled
numbers in (the input row)
↓ every slider, W, blends all the inputs
nudge, then clip anything negative to zero (keeps it from just being a straight line)
↓
numbers out — feeds the next console

Chain a few dozen consoles together and the whole pipeline can approximate absurdly complicated things — "is this a cat" or "what word comes next". Training is the loop you already know, just at scale: run an example through, see how wrong the answer was, then figure out which sliders to nudge and in which direction to make it less wrong.

Working out "which direction" for a slider buried in the third console, when the wrongness is only measured at the very end, is the one genuinely new piece. It's solved by working backwards: a coach reviewing game footage in reverse, starting from the missed goal, and tracing back through each player's touch to say exactly how much each one contributed and which way they should have moved. That backward pass is called backpropagation. Once every slider has its "which way, how much", that's a gradient — and you already know what to do with a gradient.

A 4-billion-parameter model like Qwen3-4B is this exact console-chain, with 4 billion sliders instead of a few hundred. Nothing more exotic is happening underneath — just scale.
Step 2

The one new idea in a transformer: attention

Plain mixing consoles have a blind spot: a word's meaning depends on other words far away in the sentence, but a slider only ever sees its own fixed input slot — it has no way to reach across and check a different word. You do this reaching-across constantly without noticing: reading a mystery novel and hitting "he did it", your brain instantly flips back through everyone introduced so far and weighs which one "he" most plausibly points at. Attention is a layer built to do exactly that flip-back-and-check, for every word, automatically.

Worked example — what "it" attends to

The trophy didn’t fit in the suitcase because it was too big.

other wordraw scoreattention weight
trophy4.8
0.61
suitcase3.1
0.23
big2.6
0.14
because0.4
0.02

Swap “big” for “small” and every one of these weights recomputes — now suitcase wins. Same word “it”, same position, different answer, because the answer was never stored anywhere — it’s recomputed from context every time.

Where the scores come from — think of it as a search engine, inside the sentence

Every word does two things at once: it puts out a search request saying what it's looking for, and it wears a tag advertising what it's about. "It" puts out a search request that basically reads "I'm a thing that can be too big or too small, find me a noun." Every other word's tag gets checked against that request — "trophy" wears a tag like "physical object, has a size" and matches well; "because" wears a tag like "connector word" and matches badly. The match strength is the score in the table above. Once you know the scores, you also pull each word's actual content — its value, roughly "what I mean" — and blend it in proportion to how well it matched.

In the jargon: the search request is called a Query, the tag is a Key, the content is a Value. Same idea as typing into a search box (your Query) and it matching against every webpage's index terms (their Keys), then returning the actual page content (the Value) for the best matches.

Query, Key, and Value are just three more sliders-consoles from Step 1. They are ordinary knobs, trained by the exact same backwards-review-and-nudge loop as every other layer. Attention isn't a different training method — it's one more mixing-console shape, whose particular wiring happens to let words check each other.

Stack a run of attention layers and plain layers together, and that stack is what people call a transformer.

Step 3

The problem LoRA solves

Qwen3-4B has roughly four billion sliders. Fine-tuning the ordinary way means figuring out, storing, and nudging a correction for every single one — far more memory than an M-series Mac has.

Think of the frozen model as a huge, expensively-printed textbook. You're not allowed to reprint it — too big, too expensive, and you'd need to redo all four billion words in it. But you're allowed to clip a thin insert of sticky notes to it. When reading, you read the original page and your sticky note together, and mentally add the correction on top. That's the whole trick: freeze the textbook (W), never touch it again, and only ever write and rewrite the sticky notes (A and B below).

One layer, with LoRA attached — real dimensions for a 4096-wide layer
W4096×4096, frozen
×
input4096 numbers
+
B4096×8, trained
×
A8×4096, trained
×
inputsame 4096

output = W·input + B·(A·input)

Instead of writing a full, detailed sticky note for all 4096 outputs individually, A forces the correction to first get summarised down into just 8 core adjustments — that “8” is the rank in Low-Rank Adaptation, and it's like writing a short summary note instead of retyping the whole page. B then expands those 8 core adjustments back out to touch all 4096 outputs. Squeezing through that narrow 8-wide bottleneck is the entire reason it's cheap: you're writing roughly 4096×8 + 8×4096 numbers of sticky note instead of retyping 4096×4096 — about two orders of magnitude less to write, for this one layer.

Only the sticky notes, A and B, ever get corrected during training. The textbook, W, sits there frozen — you read from it every single time, but you never once pick up a pen and change it.

Step 4

Why “fuse” is legitimate, not a hack

Matrix multiplication distributes over addition. That means:

(W + B·A) · input  =  W·input + B·(A·input)

The left side and the right side land on the exact same number. So the sticky-note correction behaves exactly as if it had been written straight into the textbook, even though you never touched the original pages — you only ever read the page and the note side by side and added them in your head.

Fuse takes that literally: it computes what the sticky-note correction would look like written out in full (the same shape as a whole page, 4096×4096), and writes a brand new page that already has the correction baked in. After fusing, there's no separate sticky note left to flip to — you get back one ordinary textbook, at the original textbook's normal reading speed, with the correction now permanent and invisible.

The “Q”
in QLoRA is separate — it means the frozen W is stored in 4-bit instead of 16-bit to save memory. It has nothing to do with the A/B trick itself.
Dataset
the small file of example conversations that got run through the frozen model plus A/B, over and over, each pass nudging only A and B toward lower loss on those examples.
200 iterations
200 gradient descent steps. Same loop as every other model you’ve trained, just 200 steps of it, on a tiny number of knobs.

What this means for a voice model

None of the above changes for audio. Orpheus-TTS is the same decoder-only shape as Qwen3-4B, trained with the same LoRA trick, on the same kind of dataset-→forward-pass-→loss-→backprop-→gradient-descent loop. The one genuinely new piece is what the tokens are: for text they’re word-pieces, for Orpheus they’re integers produced by a neural audio codec (SNAC) that turns a sound wave into a sequence, the same way a tokenizer turns a sentence into one. You’re not implementing that codec — you’re fine-tuning on top of it, exactly the way this fine-tune sat on top of Qwen3 without you writing any of Qwen3’s code.

Published with byagent