A path, not a tutorial

Build a TTS model from scratch

Text to speech you write yourself, trained on one Apple M5 Pro with MLX. No pretrained checkpoints. No rented GPUs.

Everything below was checked against the machine it is for, before it was written.

MeasuredResult
MLX version and device0.32.2, GPU
rfft, irfft, conv1d, transposed conv, autodiff, attentionall present
13.5M param vocoder-shaped model, forward and backward, batch 1626 ms / step
Training throughput617x realtime
One epoch over 24 h of audio2.3 min
Peak memory0.97 of 48 GB

Compute is not the constraint. Understanding is. So this path uses small models and fast loops, and spends the saved time on experiments instead of waiting.

The one rule: never debug two unknowns at once.

TTS is three systems stacked. Build text to speech in one go, get noise out, and you cannot tell whether the text encoder, the alignment, or the audio reconstruction is broken. People lose weeks there.

So every stage ends with a gate: something specific you must hear or see before moving on. If a gate fails, the bug is in the stage you just built and nowhere else. That property is the entire value of this ordering.

What you are building

text phonemes acoustic model mel spectrogram vocoder waveform build order
You build right to left. The vocoder first, because it needs no text and no alignment, and because it gives you ears for everything that follows.

The ladder

Stage 0

Audio is just numbers

2 to 3 evenings, no ML at all

The stage people skip, and pay for later. Nearly every "my model outputs garbage" bug in TTS is a spectrogram convention mismatch, not a model bug.

Implement by hand, with mx.fft: framing and a Hann window; STFT via rfft; a mel filterbank you derive yourself from 2595 ยท log10(1 + hz/700); log compression; and Griffin-Lim to invert a mel back to audio in about 30 lines.

Gate. Your own mel, through your own Griffin-Lim, sounds like a muffled but clearly intelligible version of the original. Save that file. It is your quality floor for the next two months.

You finish this stage holding a working non-neural vocoder. That is what will let you test the acoustic model independently in Stage 2.

Stage 1

Vocoder: mel to waveform

1 to 2 weeks

No text, no alignment, pure audio-to-audio supervision, and unlimited data because any speech will do. The easiest learning problem in the stack, which is exactly why it goes first.

The architecture choice that matters on a Mac: predict STFT coefficients and run one inverse STFT, rather than upsampling to raw samples through stacked transposed convolutions. That is the Vocos idea, and it is the shape benchmarked above at 26 ms per step. HiFi-GAN style raw-waveform upsampling costs several times more for the same quality.

Backbone of ConvNeXt-style 1D blocks. Loss is a multi-resolution STFT loss, magnitude L1 plus spectral convergence, at three or four FFT sizes. No GAN yet.

Work in this order, strictly: overfit one 3 second clip until reconstruction is near perfect; then 100 clips; then the full set with 50 clips held out. If the model cannot memorise three seconds, nothing later will work, and at 617x realtime that check costs you minutes.

Gate. A held-out clip, reconstructed from its mel, sounds clearly better than your Stage 0 Griffin-Lim version of the same clip. Listen to them back to back.

Only then, add a discriminator, and run it as a separate experiment against your non-GAN baseline. You will learn what the GAN actually buys, which is crispness and phase realism, and what it costs, which is training stability. Knowing that first hand is worth more than the quality bump.

Stage 2

Text to mel, and the alignment problem

2 to 3 weeks, the heart of it

The hard part is not generating spectrograms. It is that you have six words and four hundred frames and nobody tells you which frames belong to which sound. Every interesting idea in TTS architecture history is a different answer to that one problem.

Go to phonemes, not letters. "read" has two pronunciations and "gh" has three sounds. Use espeak-ng or g2p_en and move on. Keep punctuation and silence as real tokens, because that is where prosody lives.

Build it autoregressive with attention, Tacotron 2 shaped and deliberately small. Modern TTS is mostly non-autoregressive and you will get there next, but you start here because attention makes alignment visible. You can plot it.

step 500: noise it snaps step 12k: aligned text position runs left to right in time.
Plot the attention matrix every few hundred steps. When it snaps from noise into a diagonal, the model has learned that text runs left to right in time. Synthesis starts working within the hour. It is the best teaching moment in the whole project.

Gate. A sentence the model has never seen, synthesized to mel, run through your Stage 1 vocoder, and intelligible to another person who does not know what it is supposed to say. That is a complete TTS system you built from FFT to waveform.

Stage 3

Non-autoregressive, taught by your own model

1 week

Your Stage 2 model already contains the alignment, inside its attention matrix. Extract it: for each phoneme, count the frames that attended to it. Those counts are durations.

Then train a duration predictor from phoneme to frame count, and a parallel decoder that expands each phoneme and emits every mel frame in one pass. That is the FastSpeech idea, and the elegance is that you need no external aligner. Your slow model taught your fast one.

Gate. Same quality as Stage 2, one forward pass instead of hundreds, and the repeated and skipped words are gone permanently, because there is no autoregressive drift left to have.

Add pitch and energy prediction on the same pattern and you have FastSpeech 2, with per-phoneme control over how it sounds.

Stage 4

Catch up to the present

2 to 4 weeks, pick one

By now you can read current papers as variations rather than magic.

Stage 5

Your own voice

Condition on a speaker embedding taken from a reference clip, or record 20 to 30 minutes of yourself and adapt with LoRA. You already run Chatterbox locally, so use it as the target to beat and as a sanity check on your data pipeline.

Data

LJSpeech 1.1. 13,100 clips, about 24 hours, one speaker, 22.05 kHz, public domain, 2.6 GB. It is the canonical single-speaker set, which means when you are stuck, everyone else's numbers are comparable to yours. Do not start multi-speaker: speaker variation is a second unknown you do not need yet. VCTK comes later, and your own recordings at Stage 5.

How to know you are not fooling yourself

Listening to your own model is unreliable. You know what it is meant to say, so you hear words that are not there.

Reading, in the order it becomes useful

Do not read ahead. Each of these is far easier after the matching stage.

AfterReadWhy then
Stage 0Any DSP primer on STFT and the mel scaleYou want the maths after you have felt the problem
Stage 1HiFi-GAN, then VocosVocos only makes sense as a reaction to HiFi-GAN's cost
Stage 2Tacotron 2You will recognise every component you just built
Stage 2Attention Is All You NeedAlignment is attention, now seen concretely
Stage 3FastSpeech 2Reads as an obvious fix to problems you personally hit
Stage 4VALL-E, then F5 TTSNow you can judge them rather than admire them

Also read the mlx-audio source you already have cloned. Not to copy, but to compare against your own choices once you have made them.

Your first session

  1. uv pip install soundfile numpy matplotlib in .venv.
  2. Download LJSpeech, load one clip, print its shape, plot the waveform.
  3. Write mel.py: framing, Hann window, mx.fft.rfft, mel filterbank, log.
  4. Write griffin_lim.py. Invert your mel. Listen.
  5. Put the original and the reconstruction side by side in stage0/.

Muffled but intelligible means Stage 0 is passed and you have an audio pipeline you understand completely.

On pace

An epoch over 24 hours of audio takes 2.3 minutes on your machine, and a vocoder wants a few hundred epochs, so Stage 1 is hours of compute rather than days. The limiting factor is your thinking time, not the hardware.

Resist starting big runs. Run the smallest experiment that answers the question in front of you, look at the result, then decide what to ask next. Re-run the benchmark at each stage instead of trusting the extrapolation above: the numbers change with model shape, and measuring takes thirty seconds.

Published with byagent