A path, not a tutorial
Text to speech you write yourself, trained on one Apple M5 Pro with MLX. No pretrained checkpoints. No rented GPUs.
Everything below was checked against the machine it is for, before it was written.
| Measured | Result |
|---|---|
| MLX version and device | 0.32.2, GPU |
rfft, irfft, conv1d, transposed conv, autodiff, attention | all present |
| 13.5M param vocoder-shaped model, forward and backward, batch 16 | 26 ms / step |
| Training throughput | 617x realtime |
| One epoch over 24 h of audio | 2.3 min |
| Peak memory | 0.97 of 48 GB |
Compute is not the constraint. Understanding is. So this path uses small models and fast loops, and spends the saved time on experiments instead of waiting.
The one rule: never debug two unknowns at once.
TTS is three systems stacked. Build text to speech in one go, get noise out, and you cannot tell whether the text encoder, the alignment, or the audio reconstruction is broken. People lose weeks there.
So every stage ends with a gate: something specific you must hear or see before moving on. If a gate fails, the bug is in the stage you just built and nowhere else. That property is the entire value of this ordering.
The stage people skip, and pay for later. Nearly every "my model outputs garbage" bug in TTS is a spectrogram convention mismatch, not a model bug.
Implement by hand, with mx.fft: framing and a Hann window; STFT via rfft; a mel filterbank you derive yourself from 2595 ยท log10(1 + hz/700); log compression; and Griffin-Lim to invert a mel back to audio in about 30 lines.
Gate. Your own mel, through your own Griffin-Lim, sounds like a muffled but clearly intelligible version of the original. Save that file. It is your quality floor for the next two months.
You finish this stage holding a working non-neural vocoder. That is what will let you test the acoustic model independently in Stage 2.
No text, no alignment, pure audio-to-audio supervision, and unlimited data because any speech will do. The easiest learning problem in the stack, which is exactly why it goes first.
The architecture choice that matters on a Mac: predict STFT coefficients and run one inverse STFT, rather than upsampling to raw samples through stacked transposed convolutions. That is the Vocos idea, and it is the shape benchmarked above at 26 ms per step. HiFi-GAN style raw-waveform upsampling costs several times more for the same quality.
Backbone of ConvNeXt-style 1D blocks. Loss is a multi-resolution STFT loss, magnitude L1 plus spectral convergence, at three or four FFT sizes. No GAN yet.
Work in this order, strictly: overfit one 3 second clip until reconstruction is near perfect; then 100 clips; then the full set with 50 clips held out. If the model cannot memorise three seconds, nothing later will work, and at 617x realtime that check costs you minutes.
Gate. A held-out clip, reconstructed from its mel, sounds clearly better than your Stage 0 Griffin-Lim version of the same clip. Listen to them back to back.
Only then, add a discriminator, and run it as a separate experiment against your non-GAN baseline. You will learn what the GAN actually buys, which is crispness and phase realism, and what it costs, which is training stability. Knowing that first hand is worth more than the quality bump.
The hard part is not generating spectrograms. It is that you have six words and four hundred frames and nobody tells you which frames belong to which sound. Every interesting idea in TTS architecture history is a different answer to that one problem.
Go to phonemes, not letters. "read" has two pronunciations and "gh" has three sounds. Use espeak-ng or g2p_en and move on. Keep punctuation and silence as real tokens, because that is where prosody lives.
Build it autoregressive with attention, Tacotron 2 shaped and deliberately small. Modern TTS is mostly non-autoregressive and you will get there next, but you start here because attention makes alignment visible. You can plot it.
Gate. A sentence the model has never seen, synthesized to mel, run through your Stage 1 vocoder, and intelligible to another person who does not know what it is supposed to say. That is a complete TTS system you built from FFT to waveform.
Your Stage 2 model already contains the alignment, inside its attention matrix. Extract it: for each phoneme, count the frames that attended to it. Those counts are durations.
Then train a duration predictor from phoneme to frame count, and a parallel decoder that expands each phoneme and emits every mel frame in one pass. That is the FastSpeech idea, and the elegance is that you need no external aligner. Your slow model taught your fast one.
Gate. Same quality as Stage 2, one forward pass instead of hundreds, and the repeated and skipped words are gone permanently, because there is no autoregressive drift left to have.
Add pitch and energy prediction on the same pattern and you have FastSpeech 2, with per-phoneme control over how it sounds.
By now you can read current papers as variations rather than magic.
Condition on a speaker embedding taken from a reference clip, or record 20 to 30 minutes of yourself and adapt with LoRA. You already run Chatterbox locally, so use it as the target to beat and as a sanity check on your data pipeline.
LJSpeech 1.1. 13,100 clips, about 24 hours, one speaker, 22.05 kHz, public domain, 2.6 GB. It is the canonical single-speaker set, which means when you are stuck, everyone else's numbers are comparable to yours. Do not start multi-speaker: speaker variation is a second unknown you do not need yet. VCTK comes later, and your own recordings at Stage 5.
Listening to your own model is unreliable. You know what it is meant to say, so you hear words that are not there.
Do not read ahead. Each of these is far easier after the matching stage.
| After | Read | Why then |
|---|---|---|
| Stage 0 | Any DSP primer on STFT and the mel scale | You want the maths after you have felt the problem |
| Stage 1 | HiFi-GAN, then Vocos | Vocos only makes sense as a reaction to HiFi-GAN's cost |
| Stage 2 | Tacotron 2 | You will recognise every component you just built |
| Stage 2 | Attention Is All You Need | Alignment is attention, now seen concretely |
| Stage 3 | FastSpeech 2 | Reads as an obvious fix to problems you personally hit |
| Stage 4 | VALL-E, then F5 TTS | Now you can judge them rather than admire them |
Also read the mlx-audio source you already have cloned. Not to copy, but to compare against your own choices once you have made them.
uv pip install soundfile numpy matplotlib in .venv.mel.py: framing, Hann window, mx.fft.rfft, mel filterbank, log.griffin_lim.py. Invert your mel. Listen.stage0/.Muffled but intelligible means Stage 0 is passed and you have an audio pipeline you understand completely.
An epoch over 24 hours of audio takes 2.3 minutes on your machine, and a vocoder wants a few hundred epochs, so Stage 1 is hours of compute rather than days. The limiting factor is your thinking time, not the hardware.
Resist starting big runs. Run the smallest experiment that answers the question in front of you, look at the result, then decide what to ask next. Re-run the benchmark at each stage instead of trusting the extrapolation above: the numbers change with model shape, and measuring takes thirty seconds.