Making a 0.6B TTS model 6× faster — without touching the model
Audio8 released Audio8-TTS-Preview-0.6b (Apache-2.0) — a small, nice TTS checkpoint. We built a family of voices on it and shipped them as Warble. We never changed a weight. This is just the systems story — a few lessons that generalize past one model.
Find the bottleneck before you optimize
We ran the exact same code on a laptop RTX 5090 and a datacenter RTX PRO 6000 Blackwell. The big card was only 1.07× faster.
That ratio is the diagnosis. A compute-bound job pulls away hard on a bigger GPU. When it barely moves, you're launch-bound — the GPU is starving between kernels while the CPU dispatches the next tiny op. This model fires ~1500 kernels per utterance (10 sequential codebooks × ~150 frames); the fast-AR step spent 14 ms/call doing almost no math. When your big GPU and your laptop tie, stop optimizing FLOPs and start optimizing dispatch.
Compiling more made it worse
torch.compile on the fast-AR chain took RTF 0.51 → 0.27. So we compiled the slow step too — and it got worse (0.37). Its mask grows via torch.cat each step, so shapes change every iteration and Dynamo recompiles in a loop. torch.compile only pays over a fixed-shape region. Wrap the static part; leave the growing part eager.
The rewrite
A full static-shape rebuild (preallocated buffers, hoisted constants, a lagged EOS check, persistent caches, whole frame compiled) got the laptop 5090 to RTF 0.091 — ~6× over the eager 0.54 we started at. At that point we measured ~82% of the card's memory bandwidth: that's the bf16 floor. Knowing you've hit the wall is what tells you to stop.
One trap worth a warning
On transformers 5.x this custom-code model silently emitted all-zero codes — no error, confident garbage, a waveform that looked fine. We caught it only by checking token stats, not by listening. Pin your deps (transformers==4.57.5) and validate the intermediate representation, not just the final artifact.
Serving + phone
On consumer Blackwell (sm_120), the serving path tripped over Hopper-only FlashAttention-3 in two spots; swapping to FlashInfer + a short-sequence SDPA path fixed it, and 8 concurrent streams ran at aggregate RTF 0.053. We're contributing those fixes back to Audio8. Packaged at 4-bit (909 MB), Warble runs on an iPhone-class A18 Pro at real-time — quantize only the transformer Linears, keep the tied embedding in bf16, and it's clean.
We do this kind of work
Diagnosis before optimization, a disciplined rewrite to the hardware floor, a dependency pin that stopped confident garbage. Every number here was measured on our own hardware and re-verifies independently. If you've got a model that's slower or larger than it should be and you'd rather ship it than rewrite it — that's the work we do. Thanks to Audio8 for a lovely little model.
Warble is open (Apache-2.0): huggingface.co/scrappylabsai/warble · ScrappyLabs