It is not the cloning model
Different architecture, different Space, different checkpoint. Tacotron2 rather than Tacotron, and HiFi-GAN rather than WaveRNN. The two demos share this site and a text front-end; they share no weights.
Tacotron2 · HiFi-GAN · one voice
The other model on this site clones whoever you give it. This one does the opposite: a single female voice, baked in, reading whatever you type.
The studio
No reference clip to pick — the voice is fixed. Just the text.
Output
Stage by stage, drawn in your browser from the arrays the server sends. Click any plot to enlarge it.
The decoder ran to its step ceiling ( steps) without the stop token firing, so the tail of this clip is artefact rather than speech. It is a quirk of the checkpoint, not of your text — you can see it in the alignment plot as a diagonal that runs flat into the right-hand edge.
01
Tacotron2 decodes the text into 80 mel bands over time, attending to characters as it goes. Both plots come from the same forward pass.
A clean diagonal in the alignment means the model tracked the text and stopped where the words did; one that runs flat into the right-hand edge means it never found the end.
02
HiFi-GAN turns the spectrogram back into pressure over time, in a single pass rather than sample by sample. Click the plot to seek.
0:00 / 0:00
DownloadServer timings
Under the hood
Different architecture, different Space, different checkpoint. Tacotron2 rather than Tacotron, and HiFi-GAN rather than WaveRNN. The two demos share this site and a text front-end; they share no weights.
There is no speaker embedding anywhere in this model — the decoder takes the text and nothing else. The voice is a property of the checkpoint, which is why there is no reference clip to choose and no way to ask it for a different voice.
HiFi-GAN is non-autoregressive: it produces the whole waveform in one pass instead of sampling 16,000 values per second in sequence. That is why this answers in seconds where the cloning studio, on comparable hardware, takes considerably longer for the same length of speech.
The same constraint applies: the symbol set is ASCII, unidecode
runs before the network sees anything, and digits are dropped rather than
spoken. The keyboards above are aids for you, not for the model.
Tacotron2 emits a stop token when it thinks the utterance is finished. When that gate fails to fire it runs to its 1000-step ceiling instead, which is about 11.6 s of audio regardless of how short the input was — the speech, then artefact. Watch the alignment panel: a diagonal that runs flat into the right-hand edge is this happening.
The server sends mel and alignment as raw arrays and your browser draws them, which is why they follow the theme. It used to render them as matplotlib PNGs — roughly 29% of a 705 KB response — and rebuild both networks from disk on every request while it was at it.