Nepali Voice Cloning

checking model server


SV2TTS · Nepali · 10 speakers

Clone a voice
speaking Nepali
from five seconds of audio.

A three-stage neural pipeline that learns a speaker’s timbre from one short reference clip, then reads arbitrary Nepali text back in that voice — without ever being trained on the speaker.

Built on the SV2TTS architecture and adapted to a low-resource language. The model server runs on two CPU cores, so everything you hear is synthesised live.

Reference needed
~5s
Speaker embedding
256-d
Output rate
16kHz
Trained on speaker
No, zero-shot

Try it ↓


How it works

Three networks, in series.

The trick behind zero-shot cloning: the speaker’s identity is separated from the words entirely. Stage one answers who, stage two answers what, stage three makes it audible. Watch them light up when you synthesise below.

  1. 01

    Speaker encoder

    GE2E · 3-layer LSTM

    Hears ~5s of a voice and distils it to a single fixed-length vector that captures timbre, independent of what was said.

    reference wav256-d embedding

  2. 02

    Synthesizer

    Tacotron · attention

    Conditions on that vector and decodes the text into a mel spectrogram, one frame at a time, attending to characters as it goes.

    text + embedding80 × T mel

  3. 03

    Vocoder

    WaveRNN · autoregressive

    Turns the spectrogram back into pressure over time, sampling each of 16,000 values per second in sequence.

    80 × T mel16 kHz audio

The studio

Make it say something.

01 Reference voice

The clip whose timbre gets copied. Five seconds is plenty.

02 Nepali text

Type in Devanagari, or type romanised and it converts as you go.

0 / 500

Or start from one of these

Under the hood

Things worth knowing.

It is a romanised Nepali model

The synthesizer’s vocabulary is 68 ASCII symbols — the stock Tacotron set. Devanagari never reaches the network: the text front-end runs unidecode first, so नमस्ते becomes namaste before embedding lookup. The Preeti and romanised keyboards above are typing aids for the human, not the model.

Numbers get dropped

Digits aren’t in those 68 symbols and nothing expands them, so २०२४ transliterates to 2024 and then vanishes character by character. The studio warns you rather than letting it fail silently. Proper Nepali numeral expansion is genuinely hard — the words above twenty are irregular — so it is listed as future work rather than half-done.

Nothing was trained on these speakers

The encoder was trained for speaker verification, not synthesis. It learned to map any voice to a point in 256-d space where the same person lands in the same place. That generalises to voices it has never heard, which is what makes five seconds enough.

Two CPU cores, no GPU

The model server is a free Hugging Face Space. The vocoder is autoregressive — it samples 16,000 values per second of audio, in sequence — so longer text costs proportionally more. A daily cron keeps the Space awake, since a cold start means loading roughly 490 MB of weights.

Quality is honest, not cherry-picked

Output is audibly synthetic, and attention sometimes loses the diagonal on long or unusual input — you can watch that happen in the alignment plot. A listening study put it at 3.45/5 for naturalness and 3.28/5 for speaker similarity across 57 raters, with every clip and caveat published.

Output is non-deterministic

The decoder’s prenet applies dropout at inference time, not just during training. Synthesising the same sentence twice gives two slightly different readings — inherited from the original architecture, and audible if you run it back to back.