Nepali Voice Cloning

MOS · 10 speakers · 50 clips

How close does a cloned Nepali voice actually get?

Objective loss numbers say very little about synthesised speech. The only measure that counts is whether a listener believes it. This page holds the stimuli from a mean-opinion-score study of the cloning system, the scores raters gave, and the scale they used to judge them.

Try the model yourself →

Results

What the raters said.

57 raters · 10 speakers · 570 ratings per axis · 1–5 scale · collected 2024-02-28 to 2024-03-01

Speech quality

3.45/ 5

±0.08 at 95% confidence · SD 1.00

Does it sound like natural speech?

Speaker similarity

3.28/ 5

±0.09 at 95% confidence · SD 1.15

Does it sound like the target person?

By speaker

Both axes on the same 1–5 scale, bars measured from zero.

  • Speech quality
  • Speaker similarity

Speaker 1

3.19
2.91

Speaker 2

3.67
3.35

Speaker 3

3.60
3.25

Speaker 4

3.25
3.40

Speaker 5

3.42
3.47

Speaker 6

3.26
2.89

Speaker 7

3.49
3.56

Speaker 8

3.18
2.89

Speaker 9

3.72
3.46

Speaker 10

3.77
3.65

How the ratings fell

Every individual rating, bucketed. Quality clusters on 3–4; similarity is flatter and carries a heavier tail at 1–2 — raters were harsher about who it sounded like than about whether it sounded human.

  • Speech quality
  • Speaker similarity
21
36

1

67
120

2

197
150

3

202
174

4

83
90

5

rating →

View as a table
Mean opinion scores by speaker, 1–5 scale, n = 57 raters.
SpeakerQualitySDSimilaritySD
Speaker 13.191.012.911.04
Speaker 23.670.813.351.04
Speaker 33.600.803.251.21
Speaker 43.251.073.401.05
Speaker 53.420.943.471.12
Speaker 63.261.092.891.21
Speaker 73.490.933.561.09
Speaker 83.181.232.891.22
Speaker 93.720.943.461.07
Speaker 103.770.933.651.19
All speakers3.451.003.281.15

Collected with a Google Form: raters listened to each speaker’s original recording, then rated the cloned audio on both axes. Responses were anonymous; only aggregates are published here.

Method

Two axes, five points.

Every clip was rated twice: once for how natural the speech sounds on its own, and once for how much it sounds like the target speaker. A system can score well on one and badly on the other — clean but generic, or characterful but garbled.

The five-point scale raters were given, for both axes.
Score Speech quality Speaker similarity
5 Broadcasting level: indistinguishable from a human recording. Definitely the same person; tone and speaking style match.
4 Natural, clear and understandable. Sounds like the same person, but tone and speaking style don’t match.
3 Understandable and acceptable, but the rhythmic pauses are not good enough. High chance of being the same person; some similarity.
2 Some words are unclear; pronunciation issues. Low chance of being the same person; much difference.
1 Not understandable at all. Definitely not the same person — even the gender differs.

Known limitations of this study

  • Bandwidth confound. Six of the ten original recordings are 22.05 kHz or 48 kHz while every clone is 16 kHz. A listener can pick the synthetic sample from bandwidth alone, before judging the voice. Affected speakers are flagged below. The audio is left untouched because the scores were collected against these exact files.
  • Fixed presentation order. Speakers always appear in the same sequence, original first, then short, medium, long. No randomisation or counterbalancing, so order effects are not controlled.
  • Unblinded. Clips are labelled as original or cloned, which will bias a similarity judgement.
  • Untraceable stimuli. Only two of the ten speaker slots can still be matched to a source recording, and three commits in September 2025 silently reassigned which recording sits behind which number.
  • No attention checks and no per-rater metadata, so inattentive raters cannot be identified after the fact.

None of this invalidates the samples — you can listen and judge for yourself. It does mean the aggregate numbers should be read as indicative rather than publishable.

Samples

Listen.

Audio loads only when you press play — there are 50 clips here.

Speaker 1

original 22.05 kHz vs clones 16 kHz

Original recording

4.36s · 22.05 kHz

Same text as the reference

3.93s · synthesised

Unseen text — short

2.7s · synthesised

Unseen text — medium

3.15s · synthesised

Unseen text — long

8.76s · synthesised

Speaker 2

Original recording

8.63s · 16 kHz

Same text as the reference

8.25s · synthesised

Unseen text — short

3s · synthesised

Unseen text — medium

3.58s · synthesised

Unseen text — long

8.66s · synthesised

Speaker 3

original 22.05 kHz vs clones 16 kHz

Original recording

3.55s · 22.05 kHz

Same text as the reference

2.88s · synthesised

Unseen text — short

2.76s · synthesised

Unseen text — medium

3.42s · synthesised

Unseen text — long

8.61s · synthesised

Speaker 4

original 22.05 kHz vs clones 16 kHz

Original recording

7.64s · 22.05 kHz

Same text as the reference

5.22s · synthesised

Unseen text — short

2.9s · synthesised

Unseen text — medium

3.42s · synthesised

Unseen text — long

8.55s · synthesised

Speaker 5

Original recording

4.55s · 16 kHz

Same text as the reference

4.86s · synthesised

Unseen text — short

2.7s · synthesised

Unseen text — medium

3.18s · synthesised

Unseen text — long

7.38s · synthesised

Speaker 6

Original recording

4.87s · 16 kHz

Same text as the reference

4.92s · synthesised

Unseen text — short

2.85s · synthesised

Unseen text — medium

3.3s · synthesised

Unseen text — long

8.52s · synthesised

Speaker 7

Original recording

5.36s · 16 kHz

Same text as the reference

4.53s · synthesised

Unseen text — short

2.6s · synthesised

Unseen text — medium

2.95s · synthesised

Unseen text — long

7.53s · synthesised

Speaker 8

original 22.05 kHz vs clones 16 kHz

Original recording

9.75s · 22.05 kHz

Same text as the reference

7.98s · synthesised

Unseen text — short

2.73s · synthesised

Unseen text — medium

3.21s · synthesised

Unseen text — long

9.03s · synthesised

Speaker 9

original 22.05 kHz vs clones 16 kHz

Original recording

3.96s · 22.05 kHz

Same text as the reference

4.74s · synthesised

Unseen text — short

3.3s · synthesised

Unseen text — medium

3.03s · synthesised

Unseen text — long

8.73s · synthesised

Speaker 10

original 48 kHz vs clones 16 kHz

Original recording

7.88s · 48 kHz

Same text as the reference

6.03s · synthesised

Unseen text — short

2.75s · synthesised

Unseen text — medium

3.25s · synthesised

Unseen text — long

6.92s · synthesised