Speaker 1
original 22.05 kHz vs clones 16 kHz
MOS · 10 speakers · 50 clips
Objective loss numbers say very little about synthesised speech. The only measure that counts is whether a listener believes it. This page holds the stimuli from a mean-opinion-score study of the cloning system, the scores raters gave, and the scale they used to judge them.
Results
57 raters · 10 speakers · 570 ratings per axis · 1–5 scale · collected 2024-02-28 to 2024-03-01
Speech quality
3.45/ 5
±0.08 at 95% confidence · SD 1.00
Does it sound like natural speech?
Speaker similarity
3.28/ 5
±0.09 at 95% confidence · SD 1.15
Does it sound like the target person?
Both axes on the same 1–5 scale, bars measured from zero.
Every individual rating, bucketed. Quality clusters on 3–4; similarity is flatter and carries a heavier tail at 1–2 — raters were harsher about who it sounded like than about whether it sounded human.
1
2
3
4
5
rating →
| Speaker | Quality | SD | Similarity | SD |
|---|---|---|---|---|
| Speaker 1 | 3.19 | 1.01 | 2.91 | 1.04 |
| Speaker 2 | 3.67 | 0.81 | 3.35 | 1.04 |
| Speaker 3 | 3.60 | 0.80 | 3.25 | 1.21 |
| Speaker 4 | 3.25 | 1.07 | 3.40 | 1.05 |
| Speaker 5 | 3.42 | 0.94 | 3.47 | 1.12 |
| Speaker 6 | 3.26 | 1.09 | 2.89 | 1.21 |
| Speaker 7 | 3.49 | 0.93 | 3.56 | 1.09 |
| Speaker 8 | 3.18 | 1.23 | 2.89 | 1.22 |
| Speaker 9 | 3.72 | 0.94 | 3.46 | 1.07 |
| Speaker 10 | 3.77 | 0.93 | 3.65 | 1.19 |
| All speakers | 3.45 | 1.00 | 3.28 | 1.15 |
Collected with a Google Form: raters listened to each speaker’s original recording, then rated the cloned audio on both axes. Responses were anonymous; only aggregates are published here.
Method
Every clip was rated twice: once for how natural the speech sounds on its own, and once for how much it sounds like the target speaker. A system can score well on one and badly on the other — clean but generic, or characterful but garbled.
| Score | Speech quality | Speaker similarity |
|---|---|---|
| 5 | Broadcasting level: indistinguishable from a human recording. | Definitely the same person; tone and speaking style match. |
| 4 | Natural, clear and understandable. | Sounds like the same person, but tone and speaking style don’t match. |
| 3 | Understandable and acceptable, but the rhythmic pauses are not good enough. | High chance of being the same person; some similarity. |
| 2 | Some words are unclear; pronunciation issues. | Low chance of being the same person; much difference. |
| 1 | Not understandable at all. | Definitely not the same person — even the gender differs. |
None of this invalidates the samples — you can listen and judge for yourself. It does mean the aggregate numbers should be read as indicative rather than publishable.
Samples
Audio loads only when you press play — there are 50 clips here.
original 22.05 kHz vs clones 16 kHz
original 22.05 kHz vs clones 16 kHz
original 22.05 kHz vs clones 16 kHz
original 22.05 kHz vs clones 16 kHz
original 22.05 kHz vs clones 16 kHz
original 48 kHz vs clones 16 kHz