HD delivers the best pronunciation result on the noisy sample, with no misread or garbled words noted during the listening pass.
What was measured
Pronunciation Accuracy
How well the tool handles names, technical terms, numbers, acronyms, and other tricky pronunciations.
decisive for this rankingtransformation
If a cloned voice cannot handle names, numbers, acronyms, and technical terms correctly, it fails at producing usable voiceover. (3 of 3 judges)
What was given, what came back
Test input: Low-Quality Voice Sample · mixed · group: voice-cloning
Input — what we sent
Input, verbatim
Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.
0:00 / 0:00
Loading audio...
Research media topmediai lowquality input.wav
0:00 / 0:00
Loading audio...
Low-Quality Voice Sample
A noisy voice recording with background noise, room ambience, and minor disturbances, used to test whether voice-cloning tools can preserve speaker identity when the source audio is imperfect.
Why this input is hard
- · Cloning accuracy from degraded audio
- · Noise and ambience robustness
- · Speaker identity preservation under poor recording conditions
- · Distinguishing enhancement from true cloning
Output — unretouched
0:00 / 0:00
Loading audio...
Also checked on this input — same tool, 9 other criteria
Long-Form Consistency✓ WorkedHD stays consistent through long-form narration on the noisy sample, with no abrupt changes in voice quality or pronunciation.Long-Form Consistency✓ WorkedGen+ remains stable through long-form generation on the noisy sample, with no noticeable interruptions or voice degradation.Long-Form Consistency✓ WorkedGen stays consistent through longer passages on the noisy sample, with no major voice breaks, glitches, or abrupt tonal shifts.Naturalness & Human Quality◐ MixedGen+ is more natural than Gen, but the gender inconsistency keeps the overall cloning quality from feeling convincing.Naturalness & Human Quality⚠ StruggledGen sounds noticeably robotic and lacks emotional depth and natural speech rhythm on the noisy sample.Naturalness & Human Quality✓ WorkedHD is the most human-like output on the noisy sample, with better emotional tone, speech rhythm, and vocal realism than Gen or Gen+.Voice Match Accuracy◐ MixedGen keeps some resemblance to the original speaker on the noisy source sample, but it only partially preserves identity and still sounds robotic.Voice Match Accuracy✗ FailedGen+ loses speaker identity on the noisy sample by drifting toward a feminine vocal tone instead of the source male voice.Voice Match Accuracy✓ WorkedHD is the closest match on the noisy sample and preserves vocal identity best among the three variants.
Provenance
- Observation
- e6e02fc5-bd69-4bad-a81f-c679c77f1540
- Evidence run
- 46222c41-0046-41cc-bfaa-5f7ba6aa4933
- Study
- Clone Your Voice and Generate Voiceover from Text
- Research task
- 86ba42bx1
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "topmediai-voice-cloning",
scenario: "voice-cloning"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 6 other tools
measured on Pronunciation Accuracy
ElevenLabs✓ WorkedKeeps pronunciation stable even when the source audio is noisy, with no major pronunciation breakdown reported in the generated narration.Inworld✓ WorkedAcross the low-quality pass, all words stayed intelligible and no misread or garbled words were noted; no dedicated names/numbers/acronyms stress test was run on this scenario.MiniMax◐ MixedThe low-quality English generation mispronounced a small number of words, with the report calling out 2–3 misreads in this clip.Speechify✓ WorkedThe noisy-output pass stayed pronunciation-clear, and the researcher flagged no specific mispronounced words.Uberduck✓ WorkedStandard-script English words stayed intelligible; the observed problem was delivery/prosody, not misread or garbled words.VocalAI✓ WorkedThe noisy-sample run did not show major pronunciation problems.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com