Gen keeps the script intelligible on the clean sample; the robotic character is a delivery issue, not misread words.
What was measured
Pronunciation Accuracy
How well the tool handles names, technical terms, numbers, acronyms, and other tricky pronunciations.
decisive for this rankingtransformation
If a cloned voice cannot handle names, numbers, acronyms, and technical terms correctly, it fails at producing usable voiceover. (3 of 3 judges)
What was given, what came back
Test input: High-Quality Voice Sample · mixed · group: voice-cloning
Input — what we sent
Input, verbatim
Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.
0:00 / 0:00
Loading audio...
Research media topmediai highquality input.wav
0:00 / 0:00
Loading audio...
High-Quality Voice Sample
A clean studio-quality voice recording without background noise, used to test the best-case ceiling for voice cloning, pronunciation stability, and naturalness.
Why this input is hard
- · Maximum voice-cloning accuracy
- · Naturalness with optimal source quality
- · Long-form consistency
- · Pronunciation stability
- · Voice preservation under ideal conditions
Output — unretouched
0:00 / 0:00
Loading audio...
Also checked on this input — same tool, 9 other criteria
Long-Form Consistency✓ WorkedGen+ stays stable through long-form generation on the clean sample, with no voice breaks or instability observed.Long-Form Consistency✓ WorkedHD stays consistent across longer scripts on the clean sample, with no noticeable pronunciation issues, interruptions, or degradation.Long-Form Consistency✓ WorkedGen remains consistent through the generated narration on the clean sample, with no interruptions or instability observed.Naturalness & Human Quality◐ MixedGen+ sounds more natural than Gen, but the weak identity preservation keeps the output from feeling fully convincing.Naturalness & Human Quality✓ WorkedHD is the most human-like and realistic output on the clean sample, with better emotional delivery and conversational flow.Naturalness & Human Quality⚠ StruggledGen still sounds robotic on the clean source sample, making it less convincing than HD.Voice Match Accuracy◐ MixedGen keeps some similarity to the clean source voice, but speaker identity remains only moderately accurate and still sounds robotic.Voice Match Accuracy✓ WorkedHD has the highest similarity to the clean source voice and preserves speaker identity best among the high-quality outputs.Voice Match Accuracy✗ FailedGen+ reduces similarity to the clean source voice by shifting toward a feminine tone, so speaker representation is inaccurate.
Provenance
- Observation
- b62a2ccb-2b14-4585-b4de-0b510c9887c2
- Evidence run
- 46222c41-0046-41cc-bfaa-5f7ba6aa4933
- Study
- Clone Your Voice and Generate Voiceover from Text
- Research task
- 86ba42bx1
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "topmediai-voice-cloning",
scenario: "voice-cloning"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 8 other tools
measured on Pronunciation Accuracy
AICloneVoiceFree.com◐ MixedPronunciation is not perfect even when identity matching is strong: the report notes minor pronunciation issues in the clean-voice output.ElevenLabs✓ WorkedPronounces words correctly in the long-form English output, with pronunciation reported as stable across the generated script.Fish Audio◐ MixedCan still make isolated pronunciation mistakes on clean English input: one high-quality output had a single mispronounced word, about 2–3% of the generation, while the paired output had no such issue.Inworld✓ WorkedNo misread or garbled words were noted in the high-quality pass, and no dedicated pronunciation stress test was run on that scenario.MiniMax✓ WorkedNo misread or garbled words were noted on the clean-input generation.Speechify✓ WorkedThe clean-output pass also stayed pronunciation-clear, with no flagged mispronounced words.Uberduck✓ WorkedEnglish words remained intelligible, and the cleaner source did not change that ceiling on this run.VocalAI✓ WorkedThe clean-sample run showed no major pronunciation errors.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com