Produces generally good speech, but pacing swings between noticeably too fast and noticeably too slow, making the delivery feel less natural and slightly artificial.
What was measured
Naturalness & Human Quality
How realistic and human-like the output sounds, including pacing, breathing, and micro-pauses.
decisive for this rankingtransformation
A voiceover tool has to sound human and listenable; unnatural pacing or robotic delivery undermines the main job. (3 of 3 judges)
What was given, what came back
Test input: Low-Quality Voice Sample · mixed · group: voice-cloning
Input — what we sent
Input, verbatim
Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.
0:00 / 0:00
Loading audio...
0:00 / 0:00
Loading audio...
Low-Quality Voice Sample
A noisy voice recording with background noise, room ambience, and minor disturbances, used to test whether voice-cloning tools can preserve speaker identity when the source audio is imperfect.
Why this input is hard
- · Cloning accuracy from degraded audio
- · Noise and ambience robustness
- · Speaker identity preservation under poor recording conditions
- · Distinguishing enhancement from true cloning
Output — unretouched
0:00 / 0:00
Loading audio...
0:00 / 0:00
Loading audio...
Also checked on this input — same tool, 3 other criteria
Long-Form Consistency✓ WorkedHolds voice quality and pronunciation stable across extended narration, with no major degradation reported during long-form generation.Pronunciation Accuracy✓ WorkedKeeps pronunciation stable even when the source audio is noisy, with no major pronunciation breakdown reported in the generated narration.Voice Match Accuracy◐ MixedClones only about half of the source-speaker identity in the noisy sample: the report rates the match at approximately 50% and says the output sounds heavily polished, which lowers resemblance instead of faithfully reproducing the original voice.
Provenance
- Observation
- d43f3053-c54e-42c7-8c9d-b35187ad8316
- Evidence run
- 46222c41-0046-41cc-bfaa-5f7ba6aa4933
- Study
- Clone Your Voice and Generate Voiceover from Text
- Research task
- 86ba42bx1
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "elevenlabs",
scenario: "voice-cloning"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 9 other tools
measured on Naturalness & Human Quality
AICloneVoiceFree.com✓ WorkedThe generated speech sounds natural and human-like, with smooth flow, pleasant pacing, and no major robotic artifacts noticed in the preview.Fish Audio✓ WorkedKeeps noisy-source speech sounding human-like rather than robotic: both low-quality outputs were reported as non-AI-sounding and close to the input voice's character.Heygen◐ MixedA mid-tier low-quality clone improved rhythm and speech delivery and sounded more human-like than Output 1, but it still was not fully natural.Inworld✓ WorkedThe low-quality output still sounded human rather than robotic, coming across as a genuine speaker instead of a synthetic one.MiniMax✓ WorkedThe low-quality run still sounded human rather than robotic or AI-generated, despite the accuracy gap.Speechify✓ WorkedThe noisy source still yielded a natural-sounding voice in isolation.TopMediai Voice Cloning⚠ StruggledGen sounds noticeably robotic and lacks emotional depth and natural speech rhythm on the noisy sample.Uberduck✗ FailedThe output sounded heavily robotic, with frequent unnatural pauses that made it immediately identifiable as AI-generated.VocalAI✓ WorkedThe generated speech sounded natural and human-like, with clean, easy-to-listen-to audio quality despite the weak voice match.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com