The clean source only matched modestly, at about 40% similarity, with roughly 60% of the output sounding noticeably different from the original.
What was measured
Voice Match Accuracy
Whether the generated voice actually sounds like the original speaker, including tone, pitch, rhythm, and identity.
decisive for this rankingtransformation
This ranking is fundamentally about whether the generated speech still sounds like the target speaker, so identity match is core. (3 of 3 judges)
What was given, what came back
Test input: High-Quality Voice Sample · mixed · group: voice-cloning
Input — what we sent
Input, verbatim
Removing objects from videos used to take hours of manual editing. Now AI tools claim to do it in minutes. So we tested five AI video object removers to find the most reliable one. We used the same three inputs across all the tools for a fair comparison. ABC Labs showed unstable tracking and heavy distortion. Media.io offered fast processing but unusable outputs. PhotoRoom mostly relied on blur masking instead of real reconstruction. Runway delivered the cleanest removals with the most stable tracking and realistic scene reconstruction. Here's exactly how we tested it.
0:00 / 0:00
Loading audio...
0:00 / 0:00
Loading audio...
High-Quality Voice Sample
A clean studio-quality voice recording without background noise, used to test the best-case ceiling for voice cloning, pronunciation stability, and naturalness.
Why this input is hard
- · Maximum voice-cloning accuracy
- · Naturalness with optimal source quality
- · Long-form consistency
- · Pronunciation stability
- · Voice preservation under ideal conditions
Output — unretouched
0:00 / 0:00
Loading audio...
Also checked on this input — same tool, 2 other criteria
Naturalness & Human Quality◐ MixedThe clean source sounded fairly natural, but the researcher still heard noticeable robotic coloration, estimating it at about 70% natural and 30% AI-sounding.Pronunciation Accuracy✓ WorkedThe clean-output pass also stayed pronunciation-clear, with no flagged mispronounced words.
Provenance
- Observation
- a928a78c-e71f-4cc2-9000-714d3622ff79
- Evidence run
- 46222c41-0046-41cc-bfaa-5f7ba6aa4933
- Study
- Clone Your Voice and Generate Voiceover from Text
- Research task
- 86ba42bx1
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "speechify",
scenario: "voice-cloning"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 9 other tools
measured on Voice Match Accuracy
AICloneVoiceFree.com✓ WorkedThe cloned voice closely matches the clean source speaker, with the report calling the match about 95% and saying tone and vocal characteristics were closely preserved.ElevenLabs◐ MixedImproves source resemblance only modestly: the report rates the clone at approximately 40–50% similarity, says the voice remains noticeably polished and processed, and notes that speaker identity is only partially preserved.Fish Audio✓ WorkedRecreates a clean studio voice with full identity match: the first high-quality output was described as fully matching the uploaded input, and the second stayed equally strong.Heygen◐ MixedThe best high-quality clone still reached only about 70% similarity to the original voice, so it remained below a strong match despite being the best of the three.Inworld◐ MixedOn the clean studio sample, the clone did not match the source precisely and sounded more polished and balanced, suggesting the model optimized for pleasant delivery over exact identity.MiniMax◐ MixedEven with a clean source recording, matching accuracy still was not fully there; the report rates this run as Fair (~35–45%) and says the output leaned softer than the original voice.TopMediai Voice Cloning✗ FailedGen+ reduces similarity to the clean source voice by shifting toward a feminine tone, so speaker representation is inaccurate.Uberduck✗ FailedCleaner source audio did not materially improve identity matching; the clone still did not closely track the original voice.VocalAI⚠ StruggledEven with a clean studio sample, similarity stayed at only about 10–15%, so the cleaner input produced minimal improvement in identity preservation.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com