Lip sync works on the original face, but slight delay and mismatch are visible during fast movements.
What was measured
Lip Sync Accuracy
Does the dubbed audio visually match lip movements in the original video?
decisive for this rankingtransformation
The ranking is specifically about lip sync, so visual alignment of speech to mouth movement is a core success criterion. (3 of 3 judges)
What was given, what came back
Test input: Fitness instructor short (English → Hindi) · video · group: video-translation-voice-clone-lip-sync
Input — what we sent
Fitness instructor short (English → Hindi)
A short fitness/coaching video with energetic single-speaker English speech, used to test whether tools can translate into Hindi while preserving fast delivery, motivational tone, and original-face lip sync.
Why this input is hard
- · Fast, energetic speech transcription
- · Tone preservation for motivational fitness content
- · English-to-Hindi translation quality
- · Lip-sync accuracy on a moving face
- · End-to-end dubbing workflow automation
Output — unretouched
Also checked on this input — same tool, 4 other criteria
Automation Level◐ MixedOnce uploaded, the pipeline from speech through translation to lip sync is automated, but manual input preparation keeps the end-to-end effort at medium to high.Input Handling⚠ StruggledDirect video upload works, but YouTube Shorts links are not reliably ingested, so the clip had to be downloaded and uploaded manually.Translation Accuracy◐ MixedThe Hindi translation is moderately accurate, so the meaning is mostly preserved but not at a strong or flawless level.Voice Cloning Quality◐ MixedThe dubbed voice is decent but still sounds less natural than a stronger reference voice such as ElevenLabs.
Provenance
- Observation
- 15a284a2-8d1e-4429-a8d5-15d6fb4aa0f1
- Evidence run
- ec1dd100-6af8-4f49-8f19-b78492702b14
- Study
- Translate Videos with Voice Cloning and Lip Sync Using AI
- Research task
- 86b96dfpf
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "sync-labs",
scenario: "video-translation-voice-clone-lip-sync"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 8 other tools
measured on Lip Sync Accuracy
Camb AI✗ FailedOn the fitness clip, the dubbed video shows no mouth regeneration; whole-clip pixel-diff stayed around 1.7-2.3/255 across 5-frame intervals, and the split-screen mouth crops are frame-identical.D-ID✗ FailedLip sync is only aligned to the avatar output and not to the original live-action footage.Dubverse◐ MixedLip sync is good for slow to medium speech, but the report says it shows a slight delay in fast instruction segments.ElevenLabs✗ FailedThe tool provides no built-in lip sync, so the fitness output does not visually track mouth movements and would require external editing for sync.HeyGen✗ FailedAt a matched content point at t=12.0s, the output face was frame-for-frame identical to the input, so no visible lip-sync or face regeneration occurred.Rask AI✗ FailedAt matched timestamps the output stays frame-identical to the source, with no visible mouth or face regeneration, so the free-tier export does not perform lip sync.Synthesia◐ MixedThis fitness run did not receive an independent lip-sync judgment in the report, so there is no direct evidence here for how closely the dubbed mouth movement matched the source.VEED✗ FailedNo lip-sync regeneration happened in Test 1; the report says the mouth region was essentially identical at about 1-2/255 pixel difference, so VEED swapped the audio but left the original mouth movement unchanged.
This evidence is published in
From the same study (page rebuilt from a later run)
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com