Lip sync is only aligned to the avatar output and not to the original live-action footage.

✗ Failedinput onlyTest date not recordedD-ID
What was measured
Lip Sync Accuracy

Does the dubbed audio visually match lip movements in the original video?

decisive for this rankingtransformation

The ranking is specifically about lip sync, so visual alignment of speech to mouth movement is a core success criterion. (3 of 3 judges)

What was given, what came back

Test input: Fitness instructor short (English → Hindi) · video · group: video-translation-voice-clone-lip-sync
Input — what we sent

A short fitness/coaching video with energetic single-speaker English speech, used to test whether tools can translate into Hindi while preserving fast delivery, motivational tone, and original-face lip sync.

Why this input is hard
  • · Fast, energetic speech transcription
  • · Tone preservation for motivational fitness content
  • · English-to-Hindi translation quality
  • · Lip-sync accuracy on a moving face
  • · End-to-end dubbing workflow automation
Output — unretouched
No output artifact
The verdict rests on the tester's written observation alone — no file was captured for this cell.
Provenance
Observation
f20300e9-63fd-4fd5-bbc7-8a8b015f31a0
Evidence run
ec1dd100-6af8-4f49-8f19-b78492702b14
Study
Translate Videos with Voice Cloning and Lip Sync Using AI
Research task
86b96dfpf
Tested at
not recorded
Source
first-party
Evidence state
observed
Proof shown
input only
Cost / latency
not captured
Repeat run
not captured
Tester
not captured

The last three rows are honest blanks, not placeholders — our capture has no field for them yet.

Query this
get_evidence({
  tool: "d-id",
  scenario: "video-translation-voice-clone-lip-sync"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 8 other tools
measured on Lip Sync Accuracy
From the same study (page rebuilt from a later run)
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com