The tool does not support direct video translation for this input; it required manual script entry first.
What was measured
Input Handling
Accepts plain-language prompts without detailed specs.
context, not decisivecapability
Accepting plain-language prompts affects usability and onboarding, but it is not a measure of the translation, dubbing, or lip-sync result itself. (3 of 3 judges)
What was given, what came back
Test input: Hindi vlog-style talking head (Hindi → English) · video · group: video-translation-voice-clone-lip-sync
Input — what we sent
D-ID — Input 2 Educational.mp4
Input not captured
This run recorded no prompt or input file for the test, so we cannot show you what produced the result below. Capture gaps are tracked, not hidden.
A casual Hindi/Hinglish talking-head or vlog-style creator video translated into English to stress informal language handling, slang, and natural-sounding voice cloning on real creator content.
Output — unretouched
D-ID — D-ID_EducationalVideo_EN-ES.mp4.mp4
Also checked on this input — same tool, 12 other criteria
Automation Level◐ MixedAutomates translation, but the setup remains manual, so the workflow is only partially automated.Automation Level◐ MixedThe workflow stayed only partially automated because it depended on manual setup even though translation itself worked.Automation Level◐ MixedCan translate the vlog, but the workflow still depends on manual setup, so automation is only medium.Automation Level⚠ StruggledRequires manual script extraction and avatar setup, so the pipeline is low automation rather than end-to-end.Lip Sync Accuracy✗ FailedLip sync was stable on the avatar, but it did not reflect the original speaker’s expressions.Lip Sync Accuracy✗ FailedLip sync only matched the generated AI avatar; the tool does not lip-sync the original fitness footage.Lip Sync Accuracy✗ FailedLip sync worked well on the AI avatar, but it did not apply to the original video.Output Quality & Export◐ MixedExport was available, but the free plan added limitations such as credits and watermarks.Output Quality & Export◐ MixedExport was available, but it was limited in the free tier.Translation Accuracy◐ MixedEnglish output is understandable, but it becomes slightly formal and does not fully preserve the casual vlog tone.Translation Accuracy✓ WorkedThe English→Spanish translation was accurate.Translation Accuracy✓ WorkedCan produce an acceptable English-to-Hindi translation.
Provenance
- Observation
- c8fb4db4-7880-4b69-bf4d-2059449a083e
- Evidence run
- 9feca581-b454-4bda-a610-a646cd0092f0
- Study
- Translate Videos with Voice Cloning and Lip Sync Using AI
- Research task
- 86b96dfpf
- Tested at
- Jun 24, 2026
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "d-id",
scenario: "video-translation-voice-clone-lip-sync"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 2 other tools
measured on Input Handling
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com