InVideo AI
Creates original vertical shorts and avatar-led ads, but often needs QA and follow-up prompts
Our take
- You want original AI scenes for a text-prompted short instead of a stock-footage montage.
- You need a recurring character or believable presenter across scenes in a vertical video.
- You want avatar-led UGC-style ads from a script and are willing to review each render.
- You need a guaranteed one-shot finished short or ad on the first render.
Our take
InVideo AI is strong when you want text-to-video output that looks original rather than stock-heavy: it produced original-looking scenes, consistent characters, and believable avatar-led vertical ads. It also supports scene switching, B-roll insertion, and conversational follow-up prompts, which makes refinement possible after the first pass. The tradeoff is reliability: captions, voice, and music may need extra prompting, some renders can stop short or drift in continuity, and embedded UI text can come out garbled.
In-Depth Review
Our detailed analysis of InVideo AI — features, performance, and real-world testing.
Feature-by-Feature Breakdown
Vertical MP4 ExportStrong▾
Feature tested: Vertical MP4 Export
Result: Passed
Verdict: Strong
Expected behavior: Exports a ready-to-download 9:16 MP4, including full-bleed vertical delivery and watermark-free output on paid plans. The two cards differ only in wording and plan detail.
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): Clean 9:16 export with no visible watermark. — InVideo output 1.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): Clean 9:16 export with no visible watermark. — InVideo output 1.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): Clean 9:16 export with no visible watermark. — invideo output 2.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): Clean 9:16 export with no visible watermark. — invideo output 2.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): Clean 9:16 export with no visible watermark. — invideo output 3.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): Clean 9:16 export with no visible watermark. — invideo output 3.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): Exported as a vertical MP4 on the paid Max plan, with the report noting a correct 1080x1920 format and no watermark. — InVideoAI_Anchor1_OutputVideo.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): Exported as a vertical MP4 on the paid Max plan, with the report noting a correct 1080x1920 format and no watermark. — InVideoAI_Anchor1_OutputVideo.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): Exported as a watermark-free vertical MP4, with final duration close to the requested 30 seconds. — InVideoAI_Anchor2_OutputVideo.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): Exported as a watermark-free vertical MP4, with final duration close to the requested 30 seconds. — InVideoAI_Anchor2_OutputVideo.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): Vertical 1080x1920 MP4 export from the paid plan, watermark-free. — InVideoAI_Anchor1_OutputVideo.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): Vertical 1080x1920 MP4 export from the paid plan, watermark-free. — InVideoAI_Anchor1_OutputVideo.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): Vertical 1080x1920 MP4 export from the paid plan, watermark-free. — InVideoAI_Anchor2_OutputVideo.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): Vertical 1080x1920 MP4 export from the paid plan, watermark-free. — InVideoAI_Anchor2_OutputVideo.mp4
What changed: Text prompt transformed into Video file
Why it matters / Conclusion: The paid Max export is clean and publication-ready on format and watermarking.
Exports a ready-to-download 9:16 MP4, including full-bleed vertical delivery and watermark-free output on paid plans. The two cards differ only in wording and plan detail.
Avatar-led UGC video generationStrong▾
Feature tested: Avatar-led UGC video generation
Result: Partial
Verdict: Strong
Expected behavior: Turns a supplied script into a vertical ad with a realistic on-camera AI presenter. Across the SaaS, physical-product, and app tests, the presenter stayed believable, and the main talking-head shots kept the same presenter identity.
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): Finished vertical talking-head clip with a realistic presenter, but it stops at 'worth' instead of completing the sentence and contains no ad structure or product cutaways. — InVideo output 1.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): Finished vertical talking-head clip with a realistic presenter, but it stops at 'worth' instead of completing the sentence and contains no ad structure or product cutaways. — InVideo output 1.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): Complete vertical product-ad clip with presenter-held shoe footage and running B-roll, but the runner in B-roll is a different person than the on-camera reviewer. — invideo output 2.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): Complete vertical product-ad clip with presenter-held shoe footage and running B-roll, but the runner in B-roll is a different person than the on-camera reviewer. — invideo output 2.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): Complete vertical promo with presenter, phone cutaway, and branded Duo owl outro; the B-roll is generic phone UI rather than a specific Duolingo lesson screen. — invideo output 3.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): Complete vertical promo with presenter, phone cutaway, and branded Duo owl outro; the B-roll is generic phone UI rather than a specific Duolingo lesson screen. — invideo output 3.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): The captions advance word by word, but the render cuts off at 'worth' and never reaches the full closing sentence. — InVideo output 1.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): The captions advance word by word, but the render cuts off at 'worth' and never reaches the full closing sentence. — InVideo output 1.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): The script completes cleanly through the final word 'considering' with no truncation. — invideo output 2.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): The script completes cleanly through the final word 'considering' with no truncation. — invideo output 2.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): The script is rendered verbatim end to end, including the closing CTA line. — invideo output 3.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): The script is rendered verbatim end to end, including the closing CTA line. — invideo output 3.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): The opening and closing presenter remain the same, but the running B-roll shows a visibly different person, which breaks the testimonial premise. — invideo output 2.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): The opening and closing presenter remain the same, but the running B-roll shows a visibly different person, which breaks the testimonial premise. — invideo output 2.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): The on-camera presenter stays visually consistent within the clip, with clean face and hand rendering. — InVideo output 1.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): The on-camera presenter stays visually consistent within the clip, with clean face and hand rendering. — InVideo output 1.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): The presenter remains consistent through the core ad shots and the branded outro. — invideo output 3.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): The presenter remains consistent through the core ad shots and the branded outro. — invideo output 3.mp4
What changed: Text prompt transformed into Video file
Why it matters / Conclusion: Strong for believable avatar delivery in the right format, but one render still needed QA because the ad can stop early or lose structural polish.
Turns a supplied script into a vertical ad with a realistic on-camera AI presenter. Across the SaaS, physical-product, and app tests, the presenter stayed believable, and the main talking-head shots kept the same presenter identity.
Scene switching and B-roll insertionMixed▾
Feature tested: Scene switching and B-roll insertion
Result: Partial
Verdict: Mixed
Expected behavior: Adds cutaways, product shots, and end-card style scenes instead of relying on a single static talking-head shot. In the tested ads, this sometimes created real pacing and scene changes, though not every render used them consistently.
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): The result stays as one unbroken talking-head shot with no cutaways, CTA end card, logo, or product visual. — InVideo output 1.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): The result stays as one unbroken talking-head shot with no cutaways, CTA end card, logo, or product visual. — InVideo output 1.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): The result includes running B-roll timed to the script, creating a real ad structure instead of a static monologue. — invideo output 2.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): The result includes running B-roll timed to the script, creating a real ad structure instead of a static monologue. — invideo output 2.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): The result uses a phone-use cutaway and a branded green Duo outro card, then returns to the presenter for the CTA. — invideo output 3.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): The result uses a phone-use cutaway and a branded green Duo outro card, then returns to the presenter for the CTA. — invideo output 3.mp4
What changed: Text prompt transformed into Video file
Why it matters / Conclusion: This capability is useful when it appears, but it is not consistent enough to trust without review.
Adds cutaways, product shots, and end-card style scenes instead of relying on a single static talking-head shot. In the tested ads, this sometimes created real pacing and scene changes, though not every render used them consistently.
Prompt-based refinement workflowUntested▾
Feature tested: Prompt-based refinement workflow
Result: Failed
Verdict: Untested
Expected behavior: Provides a browser workflow with a prompt canvas and slate settings around generation, suggesting an interactive surface for iterating on an ad after the initial render. The report did not directly exercise edits or regeneration loops.
Why it matters / Conclusion: Carried forward from prior research, but this report did not exercise edits or regeneration loops directly.
Provides a browser workflow with a prompt canvas and slate settings around generation, suggesting an interactive surface for iterating on an ad after the initial render. The report did not directly exercise edits or regeneration loops.
Text-to-Short-Form Video Generation▾
Feature tested: Text-to-Short-Form Video Generation
Result: Partial
Expected behavior: Turns a text prompt or short script into a structured vertical short with multiple scenes. The candidate cards exercise this on the customer-messages dashboard explainer, the tiny robot intern story, and benchmark prompts about messy support channels and a fictional startup narrative.
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): The final file includes burned-in captions and a voiceover track as part of the rendered short. — InVideoAI_Anchor1_OutputVideo.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): The final file includes burned-in captions and a voiceover track as part of the rendered short. — InVideoAI_Anchor1_OutputVideo.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): INPUT
Observed output: Output artifact (Video file): Finished vertical story in a startup office with a tiny robot intern, showing the prompt-to-video workflow produced a complete short rather than only a static storyboard. — InVideoAI_Anchor2_OutputVideo.mp4
Input artifact: Input artifact (Text prompt): INPUT
Output artifact: Output artifact (Video file): Finished vertical story in a startup office with a tiny robot intern, showing the prompt-to-video workflow produced a complete short rather than only a static storyboard. — InVideoAI_Anchor2_OutputVideo.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): Original human presenter and custom motion graphics were generated for the concept, but the in-scene phone UI text is garbled and the phone design changes across scenes. — InVideoAI_Anchor1_OutputVideo.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): Original human presenter and custom motion graphics were generated for the concept, but the in-scene phone UI text is garbled and the phone design changes across scenes. — InVideoAI_Anchor1_OutputVideo.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): The robot intern stayed visually consistent across the sampled frames, and the startup office setting also remained stable. — InVideoAI_Anchor2_OutputVideo.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): The robot intern stayed visually consistent across the sampled frames, and the startup office setting also remained stable. — InVideoAI_Anchor2_OutputVideo.mp4
What changed: Text prompt transformed into Video file
Why it matters / Conclusion: Works on both benchmark prompts, but the first render was not always complete enough to ship without follow-up.
Turns a text prompt or short script into a structured vertical short with multiple scenes. The candidate cards exercise this on the customer-messages dashboard explainer, the tiny robot intern story, and benchmark prompts about messy support channels and a fictional startup narrative.
Caption, Voice, and Music Assembly▾
Feature tested: Caption, Voice, and Music Assembly
Result: Passed
Expected behavior: Assembles burned-in captions plus voice and music into the final export. The evidence shows these audio/text elements being added or completed during rendering so the short becomes usable.
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): Final export includes burned-in captions and audio, but the workflow required follow-up prompting to get the full package. — InVideoAI_Anchor1_OutputVideo.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): Final export includes burned-in captions and audio, but the workflow required follow-up prompting to get the full package. — InVideoAI_Anchor1_OutputVideo.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Video file): Final export includes burned-in captions and audio for the robot story, after iterative finishing. — InVideoAI_Anchor2_OutputVideo.mp4
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Video file): Final export includes burned-in captions and audio for the robot story, after iterative finishing. — InVideoAI_Anchor2_OutputVideo.mp4
What changed: Text prompt transformed into Video file
Why it matters / Conclusion: The capability works, but not reliably in a single pass.
Assembles burned-in captions plus voice and music into the final export. The evidence shows these audio/text elements being added or completed during rendering so the short becomes usable.
Conversational Video Editing▾
Feature tested: Conversational Video Editing
Result: Partial
Expected behavior: Accepts typed follow-up commands to revise a rendered video after the first render, including changing voice, adding captions or music, and regenerating scenes. The cards cover iterative review-and-fix workflows rather than first-pass creation.
Test case: Text prompt → Text prompt
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Text prompt): Output
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Text prompt): Output
What changed: Text prompt transformed into Text prompt
Test case: Text prompt → Text prompt
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Text prompt): Output
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Text prompt): Output
What changed: Text prompt transformed into Text prompt
Test case: Text prompt → Text prompt
Input type: Text prompt
Input used: Input artifact (Text prompt): Input
Observed output: Output artifact (Text prompt): Output
Input artifact: Input artifact (Text prompt): Input
Output artifact: Output artifact (Text prompt): Output
What changed: Text prompt transformed into Text prompt
Why it matters / Conclusion: Useful for iterative fixes, though the first retry was not always accurate.
Accepts typed follow-up commands to revise a rendered video after the first render, including changing voice, adding captions or music, and regenerating scenes. The cards cover iterative review-and-fix workflows rather than first-pass creation.
Character-Consistent Scene GenerationMostly working▾
Feature tested: Character-Consistent Scene Generation
Result: Partial
Verdict: Mostly working
Expected behavior: Generates scenes that preserve recurring characters and recognizable environments across multiple shots. The evidence focuses on keeping the robot intern and dashboard-explainer worlds coherent while scene details vary.
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): INPUT
Observed output: Output artifact (Video file): The robot intern keeps the same white/cream capsule design with an INTERN badge across the sampled frames, and the open-plan startup office stays visually consistent. — InVideoAI_Anchor2_OutputVideo.mp4
Input artifact: Input artifact (Text prompt): INPUT
Output artifact: Output artifact (Video file): The robot intern keeps the same white/cream capsule design with an INTERN badge across the sampled frames, and the open-plan startup office stays visually consistent. — InVideoAI_Anchor2_OutputVideo.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): INPUT
Observed output: Output artifact (Video file): The recurring phone prop does not stay perfectly consistent across the short, which shows that continuity is better for central characters than for every repeated object. — InVideoAI_Anchor1_OutputVideo.mp4
Input artifact: Input artifact (Text prompt): INPUT
Output artifact: Output artifact (Video file): The recurring phone prop does not stay perfectly consistent across the short, which shows that continuity is better for central characters than for every repeated object. — InVideoAI_Anchor1_OutputVideo.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): INPUT
Observed output: Output artifact (Video file): Photoreal human presenter and custom motion graphics visualizing multiple message sources converging into one AI hub; not stock B-roll. — InVideoAI_Anchor1_OutputVideo.mp4
Input artifact: Input artifact (Text prompt): INPUT
Output artifact: Output artifact (Video file): Photoreal human presenter and custom motion graphics visualizing multiple message sources converging into one AI hub; not stock B-roll. — InVideoAI_Anchor1_OutputVideo.mp4
What changed: Text prompt transformed into Video file
Test case: Text prompt → Video file
Input type: Text prompt
Input used: Input artifact (Text prompt): INPUT
Observed output: Output artifact (Video file): Consistent white/cream robot intern with an INTERN badge in the same open-plan office across the short. — InVideoAI_Anchor2_OutputVideo.mp4
Input artifact: Input artifact (Text prompt): INPUT
Output artifact: Output artifact (Video file): Consistent white/cream robot intern with an INTERN badge in the same open-plan office across the short. — InVideoAI_Anchor2_OutputVideo.mp4
What changed: Text prompt transformed into Video file
Why it matters / Conclusion: Character continuity is a real strength here, but prop continuity is not perfect.
Generates scenes that preserve recurring characters and recognizable environments across multiple shots. The evidence focuses on keeping the robot intern and dashboard-explainer worlds coherent while scene details vary.
Access & pricing
The tested export came from the watermark-free Max tier.
Pricing was reported in the research task; verify live before publishing.
Banner Preview
How the embed badge will look on your site

Embed HTML
Copy this code to your website source
Quick Integration Guide
- 1Copy the HTML code block above.
- 2Paste it into your site's HTML or CMS editor.
- 3Banner appears instantly on your page.
- 4Links back to your tool profile here.
Similar Tools
Discover more AI tools like InVideo AI to enhance your workflow.
Comments (0)
Need a custom AI solution for this use case?
If you are looking to build a custom AI video generation, UGC ad creation, or avatar video production workflow for your business or internal workflow, email us at contact@futuresmart.ai.
Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at collaborate@aidemos.com.