Kapwing icon
video-generator

Kapwing

Good for editable AI shorts and clean vertical cutouts, but first-pass quality can be uneven

Visit Kapwing
Removal onlyNative 1080×1920Multi-person breakdownPaid plan tested
TL;DR — our verdictUpdated September 2026 · 42 test artifacts

Our take

Where it wins
  • You want to turn a text prompt into an editable vertical short with original AI visuals.
  • You plan to polish the draft inside the editor with scene regeneration, timeline changes, subtitles, and audio tweaks.
  • You can tolerate a first pass that may run short, look blurry, or only partially match the script.
Main limitation
  • You need crisp, publication-ready visuals on the first render.
Pricing (verified plans)
Free $0Pro $16/user/mo billed annuallyBusiness $50/user/mo billed annuallyEnterprise Custom
Strongest test artifacts

Our take

Kapwing is strongest when you want an editable AI short or a clean vertical subject cutout that you can refine after the first pass. In the shorts workflow, it produced original AI visuals and narration with useful regeneration controls, but the outputs were blurry, only moderately aligned to the prompt, and shorter than requested. In background removal, it handled simple talking-head footage well and exported watermark-free MP4s, but it never generated a true replacement background and became less reliable on busier scenes.

Demos by use case
Walkthrough of Kapwing's editor and the Remove Background workflow, including the project grid, loading state, and the background-removal panel on busystreet_input.mp4. · From our Remove or Replace Video Backgrounds Using AI ranking →

In-Depth Review

Our detailed analysis of Kapwing — features, performance, and real-world testing.

AD
AI Demos Team
Expert Reviewer
Verified Review

Feature-by-Feature Breakdown

Video Export
Mixed
Test Summary
Feature tested: Video Export
Result: Partial — Mixed

Feature tested: Video Export

Result: Partial

Verdict: Mixed

Expected behavior: Kapwing exports finished clips as downloadable MP4s, including native vertical outputs and watermark-free paid exports. The evidence spans continuous renders and a stitched Oregon coast/beach export that needed spot-checking.

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): Input — oregoncoast_input.mp4.mp4

Observed output: Output artifact (Video file): Native vertical MP4 export with full-bleed framing and no watermark or player UI chrome visible. — kapwing output 1.mp4

Input artifact: Input artifact (Video file): Input — oregoncoast_input.mp4.mp4

Output artifact: Output artifact (Video file): Native vertical MP4 export with full-bleed framing and no watermark or player UI chrome visible. — kapwing output 1.mp4

What changed: Video file transformed into Video file

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): Input — talkinghead_input.mp4.mp4

Observed output: Output artifact (Video file): Native vertical MP4 export with full-bleed framing and no watermark or player UI chrome visible. — kapwing output 2.mp4

Input artifact: Input artifact (Video file): Input — talkinghead_input.mp4.mp4

Output artifact: Output artifact (Video file): Native vertical MP4 export with full-bleed framing and no watermark or player UI chrome visible. — kapwing output 2.mp4

What changed: Video file transformed into Video file

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): Input — busystreet_input.mp4.mp4

Observed output: Output artifact (Video file): Native vertical MP4 export with full-bleed framing and no watermark or player UI chrome visible. — kapwing output 3.mp4

Input artifact: Input artifact (Video file): Input — busystreet_input.mp4.mp4

Output artifact: Output artifact (Video file): Native vertical MP4 export with full-bleed framing and no watermark or player UI chrome visible. — kapwing output 3.mp4

What changed: Video file transformed into Video file

Test case: Text prompt → Video file

Input type: Text prompt

Input used: Input artifact (Text prompt): INPUT

Observed output: Output artifact (Video file): Kapwing exported the short successfully as a vertical MP4, but the render still showed blur and ran much shorter than the requested 30 seconds. — Kapwing_Anchor1_OutputVideo.mp4

Input artifact: Input artifact (Text prompt): INPUT

Output artifact: Output artifact (Video file): Kapwing exported the short successfully as a vertical MP4, but the render still showed blur and ran much shorter than the requested 30 seconds. — Kapwing_Anchor1_OutputVideo.mp4

What changed: Text prompt transformed into Video file

Test case: Text prompt → Video file

Input type: Text prompt

Input used: Input artifact (Text prompt): INPUT

Observed output: Output artifact (Video file): Kapwing exported the story as a vertical MP4, and the report notes that paid users can export without a watermark. — Kapwing_Anchor2_OutputVideo.mp4

Input artifact: Input artifact (Text prompt): INPUT

Output artifact: Output artifact (Video file): Kapwing exported the story as a vertical MP4, and the report notes that paid users can export without a watermark. — Kapwing_Anchor2_OutputVideo.mp4

What changed: Text prompt transformed into Video file

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Observed output: Output artifact (Video file): The edited talking-head project exported successfully as a finished MP4 file. — Kapwing Output 1 - Talking Head with Dead Air.mp4

Input artifact: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Output artifact: Output artifact (Video file): The edited talking-head project exported successfully as a finished MP4 file. — Kapwing Output 1 - Talking Head with Dead Air.mp4

What changed: Video file transformed into Video file

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): Input 2 — Input 2 - Low-Quality Audio & Lighting.mp4

Observed output: Output artifact (Video file): The low-quality webcam-style project also exported successfully as MP4 after AI editing. — Kapwing Output 2 - Low-Quality Audio & Lighting.mp4

Input artifact: Input artifact (Video file): Input 2 — Input 2 - Low-Quality Audio & Lighting.mp4

Output artifact: Output artifact (Video file): The low-quality webcam-style project also exported successfully as MP4 after AI editing. — Kapwing Output 2 - Low-Quality Audio & Lighting.mp4

What changed: Video file transformed into Video file

Test case: Text prompt → Video file

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Video file): The export produced a vertical MP4, but the final render was still visibly blurry and shorter than requested. — Kapwing_AnchorTask1_Dashboard_Output.mp4

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Video file): The export produced a vertical MP4, but the final render was still visibly blurry and shorter than requested. — Kapwing_AnchorTask1_Dashboard_Output.mp4

What changed: Text prompt transformed into Video file

Test case: Text prompt → Video file

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Video file): The export produced a ready-to-upload vertical MP4, but the clip still came out shorter than requested and needed cleanup. — Kapwing_AnchorTask2_RobotIntern_Output - tiny robot.mp4

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Video file): The export produced a ready-to-upload vertical MP4, but the clip still came out shorter than requested and needed cleanup. — Kapwing_AnchorTask2_RobotIntern_Output - tiny robot.mp4

What changed: Text prompt transformed into Video file

Why it matters / Conclusion: Export hygiene is strong overall — the canvas stays native vertical and watermark-free — but the beach test shows you still need to spot-check continuity before publishing.

Kapwing exports finished clips as downloadable MP4s, including native vertical outputs and watermark-free paid exports. The evidence spans continuous renders and a stitched Oregon coast/beach export that needed spot-checking.

video
Native vertical MP4 export with full-bleed framing and no watermark or player UI chrome visible.
video
Native vertical MP4 export with full-bleed framing and no watermark or player UI chrome visible.
video
Native vertical MP4 export with full-bleed framing and no watermark or player UI chrome visible.
text
Export the finished customer-support dashboard short as a vertical social video.
video
Kapwing exported the short successfully as a vertical MP4, but the render still showed blur and ran much shorter than the requested 30 seconds.
text
Export the robot-intern story as a vertical social video.
video
Kapwing exported the story as a vertical MP4, and the report notes that paid users can export without a watermark.
video
The edited talking-head project exported successfully as a finished MP4 file.
video
The low-quality webcam-style project also exported successfully as MP4 after AI editing.
INPUT
INPUT — Exporting the dashboard short after generation and cleanup.
OUTPUT
The export produced a vertical MP4, but the final render was still visibly blurry and shorter than requested.
INPUT
INPUT — Exporting the robot intern short after generation and cleanup.
OUTPUT
The export produced a ready-to-upload vertical MP4, but the clip still came out shorter than requested and needed cleanup.
Bottom Line
Export hygiene is strong overall — the canvas stays native vertical and watermark-free — but the beach test shows you still need to spot-check continuity before publishing.
From our researchRemove or Replace Video Backgrounds Using AIearlier researchGenerate AI Shorts from Text Descriptions — Using Tools That Create Original Visuals, Not Stock Footage
Audio Editing and Cleanup
Test Summary
Feature tested: Audio Editing and Cleanup
Result: Passed

Feature tested: Audio Editing and Cleanup

Result: Passed

Expected behavior: Kapwing provides in-editor audio controls such as waveform editing, trim, speed, volume, auto-level, clean audio, smart cut, enhance voice, split vocals, fades, and noise reduction. The cards show these tools used on generated projects and rough creator footage.

Test case: Text prompt → Image

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Image): The editor screenshot shows extensive audio controls alongside the subtitle tools, including waveform, trim, speed, volume, auto-level volume, and clean-audio options. — kapwing_Anchor2_Manual-Subtitle-Generation.png

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Image): The editor screenshot shows extensive audio controls alongside the subtitle tools, including waveform, trim, speed, volume, auto-level volume, and clean-audio options. — kapwing_Anchor2_Manual-Subtitle-Generation.png

What changed: Text prompt transformed into Image

Test case: Text prompt → Text prompt

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Text prompt): Output

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Text prompt): Output

What changed: Text prompt transformed into Text prompt

Test case: Text prompt → Image

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Image): The editor exposes audio and voice controls alongside subtitle tools, including waveform, trim, speed, volume, auto level volume, clean audio, smart cut, enhance voice, split vocals, and fade controls. — kapwing_Anchor2_Manual-Subtitle-Generation.png

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Image): The editor exposes audio and voice controls alongside subtitle tools, including waveform, trim, speed, volume, auto level volume, clean audio, smart cut, enhance voice, split vocals, and fade controls. — kapwing_Anchor2_Manual-Subtitle-Generation.png

What changed: Text prompt transformed into Image

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Observed output: Output artifact (Video file): The edited talking-head video sounded cleaner and more consistent than the source, with noticeably improved clarity after the AI pass. — Kapwing Output 1 - Talking Head with Dead Air.mp4

Input artifact: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Output artifact: Output artifact (Video file): The edited talking-head video sounded cleaner and more consistent than the source, with noticeably improved clarity after the AI pass. — Kapwing Output 1 - Talking Head with Dead Air.mp4

What changed: Video file transformed into Video file

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): Input 2 — Input 2 - Low-Quality Audio & Lighting.mp4

Observed output: Output artifact (Video file): The low-quality recording was cleaned up with reduced room noise and better volume balance, making the narration easier to follow than the raw input. — Kapwing Output 2 - Low-Quality Audio & Lighting.mp4

Input artifact: Input artifact (Video file): Input 2 — Input 2 - Low-Quality Audio & Lighting.mp4

Output artifact: Output artifact (Video file): The low-quality recording was cleaned up with reduced room noise and better volume balance, making the narration easier to follow than the raw input. — Kapwing Output 2 - Low-Quality Audio & Lighting.mp4

What changed: Video file transformed into Video file

Why it matters / Conclusion: The editor has real audio tools, but the benchmark runs still needed manual attention on music and timing.

Kapwing provides in-editor audio controls such as waveform editing, trim, speed, volume, auto-level, clean audio, smart cut, enhance voice, split vocals, fades, and noise reduction. The cards show these tools used on generated projects and rough creator footage.

INPUT
INPUT — Kapwing editor during post-generation cleanup.
OUTPUT
Output artifact for "Audio Editing and Cleanup" test: The editor screenshot shows extensive audio controls alongside the subtitle tools, including waveform, trim, speed, volume, auto-level volume, and clean-audio options., kapwing_Anchor2_Manual-Subtitle-Generation.png
The editor screenshot shows extensive audio controls alongside the subtitle tools, including waveform, trim, speed, volume, auto-level volume, and clean-audio options.
INPUT
INPUT — Background audio controls visible in the project editor.
OUTPUT
The report says background music and voice tracks were editable, but they still needed manual attention during cleanup.
INPUT
Post-generation editor view with subtitle and audio controls.
image
Output artifact for "Audio Editing and Cleanup" test: The editor exposes audio and voice controls alongside subtitle tools, including waveform, trim, speed, volume, auto level volume, clean audio, smart cut, enhance voice, split vocals, and fade controls., kapwing_Anchor2_Manual-Subtitle-Generation.png
The editor exposes audio and voice controls alongside subtitle tools, including waveform, trim, speed, volume, auto level volume, clean audio, smart cut, enhance voice, split vocals, and fade controls.
video
The edited talking-head video sounded cleaner and more consistent than the source, with noticeably improved clarity after the AI pass.
video
The low-quality recording was cleaned up with reduced room noise and better volume balance, making the narration easier to follow than the raw input.
Bottom Line
The editor has real audio tools, but the benchmark runs still needed manual attention on music and timing.
From our researchearlier researchEdit Videos Using AI — No Editing Skills Required
Video Background Removal
Mixed
Test Summary
Feature tested: Video Background Removal
Result: Partial — Mixed

Feature tested: Video Background Removal

Result: Partial

Verdict: Mixed

Expected behavior: Kapwing can isolate a subject from uploaded video and remove the background to a plain white cutout. It was exercised on an indoor talking-head clip, an outdoor walking clip, and a busy street scene with another person entering.

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): INPUT — oregoncoast_input.mp4.mp4

Observed output: Output artifact (Video file): Background removal is applied to the walking subject, but only for the first half of the file. Around the midpoint, the export hard-cuts back to the untouched beach footage, so this run never produces a replacement scene. — kapwing output 1.mp4

Input artifact: Input artifact (Video file): INPUT — oregoncoast_input.mp4.mp4

Output artifact: Output artifact (Video file): Background removal is applied to the walking subject, but only for the first half of the file. Around the midpoint, the export hard-cuts back to the untouched beach footage, so this run never produces a replacement scene. — kapwing output 1.mp4

What changed: Video file transformed into Video file

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): INPUT — talkinghead_input.mp4.mp4

Observed output: Output artifact (Video file): Clean white cutout throughout with crisp edges and preserved fine detail; this is removal only, not replacement. — kapwing output 2.mp4

Input artifact: Input artifact (Video file): INPUT — talkinghead_input.mp4.mp4

Output artifact: Output artifact (Video file): Clean white cutout throughout with crisp edges and preserved fine detail; this is removal only, not replacement. — kapwing output 2.mp4

What changed: Video file transformed into Video file

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): INPUT — busystreet_input.mp4.mp4

Observed output: Output artifact (Video file): The subject is removed to white, but the cutout breaks down when a second pedestrian enters frame, leaving bleed-through and ghost smudges. — kapwing output 3.mp4

Input artifact: Input artifact (Video file): INPUT — busystreet_input.mp4.mp4

Output artifact: Output artifact (Video file): The subject is removed to white, but the cutout breaks down when a second pedestrian enters frame, leaving bleed-through and ghost smudges. — kapwing output 3.mp4

What changed: Video file transformed into Video file

Why it matters / Conclusion: Strong on the simple indoor talking-head shot, acceptable on the outdoor walk for the processed segment, and much weaker in the busy multi-person scene. Across all three runs, it only removed to white; no replacement scene ever appeared.

Kapwing can isolate a subject from uploaded video and remove the background to a plain white cutout. It was exercised on an indoor talking-head clip, an outdoor walking clip, and a busy street scene with another person entering.

video
Background removal is applied to the walking subject, but only for the first half of the file. Around the midpoint, the export hard-cuts back to the untouched beach footage, so this run never produces a replacement scene.
video
Clean white cutout throughout with crisp edges and preserved fine detail; this is removal only, not replacement.
video
The subject is removed to white, but the cutout breaks down when a second pedestrian enters frame, leaving bleed-through and ghost smudges.
Bottom Line
Strong on the simple indoor talking-head shot, acceptable on the outdoor walk for the processed segment, and much weaker in the busy multi-person scene. Across all three runs, it only removed to white; no replacement scene ever appeared.
From our researchRemove or Replace Video Backgrounds Using AI
Prompt-to-Editable Video Project Creation
Previously observed, not re-tested in this pass.
Test Summary
Feature tested: Prompt-to-Editable Video Project Creation
Result: Partial — Previously observed, not re-tested in this pass.

Feature tested: Prompt-to-Editable Video Project Creation

Result: Partial

Verdict: Previously observed, not re-tested in this pass.

Expected behavior: Kapwing turns natural-language prompts into complete editable vertical video projects with generated scenes, narration, and export-ready structure. The evidence includes the customer-support dashboard concept, the robot-intern story, and an earlier broader AI editing workspace example.

Test case: Text prompt → Video file

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Video file): Kapwing converted the customer-support-dashboard prompt into a complete vertical short, but the first render was only about 15 seconds long instead of 30 seconds. — Kapwing_AnchorTask1_Dashboard_Output.mp4

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Video file): Kapwing converted the customer-support-dashboard prompt into a complete vertical short, but the first render was only about 15 seconds long instead of 30 seconds. — Kapwing_AnchorTask1_Dashboard_Output.mp4

What changed: Text prompt transformed into Video file

Test case: Text prompt → Video file

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Video file): Kapwing converted the robot-intern prompt into a complete vertical short, but the first render was only about 14 seconds long instead of 30 seconds. — Kapwing_AnchorTask2_RobotIntern_Output - tiny robot.mp4

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Video file): Kapwing converted the robot-intern prompt into a complete vertical short, but the first render was only about 14 seconds long instead of 30 seconds. — Kapwing_AnchorTask2_RobotIntern_Output - tiny robot.mp4

What changed: Text prompt transformed into Video file

Why it matters / Conclusion: Useful to keep in mind as prior context, but not re-validated in this background-removal task.

Kapwing turns natural-language prompts into complete editable vertical video projects with generated scenes, narration, and export-ready structure. The evidence includes the customer-support dashboard concept, the robot-intern story, and an earlier broader AI editing workspace example.

INPUT
INPUT — Anchor Task 1: Create a 30-second vertical short about an AI assistant organizing messy customer support messages from email, chat, and WhatsApp into one clean dashboard.
OUTPUT
Kapwing converted the customer-support-dashboard prompt into a complete vertical short, but the first render was only about 15 seconds long instead of 30 seconds.
INPUT
INPUT — Anchor Task 2: Create a 30-second vertical short about a tiny robot intern joining a startup team, making mistakes, and learning to read documentation before asking questions.
OUTPUT
Kapwing converted the robot-intern prompt into a complete vertical short, but the first render was only about 14 seconds long instead of 30 seconds.
Bottom Line
Useful to keep in mind as prior context, but not re-validated in this background-removal task.
From our researchRemove or Replace Video Backgrounds Using AIearlier researchGenerate AI Shorts from Text Descriptions — Using Tools That Create Original Visuals, Not Stock Footage
Narration Script Generation
Test Summary
Feature tested: Narration Script Generation
Result: Passed

Feature tested: Narration Script Generation

Result: Passed

Expected behavior: Kapwing generates a narration script from prompt inputs so the video has a story structure. In the benchmark runs, the script captured the gist of the idea, with moderate scene-level precision and visual sync.

Test case: Text prompt → Video file

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Video file): The generated script conveyed the messy-messages-to-clean-dashboard idea, but the visuals only matched the intended concept moderately well. — Kapwing_AnchorTask1_Dashboard_Output.mp4

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Video file): The generated script conveyed the messy-messages-to-clean-dashboard idea, but the visuals only matched the intended concept moderately well. — Kapwing_AnchorTask1_Dashboard_Output.mp4

What changed: Text prompt transformed into Video file

Test case: Text prompt → Video file

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Video file): The generated script followed the robot-intern storyline, but several scenes still drifted from the narration. — Kapwing_AnchorTask2_RobotIntern_Output - tiny robot.mp4

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Video file): The generated script followed the robot-intern storyline, but several scenes still drifted from the narration. — Kapwing_AnchorTask2_RobotIntern_Output - tiny robot.mp4

What changed: Text prompt transformed into Video file

Why it matters / Conclusion: Useful for getting a narrative draft quickly, but not accurate enough to trust without review.

Kapwing generates a narration script from prompt inputs so the video has a story structure. In the benchmark runs, the script captured the gist of the idea, with moderate scene-level precision and visual sync.

INPUT
INPUT — Anchor Task 1 prompt for the customer support dashboard short.
OUTPUT
The generated script conveyed the messy-messages-to-clean-dashboard idea, but the visuals only matched the intended concept moderately well.
INPUT
INPUT — Anchor Task 2 prompt for the robot intern story short.
OUTPUT
The generated script followed the robot-intern storyline, but several scenes still drifted from the narration.
Bottom Line
Useful for getting a narrative draft quickly, but not accurate enough to trust without review.
From our researchGenerate AI Shorts from Text Descriptions — Using Tools That Create Original Visuals, Not Stock Footageearlier research
AI Visual Generation
Test Summary
Feature tested: AI Visual Generation
Result: Partial

Feature tested: AI Visual Generation

Result: Partial

Expected behavior: Kapwing creates original AI-generated scenes and characters from short-video prompts instead of relying on stock footage. The tested inputs included dashboard-style business scenes and a 3D robot intern.

Test case: Text prompt → Video file

Input type: Text prompt

Input used: Input artifact (Text prompt): INPUT

Observed output: Output artifact (Video file): Kapwing generated original AI visuals for the dashboard concept, but the dashboard interface was blurry and one laptop angle looked visually inconsistent. — Kapwing_Anchor1_OutputVideo.mp4

Input artifact: Input artifact (Text prompt): INPUT

Output artifact: Output artifact (Video file): Kapwing generated original AI visuals for the dashboard concept, but the dashboard interface was blurry and one laptop angle looked visually inconsistent. — Kapwing_Anchor1_OutputVideo.mp4

What changed: Text prompt transformed into Video file

Test case: Text prompt → Video file

Input type: Text prompt

Input used: Input artifact (Text prompt): INPUT

Observed output: Output artifact (Video file): Kapwing generated original 3D-style robot scenes rather than stock footage, but several frames were blurry and one scene did not clearly communicate the intended action. — Kapwing_Anchor2_OutputVideo.mp4

Input artifact: Input artifact (Text prompt): INPUT

Output artifact: Output artifact (Video file): Kapwing generated original 3D-style robot scenes rather than stock footage, but several frames were blurry and one scene did not clearly communicate the intended action. — Kapwing_Anchor2_OutputVideo.mp4

What changed: Text prompt transformed into Video file

Test case: Text prompt → Image

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Image): The dashboard interface is clearly AI-generated, but it is blurry enough that the customer support messages and UI elements are hard to read. — kapwing_Anchor1_Blurry-Dashboard-Interface.png

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Image): The dashboard interface is clearly AI-generated, but it is blurry enough that the customer support messages and UI elements are hard to read. — kapwing_Anchor1_Blurry-Dashboard-Interface.png

What changed: Text prompt transformed into Image

Test case: Text prompt → Image

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Image): The laptop scene is original AI output, but the screen perspective looks awkward and visually inconsistent with the camera angle. — kapwing_Anchor1_Incorrect-Laptop-Screen-Perspective.png

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Image): The laptop scene is original AI output, but the screen perspective looks awkward and visually inconsistent with the camera angle. — kapwing_Anchor1_Incorrect-Laptop-Screen-Perspective.png

What changed: Text prompt transformed into Image

Test case: Text prompt → Image

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Image): The robot scene is original AI output, but the visual does not cleanly match the narration about workplace mistakes. — kapwing_Anchor2_Scene-Does-Not-Match-Narration.png

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Image): The robot scene is original AI output, but the visual does not cleanly match the narration about workplace mistakes. — kapwing_Anchor2_Scene-Does-Not-Match-Narration.png

What changed: Text prompt transformed into Image

Test case: Text prompt → Image

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Image): The monitor content is blurred, so the project-documentation scene loses readability and weakens the story. — kapwing_Anchor2_Blurry-Dashboard-Content.png

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Image): The monitor content is blurred, so the project-documentation scene loses readability and weakens the story. — kapwing_Anchor2_Blurry-Dashboard-Content.png

What changed: Text prompt transformed into Image

Why it matters / Conclusion: Kapwing does generate original visuals, but blur and occasional scene mismatch kept the outputs from feeling polished.

Kapwing creates original AI-generated scenes and characters from short-video prompts instead of relying on stock footage. The tested inputs included dashboard-style business scenes and a 3D robot intern.

text
Customer support dashboard prompt with a modern, simple, slightly futuristic style.
video
Kapwing generated original AI visuals for the dashboard concept, but the dashboard interface was blurry and one laptop angle looked visually inconsistent.
text
Robot intern story prompt with a playful but professional tone.
video
Kapwing generated original 3D-style robot scenes rather than stock footage, but several frames were blurry and one scene did not clearly communicate the intended action.
INPUT
INPUT — Anchor Task 1 customer support dashboard prompt, focusing on the messy dashboard concept.
OUTPUT
Output artifact for "AI Visual Generation" test: The dashboard interface is clearly AI-generated, but it is blurry enough that the customer support messages and UI elements are hard to read., kapwing_Anchor1_Blurry-Dashboard-Interface.png
The dashboard interface is clearly AI-generated, but it is blurry enough that the customer support messages and UI elements are hard to read.
INPUT
INPUT — Anchor Task 1 customer support dashboard prompt, focusing on the laptop scene.
OUTPUT
Output artifact for "AI Visual Generation" test: The laptop scene is original AI output, but the screen perspective looks awkward and visually inconsistent with the camera angle., kapwing_Anchor1_Incorrect-Laptop-Screen-Perspective.png
The laptop scene is original AI output, but the screen perspective looks awkward and visually inconsistent with the camera angle.
INPUT
INPUT — Anchor Task 2 robot intern story prompt, focusing on the mistaken workplace moment.
OUTPUT
Output artifact for "AI Visual Generation" test: The robot scene is original AI output, but the visual does not cleanly match the narration about workplace mistakes., kapwing_Anchor2_Scene-Does-Not-Match-Narration.png
The robot scene is original AI output, but the visual does not cleanly match the narration about workplace mistakes.
INPUT
INPUT — Anchor Task 2 robot intern story prompt, focusing on the project documentation monitor.
OUTPUT
Output artifact for "AI Visual Generation" test: The monitor content is blurred, so the project-documentation scene loses readability and weakens the story., kapwing_Anchor2_Blurry-Dashboard-Content.png
The monitor content is blurred, so the project-documentation scene loses readability and weakens the story.
Bottom Line
Kapwing does generate original visuals, but blur and occasional scene mismatch kept the outputs from feeling polished.
From our researchearlier researchGenerate AI Shorts from Text Descriptions — Using Tools That Create Original Visuals, Not Stock Footage
Automatic Background Music Addition
Test Summary
Feature tested: Automatic Background Music Addition
Result: Passed

Feature tested: Automatic Background Music Addition

Result: Passed

Expected behavior: Kapwing adds background music automatically as part of the AI editing workflow, removing the need for manual music selection in the first pass. The member cards show it working consistently across benchmark inputs as a built-in enhancement.

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Observed output: Output artifact (Video file): Background music was added automatically during the AI edit, alongside captions and cleanup. — Kapwing Output 1 - Talking Head with Dead Air.mp4

Input artifact: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Output artifact: Output artifact (Video file): Background music was added automatically during the AI edit, alongside captions and cleanup. — Kapwing Output 1 - Talking Head with Dead Air.mp4

What changed: Video file transformed into Video file

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): Input 2 — Input 2 - Low-Quality Audio & Lighting.mp4

Observed output: Output artifact (Video file): Background music was also applied automatically to the second benchmark edit. — Kapwing Output 2 - Low-Quality Audio & Lighting.mp4

Input artifact: Input artifact (Video file): Input 2 — Input 2 - Low-Quality Audio & Lighting.mp4

Output artifact: Output artifact (Video file): Background music was also applied automatically to the second benchmark edit. — Kapwing Output 2 - Low-Quality Audio & Lighting.mp4

What changed: Video file transformed into Video file

Test case: Text prompt → Video file

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Video file): The dashboard short included AI narration and background audio, but the voiceover did not fully solve the scene timing issues. — Kapwing_AnchorTask1_Dashboard_Output.mp4

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Video file): The dashboard short included AI narration and background audio, but the voiceover did not fully solve the scene timing issues. — Kapwing_AnchorTask1_Dashboard_Output.mp4

What changed: Text prompt transformed into Video file

Test case: Text prompt → Video file

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Video file): The robot-intern short also included narration, but the voiceover did not fix the scene-to-script mismatch. — Kapwing_AnchorTask2_RobotIntern_Output - tiny robot.mp4

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Video file): The robot-intern short also included narration, but the voiceover did not fix the scene-to-script mismatch. — Kapwing_AnchorTask2_RobotIntern_Output - tiny robot.mp4

What changed: Text prompt transformed into Video file

Test case: Text prompt → Video file

Input type: Text prompt

Input used: Input artifact (Text prompt): INPUT

Observed output: Output artifact (Video file): The export included generated narration and background music, producing a complete short even though the visual quality and runtime were below the requested target. — Kapwing_Anchor1_OutputVideo.mp4

Input artifact: Input artifact (Text prompt): INPUT

Output artifact: Output artifact (Video file): The export included generated narration and background music, producing a complete short even though the visual quality and runtime were below the requested target. — Kapwing_Anchor1_OutputVideo.mp4

What changed: Text prompt transformed into Video file

Test case: Text prompt → Video file

Input type: Text prompt

Input used: Input artifact (Text prompt): INPUT

Observed output: Output artifact (Video file): The export included AI narration and completed the short-form structure, but the final video still needed manual cleanup because subtitles and music were not fully automated in the generation flow. — Kapwing_Anchor2_OutputVideo.mp4

Input artifact: Input artifact (Text prompt): INPUT

Output artifact: Output artifact (Video file): The export included AI narration and completed the short-form structure, but the final video still needed manual cleanup because subtitles and music were not fully automated in the generation flow. — Kapwing_Anchor2_OutputVideo.mp4

What changed: Text prompt transformed into Video file

Why it matters / Conclusion: This worked consistently across both tests and was part of the baseline AI edit rather than a manual add-on.

Kapwing adds background music automatically as part of the AI editing workflow, removing the need for manual music selection in the first pass. The member cards show it working consistently across benchmark inputs as a built-in enhancement.

video
Background music was added automatically during the AI edit, alongside captions and cleanup.
video
Background music was also applied automatically to the second benchmark edit.
INPUT
INPUT — Anchor Task 1 generated short after the first render.
OUTPUT
The dashboard short included AI narration and background audio, but the voiceover did not fully solve the scene timing issues.
INPUT
INPUT — Anchor Task 2 generated short after the first render.
OUTPUT
The robot-intern short also included narration, but the voiceover did not fix the scene-to-script mismatch.
text
Generate the customer-support dashboard short with visuals, voiceover or audio, and captions.
video
The export included generated narration and background music, producing a complete short even though the visual quality and runtime were below the requested target.
text
Generate the robot-intern story with scenes, captions, and audio/voice.
video
The export included AI narration and completed the short-form structure, but the final video still needed manual cleanup because subtitles and music were not fully automated in the generation flow.
Bottom Line
This worked consistently across both tests and was part of the baseline AI edit rather than a manual add-on.
From our researchEdit Videos Using AI — No Editing Skills Required
Subtitle Creation
Test Summary
Feature tested: Subtitle Creation
Result: Partial

Feature tested: Subtitle Creation

Result: Partial

Expected behavior: Kapwing supports subtitle workflows in the editor, including Auto subtitles, Upload SRT/VTT, and starting from scratch. In testing, subtitles were usually a follow-up step rather than a fully automatic part of generation.

Test case: Text prompt → Image

Input type: Text prompt

Input used: Input artifact (Text prompt): INPUT

Observed output: Output artifact (Image): The Subtitles panel exposes Auto subtitles, Upload SRT / VTT, and Start from scratch, confirming that captioning requires a manual editor step after generation. — Kapwing_Anchor2_ManualSubtitleGeneration.png

Input artifact: Input artifact (Text prompt): INPUT

Output artifact: Output artifact (Image): The Subtitles panel exposes Auto subtitles, Upload SRT / VTT, and Start from scratch, confirming that captioning requires a manual editor step after generation. — Kapwing_Anchor2_ManualSubtitleGeneration.png

What changed: Text prompt transformed into Image

Test case: Text prompt → Image

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Image): The subtitle editor exposes Auto subtitles, Upload SRT / VTT, and Start from scratch, confirming that subtitles are handled as an editor step. — kapwing_Anchor2_Manual-Subtitle-Generation.png

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Image): The subtitle editor exposes Auto subtitles, Upload SRT / VTT, and Start from scratch, confirming that subtitles are handled as an editor step. — kapwing_Anchor2_Manual-Subtitle-Generation.png

What changed: Text prompt transformed into Image

Test case: Text prompt → Text prompt

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Text prompt): Output

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Text prompt): Output

What changed: Text prompt transformed into Text prompt

Why it matters / Conclusion: Subtitle support exists, but it behaves like a manual finish-pass rather than a fully automatic final step.

Kapwing supports subtitle workflows in the editor, including Auto subtitles, Upload SRT/VTT, and starting from scratch. In testing, subtitles were usually a follow-up step rather than a fully automatic part of generation.

text
Open the subtitle tools after generating the robot-intern short.
image
Output artifact for "Subtitle Creation" test: The Subtitles panel exposes Auto subtitles, Upload SRT / VTT, and Start from scratch, confirming that captioning requires a manual editor step after generation., Kapwing_Anchor2_ManualSubtitleGeneration.png
The Subtitles panel exposes Auto subtitles, Upload SRT / VTT, and Start from scratch, confirming that captioning requires a manual editor step after generation.
INPUT
INPUT — Post-generation cleanup on the robot intern project.
OUTPUT
Output artifact for "Subtitle Creation" test: The subtitle editor exposes Auto subtitles, Upload SRT / VTT, and Start from scratch, confirming that subtitles are handled as an editor step., kapwing_Anchor2_Manual-Subtitle-Generation.png
The subtitle editor exposes Auto subtitles, Upload SRT / VTT, and Start from scratch, confirming that subtitles are handled as an editor step.
INPUT
INPUT — Generated robot intern video before final publishing.
OUTPUT
The report says subtitles were not added automatically during creation and needed manual generation in the editor.
Bottom Line
Subtitle support exists, but it behaves like a manual finish-pass rather than a fully automatic final step.
From our researchearlier research
Draft Editing and Regeneration
Test Summary
Feature tested: Draft Editing and Regeneration
Result: Partial

Feature tested: Draft Editing and Regeneration

Result: Partial

Expected behavior: Kapwing keeps generated projects editable so users can inspect cuts, edit scripts, adjust timing, rearrange scenes, regenerate individual scenes, and replace visuals or media assets. The benchmark exercised these repair controls on generated projects.

Test case: Video file → Image

Input type: Video file

Input used: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Observed output: Output artifact (Image): The transcript review view exposed AI-selected cuts, but the user still had to inspect the edit to make sure pacing and flow stayed natural. — kapwing-input1-failure-2-transcript-review-required-1.png

Input artifact: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Output artifact: Output artifact (Image): The transcript review view exposed AI-selected cuts, but the user still had to inspect the edit to make sure pacing and flow stayed natural. — kapwing-input1-failure-2-transcript-review-required-1.png

What changed: Video file transformed into Image

Test case: Video file → Image

Input type: Video file

Input used: Input artifact (Video file): Input 2 — Input 2 - Low-Quality Audio & Lighting.mp4

Observed output: Output artifact (Image): Even after the AI pass, the timeline still needed manual refinement to remove remaining imperfections before publication. — kapwing-input2-failure-3-manual-timeline-refinement-required.png

Input artifact: Input artifact (Video file): Input 2 — Input 2 - Low-Quality Audio & Lighting.mp4

Output artifact: Output artifact (Image): Even after the AI pass, the timeline still needed manual refinement to remove remaining imperfections before publication. — kapwing-input2-failure-3-manual-timeline-refinement-required.png

What changed: Video file transformed into Image

Test case: Text prompt → Text prompt

Input type: Text prompt

Input used: Input artifact (Text prompt): INPUT

Observed output: Output artifact (Text prompt): Observation

Input artifact: Input artifact (Text prompt): INPUT

Output artifact: Output artifact (Text prompt): Observation

What changed: Text prompt transformed into Text prompt

Test case: Text prompt → Text prompt

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Text prompt): Output

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Text prompt): Output

What changed: Text prompt transformed into Text prompt

Test case: Text prompt → Text prompt

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Text prompt): Output

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Text prompt): Output

What changed: Text prompt transformed into Text prompt

Test case: Text prompt → Text prompt

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Text prompt): Output

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Text prompt): Output

What changed: Text prompt transformed into Text prompt

Test case: Text prompt → Text prompt

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Text prompt): Output

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Text prompt): Output

What changed: Text prompt transformed into Text prompt

Test case: Text prompt → Text prompt

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Text prompt): Output

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Text prompt): Output

What changed: Text prompt transformed into Text prompt

Test case: Text prompt → Text prompt

Input type: Text prompt

Input used: Input artifact (Text prompt): Input

Observed output: Output artifact (Text prompt): Output

Input artifact: Input artifact (Text prompt): Input

Output artifact: Output artifact (Text prompt): Output

What changed: Text prompt transformed into Text prompt

Why it matters / Conclusion: Useful as a fix-up tool, but the report does not show it magically repairing a bad first pass.

Kapwing keeps generated projects editable so users can inspect cuts, edit scripts, adjust timing, rearrange scenes, regenerate individual scenes, and replace visuals or media assets. The benchmark exercised these repair controls on generated projects.

image
Output artifact for "Draft Editing and Regeneration" test: The transcript review view exposed AI-selected cuts, but the user still had to inspect the edit to make sure pacing and flow stayed natural., kapwing-input1-failure-2-transcript-review-required-1.png
The transcript review view exposed AI-selected cuts, but the user still had to inspect the edit to make sure pacing and flow stayed natural.
image
Output artifact for "Draft Editing and Regeneration" test: Even after the AI pass, the timeline still needed manual refinement to remove remaining imperfections before publication., kapwing-input2-failure-3-manual-timeline-refinement-required.png
Even after the AI pass, the timeline still needed manual refinement to remove remaining imperfections before publication.
text
After generation, adjust scenes, timing, subtitles, music, and replace weak visuals without rebuilding the project.
text
Kapwing exposes a flexible editor workflow: script edits, individual scene regeneration, visual replacement, subtitle changes, background-music adjustment, and timing tweaks are all available after generation. The report also notes that regenerated scenes still suffered from quality limits.
INPUT
INPUT — First-render dashboard short that needed cleanup after generation.
OUTPUT
Users can edit the script, change scene timing, rearrange clips, modify subtitles, change background music, and export the updated video without recreating the whole project.
INPUT
INPUT — First-render robot intern short that needed cleanup after generation.
OUTPUT
The same in-editor workflow applies to the second test: scenes, timing, subtitles, visuals, and audio can all be refined in place.
INPUT
INPUT — A weak scene from the dashboard short that needed improvement.
OUTPUT
Kapwing's editing workflow lets users regenerate individual scenes and replace visuals instead of starting over from scratch.
INPUT
INPUT — A weak scene from the robot intern short that needed improvement.
OUTPUT
The report notes that regenerated visuals still did not always resolve the blur, awkward framing, or prompt mismatch.
INPUT
INPUT — A weak scene after the first render of either benchmark project.
OUTPUT
The report says scenes can be replaced with AI-generated visuals, stock assets, or uploaded media files.
INPUT
INPUT — A scene with blur or framing issues that needed cleanup.
OUTPUT
The source evidence does not show this feature by itself eliminating the visual blur or prompt mismatch.
Bottom Line
Useful as a fix-up tool, but the report does not show it magically repairing a bad first pass.
From our researchGenerate AI Shorts from Text Descriptions — Using Tools That Create Original Visuals, Not Stock Footageearlier research
AI-Assisted Automatic Editing
Test Summary
Feature tested: AI-Assisted Automatic Editing
Result: Partial

Feature tested: AI-Assisted Automatic Editing

Result: Partial

Expected behavior: Kapwing automatically edits raw talking-head and webcam-style recordings by trimming silences, filler words, short pauses, and other unnecessary segments into a cleaner draft. The tested footage produced a better-paced first pass, though not a fully autonomous final edit.

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Observed output: Output artifact (Video file): Kapwing removed many pauses and filler words from the talking-head clip and made the pacing noticeably cleaner, but a few unnecessary pauses still remained and the result still benefited from manual review. — Kapwing Output 1 - Talking Head with Dead Air.mp4

Input artifact: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Output artifact: Output artifact (Video file): Kapwing removed many pauses and filler words from the talking-head clip and made the pacing noticeably cleaner, but a few unnecessary pauses still remained and the result still benefited from manual review. — Kapwing Output 1 - Talking Head with Dead Air.mp4

What changed: Video file transformed into Video file

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): Input 2 — Input 2 - Low-Quality Audio & Lighting.mp4

Observed output: Output artifact (Video file): Kapwing cleaned up the low-quality webcam-style recording by removing many pauses and filler words, but some repeated words and brief pauses were still preserved in the first pass. — Kapwing Output 2 - Low-Quality Audio & Lighting.mp4

Input artifact: Input artifact (Video file): Input 2 — Input 2 - Low-Quality Audio & Lighting.mp4

Output artifact: Output artifact (Video file): Kapwing cleaned up the low-quality webcam-style recording by removing many pauses and filler words, but some repeated words and brief pauses were still preserved in the first pass. — Kapwing Output 2 - Low-Quality Audio & Lighting.mp4

What changed: Video file transformed into Video file

Why it matters / Conclusion: Good at producing a cleaner first draft, but it is not fully autonomous and still leaves some pauses or repeated phrases behind.

Kapwing automatically edits raw talking-head and webcam-style recordings by trimming silences, filler words, short pauses, and other unnecessary segments into a cleaner draft. The tested footage produced a better-paced first pass, though not a fully autonomous final edit.

video
Kapwing removed many pauses and filler words from the talking-head clip and made the pacing noticeably cleaner, but a few unnecessary pauses still remained and the result still benefited from manual review.
video
Kapwing cleaned up the low-quality webcam-style recording by removing many pauses and filler words, but some repeated words and brief pauses were still preserved in the first pass.
Bottom Line
Good at producing a cleaner first draft, but it is not fully autonomous and still leaves some pauses or repeated phrases behind.
From our researchEdit Videos Using AI — No Editing Skills Required
Smart B-Roll Suggestions
Test Summary
Feature tested: Smart B-Roll Suggestions
Result: Partial

Feature tested: Smart B-Roll Suggestions

Result: Partial

Expected behavior: Kapwing can generate contextually relevant Smart B-roll suggestions from the transcript for a project. The user still has to insert, reposition, and adjust the suggested visuals manually.

Test case: Video file → Image

Input type: Video file

Input used: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Observed output: Output artifact (Image): After the AI edit finished, Smart B-roll still had not been inserted into the project, so it remained a separate manual step. — kapwing-input1-failure-1-smart-broll-not-applied.png

Input artifact: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Output artifact: Output artifact (Image): After the AI edit finished, Smart B-roll still had not been inserted into the project, so it remained a separate manual step. — kapwing-input1-failure-1-smart-broll-not-applied.png

What changed: Video file transformed into Image

Test case: Video file → Image

Input type: Video file

Input used: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Observed output: Output artifact (Image): The generated B-roll was relevant to the narration, but it had to be manually positioned and timed to match the spoken content. — kapwing-input1-failure-3-manual-broll-positioning.png

Input artifact: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Output artifact: Output artifact (Image): The generated B-roll was relevant to the narration, but it had to be manually positioned and timed to match the spoken content. — kapwing-input1-failure-3-manual-broll-positioning.png

What changed: Video file transformed into Image

Why it matters / Conclusion: The suggestions are useful, but the timing and placement burden stays on the user.

Kapwing can generate contextually relevant Smart B-roll suggestions from the transcript for a project. The user still has to insert, reposition, and adjust the suggested visuals manually.

image
Output artifact for "Smart B-Roll Suggestions" test: After the AI edit finished, Smart B-roll still had not been inserted into the project, so it remained a separate manual step., kapwing-input1-failure-1-smart-broll-not-applied.png
After the AI edit finished, Smart B-roll still had not been inserted into the project, so it remained a separate manual step.
image
Output artifact for "Smart B-Roll Suggestions" test: The generated B-roll was relevant to the narration, but it had to be manually positioned and timed to match the spoken content., kapwing-input1-failure-3-manual-broll-positioning.png
The generated B-roll was relevant to the narration, but it had to be manually positioned and timed to match the spoken content.
Bottom Line
The suggestions are useful, but the timing and placement burden stays on the user.
From our researchEdit Videos Using AI — No Editing Skills Required
Caption Generation and Styling
Test Summary
Feature tested: Caption Generation and Styling
Result: Passed

Feature tested: Caption Generation and Styling

Result: Passed

Expected behavior: Kapwing generates synchronized captions from spoken narration and lets users style them after generation with changes to fonts, colors, sizing, positioning, and animations. The evidence comes from captioned talking-head style videos where the captions stayed aligned while being customizable.

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Observed output: Output artifact (Video file): Captions were accurately synchronized to the talking-head narration and required little or no correction in the edited output. — Kapwing Output 1 - Talking Head with Dead Air.mp4

Input artifact: Input artifact (Video file): Input 1 — Input 1 - Talking Head with Dead Air.mp4

Output artifact: Output artifact (Video file): Captions were accurately synchronized to the talking-head narration and required little or no correction in the edited output. — Kapwing Output 1 - Talking Head with Dead Air.mp4

What changed: Video file transformed into Video file

Test case: Video file → Video file

Input type: Video file

Input used: Input artifact (Video file): Input 2 — Input 2 - Low-Quality Audio & Lighting.mp4

Observed output: Output artifact (Video file): Captions stayed well synchronized on the low-quality webcam clip and remained accurate enough that no major correction was needed. — Kapwing Output 2 - Low-Quality Audio & Lighting.mp4

Input artifact: Input artifact (Video file): Input 2 — Input 2 - Low-Quality Audio & Lighting.mp4

Output artifact: Output artifact (Video file): Captions stayed well synchronized on the low-quality webcam clip and remained accurate enough that no major correction was needed. — Kapwing Output 2 - Low-Quality Audio & Lighting.mp4

What changed: Video file transformed into Video file

Why it matters / Conclusion: One of Kapwing's strongest capabilities in this benchmark: accurate, synchronized captions with flexible styling control.

Kapwing generates synchronized captions from spoken narration and lets users style them after generation with changes to fonts, colors, sizing, positioning, and animations. The evidence comes from captioned talking-head style videos where the captions stayed aligned while being customizable.

video
Captions were accurately synchronized to the talking-head narration and required little or no correction in the edited output.
video
Captions stayed well synchronized on the low-quality webcam clip and remained accurate enough that no major correction was needed.
Bottom Line
One of Kapwing's strongest capabilities in this benchmark: accurate, synchronized captions with flexible styling control.
From our researchEdit Videos Using AI — No Editing Skills Required

How it scored on the research's own criteria

The 9 evaluation dimensions from our hands-on research on Kapwing, each judged from recorded runs on 3 test inputs — the same verdicts the ranking page ranks on.

held up  partial  failed  not exercised by this input

CriterionVerdictWhat the runs showedPer inputProof
Edge qualityStrong4/5The edges are mostly strong, but the backlit beach shot shows a noticeable halo that keeps this from being top-tier. The indoor talking head is clean enough to show the tool can do very good matte edges, just not consistently across lighting conditions.open proof ↗
Hair and fine detailStrong5/5The indoor talking-head result keeps the hardest small details intact, not just the main silhouette. That level of preservation is strong enough to treat this as a top result, even though it was only demonstrated on one relatively easy scene.open proof ↗
Lighting adaptationMixedNo run produced a replacement background or new scene to compare against the subject, so there is nothing to judge lighting match from. A new test that actually applies a replacement background is needed.open proof ↗
Motion handlingMixed3/5Simple walking motion is handled well, but the moment the scene gets busier the mask starts leaking. Because it can track motion in one scenario yet break down in another, this is clearly mixed rather than fully reliable.open proof ↗
Temporal consistencyMixed3/5One clip stays stable the whole way through, but the beach export is broken by a hard switch to untouched source footage. That makes the tool reliable on one test and clearly unstable on another, so this lands in the middle rather than as a clean pass.open proof ↗
Background optionsMixed3/5The interface clearly offers some background-related tools, but the tested workflow never showed actual background replacement choices or transparent export. That gives it partial support, not full background-option coverage.open proof ↗
Format supportStrong4/5MP4 input and MP4 output are clearly supported, and that worked consistently across all three scenarios. The score stops short of a perfect mark because no limits for other formats, durations, or edge-case upload types were actually demonstrated.open proof ↗
Output resolutionStrong5/5All three exports keep the native vertical canvas and do not show downscaling or watermark issues. That is exactly what a strong export-quality score should look like.open proof ↗
Processing speedMixedNo timed run was recorded for any clip, so there is no basis for a 60-second processing-speed score. A fresh timed test is needed.

Verdicts come verbatim from the study's recorded observations, never re-derived at render; a criterion with no recorded run shows Not exercised — this section cannot invent a score.

Pricing & access

Live plans were listed on Kapwing's pricing page.

Free
$0
10 credits, watermarked exports, videos up to 4 min, 720p
Pro
$16/user/mo billed annually ($192/yr), or $24/mo billed monthly
1,000 credits/mo, no watermark, 4K export, uploads up to 6 GB, videos up to 2 hrs, Brand Kit, 100 GB storage
Business
$50/user/mo billed annually ($600/yr), or $64/mo billed monthly
4,000 credits/mo, everything in Pro + voice cloning, lip-sync, 500 GB storage
Enterprise
Custom
Everything in Business + SSO, dedicated account manager, custom storage/limits

The test notes only say a paid version was used; the exact paid tier was not explicitly logged.

✓ Use This If
You want to turn a text prompt into an editable vertical short with original AI visuals.
You plan to polish the draft inside the editor with scene regeneration, timeline changes, subtitles, and audio tweaks.
You can tolerate a first pass that may run short, look blurry, or only partially match the script.
You need clean background removal for a simple talking-head or lightly moving vertical clip.
You need native vertical, watermark-free MP4 exports.
You can spot-check outputs for stitching, drift, or other integrity bugs.
You are willing to handle any true background replacement in another tool.
✕ Skip This If
You need crisp, publication-ready visuals on the first render.
You need the video to reliably hit a 30-second runtime without manual extension.
You need subtitles to appear automatically during generation.
You need every scene to match technical or story-specific narration exactly.
You need an actual new background or generated replacement scene.
Your footage has multiple people or other busy scene complexity and you need consistently clean segmentation.
You need every export to be a single continuous file without manual verification.
video-generatorvideo-bg-removervideoCreatorEditor
In the short-form video benchmark, Kapwing generated original AI visuals and 3D-style scenes rather than stock footage. The tradeoff was that many frames were blurry or visually weak.
Prompt relevance was only moderate, with the report estimating about 60–70% alignment in both tests. Both videos also came in much shorter than requested: one was around 19 seconds and the other around 14 seconds, even though both prompts asked for 30 seconds.
No. Subtitles were not attached automatically during generation, and the editor's subtitle tools had to be used as a manual follow-up step.
Yes. The editor supports scene regeneration and visual replacement, which is one of Kapwing's strongest workflow advantages. The report also notes that regenerated scenes did not always fix blur or prompt drift.
The benchmark exported vertical MP4 videos. The research also says paid users can export without a watermark.
No. In the background-removal tests, Kapwing only removed the background to solid white. No generated or supplied replacement scene appeared in the tested workflow.
It was excellent on the indoor talking-head clip, with crisp edges and fine detail preserved. On the backlit beach clip the edge got softer and slightly fringed, and on the busy-street clip the cutout broke down once a second pedestrian entered frame.
Yes. The Oregon coast output ran as a processed white-background first half, then hard-cut back to untouched original footage for the second half, so it behaved like two clips concatenated into one file.
The research listed Free, Pro, Business, and Enterprise. Enterprise was the tested plan in one benchmark, while the background-removal research only noted that a paid version was used and did not explicitly log the exact tier.

Banner Preview

How the embed badge will look on your site

Kapwing featured on AI Demos

Embed HTML

Copy this code to your website source

<a target="_blank" href="https://aidemos.com/tools/kapwing?utm_source=kapwing_embed" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> <img src="https://aidemos-website-images.s3.amazonaws.com/featured.png" alt="Kapwing | Featured on AI Demos" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> </a>

Quick Integration Guide

  • 1Copy the HTML code block above.
  • 2Paste it into your site's HTML or CMS editor.
  • 3Banner appears instantly on your page.
  • 4Links back to your tool profile here.
Similar Tools

Similar Tools

Discover more AI tools like Kapwing to enhance your workflow.

🤖
Cutout.Pro
Quick automatic video background removal with solid subject retention, but repeatable edge and crop-stability tradeoffs.
AI Tool
🤖
FlexClip
Browser-based editor with usable AI background removal and MP4 exports, but loose sound design
AI Tool
🤖
Descript
Descript automates editing, cleanup, and effects well enough to speed drafts, but review is still essential.
AI Tool
🤖
Bria.ai
API-first video background removal that isolates subjects well in cluttered scenes, but backlit edges and wrapper exports still need cleanup.
AI Tool
🤖
Media.io
Automatic browser-based background removal and scene swapping for creator videos, with clean isolation but visible edge and shadow limits.
AI Tool
🤖
VEED.io
Browser-based VEED covers captions, avatars, dubbing, and cleanup, but rough edges and limits stay.
AI Tool
🤖
Fotor
Reliable image-to-video and cutouts, but motion is restrained and background replacement is shaky
AI Tool
🤖
InVideo
InVideo AI turns prompts and clips into original videos, but exports still need QA
AI Tool
🤖
Picsart
Free, watermark-free video background removal for single-subject clips, with solid motion tracking but no true scene replacement.
AI Tool
🤖
FutureSmart AI
Fast prompt-to-short generation with script controls and download-ready exports, but detailed scenes and post-render fixes are limited.
AI Tool
🤖
Steve AI
Fast prompt-to-short generation with strong editing controls, but free-plan visuals are image-based and watermarked.
AI Tool
🤖
HeyGen
Fast avatar-led video drafts with strong voice cloning, but visuals and exports still need QA
AI Tool
🤖
Revid.ai
Turns text prompts into complete vertical shorts with AI visuals, voice, captions, and editing, but final export is paywalled.
AI Tool

Comments (0)

Please Log in to join the discussion.

Built by FutureSmart AI — the team behind AI Demos

Need a custom AI solution for this use case?

If you are looking to build a custom video background removal, subject segmentation, or editing tool for your business or internal workflow, email us at contact@futuresmart.ai.

Get a custom build

Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at collaborate@aidemos.com.

Back to Top