Leonardo AI icon
image-generator

Leonardo AI Review: Tested Hands-On (2026)

Leonardo AI is polished for reference edits and short clips, but likeness and prompt fidelity drift

Visit Leonardo AI
Single referenceScene-accurateIdentity drift on hard anglesFree version tested
TL;DR — our verdictUpdated September 2026 · 20 test artifacts

Our Take

Where it wins
  • You want a quick single-image-to-video workflow with visible motion.
  • You want short MP4 clips generated from one still image.
  • You want polished reference-guided scene variations from a source image.
Main limitation
  • You need exact prompt adherence or strict facial/object preservation.
Pricing (verified plans)
Free Free
Strongest test artifacts

Feature scores on this page: 4.8/10 (3 scored features)

Our take

Leonardo AI is easy to use for turning one image into polished scene variations, short MP4 clips, and multi-scene photoshoots. Across those workflows, it usually gets the overall look, environment, wardrobe, and composition right, but faces, expressions, and some details can drift or get smoothed out. It fits best when you want a clean visual interpretation from a reference image and can tolerate weaker identity continuity, prompt adherence, and audio.

Demos by use case
Hands-on workflow recording from the Leonardo AI test.

In-Depth Review

Our detailed analysis of Leonardo AI — features, performance, and real-world testing.

AD
AI Demos Team
Expert Reviewer
Verified Review

Feature-by-Feature Breakdown

Reference-Guided Image Generation
Leonardo could restage a reference image into new scenes, but it did not preserve the same face reliably.
4/10
Test Summary
Feature tested: Reference-Guided Image Generation
Result: Failed (4/10) — Leonardo could restage a reference image into new scenes, but it did not preserve the same face reliably.

Feature tested: Reference-Guided Image Generation

Result: Failed (4/10)

Verdict: Leonardo could restage a reference image into new scenes, but it did not preserve the same face reliably.

Expected behavior: Creates new realistic, scene-based portrait images from a single uploaded reference photo while varying scene, pose, outfit, composition, environment, expression, or lighting. The member cards were exercised on café close-ups, desert horse-riding, interrogation-room, restaurant, laptop-at-desk, conference/leadership, podcast-studio, coastal travel, and stage-performance variants.

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Primary reference — input 1.png

Observed output: Output artifact (Image): The tool produced a cozy café portrait with the correct sweater, braid, and relaxed pose, but the face reads as a polished lookalike rather than the same woman, and the hair is smoother and more stylized than the reference. — Leonardo_input1_warm_cafe.jpg

Input artifact: Input artifact (Image): Primary reference — input 1.png

Output artifact: Output artifact (Image): The tool produced a cozy café portrait with the correct sweater, braid, and relaxed pose, but the face reads as a polished lookalike rather than the same woman, and the hair is smoother and more stylized than the reference. — Leonardo_input1_warm_cafe.jpg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Primary reference — input 1.png

Observed output: Output artifact (Image): The tool produced a cinematic desert horse-riding scene that matches the action and scarf details, but the face, proportions, and overall appearance no longer resemble the reference in a meaningful way. — Leonardo_input1_horseride.jpg

Input artifact: Input artifact (Image): Primary reference — input 1.png

Output artifact: Output artifact (Image): The tool produced a cinematic desert horse-riding scene that matches the action and scarf details, but the face, proportions, and overall appearance no longer resemble the reference in a meaningful way. — Leonardo_input1_horseride.jpg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Secondary reference — image.png

Observed output: Output artifact (Image): The tool produced a lively street-market scene with accurate sari styling, natural walking pose, and strong environment rendering, but the face is partially turned away so identity cannot be fully verified. — Leonardo_input2_market.jpg

Input artifact: Input artifact (Image): Secondary reference — image.png

Output artifact: Output artifact (Image): The tool produced a lively street-market scene with accurate sari styling, natural walking pose, and strong environment rendering, but the face is partially turned away so identity cannot be fully verified. — Leonardo_input2_market.jpg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Near-profile stress-test reference. — input 3.webp

Observed output: Output artifact (Image): The rooftop sunset scene delivered the outfit, skyline, and golden-hour lighting, but the face was over-rotated far beyond near-profile, so identity could not be verified. — Leonardo_input3_rooftop.jpg

Input artifact: Input artifact (Image): Near-profile stress-test reference. — input 3.webp

Output artifact: Output artifact (Image): The rooftop sunset scene delivered the outfit, skyline, and golden-hour lighting, but the face was over-rotated far beyond near-profile, so identity could not be verified. — Leonardo_input3_rooftop.jpg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Primary reference — input 1.png

Observed output: Output artifact (Image): The frontal reference produced the best identity match, but the output still softened the face, kept a calm neutral expression instead of angry or guarded, and made the room feel neat rather than harsh. — Leonardo_input1_interrogation.jpg

Input artifact: Input artifact (Image): Primary reference — input 1.png

Output artifact: Output artifact (Image): The frontal reference produced the best identity match, but the output still softened the face, kept a calm neutral expression instead of angry or guarded, and made the room feel neat rather than harsh. — Leonardo_input1_interrogation.jpg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Reference used for the professional conference scene. — INPUT 1.jpg

Observed output: Output artifact (Image): The blazer, glass wall, whiteboard, and tablet all matched the conference/leadership setting, with a believable speaking gesture. — gemini-2.5-flash-image_Professional_Conference_Leadership_Setting_A_realistic_candid_photo_of_the_same_-0.jpg

Input artifact: Input artifact (Image): Reference used for the professional conference scene. — INPUT 1.jpg

Output artifact: Output artifact (Image): The blazer, glass wall, whiteboard, and tablet all matched the conference/leadership setting, with a believable speaking gesture. — gemini-2.5-flash-image_Professional_Conference_Leadership_Setting_A_realistic_candid_photo_of_the_same_-0.jpg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Reference used for the podcast thumbnail scene. — INPUT 2.jpg

Observed output: Output artifact (Image): The mic setup, lighting, jacket, and table layout matched the thumbnail-style podcast prompt. — 9c6ed78e8b584906a4d90b4371cfb541.jpeg

Input artifact: Input artifact (Image): Reference used for the podcast thumbnail scene. — INPUT 2.jpg

Output artifact: Output artifact (Image): The mic setup, lighting, jacket, and table layout matched the thumbnail-style podcast prompt. — 9c6ed78e8b584906a4d90b4371cfb541.jpeg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Reference used for the travel scene. — INPUT 3.jpg

Observed output: Output artifact (Image): The green jacket, coastal overlook, rooftops, and water all fit the travel location prompt. — fe1ca7c13d6443f686b386719e1884d3.jpeg

Input artifact: Input artifact (Image): Reference used for the travel scene. — INPUT 3.jpg

Output artifact: Output artifact (Image): The green jacket, coastal overlook, rooftops, and water all fit the travel location prompt. — fe1ca7c13d6443f686b386719e1884d3.jpeg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Reference used for the speaking-on-stage scene. — INPUT 3.jpg

Observed output: Output artifact (Image): The microphone, spotlight, audience, and dynamic hand gesture matched the stage-performance prompt. — b8405c6dd1484e8983c980326868adb2.jpeg

Input artifact: Input artifact (Image): Reference used for the speaking-on-stage scene. — INPUT 3.jpg

Output artifact: Output artifact (Image): The microphone, spotlight, audience, and dynamic hand gesture matched the stage-performance prompt. — b8405c6dd1484e8983c980326868adb2.jpeg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Reference used for the working-on-laptop scene. — INPUT 1.jpg

Observed output: Output artifact (Image): The window, desk, and mug supported the working-at-laptop scene well, and the pose looked natural. — gemini-2.5-flash-image_Working_on_Laptop_A_candid_mid-morning_photo_of_the_same_person_seated_at_a_worn-0.jpg

Input artifact: Input artifact (Image): Reference used for the working-on-laptop scene. — INPUT 1.jpg

Output artifact: Output artifact (Image): The window, desk, and mug supported the working-at-laptop scene well, and the pose looked natural. — gemini-2.5-flash-image_Working_on_Laptop_A_candid_mid-morning_photo_of_the_same_person_seated_at_a_worn-0.jpg

What changed: Image transformed into Image

Why it matters / Conclusion: Leonardo was good at making attractive scene variations from one image, but not at keeping the same person recognizably intact across those variations.

Creates new realistic, scene-based portrait images from a single uploaded reference photo while varying scene, pose, outfit, composition, environment, expression, or lighting. The member cards were exercised on café close-ups, desert horse-riding, interrogation-room, restaurant, laptop-at-desk, conference/leadership, podcast-studio, coastal travel, and stage-performance variants.

image
Input artifact for "Reference-Guided Image Generation" test: Primary reference, input 1.png
image
Output artifact for "Reference-Guided Image Generation" test: The tool produced a cozy café portrait with the correct sweater, braid, and relaxed pose, but the face reads as a polished lookalike rather than the same woman, and the hair is smoother and more stylized than the reference., Leonardo_input1_warm_cafe.jpg
The tool produced a cozy café portrait with the correct sweater, braid, and relaxed pose, but the face reads as a polished lookalike rather than the same woman, and the hair is smoother and more stylized than the reference.
image
Input artifact for "Reference-Guided Image Generation" test: Primary reference, input 1.png
image
Output artifact for "Reference-Guided Image Generation" test: The tool produced a cinematic desert horse-riding scene that matches the action and scarf details, but the face, proportions, and overall appearance no longer resemble the reference in a meaningful way., Leonardo_input1_horseride.jpg
The tool produced a cinematic desert horse-riding scene that matches the action and scarf details, but the face, proportions, and overall appearance no longer resemble the reference in a meaningful way.
image
Input artifact for "Reference-Guided Image Generation" test: Secondary reference, image.png
image
Output artifact for "Reference-Guided Image Generation" test: The tool produced a lively street-market scene with accurate sari styling, natural walking pose, and strong environment rendering, but the face is partially turned away so identity cannot be fully verified., Leonardo_input2_market.jpg
The tool produced a lively street-market scene with accurate sari styling, natural walking pose, and strong environment rendering, but the face is partially turned away so identity cannot be fully verified.
image
Input artifact for "Reference-Guided Image Generation" test: Near-profile stress-test reference., input 3.webp
Near-profile stress-test reference.
image
Output artifact for "Reference-Guided Image Generation" test: The rooftop sunset scene delivered the outfit, skyline, and golden-hour lighting, but the face was over-rotated far beyond near-profile, so identity could not be verified., Leonardo_input3_rooftop.jpg
The rooftop sunset scene delivered the outfit, skyline, and golden-hour lighting, but the face was over-rotated far beyond near-profile, so identity could not be verified.
image
Input artifact for "Reference-Guided Image Generation" test: Primary reference, input 1.png
image
Output artifact for "Reference-Guided Image Generation" test: The frontal reference produced the best identity match, but the output still softened the face, kept a calm neutral expression instead of angry or guarded, and made the room feel neat rather than harsh., Leonardo_input1_interrogation.jpg
The frontal reference produced the best identity match, but the output still softened the face, kept a calm neutral expression instead of angry or guarded, and made the room feel neat rather than harsh.
image
Input artifact for "Reference-Guided Image Generation" test: Reference used for the professional conference scene., INPUT 1.jpg
Reference used for the professional conference scene.
image
Output artifact for "Reference-Guided Image Generation" test: The blazer, glass wall, whiteboard, and tablet all matched the conference/leadership setting, with a believable speaking gesture., gemini-2.5-flash-image_Professional_Conference_Leadership_Setting_A_realistic_candid_photo_of_the_same_-0.jpg
The blazer, glass wall, whiteboard, and tablet all matched the conference/leadership setting, with a believable speaking gesture.
image
Input artifact for "Reference-Guided Image Generation" test: Reference used for the podcast thumbnail scene., INPUT 2.jpg
Reference used for the podcast thumbnail scene.
image
Output artifact for "Reference-Guided Image Generation" test: The mic setup, lighting, jacket, and table layout matched the thumbnail-style podcast prompt., 9c6ed78e8b584906a4d90b4371cfb541.jpeg
The mic setup, lighting, jacket, and table layout matched the thumbnail-style podcast prompt.
image
Input artifact for "Reference-Guided Image Generation" test: Reference used for the travel scene., INPUT 3.jpg
Reference used for the travel scene.
image
Output artifact for "Reference-Guided Image Generation" test: The green jacket, coastal overlook, rooftops, and water all fit the travel location prompt., fe1ca7c13d6443f686b386719e1884d3.jpeg
The green jacket, coastal overlook, rooftops, and water all fit the travel location prompt.
image
Input artifact for "Reference-Guided Image Generation" test: Reference used for the speaking-on-stage scene., INPUT 3.jpg
Reference used for the speaking-on-stage scene.
image
Output artifact for "Reference-Guided Image Generation" test: The microphone, spotlight, audience, and dynamic hand gesture matched the stage-performance prompt., b8405c6dd1484e8983c980326868adb2.jpeg
The microphone, spotlight, audience, and dynamic hand gesture matched the stage-performance prompt.
image
Input artifact for "Reference-Guided Image Generation" test: Reference used for the working-on-laptop scene., INPUT 1.jpg
Reference used for the working-on-laptop scene.
image
Output artifact for "Reference-Guided Image Generation" test: The window, desk, and mug supported the working-at-laptop scene well, and the pose looked natural., gemini-2.5-flash-image_Working_on_Laptop_A_candid_mid-morning_photo_of_the_same_person_seated_at_a_worn-0.jpg
The window, desk, and mug supported the working-at-laptop scene well, and the pose looked natural.
Bottom Line
Leonardo was good at making attractive scene variations from one image, but not at keeping the same person recognizably intact across those variations.
From our researchearlier researchGenerate AI Photoshoots of Yourself Without a PhotographerGenerate a cinematic AI video from a single imageGenerate Consistent AI Characters Across Different Scenes and Poses
Expression and Mood Control
Leonardo repeatedly missed emotionally intense prompts and defaulted to calm, polished portraits.
3/10
Test Summary
Feature tested: Expression and Mood Control
Result: Failed (3/10) — Leonardo repeatedly missed emotionally intense prompts and defaulted to calm, polished portraits.

Feature tested: Expression and Mood Control

Result: Failed (3/10)

Verdict: Leonardo repeatedly missed emotionally intense prompts and defaulted to calm, polished portraits.

Expected behavior: Renders requested facial expressions and emotional atmosphere in portrait scenes. The tested inputs covered focused work, active speaking, podcast energy, travel glance-back, stage presence, and harsher interrogation-room anger or guardedness across different reference images.

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Input 1 was a full frontal portrait. Prompted scene: interrogation room with formal clothing, harsh overhead lighting, and an angry, guarded expression. — Input 1-2.Input 1

Observed output: Output artifact (Image): This was Leonardo's best identity result from Input 1: the face remained broadly recognizable. But the core emotional instruction failed. The subject looks calm — best-ai-tools-to-generate-consistent-characters-ac-leonardo-input1-interrogation.jpg

Input artifact: Input artifact (Image): Input 1 was a full frontal portrait. Prompted scene: interrogation room with formal clothing, harsh overhead lighting, and an angry, guarded expression. — Input 1-2.Input 1

Output artifact: Output artifact (Image): This was Leonardo's best identity result from Input 1: the face remained broadly recognizable. But the core emotional instruction failed. The subject looks calm — best-ai-tools-to-generate-consistent-characters-ac-leonardo-input1-interrogation.jpg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Input 2 was a 3/4 warm indoor portrait. Prompted scene: the same interrogation-room setup with the same angry, guarded expression request. — Input 2-4.Input 2

Observed output: Output artifact (Image): Leonardo again returned a neutral expression instead of the requested intensity, confirming the miss was not specific to one reference image. The room reads mor — best-ai-tools-to-generate-consistent-characters-ac-leonardo-input2-interrogation.jpg

Input artifact: Input artifact (Image): Input 2 was a 3/4 warm indoor portrait. Prompted scene: the same interrogation-room setup with the same angry, guarded expression request. — Input 2-4.Input 2

Output artifact: Output artifact (Image): Leonardo again returned a neutral expression instead of the requested intensity, confirming the miss was not specific to one reference image. The room reads mor — best-ai-tools-to-generate-consistent-characters-ac-leonardo-input2-interrogation.jpg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Reference used for the focused laptop-work scene. — INPUT 1.jpg

Observed output: Output artifact (Image): The downward gaze and calm expression read as focused and matched the working-on-laptop mood well. — gemini-2.5-flash-image_Working_on_Laptop_A_candid_mid-morning_photo_of_the_same_person_seated_at_a_worn-0.jpg

Input artifact: Input artifact (Image): Reference used for the focused laptop-work scene. — INPUT 1.jpg

Output artifact: Output artifact (Image): The downward gaze and calm expression read as focused and matched the working-on-laptop mood well. — gemini-2.5-flash-image_Working_on_Laptop_A_candid_mid-morning_photo_of_the_same_person_seated_at_a_worn-0.jpg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Reference used for the speaking-in-conference scene. — INPUT 1.jpg

Observed output: Output artifact (Image): The open mouth and hand gesture read as active speaking and fit the leadership/conference prompt. — gemini-2.5-flash-image_Professional_Conference_Leadership_Setting_A_realistic_candid_photo_of_the_same_-0.jpg

Input artifact: Input artifact (Image): Reference used for the speaking-in-conference scene. — INPUT 1.jpg

Output artifact: Output artifact (Image): The open mouth and hand gesture read as active speaking and fit the leadership/conference prompt. — gemini-2.5-flash-image_Professional_Conference_Leadership_Setting_A_realistic_candid_photo_of_the_same_-0.jpg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Reference used for the high-energy podcast thumbnail scene. — INPUT 2.jpg

Observed output: Output artifact (Image): The laugh, head tilt, and eye crinkle clearly carried the high-energy podcast mood. — 9c6ed78e8b584906a4d90b4371cfb541.jpeg

Input artifact: Input artifact (Image): Reference used for the high-energy podcast thumbnail scene. — INPUT 2.jpg

Output artifact: Output artifact (Image): The laugh, head tilt, and eye crinkle clearly carried the high-energy podcast mood. — 9c6ed78e8b584906a4d90b4371cfb541.jpeg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Reference used for the travel glance-back scene. — INPUT 3.jpg

Observed output: Output artifact (Image): The wind-swept hair and half-smile glance back matched the travel mood and pose description. — fe1ca7c13d6443f686b386719e1884d3.jpeg

Input artifact: Input artifact (Image): Reference used for the travel glance-back scene. — INPUT 3.jpg

Output artifact: Output artifact (Image): The wind-swept hair and half-smile glance back matched the travel mood and pose description. — fe1ca7c13d6443f686b386719e1884d3.jpeg

What changed: Image transformed into Image

Test case: Image → Image

Input type: Image

Input used: Input artifact (Image): Reference used for the speaking-on-stage scene. — INPUT 3.jpg

Observed output: Output artifact (Image): The parted lips, torso angle, and gesturing hands conveyed strong stage presence and performance energy. — b8405c6dd1484e8983c980326868adb2.jpeg

Input artifact: Input artifact (Image): Reference used for the speaking-on-stage scene. — INPUT 3.jpg

Output artifact: Output artifact (Image): The parted lips, torso angle, and gesturing hands conveyed strong stage presence and performance energy. — b8405c6dd1484e8983c980326868adb2.jpeg

What changed: Image transformed into Image

Why it matters / Conclusion: Across two different references, Leonardo failed the same emotional prompt in the same way, which points to a tool-level limitation in expression control.

Renders requested facial expressions and emotional atmosphere in portrait scenes. The tested inputs covered focused work, active speaking, podcast energy, travel glance-back, stage presence, and harsher interrogation-room anger or guardedness across different reference images.

INPUT
Input artifact for "Expression and Mood Control" test: Input 1 was a full frontal portrait. Prompted scene: interrogation room with formal clothing, harsh overhead lighting, and an angry, guarded expression., Input 1-2.Input 1

Input 1 was a full frontal portrait. Prompted scene: interrogation room with formal clothing, harsh overhead lighting, and an angry, guarded expression.

image
Output artifact for "Expression and Mood Control" test: This was Leonardo's best identity result from Input 1: the face remained broadly recognizable. But the core emotional instruction failed. The subject looks calm, best-ai-tools-to-generate-consistent-characters-ac-leonardo-input1-interrogation.jpg

This was Leonardo's best identity result from Input 1: the face remained broadly recognizable. But the core emotional instruction failed. The subject looks calm and neutral rather than angry or guarded, the room is too neat, and the lighting lacks the harsh institutional feel described in the prompt.

INPUT
Input artifact for "Expression and Mood Control" test: Input 2 was a 3/4 warm indoor portrait. Prompted scene: the same interrogation-room setup with the same angry, guarded expression request., Input 2-4.Input 2

Input 2 was a 3/4 warm indoor portrait. Prompted scene: the same interrogation-room setup with the same angry, guarded expression request.

image
Output artifact for "Expression and Mood Control" test: Leonardo again returned a neutral expression instead of the requested intensity, confirming the miss was not specific to one reference image. The room reads mor, best-ai-tools-to-generate-consistent-characters-ac-leonardo-input2-interrogation.jpg

Leonardo again returned a neutral expression instead of the requested intensity, confirming the miss was not specific to one reference image. The room reads more like a bright office than an interrogation setting, and the reference's dense natural curls were flattened into straighter, oilier-looking hair.

image
Input artifact for "Expression and Mood Control" test: Reference used for the focused laptop-work scene., INPUT 1.jpg
Reference used for the focused laptop-work scene.
image
Output artifact for "Expression and Mood Control" test: The downward gaze and calm expression read as focused and matched the working-on-laptop mood well., gemini-2.5-flash-image_Working_on_Laptop_A_candid_mid-morning_photo_of_the_same_person_seated_at_a_worn-0.jpg
The downward gaze and calm expression read as focused and matched the working-on-laptop mood well.
image
Input artifact for "Expression and Mood Control" test: Reference used for the speaking-in-conference scene., INPUT 1.jpg
Reference used for the speaking-in-conference scene.
image
Output artifact for "Expression and Mood Control" test: The open mouth and hand gesture read as active speaking and fit the leadership/conference prompt., gemini-2.5-flash-image_Professional_Conference_Leadership_Setting_A_realistic_candid_photo_of_the_same_-0.jpg
The open mouth and hand gesture read as active speaking and fit the leadership/conference prompt.
image
Input artifact for "Expression and Mood Control" test: Reference used for the high-energy podcast thumbnail scene., INPUT 2.jpg
Reference used for the high-energy podcast thumbnail scene.
image
Output artifact for "Expression and Mood Control" test: The laugh, head tilt, and eye crinkle clearly carried the high-energy podcast mood., 9c6ed78e8b584906a4d90b4371cfb541.jpeg
The laugh, head tilt, and eye crinkle clearly carried the high-energy podcast mood.
image
Input artifact for "Expression and Mood Control" test: Reference used for the travel glance-back scene., INPUT 3.jpg
Reference used for the travel glance-back scene.
image
Output artifact for "Expression and Mood Control" test: The wind-swept hair and half-smile glance back matched the travel mood and pose description., fe1ca7c13d6443f686b386719e1884d3.jpeg
The wind-swept hair and half-smile glance back matched the travel mood and pose description.
image
Input artifact for "Expression and Mood Control" test: Reference used for the speaking-on-stage scene., INPUT 3.jpg
Reference used for the speaking-on-stage scene.
image
Output artifact for "Expression and Mood Control" test: The parted lips, torso angle, and gesturing hands conveyed strong stage presence and performance energy., b8405c6dd1484e8983c980326868adb2.jpeg
The parted lips, torso angle, and gesturing hands conveyed strong stage presence and performance energy.
Bottom Line
Across two different references, Leonardo failed the same emotional prompt in the same way, which points to a tool-level limitation in expression control.
From our researchearlier researchGenerate a cinematic AI video from a single imageGenerate Consistent AI Characters Across Different Scenes and PosesGenerate AI Photoshoots of Yourself Without a Photographer
Image-to-Video Generation
Strong — smooth motion and stable rendering
7.5/10
Test Summary
Feature tested: Image-to-Video Generation
Result: Partial (7.5/10) — Strong — smooth motion and stable rendering

Feature tested: Image-to-Video Generation

Result: Partial (7.5/10)

Verdict: Strong — smooth motion and stable rendering

Expected behavior: Animates a still image or static scene into a short playable MP4 with added motion, camera movement, environmental motion, and cinematic transitions. The member cards cover 2D illustrations, stylized 3D scenes, realistic wildlife/photo-style images, and general still-image inputs, including a dedicated 2D cinematic variant.

Test case: Image → Video file

Input type: Image

Input used: Input artifact (Image): Input — leonardo-ai-2d-image-input.png

Observed output: Output artifact (Video file): Generated a short, smooth animated clip from the 2D clover illustration, but the requested hand and environmental details did not fully land, and the result was silent. — leonardo-ai-2d-image-output.mp4

Input artifact: Input artifact (Image): Input — leonardo-ai-2d-image-input.png

Output artifact: Output artifact (Video file): Generated a short, smooth animated clip from the 2D clover illustration, but the requested hand and environmental details did not fully land, and the result was silent. — leonardo-ai-2d-image-output.mp4

What changed: Image transformed into Video file

Test case: Image → Video file

Input type: Image

Input used: Input artifact (Image): Input — leonardo-ai-realistic-image-input.png

Observed output: Output artifact (Video file): Generated a clean animated clip from the realistic tiger photo, but the tiger became more stylized than realistic, showed facial distortion, and had no audio. — leonardo-ai-realistic-image-output.mp4

Input artifact: Input artifact (Image): Input — leonardo-ai-realistic-image-input.png

Output artifact: Output artifact (Video file): Generated a clean animated clip from the realistic tiger photo, but the tiger became more stylized than realistic, showed facial distortion, and had no audio. — leonardo-ai-realistic-image-output.mp4

What changed: Image transformed into Video file

Test case: Image → Video file

Input type: Image

Input used: Input artifact (Image): source image — 3d image.png

Observed output: Output artifact (Video file): The clip preserved the sunset sky and crowd motion, but the birds requested in the prompt never appeared. Two figures on the left footpath break the scene in the first 3 seconds, and the cart rider's face is where the feature loss becomes most obvious. — output 1.mp4

Input artifact: Input artifact (Image): source image — 3d image.png

Output artifact: Output artifact (Video file): The clip preserved the sunset sky and crowd motion, but the birds requested in the prompt never appeared. Two figures on the left footpath break the scene in the first 3 seconds, and the cart rider's face is where the feature loss becomes most obvious. — output 1.mp4

What changed: Image transformed into Video file

Why it matters / Conclusion: The animation holds through the clip — but pause at the 4-second mark to check the hands against what the prompt asked for, and watch the environmental motion with sound on; the silence behind the bush movement is what actually lowers the cinematic feel here.

Animates a still image or static scene into a short playable MP4 with added motion, camera movement, environmental motion, and cinematic transitions. The member cards cover 2D illustrations, stylized 3D scenes, realistic wildlife/photo-style images, and general still-image inputs, including a dedicated 2D cinematic variant.

INPUT
Input artifact for "Image-to-Video Generation" test: Input, leonardo-ai-2d-image-input.png
OUTPUT
Generated a short, smooth animated clip from the 2D clover illustration, but the requested hand and environmental details did not fully land, and the result was silent.
INPUT
Input artifact for "Image-to-Video Generation" test: Input, leonardo-ai-realistic-image-input.png
OUTPUT
Generated a clean animated clip from the realistic tiger photo, but the tiger became more stylized than realistic, showed facial distortion, and had no audio.
image
Input artifact for "Image-to-Video Generation" test: source image, 3d image.png
video
The clip preserved the sunset sky and crowd motion, but the birds requested in the prompt never appeared. Two figures on the left footpath break the scene in the first 3 seconds, and the cart rider's face is where the feature loss becomes most obvious.
Bottom Line
The animation holds through the clip — but pause at the 4-second mark to check the hands against what the prompt asked for, and watch the environmental motion with sound on; the silence behind the bush movement is what actually lowers the cinematic feel here.
From our researchearlier research

How it scored on the research's own criteria

The 7 evaluation dimensions from our hands-on research on Leonardo AI — the same verdicts the ranking page ranks on.

held up  partial  failed  not exercised by this input

CriterionVerdictWhat the runs showedPer inputProof
Consistent patternMixed3/5The results were usable, but repeated hair shifts and the return of skin smoothing show a real pattern drift instead of a stable likeness across scenes.
Identity & LikenessMixed3/5It usually stayed close enough to read as the same person, but the harder references repeatedly lost hair colour and freckle fidelity, so the match was solid rather than exact.
Realism & AI-DetectabilityStrong4/5Most outputs looked convincingly photographic, and the only clear blemish was a slightly retouched conference image, so this lands in the mostly-worked range.
Scene Fidelity / Prompt AdherenceStrong4/5It followed the requested activities, clothing, and settings in nearly every scene, with only one prompt missing a couple of background details and a tighter crop.
Automation levelStrong5/5The job could be run with a simple upload-and-prompt flow, without extra setup beyond basic model and aspect-ratio choices.run-wide
ExportStrong5/5The interface let us download the results directly, so export was fully available.run-wide
Input handlingStrong5/5Every tested reference uploaded cleanly, so there is no sign of rejection or per-image handling trouble.run-wide

Verdicts come verbatim from the study's recorded observations, never re-derived at render; a criterion with no recorded run shows Not exercised — this section cannot invent a score.

Free plan tested

TESTED
Free
Free
150 credits/day; 40 credits per generation.
✓ Use This If
You want a quick single-image-to-video workflow with visible motion.
You want short MP4 clips generated from one still image.
You want polished reference-guided scene variations from a source image.
You want lifestyle, conference, podcast, travel, or stage photos from one reference photo.
You want strong environments, outfits, composition, and prop handling from a reference image.
You can tolerate some face, hair, or prompt drift in exchange for polished outputs.
You prefer a simple upload-prompt-generate workflow with direct downloads.
✕ Skip This If
You need exact prompt adherence or strict facial/object preservation.
You need stable identity continuity across different scenes, poses, or near-profile views.
You need angry, guarded, or otherwise intense expressions to be followed closely.
You need built-in sound or audio in the exported video.
You need exact likeness preservation from harder angles or side-profile references.
You need all requested background details to appear in every scene.
You need undistorted realism rather than a polished interpretation.
image-generatorphoto-studioimageCreatorMarketingFounder
Yes. In the tests, Leonardo AI accepted one static image and generated a short playable MP4 clip from it.
The report describes the clips as short, around 5 seconds. They had visible motion and a cinematic feel, so they read as more than a basic slideshow.
No. The video output was described as silent, with no background sound or audio.
Only partially. Prompt instructions were not followed accurately, and some requested details or elements were missed.
Not reliably in this research. It produced polished scene variations, but the face drifted enough that several outputs felt like lookalikes rather than the same person.
Poorly when the prompt called for intensity. In the interrogation tests, angry or guarded expressions came back calm and neutral instead.
Yes. In this test it produced laptop-at-desk, conference/leadership, podcast-studio, coastal travel, and stage-performance images from the uploaded references.
They were generally accurate. The conference, podcast, travel, and stage scenes matched the requested clothing, props, locations, and actions well, and the report noted natural-looking hands, fingers, wrists, and body poses.
Identity was the weak spot on harder inputs. Hair color drifted on some tests, skin texture was over-smoothed in at least one scene, and the laptop scene was the closest likeness match even though it still missed some requested background items.

Banner Preview

How the embed badge will look on your site

Leonardo AI featured on AI Demos

Embed HTML

Copy this code to your website source

<a target="_blank" href="https://aidemos.com/tools/leonardo-ai?utm_source=leonardo-ai_embed" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> <img src="https://aidemos-website-images.s3.amazonaws.com/featured.png" alt="Leonardo AI | Featured on AI Demos" style="width: 250px; height: 80px; border-radius:4px;" width="250" height="80"> </a>

Quick Integration Guide

  • 1Copy the HTML code block above.
  • 2Paste it into your site's HTML or CMS editor.
  • 3Banner appears instantly on your page.
  • 4Links back to your tool profile here.
Similar Tools

Similar Tools

Discover more AI tools like Leonardo AI to enhance your workflow.

Comments (0)

Please Log in to join the discussion.

Built by FutureSmart AI — the team behind AI Demos

Need a custom AI solution for this use case?

If you are looking to build a custom AI image generation, reference-photo photoshoots, or image styling system for your business or internal workflow, email us at contact@futuresmart.ai.

Get a custom build

Found something inaccurate or missing? We try to keep our AI research accurate and useful. If you found outdated information, an issue, or have a suggestion, email us at collaborate@aidemos.com.

Back to Top