The tool can capture a prompted angry and guarded mood, with direct eye contact and an intense expression in the interrogation shot.
What was measured
Expression accuracy
Whether the emotional tone and facial expression match what was explicitly prompted.
decisive for this rankingtransformation
If the character’s intended emotion or facial expression does not match the prompt, the generated character is not being controlled reliably across scenes. (3 of 3 judges)
What was given, what came back
Test input: Full frontal portrait · image
Input — what we sent

Chatgpt input 1.png
Input not captured
This run recorded no prompt or input file for the test, so we cannot show you what produced the result below. Capture gaps are tracked, not hidden.
Full frontal portrait reference image with fair skin, curly dark hair, bindi, gold jhumka earrings, and a green stone necklace. All features are clearly visible in good natural lighting, making it the easiest identity anchor for the tools.
Why this input is hard
- · Baseline identity preservation
- · Accessory retention
- · Best-case frontal face matching
- · Consistent character reuse across varied scenes
Output — unretouched

Also checked on this input — same tool, 6 other criteria
Identity preservation✗ FailedAgainst the full-frontal reference, the warm cafe output can drift into a different character: it was judged very weak and the report says multiple facial features changed.Identity preservation✗ FailedIn the horse-riding scene, the tool can erase the reference face entirely: the report calls the identity match very weak and says the output is a completely different character with changed face shape, eyes, and eyebrows.Identity preservation✓ WorkedThe plain interrogation setup is the best identity anchor for this input, keeping the eyes, face shape, nose, and overall facial structure close to the reference.Scene compliance✓ WorkedThe horse-riding prompt is followed strongly, including the dusty desert setting, sunset lighting, horse motion, riding costume, gloves, boots, scarf, and believable action pose.Scene compliance✓ WorkedThe interrogation-room prompt is rendered cleanly, with a plain room, metal table, overhead lighting, formal clothing, and uncluttered environment all matching the scene.Scene compliance✓ WorkedThe warm cafe prompt is followed strongly, with the cozy cafe setting, warm lighting, background blur, sweater, braided hairstyle, and natural pose all rendered correctly.
Provenance
- Observation
- d08a8783-fc20-4cf5-8b89-47e25f9f73b7
- Evidence run
- 1dfb8fa4-f7f0-47dc-911b-7db3af467e9a
- Study
- Generate Consistent AI Characters Across Different Scenes and Poses
- Research task
- 86b96df11
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "gemini"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 4 other tools
measured on Expression accuracy
ChatGPT✗ FailedMisses the prompted brave/determined emotional tone and instead outputs a soft neutral expression, leaving the scene with essentially no emotional alignment.ImagineArt✓ WorkedThe interrogation-room output matches the requested serious, guarded expression.Leonardo AI✗ FailedIt does not translate an explicitly angry or guarded prompt into facial expression, defaulting to calm neutrality.Scenario✗ FailedIt fails to map an angry or guarded prompt onto the face; the output stays neutral or calm and even reads with a slight smile.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com