The interrogation-room output preserves the reference expression well.
What was measured
Expression accuracy
Whether the emotional tone and facial expression match what was explicitly prompted.
decisive for this rankingtransformation
If the character’s intended emotion or facial expression does not match the prompt, the generated character is not being controlled reliably across scenes. (3 of 3 judges)
What was given, what came back
Test input: Three-quarter face portrait · image
Input — what we sent
Input not captured
This run recorded no prompt or input file for the test, so we cannot show you what produced the result below. Capture gaps are tracked, not hidden.
Three-quarter face reference image with medium-dark skin, tight curly hair in an updo, bindi, and a floral dress. The partial angle and softer lighting make it a harder reference than the frontal portrait and are meant to stress identity consistency.
Why this input is hard
- · Identity preservation with partial face angle
- · Hair texture and updo retention
- · Skin tone fidelity
- · Reference-image difficulty under softer lighting
Output — unretouched

Also checked on this input — same tool, 6 other criteria
Accessory & detail retention⚠ StruggledThe interrogation-room output flattens the dense curly hair substantially and makes the eyebrows look shorter and slightly uneven.Accessory & detail retention⚠ StruggledThe market output lightens the skin tone, removes visible skin marks and scars through smoothing, and makes the lips look less full and more evenly coloured.Identity preservation✓ WorkedThe interrogation-room output stays close to the reference in face shape, skin tone, expression, posture, and overall appearance.Identity preservation◐ MixedThe market output keeps the subject reasonably close overall, but the report still calls the identity match moderate because the skin tone shifts and other facial details change slightly.Scene compliance✓ WorkedThe interrogation-room output follows the clothing, lighting, pose, and environment accurately.Scene compliance✓ WorkedThe market output renders the market environment realistically and follows the prompt well.
Provenance
- Observation
- 9f6dd6f3-855d-4f0a-9535-a5219634595e
- Evidence run
- 1dfb8fa4-f7f0-47dc-911b-7db3af467e9a
- Study
- Generate Consistent AI Characters Across Different Scenes and Poses
- Research task
- 86b96df11
- Tested at
- not recorded
- Source
- first-party
- Evidence state
- verified
- Proof shown
- output only
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "imagineart"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 3 other tools
measured on Expression accuracy
ChatGPT✓ WorkedDelivers the requested cold, guarded, unsmiling expression with the exact vibe the prompt asked for.Gemini✗ FailedThe tool can miss the requested emotion entirely, producing a neutral and emotionless face instead of the prompted angry and guarded look.Leonardo AI✗ FailedThe same neutral-expression failure repeats on a second reference, so the tool still misses the prompt's angry, guarded tone.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com