Evidence · first-party tested

Expression accuracy is a hard ceiling: across two different reference images, intense interrogation prompts still produced flat, neutral-to-unsmiling faces instead of the requested angry/guarded intensity, and the report rates this 4/10.

✗ Failed4 / 10Scenario
Output — unretouched
tool output
Output — unretouched
tool output
Output — unretouched
tool output
The observation

Expression accuracy is a hard ceiling: across two different reference images, intense interrogation prompts still produced flat, neutral-to-unsmiling faces instead of the requested angry/guarded intensity, and the report rates this 4/10.

Criterion: Expression accuracy ·
Query this
get_evidence({
  tool: "scenario"
})
MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Real outputs, no retouching · every cell queryable via API & MCP · aidemos.com