Does not fully realize the requested developer-themed UI, leaving the search animation visually plain.
What was measured
Output Quality
Conceptually matches the prompt and is visually usable.
decisive for this rankingtransformation
The tool is being chosen to produce usable code-based animations, so conceptual match to the prompt and visual usability are core success criteria. (3 of 3 judges)
What was given, what came back
Test input: Search engines basics explainer · text
Input — what we sent
The exact prompt
Create an animation video explaining how search engines work.
A vague, linear educational animation prompt asking the tool to explain how search engines work with minimal guidance. Designed to stress inference, scene structuring, pacing, and coherent visual flow from an underspecified request.
Output — unretouched
Gemini Canvas — Screen Recording 2026-05-04 120930.mp4
Also checked on this input — same tool, 6 other criteria
Code Generation Quality✓ WorkedProduces syntactically correct HTML/CSS/JavaScript on the first attempt for a basic explainer animation.Code Generation Quality✓ WorkedGenerates syntactically correct HTML/CSS/JavaScript on the first attempt; the page ran immediately as a self-contained animation.Input Handling✓ WorkedThe tool can accept a plain-language animation prompt and infer a complete implementation stack and scene structure without needing the framework to be specified.Input Handling✓ WorkedAccepted a vague animation brief and still selected a complete HTML/CSS/JavaScript implementation without the user specifying a framework, showing it can infer a workable structure from underspecified input.Input Handling✓ WorkedAccepts a minimal search-engine prompt and infers a complete 3-step structure—crawling, indexing, and ranking—without needing the framework spelled out.Input Handling✓ WorkedAccepts a plain-language animation prompt and infers the full search workflow without needing a detailed implementation spec.
Provenance
- Observation
- f3c4669a-f5e2-401f-8f82-a14ae83910ea
- Evidence run
- c751a76b-a105-4b3e-9ff3-6c4d6adc9f04
- Study
- Generate Code-based Animations from text inputs
- Research task
- 86b9nnzmy
- Tested at
- Jun 24, 2026
- Source
- first-party
- Evidence state
- observed
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "gemini-canvas"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 7 other tools
measured on Output Quality
AnimG◐ MixedMatches the approved search-engine workflow concept, but the final animation stays visually basic and shows minor clutter even on the low-complexity scene.Antigravity✓ WorkedProduces a visually readable 15-second explainer that presents crawling, indexing, ranking, and serving in sequence.ChatGPT Canvas◐ MixedCan match the requested topic, but the first-pass visual output is only a plain horizontal flowchart with text-only nodes, overflowing labels, and no icons or hierarchy, so it is not visually polished.Claude Artifacts✓ WorkedDelivered a polished animation with smooth motion, clean hierarchy, and a recognizable search-engine workflow.Grok✗ FailedProduces a cluttered, webpage-like result with overlapping elements, cramped spacing, and no clear motion-graphics hierarchy.RemotionVideo✓ WorkedProduces a clean, polished explainer that visually communicates a 3-step crawl → index → rank flow with no overlap or layout clutter.Replit✓ WorkedDelivers clean motion graphics with clear sequencing and visual hierarchy, making the explanation immediately understandable.
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com