For the low-complexity search-engine explainer, the animation stays conceptually aligned with crawling, indexing, ranking, and results display, but the visuals are basic, flat, and show minor overlap or clutter.
What was measured
Output Quality
Conceptually matches the prompt and is visually usable.
decisive for this rankingtransformation
The tool is being chosen to produce usable code-based animations, so conceptual match to the prompt and visual usability are core success criteria. (3 of 3 judges)
What was given, what came back
Test input: Search engines basics explainer · text
Input — what we sent
The exact prompt
Create an animation video explaining how search engines work.
A vague, linear educational animation prompt asking the tool to explain how search engines work with minimal guidance. Designed to stress inference, scene structuring, pacing, and coherent visual flow from an underspecified request.
Output — unretouched
AnimG — 20260419_090730_SearchEngineWorkflow.mp4
Also checked on this input — same tool, 3 other criteria
Code Generation Quality✓ WorkedGenerates syntactically correct Manim code on the first pass.Export Experience✓ WorkedFinishes by producing an MP4 inside the web app and surfacing the result in a built-in video player, so getting the rendered video is straightforward.Export Experience◐ MixedGetting an MP4 is straightforward, but the free output is watermarked and the UI prompts an upgrade for watermark-free export.
Provenance
- Observation
- c25c8cdf-0d93-46a1-a835-4cc0824146f1
- Evidence run
- c751a76b-a105-4b3e-9ff3-6c4d6adc9f04
- Study
- Generate Code-based Animations from text inputs
- Research task
- 86b9nnzmy
- Tested at
- Jun 24, 2026
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "animg"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 7 other tools
measured on Output Quality
Antigravity✓ WorkedProduces a visually readable 15-second explainer that presents crawling, indexing, ranking, and serving in sequence.ChatGPT Canvas◐ MixedCan match the requested topic, but the first-pass visual output is only a plain horizontal flowchart with text-only nodes, overflowing labels, and no icons or hierarchy, so it is not visually polished.Claude Artifacts✓ WorkedDelivered a polished animation with smooth motion, clean hierarchy, and a recognizable search-engine workflow.Gemini Canvas◐ MixedKeeps the animation inside a macOS-style tab frame with a very small aspect ratio, which makes the result feel presentation-like rather than like a full motion-graphics video.Grok✗ FailedProduces output that reads more like an interactive webpage than a video-like animation, with cramped spacing, overlapping elements, weak visual hierarchy, and visible controls.RemotionVideo✓ WorkedProduces a clean, polished explainer that visually communicates a 3-step crawl → index → rank flow with no overlap or layout clutter.Replit✓ WorkedDelivers clean motion graphics with clear sequencing and visual hierarchy, making the explanation immediately understandable.
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com