Evidence · first-party tested/Best AI Tools to Scrape Web Pages Into Clean Markdown or Structured Data
Extracted the product title, category, and price cleanly, but the result remained incomplete because the dynamic size inventory was not captured.
What was measured
Output Quality
Conceptually matches the prompt and is visually usable.
decisive for this rankingtransformation
Whether the result matches the page and is actually usable is the main outcome the ranking is trying to measure. (3 of 3 judges)
What was given, what came back
Test input: Nike Air Force 1 '07 size options extraction · mixed
Input — what we sent
Input, verbatim
https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111 — Wait for the size selection options to fully render. Extract the product name, price, and a list of all available shoe sizes.
A Nike product page with client-side JavaScript hydration used to test whether a headless scraper waits for dynamic DOM content before extracting product details and all available shoe sizes.
Output — unretouched

Spider Playground showing scraped Nike product results
Also checked on this input — same tool, 4 other criteria
JS DOM Hydration✓ WorkedIt captured the rendered product headline and basic metadata, including the Nike Air Force 1 '07 title, the Men's Shoes label, and the $115 price.JS DOM Hydration✗ FailedIt failed to wait for client-side hydration of the size picker, leaving the size-selection area empty and returning zero available sizing attributes.JS DOM Hydration✗ FailedDoes not wait for client-side hydration long enough; it can capture the title and price but misses the size-selection UI entirely, leaving zero available sizing attributes.Schema Extraction Integrity◐ MixedIt can still extract static metadata cleanly, such as structural description definitions and basic marketing attributes, even when the dynamic transactional section is missing.
Provenance
- Observation
- 4c241c05-718a-4c32-acda-4ce23a4712cf
- Evidence run
- 06e1dbd6-5518-4af8-aa1a-735259a75b4f
- Study
- Scrape Web Pages Into Clean Markdown or Structured Data Using AI
- Research task
- 86b9jm3a3
- Tested at
- Jun 23, 2026
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "spider"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 3 other tools
measured on Output Quality
Firecrawl◐ MixedCaptures the title, pricing, and inventory layout, but the output still includes raw backend code artifacts, media-matrix noise, and uncleaned global link trees.Jina AI Reader◐ MixedIt recovered the product title and the $115 price, but the result was padded with image/link markup and navigation artifacts rather than a clean product-only extract.Skyvern✓ WorkedReturns accurate product data for the hydrated page, including the correct product name Nike Air Force 1 '07 and a long size list; the visible output shows 14 fully readable size pairs from M 7 / W 8.5 through M 14 / W 15.5.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com