Evidence · first-party tested/Best AI Tools to Scrape Web Pages Into Clean Markdown or Structured Data
It fails to wait for client-side hydration on the size-selector grid: the response leaves the selector as empty layout nodes and instead spills the site-wide international menu and regional index into the output.
What was measured
JS DOM Hydration
How well the tool waits for and captures content rendered by client-side JavaScript on dynamic pages.
decisive for this rankingtransformation
Capturing JavaScript-rendered content is a core scraping capability for modern web pages. (3 of 3 judges)
What was given, what came back
Test input: Nike Air Force 1 '07 size options extraction · mixed
Input — what we sent
Input, verbatim
https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111 — Wait for the size selection options to fully render. Extract the product name, price, and a list of all available shoe sizes.
A Nike product page with client-side JavaScript hydration used to test whether a headless scraper waits for dynamic DOM content before extracting product details and all available shoe sizes.
Output — unretouched

Jina Reader extracting a Nike product page
Also checked on this input — same tool, 3 other criteria
Interaction Stability✗ FailedIt did not wait for the client-side size selector to hydrate; the response shows the product title and $115 price but no concrete size inventory values.Output Quality◐ MixedIt recovered the product title and the $115 price, but the result was padded with image/link markup and navigation artifacts rather than a clean product-only extract.Schema Extraction Integrity◐ MixedCan still extract some static fields correctly on a dynamic product page, including the SEO header and price markers, but it corrupts the requested product-specific output by substituting broad site-directory text for the size data.
Provenance
- Observation
- ee913a05-fce4-4910-b38e-711b6e632b06
- Evidence run
- 06e1dbd6-5518-4af8-aa1a-735259a75b4f
- Study
- Scrape Web Pages Into Clean Markdown or Structured Data Using AI
- Research task
- 86b9jm3a3
- Tested at
- Jun 23, 2026
- Source
- first-party
- Evidence state
- verified
- Proof shown
- input + output shown
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "jina-ai-reader"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 3 other tools
measured on JS DOM Hydration
Firecrawl◐ MixedCaptures hydrated product content but leaves substantial non-content noise in the output, including raw backend code artifacts and raw media-attachment matrices.Skyvern✓ WorkedWaits for the client-rendered product page to hydrate and extracts the size grid instead of stopping at the initial shell.Spider✓ WorkedIt captured the rendered product headline and basic metadata, including the Nike Air Force 1 '07 title, the Men's Shoes label, and the $115 price.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com