Evidence · first-party tested/Best AI Tools to Scrape Web Pages Into Clean Markdown or Structured Data
The extraction completed, but the screen-capture recorder fell out of sync and froze on an early page state, so runtime observability degraded.
What was measured
Interaction Stability
How reliably the tool executes through dynamic loading and interaction steps without losing sync or failing at runtime.
decisive for this rankingcapability
Reliably completing dynamic loading and interaction steps is part of successfully scraping pages, not just convenience. (3 of 3 judges)
What was given, what came back
Test input: Nike Air Force 1 '07 size options extraction · mixed
Input — what we sent
Input, verbatim
https://www.nike.com/t/air-force-1-07-mens-shoes-jBrhbr/CW2288-111 — Wait for the size selection options to fully render. Extract the product name, price, and a list of all available shoe sizes.
A Nike product page with client-side JavaScript hydration used to test whether a headless scraper waits for dynamic DOM content before extracting product details and all available shoe sizes.
Output — unretouched
No output artifact
The verdict rests on the tester's written observation alone — no file was captured for this cell.
Also checked on this input — same tool, 7 other criteria
JS DOM Hydration✓ WorkedThe tool can wait for client-side hydration and capture the rendered product state, including a populated size-selection grid; in this run it extracted a structured schema of the Nike Air Force 1 '07 size options and showed multiple size variants rather than an empty shell.JS DOM Hydration◐ MixedThe extraction pipeline recovered the hydrated Nike size data, but the screen-recording/visual trace subsystem was out of sync and froze on an initial page view, so the captured recording did not reflect the final dynamic state.JS DOM Hydration✓ WorkedWaits for the client-rendered product page to hydrate and extracts the size grid instead of stopping at the initial shell.Output Quality✓ WorkedReturns accurate product data for the hydrated page, including the correct product name Nike Air Force 1 '07 and a long size list; the visible output shows 14 fully readable size pairs from M 7 / W 8.5 through M 14 / W 15.5.Schema Extraction Integrity✓ WorkedCan accurately preserve a hydrated page’s structured output, extracting a complete schema of all 22 shoe sizes without corrupting the requested JSON structure.Schema Extraction Integrity✓ WorkedAccurately extracts hydrated product-size data into structured output; the visible result enumerates multiple size pairs, including M 7 / W 8.5, M 7.5 / W 9, M 8 / W 9.5, and through M 15 / W 16.5 in the shown excerpt.Schema Extraction Integrity✓ WorkedOutputs a structured size list with many men/women pairs, with the visible payload spanning at least M 6 / W 7.5 through M 14 / W 15.5.
Provenance
- Observation
- 55dd439d-f92d-466b-8207-4755b499332d
- Evidence run
- 06e1dbd6-5518-4af8-aa1a-735259a75b4f
- Study
- Scrape Web Pages Into Clean Markdown or Structured Data Using AI
- Research task
- 86b9jm3a3
- Tested at
- Jun 23, 2026
- Source
- first-party
- Evidence state
- observed
- Proof shown
- input only
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "skyvern"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 2 other tools
measured on Interaction Stability
Firecrawl✓ WorkedReliably waited for client-side hydration and captured the complete size set, from M 5 / W 6.5 through M 18 / W 19.5.Jina AI Reader✗ FailedIt did not wait for the client-side size selector to hydrate; the response shows the product title and $115 price but no concrete size inventory values.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com