Evidence · first-party tested/Best AI Tools to Scrape Web Pages Into Clean Markdown or Structured Data
It fails to wait for client-side hydration on a dynamic product page: the size selector grid comes back as empty layout nodes while the extractor instead pulls the site's international menu and regional index.
What was measured
JS DOM Hydration
How well the tool waits for and captures content rendered by client-side JavaScript on dynamic pages.
decisive for this rankingtransformation
Capturing JavaScript-rendered content is a core scraping capability for modern web pages. (3 of 3 judges)
What was given, what came back
Test input: Glassdoor software engineer jobs behind sign-in modal · mixed · group: web-scraping-benchmark
Input — what we sent
Input, verbatim
https://www.glassdoor.com/Job/software-engineer-jobs-SRCH_KO0,17.htm — Dismiss any immediate sign-in or signup modal overlays that block the view. Once cleared, extract the top 5 job listings, including job title, company name, location, and the short summary snippet.
A Glassdoor jobs listing page protected by Cloudflare and a sign-in/interstitial overlay, used to test proxy evasion, anti-bot handling, and the ability to dismiss blocking modal UI before extracting listings.
Output — unretouched
No output artifact
The verdict rests on the tester's written observation alone — no file was captured for this cell.
Also checked on this input — same tool, 14 other criteria
Interaction Stability✓ WorkedCan process a protected page through edge security without being dropped, successfully bypassing the firewall and recovering plain-text markers from the page flow.Noise Filtering◐ MixedIt can recover the page text layer, but the extraction still leaves job data interleaved with French and German translation strings and header redirect text, so heavy post-processing cleanup is still required.Output Quality✗ FailedWhen blocked, it emits an anti-bot 'Humans only' text wall in multiple languages rather than job listings, so the returned text is not usable as a listing extract.Output Quality◐ MixedIts output quality is partial on the hydrated ecommerce page: it preserves the SEO header and static price markers, but replaces the transactional product data with a giant global link directory.Output Quality✗ FailedWhen this run collapses, the extractor returns only a standard HTTP 404 text string wrapped inside a global website layout frame, leaving the primary output unusable for data compilation.Proxy Evasion⚠ StruggledAlthough the anti-bot wall is bypassed, the recovered Glassdoor output is still a raw DOM dump with sign-in notices, framework noise, and header redirects interleaved, so the target listings require heavy downstream cleanup.Proxy Evasion✓ WorkedThe tool can bypass standard edge security filters on the Glassdoor target and is not dropped by the firewall checks in this run sequence.Proxy Evasion✓ WorkedThe backend can bypass a standard Glassdoor "Humans only" interstitial and return page text in about 3.6 seconds, indicating that basic anti-bot and proxy barriers were cleared in this run.Proxy Evasion✗ FailedIt failed to get past the anti-bot barrier; the 3.6 s response is the "Humans only" block page instead of the target job listings.Schema Extraction Integrity✗ FailedCan break entirely when URL handling loops or duplicates the target path: the report says the primary output became completely useless for data compilation after an address-resolution bug.Visual Spatial Awareness✗ FailedIt fails to isolate the content region and instead emits the sign-in overlay text in at least 5 languages (English, French, German, Dutch, Spanish).Visual Spatial Awareness✓ WorkedCan isolate basic boilerplate elements from a broken page render: it still extracted header navigation and privacy-disclosure text into markdown even when the target resolved into a 404 layout.Visual Spatial Awareness✗ FailedFails to structurally filter interstitial noise, leaving raw framework and security-page text instead of isolating meaningful job content.Visual Spatial Awareness⚠ StruggledProduces poor structural filtering on guarded pages: the recovered job text is heavily interleaved with framework noise, sign-in alerts, and header redirects, requiring substantial post-processing to clean.
Provenance
- Observation
- 9ea0b03c-f458-4cb0-827b-c31ddd8edde7
- Evidence run
- 06e1dbd6-5518-4af8-aa1a-735259a75b4f
- Study
- Scrape Web Pages Into Clean Markdown or Structured Data Using AI
- Research task
- 86b9jm3a3
- Tested at
- Jun 23, 2026
- Source
- first-party
- Evidence state
- observed
- Proof shown
- input only
- Cost / latency
- not captured
- Repeat run
- not captured
- Tester
- not captured
The last three rows are honest blanks, not placeholders — our capture has no field for them yet.
Query this
get_evidence({
tool: "jina-ai-reader",
scenario: "web-scraping-benchmark"
})MCP · mcp.aidemos.com/api/mcp
Free with attribution.
Same input, same check — 0 other tools
measured on JS DOM Hydration
No other tool was measured on this criterion for this input.
This evidence is published in
Real inputs and real outputs, no retouching · every cell queryable via API & MCP · aidemos.com