The Extract API scrapes any web page to clean Markdown, HTML or plain text through residential proxies with real-browser TLS fingerprints. It uses a fast TLS tier by default and only spins up a headless browser when a page is bot-challenged.
The auto engine resolves most pages on the fast TLS tier, retries on a fresh exit IP if blocked, and escalates to a stealth headless browser only on a real challenge — maximum speed, minimum cost.
Pull specific fields with CSS selectors, or describe what you want in natural language and let AI extraction return structured JSON. Mobile emulation, screenshots and page actions are supported.
Discover a site's URLs with map, crawl a whole site to Markdown per page, or scrape many URLs asynchronously with batch — all through the same residential-proxy pipeline.