# Web Scraping & Extract API – Markdown, HTML & AI Extraction | Quantum Proxies

Scrape any page to Markdown/HTML/JSON through residential proxies with real-browser TLS fingerprints. Structured and AI extraction, crawl, map and batch.

The Extract API scrapes any web page to clean Markdown, HTML or plain text through residential proxies with real-browser TLS fingerprints. It uses a fast TLS tier by default and only spins up a headless browser when a page is bot-challenged.

## Smart engine

The auto engine resolves most pages on the fast TLS tier, retries on a fresh exit IP if blocked, and escalates to a stealth headless browser only on a real challenge — maximum speed, minimum cost.

## Structured & AI extraction

Pull specific fields with CSS selectors, or describe what you want in natural language and let AI extraction return structured JSON. Mobile emulation, screenshots and page actions are supported.

## Crawl, map & batch

Discover a site's URLs with map, crawl a whole site to Markdown per page, or scrape many URLs asynchronously with batch — all through the same residential-proxy pipeline.

## Sources

- [CommonMark Spec](https://spec.commonmark.org/0.31.2/)
- [The /llms.txt file (llmstxt.org)](https://llmstxt.org/)
- [RFC 8259: The JSON Data Interchange Format](https://www.rfc-editor.org/rfc/rfc8259.html)
- [Web scraping (Wikipedia)](https://en.wikipedia.org/wiki/Web_scraping)
- [Intro to How Structured Data Markup Works (Google Search Central)](https://developers.google.com/search/docs/appearance/structured-data/intro-structured-data)
