Website Scraper for a Whole Domain
- Status
- ...
- Elapsed
- ... ms
- Exit country
- ...
- Browser used
- ...
This console sends a real request and shows you the real response. Five live fetches per session, and the sample targets run as often as you like.
From one seed URL to a structured set
You give a starting address and the rules. The crawler follows links from there, obeys the limits you set, discards what does not match and returns every page it kept with your fields already extracted.
- Depth limit, so a category tree does not turn into the entire internet
- URL patterns to include and exclude, by path, by prefix or by regular expression
- Page count cap per job, so the cost of a crawl is known before it starts
- Deduplication by canonical URL and by content hash, so mirrors and session parameters do not double the bill
- Sitemap discovery first, with link following as the fallback
- One schema applied to every page, with per pattern schemas when sections differ
Where the result goes
| Mode | How it works | Good for |
|---|---|---|
| Synchronous | A single page comes back in the response | Testing a schema, single URLs |
| Crawl job | You start a job and poll it, or receive a webhook when it finishes | Whole domains, nightly refreshes |
| Scheduled crawl | The same job repeats on a schedule and posts each run to your endpoint | Catalogs and listings that change daily |
Crawl jobs are available on Growth and above: 25 running in parallel on Growth, 500 on Scale.
Common crawls
- Catalog extraction across a manufacturer or distributor site, feeding product catalog enrichment.
- Content inventories before a migration, with every URL, title and canonical in one table.
- Documentation and knowledge base collection for an internal search index.
- Listing sites where the useful data is spread over thousands of detail pages.
- Regular refreshes of a competitor site, where the change between runs is the finding.
How the crawler behaves on a target site
Concurrency is paced per host rather than fired all at once, retries back off, and a crawl obeys the page count and depth you set instead of expanding on its own. Requests come from the country you choose, so a regional site is read the way its own visitors read it. The proxy and unblocking layer behind this is described on the web scraping proxy page, and the legal boundaries are in is web scraping legal.
The request behind it, with every parameter, the page limit per plan and the webhook payload, is documented on the website crawler API page.