Skip to content
WebScrap

Website Scraper for a Whole Domain

One page is a request. A whole site is a crawl: discovery, depth, filters, deduplication and one schema applied to everything worth keeping.
See pricing

Run a real request

Render
Extract
Status
...
Elapsed
... ms
Exit country
...
Browser used
...

This console sends a real request and shows you the real response. Five live fetches per session, and the sample targets run as often as you like.

From one seed URL to a structured set

You give a starting address and the rules. The crawler follows links from there, obeys the limits you set, discards what does not match and returns every page it kept with your fields already extracted.

  • Depth limit, so a category tree does not turn into the entire internet
  • URL patterns to include and exclude, by path, by prefix or by regular expression
  • Page count cap per job, so the cost of a crawl is known before it starts
  • Deduplication by canonical URL and by content hash, so mirrors and session parameters do not double the bill
  • Sitemap discovery first, with link following as the fallback
  • One schema applied to every page, with per pattern schemas when sections differ

Where the result goes

ModeHow it worksGood for
SynchronousA single page comes back in the responseTesting a schema, single URLs
Crawl jobYou start a job and poll it, or receive a webhook when it finishesWhole domains, nightly refreshes
Scheduled crawlThe same job repeats on a schedule and posts each run to your endpointCatalogs and listings that change daily

Crawl jobs are available on Growth and above: 25 running in parallel on Growth, 500 on Scale.

Common crawls

  • Catalog extraction across a manufacturer or distributor site, feeding product catalog enrichment.
  • Content inventories before a migration, with every URL, title and canonical in one table.
  • Documentation and knowledge base collection for an internal search index.
  • Listing sites where the useful data is spread over thousands of detail pages.
  • Regular refreshes of a competitor site, where the change between runs is the finding.

How the crawler behaves on a target site

Concurrency is paced per host rather than fired all at once, retries back off, and a crawl obeys the page count and depth you set instead of expanding on its own. Requests come from the country you choose, so a regional site is read the way its own visitors read it. The proxy and unblocking layer behind this is described on the web scraping proxy page, and the legal boundaries are in is web scraping legal.

The request behind it, with every parameter, the page limit per plan and the webhook payload, is documented on the website crawler API page.