Skip to content
WebScrap

Website Crawler API to Crawl a Whole Site into JSON

Give the web crawler API one URL, a depth and the fields you want. It follows the links on that site, fetches every page through rotating proxies with a real browser when needed, and returns each page as clean JSON. You pay 0.98 USD per 1,000 successful pages on Starter, less on larger plans.

See pricing

Run a real request

Render
Extract
Status
...
Elapsed
... ms
Exit country
...
Browser used
...

This console sends a real request and shows you the real response. Five live fetches per session, and the sample targets run as often as you like.

One request in, a whole site of structured pages out

A website crawler API finds the pages for you. You send a seed URL such as a category or a sitemap index page, set how many links deep to go and which paths to keep, and the API returns every matching page with the fields you asked for. No URL list to build first, no queue, no headless browser fleet to run.

Most teams that look for a crawler already have a scraper that works on one page. The problem is everything around it. Discovering the 1,800 product pages behind 40 category pages. Keeping the crawl on the right host and out of the cart, the login and the faceted filter URLs that multiply forever. Retrying the pages that come back as a challenge. Knowing when the job is finished so the next step in the pipeline can start.

The crawl endpoint handles that loop. Each page is an ordinary request through the same proxy pool, retries and optional browser rendering as the web scraping API, so a crawled page costs exactly what a single scraped page costs, and a failed one costs nothing.

Web crawler API options compared

Rules and rates as published by each provider, checked on October 2, 2026. Prices change, so confirm them on the vendor's own page before you budget.

OptionWhat it doesProtected sitesHow it bills
Cloudflare Browser Rendering crawl endpointCrawls from a start URL, returns HTML, Markdown or JSONIdentifies itself as a bot, follows robots.txt and AI Crawl Control by default, cannot pass Cloudflare bot detection or captchasBrowser hours, 10 a month on Workers Paid then 0.09 USD an hour
Firecrawl crawlCrawls and returns Markdown, or JSON per pageManaged proxies1 credit per page, 5 with JSON extraction, and a 403 or 404 from the target still costs a credit
Browserbase Fetch APIFetches one URL per call, no link discoveryProxies and captcha solving on paid plans1 USD per 1,000 calls, 4 with proxies, 5 requests a second per project
Scrapy on your own serversAny crawl logic you writeOnly what you build or buy separatelyServers, proxies and the engineer who keeps it running
WebScrap crawlCrawls one host to depth 5 and returns your fields as JSONRotating datacenter and residential proxies, browser rendering, automatic retriesOne request per successful page, 0.98 USD per 1,000 on Starter, 0.60 on Growth, 0.33 on Scale

Where the others win, honestly. Cloudflare's endpoint is the cheapest way to crawl sites that welcome bots, such as your own documentation. Firecrawl is built around Markdown for language models and adds map and search endpoints we do not have; our Firecrawl pricing breakdown has the per page numbers. Scrapy gives total control if you already run the infrastructure. We fit the team that needs typed fields from a site that pushes back, on a schedule, without running browsers.

Start a crawl with one request

The seed sets the host. max_depth counts link hops from the seed, 0 to 5. max_pages stops the crawl at your limit. Patterns in include and exclude match the full URL or the path, with * as a wildcard, so /p/* keeps product pages and *?sort=* drops the endless sort variants.

Everything else is the same as a single request: render for pages built in JavaScript, wait_for a selector, a CSS or XPath schema, or on Growth and Scale a plain English describe of the fields. The response is a crawl ID straight away, with HTTP 202.

curl "https://webscrap.com/api/v1/crawl" \
  -H "Authorization: Bearer $WEBSCRAP_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://shop.example.com/shop",
    "max_depth": 2,
    "max_pages": 500,
    "include": ["/shop/*", "/p/*"],
    "exclude": ["*?sort=*", "/cart*", "/login*"],
    "render": "browser",
    "schema": {
      "title": "h1",
      "price": {"selector": ".price", "type": "number"},
      "sku": "[itemprop=sku]"
    },
    "webhook": "https://hooks.example.com/crawls"
  }'

# later, or when the webhook arrives
curl "https://webscrap.com/api/v1/crawl/4812?page=1" \
  -H "Authorization: Bearer $WEBSCRAP_KEY"

What a crawl costs per month

One successful page is one request, with rendering and extraction included. Monthly billing, 30 day months. The Firecrawl column counts credits with JSON extraction on every page, 5 per page.

JobPages per monthWebScrap planFirecrawl credits for the same pages as JSON
40 competitor sites, 100 pages each, weeklyabout 17,20049 USD, Starterabout 86,000
One retailer, 2,000 product pages, daily60,000149 USD, Growth, as 4 crawls of 500300,000
300 supplier sites, 500 pages each, monthly150,000149 USD, Growth750,000
Five retailers, 6,000 pages each, daily900,000499 USD, Scale4,500,000

If you only need Markdown for a language model, Firecrawl's 1 credit per page brings its numbers down by five and it becomes the cheaper choice at small volume. Yearly billing halves every price on our side, which takes Scale to 0.17 USD per 1,000 pages. All plans and limits are on the pricing page.

How the crawl works in four steps

01

Pick a seed that lists what you want

A category page, a store directory or a sitemap index page. The better the seed links to your targets, the shallower the crawl and the fewer pages you pay for.

02

Fence it with patterns

Keep the paths that hold data, drop carts, logins, sort and filter parameters. The crawl never leaves the seed's host, so ad and social links are ignored without any rule.

03

Name the fields once

One schema applies to every crawled page. Pages that do not match simply return empty fields, which makes listing pages easy to filter out afterwards.

04

Receive it, or schedule it

A signed webhook fires when the crawl ends. On Growth and Scale the same crawl becomes a scheduled job that runs on the interval you set, hourly or daily for most teams.

Who buys a website crawler API and what they run on it

Ecommerce pricing teams

Crawl each competitor's catalog nightly and land price, stock and SKU per product in a table. The Amazon scraper API covers the marketplace side.

Data and AI teams

Build a fresh, structured corpus from a documentation site, a regulator's publication list or a supplier portal for retrieval and fine tuning.

Recruiting and HR analytics

Walk a company's careers section every morning and catch new openings the day they go up, alongside the job scraping API.

SEO and site audit tools

Collect titles, headings, canonicals and status codes across a client's site from US addresses, then compare with positions from the rank tracking API.

Four reasons a crawl comes back incomplete

Faceted URLs eat the page budget. A store with 12 filters can link to millions of filter combinations. Without an exclude rule for the query string, a 500 page crawl fills up with the same 20 products in different orders. Exclude *?* first, then allow back only the parameters you need, such as pagination.

Links that only exist after JavaScript runs. Many category pages render the product grid in the browser. Fetched as plain HTML, they show no product links and the crawl stops at depth 1. Set render to browser and, if the grid loads late, wait_for its selector.

A challenge page that looks like a page. Anti-bot systems often answer with a 200 and a challenge body. A naive crawler stores it as content and finds no links on it. Here a recognised challenge counts as a failed attempt and is retried from another address; the patterns for Cloudflare and Imperva and Kasada are documented on their own pages.

Infinite scroll and load more buttons. If a listing only reveals products as you scroll, a link crawler sees the first screen. Seed from the paginated URL the site uses behind the button, or from its sitemap, instead of the scrolling page.

Website crawler API questions

How do I crawl an entire website with an API?

Send one POST request with the seed URL, a maximum depth, a page limit and the fields you want from each page. The crawler fetches the seed, collects the links on the same host, filters them through your include and exclude patterns and keeps going level by level. You read the results by crawl ID or receive them on a webhook.

What is the difference between a web crawler and a web scraper?

A crawler discovers pages by following links; a scraper reads data out of a page you already know. A crawler API does both in one job: it finds the product, listing or article pages on a site and returns the fields from each one, so you do not have to build the URL list yourself first.

How much does a web crawler API cost?

On WebScrap every successfully crawled page is one request: 0.98 USD per 1,000 pages on Starter, 0.60 on Growth and 0.33 on Scale, with rendering and extraction included. Credit based crawlers charge more per page once you ask for JSON, and some bill pages that returned a 403 or 404.

How many pages can one crawl cover?

Up to 100 pages per crawl on Starter, 500 on Growth and 2,000 on Scale, at a depth of up to 5 links from the seed. Larger sites are split into several crawls by section with include patterns, or you send the URLs from the sitemap straight to the scrape endpoint.

Can a crawler API get past Cloudflare?

Some can and some cannot. Cloudflare's own crawl endpoint identifies itself as a bot and cannot pass Cloudflare bot detection or captchas. WebScrap fetches every crawled page through rotating datacenter and residential addresses with automatic retries and a real browser when you set render, and a blocked page is reported as failed rather than billed.

Does the crawler stay on one domain?

Yes. A crawl only follows links on the exact host of the seed URL, so it never wanders onto social networks, ad servers or other subdomains. To cover a blog subdomain or a regional store, start a second crawl from that host.

How do I get the results of a crawl?

Read the crawl by its ID, 50 pages per response, each with the URL, depth, status and the extracted fields. Or pass a webhook URL when you start it: when the crawl ends you receive a signed crawl.finished event with the summary and the first 200 pages, and page through the rest by ID.

Can I schedule a website crawl to run every day?

Yes, on Growth and Scale. Save the crawl as a scheduled job with an interval in minutes, such as every hour or every 1,440 minutes for daily, and a webhook, and each run posts its results when it finishes. Growth runs 25 scheduled jobs and Scale 500. On Starter you trigger crawls from your own cron.

Do failed pages count against my plan?

No. Only pages that come back with a 2xx response count as requests. A timeout, a block page or a 5xx from the target is retried inside the request and, if it still fails, marked failed in the crawl and not billed.

Is crawling a website legal?

Crawling public pages is common practice for price monitoring, research and search, but the rules depend on what you collect and how you use it. Read the site's terms and robots.txt, put disallowed paths in your exclude patterns, keep the request rate reasonable and avoid storing personal data you do not need. More in is web scraping legal.

Crawl the site once, get every page as data

One seed URL, your depth and your fields. Pages come back as JSON, failed pages are never billed, and Growth runs the crawl on a schedule.

Billed only on successful requests. No card required to create your account.