Web Data Scraping, One Request at a Time
The lifecycle
-
Step 1. The request arrives
You send one POST with the target URL, the country, the render mode and the extraction schema. The request is validated, your key is checked against the plan and the job is queued at the concurrency your plan allows.
-
Step 2. Routing and the exit address
An exit address is chosen from the pool you asked for, residential or datacenter, in the country and, in the United States, the city you named. The identity attached to it is consistent: TLS fingerprint, header order, browser version and viewport all agree with each other, because a mismatch is what most blocks actually detect. The pool is described on the web scraping proxy page.
-
Step 3. Fetch or render
Fast mode takes the document as it is served. Browser mode loads the page in a real browser, runs the scripts, waits for the selector you named or for the network to settle, and can scroll or click where content only appears after an interaction. A screenshot is returned when you ask for one.
-
Step 4. Unblocking and retries
A challenge, a redirect chain or a rate limit is handled inside the request. A failed attempt is retried with a new address and a new identity, up to the limit for your plan. None of those attempts reach your code, and none of them touch your counter.
-
Step 5. Extraction and response
Your schema turns the page into fields before it leaves us: CSS or XPath, or a plain English description through the ai web scraper mode. The response carries the status, elapsed milliseconds, exit country, whether a browser was used, and the typed JSON. Reading what a rendered page shows on screen is described on the screen scraper page.
The shape of a response
| Field | What it contains |
|---|---|
status | The HTTP status the target returned on the successful attempt |
elapsed_ms | Total time from your call to the response leaving us |
country | The exit location the request went out through |
rendered | Whether a browser was used for this page |
attempts | How many attempts it took, so you can see which targets are expensive |
data | Your fields, typed, under the keys you chose |
html | The document, when you asked for Raw HTML |
screenshot_url | A time limited link, when you asked for a screenshot |
The same shape comes back from every endpoint, including the serp api and the website scraper, so one response handler covers the whole integration.
Run a requestThe failure modes and what happens to each
- The target is down or returns a 5xx: retried, then reported with the status. Not billed.
- The target blocks every attempt: reported as blocked with the attempt count. Not billed.
- The selector matches nothing: a successful response with an empty field, so you can tell a missing value from a failed fetch.
- The page takes longer than the timeout: reported as a timeout. Not billed.
- Your monthly limit is reached: a quota error your code can read, never an extra charge.
Only a 2xx response with a payload moves the counter, which is what makes the monthly cost arithmetic rather than a forecast.
When the crawl runs without you
Scheduled jobs on Growth and above run a saved request or a crawl on a schedule and deliver each run to your webhook, with retries on delivery failure. That covers the daily catalog refresh, the hourly job board sweep and the morning rank report, described on the web crawling tools page.