Skip to content
WebScrap

Web Data Scraping, One Request at a Time

What happens between the moment your code sends a URL and the moment typed JSON comes back. Five steps, none of which you have to build.

The lifecycle

  1. Step 1. The request arrives

    You send one POST with the target URL, the country, the render mode and the extraction schema. The request is validated, your key is checked against the plan and the job is queued at the concurrency your plan allows.

  2. Step 2. Routing and the exit address

    An exit address is chosen from the pool you asked for, residential or datacenter, in the country and, in the United States, the city you named. The identity attached to it is consistent: TLS fingerprint, header order, browser version and viewport all agree with each other, because a mismatch is what most blocks actually detect. The pool is described on the web scraping proxy page.

  3. Step 3. Fetch or render

    Fast mode takes the document as it is served. Browser mode loads the page in a real browser, runs the scripts, waits for the selector you named or for the network to settle, and can scroll or click where content only appears after an interaction. A screenshot is returned when you ask for one.

  4. Step 4. Unblocking and retries

    A challenge, a redirect chain or a rate limit is handled inside the request. A failed attempt is retried with a new address and a new identity, up to the limit for your plan. None of those attempts reach your code, and none of them touch your counter.

  5. Step 5. Extraction and response

    Your schema turns the page into fields before it leaves us: CSS or XPath, or a plain English description through the ai web scraper mode. The response carries the status, elapsed milliseconds, exit country, whether a browser was used, and the typed JSON. Reading what a rendered page shows on screen is described on the screen scraper page.

The shape of a response

FieldWhat it contains
statusThe HTTP status the target returned on the successful attempt
elapsed_msTotal time from your call to the response leaving us
countryThe exit location the request went out through
renderedWhether a browser was used for this page
attemptsHow many attempts it took, so you can see which targets are expensive
dataYour fields, typed, under the keys you chose
htmlThe document, when you asked for Raw HTML
screenshot_urlA time limited link, when you asked for a screenshot

The same shape comes back from every endpoint, including the serp api and the website scraper, so one response handler covers the whole integration.

Run a request

The failure modes and what happens to each

  • The target is down or returns a 5xx: retried, then reported with the status. Not billed.
  • The target blocks every attempt: reported as blocked with the attempt count. Not billed.
  • The selector matches nothing: a successful response with an empty field, so you can tell a missing value from a failed fetch.
  • The page takes longer than the timeout: reported as a timeout. Not billed.
  • Your monthly limit is reached: a quota error your code can read, never an extra charge.

Only a 2xx response with a payload moves the counter, which is what makes the monthly cost arithmetic rather than a forecast.

When the crawl runs without you

Scheduled jobs on Growth and above run a saved request or a crawl on a schedule and deliver each run to your webhook, with retries on delivery failure. That covers the daily catalog refresh, the hourly job board sweep and the morning rank report, described on the web crawling tools page.