Skip to content
WebScrap

Web Scraping API Python Docs and Reference

Everything the API accepts and everything it returns, on one page. Examples are shown for curl and for Python, because those are the two ways almost everyone starts.

Your key and where it goes

Every request carries your API key in an Authorization header as a bearer token. Keys are created, scoped, rotated and revoked in the account, shown once at creation and stored hashed afterwards. A key can be limited to a project, given its own rate limit and given a budget, so a test key cannot spend a production allowance.

Requests without a valid key return 401. A key that is valid but out of scope for the endpoint returns 403.

Your first request

curl "https://webscrap.com/api/v1/scrape" \
  -H "Authorization: Bearer $WEBSCRAP_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/products/42",
    "render": "browser",
    "country": "us",
    "schema": {
      "title": "h1",
      "price": {"selector": ".price", "type": "number"}
    }
  }'
import os
import requests

response = requests.post(
    "https://webscrap.com/api/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['WEBSCRAP_KEY']}"},
    json={
        "url": "https://example.com/products/42",
        "render": "browser",
        "country": "us",
        "schema": {
            "title": "h1",
            "price": {"selector": ".price", "type": "number"},
        },
    },
    timeout=60,
)
response.raise_for_status()
print(response.json()["data"])

Endpoints

EndpointMethodWhat it does
/v1/scrapePOSTFetch one URL, render it if asked, extract the schema, return the result
/v1/crawlPOSTStart a crawl job from a seed URL with depth and pattern rules
/v1/crawl/{id}GETRead the status and the results of a crawl job
/v1/serpPOSTSearch results for a query, country, language and device
/v1/productPOSTA product page as named fields: name, price, currency, availability, brand, SKU, rating, images
/v1/jobPOSTA job listing, or a page of them, as named fields: title, company, location, salary, dates
/v1/jobsPOSTCreate a scheduled job from a saved request or crawl
/v1/jobs/{id}GET, PATCH, DELETERead, change or remove a scheduled job
/v1/usageGETRequests used, requests remaining and the current period
/v1/responsesDELETEDelete every stored response of the account in one call

Request parameters

ParameterTypeMeaning
urlstringThe target address, http or https
renderstringfast for a plain fetch, browser to run the page in a real browser
wait_forstringA CSS selector the browser waits for before reading the page
countrystringTwo letter country code, plus an optional city in the United States
proxy_typestringdatacenter or residential
sessionstringAn identifier that keeps the same exit address across calls, up to 30 minutes
schemaobjectField name to CSS or XPath selector, with an optional type
describestringA plain English list of the fields you want, instead of a schema
screenshotbooleanReturn a time limited link to a screenshot of the rendered page
timeoutintegerSeconds to wait before the attempt is abandoned, up to 60
webhookstringAn address that receives the result instead of the response body
pageintegerThe results page for listings, catalogs and search results, starting at 1; the response carries result_count

The response shape

Every endpoint returns the same envelope: status, elapsed_ms, country, rendered, attempts, and then data with your typed fields, plus html or screenshot_url when you asked for them. One response handler covers every endpoint, which is described in full on the web data scraping page.

Error codes

CodeMeaningWhat to do
200The target answered and your schema ranRead data and continue
400The request is malformed, usually the schema or the URLFix the body, the message names the field
401Missing or invalid API keyCheck the header and the key
403The key is out of scope for this endpoint or projectUse a key with the right scope
422The target was reached but every attempt was blockedSwitch to residential, add browser rendering, or change country
429Concurrency or rate limit for your planSlow down, or move up a plan
402The monthly request limit is used upMove up a plan, the new limit applies immediately
504The target did not answer within the timeoutRetry later, or raise the timeout

Codes 4xx and 5xx that come from us do not count against your monthly limit. Only a 200 with a payload does.

What your plan allows at once

Concurrency is 5 requests on Starter, 50 on Growth, 200 on Scale and 400 on Enterprise. Exceeding it returns 429 rather than queueing forever, so your own backpressure stays in your control. Scheduled jobs run inside the same concurrency, and crawl jobs pace themselves per host.

The Python quickstart with pagination and a pandas load is on the python scraping tool page. If you would rather not write any of it, the no code web scraper builds the same requests by clicking.