Web Scraping API Python Docs and Reference
Your key and where it goes
Every request carries your API key in an Authorization header as a bearer token. Keys are created, scoped, rotated and revoked in the account, shown once at creation and stored hashed afterwards. A key can be limited to a project, given its own rate limit and given a budget, so a test key cannot spend a production allowance.
Requests without a valid key return 401. A key that is valid but out of scope for the endpoint returns 403.
Your first request
curl "https://webscrap.com/api/v1/scrape" \
-H "Authorization: Bearer $WEBSCRAP_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/products/42",
"render": "browser",
"country": "us",
"schema": {
"title": "h1",
"price": {"selector": ".price", "type": "number"}
}
}'
import os
import requests
response = requests.post(
"https://webscrap.com/api/v1/scrape",
headers={"Authorization": f"Bearer {os.environ['WEBSCRAP_KEY']}"},
json={
"url": "https://example.com/products/42",
"render": "browser",
"country": "us",
"schema": {
"title": "h1",
"price": {"selector": ".price", "type": "number"},
},
},
timeout=60,
)
response.raise_for_status()
print(response.json()["data"])
Endpoints
| Endpoint | Method | What it does |
|---|---|---|
/v1/scrape | POST | Fetch one URL, render it if asked, extract the schema, return the result |
/v1/crawl | POST | Start a crawl job from a seed URL with depth and pattern rules |
/v1/crawl/{id} | GET | Read the status and the results of a crawl job |
/v1/serp | POST | Search results for a query, country, language and device |
/v1/product | POST | A product page as named fields: name, price, currency, availability, brand, SKU, rating, images |
/v1/job | POST | A job listing, or a page of them, as named fields: title, company, location, salary, dates |
/v1/jobs | POST | Create a scheduled job from a saved request or crawl |
/v1/jobs/{id} | GET, PATCH, DELETE | Read, change or remove a scheduled job |
/v1/usage | GET | Requests used, requests remaining and the current period |
/v1/responses | DELETE | Delete every stored response of the account in one call |
Request parameters
| Parameter | Type | Meaning |
|---|---|---|
url | string | The target address, http or https |
render | string | fast for a plain fetch, browser to run the page in a real browser |
wait_for | string | A CSS selector the browser waits for before reading the page |
country | string | Two letter country code, plus an optional city in the United States |
proxy_type | string | datacenter or residential |
session | string | An identifier that keeps the same exit address across calls, up to 30 minutes |
schema | object | Field name to CSS or XPath selector, with an optional type |
describe | string | A plain English list of the fields you want, instead of a schema |
screenshot | boolean | Return a time limited link to a screenshot of the rendered page |
timeout | integer | Seconds to wait before the attempt is abandoned, up to 60 |
webhook | string | An address that receives the result instead of the response body |
page | integer | The results page for listings, catalogs and search results, starting at 1; the response carries result_count |
The response shape
Every endpoint returns the same envelope: status, elapsed_ms, country, rendered, attempts, and then data with your typed fields, plus html or screenshot_url when you asked for them. One response handler covers every endpoint, which is described in full on the web data scraping page.
Error codes
| Code | Meaning | What to do |
|---|---|---|
| 200 | The target answered and your schema ran | Read data and continue |
| 400 | The request is malformed, usually the schema or the URL | Fix the body, the message names the field |
| 401 | Missing or invalid API key | Check the header and the key |
| 403 | The key is out of scope for this endpoint or project | Use a key with the right scope |
| 422 | The target was reached but every attempt was blocked | Switch to residential, add browser rendering, or change country |
| 429 | Concurrency or rate limit for your plan | Slow down, or move up a plan |
| 402 | The monthly request limit is used up | Move up a plan, the new limit applies immediately |
| 504 | The target did not answer within the timeout | Retry later, or raise the timeout |
Codes 4xx and 5xx that come from us do not count against your monthly limit. Only a 200 with a payload does.
What your plan allows at once
Concurrency is 5 requests on Starter, 50 on Growth, 200 on Scale and 400 on Enterprise. Exceeding it returns 429 rather than queueing forever, so your own backpressure stays in your control. Scheduled jobs run inside the same concurrency, and crawl jobs pace themselves per host.
The Python quickstart with pagination and a pandas load is on the python scraping tool page. If you would rather not write any of it, the no code web scraper builds the same requests by clicking.