Python Scraping Tool You Call in Ten Lines
import os
import requests
response = requests.post(
"https://webscrap.com/api/v1/scrape",
headers={"Authorization": f"Bearer {os.environ['WEBSCRAP_API_KEY']}"},
json={"url": "https://example.com/products/1", "country": "us", "render": "browser",
"schema": {"title": "h1", "price": {"selector": ".price", "type": "number"}}},
)
print(response.json()["data"])
Install nothing new
The only dependency is the HTTP client you already use. There is no SDK to pin, no browser driver to install on the build machine and no version of a headless binary to chase when the operating system moves.
- 1
Confirm your work email and copy the API key from your account.
- 2
Send a POST to the endpoint with the URL, the country, the render mode and your field schema.
- 3
Read the typed JSON out of the response and write it wherever your pipeline puts rows.
The parameter reference, the response shape and every error code are in the web scraping api python docs.
Read the docsThe four things every Python crawl needs
Concurrency without a thread pool of your own
Your plan sets the concurrency: 5 on Starter, 50 on Growth, 200 on Scale. Fire requests with your usual async client or a thread pool and let the limit be the limit, instead of tuning a semaphore against a proxy provider.
Retries you do not write
Blocks, challenges and timeouts are retried on our side with a fresh identity. Your code keeps a retry only for network faults between you and us, which is a much shorter piece of code.
Pagination
Search results, listings and catalogs come back with a page parameter and a result count, so a loop over pages is a loop, not a state machine that learns where it stopped.
Straight into pandas
Typed JSON means a list of dictionaries with stable keys, which loads into a data frame without a cleaning pass. Numbers arrive as numbers and dates as dates, so the first thing you do with the data is analysis rather than parsing.
import os
import pandas as pd
import requests
API = "https://webscrap.com/api/v1/scrape"
HEADERS = {"Authorization": f"Bearer {os.environ['WEBSCRAP_API_KEY']}"}
rows, page = [], 1
while True:
body = {"url": "https://example.com/category/shoes", "country": "us", "page": page,
"schema": {"name": ".product h2", "price": {"selector": ".product .price", "type": "number"}}}
result = requests.post(API, headers=HEADERS, json=body, timeout=90).json()
rows.extend(result["data"])
if not result["data"] or len(rows) >= result["result_count"]:
break
page += 1
frame = pd.DataFrame(rows)
print(frame.describe())
Code you can remove from the repository
- The proxy rotation module and its list of addresses
- The user agent and header randomisation helper
- The headless browser launcher and its wait conditions
- The ban detector and the exponential backoff around it
- Most of the parser, which becomes a schema in the request body
What stays is the part that is actually yours: which targets, how often and what you do with the rows.
Related pages
- Web scraping api python docs for the full reference
- AI web scraper when you would rather describe fields than write selectors
- Website scraper for crawling a whole domain from one seed
- Website crawler API when Scrapy is more crawler than you want to run
- No code web scraper for the colleague who does not write Python