Skip to content
WebScrap

Python Scraping Tool You Call in Ten Lines

The Python you keep is the loop over your own targets. Everything below it, the session pool, the retry policy, the rotating addresses and the headless browser, is one HTTP call.
Read the docs
quickstart.py POST /v1/scrape
import os
import requests

response = requests.post(
    "https://webscrap.com/api/v1/scrape",
    headers={"Authorization": f"Bearer {os.environ['WEBSCRAP_API_KEY']}"},
    json={"url": "https://example.com/products/1", "country": "us", "render": "browser",
          "schema": {"title": "h1", "price": {"selector": ".price", "type": "number"}}},
)
print(response.json()["data"])

Install nothing new

The only dependency is the HTTP client you already use. There is no SDK to pin, no browser driver to install on the build machine and no version of a headless binary to chase when the operating system moves.

  1. 1

    Confirm your work email and copy the API key from your account.

  2. 2

    Send a POST to the endpoint with the URL, the country, the render mode and your field schema.

  3. 3

    Read the typed JSON out of the response and write it wherever your pipeline puts rows.

The parameter reference, the response shape and every error code are in the web scraping api python docs.

Read the docs

The four things every Python crawl needs

Concurrency without a thread pool of your own

Your plan sets the concurrency: 5 on Starter, 50 on Growth, 200 on Scale. Fire requests with your usual async client or a thread pool and let the limit be the limit, instead of tuning a semaphore against a proxy provider.

Retries you do not write

Blocks, challenges and timeouts are retried on our side with a fresh identity. Your code keeps a retry only for network faults between you and us, which is a much shorter piece of code.

Pagination

Search results, listings and catalogs come back with a page parameter and a result count, so a loop over pages is a loop, not a state machine that learns where it stopped.

Straight into pandas

Typed JSON means a list of dictionaries with stable keys, which loads into a data frame without a cleaning pass. Numbers arrive as numbers and dates as dates, so the first thing you do with the data is analysis rather than parsing.

pages_to_pandas.py Python 3
import os
import pandas as pd
import requests

API = "https://webscrap.com/api/v1/scrape"
HEADERS = {"Authorization": f"Bearer {os.environ['WEBSCRAP_API_KEY']}"}

rows, page = [], 1
while True:
    body = {"url": "https://example.com/category/shoes", "country": "us", "page": page,
            "schema": {"name": ".product h2", "price": {"selector": ".product .price", "type": "number"}}}
    result = requests.post(API, headers=HEADERS, json=body, timeout=90).json()
    rows.extend(result["data"])
    if not result["data"] or len(rows) >= result["result_count"]:
        break
    page += 1

frame = pd.DataFrame(rows)
print(frame.describe())

Code you can remove from the repository

  • The proxy rotation module and its list of addresses
  • The user agent and header randomisation helper
  • The headless browser launcher and its wait conditions
  • The ban detector and the exponential backoff around it
  • Most of the parser, which becomes a schema in the request body

What stays is the part that is actually yours: which targets, how often and what you do with the rows.