Skip to content
WebScrap

Data Scraping FAQ

The questions engineers and the people who sign their invoices actually ask before they put a data scraping pipeline into production. Answers in the present tense, with the uncomfortable ones kept first.

Legality and ownership

Is data scraping legal

Collecting information from pages that anyone can open without logging in is generally lawful in the jurisdictions our customers work in, and courts have repeatedly declined to treat reading public data as unauthorised access to a computer system. The obligations attach to what you collect and what you do next: personal data falls under GDPR, CCPA and their equivalents, copyrighted text and images stay protected, and terms you actively accepted can bind you by contract. The full discussion with the case law is in is web scraping legal. None of this is legal advice.

Who owns the data I collect through the API

You do. WebScrap is the pipe, not the owner. We never resell responses, never share them with another customer, and never use your targets or payloads to build a dataset for anyone else.

Can I scrape pages behind a login

No. The API fetches public pages only and forwards no credentials of yours. Logged in areas, paywalled content and anything requiring an account are out of scope, and that limit is deliberate.

Do you collect personal data on my behalf

Only when you point the API at a page that contains it, and then you are the controller of that data, not us. Our own retention is short by design and you can set it shorter. The details live on the enterprise web scraping page.

Reliability and blocking

Why do my requests get blocked without an API

Sites read far more than your IP address. TLS fingerprint, header order, browser version, cookie behaviour, request rhythm and the absence of JavaScript execution all mark a plain client. Changing your address alone fixes the smallest part of that.

How does WebScrap get past a block

Every request goes out through a rotating web scraping proxy with a consistent browser identity behind it: matching TLS fingerprint, header set and viewport, with a real browser rendering the page when the site needs one. Challenges and redirects are handled inside the request, and a retry goes out with a fresh identity.

What happens when a site changes its layout

Your extraction schema stops matching, so you change one selector or one line of the plain English description in the request body. Nothing is redeployed on your side, because the parser is a parameter, not a piece of your code.

What is your uptime

Scale carries a 99.9 percent uptime commitment and a four business hour response time. Enterprise adds an SLA with credits and a named technical contact. Starter and Growth run on the same infrastructure without the contractual commitment.

Cost and limits

What does a database scraper cost to run yourself

Two proxy providers, a headless browser cluster, a retry layer and an engineer maintaining parsers is the honest shape of it, and none of those line items stop once they start. The comparison with real numbers is in open source scraping tools versus an API.

Do I pay for failed requests

No. Only 2xx responses move the counter. Blocks, timeouts and retries are absorbed on our side, which is what makes the monthly cost predictable.

What happens when I hit my monthly limit

The API returns a quota error that your code can read, not an overage invoice. You move up a plan from inside the account and the higher limit applies immediately.

Is there a way to try it before paying

Yes. The demo on the web scraping api homepage runs five live requests per session against any URL you type, plus unlimited runs on the four bundled targets. No account and no card to see a real response.

Integration

Can I use this from Python

Yes, and most customers do. The quickstart with requests, pandas and a working pagination loop is on the python scraping tool page, and the full reference is in the web scraping api python docs.

Can I use it without writing code

Yes. The no code web scraper builds a schema by pointing at fields on the page, runs it on a schedule and posts the result to a spreadsheet or a webhook.

Do you have ready made endpoints for common targets

Yes, on Growth and above: search results through the serp api, product pages, job listings, plus dedicated endpoints for amazon scraping, twitter scraping, scrape linkedin and scrap google maps.

Can I crawl an entire site rather than single pages

Yes. The website scraper endpoint walks a domain by depth and URL pattern, deduplicates what it finds, and applies your schema to every page it keeps.

Data handling

What happens to my data after the request

Responses are deleted when the retention window of your plan ends: 24 hours on Starter, 7 days on Growth, 30 days on Scale and 90 days on Enterprise, and you can set it shorter on any plan. Deletion covers the payload, the rendered HTML and any screenshot.

Where are the requests processed

In the region you choose. Enterprise selects data residency in the EU or the US, and requests never leave the region they were routed to.

Can I delete everything on demand

Yes. A single call purges every stored response for your account, and closing your account purges the rest within 30 days.

Create your account and answer the last question yourself

Run the request that matters to you and read the response. That settles more than any answer on this page.

Billed only on successful requests. No card required to create your account.