Skip to content
WebScrap

Web Scraping Legal Limits and What the Courts Have Actually Said

The question arrives as one sentence and splits into four. Reading a public page, agreeing to terms, copying protected content and holding personal data are separate legal problems, and confusing them is how a data project gets stopped in month three.

Nobody is asking whether reading is allowed

Every browser on the internet scrapes. It requests a document, parses it and renders the result. The law has never taken an interest in that, and it does not start caring because the parser is yours rather than a vendor's. What the law takes an interest in is access to systems you were not allowed into, agreements you made, works someone owns and facts about people.

So the useful version of the question is not whether scraping is legal. It is: which of those four am I touching.

Unauthorised access and the computer crime question

The strongest fear in the room is usually criminal: that a crawler equals hacking. The line the courts have drawn repeatedly is authorisation. Where information sits on a public server, with no login and no technical barrier to entry, courts in the United States have declined to treat reading it as access without authorisation under the Computer Fraud and Abuse Act, and the Supreme Court read that statute narrowly in 2021, rejecting the theory that using data for a purpose the owner dislikes turns permitted access into a crime. European courts have approached it similarly: the offence is circumventing protection, not reading what was published.

The practical consequence is a line you can hold. Public page, no login, no bypassed control: on the right side. Credentials you were not given, a paywall you got around, a technical block you defeated by pretending to be an authorised user: on the wrong side, and no volume of business justification moves it back.

Contract terms are where people actually lose

This is where the surprises live. Terms of service are a contract when you accept them, and a contract can forbid what the law permits. Courts distinguish between terms you clicked to accept, which usually bind you, and terms sitting in a link at the bottom of a page that nobody was asked to agree to, which often do not.

The asymmetry matters more than any other rule on this page. A team that logs into a platform to collect data has agreed to that platform's terms and is in contract territory. A crawler that never logs in and never accepts anything is usually not. That is exactly why our API forwards no credentials and fetches public pages only, described on the enterprise web scraping page. It is not caution for its own sake: it keeps a contract claim off the table.

A price is a fact. A specification is a fact. A rating out of five is a fact. Facts are not protected by copyright, and a database of them can be collected and analysed. Article text, photographs, product descriptions written to persuade and reviews written by customers are expression, and they belong to whoever wrote them.

Two rules keep most projects clean. First, collect facts and derive analysis rather than republishing prose. Second, do not rebuild the source: a service that reproduces an article in full, or a catalog that mirrors another catalog page for page, is a substitute for the original, and substitution is where infringement arguments start. In the European Union there is a further layer, the sui generis database right, which can protect a substantial investment in assembling a database even when the individual facts are open.

Personal data brings obligations that follow you home

This is the part that was quiet for years and is now the loudest. Under GDPR and its equivalents, personal data does not lose its protection because it was published. If you collect names, profiles, contact details, reviews attributable to a person or anything that can be tied back to an individual, you become a controller of that data with the full set of obligations: a lawful basis, a purpose you wrote down before collecting, a retention period, a way to answer an access request, a deletion process and often an obligation to tell people you hold data about them.

Regulators have fined companies for building people databases from public sources, and the fines were not about the collection method. They were about holding the data without a basis and without telling anyone. If your project needs personal data, involve whoever owns data protection in your company before the first request, not after the pipeline works.

The questions a lawyer will actually ask you

  • Does the collection require a login, a bypassed control or credentials you were not given.
  • Did anyone at your company accept the terms of the site you are collecting from.
  • Are you collecting facts, or copying expression that someone wrote.
  • Does the result contain data about identifiable people, and if so what is your basis for keeping it.
  • Would your crawl degrade the service of the site you read it from.
  • Would you be comfortable if the source company read your internal description of this project.

If those six answers are clean, the legal conversation is usually short. If one of them is not, that is the one to fix, and it is almost never fixed by changing scraping tool.

What a scraping API changes and what it does not

It changes the technical questions: blocking, rendering, rotation, retries and pacing stop being your problem, and pacing in particular matters legally, because a crawl that degrades a target's service invites a claim that reading alone never would. The web scraping proxy layer paces per host for exactly that reason.

It does not change the four questions above. You choose the targets, you choose the fields, you decide what is stored and for how long. We keep the responses only for the retention window of your plan and never resell them, which is written into the privacy policy, but the decision about what to collect stays yours, because it is the only part that cannot be delegated.

The short version

Reading public pages is generally lawful. The trouble comes from logins, from agreements, from copying expression and from data about people. Keep the crawler to public pages, collect facts, pace it politely, write down why you hold what you hold, and take advice for your own jurisdiction before you scale.

This article is a description of the landscape, not legal advice, and none of it is a substitute for a lawyer who knows your case.