Cloudflare Scraping API for Cloudflare Web Scraping Without the Blocks
Send a URL that sits behind Cloudflare and get the finished page back. Each request goes out on a rotating residential address with a matching browser identity, the challenge is answered before the response reaches your code, and you receive raw HTML or the fields you named as typed JSON. Blocked attempts are never billed.
- Status
- ...
- Elapsed
- ... ms
- Exit country
- ...
- Browser used
- ...
This console sends a real request and shows you the real response. Five live fetches per session, and the sample targets run as often as you like.
What a Cloudflare scraping API actually does
A Cloudflare scraping API fetches public pages served behind Cloudflare Bot Management and hands back the finished HTML or parsed JSON. It rotates residential addresses, presents a TLS fingerprint and header order that match the browser it claims to be, executes the challenge JavaScript, and retries on a fresh identity when a request is turned away.
That is the whole job. Cloudflare sits in front of a large share of the commercial web, so any crawler that covers retail, travel, real estate, ticketing or marketplace sites meets it within the first week. The teams who end up buying this are the ones who already wrote the scraper, watched it work for a month, and then watched it return challenge pages instead of data.
Cloudflare web scraping fails for a specific reason, and it is almost never the address alone. Bot Management scores every request from 1 to 99, where 1 means Cloudflare is quite certain the request was automated, and the site owner picks what happens at each score. A plain HTTP client gives away enough signals to land near the bottom of that range on its first request.
What Cloudflare checks before it serves your request
Cloudflare anti scraping is a stack of independent checks, and failing any one of them is enough to get challenged. This is what each layer looks at, what a normal script gets wrong, and what leaves our infrastructure instead.
| Signal | What a plain script sends | What the API sends |
|---|---|---|
| TLS fingerprint | A JA3 or JA4 signature belonging to Python requests, curl or Go, which no human browser produces | A handshake that matches the browser version in the user agent, cipher order included |
| HTTP/2 frame order | Header names, casing and pseudo-header order that do not match any shipped browser | The frame settings and header order of the real browser build being presented |
| Address reputation | A cloud provider ASN, usually already flagged by thousands of other crawlers | Residential and mobile addresses in 90 countries, rotated per request or held in a sticky session |
| JavaScript challenge | Nothing. The interstitial is saved to disk as if it were the page | A real browser runs the challenge, holds the clearance cookie and continues to the page |
| Turnstile | An unanswered widget and a response that never becomes the content you asked for | Handled inside the request, with the attempt retried on a new identity if it does not clear |
| Pacing and behavior | Perfectly even intervals from one address, with no referer chain and no asset loads | Concurrency spread across the pool, with session reuse where the target expects continuity |
The address pool matters, but it is the last piece, not the first. Teams who buy residential proxies and keep the same HTTP client usually see the block rate improve for a few days and then settle back. The detail on pool types is on the web scraping proxy service page.
Which Cloudflare block you are looking at
The response tells you which layer refused you, and the fix is different for each. Cloudflare publishes the meaning of the numbered errors, and the interstitials are easy to recognize once you know what they are.
| What you get back | What triggered it | What clears it |
|---|---|---|
| A page reading "Just a moment" | The JavaScript challenge, served because the request scored as automated | Execute the challenge in a real browser and keep the clearance cookie for the session |
| HTTP 403 with a Cloudflare page | A WAF or bot rule refused the request outright, often on fingerprint alone | Change the client identity, not just the address, and slow the burst |
| Error 1020: access denied | A firewall rule written by the site owner matched your request | Look at what the rule matches, commonly the country, the ASN or a header pattern |
| Error 1015: you are being rate limited | Too many requests from one address inside the window the owner configured | Spread the same volume across more addresses and add jitter between calls |
| Error 1010: browser signature banned | The client signature itself is on a block list, which catches headless defaults | Present a shipped browser build end to end, including the TLS layer |
| A Turnstile widget | Managed Challenge, usually on login, search or checkout paths | Solve inside the request, or route around it if the data exists on an unprotected path |
Our API returns the underlying status and a reason code when a target wins, so you can tell a genuine 404 from a challenge you never saw. Those attempts do not count against your plan. The full error table is in the API docs.
Who pays to scrape Cloudflare-protected sites
-
Ecommerce and pricing teams
Most US retail storefronts sit behind Cloudflare, so a price monitor that worked on five competitors in January is returning challenge pages by March. Flat per request pricing matters here, because you are checking the same SKU list every day and the cost per check decides whether the project survives review.
-
Travel, ticketing and marketplace data
Fares, inventory and resale prices change hourly and the sites defending them are the ones with the tightest bot rules. Sticky sessions keep a search context alive long enough to read results that only exist after a form submission.
-
SEO and market intelligence agencies
Client reporting breaks the moment a tracked competitor turns on Bot Management, and nobody wants to explain a gap in a monthly deck. Pair this with the SERP API when the same report needs ranking data alongside on-page changes.
-
AI and data engineering teams
Retrieval corpora and evaluation sets need current pages, not a crawl from last year. Naming the fields you want returns clean records straight away, which is what the AI web scraper endpoint does instead of handing you markup to parse.
How a protected page comes back in four steps
-
Step 1
Send the URL
One GET with your key. No proxy configuration, no browser to install, nothing to keep patched.
-
Step 2
We pick an identity
An address, a browser build and a matching fingerprint are chosen for that target, with geotargeting when the page changes by region.
-
Step 3
The challenge is answered
The browser runs the challenge, holds the clearance cookie and loads the real page. A refusal is retried on a fresh identity, and those attempts are not billed.
-
Step 4
You get HTML or JSON
Raw markup if you already have a parser, or the fields you named as typed values ready for the database.
The call is a few lines in any language. The Python scraping tool quickstart shows the loop with retries removed, and the web scraping tool overview covers the parameters that change rendering and session behavior.
Four ways to scrape a Cloudflare site, compared honestly
Each of these is the right answer for somebody. The question is how much of your week you want to spend on the arms race.
| Approach | Where it wins | Where it costs you |
|---|---|---|
| cloudscraper and similar modules | A site left on the lightest settings, a one off job, no budget line at all | No answer to Turnstile or a Managed Challenge, and nothing done about the TLS fingerprint, so results are unpredictable week to week |
| Patched Playwright or Puppeteer, self hosted | Full control of the session, useful when you need to click through a flow the API cannot describe | You run and re-patch a browser fleet after most Chrome releases, buy residential addresses separately, and pay for the memory each instance holds |
| Residential proxies on their own | Cheap address diversity, and the right buy when your client is already a real browser | The fingerprint still fails, so the challenge still appears. Most of the gain disappears within days |
| A managed Cloudflare scraping API | One call, a predictable monthly number, and somebody else keeping up with the detection changes | You pay per request and your traffic goes through a third party, which needs a look from security on a regulated team |
What Cloudflare scraping costs here
One successful request is one unit, whether the target is a static blog or a storefront behind Managed Challenge. There is no multiplier for a hard page, which is the part that decides the bill on this kind of work.
| Plan | Per month | Successful requests | Cost per 1,000 pages |
|---|---|---|---|
| Starter | 49 USD | 50,000 | 0.98 USD |
| Growth | 149 USD | 250,000 | 0.60 USD |
| Scale | 499 USD | 1,500,000 | 0.33 USD |
Compare that with credit pricing, where a protected page costs more than a plain one. At ScraperAPI's published rates, checked on 20 September 2026, a normal request is 1 credit and the Cloudflare bypass adds 10, so a Cloudflare page is 11 credits, or 21 with JavaScript rendering on top. On their 49 USD plan of 100,000 credits that is about 5.39 USD per 1,000 protected pages, rising to roughly 10.29 USD when rendering is needed. The same 1,000 pages cost 0.98 USD on Starter here. The full breakdown by target, including where credit pricing is the cheaper buy, is in ScraperAPI pricing per 1,000 pages.
What this does not do
Public pages only. The API holds no accounts, forwards no credentials and does not sign in anywhere, so anything that lives behind a login or a paywall is out of scope no matter which protection sits in front of it. It is a reader, not a way into somebody's system.
It also is not a tool for hammering a site. Concurrency is capped per plan and requests are paced, because a crawler that degrades the target is both a legal problem and a technical one: the faster you push, the faster the rules tighten for everyone. Where a site publishes an API or a data feed, use it, since it will be cheaper and more stable than any scrape.
On the law, Cloudflare sitting in front of a page does not change the analysis. US courts have declined to treat reading a page anyone can open as unauthorized access, while obligations attach to what you collect and what you accepted. The cases are summarized in is web scraping legal, and none of it is legal advice.
Cloudflare scraping questions buyers ask
Does Cloudflare block scraping?
Cloudflare does not block scraping by default, it scores every request and lets the site owner decide. Bot Management assigns a score from 1 to 99, where 1 means Cloudflare is quite certain the request was automated and 99 means it is quite certain a human sent it. The owner sets what happens in each band, which is why the same script sails through one site and is challenged instantly on another.
How does Cloudflare detect a scraper?
It reads the TLS handshake fingerprint, the HTTP/2 frame and header order, the reputation and network type of the address, whether the client executes the challenge JavaScript, and how requests are paced. A plain HTTP library fails several of those at once. That is the reason a headless browser on the same address often gets through when curl does not, and the reason a proxy alone rarely fixes anything.
Is it legal to scrape a Cloudflare-protected website?
Cloudflare being in front of a site does not change the legal question. US courts have repeatedly declined to treat reading a page anyone can open without logging in as unauthorized access under the CFAA. What does create obligations is what you collect and what you agreed to: personal data carries duties under state privacy law, and terms you actively accepted can bind you by contract. Ask your counsel about your specific targets, because this is not legal advice.
Does cloudscraper still work?
On sites left on the lightest settings, sometimes. The cloudscraper and cfscrape family was written against the old JavaScript challenge, and it has no answer for Turnstile or a Managed Challenge. It also leaves the TLS fingerprint alone, so a site checking JA3 or JA4 refuses the connection before any challenge page is rendered. Teams usually discover this the week a target upgrades its plan.
Can Playwright or Puppeteer get past Cloudflare?
A stock instance is usually challenged within a few requests, since automation flags and a datacenter address are both visible. Patched builds do better and are a reasonable choice at low volume. The cost arrives later: you own a browser fleet, you re-patch after most Chrome releases, and you buy residential addresses separately. That recurring maintenance is what moves most teams to an API.
What is Cloudflare Turnstile and can a scraper solve it?
Turnstile is Cloudflare's CAPTCHA replacement, a widget that scores the visitor in the background and only shows an interaction when it is unsure. It is handled inside the request here, and an attempt that does not clear is retried on a new identity at no charge to you. If the data also exists on a path without the widget, reading that path is faster and steadier than fighting the challenge.
How much does a Cloudflare scraping API cost?
Here it is 49 USD a month for 50,000 successful requests, which is 0.98 USD per 1,000 pages, and 0.33 USD per 1,000 on the Scale plan. The number does not change because the target is protected. Credit based providers add a multiplier for a bypass, so on comparable entry plans a Cloudflare page there lands several times higher per 1,000. Blocked and timed out attempts are not billed on any plan.
Can I scrape a Cloudflare site on a schedule?
Yes. Scheduled jobs come with the Growth plan and above, so a price list or a competitor page can be checked hourly or daily and the rows delivered to a webhook. Scheduling also helps with block rates, because a steady drip across the address pool looks less like a crawler than the same volume fired in one burst.
Point it at the page that keeps returning a challenge
Create your account, send the URL that has been failing, and see the finished page come back on the first call.
Related endpoints
- DataDome bypass API for sites protected by DataDome or PerimeterX
- Akamai bypass API for pages behind Akamai Bot Manager
- Imperva Incapsula bypass API for Imperva and Kasada protected pages
- Web scraping proxy service: rotating residential and datacenter addresses
- Amazon scraping API for product, price and availability data
- Screen scraper API that reads the rendered page
- Datacenter proxies versus residential: which one your crawler needs
- Website crawler API for whole sites behind the same unblocking
- Web scraping API pricing and plans