AI Web Scraper That Takes Plain English
- Status
- ...
- Elapsed
- ... ms
- Exit country
- ...
- Browser used
- ...
This console sends a real request and shows you the real response. Five live fetches per session, and the sample targets run as often as you like.
Three modes, one request
| Mode | What you send | When to use it |
|---|---|---|
| Raw HTML | Nothing but the URL | You already have a parser and want the document |
| Schema | CSS or XPath per field | A page you crawl often, where the markup is stable |
| Describe it | A sentence listing the fields | A page you crawl once, or a hundred sites whose markup all differs |
Describe it reads the rendered page, finds the values you named, types them as text, number, boolean or date, and returns them under the keys you used. The same request also returns the selectors it matched, so you can promote a good result to Schema mode and stop paying the extraction step on every call.
The case it solves that selectors do not
One schema across a hundred suppliers
A marketplace enriching its catalog reads manufacturer sites that share nothing but a subject. Writing a hundred parsers is a quarter of engineering time. Describing the six fields once covers all of them.
Pages you touch once
Research crawls, one off audits and long tail sources never justify a parser. A sentence does.
Markup that changes under you
When a redesign moves a price into a different element, a description of the field still describes the field. You change nothing, and the web scraping tool keeps returning the same keys.
Fields that are not in the markup at all
Availability expressed as a sentence, a delivery promise buried in a paragraph, a discount stated only in words: those become typed fields rather than text you post process.
Honest limits
- It reads the page in front of it. If a value is not on the rendered page, no description will conjure it.
- It is slower and costs more per call than a CSS selector, so high volume crawls should be promoted to Schema mode once the selectors are known.
- It returns what it found, with the matched selector next to it, so you can verify rather than trust.
- It does not invent a value to fill a field. A field that is not present comes back empty.
Where the mode is available
Plain English extraction is included on Growth, Scale and Enterprise. Starter uses CSS and XPath schemas, which cover the pages you crawl repeatedly. The demo on the web scraping api homepage runs the Describe it mode once so you can see the result before you pick a plan. If you are comparing this with credit based LLM scrapers, our Firecrawl pricing breakdown shows what JSON extraction does to a per page price.