An AI web scraper for sites without a dedicated tool
Describe what you want, point it at any URL, get JSON.
The Kavex AI web scraper handles the pages no fixed scraper was built for. Paste a list of URLs, describe the fields you want in plain language, and it returns structured data per page, no selectors, no XPath, no code. It is the tool for long-tail directories, supplier catalogues and niche sites where building a custom scraper would never pay off. You pay per page processed.
A RUN OF 1,000 PAGES ON THIS SOURCE, AT 4 PER PAGE
Five columns on every row.
These come with the run and cost nothing extra. Add-ons widen the same sheet, and you pay only for the ones you switched on.
| URL | Product | Price | Lead time | MOQ |
|---|---|---|---|---|
| supplier-a.com/p/1042 | Steel bracket M8 | listed on request | 3 weeks | 500 |
| supplier-b.de/katalog/77 | Aluminium profile 40x40 | per metre | 10 days | 100 |
| supplier-c.io/items/std-9 | Nylon spacer kit | bundle | 2 weeks | · |
| supplier-d.fr/ref/x21 | Brass fitting 1/2" | per unit | 4 weeks | 250 |
Say what you want. Read the price first.
Everything else is a ceiling or a filter. The estimate moves with the numbers you type, so the run never surprises you at the end.
The AI web scraper lets you define a dataset in a sentence. Instead of pointing at HTML elements, you write what you want, for example "product name, price, lead time and minimum order quantity", and the tool extracts those fields from every URL you give it.
It returns clean, structured rows that match the fields you asked for, so the output of a hundred different pages lines up in one consistent table. Pages that genuinely do not contain a field return it empty rather than guessing, so you can trust the gaps.
This is what makes it a fit for the long tail. A dedicated scraper exists for Google Maps or LinkedIn because those sites are worth the engineering. For a regional manufacturer directory or a one-off research list, the AI web scraper gives you the same structured result without any of that build.
The trade-off to understand is consistency. A dedicated scraper for Google Maps returns identical fields every time because the site is known; the AI web scraper works across pages it has never seen, so the structure of a site still matters. Clear, visible data extracts cleanly, while information buried in images or behind interactions is harder to reach. In practice that makes it ideal for text-rich directories, catalogues and listing pages, and less suited to heavily interactive apps. Used for what it is good at, it removes the main reason most niche datasets never get built, the cost of engineering a one-off scraper, and turns a list of awkward URLs into a normal spreadsheet.
The AI web scraper fetches each page live, condenses its content down to the meaningful text and structure, and passes that to Google Gemini along with your plain-language field description. Gemini returns the values as structured JSON, which is mapped into the columns you asked for.
Because extraction is driven by an instruction rather than a fixed template, the same job works across pages with completely different layouts. Sites that block plain requests are reached through rotating residential proxies, so a varied list of real-world URLs comes back complete.
Charged per delivered row. If the source holds fewer than you asked for, the rest of the credits go back to your balance the second the run ends.
One thousand pages, itemised.
One line, because this source has nothing bundled on top. You pay the rate times the rows the run actually delivered.
| Line | Rate | Units | Credits |
|---|---|---|---|
| AI Scraper | 4 / page | 1,000 pages | 4,000 |
| 1,000 pages | 4 / page | 4,000 credits | $4.00 |
$0.004 PER FINISHED PAGE · ONLY DELIVERED ROWS ARE BILLED
What people point it at.
Four ways this source is used today. Every one of them is the same run, the same export and the same credit pool.
- Researchers pulling structured data from a niche industry directory that no standard tool covers.
- Operations teams building a one-off supplier or catalogue dataset from a list of obscure URLs.
- Analysts assembling a custom comparison set from pages that all present data differently.
- Founders prototyping a new scraping idea quickly before committing to a dedicated vertical tool.
The rest of the web, contact & ai stack.
Same credit pool, same export. Chain the output of one run into the next without paying for the list twice.
Asked before the first run.
Short answers, no accordion to click through. If something is still missing, the help pages go deeper.
What kinds of sites does it work on?
It works on almost any page with readable content, directories, catalogues, listings and profile pages. It is most useful for long-tail sites where no dedicated scraper exists and building one would not be worth it.
How do I tell it what to extract?
You write the fields you want in plain language, like "company name, founding year and main product". The clearer and more specific the description, the more precise the extracted columns.
What format is the output?
Each page returns structured data that maps to the fields you described. The job downloads as a CSV with one row per URL and one column per field you asked for.
What happens if a page is missing a field?
If a page genuinely does not contain a requested field, that cell is left empty rather than filled with a guess, so the gaps in your dataset are real and trustworthy.