Zillusion · AI web-scraping & data agent

The web is the world's largest database. We make it queryable.

Describe the data you need in one sentence; AI agents do the web scraping — finding sources, building and verifying scrapers, keeping datasets fresh. Every result passes an independent verification gate first.

zillusion — replay of a real recorded runreplay
you ▸Global CO₂ emissions by country — CSV or JSON
discoverfan-out: 4 search engines · vertical registries · 23k-row API catalog (semantic)
discover13 sources found → scored on 5 dimensions · grouped APIs / files / embedded
top hitGovernment of Canada — Greenhouse Gas Emissions (CESI) · score 8.6
exploreprobing access paths (API? CSV? page table?) → writing workflow.py
validateisolated validator re-runs workflow.py in a fresh session…
validateresampled rows re-fetched → values match
validatefield meanings re-proven against the live page
gatescorecard complete → verdict: PASS
runfull crawl: 63 records · 3 CSV files · exit 0 · 9.5s clean
warehousedataset v1 + manifest saved · monitor armed → diff & notify

Web data is a maintenance treadmill.

Anyone who needs web data pays the same three taxes.

Most of the cost is discovery and maintenance — not analysis.

Discovery is a research project

The right API, file, or table is rarely the first search result. Often it's an undocumented portal API that takes days to find.

Scrapers rot

Site redesigns silently break selectors. The worst failures aren't crashes — a selector wired to the wrong field keeps returning plausible, wrong data.

“Looks right” isn't validation

Spot checks don't scale, and a model reviewing its own scraper approves it. Nobody re-checks the data against the live page.

One sentence in. A monitored dataset out.

The whole lifecycle runs as one conversation.

01

Discover

Searches 4 engines, vertical registries, and a 23,000-row API catalog with bilingual semantic retrieval. Candidates are scored on 5 dimensions and grouped as APIs, files, or embedded data.

dense embeddings + BM25 · reciprocal rank fusion
02

Build & validate

Agents probe the live site and write workflow.py, preferring APIs over brittle selectors. An isolated validator re-runs it; a gate computes the verdict.

terminal routes: deterministic · agentic · inline · infeasible
03

Warehouse & monitor

Datasets are versioned into a warehouse with incremental diffs. Schedules re-run robots; drift and failures notify you by webhook or email.

versioned datasets · scheduled re-runs · diff + notify

Engineered so it cannot lie to you.

Four things you'll notice from your first run.

Every dataset arrives with the evidence it was verified against — you can defend it, not just trust it.

isolated · adversarialVerification a model can't sweet-talk

A separate read-only agent re-runs every scraper in a fresh session; it cannot edit what it grades. Most tools let the same agent write and approve the scraper.

gate-computedThe verdict is computed, not argued

PASS or FAIL comes from a scorecard: re-runs clean, reproducible, resample matches, field meanings re-checked against the live page. Not from a model saying it looks good.

discover + steerIt finds the source, you keep the wheel

One sentence in, and it ranks every public path to the data: API, file, or page. Most tools need you to supply the URL. At a login or captcha wall, you take over in an embedded browser, then hand back.

compounds with useEvery solved job makes the next one cheaper

Access playbooks and per-site extractors persist and become cross-site skills; a shared data layer makes the hundredth request for a dataset far cheaper than the first. The skill reuse is live today.

zillusion · workbench
Zillusion workbench: the agent asks what kind of Beijing hotel data is needed before crawling
a live task in the workbench

How it compares to tools you may know.

Different tools solve different jobs. Here's what changes in practice with Zillusion.

Research & coding agentsPerplexity · ChatGPT · Claude · Manus
Answer a question or produce a one-off file. None runs a repeatable, monitored data pipeline.
where they winGeneral reasoning, zero setup.
No-code scrapersBrowse AI · Octoparse · Apify
You point at a site and configure each robot; page changes mean re-tuning.
where they winMature point-and-click, huge template libraries.
Crawl & extract APIsFirecrawl · Zyte · Diffbot
Fetch or auto-extract from a URL you supply. Choosing the source, verifying, and hosting stay your job.
where they winManaged proxies, JS rendering, scale.
Data infrastructureBright Data & co.
Sell proxies and pre-collected dataset catalogs. You still pick what to buy or configure.
where they winIndustrial unblocking, compliance, delivery SLAs.
Zillusionby Xscent Global Inc.
From one sentence: find and rank sources, build the extractor, verify it against the live site, warehouse and monitor. Steerable at every step.
what's differentDiscovery + isolated verification + monitoring, in one loop.

Firecrawl and SearXNG run as open-source components inside Zillusion — we build on them rather than against them.

Read the detailed comparisons: Zillusion vs Octoparse · vs Browse AI · vs Firecrawl · vs Bright Data · all comparisons

Live in production.

The numbers behind the pipeline you just read about.

23,000+ rows in the curated API catalog5 scoring dimensions per source4 gate checks before a PASS500+ backend tests green
zillusion · clarify
Close-up of the clarify step: the agent asks what kind of Beijing hotel data is needed and the user picks an option
it asks before it crawls

Credits, priced plainly

One-time prepaid credit packs for verified web-data work. No subscription and no automatic renewal.

Loading current pricing…

Talk to us.

Stuck on a hard dataset, blocked by a site, or wondering if Zillusion fits your use case? Write us — the founder reads every email.

or write directly: support@zillusion.app