Zillusion · agentic web-data platform

The web is the world's largest database. We make it queryable.

Xscent Global Inc. builds Zillusion. Describe the data you need in one sentence; agents find the sources, build and verify the scrapers, and keep the datasets fresh. Every result passes an independent verification gate first.

zillusion — replay of a real recorded runreplay
you ▸Global CO₂ emissions by country — CSV or JSON
discoverfan-out: 4 search engines · vertical registries · 23k-row API catalog (semantic)
discover13 sources found → scored on 5 dimensions · grouped APIs / files / embedded
top hitGovernment of Canada — Greenhouse Gas Emissions (CESI) · score 8.6
exploreprobing access paths (API? CSV? page table?) → writing workflow.py
validateisolated validator re-runs workflow.py in a fresh session…
validateresampled rows re-fetched → values match
validatefield meanings re-proven against the live page
gatescorecard complete → verdict: PASS
runfull crawl: 63 records · 3 CSV files · exit 0 · 9.5s clean
warehousedataset v1 + manifest saved · monitor armed → diff & notify

Web data is a maintenance treadmill.

Anyone who needs web data pays the same three taxes.

Most of the cost is discovery and maintenance — not analysis.

Discovery is a research project

The right API, file, or table is rarely the first search result. Often it's an undocumented portal API that takes days to find.

Scrapers rot

Site redesigns silently break selectors. The worst failures aren't crashes — a selector wired to the wrong field keeps returning plausible, wrong data.

“Looks right” isn't validation

Spot checks don't scale, and a model reviewing its own scraper approves it. Nobody re-checks the data against the live page.

One sentence in. A monitored dataset out.

The whole lifecycle runs as one conversation.

01

Discover

Searches 4 engines, vertical registries, and a 23,000-row API catalog with bilingual semantic retrieval. Candidates are scored on 5 dimensions and grouped as APIs, files, or embedded data.

dense embeddings + BM25 · reciprocal rank fusion
02

Build & validate

Agents probe the live site and write workflow.py, preferring APIs over brittle selectors. An isolated validator re-runs it; a gate computes the verdict.

terminal routes: deterministic · agentic · inline · infeasible
03

Warehouse & monitor

Datasets are versioned into a warehouse with incremental diffs. Schedules re-run robots; drift and failures notify you by webhook or email.

versioned datasets · scheduled re-runs · diff + notify

Engineered so it cannot lie to you.

Four bets: one hard to copy, three that deepen with use.

The interface can be copied. The verified datasets and solved walls underneath can't.

isolated · adversarialVerification a model can't sweet-talk

A separate read-only agent re-runs every scraper in a fresh session; it cannot edit what it grades. Most tools let the same agent write and approve the scraper.

gate-computedThe verdict is computed, not argued

PASS or FAIL comes from a scorecard: re-runs clean, reproducible, resample matches, field meanings re-checked against the live page. Not from a model saying it looks good.

discover + steerIt finds the source, you keep the wheel

One sentence in, and it ranks every public path to the data: API, file, or page. Most tools need you to supply the URL. At a login or captcha wall, you take over in an embedded browser, then hand back.

compounds with useEvery solved job makes the next one cheaper

Access playbooks and per-site extractors persist and become cross-site skills; a shared data layer makes the hundredth request for a dataset far cheaper than the first. The skill reuse is live today.

zillusion · workbench
Zillusion workbench: the agent asks what kind of Beijing hotel data is needed before crawling
a live task in the workbench

An honest map of the landscape.

A big, crowded, fast-moving market. Who does what, and where we differ.

~$25BData-as-a-Service market · 2025Mordor Intelligence
≈17%/yrAI web-scraping segment growthFuture Market Insights, 2026
33% by 2028of enterprise apps will use agentic AI (from <1% in 2024)Gartner, 2025
Research & coding agentsPerplexity · ChatGPT · Claude · Manus
Answer a question or produce a one-off file. None runs a repeatable, monitored data pipeline.
where they winGeneral reasoning, zero setup.
No-code scrapersBrowse AI · Octoparse · Apify
You point at a site and configure each robot; page changes mean re-tuning.
where they winMature point-and-click, huge template libraries.
Crawl & extract APIsFirecrawl · Zyte · Diffbot
Fetch or auto-extract from a URL you supply. Choosing the source, verifying, and hosting stay your job.
where they winManaged proxies, JS rendering, scale.
Data infrastructureBright Data & co.
Sell proxies and pre-collected dataset catalogs. You still pick what to buy or configure.
where they winIndustrial unblocking, compliance, delivery SLAs.
Zillusionby Xscent Global Inc.
From one sentence: find and rank sources, build the extractor, verify it against the live site, warehouse and monitor. Steerable at every step.
what's differentDiscovery + isolated verification + monitoring, in one loop.

Kadoa, Nimble, and Reworkd are racing at the same loop. Our bet: plain-language source discovery across every public path, and a PASS/FAIL computed by an isolated gate rather than claimed by the model. Firecrawl and SearXNG run as components inside Zillusion.

Where we are.

The platform runs end-to-end in production. What's left is productization, not research.

Shipped
  • End-to-end pipeline live: discover → build → validate → run → warehouse → monitor
  • Conversational orchestrator with a per-user robot library
  • 23k-row API catalog, bilingual semantic retrieval
  • Anti-bot memory reused across sites
  • Bilingual UI (EN / 中文), deployed on cloud infra
Now
  • Hardening against live sites at scale
  • Usage-metered billing on prepaid credits (design finalized)
  • Early design-partner conversations
Next
  • Hosted multi-tenant offering
  • Data-product layer: audited cleaning, reports, derived datasets
  • Managed proxy pool
zillusion · clarify
Close-up of the clarify step: the agent asks what kind of Beijing hotel data is needed and the user picks an option
it asks before it crawls
23,000+ rows in the curated API catalog5 scoring dimensions per source4 gate checks before a PASS500+ backend tests green

Xscent Global Inc.

We're building the query layer for the open web. It's early, and we're talking to investors who back infrastructure.

or write directly: baibaze35@gmail.com