Discovery is a research project
The right API, file, or table is rarely the first search result. Often it's an undocumented portal API that takes days to find.
Xscent Global Inc. builds Zillusion. Describe the data you need in one sentence; agents find the sources, build and verify the scrapers, and keep the datasets fresh. Every result passes an independent verification gate first.
Anyone who needs web data pays the same three taxes.
Most of the cost is discovery and maintenance — not analysis.
The right API, file, or table is rarely the first search result. Often it's an undocumented portal API that takes days to find.
Site redesigns silently break selectors. The worst failures aren't crashes — a selector wired to the wrong field keeps returning plausible, wrong data.
Spot checks don't scale, and a model reviewing its own scraper approves it. Nobody re-checks the data against the live page.
The whole lifecycle runs as one conversation.
Searches 4 engines, vertical registries, and a 23,000-row API catalog with bilingual semantic retrieval. Candidates are scored on 5 dimensions and grouped as APIs, files, or embedded data.
dense embeddings + BM25 · reciprocal rank fusionAgents probe the live site and write workflow.py, preferring APIs over brittle selectors. An isolated validator re-runs it; a gate computes the verdict.
terminal routes: deterministic · agentic · inline · infeasibleDatasets are versioned into a warehouse with incremental diffs. Schedules re-run robots; drift and failures notify you by webhook or email.
versioned datasets · scheduled re-runs · diff + notifyFour bets: one hard to copy, three that deepen with use.
The interface can be copied. The verified datasets and solved walls underneath can't.
A separate read-only agent re-runs every scraper in a fresh session; it cannot edit what it grades. Most tools let the same agent write and approve the scraper.
PASS or FAIL comes from a scorecard: re-runs clean, reproducible, resample matches, field meanings re-checked against the live page. Not from a model saying it looks good.
One sentence in, and it ranks every public path to the data: API, file, or page. Most tools need you to supply the URL. At a login or captcha wall, you take over in an embedded browser, then hand back.
Access playbooks and per-site extractors persist and become cross-site skills; a shared data layer makes the hundredth request for a dataset far cheaper than the first. The skill reuse is live today.

A big, crowded, fast-moving market. Who does what, and where we differ.
Kadoa, Nimble, and Reworkd are racing at the same loop. Our bet: plain-language source discovery across every public path, and a PASS/FAIL computed by an isolated gate rather than claimed by the model. Firecrawl and SearXNG run as components inside Zillusion.
The platform runs end-to-end in production. What's left is productization, not research.

We're building the query layer for the open web. It's early, and we're talking to investors who back infrastructure.
or write directly: baibaze35@gmail.com