Data Sources & Methodology
Last updated June 24, 2026
We believe a research tool should be transparent about where its numbers come from and how confident they are. This page explains exactly that. The short version: our figures are estimates, clearly labelled, built to reveal direction and relative size — not penny-accurate counts.
Where the data comes from
- Anonymized web-activity data — large-scale, aggregated browsing signals licensed from third-party data partners. We use it to estimate AI-referral traffic to products and sites.
- AI-answer analysis — we continuously observe what AI assistants surface in shopping and research answers: which brands appear, and organic vs. paid placement.
- Public search-interest signals — relative, normalized indices of topic demand, used as a directional signal.
We deliberately don’t publish the exact list of providers, feeds, and models behind each figure — that combination is our own work. What we do commit to is being clear about how confident any given number is (see the labels below) and where its limits are.
How AI-referral estimates are produced
This is the most important thing to understand: AI assistants do not publish a log of which products they send shoppers to. The underlying web-activity data does not record a direct “ChatGPT → store” click either. Instead, AI-referral figures are inferred through our own attribution models applied to anonymized, aggregated browsing behaviour. The result is a modelled estimate of relative referral volume — useful for spotting which niches and products are gaining attention, but not a literal click count.
The ChatGPT shopping shelf
ChatGPT’s shopping answers return their own list of brands and products, which is separate from our product catalog. We match the two by brand. When a catalog product is found on the shelf, we show its organic vs. paid placement. When it is not, we say so honestly — “not on the shelf” — rather than inventing a zero. The shelf can also surface brands we don’t track in the catalog; those are labelled as such.
What the confidence labels mean
- L1 — Measured: a directly observed value (e.g., a count we can read straight from a source).
- L2 — Estimated: derived via attribution or modelling from panel/keyword data (most AI-referral and share-of-voice figures).
- L3 — Indexed / Computed: a normalized relative index (e.g., search-interest) or a figure computed from other numbers (e.g., estimated revenue = visits × conversion × AOV). Good for direction, not an absolute measured count.
We show the appropriate label next to figures so you always know how much weight to put on a number.
Cadence & coverage
- Refresh: monthly. We ingest after the closed-month dataset finalizes (around the 10th) and recompute growth, saturation, and engine mix.
- AI engines covered: ChatGPT, Gemini, Claude, Perplexity, and Copilot. Others (Grok, DeepSeek, Google AI Overview, AI Mode) are on the roadmap.
- Geographic coverage: United States and United Kingdom.
- Catalog: 25+ niches.
Limitations
- Estimates carry uncertainty and are best used for relative comparison and trend direction, not exact accounting.
- Coverage is limited to the engines, countries, and niches listed above; absence from our data is not proof of absence in the market.
- Attribution models and source coverage can change, which can shift figures month to month.
Early-access note
During early access, while we finalize data-licensing agreements with our providers, parts of the dashboard may show representative sample data to demonstrate the experience. Any such figures are part of a preview and are labelled as estimates. We will update this page when live data is fully enabled. Questions: hello@aitrends.so.