AI search research: original studies on who AI engines cite
This is our own AI search research, run in-house and published in full. We tell clients that
original data is the most reliably cited thing a brand can publish, so we publish ours. Every
study here states its sample, its method, its dates and what it cannot tell you. All of it is
CC BY 4.0, so you can reuse the numbers with attribution.
of the Houston real estate agents AI names appear in exactly one of 40 areas
We put 500 real estate queries about Houston to ChatGPT, Gemini (API-based), Gemini (web) and Perplexity and extracted every agent and firm each answer named. Across 1,940 answers the surfaces named roughly 4,800 firms, and 87.2% of them were named by a single surface. Full method, row-level data and a searchable list of every name.
domains cited, published as an open list with per-engine counts
The open list of every website five AI engines cited across 747 measured answers, with per-engine counts, engine reach and site type for each one. All four property portals together drew 35 citations, fewer than YouTube alone. Filterable on the page, complete as a CSV under CC BY.
cache-miss requests GPTBot got through, while every other crawler got 8 of 8
A platform-level CDN rule returns HTTP 429 to GPTBot before a request reaches origin, confirmed in writing by the host. The measured evidence, the three wrong turns it took to diagnose, the ticket, and a reproducible method for testing any host.
We put 150 real estate search questions to ChatGPT, Perplexity, Gemini, Google AI Overviews and Copilot and recorded every source each one cited. Across 747 measurements the engines cited 1297 domains, and only 3 were cited by all five. Full method, row-level data and limitations.
We asked ChatGPT, Perplexity, Google AI Overviews and Gemini a buying question about 181 United States ecommerce brands, using each brand's own category. 158 of them were named by none of the four. Full method, sample, data and limitations.
Every study starts with a real buying question rather than a brand name, because that is the
moment an engine decides whose name to hand a customer. We put that question to each engine in
turn, record what comes back, and classify the brand as cited, mentioned or absent.
Then we publish the sample, the dates, the exact wording and the parts of the method we do not
trust. Where the data is weak we say so on the page instead of burying it in a footnote.
03 Free, no call required
Run the same test on your own brand.
Two fields. We ask your buyers' real questions across ChatGPT, Gemini, Perplexity and Google AI Overviews, screenshot every answer and email the report inside 48 hours. Same method as the studies above, pointed at your site.