← All articlesHow to Track Competitor Rankings in AI Search

Images in this guide are free to reuse (CC BY 4.0). Credit CiteVantage with a link to this page.

How to Track Competitor Rankings in AI Search Results Effectively

How to track competitor rankings in AI search results effectively, in one paragraph: name a panel of five to eight rivals, write a fixed prompt set, run every prompt several times per engine inside a clean session, and log which brands get mentioned and which URLs get cited. Then compare share of voice month over month, not day to day.

45% of marketing leaders cannot accurately measure brand visibility in AI answers, and only 9% have the tools to track every relevant metric across platforms, per the Semrush 2026 AI Visibility Index, built on 126 million U.S. AI search prompts. Most teams are guessing about themselves. About their rivals, they are guessing harder.

Meanwhile your rank tracker still says position three. The answer sitting above that result names three other brands and never yours. Organic CTR for informational queries with AI Overviews fell 61%, per Seer Interactive’s study of 3,119 queries reported by Search Engine Land. The old scoreboard keeps measuring a thing that keeps shrinking.

What follows is the method we use to track competitor rankings in AI search results for our own clients: the prompt suite, the run count that makes a number real, the clean-room rules, and the log schema.

Key takeaways

  • A competitor number is only real if it beats your own run-to-run noise. 34.8% of the movement in LLM brand answers is pure resampling.
  • Fixed panel, fixed prompts, fixed model version. Change one thing at a time or the comparison means nothing.
  • Engines do not agree with each other. Only 11% of cited domains are shared by ChatGPT and Perplexity. Track each engine on its own line.
  • Five runs per prompt is the floor. Past the fifth repeat, the payoff collapses.
  • Log the schema, not the screenshot. A number you cannot re-run in six months is a story.

How to track competitor rankings in AI search results effectively: the six-step monthly loop, name a panel of five to eight rivals, build 40 prompts across six archetypes, set the clean room, run five times per prompt per engine, log mentions and cited URLs, then compare share of voice month over month

How do you track competitor rankings in AI search results effectively?

You run a fixed set of buyer questions through each engine, several times, in a clean session, and count who gets named. Six steps: pick the panel, write the prompt set, set the clean room, run and repeat, log mentions plus cited URLs, then compare across a fixed window.

  1. Pick the panel. Five to eight rivals, chosen once and frozen. Include the two brands you lose deals to and one smaller brand that keeps showing up. Swapping the panel mid-quarter destroys the comparison.
  2. Write 40 prompts your buyer would actually type, split across archetypes (the archetype table below has the split).
  3. Set the clean room before the first run. Fresh session, memory off, locale written down.
  4. Five runs per prompt, per engine. Same day, same conditions.
  5. Log two things every time: which brands appear, and which URLs the engine cites.
  6. Compare month over month. Never day over day.

One word to drop: rank. There is no position one in an AI answer. There is mentioned or not mentioned, cited or not cited, named first or named fifth. That is the whole scoreboard, and it behaves nothing like a ten-blue-links report. Track mentions and citations. Not positions.

Start there. The rest is bookkeeping.

What is share of voice in AI search, and how do you calculate it?

Share of voice is your brand’s mentions divided by every panel brand’s mentions, calculated per engine, per window. If your panel got named 283 times across 200 logged answers and you were named 34 of those times, your share of voice is 12%. Simple math. The value sits entirely in what you feed it.

The formula in plain words:

Your mentions ÷ all panel-brand mentions, for one engine, over one fixed window.

Here is a worked example. The numbers below are illustrative, not measured client data, and they exist to show the shape of the sheet.

BrandMentions (40 prompts x 5 runs = 200 answers)Share of voiceDistinct cited URLs
Your brand3412%6
Competitor A9634%21
Competitor B7125%17
Competitor C5218%11
Competitor D3011%5
Panel total283100%

Notice the fourth column. Competitor A is not just mentioned more, it is being pulled from three times as many sources. That gap is the actual work order, and it is far more useful than the percentage next to it.

A share-of-voice number with no run count attached is decoration. Always write the denominator on the chart: 40 prompts, 5 runs, 4 engines, this month. Without it, nobody can check you. Including you, in six months.

Share of voice is the headline metric when you benchmark AI visibility against competitors across GEO platforms, and the first number to report when you track competitor rankings in AI search results. It is also the easiest one to fake by accident, which brings us to sample size.

How many prompts and how many runs make a competitor number real?

Prompt count sets your coverage. Run count sets your confidence. Forty prompts, five runs each, per engine, is the working floor. Almost every tracking setup gets the second number wrong, which is why so many AI visibility dashboards show dramatic weekly swings that mean nothing at all.

Start with the evidence. A July 2026 variance-components study on non-determinism in LLM brand answers decomposed 12,933 responses and found that within-prompt resampling accounts for 34.8% of the variance, while the brand’s context-free true score accounts for 0.7%. Ask the same question twice and the answer changes, all by itself. The study also found that a repeat past the fifth cuts relative-error variance by about 0.0003, which is the practical argument for stopping at five.

That gives you one rule worth taping to the wall:

Do not act on a share-of-voice change smaller than your own run-to-run spread.

Finding that spread takes an afternoon. Run your existing prompt set twice on the same day, same engine, same conditions, and measure the gap between the two results. That gap is your noise floor. Anything under it is weather.

Prompt count is a coverage decision

The 50 to 100 prompt figure is coverage advice. Onely’s tracking framework sets it at 50 to 100 queries per test suite, and most vendor pages repeat some version of that number. It tells you how much of your category you can see, not how much you can trust. Forty well-chosen prompts run five times beats 100 prompts run once, every time.

Prompt wording moves the number as much as prompt count

Wording is not a detail. In a Peec AI study of 37,804 AI responses across five engines, reported by Search Engine Journal, brand mentions dropped roughly 50% when prompts fell to 0.35 to 0.39 cosine similarity with the core topic, and list or ranking prompts produced up to 20% higher average visibility. Change the phrasing and you change the result. So freeze the phrasing.

Here is the suite we build, sized at 40 prompts. It is a template, not a measured benchmark, and you rewrite the brackets for your own category.

Prompt archetypeExample prompt shapeWhat it testsHow many
Category discovery“best [product] for [buyer type]”whether you exist in the default list at all12
Head-to-head“[Competitor A] vs [Competitor B]”whether you get pulled into rival comparisons6
Buying constraint“best [product] under [budget]”, “for [condition]”filtered lists where smaller brands can win8
Problem-first“how do I fix [problem]”upper-funnel answers that name brands as solutions6
Brand-direct“is [your brand] any good”what the engine says about you unprompted4
Geo or local“best [product] in [city]”location-shaped answers, only where relevant4

“This sounds like a lot of work” is the fair objection to any honest way to track competitor rankings in AI search results. The smallest version is 40 prompts, 5 runs, 4 engines, one afternoon a month. That is 800 logged answers a month from one person with a spreadsheet and a private window.

Where the movement in an AI brand answer comes from: within-prompt resampling accounts for 34.8% of the variance, the brand's context-free true score for 0.7%, and other components for the remaining 64.5%, from an arXiv variance-components study of 12,933 responses

How do you set up a clean room so the answer you log is real?

Your own account contaminates the result. Chat memory, past conversations, saved preferences, account history and location all personalize what an engine says, so a logged-in run tells you about you, not about the market. The clean room removes the personalization so the answer you log is closer to what a stranger sees.

The protocol, every single time:

  • Fresh session or a private window. No exceptions, no “it is probably fine”.
  • Memory, personalization and custom instructions switched off.
  • One prompt per session. No prior turns in the thread, because the previous answer steers the next one.
  • Write the locale into the log. US English and AU English return different brands.
  • Same model version across the whole panel, for the whole batch.
  • Run the entire batch inside one day.

Record the model name and version on every row. Engines update silently and nobody sends you a memo, so a share-of-voice drop in March can turn out to be a model change rather than a competitor win. Without the version column you will never know which one it was.

This is the difference between “our competitor is winning in AI search” and “my browser likes my competitor”. Two very different meetings.

Clean-room checklist to run before every AI search tracking batch: fresh session or private window, memory and personalization off, one prompt per session with no prior turns, locale written into the log, same model version across the whole panel, entire batch inside one day, model name and version recorded on every row

People skip this step, then wonder why their numbers disagree with their agency’s. You cannot track competitor rankings in AI search results from a logged-in tab. Same prompts, different rooms.

What log schema makes your comparison repeatable in six months?

The schema is the deliverable. Everything else about how you track competitor rankings in AI search results is method, and method without a log is memory. If a column is missing, the comparison cannot be re-run, and a comparison you cannot re-run is an anecdote with a chart on top.

These are the exact fields we log for every single answer, one row per run.

FieldExample valueWhy it is required
date2026-07-12anchors the window and catches model changes
engineChatGPTengines disagree, so they never share a line
model name and versioncopied exactly from the model picker that daya silent update invalidates the comparison
localeUS, Englishlocation changes which brands appear
session stateclean: private window, memory offseparates market signal from personalization
prompt_idP-07lets you group runs and compare across months
prompt_text“best merino base layers for winter running”wording changes results, so store the wording
run_number3 of 5your confidence lives in this column
brands_mentionedCompetitor A, Competitor Cthe raw material for share of voice
mention_orderCompetitor A first, Competitor C thirdfirst-named is worth more than also-mentioned
cited_domainsreddit.com, wirecutter.comshows which surfaces the engine trusts here
cited_urlsfull URLs copied from the answeryour Fix / Build / Influence work order
notesrefused to rank, listed alphabeticallycatches behavior a number cannot hold

Those thirteen columns give you every view you actually need: share of voice by engine, mention rate by prompt archetype, citation source overlap between you and each rival, and month-over-month movement with the noise floor drawn on it. Same discipline we apply on the other side of the fence, in your own AI visibility audit.

Screenshots age badly. A sheet re-runs.

Why do competitors appear in AI search results when your brand does not?

Usually not authority. Usually sources. The engine is quoting places your competitor shows up and you do not, and those places barely overlap between engines, so a rival can own Perplexity while being invisible in ChatGPT. Tracking tells you which pocket you are missing.

The overlap evidence is blunt. The Digital Bloom’s 2025 AI Visibility Report analyzed 680 million-plus citations and found only 11% of cited domains are shared by ChatGPT and Perplexity. Same report: Reddit accounts for 46.7% of Perplexity’s top citations, and Wikipedia accounts for 47.9% of ChatGPT’s citations. Those are two different internets.

EngineSources cited per answerWhat it leans onWhat that means for your tracking
ChatGPTabout 15 (Semrush)Wikipedia at 47.9% of citations (The Digital Bloom)wide source net, so log more domains per answer
Perplexitynot measured in these studiesReddit at 46.7% of top citations (The Digital Bloom)forum and community presence shows up here first
Geminiabout 3 (Semrush)fewer sources per responseeach single citation carries far more weight

So the tracking output turns into a work order. For every domain citing a competitor and not you, sort it into one of three buckets:

  • Fix. You have a page on that topic and it is weak, thin, or buried under a hero image.
  • Build. No asset exists yet. Write the thing the engine keeps quoting from someone else.
  • Influence. Third-party surface you do not own: a roundup, a subreddit, a review site, a comparison page.

We see the gap in our own numbers. In our audit of 181 ecommerce brands, 72% were cited zero times across four AI engines. We have also watched DA-10 sites beat DA-90 sites inside ChatGPT answers during our ongoing citation audits, which is why “they are just bigger than us” is rarely the real explanation. Our case studies show what closing that source gap looks like in practice, and why AI skips a brand that ranks fine on Google covers the diagnosis end.

One more from Semrush’s index: only 36 brands maintained visibility across every platform. Nobody wins everywhere. Pick your engine and win there first.

ChatGPT versus Perplexity citation sources: Wikipedia is 47.9% of ChatGPT citations with about 15 sources cited per answer, Reddit is 46.7% of Perplexity top citations, and only 11% of cited domains are shared across 680 million-plus citations analyzed

Why do two AI tracking tools report different numbers?

Different prompt sets, different run counts, different regions, different model versions. Two vendors can both be honest and still hand you numbers that disagree by twenty points. Before you pick a side, check whether the disagreement is a measurement gap or just resampling.

The ten-minute spot check, before you trust any tool that claims to track competitor rankings in AI search results:

  1. Pull five prompts from the tool’s own prompt list.
  2. Run each one by hand, clean session, five times.
  3. Log mentions in your sheet, using the log schema columns above.
  4. Compare against what the tool recorded for the same day and region.
  5. If the tool cannot show you its prompt list, its run count and its region, stop. That is the finding.

Some of the gap is not error. Remember the 34.8%: a chunk of any two-tool disagreement is the model answering itself differently, not a vendor being wrong. Which is exactly why the noise floor comes first and the tool comparison comes second. For the checking questions to ask a vendor, our guide on how to verify what a tracking tool measures goes deeper on that.

How often should you track competitor rankings in AI search results?

Monthly for the scoreboard. Weekly only during a live test, and only on the prompts that test touches. A single-day swing is almost never a trend, because the mechanism that moves AI answers is slow and the noise that moves them is fast.

The reason is mechanical. Content, schema and third-party mentions take weeks to work their way into what engines quote, so a weekly share-of-voice number is mostly reading resampling. Monthly gives the slow thing time to show up. Quarterly is too slow to catch a competitor move you could still respond to.

Weekly earns its place in three cases:

  • During an active sprint, when you changed one thing and want to watch it.
  • On a subset of prompts, never the full 40.
  • When a rival has just launched something and you need the before picture fast.

Three events override the calendar entirely: a model version change, a competitor launch, and a large third-party roundup going live in your category. Any of those, re-run the batch.

The trap is reporting weekly to a client. You teach them to react to noise, then you spend the next call explaining a four-point drop that was never real. Monthly, with the noise floor on the chart.

What tools track competitor rankings in AI search results effectively, and can you do it for free?

Tools save time, not judgment. Every tool built to track competitor visibility in LLMs is doing what you can do by hand: running prompts in a session and counting mentions. So the real question is not which tool is best. Ask whether it will show you its prompt list, its run count and its region.

Free is genuinely possible. Just slower.

Tool categoryWhat it does wellWhat you still have to check yourself
Dedicated AI visibility trackers (Otterly.ai, Profound, Peec AI)run prompt sets on a schedule across several enginesthe prompt list, run count and region behind the number
SEO platforms with AI modules (SE Ranking)sit beside your existing rank and keyword datawhether their competitor set matches your real panel
Google Search Console Performance reportfirst-party, free, real clicks from AI surfaces, since AI Overviews and AI Mode traffic sits inside the Web search typeit shows your traffic, never a rival’s mentions
Manual log (the log schema above)full control of prompts, runs, locale and versionyour own time, roughly one afternoon a month

Best practices for benchmarking AI answer visibility against competitors

The zero-spend stack is three pieces: the Performance report in Google Search Console, a spreadsheet built on that log schema, and one clean-room afternoon a month. That combination will tell you more about how to track SEO effectiveness in AI search engines than most dashboards, because you control every variable that produced the number. If you want the brand-side tooling instead, our roundup of tools that track your own brand mentions in ChatGPT covers that half.

Paid tools are worth it when the time cost of the manual log exceeds the subscription. Not before. If you get to that point, the four categories above price very differently, and our breakdown of the most popular AI visibility products for SEO compares them on price per tracked prompt.

Want the competitor panel run for you first?

Building the sheet is the right long-term move if you plan to track competitor rankings in AI search results month after month. Getting a baseline this week is faster. Our free AI Visibility Snapshot checks your brand across ChatGPT, Perplexity, Gemini, Copilot, Grok and Google AI Overviews, and sends back screenshots of what those engines actually say on your buyer questions, plus which competitor gets named instead of you. No spam, and no sales call unless you ask. Start with the free AI Visibility Snapshot, or see how the same work runs for ecommerce brands end to end. Worth noting from Semrush’s AI SEO statistics: the average AI search visitor is worth 4.4x more than a traditional organic search visitor. Which is why the gap on that sheet costs more than it looks like it should.

Frequently asked questions

How do you track competitor rankings in AI search results effectively without paid tools?+

Forty prompts, five runs each, one clean session, one sheet. Name a panel of five to eight rivals, run the prompts in a fresh private window with memory off, and log every brand mentioned plus every URL cited. Add the Performance report in Google Search Console for your own side of the picture, where AI Overviews and AI Mode traffic sits inside the Web search type. That is the whole free stack.

How many queries do I need to track AI search competitors?+

Coverage and confidence are two different numbers. The 50 to 100 prompt figure vendors repeat is coverage advice, which decides how much of your category you see. Confidence comes from repeats, and five runs per prompt is the practical floor. A repeat past the fifth cuts relative-error variance by about 0.0003, per an arXiv variance study of 12,933 responses.

How long until I see results from AI visibility tracking?+

The picture arrives on day one. Movement in the picture takes weeks, because content and third-party mentions have to work their way into what engines quote before your share of voice shifts. Track monthly for the scoreboard. Weekly numbers mostly measure resampling noise, not progress.

Why do two AI tracking tools report different results?+

Different prompt sets, different run counts, different regions, different model versions. Two tools can both be honest and still disagree by a wide margin. Some of the gap is not error at all: within-prompt resampling accounts for 34.8% of variance in LLM brand answers. Ask any vendor for its prompt list, run count and region.

Why do competitors appear in AI search but my brand does not?+

Usually sources, not authority. The engines are quoting places your rival appears and you do not, and those places barely overlap. Only 11% of cited domains are shared by ChatGPT and Perplexity across 680 million-plus citations analyzed by The Digital Bloom. Being cited in one pocket of the web is not being cited everywhere.

See where AI is hiding your brand

Free multi-engine audit across ChatGPT, Gemini, Google AI & Perplexity.

Get your free audit