
Images in this guide are free to reuse (CC BY 4.0). Credit CiteVantage with a link to this page.
How to Run an AI Visibility Audit Yourself, Step by Step
An AI visibility audit checks whether ChatGPT, Perplexity, Gemini, and Google AI Overviews name your brand when buyers ask. You build a prompt set, run it across each engine, log every appearance and citation, score it, and read the gaps. The whole thing runs free. This is the exact process we use.
Most guides that promise this quietly turn into “sign up for our tool” by step two. This one does not. You will need a spreadsheet, a browser, and about an afternoon.
The short version: Build 15 to 25 buyer questions. Run each across four engines while logged out. Log who appears, who gets cited, and which sources the engines pull. Grade your llms.txt and crawler access. Turn the notes into one citability score and a share-of-voice number. Then read the gap by engine, because the fix for ChatGPT is not the fix for Perplexity.
The stakes are real. When a Google AI Overview appears, users click a traditional result in only 8% of visits, versus 15% when no AI summary shows, according to Pew Research Center’s July 2025 study of 68,879 searches. Nearly half the clicks, gone. If AI answers your buyer’s question and never names you, that traffic was never yours to lose.


And a bad first result is normal. In our own audit of 181 ecommerce brands across four AI engines, 72% were cited zero times when AI answered questions about their own product category. Most had fine SEO. They just were not built to be quoted. So if your audit comes back ugly, you are in the majority, not the exception.
Key Takeaways
- An AI visibility audit measures four things: does your brand appear, does it get cited, is the framing accurate, and what share of the answer do you own.
- The entire audit runs free. A paid tool is an optional scale-up, not a requirement to start.
- Test each engine separately. They cite differently. ChatGPT pulls around 15 sources per answer; Gemini pulls around 3.
- Include the checkpoint nobody else does: grade your llms.txt file and confirm AI crawlers can reach you.
- Low Google rankings are not why ChatGPT ignores you. Off-site authority and structured facts move that engine, not another ranking push.
What is an AI visibility audit, and how is it different from an SEO audit?
An AI visibility audit checks whether AI engines name and cite your brand in their answers. An SEO audit checks whether Google can crawl, index, and rank your pages. The engines overlap, but the win conditions do not. One rewards pages. The other rewards quotable, trusted facts about your brand.
Here is the split, side by side.
| SEO audit | AI visibility audit | |
|---|---|---|
| Core question | Can Google rank this page? | Will AI name and cite this brand? |
| Where you check | Google Search, GSC, crawlers | ChatGPT, Perplexity, Gemini, AI Overviews |
| Win condition | Position 1 to 10 | Named in the answer, with a citation |
| Main lever | On-page + backlinks | Off-site authority + structured facts |
| How you measure | Rank tracking | Appearance rate + share of voice |
The two used to move together. They are diverging now. Search Engine Journal, citing position.digital’s 2026 roundup, reports AI Overview coverage expanded roughly 58% over twelve months. The surface where buyers get answers is shifting under you, and a rank tracker cannot see it.
Step 1: Build your prompt library
Write 15 to 25 questions your buyers actually ask, in their words, not yours. Group them by intent so you can read the results cleanly later: definitional (“what is X”), comparison (“X vs Y”), and recommendation (“best X for Y”). Weight the set toward recommendation prompts. That is where a buyer meets a brand name inside an AI answer.
Skip your own brand name in most prompts. You are testing whether AI reaches for you unprompted, the way a real buyer’s session would.
Some examples for a Shopify skincare brand:
- Definitional: “what is a good vitamin C serum for sensitive skin”
- Recommendation: “best clean skincare brands for oily skin 2026”
- Comparison: “is [competitor] or [competitor] better for acne”
- Buying: “where to buy affordable vegan moisturizer online”
- Problem-led: “how do I fix flaky skin without heavy creams”
How many prompts is enough?
15 to 25 for a first pass. Below 15, one strange answer distorts your whole read. Above 25, you spend an afternoon logging before you learn much. If one intent group matters more to your revenue, load more prompts there. A DTC brand should skew recommendation and comparison, since that is the moment a purchase decision forms.
Step 2: Run every prompt across the engines
Run each prompt through ChatGPT, Perplexity, Gemini, and Google AI Overviews or AI Mode. Do it logged out or in an incognito window so your history does not shape the answer. Then run each prompt a second time on a different day, because these systems are non-deterministic and one response is a sample, not a verdict.
Testing all four is not optional busywork. The engines cite from different pools. Per a Semrush citation study reported by Search Engine Journal, ChatGPT surfaces around 15 sources per answer while Gemini surfaces around 3. Optimize for one and you can still be invisible on the rest.
Treat AI Overviews and AI Mode as two separate runs, not one. If you would rather automate that half of the work, our review of the best AI Mode SEO tracking tools covers which ones query AI Mode directly and which simply relabel AI Overviews data.

Controlling for personalization and randomness
Two traps sink most DIY audits. First, personalization: if you are logged in, the engine leans on your past chats and account signals, and you see a flattering answer no stranger would get. Log out. Second, non-determinism: the same prompt can return different brands on Tuesday and Thursday. So run twice, note what holds steady, and treat anything that appears once as unconfirmed. Our own checks re-verify facts live before we quote them, for exactly this reason. Answers drift.
Step 3: Log appearance, position, and citations in one sheet
Record one row per prompt-and-engine pair. For each, note whether your brand appeared, roughly where in the answer, whether the engine cited one of your own pages, which competitors got named, how you were framed, and every source URL the answer linked. This sheet is the audit. Everything after it is reading the sheet.
Build these columns and you have a reusable template you can rerun every month:
| Column | What it captures |
|---|---|
| Prompt | The exact question tested |
| Intent | Definitional / comparison / recommendation |
| Engine | ChatGPT, Perplexity, Gemini, AIO |
| Date | Run date (you run twice) |
| Appeared? | Yes / No |
| Position | Top, middle, or buried in the answer |
| Own page cited? | Yes / No, plus the URL |
| Competitors named | Every rival the answer mentioned |
| Sentiment | Positive, neutral, or wrong |
| Source URLs | Every link the engine cited |

Reading the citations
Pull every URL from the “Source URLs” column into one list, then sort it into three buckets: your own pages, neutral third parties (Reddit, review sites, editorial roundups), and competitor pages. That split tells you where the engines’ trust actually lives. Most brands find their own domain barely shows, and third-party pages do the heavy lifting. If you want a quick sanity check on one engine before the full run, our guide on how to check whether ChatGPT recommends your brand walks a single-engine version.
Step 4: Grade your AI-crawler access and llms.txt
Before you blame your content, confirm the engines can even reach it. A blocked crawler makes the best page on the internet invisible. This is the checkpoint most audit guides skip entirely, and it is often the whole problem.
Run this short list:
- Open robots.txt and confirm it does not block GPTBot, Google-Extended, ClaudeBot, or PerplexityBot.
- Load a key page with JavaScript disabled. If the content vanishes, crawlers may see nothing either.
- Check that product, FAQ, and review schema are present and valid on your money pages. If you are not sure what to look for, here is how to audit your schema markup for AI search.
- Confirm your pages return real HTML, not a client-side shell.
Fetch and grade your llms.txt
Go to yourdomain.com/llms.txt and see what loads. This file is a plain-text map of your most important pages, written for AI systems. Almost no competitor audit checks it. We grade it on every audit as a standing rule. A good one lists your key pages with short, accurate descriptions. A missing or empty one is a fast, cheap fix. It is not a silver bullet, but leaving it blank is a signal you skipped the basics. Our explainer on what an llms.txt file is and how to build one covers the format.
Step 5: Score it into one citability number
Turn the scattered notes into a single 0 to 100 score you can track month over month. Weight five inputs from your sheet: appearance rate, citation rate, accuracy of framing, position in the answer, and share of voice. One number per audit, one trend line over time. That is how you tell progress from noise.
Here is the rubric we use.
| Input | What it measures | Weight |
|---|---|---|
| Appearance rate | % of prompts where you showed at all | 30 |
| Citation rate | % where your own page was cited | 25 |
| Share of voice | Your mentions vs all brand mentions | 20 |
| Accuracy / sentiment | Right facts, positive framing | 15 |
| Position | How high in the answer you land | 10 |

How to measure AI share of voice
Count how many times your brand is named across the whole prompt set. Divide that by the total brand mentions, yours plus every competitor. That percentage is your share of voice. Triple Whale’s published benchmarks put a challenger brand at 10 to 20%, a growing brand at 15 to 30%, and an established category leader above 40%. Track the number, not the feeling.
Step 6: Read the gap by engine logic
The same low score means a different fix on each engine. That is the step every generic audit misses. Because the engines source their answers so differently, chasing Google rankings will move one of them and do nothing for another. Read the gap through each engine’s real logic before you spend a dollar fixing it.
The clearest proof is overlap with Google. Per Semrush’s July 2025 study of 5,000 keywords and over 150,000 citations, Perplexity’s cited domains overlap Google’s top 10 by about 91%, Google AI Mode by about 54%, and ChatGPT the lowest of the four. So better rankings can lift Perplexity while barely touching ChatGPT. We have watched DA-10 sites beat DA-90 sites inside ChatGPT answers, which is the same lesson from the other direction.

Why isn’t my brand showing up in ChatGPT?
Usually because ChatGPT does not lean on Google rankings the way you assume. It rewards structured facts about your brand and mentions on sources it already trusts. If you rank well but ChatGPT ignores you, the lever is off-site authority and clean entity data, not another push for position one. Our diagnosis of why brands go invisible to AI breaks down the common causes.
Step 7: Turn findings into a Fix / Build / Influence list
Every finding maps to one of three actions. Fix what is broken, build what is missing, influence what lives off your site. Sort the list by payoff, not by ease, and work the top of it first.
- Fix: unblock crawlers in robots.txt, add missing product and FAQ schema, publish or repair your llms.txt. Fast wins that remove hard blocks.
- Build: write answer-shaped capsules on money pages, publish comparison pages, structure your facts so AI can lift them cleanly.
- Influence: earn third-party mentions and reviews. These engines lean on what neutral sources say about you, so your presence on review sites, roundups, and community threads is an AI visibility lever, not just a storefront number.
How often should you re-run the audit?
Monthly. These metrics move faster than SEO ever did. BrightEdge tracked AI Overview overlap with Google’s organic top 10 climbing from 32% to about 54% over 16 months. A quarterly cadence misses shifts you could have acted on. Monthly keeps your score honest and your fixes current. If you would rather have that re-run happen on a schedule, our independent review of the AEO tools that track brand mentions in ChatGPT covers what the free checkers and the paid trackers each buy you.
Done right, this pays off. The Princeton-led GEO paper (Aggarwal et al., arXiv 2311.09735) found that generative-engine optimization can lift a page’s visibility in AI responses by up to 40%. The audit is how you find which pages have that room to gain.
If you would rather not run all seven steps by hand, our free AI Visibility Snapshot does steps one through five in one pass and shows you where you stand across all four engines. Or you can hand it to a GEO agency instead and skip the afternoon entirely. No price, no pitch, just the baseline. And if ecommerce is your lane, our ecommerce AI visibility service picks up where the audit leaves off. The RIPT Apparel case study shows what that next step produced for one store.
Frequently asked questions
Can you run an AI visibility audit for free?
Yes, the whole thing. Every engine we test answers on a free tier, and a spreadsheet holds the log. A paid tool earns its place only once you are tracking hundreds of prompts across many competitors on a schedule. For a first audit, free covers it.
How many prompts should you test?
15 to 25 for a first pass. Fewer, and one odd answer skews the picture. More, and you log for days before learning much. Weight the set toward recommendation prompts, because that is where buyers actually meet your brand in an answer.
How do you avoid personalization bias when testing AI answers?
Log out or use incognito, so past chats do not color the result. Run each prompt twice on different days. AI answers are non-deterministic, so one response is a sample, not the truth. Two runs separate what is stable from what is noise.
How often should you run an AI visibility audit?
Monthly. These numbers move fast. BrightEdge tracked AI Overview overlap with Google’s top 10 rising from 32% to about 54% over 16 months. The engines change what they cite week to week, so monthly catches shifts a quarterly check would miss.
How do you measure AI share of voice?
Count your brand mentions across the whole prompt set, then divide by total brand mentions, yours plus every competitor. That percentage is your share of voice. For a challenger, 10 to 20% is a normal band. Established leaders often sit above 40%, per Triple Whale benchmarks.
Frequently asked questions
Can you run an AI visibility audit for free?+
Yes, the whole thing. Every engine we test, ChatGPT, Perplexity, Gemini, and Google AI Overviews, answers questions on a free tier. A spreadsheet holds the log. A paid tool only helps once you are tracking hundreds of prompts across many competitors on a schedule. For a first audit, free is enough.
How many prompts should you test?+
15 to 25 for a first pass. Fewer and one odd answer skews the whole picture. More and you are logging for days before you learn anything. Weight the set toward recommendation-intent prompts ("best X for Y"), because that is where buyers actually meet your brand inside an AI answer.
How do you avoid personalization bias when testing AI answers?+
Log out, or use an incognito window, so past chats and account history do not color the result. Run each prompt twice on different days too. AI answers are non-deterministic, so a single response is one sample, not the truth. Two runs tell you what is stable versus what is noise.
How often should you run an AI visibility audit?+
Monthly. These numbers move fast. BrightEdge tracked AI Overview overlap with Google's top 10 climbing from 32% to about 54% over 16 months. The engines change what they cite week to week, so a quarterly check misses shifts that a monthly one catches while you can still act on them.
How do you measure AI share of voice?+
Count your brand mentions across the whole prompt set, then divide by the total brand mentions (yours plus every competitor named). That percentage is your share of voice. For a challenger, 10 to 20% is a normal starting band. Established category leaders often sit above 40%, per benchmarks Triple Whale publishes.
What is the difference between an SEO audit and an AI visibility audit?+
An SEO audit asks whether Google can rank your pages. An AI visibility audit asks whether ChatGPT, Perplexity, and Gemini will name and cite your brand when someone asks. Different win conditions. You can rank on page one and still never appear in a single AI answer about your own category.
See where AI is hiding your brand
Free multi-engine audit across ChatGPT, Gemini, Google AI & Perplexity.
Get your free audit