SEPTEMBER 18, 2026

AI Visibility Audit: What You Can Actually Measure

Marketing & Growth
Dhawal Shah
Dhawal Shah

14 years building businesses across Asia. Co-founded 2Stallions (40+ person agency), launched ChutneyAds (AI-powered ad network), and has worked with 30+ startups as advisor and investor. SID Accredited Director (Singapore Institute of Directors). He writes from the operator side of the table.

The Number Your AI Visibility Tool Gives You Is Mostly Noise

Ask an AI engine the same question twice in a day and most of its sources will change. A 2026 review of the research measured this across four engines over 45 days. Repeat a question, even within 24 hours, and only about a third of the cited sources come back, a source overlap of 0.34 to 0.42 (arXiv, 2026). Every AI visibility score on sale is built on top of that movement. The tool reporting it is usually working fine.

TL;DR: No AI visibility score can be compared between tools, because each one asks different questions of different models. A repeated question returns only about a third of the same sources (arXiv, 2026). Google and Bing now publish their own AI data free. Most of what AI engines cite about a brand is written by someone else. Start with the free reports and with where you are being discussed. A paid seat can wait.

Two words come up throughout, so here is what I mean by them. An engine “cites” you when it lists your page as a source under its answer. It “mentions” you when your name appears in the answer itself. The two can happen together or apart. A good deal of the muddle in this category comes from treating them as one thing.

I run a 40-person digital agency in Singapore and I build the measurement systems we use for clients. I also publish my own site’s numbers below, zeroes included. A small site with honest data is more use to you than a case study with a happy ending. This is the measurement companion to my piece on moving from SEO to AI engine optimisation. Nothing here is a product pitch.

What Is an AI Visibility Audit?

An AI visibility audit checks whether AI assistants mention your business when people ask about what you do. The paid versions add prompt tracking, competitor comparison, sentiment scores and a headline number. I have built one for my own agency, so none of this is criticism from the outside. Strip the dashboard away and a useful audit answers four questions.

  • Do the engines mention you at all?
  • Which of your pages do they cite?
  • Which questions trigger those citations?
  • How does that compare with the companies you compete against?

AI brand visibility is the part most buyers care about, and it behaves least like ordinary search. In Google, a ranking is a position on a page. You are third, or you are eleventh. In an AI answer there is no page. You are named in the text, listed as a source underneath, both, or absent, and those are four different states. Most tools squash them into one score.

Two-by-two grid of the four states a brand can be in inside an AI answer: named and cited, cited only, named only, and absent, shaded darker for more visible
The four states an AI visibility audit should keep separate. Darker means more visible.

A fifth question sits outside the usual audit template, and it has its own section further down. When an engine does mention you, where did it find you? On my own site the answer was mostly somewhere other than my own site.

Whether any of this is worth paying for also comes up below. My short answer is to read the two free reports first, because they cover three of the four questions at no cost. The fourth, the competitor comparison, is the one a paid tool genuinely adds. Even there, Bing now gives you a share-of-citations figure for nothing.

Citation capsule

An AI visibility audit measures whether AI assistants name a brand, cite its pages, or leave it out when answering questions about its category. Being named and being cited are separate states. Most commercial tools merge them into one score, so two audits of the same site rarely agree. (Dhawal Shah, 2026)

What Is a Good AI Visibility Score?

There is no good AI visibility score, because no two tools measure the same thing. A 42 from one vendor and a 71 from another are readings from two different instruments. Neither vendor can honestly tell you what number to aim for.

A score is an average across a set of prompts, the questions the tool types into the engines on your behalf. No two vendors use the same set. Swap the prompts and the number moves, with nothing about your site having changed. The vendors also differ on which models they ask. ChatGPT, Claude, Gemini, Perplexity and Copilot each find and cite sources in their own way, so a score blended across four of them hides which one actually moved. On top of that sits the sampling problem from the top of this article. Ask the same engine the same question twice and two-thirds of the sources are different.

Building the measurement model for our own internal system forced me to write this down as a rule. Our definitions now say that changing the provider, the model, the source or the prompt wording starts a new measurement series. The old series cannot be compared with the new one unless someone writes an explicit rule for how. It sounded fussy when we drafted it. The alternative is drawing a trend line through two unrelated numbers and presenting it to a client as progress, and I have seen that done.

What you can compare is your own count, on your own fixed prompts, on one engine, over months. That is a much smaller claim than a score. In my experience it is also the only one that survives a second look.

Why Do AI Visibility Tools Disagree on LLM Visibility?

Because LLM visibility is not a fixed property of your website in the way a Google ranking is. It is the outcome of a conversation between a model, a search step and a question, and two of those three change every time you ask.

A single 100% bar: the first 34% to 42% is filled purple for sources that came back when the same question was repeated; the remaining 58% to 66% is grey for sources that changed
Source overlap on repeated questions, four generative engines over 45 days. Source: arXiv, 2026.

Ordinary rank tracking works because a Google results page is close to fixed. Two people in the same city searching the same term see nearly the same page, in nearly the same order. The whole rank-tracking industry was built on that stability without anyone having to think about it.

AI answers have no such stability. A large language model, the software behind ChatGPT or Gemini, picks its words with a small amount of randomness built in on purpose. The same question produces slightly different answers by design. The search step that feeds it runs fresh each time, and the index underneath shifts daily. Ask on Tuesday and again on Wednesday and you have two draws from a hat rather than two points on a line.

The research puts a number on it. Across four engines over 45 days, only about a third of the sources cited for a question came back when it was asked again, a source overlap of 0.34 to 0.42 (arXiv, 2026). The figure held even for repeats within a single day.

So when a client shows me a chart where their AI visibility fell 20% in a week, my first question is what changed structurally. Usually nothing did. Which pages get cited, and by which engine, is a readable signal. Week-on-week percentages are not, and I would rather say so than invent an explanation for them.

Citation capsule

Research across four generative engines over 45 days found that a repeated question returned only about a third of the same cited sources, a source overlap of 0.34 to 0.42, even within 24 hours. Generative answers are sampled rather than fixed, so a week-on-week change in an AI visibility score is usually sampling noise. (arXiv, 2026)

Why Does Your AI Visibility Change When Your Website Has Not?

Because the answer you see is built from a live web search, not from the model’s memory. The memory is a snapshot taken months ago. The search happens in the second before the engine replies, and the search is what usually decides whether you appear.

Every model has a training cutoff, the date after which it learned nothing new. The cutoffs are real and they are recent. Claude Opus 5 stops at May 2026 (Anthropic, 2026). Gemini 3.7 Flash stops at March 2026 (Google DeepMind, 2026). GPT-5.6 Sol stops at 16 February 2026 (RankScope, 2026).

If those dates decided visibility, a page published in June could not appear in a ChatGPT answer until the next model shipped. It does appear, because the engines fetch live information before they answer. Each does it differently. Gemini grounds its answers in live Google Search by default, so it runs a Google search and reads the results before it writes. ChatGPT browses selectively, through Bing. Claude reaches for a web search tool when the question calls for it rather than browsing every time. Perplexity is search-first by design, so every answer starts with a search.

Your visibility therefore depends mostly on each engine’s live retrieval, and on what it happens to pull in that moment. The two free reports covered below come from the search layers of Google and Bing, the layer where that decision gets made.

It also tells you which other platforms are worth watching. There is no first-party report for ChatGPT, Perplexity or Claude, so any number you see for them is an estimate from sampling prompts, never a reading from a log. If you do sample, I would go in this order. ChatGPT first, because it is the most used consumer AI app. Perplexity next if your audience behaves like searchers. Claude if your audience skews technical. Treat it as a priority order rather than a shopping list, and label every number that comes out of it as an estimate.

Which Free AI Visibility Tools Actually Exist?

Two free AI visibility tools exist, and both come from the two largest search engines. The tool listicles tend to skip them. I suspect that is because neither one pays an affiliate commission.

Bing Webmaster Tools shipped its AI Performance report as a public preview on 10 February 2026. No major engine had published its own AI citation data to site owners before. In June it added Intents, Topics, Citation Share and Compare (Bing Webmaster Blog, 2026). It is free, and there is no paid tier to be upsold into.

Google announced its generative AI performance report on 3 June 2026 and finished rolling it out worldwide on 31 August (Google, 2026). Also free, and it sits inside the Search Console account you already have.

“First-party” needs a plain definition, because the whole argument rests on it. A first-party report is the engine telling you what it did with your pages. A third-party tool cannot see inside the engine, so it works the other way round. It types sample questions in, counts what comes back, and infers the rest. You would not run a customer survey if you already had the sales ledger.

Before you buy a seat on an AI visibility tool, verify your site in both reports and read three months of data. If a question is still unanswered after that, you now know exactly what you are paying to find out, and you can hold the vendor to it. Most teams I have watched do it the other way round. They buy a seat first and never open either report, so they end up paying for a survey of a system whose own logs were sitting in a free tab.

Citation capsule

Google and Bing publish first-party AI visibility data free. Bing Webmaster Tools’ AI Performance report has been in public preview since 10 February 2026 and added Intents, Topics, Citation Share and Compare in June. Google’s generative AI performance report, announced 3 June 2026, reached all Search Console users worldwide on 31 August. Neither has a paid tier. (Bing Webmaster Blog, 2026)

What Do Google and Bing’s Free AI Reports Give You?

Google’s report gives you impressions only; Bing’s gives you cited pages, the queries behind them and your citation share. Both give you less than you want and more than the free-tools lists admit, and their limits are honest ones.

An impression, in Google’s report, means one of your pages was shown inside an AI answer. Just that. No clicks, no click-through rate, no record of which question triggered it, no position. You learn that you appeared, not whether anyone did anything about it. Google has said more metrics will follow. Until then I treat it as a presence check.

Bing’s is the more useful of the two today. It names the pages Copilot cited, lists the search queries that led to each citation, and shows your share of citations against every other site named in the same answers. That is closer to a real audit than most paid dashboards manage, and it costs nothing.

Matrix comparing Google Search Console and Bing Webmaster Tools AI reports: both show which pages appeared and how often; only Bing shows the queries behind them and citation share; Google reports no clicks, click-through rate or position but can be grouped by country; Bing has no country view
What the two free first-party reports give and withhold. Google Search Console generative AI report; Bing Webmaster Tools AI Performance report.

The two also differ on geography, and that matters if you sell in more than one country. Google’s report can be grouped by the country the search came from, so you can ask “how do I look in Malaysia versus Singapore”, at least for impressions. Bing’s report has no country breakdown that I can find, and the consumer apps for ChatGPT, Perplexity and Claude offer nothing at all. The odd part is that the localisation exists under the hood. At the developer level, OpenAI, Anthropic, Google and Perplexity all let a query carry the user’s location, so the engines know where the asker is. Only Google turns that into a report a brand can read. For every other engine, a business selling in several markets is guessing about all but one of them, and I include my own agency in that.

Citation capsule

Google’s generative AI performance report, worldwide since 31 August 2026, carries impressions only: no clicks, no click-through rate, no queries and no position. Google’s report can be grouped by country; Bing Webmaster Tools, live since 10 February 2026, reports cited pages, the queries behind them and citation share, with no country view. (Google, 2026)

Can Google and Bing’s Reports Tell You Anything About ChatGPT Visibility?

No. Neither report covers ChatGPT, and the two do not even agree with each other about my own website. The size of that disagreement changed how I read all of this data.

The table below is the same site over the same 92 days, 23 May to 22 August 2026, AI surfaces only. Bing counts the times Copilot cited a page as a source. Google counts the times a page was shown inside one of its AI features.

PageBing AI citationsGoogle AI impressions
/article/ai-fluency-for-directors/6311
/hermes-agent/017
/ (homepage)1016
/article/openclaw-vs-claude-managed-agents/126
/apac-ai-adoption-2026/106
/about/211
/services/digital-marketing-agency/09
/article/ai-agent-board-governance-singapore/81
/article/ai-agents-for-business-2026/60
/article/ad-platform-mcp-claude/40

Six pages are cited only by Bing. Twelve appear only on Google’s AI surface. The top page on each engine is close to invisible on the other. My most-cited page on Bing, an explainer on AI fluency for directors, has 63 citations and 11 Google AI impressions. Google’s favourite, the Hermes agent page, has 17 impressions and zero Bing citations.

I do not think that is random. Google’s AI surface leans towards pages about who I am and what I sell: the homepage, the about page, the services page, the agent page. Copilot cites the articles. Two engines read the same site and came away with opposite views on which parts were worth quoting, and I cannot tell you which of them is right.

A single site-wide number would average those two views into something that describes neither. I track them as two separate figures. When a client asks for one blended AI visibility number, I show them this table and explain why they are not getting one.

Grouped bar chart of ten dhawalshah.net pages, Bing Copilot citations against Google AI impressions, 23 May to 22 August 2026: the AI fluency explainer has 63 Bing citations and 11 Google impressions while the Hermes agent page has 0 Bing citations and 17 Google impressions
The same ten pages, two engines, 23 May to 22 August 2026. Sorted by Bing citations.
Citation capsule

On dhawalshah.net over the 92 days to 22 August 2026, six pages were cited only by Copilot and twelve appeared only on Google’s AI surfaces. Bing’s most-cited page (63 citations) earned 11 Google AI impressions; Google’s top page (17 impressions) earned zero Bing citations. Google favoured the homepage, about and services pages; Copilot cited articles. (Dhawal Shah, 2026)

What Does AI Brand Visibility Look Like on a Small Site?

On a small site, AI brand visibility is small and uneven, nothing like the case studies. Every other page competing for this term shows a win, so here is the working.

Over the same 92 days, my Hermes agent page earned 1,479 ordinary Google search impressions, the most on the site. On AI surfaces it earned 17 Google AI impressions and zero Bing citations. The AI fluency explainer earned 78 ordinary search impressions, the lowest of the ten pages in the table above. It earned 63 Bing citations, 52% of the site’s total, plus 11 Google AI impressions. My most-seen page in ordinary search is my least-cited in AI, and the reverse.

Site-wide, the shape is the same: 5,932 ordinary search impressions in those 92 days against 89 Google AI impressions. That is a ratio of 1.5%, and it is a ratio of two kinds of impression, not a conversion rate.

Bing’s page report recorded 122 Copilot citations for roughly the same window; its overview report says 76. I have no explanation and treat the total as uncertain.

Per day the two engines look alike, and both slowed over the summer:

MonthGoogle AI impressions per dayBing citations per day
June1.271.17
July1.101.13
August (22 days)0.680.27
Grouped bar chart of AI appearances per day on dhawalshah.net for June, July and August 2026: Bing Copilot citations 1.17, 1.13 and 0.27; Google AI impressions 1.27, 1.10 and 0.68
AI appearances per day, both engines, June to 22 August 2026. Bing Webmaster Tools; Google Search Console.

Most days nothing happens. Google showed nothing on 42 of 92 days, Bing on 66 of 91. Bing arrives in bursts: 10 citations of one page in a day, then silence.

One small site cannot prove a rule and I am not claiming one. The narrower point holds: ranking well in ordinary search did not earn AI citations here, or predict them. The next section is my best explanation.

These are small numbers. I publish them because the zeroes carry as much information as the counts, and nobody else in this category shows theirs.

Citation capsule

Over 92 days to 22 August 2026, dhawalshah.net recorded 89 Google AI impressions and 122 Copilot citations (Bing page report), against 5,932 ordinary Google search impressions in the same window. The top search page (1,479 impressions) earned zero Bing citations; a page with 78 earned 63, 52% of the total. (Dhawal Shah, 2026)

How Much AI Brand Visibility Comes From Your Own Website?

Very little of a brand’s AI visibility comes from its own website, somewhere between 3% and 10%. The rest is written by other people on sites you do not control.

Muck Rack analysed more than 25 million links cited by ChatGPT, Claude and Gemini in 2026. 84% pointed to earned media, meaning news coverage and other writing the brand does not own. Paid and advertorial placements were 0.3% (Muck Rack, 2026).

How much points at the brand’s own site depends on who counts. An analysis of 149,912 citations put it at 2.9% (Ranqo, 2026). The June 2026 Cited Index, 226 Indian brands across five platforms on 265 non-branded prompts (questions naming no company), put it at 4.1%; its August edition put the median brand below 1% (Cited, 2026). A study of 210 businesses across ChatGPT, Gemini and Perplexity found a median of 3.0% per business and 7.7% pooled (BizLoc8, 2026). Call it 3% to 10%, and do not quote any one figure as settled.

Reddit is the most-cited domain, around 40% of citation frequency across engines. Wikipedia is second, in 26 to 48% of ChatGPT’s top ten citations. YouTube, LinkedIn and Forbes follow (Everything-PR, 2026). Review platforms count too: domains listed on G2, Capterra, Trustpilot, Sitejabber or Yelp averaged 4.6 to 6.3 ChatGPT citations against 1.8 for domains not listed, roughly three times as many (SE Ranking, 2025).

My own data matches this in miniature: one article carries half my Bing citations, and most of my pages never appear on either engine. It also explains the inversion above. If only a few per cent of citations touch a brand’s own domain, ranking well there has little pull over whether an engine cites you. Fixing my own pages is still worth doing, but it is the smaller half of the job.

Citation capsule

Of more than 25 million links cited by ChatGPT, Claude and Gemini, 84% pointed to earned media rather than brand-owned content; paid placements were 0.3%. Separate studies put a brand’s own website at roughly 3% to 10% of its citations; Reddit is the most-cited domain at around 40%. (Muck Rack, 2026)

How Do You Get Cited by ChatGPT?

Adding statistics, named sources and direct quotations to a page is the one change with measured evidence behind it. The original generative engine research, presented at KDD 2024, tested which content changes made a page appear more often in AI answers. The largest gains came from adding statistics, citing sources and including quotations, a lift of around 40% (Aggarwal et al., 2024).

The “citation capsule” boxes on this page exist because of that finding. Each is a short, self-contained, attributed passage an engine can lift whole without the paragraphs around it. It is cheap, and as far as anyone has measured it is the highest-return formatting change available. It works as well in a guest post or forum answer as on your own page. Given where most citations come from, that matters.

The less comfortable finding arrived in August 2026. Press Ranger and OtterlyAI compared news pages by whether the publisher had a licensing deal with OpenAI. Pages from publishers with a deal averaged 10.2 ChatGPT citations each; those without averaged 6.9, a 48% gap. Publishers signed exclusively with OpenAI earned 112% more (Press Ranger and OtterlyAI, 2026).

The study does not establish cause, and its authors say so. OpenAI signed prominent English-language publishers who may already have been more citable. The effect was also specific to ChatGPT, with no equivalent advantage for Google or Perplexity licensees on their own platforms.

Even with those caveats, the largest measured driver of ChatGPT citations in public data is a commercial contract. No audit detects it and no content work replicates it. A small business in Singapore will not be offered one, and I would rather you knew that before reading a vendor’s promise about their optimisation.

Citation capsule

Research at KDD 2024 measured roughly 40% higher visibility in generative engine answers from adding statistics, citing sources and including quotations. In August 2026, pages from publishers with OpenAI licensing deals averaged 10.2 ChatGPT citations against 6.9 for unlicensed publishers, a 48% gap; exclusive licensees earned 112% more. The licensing effect is correlational and ChatGPT-specific. (Aggarwal et al., 2024)

What AI Visibility Tracking and Monitoring Is Worth Running?

The AI visibility tracking worth running is a loop you repeat rather than a report you file. A report tells you where you stand today. A loop tells you whether anything you did afterwards made a difference, the question a budget holder actually cares about.

Measure what the engines actually did with your pages, decide the next action from the gaps, do the work, then measure again to see whether it moved.

In practice, watch:

  • Whether you appear when someone asks about your category, per engine, never blended.
  • How many of your pages carry short, attributed passages an engine can lift whole.
  • Where your category gets discussed outside your own site, the larger half of the job by the earned-media numbers.
  • Movement over months, since a single week is inside the noise.
Four-step loop diagram: measure what the engines did, decide the next action, do the work, measure again, with an arrow returning to the first step
The loop: measure, decide, do the work, measure again. Monthly, on fixed prompts, one engine at a time.

For AI visibility monitoring, cadence beats precision. A monthly read on a fixed set of prompts beats a weekly read on a shifting one, because the noise swallows a week. It is the argument I made about marketing analytics generally: fewer numbers, read more carefully. I run the whole loop from the terminal, alongside the rest of the channels, which I have written up in running marketing from one surface.

We run AI visibility and AI SEO for agency clients on this loop rather than a vendor dashboard; it is how the SEO practice at 2Stallions is set up. The loop carries from one client to the next, where a dashboard tuned to one client’s prompts usually does not.

Two numbers our own system deliberately leaves out say more about the category than anything it includes. We did not ship a peer benchmark score, because with our data it would be too thin to stand behind in front of a client. We also refuse to state a financial return from citation movement, because nobody can currently prove that link. When a tool hands you both without a caveat, ask how it got them.

This loop is packaged as a check you can run yourself: one skill file plus a citation readiness scorecard, with "unseen" treated as a real state rather than a blank. Get the AI Visibility Check.

Frequently Asked Questions

What is an AI visibility audit?

An AI visibility audit checks whether AI assistants mention your brand, cite your pages, or leave you out when answering questions about your category. A useful one separates being named from being cited, and reports each engine on its own rather than blending them into a single score.

How can I check my AI visibility?

Start with the two free first-party reports. Bing Webmaster Tools has published its AI Performance report since 10 February 2026, showing cited pages, the search queries behind them and citation share. Google Search Console’s generative AI report, worldwide since 31 August 2026, shows impressions only. Both are free.

What is a good AI visibility score?

There is no comparable answer, because no two tools measure the same thing. Each vendor uses its own set of prompts and its own mix of models, so the numbers describe different measurements rather than different results. Track your own citations per engine over time instead.

Is there a free AI visibility tool?

Yes, two, and both come from search engines rather than vendors. Bing Webmaster Tools and Google Search Console both publish first-party AI reporting at no cost, with no paid tier. Verify your site in both and read three months of data before buying a commercial seat.

How do you get cited by ChatGPT?

Research presented at KDD 2024 found roughly 40% higher visibility in LLM answers from adding statistics, citing sources and including quotations. Write short, self-contained, attributed passages an engine can quote whole. Publisher licensing deals correlate with 48% more ChatGPT citations, and no amount of content work replicates that.

Can you track ChatGPT citations the way you track Copilot citations?

No. Microsoft publishes Copilot citation data directly in Bing Webmaster Tools. OpenAI publishes no equivalent report, so ChatGPT visibility can only be estimated by running sample prompts and recording what comes back. Treat any ChatGPT number as an estimate, not a measurement.

What To Do With This

Start by opening the two free reports before you spend a dollar on anything else. Verify your site in Bing Webmaster Tools and Google Search Console, find the AI reports, and read them for three months. They are free and first-party, and most teams I meet have never opened them. If a paid tool still looks necessary after that, you will at least know which question you are paying to answer.

Then track each engine on its own. A blended score would have hidden that my most-cited page on Bing barely registers on Google, and that Google’s favourite page has no Bing citations at all. Two columns in a spreadsheet, one per engine, updated monthly, is enough to start. Add ChatGPT, Perplexity or Claude as sampled estimates only, clearly labelled, and only once the free data has stopped teaching you anything new.

Spend at least as much effort on where you are discussed as on your own pages. The studies above put a brand’s own site at a few per cent of its citations, and on mine a strong search ranking did not predict a single citation. Where your category gets written about and argued over is the bigger lever. It is also the one most marketing teams are least set up to pull, because it means working on pages they do not own.

Judge changes over months rather than weeks, because a week is inside the noise. Expect zeroes, and record them. A page that never appears is telling you something that a dashboard hiding its blanks will not. The audit is only the first lap. The useful part is the second and third, once you can see what actually moved.


More on marketing and AI strategy: Marketing & Growth | Related: Building an AI Tool Without Code | Claude Code for Marketing | From SEO to AI Engine Optimisation

Subscribe
Subscribe to the newsletter

Get Practical Insights Every 2 Weeks.

No spam. Ever.