New Run the 150-Point Growth Audit on your funnel
Back to Blog

How to Tell Whether AI Engines Are Citing You

Marketing Akif Kartalci 16 min read
ai citation monitoringtrack AI brand mentionsAEO monitoringAI brand visibilityLLM brand monitoringshare of voice AI searchanswer engine optimization
How to Tell Whether AI Engines Are Citing You

Here is the number that should bother you: 70.6% of AI referrals land in your analytics as direct traffic. No source, no referrer, no trace. Your GA4 dashboard shows a bump in “direct” conversions and you have no idea that ChatGPT sent those visitors. You keep optimizing for organic rankings while the buyers who actually converted came through an AI citation you never knew existed.

AI citation monitoring is the gap between what your analytics show and what actually drives demand. Verifying whether AI engines are citing your brand, which engines, how often, and what they say, is not yet on most growth teams’ radars. That is a mistake. We began auditing our own citation status at Momentum Nexus in early 2025, and what we found was not reassuring. The gap between our perceived AI visibility and our actual citation rate was large enough to restructure our content priorities.

This post walks through the exact process: a manual audit you can run in two hours, the metrics that actually matter, how to interpret your results, and when to graduate to automated monitoring. It pairs with our practitioner’s guide to answer engine optimization, which covers how to earn citations once you know what your baseline looks like.

The Dark Traffic Problem: Why Your Analytics Lie About AI

Before running any audit, it helps to understand why you need one at all. AI-driven demand is largely invisible to standard analytics tools, and the scale of that invisibility is larger than most people realize.

A 2026 analysis of 446,405 sessions by Loamly found that 70.6% of AI-referred traffic arrives without referrer headers and gets misclassified as “direct” in GA4. Google added a native AI Assistant channel to GA4 in May 2026, but that only captures sessions where the referrer header was preserved, which excludes most copy-paste behavior and many AI assistant interactions. The practical result: you cannot judge your AI visibility from your traffic reports.

The gap becomes more frustrating when you look at what AI-influenced sessions actually do. Seer Interactive found that ChatGPT referral traffic converts at 15.9% versus 1.76% for Google organic. An 11-to-1 conversion advantage, hidden inside your “direct” bucket. For one enterprise brand in their dataset, AI referrals accounted for only 0.5% of traffic but drove 12.1% of signups. If you are measuring AEO effectiveness by referral traffic volume, you are looking at a number that undercounts the channel by roughly 70%.

There is also a category of AI interaction that generates no click at all. These are answers that name your brand or describe your product without the user ever visiting your site. Research from Superlines illustrates this clearly. During a 30-day self-audit in early 2026, they found Gemini cited their domain 182 times but mentioned the brand name “Superlines” zero times. Perplexity generated 20 times more website links than actual brand name mentions. Their overall AI presence was 73% ghost citations: the domain appeared as a source but the brand identity never surfaced in the answer text.

The other direction of the problem: 93% of AI-driven sessions end without clicks to any website at all. The buyer read the answer, formed an impression, and moved on. This means your AI visibility builds (or fails to build) brand awareness entirely inside the answer, before any site visit happens. Citation monitoring is how you find out what that answer says.

We covered the conversion measurement challenge in our content marketing ROI measurement framework, but for AI visibility specifically, measurement has to start before the traffic. It starts with whether you appear in the answer.

What AI Citation Monitoring Actually Measures

The four metrics worth tracking work as a cascade, not a single number.

MetricWhat it measuresWhy it matters
MentionsHow often your brand name appears in AI answers, with or without a linkTop of funnel: are you in the conversation at all
CitationsHow often those mentions include a clickable source link to your domainAre you driving navigable traffic from AI answers
Share of voiceYour mention rate versus competitors across a fixed prompt setCompetitive position in the AI layer
AI referral trafficVisitors arriving from answer engines, filtered by known AI domainsRevenue-facing outcome, severely undercounted

The cascade matters because these metrics move on different timelines. Share of voice and mentions are leading indicators: they shift within weeks of a structural content change. Referral traffic is lagging: it trails the leading metrics by a full reporting cycle or two. The mistake I see most often is teams watching only the lagging number, finding it too small to care about, and killing an AEO program that was actually gaining share of voice. The first 90 days of any monitoring program should be judged on share-of-voice movement, not referral traffic.

Platform differences matter here too. ChatGPT includes citations in only about 31% of responses. Perplexity cites sources in 77% or more of responses with numbered inline links. Claude mentions brands in 97.3% of responses but often without links, requiring web search activation for source attribution. Google AI Overviews appear in roughly 30% of all searches but 75% of problem-solving queries. These numbers mean that a brand with zero ChatGPT citations may still have significant Perplexity presence, and vice versa. Checking one engine and concluding you have no AI visibility is one of the most common audit errors.

How to Run a Manual Citation Audit

This is the section most guides skip. The monitoring tool comparisons get written, the strategy frameworks get published, but the manual “here is how you check right now, for free, this afternoon” guide is rare. Here is ours.

The full audit takes about two hours the first time. After that, the weekly maintenance cadence runs in 45 to 60 minutes.

Step 1: Build Your Prompt Set

Before touching any AI engine, define the questions your buyers actually ask. Your prompt set determines everything about whether the data is useful, and getting it wrong at this step ruins every subsequent run.

Three prompt categories produce the right coverage:

Category prompts: “What are the best [your category] tools for [core use case]?” This tests whether your brand appears in a standard “who are the players” response.

Comparison prompts: “[Competitor A] versus [Competitor B] for [job to be done]?” These reveal whether AI engines include you in competitive conversations or route buyers entirely to your competitors without mentioning you.

Use-case prompts: “How do I [core task]?” or “Who are the leading providers of [service]?” These catch queries where buyers describe their problem in their own language rather than naming a category.

Source your prompt language from Google Search Console, not from how your team talks internally. The queries buyers actually type are different from the vocabulary in your marketing materials. “Best project management software for remote startups” is a real search query. “Top work OS platforms for distributed agile teams” is a marketing team prompt.

Twenty-five prompts is enough for a first baseline. Fifty gives you statistical confidence across a full month of tracking. Industry consensus from mid-2026 research puts 30 to 100 prompts as the range for a serious monitoring program.

Two rules that determine whether your audit data is usable:

Lock the exact wording of your prompts. Do not rephrase them between runs. AI responses are nondeterministic, so changing prompt wording makes it impossible to distinguish signal from variation.

Build the complete prompt set before running anything. Do not start running prompts and add more as you think of them. The set needs to be stable from day one.

Step 2: Run Each Engine

Each engine requires different setup, and the differences matter for what you can see.

EngineSetup RequiredCitation BehaviorWhat to Note
ChatGPTEnable “Browse the Web” in settingsCites in ~31% of responsesSource domain, brand name in text, position in answer
PerplexityNo setup neededCites in 77%+ of responses, numbered inlineFull source URL list visible immediately
ClaudeActivate web search in settingsMentions brands in 97.3% of responsesWhether brand name appears, links attached or not
Google AI OverviewsUse incognito browserAppears in ~30% of queriesScreenshot overview box, note source links at bottom

Run each prompt in all four engines. Use an incognito window for Google to strip personalization. For ChatGPT, “Browse the Web” is required to get source citations: without it, you are testing training corpus knowledge, which updates on a 6 to 12 month cycle rather than 24 to 72 hours. Training corpus results tell you about your historical reputation, not your current citation status.

Practical note on timing: results rotate 40 to 60% month over month across most engines. A single run is a data point, not a conclusion. Month one establishes a baseline to compare against month two. Do not draw strategic conclusions from a single audit pass.

Step 3: Document Your Findings

Log results in a Google Sheet before moving on. Memory and screenshots are not a system. The specific columns matter.

ColumnWhat to capture
PromptExact wording, locked
PlatformChatGPT / Perplexity / Claude / Google AIO
DateRun date
Cited?Y / N — your domain appeared as a source link
Mentioned?Y / N — your brand name appeared in the answer text
PlacementHeadline / Body / Footnote / None
SentimentPositive / Neutral / Negative / Mixed
Competitors citedEvery competitor that appeared in the same answer
Source URLExact URL cited, if available
NotesGhost citation? Factual error? Surprising framing?

The distinction between “cited” and “mentioned” is worth taking seriously. A ghost citation, where your domain appears as a source link but your brand name never appears in the answer text, gives you a technical attribution your audience will never see. The buyer reads the synthesized answer, which does not name you, and moves on. Ghost citations validate that your content structure is extractable. They do not mean buyers know you exist.

Log both metrics separately. Your ghost citation rate is a distinct signal from your brand mention rate.

Step 4: Calculate Your Baseline

With your full prompt set documented, the math is simple.

Citation rate per engine: Citation rate = (prompts where your domain appeared as a source / total prompts run) × 100

Share of voice: Share of voice = your brand mentions / (your mentions + all competitor mentions across same answers) × 100

Run both calculations separately per engine. Blended averages hide important platform-level differences that require different responses.

For B2B SaaS companies, here are citation rate benchmarks by company stage from Data-Mania’s 2026 AI Search Visibility analysis:

Company StageTypical Citation RateShare of Voice Signal
Pre-seed / seed0 to 8%Below 5% = no AI presence yet
Series A8 to 20%10 to 15% = building position
Series B+20 to 35%20%+ = competitive presence
Category leader35 to 50%30%+ = strong category authority

In categories with 5 to 8 competitors, share of voice below 10% is a warning sign. The same research found that 89% of ChatGPT citations concentrate among the top three brands in a category. The long tail of AI citations is thin. If you are not in the top three, you are competing for what is left after the category leaders take the lion’s share.

What Your Results Actually Tell You

The audit produces different conclusions depending on what you find. Here is how to interpret each pattern.

Zero citations across all engines. This is more common than most brands expect. Analysis across healthcare, SaaS, and financial services found that 90% of brands have zero AI search mentions. Zero citations do not mean your content is bad. They typically mean your entity signals are too weak for engines to confidently cite you, or your content structure makes extraction difficult. Zero citations in Perplexity, which cites sources in 77% of responses, is a structural content problem. Zero citations in ChatGPT, which cites in only 31%, might just mean low priority topics. Diagnose by engine before drawing a blanket conclusion.

High ghost citation rate. If your domain appears frequently but your brand name does not, AI engines are using your content but not attributing it to you. This pattern usually means your content structure is strong enough for extraction but your entity establishment is weak: not enough third-party sites describing who you are and what you do consistently. Fix the entity layer before worrying about citation volume. We cover what drives training-corpus recognition in our GEO and AEO strategy guide.

Negative sentiment citations. This is more urgent than zero citations. An AI engine citing you as a cautionary example, or citing outdated product information that paints you unfavorably, reaches buyers before your own content does. Peec AI’s sentiment framing feature specifically tracks this pattern because a negative mention at scale can actively harm pipeline. If you find negative framing in your audit, prioritize fixing the source content and the third-party mentions that inform that framing.

Strong competitor presence, low own presence. The most common audit outcome. Across AEO programs we have run, the typical pattern is: one or two competitors appear in 6 of 10 relevant answers, you appear in 1 or 2. This is not random. Competitor citation authority has a specific cause, usually a combination of structured content, third-party mention volume, and established entity signals. It responds to targeted fixes, but it takes time to move.

Factual errors about your product. AI engines sometimes make incorrect claims about pricing, features, or positioning. If your audit surfaces these, they are the highest priority issue on your list. A buyer who gets wrong information from an AI engine may never reach your site to get the correct information. Correcting this requires updating your own source content and earning corrections on the third-party sources the engine is drawing from.

The Ongoing Monitoring Cadence

The audit above gives you a baseline. The cadence below keeps it current. Citation patterns rotate significantly month to month, so a one-time snapshot tells you where you were, not where you are.

The sustainable cadence for a growth team without a dedicated AEO tool:

Weekly (45 to 60 minutes): Run your top 20 highest-priority prompts across all four engines. Focus on prompts tied directly to buyer purchase intent. Log results in the tracking sheet, note any changes from the prior week.

Monthly (3 to 4 hours): Run the full prompt set. Calculate citation rates and share of voice for the month. Compare against prior months. Flag prompts where competitor position changed significantly or where new competitors appeared.

Quarterly: Strategic review. Add new prompts in a separate cohort to capture new buyer questions or category language shifts. Do not replace existing prompts mid-program. Treat the new cohort as a parallel track so you preserve period-over-period comparability on the original set.

When to add a paid tool: Manual tracking becomes unsustainable around 5 hours per week. At that point, analyst time costs more than basic monitoring software. We reviewed the tool landscape in our guide to answer engine optimization tools. The short version: Rankscale at $17 to 20 per month works for initial proof of concept before you commit to ongoing monitoring. Peec AI at $95 per month handles multi-country monitoring and sentiment framing once you know the channel is generating real results.

Five Mistakes That Invalidate Your Audit Data

Running prompts in your authenticated session. ChatGPT, Claude, and Google personalize results based on your history and profile. If you run audit prompts while logged in on a browser you use regularly, you are measuring personalized output, not what buyers see. Always use an incognito window. For ChatGPT, use a fresh conversation each time.

Checking only one engine. Only 11% of domains appear in both ChatGPT and Perplexity results. A brand can have strong Perplexity presence and zero ChatGPT citations, or vice versa. The engine your buyers actually use determines where your citation gap matters most. If your ICP is technical and researches tools on Perplexity, your ChatGPT citation rate is less important than your Perplexity share of voice. Find out which engines your buyers use before you decide where to focus.

Changing prompt wording between runs. AI responses are nondeterministic. Even small prompt variations produce meaningfully different outputs. If you rephrase “best B2B sales tools for startups” to “top sales software for small companies,” you are comparing different queries, not measuring movement on the same query over time. Lock the wording on day one and never touch it.

Treating ghost citations as brand visibility. A domain citation in a source list is not a brand impression. If the answer text never mentions your company name, the buyer has no idea the information came from you. Ghost citations validate that your content is structurally extractable. They do not mean buyers associate that information with your brand. Log them separately and do not let a high ghost citation rate fool you into thinking your AI brand awareness is stronger than it is.

Judging the program by referral traffic before 90 days. The referral traffic number is an unreliable short-term signal for two reasons. First, 70.6% of it is invisible in standard analytics. Second, what is visible responds slowly to changes in citation rate. I have watched teams kill working AEO programs in month three because “AI traffic is only 40 sessions this week.” Share of voice is the right signal to watch in the first 90 days. Traffic follows, but not on a four-week timeline.

Run the Audit Before You Do Anything Else

Most B2B brands running AEO programs today are optimizing without a baseline. They are publishing content, building entity signals, and restructuring pages without knowing their starting citation rate, their share of voice against competitors, or even which engines are reaching their buyers. That is like running paid ads without conversion tracking before the first campaign.

The two-hour manual audit described above is not a permanent solution. It is the starting point that turns guesswork into a number. Once you have that number, you can make actual decisions: which engines to prioritize, which content to rebuild, which entity signals need work, and whether the effort is moving at all.

Run the audit. Build your prompt set tonight. Find out where you actually stand.

If you want help building an AI visibility program that connects citation monitoring to pipeline, we have done this for B2B SaaS teams at the $50K to $150K Monthly Recurring Revenue stage. Book a free growth audit at Momentum Nexus and we will map your current citation state, identify your highest-leverage fixes, and build a 90-day program around what actually moves your share of voice.

Ready to Scale Your Startup?

Let's discuss how we can help you implement these strategies and achieve your growth goals.

Schedule a Call