NEWSFunding news: €500K for AI search visibility
Back
AEO Strategy

GEO KPIs: What to Measure and How to Report Them to Clients

Simos ChristodoulouSimos Christodoulou
·Jul 21, 2026·23 min

A GEO KPI is a measurement of whether AI chatbots name, cite, or recommend a brand inside their generated answers, replacing the rank-and-click metrics that traditional search reporting depends on. The definition and every threshold in this guide come from the AI Visibility KPI framework (v2.0, March 2026) published by Visiblie, the AI visibility company that gets brands recommended in AI answers.

Four GEO KPIs can be measured and defended: visibility score movement, prompts won, citation increase, and AI referral sessions. A fifth number that clients ask for constantly, how many people saw the brand inside an AI answer, cannot be measured by anyone, because no provider publishes it.

That gap is where most GEO (Generative Engine Optimization) reporting loses credibility. The sections below define each measurable KPI, set the thresholds that separate signal from noise, and give you the exact sentences to defend each number when a client pushes back.

Get Your Free AI Visibility Report: See how your brand appears across ChatGPT, Gemini, and Perplexity, in 60 seconds. Run the free report

What are GEO KPIs and why do they replace ranking reports?

GEO KPIs replace ranking reports because the unit of measurement changed from position to presence. A ranking report answers where a page sits in a list of ten blue links. A GEO KPI answers a different question: when a buyer describes their problem to an AI chatbot, does the brand appear in the answer at all?

That shift breaks the reporting habit built over two decades of search marketing. Position 4 means nothing inside a generated paragraph that lists three vendors and links to two of them. Evidence in generative engine optimization takes a binary form instead: a brand was named, or it was not.

Four KPIs carry that evidence and stand up to client scrutiny.

GEO KPIWhat it answersWhere it is read from
Visibility score movementIs overall presence rising across tracked answers?AI visibility platform, run on a schedule
Prompts wonWhich specific buyer questions now return the brand?AI visibility platform, prompt-level detail
Citation increaseWhich pages earn a source link inside answers?AI visibility platform, citation records
AI referral sessionsIs AI presence sending real traffic?Google Analytics, referral segment

The rest of this guide works through each one, then covers the harder half of the job: judging whether a number moved for real, admitting what cannot be measured, and building a monthly report a client trusts.

Why does GEO reporting break the Search Console habit?

GEO reporting breaks the Search Console habit because no AI provider publishes the data Search Console provides. Google Search Console reports impressions, average position, clicks, and click-through rate for every query that surfaced a page. ChatGPT (OpenAI), Google Gemini, Perplexity, and Claude (Anthropic) report none of those figures to brands.

The practical consequence reshapes what an agency actually reports. A GEO report rests on a sample of prompts, run on a schedule against tracked models, rather than a census of what real users asked. The tracked prompt set is a research instrument, not a traffic log.

Set that expectation in the first meeting. A client who discovers the sampling basis three months into a retainer stops trusting the entire report, including the parts that were solid.

SEO metricGEO equivalentAvailable to brands?
ImpressionsNoneNo. No provider publishes answer-impression counts.
Average positionProminence within the answerPartly. Order of mention is observable, not supplied as a metric.
ClicksAI referral sessionsYes, with a known undercount.
Click-through rateNoneNo. Without impressions, no rate can be calculated.
Query coveragePrompt coverage across a tracked setYes, limited to prompts you chose to track.
Ranking positionNamed or not namedYes, binary per prompt per run.

That table is the honest starting point for any conversation about AI visibility vs SEO. Two rows read "No". Naming them early converts a weakness into a credibility signal.

Which four GEO KPIs can you actually measure?

Four GEO KPIs can be measured and defended in a client report. Each one below follows the same four-part structure: what it is, how it is measured, what it proves to a client, and how to show it.

Visibility score movement

A visibility score is a composite metric combining brand mention rate, citation rate, sentiment, and accuracy into one figure that tracks overall presence across AI answers.

Measurement runs a fixed prompt set against tracked models on a schedule, then scores each answer for whether the brand appeared, whether the mention carried a source link, how the brand was characterised, and whether the description was factually correct. Brand mention rate underneath it uses a simple formula: (queries with mention / total queries) x 100.

For a client, the composite proves direction. The composite answers "is this working overall" in a single number that fits on a first slide. The same composite performs badly at diagnosis, because a score that rises on improved sentiment tells you nothing about whether more buyers now encounter the brand.

Show the score as a trend line across at least three runs, never as a standalone figure. One score without a prior period is a data point, not evidence. Pair it with the AI visibility metrics that compose it whenever a client asks why it moved.

Prominence deserves a manual log until platforms score it. When you read each answer, record whether AI named the brand first, mid-list, or last, and whether the sentence recommended the brand or merely listed it. Three runs of that log reveal a pattern the composite hides: a brand can hold a flat visibility score while climbing from last-listed to first-named, which is real progress the headline number misses entirely.

When the client asks how the brand compares rather than whether it improved, report share of voice next to the composite. Share of voice divides brand mentions by the total brand mentions appearing in the same answers: (brand mentions / total brand mentions) x 100. That conversion turns an absolute score into a competitive position, and it answers the question a founder actually cares about, which is whether the gap to the category leader narrowed.

Prompts won

Prompts won counts the distinct tracked prompts where AI names the brand, expressed as a fraction of the tracked set. Prompt coverage formalises it: (prompts with appearance / total tracked prompts) x 100.

Measurement produces a per-prompt record rather than an aggregate. Each prompt carries the answer text that produced the verdict, which makes this the most inspectable KPI in the set. A client who doubts the number reads the actual answer.

Split the count into branded and unbranded prompts before reporting it. A prompt won on the company's own name proves that AI recalls the brand correctly. A prompt won on an unbranded category question, where the buyer described a problem and never typed the brand name, proves demand capture. Merging them inflates the result, because branded prompts are far easier to win and any brand with a website wins most of them.

A worked example shows the size of the distortion. Take a tracked set of 40 prompts, 15 branded and 25 unbranded. The brand wins 13 of the 15 branded prompts and 3 of the 25 unbranded ones.

Merged, that reads as 16 of 40, or 40% prompt coverage, a number that sounds like category strength. Split, the same result reads as 87% branded and 12% unbranded, which describes a brand AI recognises well and recommends rarely. The two readings lead to opposite strategies, and only the split version survives a client asking which questions the brand actually wins.

Report the split as two numbers. "16 of 40 prompts won, 13 branded and 3 unbranded" is defensible. "40% prompt coverage" hides the only half that represents new demand.

Citation increase

Citation rate measures the share of brand mentions that carry a source link back to the domain: (mentions with citation / total mentions) x 100.

Measurement captures both the rate and the specific pages earning links. Being mentioned and being cited are different outcomes with different causes. A brand mentioned from model memory produces no link and no traffic. A brand cited from a retrieved page produces both, and identifies exactly which page did the work.

For a client, citations prove the content investment is landing. Mentions prove awareness; citations prove the published pages are the source of that awareness. The page-level detail converts the KPI into a content decision, because the pages earning citations tell you which format and depth the models retrieve from your domain.

Report per engine rather than blended. Perplexity cites heavily by design because it retrieves in real time. Other models cite sparingly. A blended citation rate mixes those behaviours into a figure that describes no engine accurately and moves whenever the model mix shifts.

AI referral sessions

AI referral sessions count visits arriving on the site from an AI chatbot, recorded in Google Analytics as referral traffic from domains such as chatgpt.com and perplexity.ai.

Measurement happens in the client's own analytics rather than a vendor platform, which makes this KPI structurally different from the other three. The client verifies it without trusting anyone.

That independence is the reason to lead with it when credibility is the problem. AI referral sessions answer "show me it is working" in a tool the client already logs into.

State the undercount every time. Sessions that arrive without a referrer land in direct traffic, and a buyer who reads an AI answer on Monday and types the domain on Thursday never appears as an AI session at all. Report AI referral sessions as the measurable floor of AI-driven traffic, never as the total.

See How Visiblie Automates This: Track prompts, citations, and referral sessions across 8+ AI models in one place. Explore the platform

How much movement counts as real movement?

Movement counts as real when it clears a defined threshold band, and each band applies only once its phase is unlocked. Threshold bands are where most GEO reporting fails: a number moves, the report calls it progress, and the client asks how much movement would have counted as failure. Without published thresholds, that question has no answer.

The framework scores every KPI against the phase it belongs to in the AI visibility maturity model.

Each phase carries two KPIs: a verification KPI that asks whether AI knows something about the brand when asked directly, and an organic KPI that asks whether AI surfaces it unprompted. Verification rates run high; organic rates run low. Judging both against one bar produces false failure.

PhaseVerification KPIGreenYellowRedOrganic KPIGreenYellowRed
Category formationCategory Placement Accuracy70%+40-69%Under 40%Organic Category Inclusion15%+5-14%Under 5%
Attribute recallAttribute Verification Rate60%+30-59%Under 30%Organic Attribute Association25%+10-24%Under 10%
Proof and trustDirect Trust Confirmation50%+25-49%Under 25%Organic Trust Mention10%+3-9%Under 3%
Competitive selectionComparison Inclusion Rate20%+8-19%Under 8%Combined named and unnamed
AmplificationMulti-Model Consistency70%+40-69%Under 40%Requires 2+ tracked models

Three rules make the table usable in a client meeting.

  1. A red number in a locked phase is expected, not failure. A brand working through category formation will score red on competitive selection. Reporting that red as underperformance misleads the client and invites a budget conversation the work does not deserve. Mark it "not evaluated yet, prerequisite phase not unlocked".

  2. Movement inside plus or minus 5 percentage points is noise. Between two runs, a change beyond that band counts as improving or declining. Anything inside it stays flat, regardless of how the number looks on a chart.

  3. A completed phase does not regress on the report. When a previously green phase drops, flag it as completed with a warning and fix it, rather than moving the brand backward through the model.

Walk a real reading through those rules. A brand posts 62% Category Placement Accuracy and 7% Organic Category Inclusion, against 58% and 6% last month. Both KPIs sit yellow, and both moved less than 5 points, so the correct report says the brand held position in category formation and cleared nothing.

A weaker report would headline "Category accuracy up 4 points" and invite the client to expect the same gain every month. The stricter reading costs you a good-news line and buys a threshold the client can hold you to next quarter.

Read every band against the brand's current stage in the maturity model, never in isolation. A 12% organic category inclusion rate is yellow progress for a brand in category formation and a warning sign for a brand that reached competitive selection two quarters ago. The same figure carries opposite meanings depending on where the brand sits, which is the reason the maturity model precedes the threshold table rather than decorating it.

These thresholds are provisional. The bands will be recalibrated once 20 or more brands complete two or more classified runs. Publishing a provisional number with its caveat attached beats publishing a confident number with no derivation, and clients who work with data respect the distinction.

Visiblie team

Want to see how AI talks about your brand?

Join 500+ companies tracking their AI visibility. Get started in 2 minutes.

Start Free Trial

What can you not measure, and what do you tell the client?

No, you cannot report how many people saw a brand inside an AI answer. No provider publishes impression, reach, or query-volume data to brands. Any metric presented as "AI impressions" is modelled from traditional search volume, and a client deserves to be told that before the invoice arrives.

Three numbers stay out of reach.

Run-to-run variance is the second limit, and it changes what counts as a measurement. The same prompt, sent to the same model twice, does not return the same answers. A single run is a sample of one. Treat one run as evidence and the report will contradict itself the following month.

Visiblie handles this with repeat runs inside each evaluation window, counting a prompt as won only when the brand appears in both runs. The two-run rule discards genuine wins occasionally, and never manufactures one, which is the correct trade when a client is deciding whether to renew.

The third limit catches brands that share a name with another company. AI describing a same-named business in a different country or category still produces a brand-name match, and a naive count records it as a win. Before reporting any result, check the citations behind each mention and confirm the answer describes the right company. A mention rate inflated by mistaken identity collapses the moment a client reads one of those answers.

How do you build the monthly GEO report?

Build the monthly report in four blocks, in this order: the numbers, the movement, the cause, and the next action.

  1. The four KPIs, each against the prior period. Visibility score, prompts won split branded and unbranded, citation rate per engine, and AI referral sessions. Every figure carries its run date, because AI answers cannot be verified retrospectively. An answer from three weeks ago no longer exists to be checked.

  2. What cleared the threshold. Apply the plus or minus 5 point band and state plainly which KPIs moved and which stayed flat. Resist narrating flat numbers as progress.

  3. Why it moved. Tie each change to work that shipped: pages published, schema deployed, review volume gained, mentions earned on third-party sources. A number without a cause reads as luck.

  4. What happens next period. Name the prompts targeted next and the phase gate the work is aimed at.

Lead the first slide with prompts won rather than the composite score. Prompts won is countable, inspectable, and survives the question "show me". The composite score requires the client to trust a calculation they cannot see.

Run this monthly for the operational report and quarterly for the business review, where the four KPIs connect to pipeline and revenue. Agencies packaging this as a service line will find the delivery model covered in AI visibility for agencies, and the competitive framing in AI share of voice.

Compare Plans: Find the right plan for your team size and monitoring needs. See pricing

How do you defend each number to a sceptical client?

Defend each number by answering the objection directly and handing over the underlying data. Five objections cover most client meetings.

ObjectionWhat to say
"Our traffic did not change.""These KPIs measure presence, not sessions. AI referral traffic is the floor of what AI sends, because sessions without a referrer land in direct. Presence moves first, traffic follows."
"You only tested a handful of questions.""We track a fixed set of prompts built from your buyer's actual questions. Here is the full list. Add any question you think is missing and we will track it from next run."
"I asked ChatGPT and got a different answer.""Expected. The same prompt returns different answers across runs, which is why we run each prompt twice per window and count a win only when it holds in both."
"How do I know you did not pick easy questions?""Branded prompts are the easy ones and we report them separately. The unbranded number is the one that matters, and it is the smaller of the two."
"This number went down.""Movement inside 5 points is noise. This moved 3, so it is flat. Here is what did clear the band."

Each response works because it hands the client something inspectable: the prompt list, the answer text, the split, the threshold. Every claim in a GEO report must trace back to an artefact the client can open.

Two supporting habits reduce objections before they arrive. Share the tracked prompt list at the start of an engagement rather than on request, which removes the cherry-picking suspicion permanently. Record the answer text alongside each verdict, so any disputed prompt resolves in seconds.

Platform-specific detail on capturing that evidence appears in the guide to track brand mentions in ChatGPT, the qualitative layer in AI brand sentiment, and the mechanics of why models diverge in our explainer on whether ChatGPT gives the same answers.

Frequently asked questions about GEO KPIs

What are the most important GEO KPIs?

Four GEO KPIs carry defensible evidence: visibility score movement, prompts won, citation increase, and AI referral sessions. Prompts won matters most in client reporting, because the client inspects the underlying answer and verifies the result without trusting the platform that produced it.

How often do you report GEO KPIs?

Report GEO KPIs monthly for operational reviews and quarterly for business reviews. Monthly cadence surfaces directional movement without overreacting to run-to-run variance. Quarterly reviews connect the four KPIs to pipeline, revenue, and competitive position.

Can you track GEO in Google Analytics?

Partly. Google Analytics records AI referral sessions when a user clicks through from an AI chatbot, and that figure is the one KPI a client verifies independently. Analytics captures nothing about mentions, citations, or prompts, because those events happen inside the answer and never reach the site.

What is a good citation rate?

No industry benchmark exists yet. Citation rate varies by engine, because Perplexity cites heavily by design while other models cite sparingly. Judge citation rate against the brand's own prior runs and against the phase thresholds in the maturity model, rather than against a published figure with no stated derivation.

Is GEO the same as SEO?

No. SEO optimises for position in a ranked list of links. GEO optimises for presence inside a generated answer. The disciplines share technical foundations, including crawlability and structured data, and diverge completely at the measurement layer, as the metric mapping earlier in this guide sets out.

Reporting AI visibility credibly starts with knowing which four numbers hold up and saying plainly which one does not exist. Brands that publish thresholds, split branded from unbranded, and hand over the prompt list build client trust that survives a flat month.

For a baseline reading of where a brand stands across AI visibility today, and the composite AI visibility metrics behind it, start with the free report.

Get Your Free AI Visibility Report: See how your brand appears across ChatGPT, Gemini, and Perplexity, in 60 seconds. Run the free report

kpisai visibilitymetricsgeoaeo
Simos Christodoulou

Simos Christodoulou

Head of SEO & GEO

Expert in search engine optimization, generative engine optimization, and AI visibility strategies. Experienced in technical SEO, structured data implementation, semantic SEO, and optimizing brand presence across AI platforms.