How to Benchmark Your AI Citation Share Against a Single Named Competitor
Aggregate share-of-voice hides the fight that matters. Here is how to benchmark your AI citation share head-to-head against one named competitor: build the prompt set, score three levels, compute a fair two-brand share per engine, and read the gap.

Key Highlights
Pick one competitor, build a fixed set of 30 to 50 real buyer prompts, and run it across ChatGPT, Perplexity, Gemini, and Claude from clean sessions. Score three levels: cited, mentioned, recommended. Compute your share as your count divided by the two-brand total, per engine, and repeat weekly so you measure a rate, not one lucky draw.
Most AI visibility dashboards give you a category-wide share-of-voice number: your brand against every competitor, every mention pooled into one percentage. That number is fine for a board slide and useless for a fight. The question a CMO actually loses sleep over is narrower and sharper. When a buyer asks ChatGPT for a tool in your category, do they hear your name or the one competitor who keeps beating you to the shortlist? This is how to measure that specific matchup, engine by engine, so the gap becomes a number you can move.
Why a head-to-head benchmark beats aggregate share of voice
Aggregate share of voice answers "how visible are we." A one-competitor benchmark answers "are we winning the deal the buyer is actually weighing." Those are different questions, and the second one drives strategy.
Consider the arithmetic. If your category has twelve named players and you hold 8 percent aggregate share, that reads as failure. But if the buyer's real shortlist is you and one incumbent, and the engine names you in four answers out of every ten where it names either of you, your head-to-head share is 40 percent. You are not invisible. You are losing a two-horse race by a bit, which is a completely different problem with a completely different fix. Aggregate share buries that signal under ten irrelevant competitors the buyer was never going to choose anyway.
The head-to-head frame also makes the work tractable. You cannot out-content an entire category at once. You can study one competitor's cited pages, find the specific prompts where they win, and target those. The whole method for reading an answer to find why the other brand got named is in why AI assistants recommend your competitor instead of you. The benchmark tells you where to point that diagnosis.
Step one: build a prompt set that reflects the real decision
The benchmark is only as honest as the prompts behind it. A set stuffed with your brand name will flatter you; a set of generic category questions will bury the matchup in noise. You want the prompts a buyer types when they are choosing between you and this specific competitor.
Build 30 to 50 prompts across four intent bands:
- Category discovery. "Best tools for X," "top X platforms for a mid-market team." These are where a buyer meets the shortlist for the first time and neither brand is named yet.
- Named comparison. "Competitor versus your brand," "alternatives to Competitor," "is your brand or Competitor better for Y." These are the highest-intent prompts in the set, and the ones where the matchup is explicit.
- Use-case and persona. "Which is better for a Series B SaaS team," "cheapest option for a solo founder." A brand can win discovery and lose every persona cut, or the reverse.
- Objection and switching. "Reasons not to use Competitor," "migrating from Competitor to your brand." These surface where the engine has absorbed a competitor's weakness you can press on.
Keep the set fixed once you start. Changing prompts between runs means you are measuring your prompt edits, not your citation share. The full method for weighting a prompt set by intent and buyer value is in how to design a prompt set that reflects how buyers actually ask AI about your category.
Step two: score three levels, not one
A brand can appear in an AI answer three ways, and collapsing them into a single "mentioned" tally throws away the most important distinction in the whole exercise. Score each brand, on each prompt, at three levels.
| Level | What it means | Why it matters |
|---|---|---|
| Cited | Your domain appears as a linked source under the answer | Proof the engine retrieved your page, even if it did not name you in the prose |
| Mentioned | Your brand name appears in the answer text | The engine considers you part of the category |
| Recommended | The engine actively suggests you as an answer to the buyer's question | The only level that maps to a shortlist slot |
The gaps between these levels are diagnostic. If you are cited but not mentioned, the engine reads your content and still names the competitor, which is a positioning and entity problem, not a crawlability one. If you are mentioned but never recommended, you make the long list and lose the cut, which points at proof and differentiation. Tracking all three tells you which fix to reach for. The mechanics of that mention-versus-source divide are covered in how to measure your brand's AI citation share across LLMs.
Step three: compute a fair two-brand share, per engine
Here is where the head-to-head frame earns its keep. For each engine, count only answers where the engine named at least one of the two brands. Within that set, your share is your count over the combined count.
Say you run 50 prompts in Perplexity. The engine recommends you in 12 answers, the competitor in 20, both in 6, and neither in 12. Discard the 12 where neither appears, since a two-horse benchmark cannot score a race nobody entered. Of the 38 answers where at least one of you was named, you appear in 18 (12 solo plus 6 shared) and the competitor in 26 (20 solo plus 6 shared). Your head-to-head recommendation share in Perplexity is 18 divided by (18 plus 26), or 41 percent. Do the same for cited and mentioned, and you have a three-by-one scorecard for that engine.
Do not blend engines into one number until you have read them separately. A ChatGPT answer, a Perplexity citation set, and a Gemini response are not equivalent impressions, and averaging them hides the platform where you are actually losing. Reported 2026 benchmarks for the same brand and query set commonly land Perplexity in the high 20s to high 30s, ChatGPT in the low-to-mid teens, Gemini in the teens, and Claude in the low single digits to high single digits, according to practitioner share-of-voice measurement guides. Those spreads mean a blended average can look healthy while your weakest engine, often ChatGPT, is where most of your buyers actually are.
Step four: measure a rate, not a moment
The single most common mistake is running the prompt set once and treating the result as truth. AI answers are non-deterministic. Ask the same question twice and the brand list reshuffles, because the engine samples, retrieval indexes refresh, and model parameters update. Profound's analysis of roughly 680 million citations, cited across share-of-voice methodology writeups, found citation distributions shifting within weeks from those forces alone, before you change a word of your content.
The fix is repetition. Run the full set from clean, logged-out sessions, at least three times per engine per cycle, and record the rate at which each brand appears rather than a single yes or no. Then re-run the whole thing on a fixed cadence, weekly or biweekly, so a real move separates from the noise floor. A jump from 41 to 44 percent in one run is noise. The same brand climbing three cycles running is a trend. Holding the prompt set, the engines, and the session hygiene constant is what turns this from a screenshot into a benchmark.
Reading the gap: four patterns and what each demands
Once you have per-engine, three-level shares tracked over time, the shape of the gap tells you what to build.
- Losing discovery, winning comparison. The engine names the competitor first in "best tools" prompts but leans your way when both are named explicitly. You have a category-entity problem: the engine does not file you as a default member of the category. The work is entity and consensus, not more blog posts.
- Winning discovery, losing use-case cuts. You make the general shortlist but lose "best for a Series B team" or "cheapest for a solo founder." You are under-serving persona-specific proof. Build the pages that answer those exact cuts.
- Cited everywhere, recommended nowhere. The engine retrieves your pages and still names the competitor. Your content is crawlable but not persuasive as an answer, or your differentiation is not legible to a model. This is where structure and proof matter more than volume.
- Strong in three engines, absent in one. Usually a single-engine access or index problem. Verify the engine can actually fetch and has ingested your pages before you assume a content gap.
Only the last pattern is a plumbing fix. The other three are strategy, which is why the benchmark is worth running before you spend a dollar on content.
Keep the inputs the engines read clean and current
A benchmark measures the result of what engines find when they read the category. If your own facts are stale or scattered, you will lose the matchup for reasons that have nothing to do with the competitor's content. Two hygiene moves keep the playing field fair. Give crawlers a tidy, canonical source of truth with a free llms.txt generator, and keep that source current as your product and pricing change through the AI Feed Engine, so the version of you the engine cites is the version that exists today. The full picture of how a benchmark, a diagnosis, and a fix loop connect is what how OnlyAEO works is built around, and the arc of a brand closing a competitor gap engine by engine is documented in the FastTrackr AI case study.
Get your free AI visibility audit
OnlyAEO runs your prompt set across ChatGPT, Perplexity, Gemini, and Claude, scores cited-mentioned-recommended per engine, and tracks the head-to-head gap over time. See what it costs to measure the matchup that matters.
View pricingFrequently Asked Questions
How many prompts do I need to benchmark against one competitor?+
Should I include prompts that name both brands?+
Why compute a two-brand share instead of using my aggregate share of voice?+
How often should I re-run the benchmark?+
What does it mean if I am cited but my competitor is recommended?+

OnlyAEO
Expert insights on Answer Engine Optimization and AI visibility strategy.
Related Articles

AEO for a Two-Sided Marketplace: Getting Cited on Both Sides of the Network
A marketplace has two audiences who ask AI opposite questions. Here is how to win AI citations on both the supply and demand sides: two prompt sets, two content tracks, per-side measurement, and the entity problem that makes one side invisible.
Read article
How to Attribute a Free Trial Signup to an AI Assistant Recommendation
AI assistants send buyers who sign up for your trial, then hide in Direct traffic. Here is how to attribute a free trial signup to a ChatGPT or Perplexity recommendation using self-report, cohort corroboration, and branded-search lift.
Read article
How a CFO Should Evaluate an AEO Investment Request
A finance-side diligence framework for an answer engine optimization budget ask: the questions to put to marketing, why standard payback math breaks on a channel with no referrer, how to stage the funding, and the guarantees to reject.
Read article