Competitive Benchmarking in AEO: OnlyAEO's Approach to Tracking Brand Visibility
A practitioner walkthrough of how OnlyAEO benchmarks brand visibility against named competitors across ChatGPT, Claude, Gemini and DeepSeek, including the Gumshoe-based measurement loop, the cadence we publish, and how teams should read the monthly delta.
Key Highlights
- Competitive AEO benchmarking measures your brand's citation share against named competitors across ChatGPT, Claude, Gemini, and DeepSeek on a fixed prompt set.
- OnlyAEO runs the measurement loop monthly using Gumshoe, tracking visibility percentage, recommendation position, and persona-specific mention rates.
- A useful benchmark requires three things: a stable prompt set, a named competitor list, and consistent persona segmentation across runs.
- The monthly delta is the actionable signal. Track which competitors gained, which lost, and on which models, then prioritize content production accordingly.
- OnlyAEO publishes 500+ articles per month per client to compound visibility gains, and guarantees measurable improvement in the first 60 days.
Why visibility benchmarks are the new market share
For two decades B2B marketing teams used share of voice in media coverage and share of search in Google as the primary external benchmarks. Both worked because both reflected where buyers were spending attention. As buyers shift discovery work into ChatGPT and Claude, those benchmarks miss the most important room.
A SaaS marketing manager who runs a competitive search for "best customer data platform for mid-market" in Google sees the same SEO leaders that have ranked there for years. The same query inside ChatGPT or Claude returns a different ordering, sometimes a different vendor entirely, and almost always a different reasoning chain. If you are not measuring that, you are not measuring the room where deals are starting.
OnlyAEO runs competitive AEO benchmarking as a standing program for every client. This piece walks through how we design the prompt set, how we measure with Gumshoe, what cadence we publish at, and how growth teams should read the monthly delta to drive content decisions. For broader context on why this matters, AI competitive benchmarking for SaaS brands covers the strategic case.
The three inputs that make a benchmark trustworthy
A stable prompt set tied to buyer intent
The benchmark is only as good as the prompts feeding it. We build a fixed prompt set of 60 to 120 questions that real buyers ask, segmented by buying stage. Early-stage prompts ask "what is" and "how do I", consideration prompts ask "compare X and Y", and decision prompts ask "best X for Y". The set stays fixed for at least one quarter so month-over-month deltas are meaningful.
If the prompt set drifts every run, you cannot tell whether your visibility went up because you got better or because the question changed. OnlyAEO locks the prompt set on engagement start and reviews it quarterly.
A named competitor list with tier discipline
We define between three and seven named competitors per client, segmented into direct competitors, adjacent solutions, and aspirational benchmarks. The aspirational tier matters because it shows you what citation share is achievable when a brand has invested heavily in AEO. Without that ceiling reference, a 5% visibility gain feels great until you realize the category leader sits at 35%.
Persona segmentation for the prompts
Every prompt is tagged with a buyer persona. A growth marketing manager asks different questions than a CFO evaluating procurement, and AI models recommend differently based on the framing of the question. OnlyAEO segments results by persona so the team can see whether they are winning with the buyer who actually signs.
The Gumshoe measurement loop
OnlyAEO uses Gumshoe as the measurement backbone for every client. The tool runs the prompt set across ChatGPT, Claude, Gemini, and DeepSeek on a recurring cadence and captures the full text of each response, the citation list, and the recommendation position of every named competitor.
We pull three primary metrics from each run. Visibility percentage measures what share of all answers in the prompt set cite or recommend the brand. Recommendation position measures, when the brand is cited, where it lands in the list (first, second, third, or buried below). Persona mention rate measures visibility broken down by buyer persona so the team can see whether they are winning the right buyers or the wrong ones.
A secondary layer tracks citation source: when an AI model cites the brand, is it citing the brand's own site, a third-party listicle, a review platform, or a competitor's comparison page? That source mix tells us whether the brand owns its narrative or whether competitors are framing it. For the deeper measurement framework, see how to track LLM citations across ChatGPT, Claude, Gemini, and DeepSeek.
A sample benchmark output
The table below is a sanitized example of what a monthly benchmark report looks like for a mid-market SaaS client. Numbers are illustrative.
| Brand | Visibility % | Avg recommendation position | Persona match score | MoM delta |
|---|---|---|---|---|
| Client (us) | 18% | 2.3 | 0.71 | +4 pp |
| Direct competitor A | 31% | 1.4 | 0.82 | +1 pp |
| Direct competitor B | 22% | 2.1 | 0.65 | -2 pp |
| Adjacent solution | 12% | 3.0 | 0.55 | 0 pp |
| Aspirational benchmark | 44% | 1.2 | 0.88 | +2 pp |
| Long-tail challenger | 6% | 3.4 | 0.42 | +3 pp |
Reading this table the right way matters. The client gained four percentage points month over month, more than any direct competitor. That is the headline. But the long-tail challenger gained three points from a tiny base, which signals that someone in the category is investing in AEO and will be in the consideration set within two quarters. The team should plan accordingly.
The cadence OnlyAEO publishes at
We publish a full benchmark report monthly. The cadence matters for two reasons. First, AI models refresh their indices on rolling schedules, and a monthly cadence smooths short-term noise. Second, the content production loop runs on roughly the same cycle, so monthly benchmarking lets us tie publishing decisions to measurable visibility change.
Between monthly reports we run lightweight check-ins on priority prompts, particularly when a competitor launches a campaign or a new model version ships. The full report includes the metrics table, a per-model breakdown, persona drill-downs, the citation source mix, and a content priority list for the next 30 days.
OnlyAEO publishes 500+ articles per month per client because passage volume is the leading indicator of citation share. The benchmark report tells us which passages to write next, which competitor narratives to counter, and which persona gaps to close. The measurement loop and the production loop are tightly coupled, and that coupling is what produces compounding citation gains month over month.
How to read the monthly delta
The single most useful number on the benchmark report is the month-over-month delta. Absolute visibility tells you where you stand; the delta tells you whether you are gaining or losing ground. Here is the framework we walk clients through.
- Look at your own delta first. Anything above +2 percentage points in a 30-day window is meaningful progress. Above +5 points is exceptional.
- Compare your delta to each direct competitor's delta. If you are growing and they are flat, you are taking share. If you are both growing, the category is expanding into AI-driven discovery and the rising tide lifts everyone.
- Watch the aspirational benchmark for ceiling movement. If the category leader is gaining, the ceiling is rising and your target should rise with it.
- Investigate negative deltas immediately. A competitor losing two points without a clear cause usually means a model index refresh changed retrieval. Your gain may not be defensible without follow-up content.
- Segment the delta by model. Sometimes the headline number masks divergent performance: up four on Claude, down one on Gemini. Each model requires different optimization. Cross-model strategy is covered in why single-model strategies fail and how to do cross-platform AI optimization.
- Tie the delta to recent publishing. Which articles shipped in the last 30 days, which prompts did they target, and did visibility on those prompts move? This closes the loop.
- Plan the next 30 days from the gaps. The articles you commission this month should target the prompts where you lost ground or where the gap to the aspirational benchmark is widest.
Common mistakes in competitive AEO benchmarking
The first mistake is benchmarking against the wrong competitors. Teams often pick the brands they wish they competed against rather than the brands that actually show up in the buyer's consideration set. Run the prompt set once and list every brand that appears in the AI's recommendations. That list is your real competitive set, regardless of what the deck says.
The second mistake is treating visibility percentage as the only metric. A brand cited 30% of the time but always in position five is doing worse than a brand cited 18% of the time and always in position one or two. Recommendation position matters as much as raw mention rate, and persona match score often matters more than both.
The third mistake is benchmarking without changing the prompt set as the category evolves. We said the prompt set should stay fixed for a quarter, and it should, but quarterly review is non-negotiable. Buyer questions evolve as the category matures, and a prompt set that was current in Q1 may be tracking the wrong intent by Q3.
How OnlyAEO Approaches This
OnlyAEO runs competitive benchmarking as a continuous program tied directly to the content production engine. The monthly report is not a slide deck the agency hands over at the end of the month. It is the input to next month's editorial calendar. The prompts where the client lost ground become the article briefs. The competitor narratives that gained traction become counter-positioning pieces.
We optimize for all four major AI platforms simultaneously because the buyer is using all four. ChatGPT and Claude dominate consideration-stage research. Gemini owns the assistant and Google-integrated surfaces. DeepSeek is rising fast in technical evaluation work. A brand that wins on ChatGPT but loses on Gemini is invisible to a meaningful slice of the buying committee. The measurement loop catches this, and the production loop closes the gap.
The 60-day guarantee is built on this tight coupling. By the end of month two we expect measurable visibility improvement on the priority prompt set, and the benchmark report shows it explicitly. Beyond that, the citation rate compounds month over month as published passage volume grows and the four models refresh.
Get your free AI visibility audit
Get a free AI visibility audit. We'll show you where your brand currently stands across ChatGPT, Claude, Gemini, and DeepSeek and what it would take to get cited.
Get Your Free AuditFrequently Asked Questions
How often should we run competitive AEO benchmarks?+
How many competitors should we benchmark against?+
What is the difference between visibility percentage and recommendation position?+
Why does OnlyAEO use Gumshoe specifically?+
How long before we see benchmark improvement?+
Can we benchmark our own brand without an agency?+

OnlyAEO
Expert insights on Answer Engine Optimization and AI visibility strategy.
Related Articles

Citation Quality vs Citation Quantity: The OnlyAEO Framework
A 10-citation week can outperform a 100-citation week if quality is right. Here is the OnlyAEO framework for citation quality vs quantity, the four quality dimensions that matter, and how to grade every AI citation that lands.
Read article
Mentioned vs Recommended: The Citation Distinction That Matters
Being named in an AI answer is not the same as being recommended. Here is how to measure recommendation rate and move from one to the other.
Read article
Share of Voice in AI Answers: How to Measure It
AI share of voice is your slice of brand mentions and citations across AI answers. Here is how to define, compute, and benchmark it properly.
Read article