An LLM competitive benchmark measures who ChatGPT, Perplexity and Gemini cite when a customer asks a question in your market, you and your competitors alike. The goal is not the Google ranking, but your share of citations in generative answers. The method has six steps: build a prompt list that mirrors the buying journey, run them on each model under standardized conditions, log the brands cited in a prompts × models matrix, calculate your share of citations against the competition, prioritize actions, then re-measure every month. The result is a precise map of your AI visibility: which questions you exist on, which a competitor occupies alone, and where the window is still open. It is the starting point of any serious GEO strategy, because you only fix what you measure. Here is the exact protocol, reproducible every month.
Why benchmark on LLMs
Because your prospects now ask a model their questions before opening Google. If ChatGPT recommends three competitors and never your brand, you lose the sale before you were even in the race. The LLM competitive benchmark measures exactly that gap.
The logic differs from classic SEO. In traditional search, you track your position on a query. In generative search, there is no results page: the model synthesizes an answer and cites a few brands. So the right metric is no longer rank, but share of citations: across all the prompts in your market, how many answers mention you, and how many mention each competitor.
This distinction is both measurable and strategic. The signals that trigger an AI citation are not the ones that drive Google rankings. Ahrefs' analysis of 200,000 domains (Dec. 2025) shows that off-site brand mentions correlate more strongly with AI citations (YouTube 0.737, Reddit, Wikipedia) than Domain Rating (0.266). You can dominate SEO and stay invisible in ChatGPT. The overlap is low: only 11% of domains are cited by both ChatGPT and AI Overviews.
An LLM benchmark does not measure your rank, but your share of citations in generative answers, you and your competitors included. It is the only metric that reflects what a prospect querying an AI actually sees.
Before measuring, clarify the scope: your market, your three to five direct competitors, and the questions a buyer is asking. This is the foundation of a structured GEO agency approach, where every number feeds a decision rather than a gut feeling.
Building the prompt list
Everything depends on the quality of your prompts. A benchmark is only as good as how representative its tested questions are: they must reproduce what a real prospect asks, not what you wish they would ask.
Structure the list along the buying journey, in three families. This split avoids the classic bias of testing only branded queries, where you always win and learn nothing.
The three prompt families
| Family | Intent | Example prompt |
|---|---|---|
| Discovery | The prospect explores a problem, with no solution in mind | “How can I improve my site's visibility in ChatGPT?” |
| Comparison | The prospect compares approaches or providers | “Best GEO agencies in France in 2026” |
| Decision | The prospect wants to validate a specific choice | “Which agency should I use to optimize my AI visibility in Albi?” |
Aim for 20 to 40 prompts total, split across these families. Below 20, a single atypical answer skews your percentages. Phrase them in natural language, the way you talk to an assistant, not in telegraphic keywords. Vary the angles: “best,” “how to choose,” “alternatives to,” “for [industry].”
Systematically include prompts where you expect to see your competitors. That is where the benchmark becomes a competitive tool and not a mere ego test. Also document each prompt in close variants: models are sensitive to phrasing, and a reworded question can make a brand appear or disappear.
Running the test on each model
Run each prompt on each model under standardized conditions, otherwise results are not comparable from one wave to the next. The protocol matters as much as the prompts.
Test in private browsing or on a dedicated account, with no memory or personalization enabled. A personal account's history biases answers toward your own past searches.
Run each prompt on ChatGPT, Perplexity and Gemini at a minimum. They rely on distinct citation mechanisms and do not return the same brands.
For each answer, note every brand or domain mentioned, its position in the answer, and whether it is cited as a source or recommended in the text.
Keep a screenshot or the raw text of each answer. Models evolve; without an archive, you will be unable to verify or compare the following month.
LLMs are not deterministic. Run each prompt twice and record a brand as present if it appears at least once.
One technical detail weighs heavily on results: LLMs do not execute JavaScript. If your page content loads client-side, the model's crawler sees only a blank page. Server-side rendering (SSR) or static HTML is therefore essential to exist in the index that feeds these answers. A competitor missing from your benchmark despite strong awareness often has this exact problem.
ChatGPT alone represents an audience volume that justifies including it in every benchmark. Ignoring this model means ignoring the leading generative search interface.
To compare the solutions that automate this collection at scale, see our 2026 GEO tools comparison. Manual collection is enough for a first wave; a tool becomes relevant as soon as frequency and volume rise.
Reading the results matrix
The matrix is a table with prompts as rows, models as columns, and each cell listing the brands cited. It is the centerpiece of the benchmark: it turns dozens of answers into a readable map of your AI visibility.
Read the matrix along three diagnostic axes. Each cell configuration tells a different competitive story and calls for a different action.
The three zones to spot
| Cell configuration | What it means | Action priority |
|---|---|---|
| You + competitors cited | You exist on this question, the market is shared | Consolidate: strengthen your relative position |
| Competitor cited alone | They own the space, you are invisible | Attack: create the missing citable content |
| No relevant brand cited | The model improvises or cites off-topic | First mover: open window to seize quickly |
The third zone is the most profitable. When no player in your market is cited on a high-intent prompt, the first brand to publish factual, structured, citable content takes it all. It is the opposite of a head-on battle: you occupy empty ground.
Also spot the gaps between models. A brand cited on Perplexity but absent from ChatGPT reveals a precise citation signal to work on: web sourcing for one, off-site awareness for the other. The matrix tells you not only where you lose, but why.
Calculating share of citations
Share of citations is the score that condenses the entire matrix into one comparable figure. You calculate it by dividing the number of answers where you appear by the total number of answers tested, then repeating the operation for each competitor.
On 30 prompts × 3 models, you have 90 answers. If your brand appears in 18 of them, your raw share of citations is 20%. Do the same calculation for your competitors: you get an AI visibility ranking that often bears no relation to your market's Google ranking. This metric is detailed in our guide on AI share of voice, which explains how to weight it by each prompt's business value.
Then refine it with two weightings. First position: a brand cited as the first recommendation weighs more than one relegated to the end of the answer. Second prompt value: a citation on a decision question is worth more than one on a generic discovery question. A raw share of citations of 20% can hide real strength if it concentrates on prompts that convert.
An immediate snapshot of your share of citations against your direct competitors is available without building the whole protocol: our AI Visibility Score gives you a first figure in a few minutes, ideal for framing your first benchmark wave.
Turning it into an action plan
A benchmark with no action plan is a dead report. Each zone of the matrix translates into a concrete GEO project, prioritized by the gap between the prompt's business value and your current absence.
Prioritize by a simple rule: start with the decision prompts where a competitor is cited alone, then the discovery prompts where the window is open. The former recover sales; the latter build foundational authority.
For each decision prompt where you are missing, create or enrich a page that answers the question directly, with a standalone citable passage of 134 to 167 words placed up front.
Add FAQPage schema on these pages: it is a strong signal for AI Overviews, making it easier for models to extract your question-answer pairs.
Where a competitor dominates with no obvious SEO edge, strengthen your mentions on YouTube, Reddit and the sources models favor. That is what correlates most with citations.
Audit each target page: if content depends on JavaScript, switch to SSR or static HTML so the model's crawler actually sees your text.
Tie these actions to a structured diagnosis; our GEO audit method and pricing details how to scope and sequence this work.
The citable passage deserves particular attention. A paragraph of 134 to 167 words, factual and self-contained, that fully answers the sub-question, is the unit models extract. Too short, it lacks substance; too long, it dilutes the answer and loses its citability. It is the optimal format observed in content that actually gets quoted.
Industrializing monthly tracking
A one-off benchmark is a snapshot; the value comes with the series. Re-run the same prompt list a month later, under the same protocol, to measure the real movement of your share of citations.
Document each wave in the same table to track the trajectory. A share of citations that climbs wave after wave is proof your GEO strategy works, well before revenue confirms it. Without a second wave, you will never know whether your actions moved anything.
To go from one-off collection to continuous monitoring, automate citation detection between waves: our method to track your mentions in ChatGPT explains how to be alerted the moment an answer cites you or a competitor, without re-running the whole benchmark. You keep the benefit of the monthly protocol while capturing intermediate movements.
A client case's progress figures remain illustrative: what matters is your own curve, measured with a stable protocol. A benchmark comparable month over month is worth a thousand impressive but non-reproducible reports.
Our free GEO audit benchmarks your visibility against your competitors on ChatGPT, Perplexity and Gemini, and hands you the prioritized action plan.
Questions fréquentes
How many prompts do you need for a reliable benchmark?+
Plan for 20 to 40 prompts per market for a representative first snapshot. Below 20, statistical noise distorts your conclusions; above 40, collection cost explodes with no major gain in information. Split them across discovery, comparison and decision questions to cover the entire buying journey.
Which models should you run the benchmark on?+
At a minimum ChatGPT, Perplexity and Gemini, because they cover most usage and rely on different citation mechanisms. ChatGPT alone has over 900 million weekly users. Add Claude and Google AI Overviews if your audience uses them. Test each model in a fresh session, with no history, to avoid personalization.
How often should you repeat an LLM competitive benchmark?+
Once a month is enough to track a trend, since model answers evolve with updates and the web index. Keep the same prompt list and the same protocol from one wave to the next, otherwise the variations are no longer comparable. Quarterly tracking remains acceptable for a stable, low-competition market.
Does the LLM benchmark replace SEO rank tracking?+
No, it complements it. Google rankings and AI citation share only partially overlap: only 11% of domains are cited by both ChatGPT and AI Overviews. Tracking both gives a complete view of your visibility, from the blue link to the generative answer.
Do you need a paid tool to benchmark your AI visibility?+
Not to get started: a spreadsheet and an hour of manual collection are enough for a first wave on 20 prompts. A dedicated tool becomes useful when you scale up (40+ prompts, several models, automated monthly tracking) because it makes collection reliable and stores wave history. The choice depends on your target volume and frequency.



