Across the buying questions we track, 78.3% of every citation in our categories goes to a corporate blog article or listicle, measured across 96,076 citation records captured between 10 June and 8 September 2026.
A first-party case study from The Rank Masters, measured with Amadora across ChatGPT, Perplexity and Gemini, and across AI Overviews and AI Mode in Search Console.
We sell AI search visibility to B2B SaaS companies, so the fairest test of our own AEO, GEO and LLM SEO playbook was to run it on ourselves, instrument every result, and publish what came back. This is that data, including the parts that did not go our way.
▶️ If your ranked pages are not showing up in AI answers and you want a content system that fixes that, book a SaaS content strategy call.
Table of Contents
- Why We Instrumented Our Own Site First
- What We Are Actually Measuring, and Why the Distinction Matters
- The Position We Hold Across 64 Commercial Prompts
- Where the Citations Actually Come From
- The Engine Split That Nobody Is Optimising For
- Why the Roundup Format Works, Proven From the Citation Universe
- How Engines Actually Resolve a Buyer Question
- What Google's AI Surfaces Show
- Does Any of This Reach Actual Humans
- What It Produced for the Business
- What We Would Tell a SaaS Team Starting This
- Methodology and Limitations
Why We Instrumented Our Own Site First
We ran a working AI search programme with no way to measure it, which meant we were arguing from screenshots rather than from data.
Here is the honest version of where we were. Search Console and GA4 were properly configured and told us a clear story about impressions, rankings and clicks. That part was fine. What neither tool could see was the layer our clients actually pay us to influence, namely whether a language model names you when a buyer asks it for a shortlist.
We knew it was happening. We kept running into our own articles quoted inside ChatGPT, Gemini and Google's AI Overviews. But the only way we ever confirmed it was opening the tools and checking by hand, maybe once a month, three or four hours across whatever queries we happened to remember that day.
We hold client programmes to a standard of continuous measurement, and we wanted our own work held to that same standard before publishing another word about it. Not because anything was broken, but because "we have seen it happen" is a weaker sentence than "here is the number, here is the date, here is the engine." Only one of those survives a conversation with a CFO.
There was a second reason, and it turned out to matter more. Manual checking means checking ChatGPT, because ChatGPT is the one you have open. Ninety days of instrumented data across every major LLM search engine eventually showed us that ChatGPT is where we perform worst. We had been pointing our only measurement habit at our weakest channel, and there was no way to know that without tracking all of them at once.
That gap between what we could see and what was actually happening is the same gap we now close for clients through our LLM SEO programme, and we were not willing to sell it without having run it on ourselves first.

What We Are Actually Measuring, and Why the Distinction Matters
We track 64 commercial-intent prompts across 15 topics, continuously, on every major LLM search engine.
Precision matters here more than it might seem, because most published AI visibility numbers fall apart the moment you ask what they counted. So before any results, here is exactly what sits behind them.
| Element | Detail |
|---|---|
| Prompts tracked | 64 |
| Topics | 15 |
| LLM search engines | ChatGPT, Perplexity, Gemini |
| Google AI surfaces | AI Overviews and AI Mode, via Search Console |
| Locale | US |
| Reporting window | 10 June to 8 September 2026, 90 days |
| Owned citations captured | 6,831 |
These 64 prompts are not informational queries. They are shortlist questions, the ones a buyer types when they are close to choosing something and want to know who the credible options are. That distinction is the whole point. Ranking for "what is generative engine optimization" is pleasant. Being named when someone asks "best AI visibility tools with citation tracking" is commercial. Building prompt sets that sit at that end of the funnel is the first thing we do on every B2B SaaS SEO engagement, because a prompt set full of informational queries will flatter you and sell nothing.

Two metrics do most of the work in this case study, and they are not the same thing.
Visibility rate is the share of captured answers to a given prompt in which The Rank Masters is named. If a prompt shows 27%, roughly one in four answers to that question mentions us by name.
Citations count the times one of our URLs is linked as a source inside an AI answer. A brand can be named without being cited, and cited without being named. Conflating those two is the most common way AI visibility reporting misleads people, so we report them separately throughout.
One clarification on coverage. Amadora tracks our prompts across ChatGPT, Perplexity and Gemini. For Google's AI Overviews and AI Mode we use Search Console's Generative AI Features report, because that is where Google reports its own AI impressions. Every number below is labelled with the tool it came from.
The Position We Hold Across 64 Commercial Prompts
We appear on 34 of the 38 prompts with a full window of data, which is 89.5%, at an average mention position of 4.2.
Here is the complete picture for the 90 days.
| Measure | Value |
|---|---|
| Prompts where we appear at all | 47 of 64 (73.4%) |
| Prompts with a full window of data | 34 of 38 (89.5%) |
| Prompts at 20% visibility or above | 21 |
| Prompts at 25% visibility or above | 14 |
| Average mention position when named | 4.2 |
| Owned citations earned | 6,831 |


Let us walk through what those numbers actually mean, because the headline figure is not the interesting one.
The gap between 73.4% and 89.5% is explained by 26 prompts added at the very start of September. They have only a few days of data each, so they pull the average down without telling us anything useful yet. The 89.5% figure covers the 38 prompts that ran the full window and had time to produce a stable sample. That is the number we would put in front of a client, because it describes prompts we have genuinely competed for rather than prompts we have only just started watching.
Average mention position of 4.2 deserves a pause. When an engine names us, we are typically in the top half of the list it produces. Anyone who has watched a buyer skim an AI answer knows the first four or five names get read and the rest get scrolled past. For an independent publisher competing for shortlist slots against funded software companies with far bigger brand signals, appearing fourth rather than eleventh is where most of the commercial value sits.
Visibility by Topic, Including the Clusters That Underperformed
Our strongest clusters are voice cloning and proofreading at 36.2% and 25.0%, not the AI visibility cluster we expected to lead.
| Topic | Prompts | Mean Visibility | Best Prompt |
|---|---|---|---|
| Best voice cloning software | 3 | 36.2% | 38.4% |
| Best AI content generator tools | 2 | 26.7% | 36.9% |
| Brand monitoring | 2 | 26.1% | 27.3% |
| AI proofreading tools | 3 | 25.0% | 26.5% |
| Website traffic analysis | 1 | 23.4% | 23.4% |
| AI visibility | 15 | 17.8% | 40.0% |
| Best site audit tools | 2 | 16.7% | 29.4% |
| Best backlink monitoring | 2 | 12.3% | 14.8% |
| Best tools, generic | 8 | 5.0% | 10.0% |
| TRM services | 5 | 4.4% | 22.2% |
| Alternatives | 13 | 3.6% | 14.6% |
Read that table slowly, because it holds three findings that changed how we work.
The method travels across subjects. Voice cloning and proofreading outperform AI visibility, which is supposed to be our flagship subject. We know AI visibility better than we know voice cloning software, yet our voice cloning roundup and our AI proofreading roundup still outperform it. That tells us the results come from how the content is built rather than from how much we happen to know about the subject. If subject-matter authority were driving this, the ranking would be reversed.

Specificity is the single biggest lever we found. Look at the generic "best tools" bucket sitting at 5.0%, then look at the AI visibility cluster at 17.8% with a top prompt at 40.0%. Same site, same domain authority, same editorial standard, same authors. The only variable that changed is the shape of the question. Every prompt in the AI visibility cluster carries a qualifier such as with Looker Studio, by country, or for competitor benchmarking. Every prompt in the generic bucket does not. A qualified question is narrow enough to own. A generic one is a popularity contest against Semrush and Ahrefs, and nobody wins that one on brand.


Some categories are not winnable, and knowing which is worth money. Our Alternatives cluster sits at 3.6%, our weakest asset class by a wide margin. On queries like "ahrefs alternative" and "clearscope alternative", the engines overwhelmingly prefer the vendors' own comparison pages, because a vendor owns its own brand entity in a way no independent publisher can out-write. Pages like our SE Ranking alternatives roundup and our Majestic alternatives roundup are well built and still lose here. We are telling you that because a case study that reports only the wins is an advert, and because recognising an unwinnable category early is worth more than writing a better article inside it.

Where the Citations Actually Come From
Five URLs carry 52.4% of every citation we earned, and twenty URLs carry 90.4%.
6,831 citations sounds like breadth. It is not. It is concentration, and this was the most useful thing the data told us all quarter.
| Rank | Page | Citations |
|---|---|---|
| 1 | Best AI proofreading tools | 906 |
| 2 | AI video software with voice cloning | 877 |
| 3 | AI visibility tools with Looker Studio | 730 |
| 4 | AI search visibility audit tools | 548 |
| 5 | Best Jasper alternatives | 521 |
| 6 | Best site audit tools | 456 |
| 7 | GEO prompt monitoring tools | 350 |
| 8 | Best AI content generator tools for SaaS | 316 |
| 9 | Best backlink monitoring tools | 294 |
| 10 | Best website traffic analysis tools | 264 |
Seven pages from the AI visibility cluster sit inside the top twenty and account for 30.2% of every citation the entire site earned.
Here is why this matters more than the total. Most content teams operate on a volume assumption, i.e. publish forty articles a quarter, hope some get picked up, repeat. This data says that is the wrong model for AI citation earning. What actually happens is that a small number of pages become the canonical answer for a whole cluster of related questions, and then they get cited repeatedly across many different phrasings of those questions. Everything else in the library contributes almost nothing.
If you are running a content calendar built on volume, the honest read of this table is that most of your output is not doing the job you think it is doing. Depth on the pages that can realistically become canonical beats breadth across pages that cannot. That is a resourcing decision rather than a writing decision, and you cannot make it without URL-level citation data.
Executing this well is exactly the gap The Rank Masters closes for B2B SaaS teams, building an ICP-led content system that maps each topic cluster to a money page and to pipeline, rather than publishing posts that never convert.

The Engine Split That Nobody Is Optimising For
Gemini supplies 64.2% of our citations. ChatGPT supplies 15.4%.
This is the finding we did not see coming, and the one we would most want another agency to check against their own data before writing another strategy deck.
Across our top twenty cited URLs, here is where those citations came from.
| Engine | Citations | Share |
|---|---|---|
| Gemini | 3,966 | 64.2% |
| Perplexity | 1,259 | 20.4% |
| ChatGPT | 948 | 15.4% |
Visibility rate tells exactly the same story from the other direction.
| Engine | Visibility Rate | Share of Voice | Average Position |
|---|---|---|---|
| Gemini | 32.6% | 4.14% | 3.8 |
| ChatGPT | 8.2% | 1.08% | 5.1 |
| Perplexity | 4.8% | 0.56% | 5.0 |
Gemini names us on roughly one in three answers and supplies close to two-thirds of our citations. ChatGPT, the engine that dominates every conversation about AI search and the one nearly every published AEO strategy is written for, cites us least of the three.
Sit with that for a second, because the implication is uncomfortable for the whole discipline. An enormous amount of AEO advice is written from ChatGPT observation, because that is the engine practitioners have open and the one clients ask about by name. If your citation profile looks anything like ours, that advice is being generated from the weakest available signal.
We want to be careful about how far we push this. We are one site in one category. Your engine mix might be completely different, and your buyers might genuinely live in ChatGPT. The actual lesson is not "optimise for Gemini." It is that engine mix is a strategic input rather than a footnote, it varies by category, and nobody can know theirs without measuring all of them at the same time. We ran a competent programme for months on an assumption that turned out to be backwards.
Establishing that mix is now the opening move on every generative engine optimization engagement we take, before a single page gets written or refreshed, because getting it wrong sends an entire quarter of effort at the wrong retrieval behaviour.
Why the Roundup Format Works, Proven From the Citation Universe
78.3% of all citations in our categories go to corporate blog articles and listicles, measured across 96,076 citation records.
This is the part of the data we find most useful strategically, because Amadora captures the full citation universe for our tracked prompts and not only our own URLs. That means we can see every source the engines pulled from, which gives an unusually direct read on what kind of page actually wins here.
| Source Type | Share of All Citations |
|---|---|
| Corporate blog articles | 42.7% |
| Corporate listicles | 35.6% |
| Corporate website pages | 9.4% |
| Commercial listicles | 2.9% |
| Community threads such as Reddit and forums | 2.0% |
| Editorial, news, academic and social combined | 3.1% |
Listicle-format pages across all site types account for 39.4% of citations. Community, social and news content combined account for roughly 3%.
So when we say our editorial roundup format is the engine of the whole programme, that is not a stylistic preference rationalised after the fact. It is what the citation data says these models reach for when a user asks them for a shortlist. We did not reason our way to the format. We measured which format gets cited and then committed to it properly.
This table also explains the Alternatives weakness flagged earlier, and it is a good example of using data to stop doing something. The corporate pages winning on "X alternative" queries are overwhelmingly the vendors' own. They hold brand-entity advantage on their own product name that an independent publisher cannot out-rank by writing a better article. That category does not need more effort from us. It needs a different play entirely, or it needs dropping.
Working out which format your category rewards, and which formats to abandon, is the diagnostic half of answer engine optimization. It is unglamorous and it saves more budget than any single page ever will.

How Engines Actually Resolve a Buyer Question
One prompt fans out into many sub-queries, and we captured 33 distinct ones the engines ran for a single cluster.
This is the mechanic most AEO advice skips, and here it is visible directly in the data rather than inferred.
When someone asks Gemini "which AI visibility tool is best for Looker Studio and citation tracking," the model does not take that string and run it as a search. It decomposes the question into several sub-queries and searches each one separately. For our 15 AI visibility prompts in this window, Amadora captured 33 distinct sub-queries the engines actually executed.
| Sub-Query the Engine Ran | Times Executed |
|---|---|
| best AI visibility tools with Looker Studio | 40 |
| AI visibility tool track prompts by country | 36 |
| tools that track sources cited by ChatGPT | 32 |
| best AI visibility tool for Looker Studio citation tracking | 32 |
| how to run AI search visibility audit for top keywords | 20 |
| AI search visibility audit top keywords methodology | 15 |
| AI visibility tools Looker Studio integrations GEO dashboards | 8 |
One buyer question turned into a dozen retrieval attempts, each phrased slightly differently, each one a separate chance to be found or missed.
Once you have seen that, the whole job changes shape. You are not optimising a page for one query. You are optimising it to be the best retrievable answer for an unpredictable family of sub-queries generated at runtime, in phrasings you never chose. That produces five concrete editorial rules, and every one of them is downstream of the table above.
1. Cover the whole fan-out inside a single page. Every plausible sub-query a model might generate should find a passage on the page that answers it cleanly. This is why our roundups carry dedicated sections for integrations, geography, pricing, use case and methodology, instead of a flat list of tools with a paragraph each. The flat list answers one query. The sectioned version answers twelve.
2. Write passages that survive being lifted out. Models retrieve chunks, not documents. A paragraph that only makes sense after you have read the three above it cannot be cited, because the retriever will never pull those three along with it. Every section has to stand alone, name its own entities, and avoid orphan references such as this, it, or the above.
3. Lead with the answer, then the evidence. Answer-first structure is what makes a passage extractable. The build-up-then-conclusion structure that reads beautifully in an essay is exactly what makes a passage invisible to a retriever, because the useful sentence ends up buried where no chunk boundary will find it.
4. State entities and attributes explicitly. "Supports Looker Studio" written as a stated attribute in a comparison table is retrievable. The same fact implied across two sentences of prose is not. This is the most mechanical change most teams can make, and it is why we use structured specification blocks and comparison matrices rather than describing features narratively.
5. Match the format the category actually rewards. For us that is the structured comparative roundup, and 78.3% of citations in our categories say so. Check yours before assuming it is the same.
We call this AEO, GEO and LLM SEO depending on which words a given buyer uses, but underneath the labels it is one job, namely making a page the most liftable answer to a question that has not been phrased yet.
Those five rules are the editorial spine of our GEO and SEO growth programme, applied cluster by cluster against a live fan-out map rather than a keyword list, which is what separates content that gets retrieved from content that merely ranks.

What Google's AI Surfaces Show
780,609 impressions came from Google's AI Overviews and AI Mode in the same 90 days, which is 17.31% of all our Search Console impressions.
Amadora covers the LLM search engines. For Google's AI surfaces we use Search Console's Generative AI Features report, and it tells a story that runs alongside the citation data rather than duplicating it.
| Measure | Value |
|---|---|
| Impressions from Google AI surfaces | 780,609 |
| Share of all Search Console impressions | 17.31% |
| June 2026 | 135,681 |
| July 2026 | 244,321 |
| August 2026 | 379,470 |
| First half of window vs second half | Up 76.1% |
| AI features as share of impressions, first half vs second half | 11.8% to 23.6% |
Almost a quarter of our Search Console impressions now come from a surface that answers the buyer's question without requiring a click. That is the structural shift this entire discipline exists to address, and here it is measured on a property we own rather than argued from a vendor's market forecast.
Now the part most reports leave out. Over this same period our classic organic clicks declined while AI surface impressions grew. Presence went up. Clicks went down. Those two moved in opposite directions at the same time, on the same site.
We are including that because it is the single most important thing for a marketing leader to internalise about this era. If clicks are your only success metric, this data set reads as failure during precisely the period when the citation footprint was compounding fastest. A team measuring clicks alone would have concluded the programme was broken and cut it, right at the point it was working. Presence and clicks are now two separate signals, and you have to watch both or you will make the wrong call with total confidence.



Does Any of This Reach Actual Humans
572 sessions arrived directly from AI assistants, engaging for 19.5 seconds on average, almost identical to organic search visitors.
This is where most AI visibility case studies go quiet, because the traffic numbers are small. Ours are small too. Here they are anyway, because the behaviour inside them is more interesting than the volume.
| Channel | Sessions | Average Engagement Time |
|---|---|---|
| Organic search | 3,460 | 21.3s |
| AI assistant referrals | 572 | 19.5s |
| Direct | Excluded | 3.6s |
Visitors sent to us by ChatGPT, Gemini and Claude behave almost exactly like organic search visitors. They arrive, and then they read. That is the confirmation the whole model needs, because it means citation earning and human arrival are two ends of the same pipeline rather than two unrelated metrics that happen to sit on the same dashboard.
We are deliberately excluding direct traffic from every figure in this case study. Direct accounts for the large majority of raw session volume on our insights pages, at an average engagement time of 3.6 seconds. Three-second sessions at that scale are not people reading articles. Counting them as an audience would inflate our numbers and teach you nothing, so every traffic figure above is non-direct only, giving us 5,018 sessions in total.
The pages pulling AI assistant referrals are the same pages pulling citations, led by the voice cloning roundup and the brand visibility tracking pieces, with the AI brand mention tracking roundup close behind. The pipeline holds end to end.

What It Produced for the Business
Earned citations generated inbound collaboration approaches from SaaS tools we had covered, and 3 to 4 of those conversations are now live prospects in our pipeline.
This took time, and we want to be clear about how much. Our AEO, GEO and LLM SEO work needed roughly two to three months of maturation before it produced anything commercial. Citations came first. Approaches followed. Nobody should expect both inside the same quarter.
Earned collaboration approaches. A number of SaaS tools in the categories we cover reached out to us directly, having found us through the roundups they were being cited alongside. These were traceable to AI search citations and to specific clusters, with the AI visibility cluster driving the most. That is a genuinely different acquisition motion from anything traditional SEO produced for us. These companies were not out searching for an agency. They encountered us as the source an AI engine cited when someone researched their own category, and drew the obvious conclusion about who understood that category.
Independent confirmation from outside our own dashboard. One inbound approach arrived from Noble, and it opened like this.
"Over the last few months, therankmasters.com was cited 432 times in AI answers, across the buying questions we monitor for our clients. Which means your comparison and category pages are the source material AI tools pull from when buyers research your space."
We did not commission that. It is somebody else's tracking, on somebody else's prompt set, monitored for their own clients, arriving independently at the same conclusion our Amadora data supports. Their 432 and our 6,831 are not the same number because they are not counting the same questions, and that is precisely why it is useful. Two different prompt sets, two different tools, one consistent finding about which pages the engines treat as source material.

Inbound demand for the service itself. Prospects now arrive already knowing what they want, which is new. One recent enquiry read simply as "I need help with GEO, want to see if you could set it up so Claude can handle it afterwards. I like to know my options to improve the GEO of our website."
That is a buyer who has skipped the entire education stage. They are not asking whether AI search matters, or what GEO is, or whether it deserves budget. They arrived at the options conversation. Anyone who has spent years opening sales calls by explaining why organic content is worth investing in will recognise how different a starting position that is.
We are not publishing revenue figures for this period. The conversations are recent, the pipeline is young, and we would rather report a real number later than an impressive one now.
What We Would Tell a SaaS Team Starting This
Instrument before you optimise, because you cannot fix an engine mix you cannot see.
Five things we would do differently starting again, each one earned from the data above rather than from opinion.
Measure every engine at once, from day one. We ran a competent programme for months while spot-checking the engine that cites us least. Every hour of that manual checking pointed at our weakest channel and we had no mechanism to discover it. This is not a small optimisation. It is the difference between working from signal and working from habit.
Expect concentration and resource for it. Five pages carry over half our citations. Plan for a small number of pages to become canonical for a cluster and fund them properly, rather than spreading the same effort thinly across a calendar that produces mostly inert pages.
Chase qualified questions, not popular ones. Generic prompts return 5.0% visibility for us. Qualified prompts return up to 40.0%. The qualifier, whether a platform, a geography, a use case or an integration, is what makes a question narrow enough to own. Popular questions are auctions you lose to whoever has the biggest brand.
Identify the categories you cannot win, then stop. Our Alternatives cluster sits at 3.6% because vendors own their own brand entities. Recognising that early saved more budget than any on-page improvement inside it would have earned.
Watch presence and clicks as separate signals. Ours moved in opposite directions this quarter. Whichever one you are not watching is the one that will mislead you.
If you would rather not spend a quarter learning these five things the way we did, they are already built into how we run the GEO and SEO growth programme, and our pricing is published so you can size it before you ever speak to us.
Frequently Asked Questions
How long does it take for AEO and GEO work to produce citations?
Is visibility rate the same thing as a citation?
Which AI search engine should a B2B SaaS company optimise for first?
Why do listicles and roundups perform so well in AI answers?
Does being cited in AI answers actually send traffic?
Should we still care about clicks if AI surfaces are growing?
What tools do we need to measure this properly?
Methodology and Limitations
Every figure above comes from Amadora, Google Search Console or GA4 for the 90 days from 10 June to 8 September 2026.
Instrumentation. 64 prompts across 15 topics, US locale, tracked continuously in Amadora across ChatGPT, Perplexity and Gemini. Google AI Overviews and AI Mode measured separately through Search Console's Generative AI Features report. Session and engagement data from GA4.
Framing. These figures describe the position we held during the window. They are not presented as gains against a prior period, because no equivalent instrumentation existed beforehand to measure against, and we would rather publish a defensible position than a flattering comparison.
Limitations we want stated plainly.
- 26 of the 64 prompts were added at the start of September and carry limited sample. Where it matters we report the 38 full-window prompts separately.
- Average mention position and visibility rate are engine-reported at capture time and will vary with model updates outside anyone's control.
- This is a single-site study of one B2B services company in one category. It demonstrates that the methodology produces a measurable citation position here. It does not establish a benchmark for your category.
- Engine mix, format performance and category winnability all appear to be category-specific. Treat our numbers as a worked example of the method rather than as targets.
If your pages rank but never get cited, the gap is measurable and it is fixable. Book a SaaS content strategy call and we will map your highest-intent buyer questions to the engines that actually cite in your category, then to the pages that can win them.




