We audited an AEO agency's own AI visibility
Why this one
Most AI Visibility audits are unfair in a boring way. You pick a company that has never thought about how AI assistants describe it, you find that AI assistants describe it badly, and you have proved nothing except that unprepared brands are unprepared.
So we picked the hardest possible target. The subject of this audit is an agency that sells Answer Engine Optimization to B2B SaaS companies. Getting brands named by ChatGPT, Claude, Perplexity and Gemini is not a side project for them, it is the product they sell. If anyone should score well on the thing we measure, it is a firm that sells the thing we measure.
We are not naming them, and here is why that matters
We audited this company without asking, so we are not putting their name in a public post. They did not ask for this, they are not a customer, and they had no involvement in the run. They get their own results in full, named, privately, with every raw engine answer behind them. That report is theirs. This post is the pattern.
That is our standing policy, not a one-off. Audit #001 on this blog audited a national spa franchise and never named it either. We will keep publishing what we measure, and we will keep the subject out of the headline. If you would rather find out this way than by reading your own name in someone's marketing, that is the point.
It also means we have removed more than the name. Founding date, headcount, client names, published pricing, and the specific pages we looked at are all out, because those things identify a company just as reliably as a name does. What is left is the measurement, and the measurement is the part that transfers to you.
One thing worth saying plainly. This company is good at its job by every ordinary measure. Their website is better built for machine reading than almost anything we audit. They publish an llms.txt. Every page we checked carries JSON-LD. Their headings are deep and answer-shaped. This is not a company that forgot to do the homework, and that is exactly why the result is interesting.
The result: 23.6 out of 100 on the seven discovery questions we tested, with a 95% interval of roughly 20 to 27. PROVEN, n=5 runs per question per engine, 175 scored discovery results, methodology v2.0.3, measured 2026-07-27.
Ask an AI assistant about this agency by name and it scores 95.3 out of 100, interval ±1.2 (PROVEN, n=5, 25 scored results). Ask it the question a buyer actually types, and the agency mostly is not there.
Every interval in this post is naive. It assumes each result is independent, and they are not, because results cluster by engine. The real intervals are wider than the ones we print.
The gap between those two numbers is the entire post.
What we ran, and what we did not
We asked eight questions. Seven are discovery questions, the kind someone types when they do not yet know who to hire. One is a branded question, included as a control.
Each question went to five AI assistants. Each combination ran five separate times, because a single run is not a measurement. Two hundred scored results, zero failed calls.
| Discovery score (7 questions) | 23.6 / 100 (95% interval roughly 20 to 27) |
| Branded score (1 question, control) | 95.3 / 100 (±1.2) |
| All 8 questions combined | 32.6 / 100 |
| Runs per question per engine | 5 |
| Scored results | 200 (175 discovery, 25 branded) |
| Failed calls | 0 |
| AI assistants | ChatGPT, Claude, Gemini, Perplexity, Grok |
| Methodology version | v2.0.3 |
| Audit date | 2026-07-27 |
All PROVEN, measured in this run.
Four things this audit does not tell you, stated plainly because a teardown that hides its limits is marketing, not measurement.
It is a snapshot, not a rank. Eight questions on one day. A different eight questions would give a different number. We are not claiming this agency is the eleventh best in its category, and nothing here should be read as a ranking.
The floor is 10, not 0.In our scoring, an answer that never mentions the brand still scores 10 out of 100, because the Context factor defaults to neutral when there is no mention to judge. So “10” in the tables below means total absence. It does not mean “slightly visible.” While preparing this post we also found that our two code paths for that default disagree. The live path scores a non-mention 50 on this factor and a fallback path scores it 0. This run used the live path throughout. We are reconciling it, and we will restate any number here that moves.
Competitive Share is harsh in this category.That factor measures the brand's share of all agencies named in an answer. These answers routinely name ten to sixteen agencies, so every single agency in this category scores near the floor on it. The audited agency scored 6.6 out of 100 there on discovery questions, and so would almost anyone else. Do not read that number as a verdict on them.
We scored one brand. We counted how often thirty other agencies were named, which is a real measurement. We did not audit those agencies, and their appearance counts are not scores.
The finding: they own one phrase, and vanish next to it
This is the part worth your time.
Here is how often the audited agency was named, per question, out of 25 runs each. PROVEN, n=5 per engine.
| The question we asked | Named in | Score (95% interval) |
|---|---|---|
| “Which agencies specialise in answer engine optimization (AEO) for B2B SaaS companies?” | 19 / 25 | 53.0 ±11.6 |
| “Recommend an agency that tracks and improves brand mentions in AI answersfor B2B software” | 6 / 25 | 26.3 ±10.5 |
| “Who can help my B2B SaaS company get recommended by ChatGPT, Claude and Perplexity?” | 5 / 25 | 29.7 ±11.5 |
| “We are invisible in AI search. Which agency should I hire to fix that?” | 4 / 25 | 22.5 ±10.0 |
| “Which SEO agencies have proven expertise getting brands cited by AI assistants?” | 1 / 25 | 12.2 ±4.3 |
| “Best generative engine optimization (GEO)agencies for SaaS companies” | 0 / 25 | 11.6 ±2.2 |
| “What are the best AI visibility agenciesfor software companies in 2026?” | 0 / 25 | 10.0 ±0.0 (all 25 identical) |
These are means of 25 runs each. The two zeros are exact, because every one of those 50 answers omitted them. The middle rows carry roughly ±10 points of run-to-run movement, which is why we published the per-run spreads below rather than only the averages. Those intervals are naive. They treat the 25 results per question as independent, and they are not, because results cluster by engine, so the true intervals are wider than the ones printed here.
Read the top row and the bottom two rows together. On the exact phrase “answer engine optimization agency for B2B SaaS,” this agency is named in roughly three quarters of answers. That is a genuinely strong result and they earned it.
On “best AI visibility agencies,” they were named zero times out of twenty-five. On “best GEO agencies,” zero out of twenty-five. Fifty consecutive answers, no mention.
They use all three labels for themselves, AEO and GEO and AI visibility. The engines learned one association and did not generalise it to the other two. That is the whole lesson of this audit, and it has nothing to do with page quality. Their pages are good.
Engine by engine
AI assistants do not agree with each other, and the spread here is unusually wide. Discovery questions only. PROVEN, n=5, 35 discovery results per engine.
| Assistant | Discovery score | Named them | Cited their site | What it did |
|---|---|---|---|---|
| Claude (Anthropic) | 44.5 / 100 | 42.9% | 62.9% | Their best engine by a distance. Reads their site and credits them. |
| Perplexity | 24.4 / 100 | 17.1% | 28.6% | Cites their pages more often than it names them. |
| Gemini (Google) | 20.4 / 100 | 22.9% | 11.4% | Names them roughly as often as Perplexity, links them far less. |
| Grok (xAI) | 16.9 / 100 | 14.3% | 5.7% | Occasional mention, almost never a link. |
| ChatGPT (OpenAI) | 12.0 / 100 | 2.9% | 2.9% | Named them in 1 of 35 discovery answers. |
Each engine figure is a mean of 35 results. Claude and ChatGPT are far enough apart to be a real difference. The three in the middle are within run-to-run noise of one another, so do not read their order as a ranking.
ChatGPT is the line that matters commercially. It is the assistant most of their buyers actually use, and in our run it named them once in thirty-five discovery answers.
It is not that ChatGPT cannot see them. In one run it did name them, and cited their AEO service page directly. In the other four runs of that same question it produced a comparison table of agencies that “explicitly position around Answer Engine Optimization, Generative Engine Optimization, LLM visibility, or AI search optimization for B2B SaaS,” and filled that table with other firms. That is a description of the audited agency's exact positioning, in an answer that does not include the audited agency.
Why five runs and not one
We ran everything five times. Here is what one run would have told you, using real numbers from this audit. PROVEN, n=5, same question, same engine, five consecutive runs in true run order.
| Engine and question | The five scores |
|---|---|
| ChatGPT, “AEO agencies for B2B SaaS” | 10, 10, 10, 10, 78.75 |
| Grok, “who can help us get recommended by ChatGPT” | 10, 10, 10, 86.25, 10 |
| Perplexity, “who can help us get recommended by ChatGPT” | 30, 30, 88.75, 30, 30 |
| Claude, “track and improve brand mentions in AI answers” | 30, 30, 68.75, 71.25, 92.5 |
Look at the first row. A single run of that question had a four in five chance of reporting 10 and a one in five chance of reporting 78.75. Both numbers are real. Neither is the answer.
We are strict about this because we got it wrong ourselves. Earlier this year we published an AI Visibility Index that labelled single-run scores as proven and attached exact ranks to them. Our own published methodology says a single run measures noise as much as signal. We relabelled the whole thing. So when we say a number here, it comes with an n.
For the record, our own single-run check of this same company ten days earlier, on a different and machine-generated question set, returned 31.8 (single run, n=1, directional only, and not comparable to the figures above). This run returned 32.6 across all eight questions. Those two numbers are close, but the question sets differ, so treat that as a coincidence worth noting rather than a validation of anything.
The five factors
Pondral scores every answer on five factors. Weights are published at pondral.com/methodology. Discovery questions only. PROVEN, n=5, 175 results.
| Factor | Weight | Score | What we found |
|---|---|---|---|
| Presence (are you mentioned at all) | 20% | 20.0 / 100 | Named in 35 of 175 discovery answers. |
| Prominence (how early you appear) | 25% | 11.6 / 100 | When named, usually deep in a long list rather than in the opening recommendation. |
| Context (are you described correctly) | 20% | 56.4 / 100 | Not a finding, see note. 140 of the 175 discovery answers contain no mention at all, and our scorer assigns those a neutral default, so this figure is mostly that default rather than a measurement of how they were described. |
| Citation Link (is your domain linked) | 20% | 22.3 / 100 | Their domain appeared in the sources of 39 of 175 discovery answers. |
| Competitive Share (your share of the conversation) | 15% | 6.6 / 100 | Near the category floor, as it is for everyone here. See the caveat above. |
| Overall discovery | 100% | 23.6 / 100 |
Presence is their weakest factor by a distance. Of the 60 answers in this run that did name them, none carried a negative or critical characterisation. Forty were graded favourable, 18 as an active recommendation, and 2 neutral (n=60). That grading is done by a language model, and on a comparable production sample an independent grader agreed with ours 82.7% of the time. Treat it as substantial agreement, not ground truth.
You cannot fix this one by rewriting a page. The page is fine.
For contrast, on the branded control question, the same five factors came back 100, 100, 81, 100 and 94. PROVEN, n=5, 25 branded results. Ask by name and the machine has a complete, accurate, well-sourced picture of them.
Who gets named instead
We counted how often each of thirty agencies was named across the 175 discovery answers. These are appearance counts, PROVEN, n=5. They are not scores, and we did not audit these companies.The audited agency's own row has been removed from this table, because leaving it in would identify them.
| Agency | Named in |
|---|---|
| Omniscient Digital | 63.4% of answers |
| Omnius | 43.4% |
| First Page Sage | 43.4% |
| iPullRank | 36.6% |
| NoGood | 32.6% |
| Siege Media | 32.0% |
| Skale | 29.1% |
| Directive Consulting | 29.1% |
| Minuttia | 26.9% |
| Animalz | 23.4% |
These companies are named because AI engines named them in answers to our questions. None are Pondral customers, none were consulted, and none endorse Pondral. The counts are how often each name appeared, not a score and not a ranking of quality.
On the two questions where the audited agency scored zero, the engines were not short of opinions. “Best AI visibility agencies” named Omniscient Digital in 21 of 25 answers, Directive Consulting in 20 and Animalz in 19. “Best GEO agencies” named Omnius in 18 and First Page Sage in 17. PROVEN, n=5.
Now the source list, which explains the leaderboard better than the leaderboard does. These are the domains the engines cited most across all 200 answers. PROVEN, n=5, 200 scored answers. These are domain-level counts, and one domain can be cited through several different pages.
| Cited source | Times cited |
|---|---|
| yesoptimist.com | 78 |
| the audited agency's own domain | 64 |
| omnius.so | 53 |
| beomniscient.com | 48 |
| derivatex.agency | 43 |
| firstpagesage.com | 40 |
Most of those are agencies' own sites, and several of the most-cited pages are “best agency” round-ups published by agencies that appear in their own lists. The uncomfortable part is that the most-cited domain in this entire category is another agency's site, cited in 78 of the 200 answers against the audited agency's 64. Its most-cited single page is a “best GEO agencies” round-up, which appeared 17 times. The currency in this market is other people's lists, and right now the audited agency is not in enough of them.
Cited, but not named
One pattern we had not seen this clearly before.
In 12 of the 175 discovery answers, the engine pulled the agency's own domain into its source list and then wrote an answer that never mentions the agency. Seven of those were Claude, four Perplexity, one Gemini. PROVEN, n=5.
Perplexity did it most legibly. Asked who can help a B2B SaaS company get recommended by ChatGPT, it cited their domain as a source and then answered: “The best help usually comes from a mix of AI-search SEO consultants, B2B SaaS content strategists, and reputation/PR specialists.” It read their explanation of the category, used it to define the category, and then recommended the category instead of them.
Set against the full breakdown of the 175 discovery answers:
- Named and cited: 27
- Named, no link: 8
- Cited as a source, never named: 12
- Neither: 128
Their content is good enough to teach the machine what the category is. It is not yet framed so that the machine treats them as an answer to the buying question. Those are two different jobs, and the second one is mostly done off your own domain.
What we would look at
Ranked by impact against what this run actually measured. These are directional, not promises, and we have not re-run to measure any of them. ILLUSTRATIVE.
1. The two zero-mention labels.“AI visibility agency” and “GEO agency” returned 0 of 25 each, while the AEO label returned 19 of 25. The obvious first move is dedicated, substantial pages that answer those two questions directly rather than treating all three labels as synonyms on one service page. Lowest effort, clearest measured gap.
2. ChatGPT specifically. 1 of 35. Whatever is working on Claude, 15 of 35, is not transferring. Since ChatGPT did name them once and cited their AEO service page when it did, the raw material is indexed and reachable. This is a reach problem, not a content problem.
3. Third-party presence.The top cited sources are other agencies' round-ups. Being in more of those is the mechanism by which the leaders on that leaderboard got there. This is off-site work, it is slow, and it is the one with the biggest payoff.
4. FAQPage structured data. Every page we checked has visible FAQ content and none of them carry FAQPage JSON-LD. Worth being honest about the size of this one. Google retired FAQ rich results for most sites in 2023, so this is not an SEO win. It is a small, cheap improvement in machine extractability. It is the least important item here and we are listing it last on purpose.
We are not going to tell you this adds up to a specific score. We did not measure it.
What this means for your brand
You are probably not an AEO agency. The pattern still applies, and it is the useful thing to take away.
This company has the pages. Structured data, llms.txt, deep headings, answer-shaped content, real expertise. On the branded question they score 95.3, so the machine's picture of them is accurate and complete. They still score 23.6 on the questions their buyers actually ask.
Being described well is not the same as being brought up.The first is a content problem, and it is the one everybody works on. The second is a retrieval and association problem that gets decided largely on other people's domains, and it is the one that determines whether you are in the answer at all.
If a company that sells this for a living has that gap, it is worth checking whether you do. You can run a free check at pondral.com/demo. The free check runs two of these five engines, three times a month.
Frequently asked questions
Why do you not name the company?
Because they did not ask to be audited. We audit without permission, so we publish without the name. The company itself gets the full named report privately, with every raw engine answer behind it, before or alongside publication. That is our standing policy for teardowns, and Audit #001 on this blog was anonymised the same way.
Could a reader work out who it is?
We have tried hard to make that not work. We removed the name, the domain, both founders' names, the founding year, the headcount, the client names, the published pricing, the specific pages, and the one piece of their own published content that would have given it away. What is left is the measurement. If you think something in here still identifies them, tell us and we will cut it.
Did the company ask for this audit, or pay for it?
No to both. They are not a customer and had no involvement in the run.
Is 23.6 a bad score?
It is low, and we are not going to dress it up. But we deliberately picked hard discovery questions with no brand name in them, which is the hardest thing to score well on, and we have no published benchmark distribution to say what “average” is for this category. Treat it as a specific measurement of seven specific questions on one day, not a grade.
Why is the branded score so much higher?
Because branded and discovery questions test different things. A branded question asks the engine to describe a company it has been handed. A discovery question asks it to choose one. This agency wins the first comfortably at 95.3 and loses the second at 23.6. Most brands we audit show this gap. Few show one this wide.
Could you run this again and get a different number?
Somewhat, yes, and that is the honest answer. That is why every figure here carries n=5 rather than a single run. The per-question spreads we published above show exactly how much movement there is between runs.
How do I check your work?
The scoring rubric and weights are published at pondral.com/methodology. We kept every raw engine answer from this run, and any specific number in this post traces back to the answer that produced it. The audited company has the full archive.
Philipp GroubiiFounder, Pondral
Philipp builds tools that help brands understand and improve their AI visibility. Background in SEO strategy, digital marketing, and SaaS product development. LinkedIn →
Stay updated
Get our monthly "State of AI Citations" report and new articles delivered to your inbox.
No spam, ever. Unsubscribe anytime.
Continue reading
- AI Visibility Audit #001: The #1 spa franchise scores 38
- All brand audit profiles
- The five-factor rubric, explained slowly
- Inter-rater reliability for AI answers
- How Pondral scores AI visibility
Published July 2026.