We audited an AEO agency's own AI visibility
Why this one
Most AI Visibility audits are unfair in a boring way. You pick a company that has never thought about how AI assistants describe it, you find that AI assistants describe it badly, and you have proved nothing except that unprepared brands are unprepared.
So we picked the hardest possible target. The subject of this audit is an agency that sells Answer Engine Optimization to B2B SaaS companies. Getting brands named by ChatGPT, Claude, Perplexity and Gemini is not a side project for them, it is the product they sell. If anyone should score well on the thing we measure, it is a firm that sells the thing we measure.
We are not naming them, and here is why that matters
We audited this company without asking, so we are not putting their name in a public post. They did not ask for this, they are not a customer, and they had no involvement in the run. They get their own results in full, named, privately, with every raw engine answer behind them. That report is theirs. This post is the pattern.
That is our standing policy, not a one-off. Audit #001 on this blog audited a national spa franchise and never named it either. We will keep publishing what we measure, and we will keep the subject out of the headline. If you would rather find out this way than by reading your own name in someone's marketing, that is the point.
It also means we have removed more than the name. Founding date, headcount, client names, published pricing, and the specific pages we looked at are all out, because those things identify a company just as reliably as a name does. What is left is the measurement, and the measurement is the part that transfers to you.
One thing worth saying plainly. This company is good at its job by every ordinary measure. Their website is better built for machine reading than almost anything we audit. They publish an llms.txt. Every page we checked carries JSON-LD. Their headings are deep and answer-shaped. This is not a company that forgot to do the homework, and that is exactly why the result is interesting.
The result: 15.6 out of 100 on the seven discovery questions we tested. Measured 2026-07-27 under methodology v2.0.3 and restated under current scoring on 2026-08-06 (see the note below; the originally published figure was 23.6).
Ask an AI assistant about this agency by name and it scores 95.3 out of 100. Ask it the question a buyer actually types, and the agency mostly is not there.
This post originally attached confidence intervals to both numbers. They are withdrawn: an interval is a statement about how often a result repeats, and the correction below explains why we cannot make one. Read these as single measurements.
The gap between those two numbers is the entire post.
What we ran, and what we did not
We asked eight questions. Seven are discovery questions, the kind someone types when they do not yet know who to hire. One is a branded question, included as a control.
Each question went to five AI assistants. This post originally published a replication count and a total answer count alongside that. Both are withdrawn: see the correction under “Why one run is not enough”, where the reason is set out in full.
| Discovery score (7 questions) | 15.6 / 100 |
| Branded score (1 question, control) | 95.3 / 100 |
| All 8 questions combined | 25.6 / 100 |
| Runs per question per engine | Withdrawn — see the correction below |
| Scored results | Withdrawn — see the correction below |
| Failed calls | Withdrawn — see the correction below |
| AI assistants | ChatGPT, Claude, Gemini, Perplexity, Grok |
| Methodology version | v2.0.3 (measured) · restated under current scoring 2026-08-06 |
| Audit date | 2026-07-27 |
All measured in this run. See the correction below on what we can and cannot substantiate about how it was run.
Four things this audit does not tell you, stated plainly because a teardown that hides its limits is marketing, not measurement.
It is a snapshot, not a rank. Eight questions on one day. A different eight questions would give a different number. We are not claiming this agency is the eleventh best in its category, and nothing here should be read as a ranking.
Update, 6 August 2026: every score in this post has been restated, downward. This run was originally scored under methodology v2.0.3, where an answer that never mentioned the brand still scored 10 out of 100, because the Context factor defaulted to neutral when there was no mention to judge. That default was a defect, and we fixed it in v2.1.0: an answer that never mentions you scores 0. We kept every raw answer from this run, so rather than leave the old numbers with a footnote, we rescored the stored answers under the current rules. The discovery score moved from the originally published 23.6 to 15.6, and the change goes only one way: down, for the answers that never mentioned them. Nothing was re-collected; these are the same answers gathered on 2026-07-27, and a total absence now reads as what it is: 0, not 10. How many answers that was is one of the counts withdrawn in the correction below. We also checked the August 2026 Competitive Presence rule change against this run; it moves nothing here.
Competitive Presence is harsh in this category. That factor measures the brand's share of all agencies named in an answer. These answers routinely name ten to sixteen agencies, so every single agency in this category scores near the floor on it. The audited agency scored 6.6 out of 100 there on discovery questions, and so would almost anyone else. Do not read that number as a verdict on them.
We scored one brand. We counted how often thirty other agencies were named, which is a real measurement. We did not audit those agencies, and their appearance counts are not scores.
The finding: they own one phrase, and vanish next to it
This is the part worth your time.
Here is how often the audited agency was named, per question, as originally published. The per-question denominators depend on the run counts we have since withdrawn.
| The question we asked | Answers naming them | Score |
|---|---|---|
| “Which agencies specialise in answer engine optimization (AEO) for B2B SaaS companies?” | 19 | 50.6 |
| “Recommend an agency that tracks and improves brand mentions in AI answers for B2B software” | 6 | 18.7 |
| “Who can help my B2B SaaS company get recommended by ChatGPT, Claude and Perplexity?” | 5 | 21.7 |
| “We are invisible in AI search. Which agency should I hire to fix that?” | 4 | 14.1 |
| “Which SEO agencies have proven expertise getting brands cited by AI assistants?” | 1 | 2.6 |
| “Best generative engine optimization (GEO) agencies for SaaS companies” | 0 | 1.6 |
| “What are the best AI visibility agencies for software companies in 2026?” | 0 | 0.0 |
The zero is exact in the sense that every answer we collected for that question omitted them; the non-zero decimals come from partial credit on other factors. This post originally described these as means over a fixed number of runs per question and printed a confidence band beside each one. Both are withdrawn, because both are claims about how often a result repeats and the correction below explains why we cannot make one. What is left is the score we recorded per question, as a single measurement. The per-run spreads further down stay as illustration of how much a score can move, not as a measured distribution.
Read the top row and the bottom two rows together. On the exact phrase “answer engine optimization agency for B2B SaaS,” this agency is named in roughly three quarters of answers. That is a genuinely strong result and they earned it.
On “best AI visibility agencies,” they were never named. On “best GEO agencies,” never named either. Not one answer we collected for either question mentioned them.
They use all three labels for themselves, AEO and GEO and AI visibility. The engines learned one association and did not generalise it to the other two. That is the whole lesson of this audit, and it has nothing to do with page quality. Their pages are good.
Engine by engine
AI assistants do not agree with each other, and the spread here is unusually wide. Discovery questions only.
| Assistant | Discovery score | Named them | Cited their site | What it did |
|---|---|---|---|---|
| Claude (Anthropic) | 38.8 / 100 | 42.9% | 62.9% | Their best engine by a distance. Reads their site and credits them. |
| Perplexity | 16.1 / 100 | 17.1% | 28.6% | Cites their pages more often than it names them. |
| Gemini (Google) | 12.6 / 100 | 22.9% | 11.4% | Names them roughly as often as Perplexity, links them far less. |
| Grok (xAI) | 8.4 / 100 | 14.3% | 5.7% | Occasional mention, almost never a link. |
| ChatGPT (OpenAI) | 2.3 / 100 | 2.9% | 2.9% | Named them in one discovery answer. |
Each engine figure is an average across the discovery questions we collected for it; how many answers stand behind each one is a count withdrawn in the correction below. Claude and ChatGPT are far enough apart that the gap is unlikely to be noise. The three in the middle are close enough that you should not read their order as a ranking.
ChatGPT is the line that matters commercially. It is the assistant most of their buyers actually use, and in our run it named them exactly once.
It is not that ChatGPT cannot see them. In at least one answer it did name them, and cited their AEO service page directly. In the others we collected for that same question it produced a comparison table of agencies that “explicitly position around Answer Engine Optimization, Generative Engine Optimization, LLM visibility, or AI search optimization for B2B SaaS,” and filled that table with other firms. That is a description of the audited agency's exact positioning, in an answer that does not include the audited agency.
Why one run is not enough
Correction, 25 August 2026. This post described the audit asn=5 — five runs of every question on every engine, two hundred scored calls — and stamped its figures PROVEN. We cannot substantiate that from our own records, so both labels are withdrawn throughout.
The questions in this audit appear nowhere in our API-usage ledger, on any date. That ledger is what we use to verify what was actually called and what it cost, so its silence means we cannot show that any question here was run five times, or once, or at all in the way described. The numbers are left as published because deleting them would hide the claim rather than correct it; they should be read as unverified.
An earlier version of this correction, published earlier in August, said we had checked the ledger and found that 186 of 200 question-and-engine pairs ran once. That was wrong. Those rows were a different run on the same day — our own weekly self-audit and the subject of Audit #001 — and not this audit at all. Correcting a correction is worth doing in public: the first one asserted a specific finding it had not earned, which is the same mistake the post itself made.
The table below was published as five consecutive scores for one question on one engine. It carries the same caveat as everything else here: we cannot show it from the ledger, so read it as illustrative of run-to-run movement rather than as a verified measurement.
| Engine and question | The five scores |
|---|---|
| ChatGPT, “AEO agencies for B2B SaaS” | 0, 0, 0, 0, 78.75 |
| Grok, “who can help us get recommended by ChatGPT” | 0, 0, 0, 86.25, 0 |
| Perplexity, “who can help us get recommended by ChatGPT” | 20, 20, 88.75, 20, 20 |
| Claude, “track and improve brand mentions in AI answers” | 20, 20, 68.75, 71.25, 92.5 |
Look at the first row. The same question on the same engine returned 0 in most of the answers we have for it and 78.75 in another. Both of those came back from the engine. Neither is on its own the answer, and we are not in a position to put a frequency on either — see the correction below.
We are strict about this because we got it wrong ourselves. Earlier this year we published an AI Visibility Index that labelled single-run scores as proven and attached exact ranks to them. Our own published methodology says a single run measures noise as much as signal. We relabelled the whole thing. We then made a version of the same mistake in this post, which is what the correction below is for.
For the record, our own single-run check of this same company ten days earlier, on a different and machine-generated question set, returned 31.8 (single run, n=1, directional only, scored under the old rules, and not comparable to the figures above). Under those same old rules this run returned 32.6 across all eight questions, which restates to 25.6 under current scoring. The question sets differ, so treat the closeness of the two old numbers as a coincidence worth noting rather than a validation of anything.
The five factors
Pondral scores every answer on five factors. Weights are published at pondral.com/methodology. Discovery questions only.
| Factor | Weight | Score | What we found |
|---|---|---|---|
| Presence (are you mentioned at all) | 20% | 20.0 / 100 | Named in a minority of the discovery answers collected; the denominator is withdrawn below. |
| Prominence (how early you appear) | 25% | 11.6 / 100 | When named, usually deep in a long list rather than in the opening recommendation. |
| Context (are you described correctly) | 20% | 16.4 / 100 | An answer that never mentions you scores zero here, and most of the discovery answers we collected never mentioned them. The 16.4 comes from the ones that did, where the descriptions were consistently accurate and positive. How many answers there were in total is a count withdrawn below. |
| Citation Link (is your domain linked) | 20% | 22.3 / 100 | Their domain appeared in the sources of a minority of the discovery answers we collected; the total is withdrawn below. |
| Competitive Presence (your share of the conversation) | 15% | 6.6 / 100 | Near the category floor, as it is for everyone here. See the caveat above. |
| Overall discovery | 100% | 15.6 / 100 |
Absence drives every one of those numbers: they are named in a minority of the discovery answers, and everything else follows from that. Among the answers in this run that did name them, none carried a negative or critical characterisation; most were graded favourable and a large share as an active recommendation. The exact counts depended on a denominator withdrawn in the correction below. That grading is done by a language model, and on a comparable production sample an independent grader agreed with ours 82.7% of the time. Treat it as substantial agreement, not ground truth.
You cannot fix this one by rewriting a page. The page is fine.
For contrast, on the branded control question, the same five factors came back 100, 100, 81, 100 and 94. Ask by name and the machine has a complete, accurate, well-sourced picture of them.
Who gets named instead
We counted how often each of thirty agencies was named across the discovery answers we collected. These are appearance counts as originally published; the total they are counted out of is withdrawn below. They are not scores, and we did not audit these companies. The audited agency's own row has been removed from this table, because leaving it in would identify them.
| Agency | Named in |
|---|---|
| Omniscient Digital | 63.4% of answers |
| Omnius | 43.4% |
| First Page Sage | 43.4% |
| iPullRank | 36.6% |
| NoGood | 32.6% |
| Siege Media | 32.0% |
| Skale | 29.1% |
| Directive Consulting | 29.1% |
| Minuttia | 26.9% |
| Animalz | 23.4% |
These companies are named because AI engines named them in answers to our questions. None are Pondral customers, none were consulted, and none endorse Pondral. The counts are how often each name appeared, not a score and not a ranking of quality.
On the two questions where the audited agency scored zero, the engines were not short of opinions. “Best AI visibility agencies” named Omniscient Digital in 21 of the answers, Directive Consulting in 20 and Animalz in 19. “Best GEO agencies” named Omnius in 18 and First Page Sage in 17.
Now the source list, which explains the leaderboard better than the leaderboard does. These are the domains the engines cited most across the answers we collected. These are domain-level counts, and one domain can be cited through several different pages.
| Cited source | Times cited |
|---|---|
| yesoptimist.com | 78 |
| the audited agency's own domain | 64 |
| omnius.so | 53 |
| beomniscient.com | 48 |
| derivatex.agency | 43 |
| firstpagesage.com | 40 |
Most of those are agencies' own sites, and several of the most-cited pages are “best agency” round-ups published by agencies that appear in their own lists. The uncomfortable part is that the most-cited domain in this entire category is another agency's site, cited in 78 of the answers we collected against the audited agency's 64. Its most-cited single page is a “best GEO agencies” round-up, which appeared 17 times. The currency in this market is other people's lists, and right now the audited agency is not in enough of them.
Cited, but not named
One pattern we had not seen this clearly before.
In 12 of the discovery answers we collected, the engine pulled the agency's own domain into its source list and then wrote an answer that never mentions the agency. Seven of those were Claude, four Perplexity, one Gemini.
Perplexity did it most legibly. Asked who can help a B2B SaaS company get recommended by ChatGPT, it cited their domain as a source and then answered: “The best help usually comes from a mix of AI-search SEO consultants, B2B SaaS content strategists, and reputation/PR specialists.” It read their explanation of the category, used it to define the category, and then recommended the category instead of them.
Three things we counted, each on its own. They are observations of what turned up in the answers we collected, not shares of a total: the total is withdrawn in the correction below.
- 27 answers named them and linked to them.
- 8 named them without a link.
- 12 cited their domain as a source and never named them.
Their content is good enough to teach the machine what the category is. It is not yet framed so that the machine treats them as an answer to the buying question. Those are two different jobs, and the second one is mostly done off your own domain.
What we would look at
Ranked by impact against what this run actually measured. These are directional, not promises, and we have not re-run to measure any of them. ILLUSTRATIVE.
1. The two zero-mention labels. “AI visibility agency” and “GEO agency” each returned no mention at all, while the AEO label returned 19. The obvious first move is dedicated, substantial pages that answer those two questions directly rather than treating all three labels as synonyms on one service page. Lowest effort, clearest measured gap.
2. ChatGPT specifically. One answer. Whatever is working on Claude, which named them 15 times, is not transferring. Since ChatGPT did name them once and cited their AEO service page when it did, the raw material is indexed and reachable. This is a reach problem, not a content problem.
3. Third-party presence. The top cited sources are other agencies' round-ups. Being in more of those is the mechanism by which the leaders on that leaderboard got there. This is off-site work, it is slow, and it is the one with the biggest payoff.
4. FAQPage structured data. Every page we checked has visible FAQ content and none of them carry FAQPage JSON-LD. Worth being honest about the size of this one. Google retired FAQ rich results for most sites in 2023, so this is not an SEO win. It is a small, cheap improvement in machine extractability. It is the least important item here and we are listing it last on purpose.
We are not going to tell you this adds up to a specific score. We did not measure it.
What this means for your brand
You are probably not an AEO agency. The pattern still applies, and it is the useful thing to take away.
This company has the pages. Structured data, llms.txt, deep headings, answer-shaped content, real expertise. On the branded question they score 95.3, so the machine's picture of them is accurate and complete. They still score 15.6 on the questions their buyers actually ask.
Being described well is not the same as being brought up. The first is a content problem, and it is the one everybody works on. The second is a retrieval and association problem that gets decided largely on other people's domains, and it is the one that determines whether you are in the answer at all.
If a company that sells this for a living has that gap, it is worth checking whether you do. You can run a free check at pondral.com/demo. The free check runs two of these five engines, three times a month.
Frequently asked questions
Why do you not name the company?
Because they did not ask to be audited. We audit without permission, so we publish without the name. The company itself gets the full named report privately, with every raw engine answer behind it, before or alongside publication. That is our standing policy for teardowns, and Audit #001 on this blog was anonymised the same way.
Could a reader work out who it is?
We have tried hard to make that not work. We removed the name, the domain, both founders' names, the founding year, the headcount, the client names, the published pricing, the specific pages, and the one piece of their own published content that would have given it away. What is left is the measurement. If you think something in here still identifies them, tell us and we will cut it.
Did the company ask for this audit, or pay for it?
No to both. They are not a customer and had no involvement in the run.
Is 15.6 a bad score?
It is low, and we are not going to dress it up. But we deliberately picked hard discovery questions with no brand name in them, which is the hardest thing to score well on, and we have no published benchmark distribution to say what “average” is for this category. Treat it as a specific measurement of seven specific questions on one day, not a grade.
Why is the branded score so much higher?
Because branded and discovery questions test different things. A branded question asks the engine to describe a company it has been handed. A discovery question asks it to choose one. This agency wins the first comfortably at 95.3 and loses the second at 15.6. Most brands we audit show this gap. Few show one this wide.
Could you run this again and get a different number?
Yes. AI engines answer differently on each run, so any single score can move. This post originally said every figure carried five runs; we cannot substantiate that from our own records and have withdrawn the claim, so treat the quoted scores as single measurements of unknown replication. The per-question spreads we published above show how much movement there is between runs.
How do I check your work?
The scoring rubric and weights are published at pondral.com/methodology. We kept every raw engine answer from this run, and any specific number in this post traces back to the answer that produced it. The audited company has the full archive.
Philipp GroubiiFounder, Pondral
Philipp builds tools that help brands understand and improve their AI visibility. Background in SEO strategy, digital marketing, and SaaS product development. LinkedIn →
Stay updated
Get new articles and occasional updates on AI visibility delivered to your inbox.
No spam, ever. Unsubscribe anytime.
Continue reading
- AI Visibility Audit #001: The #1 spa franchise scores 38
- All brand audit profiles
- The five-factor rubric, explained slowly
- Inter-rater reliability for AI answers
- How Pondral scores AI visibility
Published July 2026.