How Pondral measures AI Visibility.
AI Visibility is whether AI assistants name and recommend your brand when buyers ask them for options. Pondral measures it by asking ChatGPT, Claude, Gemini, Perplexity, and Grok a fixed set of buyer questions, storing every answer, and grading each answer on five weighted factors: Presence (20%), Prominence (25%), Context (20%), Citation Link (20%), and Competitive Presence (15%). Your score is the average of those graded answers. This page is the full method, versioned, with the raw evidence rules.
The five factors · Engines and models · Question panels · Grounding · Repeated sampling · Variability · How to cite · What changed
The five factors
Every stored answer is graded on the five factors below. Each factor produces a 0 to 100 sub-score; the weighted sum is that answer's score. A run's headline score is the mean of every graded answer in the run (one per question per engine on the self-serve plans). The weights are published and have not changed since version 2.0.0; the changelog records every change to how a factor is graded.
Presence (20%)
What it grades: Whether the brand appears in the response at all.
How it is scored: Binary. The brand, or one of the alternate names you register for it, is found in the answer text, or it is not. Alternate names matter: an engine that writes a possessive or a product nickname still counts, and an engine that names a different company with a similar name does not.
Why it is weighted 20%: Presence is the gate. A brand that is absent scores 0 on every other factor for that answer, so the other four factors already carry most of the penalty for absence. Presence itself is a moderate weight because being named is necessary but not sufficient.
Example: On the question "How to track brand mentions in ChatGPT", ChatGPT named Pondral and Perplexity did not (Pondral-on-Pondral audit, September 21, 2026). Presence is 100 on the first answer and 0 on the second.
Prominence (25%)
What it grades: How early in the response the brand appears.
How it is scored: The position of the first mention as a fraction of the response length. A brand named in the first sentence scores near 100; a brand named in the last paragraph of a long answer scores near 0. A brand that is absent scores 0.
Why it is weighted 25%: Buyers act on the first names an assistant gives them. Being named in the tenth item of a list is a different commercial outcome from being named first, and the score has to reflect that. It is the heaviest factor for that reason.
Example: "The main options are Profound, Pondral, and Peec" scores higher on Prominence for Profound than for Peec, and the gap grows with the length of the answer that follows.
Context (20%)
What it grades: How well the brand is described — portrayed favourably, described accurately, and placed in the right category for the question.
How it is scored: A grading model reads the passage where the brand is mentioned and rates three things: portrayal (recommended, positive, neutral, negative, or criticised), factual accuracy (clean, stale, wrong, or damaging), and category fit (correct, adjacent, or wrong). The three are combined by a fixed cap rule, so a glowing but factually wrong description cannot score above a neutral accurate one. A brand that is absent scores 0 on Context; there is nothing to describe. This is the one factor graded by a model rather than by a rule, and its agreement with an independent grader is measured and published below.
Why it is weighted 20%: An assistant can name you and still send buyers elsewhere by calling you the wrong kind of tool or quoting a price you retired a year ago.
Example: "Pondral is an AI Visibility platform that checks five engines weekly and publishes its scoring rubric" is recommended, clean, and correct. "Pondral is a social listening tool" is neutral, stale, and in the wrong category.
Citation Link (20%)
What it grades: Whether the response cites a link to your own domain.
How it is scored: Binary. The answer's cited sources include a URL on a domain you own, or they do not. A mention with no link scores 0 here even if the brand is described well. Only the domains you register count; a link to your G2 page or a news article about you does not.
Why it is weighted 20%: A link is the difference between being mentioned and being the source. An answer that links your site sends the buyer to you and tells the next model that your site is where the facts live.
Example: Perplexity is measured through its Agent API since version 2.7.0, which returns the pages the engine retrieved rather than only the citations shown in the answer. That retrieval set is broader, so a match on this factor is somewhat more likely on Perplexity than on the other four engines. The difference is named in the changelog, and Perplexity readings before and after September 21, 2026 carry different version stamps and are never averaged together.
Competitive Presence (15%)
What it grades: Your share of the brand mentions in the same response.
How it is scored: Every brand named in the answer is counted; your share of those names is the score. Named alone, 100. Named alongside three competitors equally, 25. Absent while competitors are named, 0. An answer that names no brand at all, yours or anyone's, also scores 0 (since version 2.5.0; before that it scored a neutral 50).
Why it is weighted 15%: It is partly a property of the question. Some questions invite a list of ten tools and some invite one answer. The factor still matters because a buyer given five names has a one-in-five chance of picking you, which is why it is the lightest weight.
Example: An answer naming Pondral, Profound, Peec, and Scrunch once each gives every one of them 25 on this factor. The same answer naming only Pondral gives Pondral 100.
Engines and models
Every paid audit runs every question on all five engines. The free check runs two engines per check. Models below are the ones in the scoring path as of September 22, 2026, read from the adapters, not from a vendor page.
| Engine | Model in the scoring path | Grounding |
|---|---|---|
| ChatGPT | GPT-5.5 | Web search forced on; an answer that skipped search is discarded and re-run |
| Claude | Claude Sonnet 4.6 | Web search on for every call |
| Gemini | Gemini 2.5 Flash | Google Search grounding on; an ungrounded answer is discarded |
| Perplexity | sonar, through Perplexity's Agent API (since version 2.7.0, September 21, 2026; sonar-pro through Chat Completions before that) | Live retrieval, always on |
| Grok | Grok 3 on self-serve monitoring audits; Grok 4.3 on consulting audits | Live search on |
The two Grok versions are a lane difference, not an inconsistency we are hiding: the self-serve monitoring lane and the consulting audit lane are separate code paths and were pinned at different times. We record the model behind every stored answer so the two are never compared as one instrument. When a vendor changes or retires a model, Pondral records the change as a methodology version, stamps every stored answer with the version it was scored under, and never averages readings across the change. The full list is at /methodology/changelog.
Question panels (fixed, versioned)
Scores are only comparable if the questions do not move. Each brand carries a fixed Question Panel: a versioned, customer-approved set of buyer questions, built in the Panel builder or added one at a time, drawn from four templates: category entry ("best tools for X"), problem first ("how do I fix X"), comparison ("X vs Y"), and brand adjacent ("is X worth it"). At least 40% are discovery questions with no brand name in the prompt, so the score reflects where new customers originate, not just branded lookups.
- Propose: Pondral seeds a draft panel from your vertical template.
- Review: You edit wording, add questions, or remove outliers in the Panel builder.
- Freeze: You lock the panel (e.g.
2026-Q2-v1). A content hash is stored; the panel is immutable until you explicitly refresh. - Run: Every audit stamps
panel_idandpanel_content_hashon the run record for comparability.
How many questions a brand carries depends on the plan: 10 on Free and SMB, 50 on Growth, and 100 on Agency. Every weekly audit re-runs the same set on all five engines. Consulting audits use larger versioned panels that are proposed, reviewed, then frozen. The free check uses a fixed 12-question mini panel from the same template taxonomy, scored once per engine: a one-time snapshot, labeled directional only. Not for budget decisions.
Grounding
An assistant answering from memory and an assistant answering with live web results are different instruments. Pondral turns web grounding on for every engine and, where the vendor lets the caller check, discards any answer that skipped it (a "grounding tripwire": ChatGPT and Gemini today). This is why a Pondral score can differ from what you see typing the same question into a chat window with search off.
Repeated sampling
AI engines are not deterministic: the same question can return a different answer a minute later. Pondral built an instrument that asks each question of each engine several times and reports a margin of error alongside the score. When it runs, it reports three metrics: Mention Rate (how often your brand appears across the sampled responses), Quality When Mentioned (how favorably and prominently your brand is represented when it does appear), and a combined Visibility Index.
Where it stands today. It is switched on by default for consulting audits since version 2.6.0 (September 21, 2026) and not switched on for self-serve monitoring audits, which grade one response per question and engine. Turning it on for self-serve would change how those scores are produced, so it would be logged in the methodology changelog before it ran. On consulting audits the defaults are three repeats on SMB-scope engagements and five on Growth and Agency scope, with Claude and Grok capped at two repeats because their answers are the most expensive to collect. The score on those runs is the mean of the per-question, per-engine averages, so an engine sampled five times carries the same weight as one sampled twice. The self-serve lane's run-to-run stability was measured on July 31, 2026 (version 2.1.4, below) and did not need repeats to be stable.
Ask one engine "what is the best tool for [your category]?" five times. Your brand appears in three of the five answers, a Mention Rate of 60%.
In those three answers, the five-factor rubric scores the mention 48, 62, and 55 out of 100. The middle value, 55, is your Quality When Mentioned (we use the middle value, not the average, so one outlier can't swing it).
Visibility Index = 60% × 55 = 33 out of 100.
For comparison: a brand mentioned in every answer at that same quality would score 55; one mentioned just once would score 11. A high Visibility Index requires both showing up consistently and being represented well when you do.
Variability, and what a score can and cannot tell you
Three published measurements describe how much a Pondral score moves on its own. Each carries its label; none is a confidence interval.
- Same audit run twice, same day (version 2.1.4, July 31, 2026; PROXY): 10 questions on all five engines, twice. The overall score moved 0.59 points; the mean per-answer difference was 1.86 points; 93.9% of answers stayed within 5 points and 91.8% kept the same grade band. Read: the headline is stable; individual answers wobble.
- Grader agreement on the Context factor (measured August 2026 on the grader now in production, methodology 2.3.0, n=150 randomly drawn real engine answers; PROXY): against a grader from a different AI provider, agreement is 87.3% (Cohen's kappa 0.624). Against a larger grader from the same provider, agreement is 72.0% (kappa 0.402). We publish both because they disagree, and which one should govern is not settled. The previous grader measured 0.645 same-provider and 0.599 cross-provider. An earlier published figure of 0.947 was withdrawn on July 27, 2026 because it was measured on sentences we wrote ourselves; the record is at /blog/methodology-integrity.
- Alert threshold: the change alerts on Growth and Agency plans (Daily change alerts (2 engines, first 2 questions)) fire only on a drop of 15 points or more on a question, because below that engine variance dominates.
What this means for reading a score: a change of a few points between two weekly runs is noise. A change of 15 or more on a named question, or a sustained move over three runs, is signal. A single-run score for any brand is a snapshot, not a ranking.
Every audit ships with a provenance ledger: the exact prompts, timestamps, model versions, and raw responses. Available for download on every paid plan from your reports page. Drift monitoring publishes per-engine changes on the methodology status page.
How to cite this page
Pondral. "How Pondral measures AI Visibility: the five-factor rubric," methodology version 2.7.0, updated September 22, 2026. https://pondral.com/methodology
Each factor has a stable anchor: /methodology#presence, #prominence, #context, #citation-link, #competitive-presence. Terms are defined at /blog/aeo-glossary and the version history is at /methodology/changelog.
How this compares to Semrush AI Visibility
Semrush is the SEO platform that recently added an AI Visibility feature. Pondral is the AI Visibility platform. That is the whole product. Both use the term "AI Visibility." They are different companies, and in 12 months they'll be solving different problems. The side-by-side comparison lives at /compare/semrush.
What changed
- 2.7.0, September 21, 2026: Perplexity moved to the Agent API and the sonar model before the Chat Completions interface shut down; Citation Link on Perplexity now matches against a broader retrieval set.
- 2.6.0, September 21, 2026: repeated sampling on by default for consulting audits, with a published margin of error; self-serve lane unchanged.
- 2.5.0, August 4, 2026: an answer naming no brand at all scores 0 on Competitive Presence, not a neutral 50.
- 2.3.0, July 31, 2026: the Context grader now rates accuracy and category fit, not portrayal alone; agreement re-measured and published.
- 2.1.0, July 27, 2026: a brand absent from an answer scores 0 on Context (a floor of 10 had been leaking in).
- Full history with reasons: /methodology/changelog.
Methodology v2.7.0 (September 22, 2026). Every figure on this page is read from stored answers or from a published measurement with its label. Generative engines are non-deterministic; scores are directional.