Pondral
MethodologyAugust 20269 min read

Our own audit prompt was role-playing a finance analyst. Here is what it cost the brands we measure.

An AI Visibility audit measures whether the five big AI engines name your brand when someone asks them a buying question. Every Pondral audit sends those engines the same instruction before it asks anything about your brand. For most of this product's life that instruction told the model it was a senior financial analyst. We sell to dental practices, law firms, home services companies and B2B software teams. Almost none of them are finance.

Nobody put that persona there on purpose for those customers. It came from the earliest version of this tool, which was built for a finance client, and it survived every rewrite after because it sat in a constant nobody reread. That is the boring way most measurement bugs happen.

We paid to find out whether it mattered. It cost $29.72 and 324 scored engine calls, and the answer was more interesting than either the yes or the no I was expecting.

PROXY throughout. Every number in this post comes from one paired test run on 17 July 2026, under methodology version 2.0.3. Pondral is now on 2.5.0. These figures are already public on our methodology changelog under entry 2.0.11, where they were first disclosed. This post is amplification of that disclosure, not a new claim.

What we ran

The same 20 questions, across 7 industries, through all five engines, twice, under two different system prompts. One arm used the finance-analyst instruction exactly as production had it. The other used a plain vertical-neutral wording. Everything else was held identical, including the engine adapters, the parser, and the scoring code.

The arms were interleaved in time inside each engine so that a slow drift in an engine's behaviour could not land on one arm and not the other. Two of the 326 attempts came back as transient Gemini errors and were retried successfully, which is why the scored total is 324.

Two questions were deliberately about finance. Those are the control. If the finance persona is doing what its name suggests, those two should move in the opposite direction from everything else.

The average said nothing was wrong

Across all 81 paired comparisons the neutral prompt moved scores by an average of +1.04 points out of 100. The median was exactly zero. Meanwhile the natural run to run wobble of these engines, measured inside each arm, is about 8.3 points.

So the average effect was comfortably smaller than the noise. If we had looked only at the headline number we would have written “no material difference”, kept the finance persona, and moved on. That is the version of this post I expected to be writing.

The average was hiding the finding

Splitting by question type is where it shows up.

Question typePairsAverage changeLargest single change
Best provider (“best X for Y”)67-0.3625.00
Comparison (“X vs Y”)5-4.0012.50
Open-ended advisory (“how do teams choose X”)9+14.3163.75

On open-ended advisory questions the finance persona was suppressing non-finance brands by an average of 14 points, with one case at nearly 64. Seven of the nine comparisons moved positive and not one came out at zero.

The mechanism is legible once you read the answers. Asked how a buyer should choose a practice management system, the finance analyst answers like a finance analyst. It lays out an evaluation framework, talks about total cost of ownership and risk, and names very few actual products. The neutral analyst just lists the vendors. Naming fewer vendors means naming your brand less often, and presence is 20 percent of the score with prominence another 25.

We predicted the wrong place. The design memo said the damage would show up in “best provider” queries, because that is where brand names are most obviously at stake. Measured, that class came in at -0.36, which is nothing. The tail was somewhere we were not looking. I am keeping that in the post because a test that only ever confirms the memo that proposed it is not doing much work.

The control behaved

The two finance questions came out at -1.44 on average. Slightly negative, meaning the finance persona helped finance brands a little, which is exactly the signature you would want if the persona is doing what the label says. One private equity firm scored 78.75 under the finance prompt and 51.25 under the neutral one on the same question.

That is the part that convinced me. A one sided result on the treatment group is a finding. A one sided result on the treatment group plus the opposite sign on the control is a mechanism.

What we changed

The prompt now says senior research analyst and names no industry. It shipped as methodology version 2.0.11 on 17 July 2026, and it is recorded on our public methodology changelog along with these figures. Nothing else moved. Same engines, same five factor rubric, same weights, one constant swapped.

Nothing got worse. Zero grounding failures, zero parse failures, and run to run consistency came out very slightly tighter under the neutral prompt than under the finance one.

What these numbers are not

I would rather say this plainly than have someone work it out later.

These are labelled PROXY, not proven. This is one test, on one panel of questions we wrote ourselves, on one day.

The headline finding rests on nine comparisons. Three questions across three engines. Nine paired measurements is enough to notice something and act on it when the mechanism is coherent and the control confirms it. It is not enough to publish “the finance prompt costs you 14 points” as a general law, and I am not claiming that.

The advisory finding covers three of the five engines. Claude and Grok measured about $0.25 and $0.23 per call against the $0.072 blended average our budget was built on, because both re-inject web search results into the context. Running the full panel on all five would have gone over the budget cap, so those two ran a smaller stratified subset and the open-ended questions were measured on the other three. That is a real limitation and not a rounding detail.

The scoring has moved twice since. This ran under methodology 2.0.3. We are now on 2.5.0, and one of the changes in between matters here specifically: a brand that is never mentioned used to score 10 out of 100 and now scores 0. The table above is full of those floor values, so these exact numbers are not reproducible under today's scoring. The direction would almost certainly be larger now, not smaller, because the floor moved down. We have not re-run it.

The raw responses are gone. We kept the summary table and the run identifiers, not the individual engine answers. So you can check our arithmetic but you cannot independently re-derive it from source. That is a weaker evidence position than I would like and it is why the retention policy changed afterwards.

Why this is public

We had a measurement instrument with a defect that flattered one industry and quietly penalised the customers we actually sell to. Nobody reported it. No customer complained, because we do not have paying customers yet. We went looking because the constant read wrong when someone finally reread it.

If you are evaluating any tool in this category, the useful question is not whether its numbers look good. It is what the tool tells the model before it asks about you, whether the vendor will show you that text, and whether they have ever tested what changing it does. We have now done that once. It cost thirty dollars.

The run in numbers

Methodology version tested2.0.3 (current: 2.5.0)
Panel20 questions, 7 verticals, 2 finance controls
Paired comparisons81
Calls324 scored, 326 attempted
Spend$29.72
Enginesgpt-5.5-2026-04-23, claude-sonnet-4-6, gemini-2.5-flash, sonar-pro, grok-4.3
LabelPROXY throughout
First disclosedMethodology changelog entry 2.0.11, July 2026

Frequently asked questions

What was wrong with Pondral's audit prompt?

Every scored engine call carried a default system prompt telling the model it was a senior financial analyst. It came from the earliest version of the tool, which was built for a finance client, and no caller ever overrode it, so it conditioned every audit for every industry. It was replaced with a neutral research-analyst prompt on 17 July 2026 and recorded as methodology changelog entry 2.0.11.

How much did the finance prompt actually change scores?

Across 81 paired comparisons the neutral prompt moved scores by an average of +1.04 points out of 100, with a median of exactly zero. That is smaller than the run-to-run variation of the engines themselves, which averaged 8.3 points pooled across both arms. Split by question type, open-ended advisory questions moved +14.31 points on average across 9 paired comparisons, with a largest single change of +63.75. The two finance control questions moved -1.44 points. Every figure here is PROXY, from a single 324-call test run on 17 July 2026 under methodology version 2.0.3.

How many measurements is the 14-point finding based on?

Nine paired comparisons, from three questions measured across three engines. That is enough to notice a pattern and act on it when the mechanism is coherent and the control group moves the other way. It is not enough to publish as a general law, and Pondral is not claiming that a finance-flavoured prompt costs any given brand 14 points.

Why does the advisory finding cover only three of the five engines?

Claude and Grok measured about $0.25 and $0.23 per call against the $0.072 blended average the budget was built on, because both re-inject web search results into the model context. Running the full panel on all five engines would have gone over the approved budget, so those two ran a smaller stratified subset and the open-ended questions were measured on the other three engines. That is a real limitation of the test, not a rounding detail.

Are these numbers reproducible under Pondral's current scoring?

No, and that is stated deliberately. The test ran under methodology version 2.0.3. Pondral is now on 2.5.0, and one change in between matters here specifically: a brand that is never mentioned used to score 10 out of 100 and now scores 0. Several values in this test sit at that old floor. The test has not been re-run under current scoring.

Can I check this work?

The figures were first disclosed on Pondral's public methodology changelog under entry 2.0.11 in July 2026, and this post amplifies that disclosure rather than making a new claim. The raw engine responses from the run were not retained, only the summary table and the run identifiers, so the arithmetic can be checked but the result cannot be independently re-derived from source. The retention policy changed afterwards.


PG

Philipp GroubiiFounder, Pondral

Philipp builds tools that help brands understand and improve their AI visibility. Background in SEO strategy, digital marketing, and SaaS product development. LinkedIn →

Stay updated

Get new articles and occasional updates on AI visibility delivered to your inbox.

No spam, ever. Unsubscribe anytime.

Continue reading

Published August 2026.