Methodology Changelog
Scoring Methodology Version History
Every change to Pondral's scoring rubric, factor weights, reproducibility approach, and rater configuration is documented here. Reproducibility means knowing how the system works today and how it has changed.
For the current methodology, see the full methodology page. For the design rationale behind our scoring decisions, read How We Built Pondral's Scoring Rubric.
v2.5.0 (August 4, 2026)
- When an AI answer names no brand at all — not you, and not any competitor — the Competitive Presence factor now scores 0 instead of the neutral 50 it had been using. Competitive Presence measures your share of the brands named in an answer. If nobody is named, you have no presence in that answer, and scoring it as a neutral half-mark was quietly lifting the overall score of a result where your brand was completely absent.
- This is the same correction we made to the Context factor on July 28. A result where the AI never mentions you is zero visibility for you on that question, and every one of the five factors should reflect that. The neutral 50 was adding about 7.5 points to the score of answers that named nobody — small, but always in the flattering direction, which is exactly the kind of drift we commit to removing.
- Nothing changes for answers that do name brands. If you or a competitor is mentioned, Competitive Presence is still your share of those mentions, scored exactly as before. Only the 'nobody was named' case moves, and it moves down.
- This reverses a choice we made on July 31 (version 2.2.0). We had scored the no-one-named case as neutral on the reasoning that nobody had won or lost; we now think consistency about what a zero-mention answer means matters more than that distinction, which is diagnostic and still available in your competitor breakdown. Stored results are not rewritten — every stored answer carries the version it was scored under, and new runs are stamped 2.5.0.
v2.4.1 (August 2, 2026)
- Your dashboard now calculates your score exactly the way our published method says it should. Our published method takes the score for each individual AI answer and averages those. The dashboard was doing something subtly different: averaging each of the five factors across all your answers first, and then combining those five averages using the weights. Those two routes sound equivalent and are not, so the dashboard could show a number our own published method would not produce.
- Why they disagree. Averaging first and combining second would give the same answer if both routes read the same per-answer figures. They did not. The dashboard rebuilt four of the five factors from raw stored data rather than reading the graded figures, and its Context average deliberately skipped answers that had never been graded while the other four factors counted every answer. Different inputs, different totals. On a worked example an answer that scores 76 under our published method read 74 under the old route.
- What this means in practice: the score on your dashboard and the score in your audit report are now the same calculation. Our whole argument is that you can check our numbers against our published method, which only works if we are running that method.
- No score will visibly change when this ships. The per-answer grading that this reads only started being recorded in the previous release, and no stored answer carries one yet, so both the old and the new calculation currently show the same empty state. Your first graded audit will show the correct figure from the start rather than showing one number and then moving.
- The five factor bars underneath the score are unchanged, and they will not add up exactly to the headline. They are a breakdown, not the calculation. Four of them are still rebuilt from raw data and are approximations of what the grader measured precisely. We would rather tell you that than quietly adjust the bars to match. Making them read the graded figures too is on our list as its own change.
- Separately, and affecting no score: our internal cost tracking was billing the Context grader at our flagship model's price rather than the cheaper model it actually runs on, overstating that line by about three times. It has never affected anything you see or pay. We are noting it because we said we would close it.
- The scoring formula, the five factors and their weights are unchanged, which is why the methodology version stays at 2.4.0. We are recording this because the route to the number you see changed, and that is worth telling you about even when the formula did not.
v2.4.0 (August 2, 2026)
- The Context factor now recognises your brand's other names. If you tell us your brand also goes by an abbreviation or a shorter trading name, we already used that list when checking whether you were mentioned at all. The Context grader, which judges how you are described, did not: it looked only for your full official name. So an AI answer that talked about you using only your shorter name scored full marks for being mentioned and zero for how you were described.
- That zero was wrong, it was worth up to 20 points, and it always pushed scores down rather than up. Context carries 20% of the total, and a score of zero on it is what we record when a brand is genuinely not mentioned anywhere in the answer. Once stored, a wrongly-zeroed answer looked identical to one where the brand really was absent.
- The rule for genuine absence has not changed. If neither your name nor any of your other names appears in an answer, Context still scores 0, exactly as it has since our July 27 correction. What changed is which names we look for, not what happens when we find none. Both the main grader and the backup grader we fall back on if our grading service is unavailable were changed together, so they cannot disagree with each other about whether you were mentioned.
- No score changes as a result of this, and we checked rather than assumed. Every account that has a list of alternative names currently has exactly one entry in it: the brand's own full name. Under that setup the new search looks for precisely the same thing the old one did, so the fix cannot move any existing number. It matters the moment anyone adds a genuinely different name.
- We are still bumping the version even though nothing moves today, and here is the honest reason. Alternative names are something you enter, not something in our code. If we left the version alone, then the day someone adds a real alternative name the scoring rule would change with nothing to mark where. Every stored result carries the version it was scored under, and that stamp is the only thing that lets us tell two different rules apart later.
- If you later add a genuinely different name, the effect can only be upward and is capped: an answer that mentions you only by that name gets its real Context grade instead of a zero, worth at most 20 points on that one answer. No answer can score lower because of this change.
- Unchanged: the grader itself, the model behind it, its temperature setting, the five factors, their weights, and how per-answer scores are combined. Stored results are not rewritten. New runs are stamped 2.4.0.
v2.3.1 (August 2, 2026)
- Correction to how your dashboard displayed the Context factor. It was not measuring context. The dashboard scored Context by checking whether a descriptive text field on each stored answer contained anything at all. That field is filled in on essentially every answer, so Context scored 100 almost every time.
- What that did to your score: Context carries 20% of the total, so a permanent 100 added a flat 20 points to every dashboard score. On our own reference brand the dashboard read 39, of which 20 came from this. Scored as our published rubric requires, the honest figure was about 19. Dashboard scores were roughly double what they should have been.
- It also contradicted the correction we published on July 27. That change made Context score 0 when a brand is not mentioned in an answer at all, because our published rules say Context cannot outrank Presence. Nearly two thirds of the answers feeding the dashboard were ones where the brand was never mentioned, and the dashboard scored every one of them 100 on Context.
- The scored audit itself was never affected. This was only the monitoring dashboard. The audit pipeline has always stored a real graded score, and the AI Visibility Index and any audit report are computed from that, not from this.
- The fix: the monitoring pipeline now grades every answer with the same rubric and the same grader the audit uses, and stores the result. The dashboard reads that stored grade instead of trying to reconstruct one from raw data.
- Where no grade exists we now show nothing rather than a number. Answers recorded before this change were never graded, and there is no way to grade them after the fact. Rather than fill the gap with a guess, they are left out of the average, and a brand with no graded answers yet sees an empty state until its next audit runs. Filling that gap with an assumed value is what caused this defect in the first place.
- No stored score was rewritten and no history was restated. The scoring formula, the five factors and their weights are all unchanged, which is why the methodology version stays at 2.3.0. We are recording this anyway, because the number you see moved materially and that is worth telling you about even when the formula did not change.
v2.3.0 (July 31, 2026)
- The Context factor now measures three things instead of one. Until this release the Context grader judged only portrayal, meaning how favourably an AI answer talks about your brand. It now also judges factual accuracy (is what the answer says about you clean, stale, wrong, or damaging) and category fit (does the answer place you in the category the question actually asked about). All three are combined into the same 0/25/50/75/100 Context bucket.
- This closes a claims gap in the honest direction. On July 28 we withdrew the published description of Context, because we had been describing it as grading whether a brand is described accurately and in the right category while the grader that actually ran graded only tone. Rather than permanently shrinking the description to match the code, we built the grader the description had promised. The published wording is true again because the code changed, not because the claim was watered down.
- How the three parts combine: the portrayal grade is lowered to the strictest applicable cap. Stale information caps the result at 50, wrong information at 25, damaging information at 0, an adjacent category at 75, and a wrong category at 25. The grader is explicitly instructed never to mark a brand down simply because it does not recognise it, so a small or new brand is not penalised for the grader's own ignorance.
- Unchanged: the rater model, its temperature setting, and the mapping from grade to bucket. The five factors, their weights, and how per-answer scores are combined are all unchanged.
- Reliability re-measured, and published in full including the half that got worse. Agreement between our production grader and an independent grader from a different vendor improved from 0.599 to 0.624 (raw agreement 82.7% to 87.3%). That clears the 0.60 threshold conventionally treated as acceptable, but only if agreement against a different vendor is the right bar to measure against, and that has not been settled (see the next point). Agreement against a grader from the same vendor got materially worse over the same test: 0.610 on a same-day re-run of the old grader against 0.402 under the new one, versus the previously published figure of 0.645. The cause is visible in the disagreement log: asked about portrayal as an explicit sub-question, our grader moves borderline "listed among peers" answers from neutral to positive while the comparison grader holds them neutral. Label: PROXY, measured with the same harness, sample design and comparison graders as the July 27 baseline below, and carrying the same corpus limits: stored responses from internal test workspaces, clustered, with no customer responses among them.
- Open question, stated rather than buried: whether our reliability bar should be measured against a grader from the same vendor or a different one has not been settled, and the answer decides whether this grader passes or fails. The previously published figures of 0.645 and 0.599 describe the OLD Context grader and are not this grader's numbers.
- Expected score movement: direction is honestly unknown, because two forces pull opposite ways. The accuracy and category caps can only lower a result, while asking about portrayal explicitly tends to raise borderline neutral answers. On the 150-response test sample, 21.3% of answers in which the brand was mentioned changed their Context bucket, 98.7% of those by no more than one grade, and the average Context bucket moved from 69.5 to 72.8, worth roughly +0.67 points on the 0-100 composite for an average mentioned answer. Label: PROXY, measured on stored test-workspace responses in which one brand accounts for 5.3% of the pool.
- Answers in which your brand is not mentioned at all are untouched. The absence rule introduced in v2.1.0 returns Context 0 before the grader is ever called.
- The free checker's Context grading moved to the same three-part construct in the same change, so the free check and the paid audit measure the same thing under the same factor name.
- This does not restore the other correction made on July 28. Citation Link had been published as also crediting an attributed quotation of your brand, not just a link back to your domain. The scored check looks for a link to your own domain in the citation list and performs no quotation detection at all, so that description stays narrowed, and it will stay narrowed until the code does what the wider wording claimed. A brand quoted with attribution but never linked still scores zero on that factor.
- Stored scores are not rewritten. Every stored result carries the methodology version it was scored under. New runs are stamped 2.3.0. Any trend line crossing the 2.2.0 / 2.3.0 boundary must carry a version marker and must not be drawn as one continuous line.
v2.2.0 (July 31, 2026)
- Competitive Presence now scores 50 (neutral), not 0, when an AI answer names no brands at all. If an answer mentions neither you nor any of your tracked competitors, nobody won and nobody lost, and our published rubric has always scored that case as neutral. The code was returning 0.
- The other zero case is unchanged, deliberately. If competitors are named and you are not, Competitive Presence still scores 0. That is the worst case in this rubric, the two situations are exactly one competitor apart, and they must never collapse into each other. Both directions are locked by tests.
- Score movement: upward, and only for answers where nobody was named. Each such answer gains exactly 7.5 points on the 0-100 composite (0.0 becomes 7.5, and 20.0 becomes 27.5 where your domain was cited without your name appearing). No answer in which you or a tracked competitor was mentioned changes by any amount.
- Scale of the effect, reasoned from the corpus snapshot taken for v2.1.0 rather than re-measured: 3,974 of 10,488 scored answers (37.9%) named nobody at all, so on a corpus of that shape the average moves up by roughly 2.84 points. Label: PROXY, and the same caveat applies as for v2.1.0, namely that those results come from internal test workspaces rather than customer audits. As with v2.1.0 the run-level effect is largest for the least visible brands. The per-answer 7.5 points does not depend on the corpus.
- Read this together with v2.1.0, four days earlier. That change removed 10.0 points from the same class of answer; this one gives back 7.5, leaving those answers a net 2.5 points lower than before either fix. Both numbers now match the worked examples in our published rubric, which neither of them matched before.
- The free checker's competitive rubric line said 0 for "none", which conflated being beaten by competitors with nobody being named at all. It now states the neutral rule, and a deterministic code guard enforces 50 when the answer names no brands whatsoever, so a grader cannot reopen the gap.
- Stored scores are not rewritten. New runs are stamped 2.2.0. Comparisons crossing the 2.1.0 / 2.2.0 boundary must carry a version marker.
v2.1.4 (July 31, 2026)
- Measurement record, no change shipped. We ran the same 10 questions through all 5 engines twice on the same day, through the production scoring path, and compared the two runs cell by cell.
- Result across 49 paired cells (one call errored): average absolute difference between runs 1.86 points, correlation 0.962, 93.9% of cells within 5 points of each other, and 91.8% landing in the same grade band. The overall composite score moved 0.59 points between the two runs (7.88 against 8.47).
- Labels and limits: PROXY. This is one day, two runs, one brand context. No margin of error is published, deliberately. Unbacked confidence intervals were the exact defect behind the withdrawal recorded in v2.0.12 below, and repeating it here would be indefensible. Any future interval needs a statistical design that genuinely supports it.
- This is not a replacement for the agreement figure withdrawn in v2.0.12. That figure measured whether two graders agree about the same stored answer. This measures whether a whole audit produces the same score when you run it again. They are different things and must never be quoted for one another.
v2.1.3 (July 31, 2026)
- A cost reduction was tested and rejected. The proposal was to cut the number of live web searches each expensive engine performs per question from 5 to 3. Paired test, 20 fixed questions, both settings, production scoring path.
- Claude failed both drift gates: an average difference of 3.62 points against a tolerance of 2 points, and a worst single-cell difference of 45 points. Label on every figure in this entry: PROXY, from that single 20-question paired test.
- The direction of the failure is the finding worth publishing. With fewer searches, Claude scored the brand HIGHER in 4 of the 5 cells where the brand was mentioned. The saving would have systematically inflated customer scores, which is the opposite of a harmless efficiency.
- The Grok half of the test is recorded as invalid rather than failed. The setting used to limit searches did not actually bind: the supposedly limited runs still fired up to 7 searches (average 4.7 against 5.2). Its apparent drift is ordinary run-to-run variance, not an effect of search depth. Any future depth test on that engine needs a control that verifiably works.
- Nothing shipped. Search depth stays at 5. One of 80 calls errored and was counted against coverage rather than retried.
- This is the second cost reduction killed by the same gate. On June 18 a proposal to route monitoring questions to a cheaper engine failed the same way (v2.0.5). The pattern is consistent: on these engines the searches are the measurement, and cutting them changes the number.
v2.1.2 (July 31, 2026)
- Question panel enforcement turned on. A manual audit run now requires a frozen, versioned question list, and returns a clear message asking you to freeze one if none exists.
- Why this matters for your score: the question set IS the score. If the questions quietly change between two audits, the month-over-month comparison is meaningless. Freezing the list makes that comparison real.
- No scoring change. No factor, weight, engine, or aggregation moved, so the methodology version stays at 2.1.0. Turning the requirement on removes a source of unexplained score movement rather than adding one.
- Before the switch was flipped we checked what it could break: it gates exactly one route, the manual audit runner, and no frozen question lists existed yet, so no existing workflow could be blocked. The recurring weekly audits do not pass through this gate.
- Repeated sampling by default remains off. Running each question several times per engine multiplies the cost of every audit, and the sample design is being priced before any change is made.
v2.1.1 (July 30, 2026)
- Dashboard display fix. The headline score at the top of the dashboard and the per-engine tiles directly beneath it used the same formula but were reading different sets of results. The headline was capped at the newest 100 results while still carrying the full date-range label, so for a brand with more results than that, the label and the number disagreed.
- A second defect found in the same review, not in the original report: leaving a query unlimited is not the same as unlimited. Our database caps any single query response at 1,000 rows, so the per-engine tiles were silently truncating too, just at a higher threshold. Simply raising the headline's cap would have left that in place. Both surfaces now read the whole range through one shared, paged helper.
- No score displayed today changes. Measured rather than assumed: the largest 30-day result count for any single brand at the time of the fix was 33 results, so nothing in production was near either threshold. It matters going forward, because at the current audit cadence one brand crosses 100 results inside a fortnight.
- Classified as restoring published behaviour, not changing it. The rubric defines the score over the range you select, and truncating to the newest 100 results was never a defined behaviour. The five factors, the five weights, and the aggregation are untouched, so the methodology version stays at 2.1.0.
- One option was considered and deliberately not taken: redefining the headline as the average of the per-engine scores rather than the average of all results. That re-weights the measurement and is a genuine rubric change, so it does not belong inside a bug fix. It remains open.
v2.1.0 (July 27, 2026)
- The Context factor now scores 0, not 50, when your brand does not appear in an AI answer at all. This is a real change to customer-facing scores.
- Why the old behaviour was wrong: our own published rubric states that Context cannot exceed the Presence bucket when Presence is below a verbatim mention. If you are not mentioned, Presence is 0, so Context must be 0. The grader path that actually runs in production was returning 50. This was a conformance defect against our own published rules, not an open judgement call.
- It was live everywhere it could be. Of 10,488 scored results in the stored corpus, 3,974 (37.9%) contained no mention of the brand. All 3,974 carried Context 50 and none carried 0, confirming that the correct branch had never once run in production.
- Score movement, measured not estimated: each affected result loses exactly 10.0 points on the 0-100 scale, which is a deterministic consequence of the formula rather than a sample estimate. Across that stored corpus the average moves from 50.23 to 46.44. No result in which your brand WAS mentioned changes by any amount.
- The run-level effect varies far more than that average suggests and is not evenly distributed. The three runs recorded in the change record dropped 6.06, 7.98 and 10.00 points. The drop is largest for the least visible brands, because the correction removes an upward bias that had been flattering exactly the brands with the least real visibility.
- Label on the corpus figures: PROXY. The 10,488 scored results all come from internal test workspaces, not from customer audits, so the 37.9% share and the 50.23 to 46.44 average describe that corpus rather than a prediction about yours. The per-result 10.0-point drop does not depend on the corpus and holds everywhere.
- Two display bugs were caught alongside it and fixed in the same change: a legitimate score of 0 would have exported blank and rendered as a dash on one report, hiding the very finding this fix exists to surface.
- Stored scores are not rewritten. Every stored result carries the methodology version it was scored under, which is precisely what that stamp exists for; rewriting history would be worse than the drift. New runs are stamped 2.1.0. Any trend view crossing the 2.0.3 / 2.1.0 boundary must carry a version marker and must not be drawn as a continuous line.
- The free checker does not use this scoring path and is unaffected by this entry.
v2.0.12 (July 27, 2026)
- Correction. Two published statistics about our Context grader were withdrawn as invalid. Neither should have been published. Both had been live on our methodology integrity page from June 2026 until they were withdrawn on July 27, 2026, and both sat inside that page's structured data, which means they were machine-readable to the same AI engines this product measures.
- Withdrawn figure 1: "Cohen's kappa 0.947 (near-perfect agreement)". It was measured on 30 short sentences written by our own team and pasted into the test script, averaging about 88 characters each. They were not AI engine responses, so the test never measured the production grader on real inputs at all. It also took the first 30 items of a list ordered by grade, which left one grade represented by a single example. The caveat on the page described the samples as hand-selected. They were hand-written, which is a materially different and more serious admission.
- Withdrawn figure 2: "correlation 0.957, 88% exact agreement, 98% within one grade". The script behind it used a different prompt and rubric from the production grader, so it never measured production either. It also counted every errored grader call as a perfect match, with a code comment describing that as conservative when the direction is backwards. That inflated all three numbers simultaneously. This one is retracted with no replacement: we currently hold no valid human-labelled ground truth for this grader.
- Replacement figure, labelled PROXY: 82.7% raw agreement, with Cohen's kappa of 0.645 against an independent grader from the same vendor and 0.599 against one from a different vendor. Measured on a uniform random sample of 150 real stored engine responses drawn from a pool of 6,222, under methodology version 2.0.3.
- Read forward before quoting those three numbers. They measure the Context grader as it stood in July 2026, and that grader was replaced on July 31, 2026 (see v2.3.0 above). They are the correct figures for the grader of that date and the wrong figures for the one running now.
- Both samples published, not only the flattering one. A second, stratified sample of 97 responses produces higher kappa (0.769 and 0.668) on LOWER raw agreement (79.4%), also labelled PROXY and measured under methodology version 2.0.3. That is a known statistical artefact of stratifying on a category the grader reproduces almost every time, which is why the uniform sample is the headline figure: its mix of cases matches what real audits actually produce.
- Limits of the corpus, disclosed rather than glossed: all 6,222 responses come from test workspaces, with zero customer responses among them. They span 201 brands, 222 runs and 548 distinct questions, but they are clustered, so 150 responses represent fewer than 150 independent observations, with a single brand accounting for 5.3% of the pool. The rarest grade appears only 17 times in the whole corpus.
- No margin of error is published, deliberately. Confidence intervals were quoted internally during this investigation and turned out to be unbacked by any artefact: there was no interval calculation in the script and no interval in the run log. Publishing them would have reproduced the exact defect being corrected. Any future interval must be calculated across brands rather than across individual responses, because of the clustering above.
- One figure that reads well and must not be misread: re-running the production grader against itself on the same stored answer gives kappa 1.000. That is a deterministic re-read of a fixed piece of text, not end-to-end audit reproducibility, and it must never be presented as the latter. For actual run-to-run reproducibility of a whole audit, see v2.1.4 above.
- One question left open rather than rounded away: the cross-vendor figure of 0.599 sits one thousandth below the 0.60 threshold conventionally treated as acceptable agreement. That threshold governs whether a change to the grader is permitted, not whether the current grader may run, so nothing was blocked operationally.
- No score changed with this entry. No formula, weight, prompt, model, or bucket mapping moved. No stored score moves and no historical run is restated. This records a measurement of an unchanged grader and the correction of a public claim about it. The methodology version stayed at 2.0.3.
- Corrected, not quietly deleted. The integrity page names both withdrawn figures, states the mechanism of each error, states how long they were live and that they were machine-readable, and says plainly that neither number should have been published. Silently removing them would have been worse than the original error.
v2.0.11 (July 17, 2026)
- The system prompt used on every scored engine call was changed from a legacy financial-analyst persona to a neutral research-analyst prompt. The finance persona was left over from an earlier consulting product and had been conditioning every scored call for every industry, because no caller ever overrode it.
- Measured with a paired A/B test before shipping: 20 questions across 5 engines, run twice under each prompt, 324 scored calls. Overall the change moved scores by +1.0 point on average (standard deviation 10.9, median 0), which sits inside ordinary run-to-run noise for this system, where the average absolute difference between repeat runs of the same prompt was 8.3 points. Label on every figure in this entry: PROXY, from that single 324-call test.
- One directional exception, disclosed rather than averaged away: on advisory, non-branded questions the neutral prompt scored +14.3 points higher on average across 9 question pairs, with a maximum of +63.8. The finance persona had been suppressing the naming of non-finance brands on exactly those questions. Finance control questions moved -1.4 points.
- No regression: zero grounding failures and zero parsing failures under both prompts, and run-to-run variance did not get worse under the neutral prompt (9.0 down to 7.6 average absolute difference).
- The five factors, their weights, and the scoring formula are unchanged, so the methodology version stayed at 2.0.3. This changes how engines are prompted, not how their answers are scored.
v2.0.10 (July 8, 2026)
- Our weekly internal reproducibility test now completes its full question set every week. It had been running out of time after roughly 52 of 250 calls, which meant it sampled only the first ten or so questions in a fixed order rather than a representative slice.
- Fix: each engine now runs in its own lane rather than behind a single shared queue, with a soft deadline that stops the run cleanly, and week-over-week drift is only compared when both weeks completed at least 95% of their calls.
- A cost-accounting bug surfaced in the same review. Per-call cost was being recorded under each engine's public name rather than its vendor's name, so three of the five engines recorded zero cost and roughly 80% of this test's spend was invisible to the weekly spending guardrail that exists to stop it running away. Fixed.
- No customer-facing score is affected. This is the internal instrument that watches for drift, not the scoring pipeline.
- One-time discontinuity: the first complete week cannot be compared against the preceding partial weeks, because the difference reflects which questions were sampled rather than any change in the engines. The completeness gate suppresses that false alarm, and the first complete run resets the baseline.
v2.0.9 (July 4, 2026)
- Disclosure about our own internal reproducibility testing. Every weekly run before July 4, 2026 was being cut short at roughly 23 of 250 calls, about 9% of the frozen question set, and always the same first questions in the same fixed order. That is a biased sample, not merely a small one.
- Consequence, stated plainly: drift measurements from before this date are not comparable with those after it, and drift alerts had to be re-baselined from the first run following the fix.
- Fixes: the time budget for the test was raised, and a sweeper now marks any run killed by the hosting platform as timed out, which is the only way to record a final status after that kind of failure. Three previously stuck records were corrected.
- No customer-facing score is affected. This entry concerns the instrument that checks our measurements, not the measurements themselves.
v2.0.8 (July 3, 2026)
- Retroactive entry closing a gap in this changelog. The measurement instrument described below shipped on June 22, 2026 without an entry of its own; this entry was filed on July 3, 2026 once that was noticed.
- What shipped: versioned question panels (a frozen, approved question list per brand), repeated sampling with a margin of error attached to the run score, and evidence-first reporting.
- It shipped switched off. Both features sat behind flags that were verified off in production, so no score changed and no run carried a margin of error. With the flags off, the pipeline scored exactly as before: one sample per question and engine, with the run score being the average of the per-answer scores.
- The methodology version stayed at 2.0.3 by design. That number moves only when a factor, a weight, or the scoring formula changes. This entry documents a capability that did not alter how any answer was scored while it lay dormant.
- Governance recorded at the time and honoured since: turning either flag on is a scoring change and requires its own signed entry in this changelog before the switch is flipped. Question panel enforcement was turned on July 31, 2026, and has its own entry (v2.1.2). Repeated sampling by default has not been turned on, and will get its own entry here if it ever is.
v2.0.7 (June 19, 2026)
- Free checker only. The scored audit pipeline is untouched by this entry.
- Second fix for the same failure as v2.0.6. Raising the token limit stopped answers being cut off, but the free checker was still discarding whole engine results when an engine returned valid content wrapped in invalid formatting, such as a sentence of preamble before the data, or a raw line break inside a quoted string.
- Fix: a single tolerant parser now recovers those replies instead of throwing them away, and the six hand-rolled copies of the old fragile parsing code across the free tools were replaced with it, removing the same latent bug from all of them.
- Effect on the free score: directional. Recovering a dropped engine moves the result toward the true average of both engines instead of collapsing onto whichever engine survived. For the specific failure that triggered this, that is upward, but in general it removes single-engine bias rather than pushing one way. Retesting the case that triggered it gave full coverage, with all six engine calls succeeding where four had before.
- No rubric, weight, or factor change. The free checker's number remains a directional snapshot, not a precise reproducible figure.
v2.0.6 (June 19, 2026)
- Free checker only, and a disclosure as much as a fix. The free checker's score was not reproducible: the same brand scored 0, then 11, then 13 within minutes, and separate instrumented runs of the same brand produced results as far apart as 0 and 76.
- Cause 1, the dominant one: one engine's reply was frequently cut off mid-answer and then discarded entirely, collapsing the score onto the single surviving engine, which often produced 0.
- Cause 2: when the checker fell back to generating questions from a page title, the brand's own name ended up inside the questions, so the AI trivially named the brand and the score inflated to 76. Generated questions are now stripped of the brand name.
- Cause 3: questions were regenerated on every run, adding drift on top of the inherent variability of a three-question live web probe. Questions are now cached per domain.
- Fixes shipped: higher token limits, lower temperature on the affected call, tolerant parsing, cached questions, brand-name stripping, and error logging that had been silently swallowing the reasons for failure.
- Claim correction: the results page had been telling users that a score of 0 meant there was not enough web coverage about them, which was false. That narrative was removed, and the free checker's number is now presented as a directional snapshot rather than a precise figure. A live web probe of this size cannot be made precise by adjusting engine settings alone, and we stopped presenting it as though it could.
- The scored audit pipeline's rubric and weights are unchanged; the methodology version stayed at 2.0.3.
v2.0.5 (June 18, 2026)
- Gemini engine fix. Google withdrew the search setting our adapter used for Gemini 2.5 Flash, and the call began failing outright. The adapter moved to the current setting, with the same intent: force a live web search on every question.
- Citation handling changed with it. Source links returned by the new setting arrive as Google redirect addresses rather than real domains. The old adapter discarded those, so citations were being lost. The new one reconstructs the real domain. Citation Link scoring matches on domain, so the resulting scores are equivalent.
- A cost reduction was tested and rejected in the same window. We evaluated routing monitoring questions to the cheaper Gemini engine instead of a more expensive one. Across 20 questions the correlation between the two engines was 0.6685 against a required 0.70, and the worst single-question difference was 77.5 points against a permitted 5.0. Both gates failed, so the change was not made. Label: PROXY, a 20-question test, and the finding is about the substitution rather than about either engine in general. Within that test the disagreements clustered on general, non-branded questions where the two engines named different sets of smaller brands, which is why we treat the two as measuring different things rather than as interchangeable at different prices.
- No rubric, weight, or factor change. The methodology version stayed at 2.0.3. This was an upstream vendor deprecation, fixed.
v2.0.4 (June 10, 2026)
- Changelog key for the ChatGPT engine cutover of June 10, 2026, which is described in full in the entry immediately below, published there under the scoring version it carried (2.0.3).
- Recorded here because it was missed at the time. The cutover shipped without an entry of its own, and this key was filed retroactively on June 18, 2026 once the gap was found. Version keys in this changelog run continuously, so an absent key would be a silent gap rather than a visible one.
- No score effect beyond what the entry below describes. The scoring version remained 2.0.3.
v2.0.3 (June 10, 2026)
- Production audit-pipeline engine update: the ChatGPT adapter moved from OpenAI GPT-4o with the deprecated web_search_preview tool to GPT-5.5 with the current web_search tool, completing the migration the free checker shipped earlier today. OpenAI retires the GPT-4o search-preview path on 2026-07-23. The 5-factor rubric and factor weights are unchanged. Audit results produced after 2026-06-10 carry methodology_version "2.0.3" in the audit_results table.
- Grounding is now enforced on the audit pipeline's ChatGPT engine: tool_choice "required" forces a live web search on every scored call, and a tripwire records any response that did not actually run one as a failed measurement instead of scoring it.
- Shadow-audit finding, disclosed for transparency: the outgoing GPT-4o adapter ran a live web search on only 25% of the 24-query validation panel, answering evergreen category questions from training data even with the search tool enabled. This is the same structural defect found and fixed in the Grok adapter in May (v2.0.1).
- Cutover validation record: correlation with the outgoing adapter was r=0.632, below the 0.75 continuity threshold (expected, because an ungrounded baseline is not a valid comparator). Against the web-grounded Claude reference panel (the gate used for the May Grok cutover) the new adapter scored r=0.777, and r=0.956 after excluding one query affected by a known validation-script substring artifact. Grounding rate 100%, zero errors. Both CTO and CAIO signed off 2026-06-10; the full artifact is at scripts/logs/openai-shadow-audit-2026-06-10T19-25-17.json with the adjudication recorded in docs/engine-adapter-health.md.
- Expected score movement: grounded answers surface more long-tail brands, so ChatGPT-engine scores for niche brands can move up after 2026-06-10. In the validation panel, every old-vs-new score difference occurred on queries the old adapter had not searched, and the reference engine sided with the new adapter on four of six. Established-brand scores were unchanged. A post-cutover drift check against the first five production audits is scheduled by 2026-06-17.
- Determinism note: GPT-5.5 does not accept temperature pinning (the old adapter requested temperature 0). Run-to-run variance expectations are unchanged from the v2.0.2 findings; repeated sampling remains available on request.
- Customer-facing engine labels updated from ChatGPT (GPT-4o) to ChatGPT (GPT-5.5) across the dashboard, audit modals, and methodology pages.
v2026-06-10 (June 10, 2026)
- Free-checker engine update (quick-check and /analyze adapters only; the production audit pipeline's engine roster is unchanged): the ChatGPT adapter moved from OpenAI GPT-4o with the deprecated web_search_preview tool to GPT-5.5 with the current web_search tool. OpenAI retires the GPT-4o search-preview path on 2026-07-23.
- Grounding is now enforced on the free checker's ChatGPT adapter: tool_choice "required" forces a live web search on every call, and a tripwire rejects any response that did not actually run one. Previously the model could answer from training data; an ungrounded response is no longer scored.
- Honest-coverage disclosure: free-check responses now report questions attempted vs. completed and engine calls attempted vs. succeeded, and the results UI flags partial scores instead of silently averaging whatever survived.
- No change to the 5-factor rubric, factor weights, or audit-pipeline scoring. METHODOLOGY_VERSION remains 2.0.1.
v2.0.2 (May 19, 2026)
- Ran the first auditable reproducibility test on the v2.0.x methodology. 5 buyer-intent queries x 5 engines (OpenAI GPT-4o, Anthropic Claude Sonnet 4.6, Google Gemini 2.5 Flash, Perplexity Sonar Pro, xAI Grok 4.3) x 3 independent trials. Target brand: Profound (cited in this query topic). Competitors: Otterly, Peec, Pondral. 71 of 75 calls succeeded.
- Run-level aeo_score variance (mean across every (query, engine) result, the customer-facing number on the dashboard): max absolute deviation 1.74 points across the 3 trials. Within normal run-to-run variance. The fixed ±5-point reproducibility target was subsequently retired; repeated sampling was added as an optional mode.
- Per-cell (query, engine) result variance: 15 of 23 measurable pairs reproduced within 5 points absolute. 8 of 23 exceeded that band; worst case 38.37 points (Q2 GROK: scores 63.8, 10.0, 71.3 across the 3 trials). Variance is intrinsic to AI engine outputs, not a scoring bug. This finding motivated adding repeated sampling as an available mode for brands that need to measure per-cell consistency.
- Claim revision: replaced the ambiguous "±5% on re-run" wording across ~22 public surfaces. The fixed-threshold target was retired; standard audits grade one response per (query, engine) pair, and repeated sampling across independent trials is available on request.
- Methodology v3 roadmap (deferred): multi-run sampling per query, t-distribution confidence intervals on per-cell scores, Cohen's kappa inter-rater statistics. These would tighten per-cell variance toward the run-level bound. Currently unscheduled.
- Audit artifact: docs/reproducibility-audit-2026-05-19T12-11-51.json (full per-trial scores, deviations, and engine health).
v2.0.1 (May 10, 2026)
- Correction to v2.0.0 entry: the Grok engine adapter listed as "grok-3" actually runs grok-4.3 in production (migrated 2026-05-04, PR #103). Shadow audit post-correction passed at r=0.866 (threshold r≥0.75); grounding rate 100%; zero errors. Both CTO and CAIO signed off 2026-05-04
- Clarified engine-weight semantics: the planned explicit weighting referenced in v2.0.0 (chatgpt 0.35, claude 0.25, gemini 0.20, grok 0.20) was scaffolding for future use. The actual production AI Visibility Score (formerly AEO Score) formula computes the arithmetic mean of per-result methodology_score values across all engines, so engine weights are implicit at 1÷n_engines today. No engine is explicitly upweighted in the code
- No customer-facing score change. Audit results produced after 2026-05-10 carry methodology_version "2.0.1" in the audit_results table
v2.0.0 (May 6, 2026)
- Reconciled platform scoring with published methodology: the customer-facing AI Visibility Score (formerly AEO Score) is now the arithmetic mean of per-observation 5-factor composite scores (Presence, Prominence, Context, Citation Link, Competitive Presence)
- Retired the interim 4-component formula (Brand SOV 15% + Generic SOV 45% + Owned Citations 30% + Grounded Mentions 10%) that had been in use since launch
- Context factor (Factor 3) graded by Claude Haiku via dedicated LLM judge per observation, replacing keyword-based sentiment heuristic
- Prominence factor uses character-ratio position (brand appearance offset / response length), consistent across all engines
- Per-result methodology scores stored in audit_results.methodology_score alongside raw data for full auditability
- SOV metrics (brand-scope, generic-scope, per-theme, per-competitor) remain unchanged and continue to appear on dashboards alongside the reconciled AI Visibility Score
v2026-04-16 (April 16, 2026)
- Published initial methodology documentation at /methodology
- Defined 5-factor scoring rubric: Presence (20%), Prominence (25%), Context (20%), Citation Link (20%), Competitive Presence (15%)
- Established ordinal bucket scoring (0, 25, 50, 75, 100) for all factors
- Published initial reproducibility approach. The original fixed-threshold target was subsequently retired; standard audits grade one response per (query, engine) pair, with repeated sampling across independent trials available on request. Drift above expected variance is logged in this changelog with a root-cause note. Per-cell variance is visible in raw evidence.
- Configured separate LLM rater for Context factor (Factor 3) to avoid self-assessment bias
- Launched "View raw" transparency feature: every score shows prompt, response, timestamp, and rater model version
- Methodology roadmap items (not in this release): multi-run sampling per query, t-distribution confidence intervals, IQR outlier detection, Cohen's kappa inter-rater agreement statistics. Targeted for methodology v3
v2026-04-23 (April 23, 2026)
- Published detailed design rationale at /blog/how-pondral-scoring-works
- Documented weight selection process: Prominence weighted highest (25%) based on click-through correlation backtests
- Documented Competitive Presence weight reduction from original 25% to 15% due to volatility concerns
- Documented scoring bucket expansion from original 3-bucket (0, 50, 100) to 5-bucket (0, 25, 50, 75, 100) system
- Added AEO glossary at /blog/aeo-glossary with DefinedTermSet schema for all scoring terminology