Pondral
← Back to methodology
Reproducibility status

Max drift, last completed run: 0.9 pp

Pondral runs the same frozen evaluation suite monthly against every production engine. We compare each run to the one before it and publish the per-engine drift. The number above is the largest engine-level change between the two most recent runs, which were 30 days apart.

Cadence changed 2026-08-03. This suite ran weekly until then and monthly since, so the runs listed here are spaced differently before and after that date, and the drift above was measured across 30 days rather than a full month. No run has been skipped: the next one is due on the first of the month.

Per-engine drift, last run vs prior

EngineLast run meanPrior run meanDeltaStatus
ChatGPT77.778.20.5 ppStable
Claude74.975.90.9 ppStable
Gemini74.373.70.7 ppStable
Perplexity76.376.00.3 ppStable
Grok74.474.50.0 ppStable

A delta under 5 percentage points is within normal run-to-run variance. 5 to 7.5 pp is flagged as moderate. Anything 7.5 pp or higher is flagged as drift and surfaced in the next methodology changelog entry with an explanation.

Recent runs

Run dateOverall meanChatGPTClaudeGeminiPerplexityGrok
Jul 12, 202675.075.075.075.075.0
Jul 19, 202674.777.275.072.374.274.9
Jul 26, 202675.478.075.073.675.075.5
Aug 2, 202675.678.275.973.776.074.5
Sep 1, 202675.677.774.974.376.374.4

How this works

Every month a cron job runs a fixed, frozen set of eval queries against every production engine. Each (query, engine) response is scored on a deterministic 4-factor subset of the rubric. Context is excluded from this eval because it uses an LLM judge whose own non-determinism would pollute the reproducibility signal. The full 5-factor rubric is the one customer scores use.

For the full rubric, weights, and rater configuration, see the methodology page. Every methodology change is logged in the changelog.

Methodology v2.5.0 · cohort v0-2026-05 · last run Sep 1, 2026 (2 weeks ago)

Last verified 2 weeks agoScore your brand free