Pondral
← Back to methodology
Reproducibility status

Last week's max drift: 1.0 pp

Pondral runs the same frozen evaluation suite every week against every production engine. We compare each run to the prior week and publish the per-engine drift. The number above is the largest engine-level change between the two most recent runs.

Per-engine drift, last run vs prior

EngineLast run meanPrior week meanDeltaStatus
ChatGPT78.278.00.2 ppStable
Claude75.975.00.8 ppStable
Gemini73.773.60.0 ppStable
Perplexity76.075.01.0 ppStable
Grok74.575.51.0 ppStable

A delta under 5 percentage points is within normal run-to-run variance. 5 to 7.5 pp is flagged as moderate. Anything 7.5 pp or higher is flagged as drift and surfaced in the next methodology changelog entry with an explanation.

Recent weekly runs

Run dateOverall meanChatGPTClaudeGeminiPerplexityGrok
Jul 5, 202674.275.073.075.073.075.0
Jul 12, 202675.277.575.573.274.875.2
Jul 19, 202674.777.275.072.374.274.9
Jul 26, 202675.478.075.073.675.075.5
Aug 2, 202675.678.275.973.776.074.5

How this works

Every week a cron job runs a fixed, frozen set of eval queries against every production engine. Each (query, engine) response is scored on a deterministic 4-factor subset of the rubric. Context is excluded from this eval because it uses an LLM judge whose own non-determinism would pollute the reproducibility signal. The full 5-factor rubric is the one customer scores use.

For the full rubric, weights, and rater configuration, see the methodology page. Every methodology change is logged in the changelog.

Methodology v2.5.0 · cohort v0-2026-05 · last run Aug 2, 2026 (1 week ago)

Last verified 1 week agoScore your brand free