Last week's max drift: 1.0 pp
Pondral runs the same frozen evaluation suite every week against every production engine. We compare each run to the prior week and publish the per-engine drift. The number above is the largest engine-level change between the two most recent runs.
Per-engine drift, last run vs prior
| Engine | Last run mean | Prior week mean | Delta | Status |
|---|---|---|---|---|
| ChatGPT | 78.2 | 78.0 | 0.2 pp | Stable |
| Claude | 75.9 | 75.0 | 0.8 pp | Stable |
| Gemini | 73.7 | 73.6 | 0.0 pp | Stable |
| Perplexity | 76.0 | 75.0 | 1.0 pp | Stable |
| Grok | 74.5 | 75.5 | 1.0 pp | Stable |
A delta under 5 percentage points is within normal run-to-run variance. 5 to 7.5 pp is flagged as moderate. Anything 7.5 pp or higher is flagged as drift and surfaced in the next methodology changelog entry with an explanation.
Recent weekly runs
| Run date | Overall mean | ChatGPT | Claude | Gemini | Perplexity | Grok |
|---|---|---|---|---|---|---|
| Jul 5, 2026 | 74.2 | 75.0 | 73.0 | 75.0 | 73.0 | 75.0 |
| Jul 12, 2026 | 75.2 | 77.5 | 75.5 | 73.2 | 74.8 | 75.2 |
| Jul 19, 2026 | 74.7 | 77.2 | 75.0 | 72.3 | 74.2 | 74.9 |
| Jul 26, 2026 | 75.4 | 78.0 | 75.0 | 73.6 | 75.0 | 75.5 |
| Aug 2, 2026 | 75.6 | 78.2 | 75.9 | 73.7 | 76.0 | 74.5 |
How this works
Every week a cron job runs a fixed, frozen set of eval queries against every production engine. Each (query, engine) response is scored on a deterministic 4-factor subset of the rubric. Context is excluded from this eval because it uses an LLM judge whose own non-determinism would pollute the reproducibility signal. The full 5-factor rubric is the one customer scores use.
For the full rubric, weights, and rater configuration, see the methodology page. Every methodology change is logged in the changelog.
Methodology v2.5.0 · cohort v0-2026-05 · last run Aug 2, 2026 (1 week ago)