Pondral
← All postsJournal

Inter-rater reliability for AI answers.

Published Feb 12, 2026By Pondral TeamRead time 14 min read

Editor's note (2026-05-18): this post describes inter-rater reliability concepts and the methodology Pondral plans to adopt at scale. Current implementation grades one response per (query, engine) pair and reports the mean across the audit. Repeated sampling (Mention Rate, Quality When Mentioned, and a combined Visibility Index) is available on request, not enabled on standard audits. Multi-run averaging by default, Cohen's kappa flagging, and 95% confidence intervals are roadmap items for further statistical rigor. See /methodology for the shipped specification.

Update, August 2026. Two things below need separating. Measuring our Context grader's agreement is done. Surfacing kappa flagging inside the product is still a roadmap item. We changed the grader in August 2026, so there are two sets of numbers and we publish both rather than the flattering one. Measured on 150 randomly drawn real AI engine responses, the grader now in production agrees with a different-provider grader on 87.3% of cases (Cohen's kappa 0.624), which clears the 0.6 convention described below. Against a larger grader from the same provider it agrees on 72.0% (kappa 0.402), which does not clear it and is a real drop from the previous grader. Which of those two comparisons is the bar we should be held to is genuinely unsettled, and we would rather say so than quote whichever number flatters us. The previous grader measured 82.7% and kappa 0.645 same-provider, 0.599 cross-provider; those are the figures that stood on this page until August 2026 and they describe a grader that is no longer running. We are stating all of this plainly rather than leaving the convention on this page and our own number on another. These are single measurements on our own benchmark responses and are directional (PROXY), not settled figures. They supersede an earlier published figure of 0.947, which was measured on short sentences written by our own team rather than on real engine responses. See the correction note for what was withdrawn and why.

A question we get often: "How do you handle the fact that AI models give different answers each run?"

The short answer is: with statistics. The long answer involves Cohen's kappa, bootstrapped confidence intervals, and a methodology paper we'll publish later this quarter.

The framework we describe below is the published end-state. Today, Pondral runs each (query, engine) pair once and reports the mean across all results in the audit. Cohen's kappa flagging and multi-run prompt variations are slated for the methodology v3 release. For a plain-English explanation of each factor we measure, see the five-factor rubric, explained slowly.

The shipped behavior is: a single sample per (query, engine), aggregated into a mean score. We publish the methodology change log so every score is traceable to a specific rubric version and run configuration. See methodology.

Last updated June 2026Run a free audit