Original HealthBench; May 2025 paper v1
HealthBench ↗
A rubric score is a balance of rewarded and penalized behavior.
HealthBench · independent analysis
Independent analysis of original HealthBench: rubric scoring, seven themes, five axes, historical results and an interactive signed-points calculator.
The original conversation distribution. Each bar is the share of all 5,000 conversations assigned to that theme.
Paper Table 2. These counts sum to 5,000; rubric-axis counts use a different denominator. [1]
Give the evaluated system the conversation context. [1]
The original paper uses GPT-4.1 to decide whether each rubric criterion is met. [1]
Keep negative points and divide by possible positive points for that example. [3]
Average normalized scores, then clip the overall mean. [1][3]
sᵢ = Σ(points × met) / Σ(positive points); score = 100 × clip(mean(sᵢ), 0, 1)
Keep rubric version, grader, sampling and subset fixed. A score of 60 is not 60% fully correct conversations. [1][3]
Selected paper-reported measurements. Dates, model configurations and scoring conditions belong to each panel.
Paper-reported results / selected rows
May 2025 paper, original HealthBench; mean across 16 evaluation runs in Table 7.
Paper-reported historical means. Standard deviations describe repeated-run variability, not confidence intervals for real-world clinical outcomes.
Source: Table 7 [1]
An original analytical tool
Toggle abstract illustrative criteria to see how signed points and the positive-point denominator interact. The example weights are invented for arithmetic; no benchmark questions or medical recommendations are reproduced.
Illustrative rubric / abstract criteria
These four made-up criteria demonstrate the scoring rule. They are not benchmark examples or clinical recommendations.
3 satisfied weight / 7 positive possible weight
Every satisfied criterion contributes its signed weight. Unmet positive criteria stay in the denominator. A satisfied negative criterion subtracts from the numerator.
This is one example’s contribution before aggregation. The published overall procedure clips the mean, not each example, to the 0–1 range. Counting criteria equally would answer a different question from applying their weights.
This illustrates the published score rule. It does not perform model grading, represent an actual HealthBench example, or turn the result into percent accuracy. Negative example scores must be retained until aggregation. [1][3]
Original HealthBench; May 2025 paper v1
A rubric score is a balance of rewarded and penalized behavior.
Original HealthBench evaluates open-ended health conversations with physician-written criteria. Its score combines rewarded behavior, penalties and a conversation-level denominator. This publication makes that measurement inspectable through sourced benchmark facts, a coverage breakdown and an abstract rubric calculator. The guides explain what a score means, how its units differ, and where grader validation applies. Arcophos provides independent analysis; OpenAI and its research collaborators created HealthBench.
Follow signed rubric points, positive-point denominators and aggregate clipping without confusing the result with accuracy.
Read conversation coverage and rubric composition without treating criteria as patients or repeated assignments as unique items.
Separate agreement with physicians, repeated-run variability and the clinical claims that neither measurement directly tests.
No. It averages conversation-level rubric scores that combine positive credit and negative penalties. Displaying the result times 100 does not turn it into the fraction of fully correct conversations.
Yes. Triggered penalties can exceed earned positive points. The published rule retains those values and clips the overall mean, rather than flooring each conversation first.
Unique criteria count distinct items; assignments count their uses across examples, including repetition. Both differ from the number of conversations.
No. Arcophos provides independent analysis and tools. OpenAI’s original paper, release, dataset and reference implementation are linked as the primary sources.
Working tool / saved on this device
Record the objects and settings needed to interpret a rubric score. The worksheet is a reporting aid, not an official benchmark run.
Every measurement has a source. Dossiers preserve benchmark versions, scoring conditions and access notes.
Download the evidence ↗