HealthBench · independent analysis

Understand the rubric behind the number.

Independent analysis of original HealthBench: rubric scoring, seven themes, five axes, historical results and an interactive signed-points calculator.

Independent analysis by Arcophos · updated

5,000 conversations, seven themes

HealthBench ↗

The original conversation distribution. Each bar is the share of all 5,000 conversations assigned to that theme.

Global health1,097
Responding under uncertainty1,071
Expertise-tailored communication919
Context seeking594
Emergency referrals482
Health data tasks477
Response depth360

Paper Table 2. These counts sum to 5,000; rubric-axis counts use a different denominator. [1]

  1. 01

    Generate the next response

    Give the evaluated system the conversation context. [1]

  2. 02

    Grade each criterion

    The original paper uses GPT-4.1 to decide whether each rubric criterion is met. [1]

  3. 03

    Normalize signed points

    Keep negative points and divide by possible positive points for that example. [3]

  4. 04

    Aggregate conversations

    Average normalized scores, then clip the overall mean. [1][3]

sᵢ = Σ(points × met) / Σ(positive points); score = 100 × clip(mean(sᵢ), 0, 1)

Keep rubric version, grader, sampling and subset fixed. A score of 60 is not 60% fully correct conversations. [1][3]

Published results

Source record ↗

Selected paper-reported measurements. Dates, model configurations and scoring conditions belong to each panel.

Paper-reported results / selected rows

Historical overall means across repeated runs

May 2025 paper, original HealthBench; mean across 16 evaluation runs in Table 7.

Rubric score × 100 · points
050100
Reported
o3Reported run SD 0.16 points; run range 59.51–60.14
59.9
GPT-4.1Reported run SD 0.22 points; run range 47.42–48.15
47.78
o1Reported run SD 0.22 points; run range 41.53–42.30
42
GPT-4o (Aug 2024)Reported run SD 0.20 points; run range 31.88–32.57
32.33

Paper-reported historical means. Standard deviations describe repeated-run variability, not confidence intervals for real-world clinical outcomes.

Source: Table 7 [1]

An original analytical tool

Build a HealthBench-style score

Inspect the arithmetic

Toggle abstract illustrative criteria to see how signed points and the positive-point denominator interact. The example weights are invented for arithmetic; no benchmark questions or medical recommendations are reproduced.

Illustrative rubric / abstract criteria

These four made-up criteria demonstrate the scoring rule. They are not benchmark examples or clinical recommendations.

Per-example weighted contribution42.86%

3 satisfied weight / 7 positive possible weight

Every satisfied criterion contributes its signed weight. Unmet positive criteria stay in the denominator. A satisfied negative criterion subtracts from the numerator.

This is one example’s contribution before aggregation. The published overall procedure clips the mean, not each example, to the 0–1 range. Counting criteria equally would answer a different question from applying their weights.

This illustrates the published score rule. It does not perform model grading, represent an actual HealthBench example, or turn the result into percent accuracy. Negative example scores must be retained until aggregation. [1][3]

The benchmark in detail

All dossiers →
Dossier01

Original HealthBench; May 2025 paper v1

HealthBench ↗

A rubric score is a balance of rewarded and penalized behavior.

UnitConversation with an example-specific rubricMeasureMean normalized rubric score

What we examine

Original HealthBench evaluates open-ended health conversations with physician-written criteria. Its score combines rewarded behavior, penalties and a conversation-level denominator. This publication makes that measurement inspectable through sourced benchmark facts, a coverage breakdown and an abstract rubric calculator. The guides explain what a score means, how its units differ, and where grader validation applies. Arcophos provides independent analysis; OpenAI and its research collaborators created HealthBench.

Score
Follow signed points through normalization and aggregation.
Coverage
Separate conversation themes from criterion-level behavior axes.
Evidence
Read historical results and grader validation within their measured scope.

Analysis & interpretation

All analyses →

Questions, answered

Is a HealthBench score percent accuracy?

No. It averages conversation-level rubric scores that combine positive credit and negative penalties. Displaying the result times 100 does not turn it into the fraction of fully correct conversations.

Can an individual HealthBench score be negative?

Yes. Triggered penalties can exceed earned positive points. The published rule retains those values and clips the overall mean, rather than flooring each conversation first.

Why do unique criteria and criterion assignments have different counts?

Unique criteria count distinct items; assignments count their uses across examples, including repetition. Both differ from the number of conversations.

Is this OpenAI’s official HealthBench website?

No. Arcophos provides independent analysis and tools. OpenAI’s original paper, release, dataset and reference implementation are linked as the primary sources.

Prepare a comparison worksheet

Working tool / saved on this device

Document a HealthBench result

Interactive worksheet

Record the objects and settings needed to interpret a rubric score. The worksheet is a reporting aid, not an official benchmark run.

Identify the measurement

Every measurement has a source. Dossiers preserve benchmark versions, scoring conditions and access notes.

Download the evidence ↗