LLM Medical Self Preference

Evaluating reliability and same-model affinity in LLM-as-a-judge across medical question answering and multi-turn clinical conversations.

Single-Turn Self-Preference

Every judge showed matched self-preference on the 620-question Real-POCQi task. Pooled Answer Rank Self-Preference was +1.219 positions, with substantial variation across models.

Model / Judge Answer Score Bias Estimate 95% CI
GPT-5.6 Sol+2.939[+2.84, +3.04]
GPT-5.6 Terra+2.224[+2.12, +2.33]
Gemini 3.7 Flash+1.236[+1.16, +1.32]
Claude Opus 5+0.969[+0.92, +1.02]
Qwen 3.5 122B+0.730[+0.61, +0.85]
Gemini 3.1 Pro+0.683[+0.58, +0.79]
Claude Sonnet 5+0.647[+0.53, +0.76]
Qwen 3.8 27B+0.324[+0.23, +0.42]

Positive values indicate that a judge ranked its own answer more favorably than outside judges ranked that same answer. Pooled Answer Rank Self-Preference is +1.219 [+1.177, +1.261]. Removing the rubric preserves 91.5% of candidate-pair orderings. Intervals are shown in the rightmost column.

Answer Score Bias

Answer Score Bias compares the rubric scores that the own-model judge and outside judges assign to the same fixed response; it is based on how models score responses, not how they rank them. Positive values indicate more generous scoring of the judge's own response.

Model / Judge Answer Score Bias Estimate 95% CI
Gemini 3.7 Flash+0.67[+0.65, +0.69]
Qwen 3.5 122B+0.64[+0.60, +0.69]
Gemini 3.1 Pro+0.55[+0.52, +0.59]
GPT-5.6 Sol+0.32[+0.30, +0.35]
Claude Opus 5+0.02[+0.00, +0.03]
Qwen 3.8 27B+0.00[−0.05, +0.05]
Claude Sonnet 5−0.09[−0.12, −0.07]
GPT-5.6 Terra−0.75[−0.79, −0.71]

Models are sorted from highest to lowest Answer Score Bias. Red bars indicate more generous own-model scoring; gray bars indicate more critical own-model scoring. The rightmost column reports 95% confidence intervals.

Own-Answer Win Rate

Own-Answer Win Rate is the percentage of competitors that a judge ranks below its own model's answer. It reflects both answer quality and how favorably the judge treats its own answer.

Model / Judge Own-Answer Win Rate Rate
Claude Opus 599.4%
GPT-5.6 Terra93.7%
GPT-5.6 Sol89.5%
Gemini 3.7 Flash73.7%
Claude Sonnet 559.8%
Gemini 3.1 Pro56.8%
Qwen 3.5 122B30.2%
Qwen 3.8 27B18.8%

The center marker denotes equal wins and losses against competitors. Unlike the fixed-answer measures, this rate can be high because the model wrote a stronger answer, because its judge favored that answer, or both.

Multi-Turn Self Preference

Matched self-preference appears at every tested conversation length. Models are ranked by their average Answer Rank Bias across the 2-, 4-, 6-, and 8-turn analyses.

Visible turns Pooled Answer Rank Bias Estimate 95% CI
2 turns+0.57[+0.49, +0.65]
4 turns+0.53[+0.45, +0.60]
6 turns+0.62[+0.54, +0.70]
8 turns+0.61[+0.53, +0.69]

The rightmost column shows 95% confidence intervals over 200 complete questions at every length.

RankJudgeAverage rank self-preference ↓2 turns4 turns6 turns8 turns
01GPT-5.6 Sol0.8930.770.730.871.20
02GPT-5.6 Terra0.8380.870.720.860.90
03Gemini 3.1 Pro0.7601.050.670.670.65
04Gemini 3.7 Flash0.7530.780.680.750.80
05Qwen 3.5 122B0.4850.620.380.530.41
06Claude Opus 50.4180.100.460.640.47
07Claude Sonnet 50.2880.040.350.390.37
08Qwen 3.8 27B0.2200.320.210.240.11
—Pooled0.5830.570.530.620.61

Positive Answer Rank Self-Preference means greater self-preference. The pooled 95% confidence intervals are [0.49, 0.65], [0.45, 0.60], [0.54, 0.70], and [0.53, 0.69] at 2, 4, 6, and 8 turns, respectively.

When identities are visible

Revealing generator names modestly weakened, rather than amplified, pooled self-preference. Every model still had positive Answer Rank Self-Preference in the identity-revealed condition.

JudgeRevealed rank self-preferenceBlinded rank self-preferenceChange ↑Revealed 95% CI
Gemini 3.1 Pro+0.536+0.777−0.241[+0.347, +0.725]
Gemini 3.7 Flash+1.109+1.232−0.123[+0.959, +1.259]
Claude Opus 5+0.877+0.982−0.106[+0.794, +0.959]
GPT-5.6 Sol+2.897+2.982−0.086[+2.723, +3.071]
GPT-5.6 Terra+2.171+2.257−0.086[+2.007, +2.334]
Claude Sonnet 5+0.638+0.714−0.075[+0.443, +0.833]
Qwen 3.8 27B+0.165+0.223−0.058[+0.032, +0.298]
Qwen 3.5 122B+0.779+0.757+0.022[+0.573, +0.985]
Pooled+1.146+1.240−0.094[+1.078, +1.215]

Paired comparison on 200 questions and all eight judges, sorted by revealed-minus-blinded change. Negative change means weaker self-preference after names are revealed. The pooled change is −0.094 [−0.148, −0.040].

How to read this benchmark

The study uses fixed-response comparisons to separate judging behavior from differences in answer quality.

Evaluation setup

Eight models answered all 620 Real-POCQi questions and completed 200 MedSP1000 scenarios. The study contains 11,360 identity-blind judgments and 13,760 judgments across all conditions.

Three measures

Answer Rank Self-Preference compares ranks for the same fixed answer. Answer Score Bias analogously compares rubric scores, and Own-Answer Win Rate reports how often a judge ranks its answer above a competitor. Larger values of all three measures point toward greater self-preference.

Adjustment

Results remove categorical presentation-position effects. Confidence intervals treat questions as independent units and average correlated judge effects within each question for pooled estimates.

Interpretation

The study measures affinity for realized answers, not conscious self-recognition. Style, reasoning conventions, response length, and family-specific quality criteria may mediate the observed effects.

Cite this work

Use this citation when referencing the benchmark or its results.

@misc{ansari2026judge,
  title   = {The Judge Is Not Impartial:
             Self-Preference in Medical LLM Evaluation},
  author  = {Ansari, Natarajan, Udash, Garvey, Yip, Lin, Fanous, Daneshjou},
  year    = {2026},
  month   = sep
}