Evaluation setup
Eight models answered all 620 Real-POCQi questions and completed 200 MedSP1000 scenarios. The study contains 11,360 identity-blind judgments and 13,760 judgments across all conditions.
Evaluating reliability and same-model affinity in LLM-as-a-judge across medical question answering and multi-turn clinical conversations.
Every judge showed matched self-preference on the 620-question Real-POCQi task. Pooled Answer Rank Self-Preference was +1.219 positions, with substantial variation across models.
Positive values indicate that a judge ranked its own answer more favorably than outside judges ranked that same answer. Pooled Answer Rank Self-Preference is +1.219 [+1.177, +1.261]. Removing the rubric preserves 91.5% of candidate-pair orderings. Intervals are shown in the rightmost column.
Answer Score Bias compares the rubric scores that the own-model judge and outside judges assign to the same fixed response; it is based on how models score responses, not how they rank them. Positive values indicate more generous scoring of the judge's own response.
Models are sorted from highest to lowest Answer Score Bias. Red bars indicate more generous own-model scoring; gray bars indicate more critical own-model scoring. The rightmost column reports 95% confidence intervals.
Own-Answer Win Rate is the percentage of competitors that a judge ranks below its own model's answer. It reflects both answer quality and how favorably the judge treats its own answer.
The center marker denotes equal wins and losses against competitors. Unlike the fixed-answer measures, this rate can be high because the model wrote a stronger answer, because its judge favored that answer, or both.
Matched self-preference appears at every tested conversation length. Models are ranked by their average Answer Rank Bias across the 2-, 4-, 6-, and 8-turn analyses.
The rightmost column shows 95% confidence intervals over 200 complete questions at every length.
| Rank | Judge | Average rank self-preference ↓ | 2 turns | 4 turns | 6 turns | 8 turns |
|---|---|---|---|---|---|---|
| 01 | GPT-5.6 Sol | 0.893 | 0.77 | 0.73 | 0.87 | 1.20 |
| 02 | GPT-5.6 Terra | 0.838 | 0.87 | 0.72 | 0.86 | 0.90 |
| 03 | Gemini 3.1 Pro | 0.760 | 1.05 | 0.67 | 0.67 | 0.65 |
| 04 | Gemini 3.7 Flash | 0.753 | 0.78 | 0.68 | 0.75 | 0.80 |
| 05 | Qwen 3.5 122B | 0.485 | 0.62 | 0.38 | 0.53 | 0.41 |
| 06 | Claude Opus 5 | 0.418 | 0.10 | 0.46 | 0.64 | 0.47 |
| 07 | Claude Sonnet 5 | 0.288 | 0.04 | 0.35 | 0.39 | 0.37 |
| 08 | Qwen 3.8 27B | 0.220 | 0.32 | 0.21 | 0.24 | 0.11 |
| — | Pooled | 0.583 | 0.57 | 0.53 | 0.62 | 0.61 |
Positive Answer Rank Self-Preference means greater self-preference. The pooled 95% confidence intervals are [0.49, 0.65], [0.45, 0.60], [0.54, 0.70], and [0.53, 0.69] at 2, 4, 6, and 8 turns, respectively.
Revealing generator names modestly weakened, rather than amplified, pooled self-preference. Every model still had positive Answer Rank Self-Preference in the identity-revealed condition.
| Judge | Revealed rank self-preference | Blinded rank self-preference | Change ↑ | Revealed 95% CI |
|---|---|---|---|---|
| Gemini 3.1 Pro | +0.536 | +0.777 | −0.241 | [+0.347, +0.725] |
| Gemini 3.7 Flash | +1.109 | +1.232 | −0.123 | [+0.959, +1.259] |
| Claude Opus 5 | +0.877 | +0.982 | −0.106 | [+0.794, +0.959] |
| GPT-5.6 Sol | +2.897 | +2.982 | −0.086 | [+2.723, +3.071] |
| GPT-5.6 Terra | +2.171 | +2.257 | −0.086 | [+2.007, +2.334] |
| Claude Sonnet 5 | +0.638 | +0.714 | −0.075 | [+0.443, +0.833] |
| Qwen 3.8 27B | +0.165 | +0.223 | −0.058 | [+0.032, +0.298] |
| Qwen 3.5 122B | +0.779 | +0.757 | +0.022 | [+0.573, +0.985] |
| Pooled | +1.146 | +1.240 | −0.094 | [+1.078, +1.215] |
Paired comparison on 200 questions and all eight judges, sorted by revealed-minus-blinded change. Negative change means weaker self-preference after names are revealed. The pooled change is −0.094 [−0.148, −0.040].
The study uses fixed-response comparisons to separate judging behavior from differences in answer quality.
Eight models answered all 620 Real-POCQi questions and completed 200 MedSP1000 scenarios. The study contains 11,360 identity-blind judgments and 13,760 judgments across all conditions.
Answer Rank Self-Preference compares ranks for the same fixed answer. Answer Score Bias analogously compares rubric scores, and Own-Answer Win Rate reports how often a judge ranks its answer above a competitor. Larger values of all three measures point toward greater self-preference.
Results remove categorical presentation-position effects. Confidence intervals treat questions as independent units and average correlated judge effects within each question for pooled estimates.
The study measures affinity for realized answers, not conscious self-recognition. Style, reasoning conventions, response length, and family-specific quality criteria may mediate the observed effects.
Use this citation when referencing the benchmark or its results.
@misc{ansari2026judge,
title = {The Judge Is Not Impartial:
Self-Preference in Medical LLM Evaluation},
author = {Ansari, Natarajan, Udash, Garvey, Yip, Lin, Fanous, Daneshjou},
year = {2026},
month = sep
}