The judge's score against people's
Each cell counts outputs by the judge's score (rows) and the human pool's average of two raters (columns). Agreement is quadratic-weighted kappa: 1 is perfect, 0 is chance. The two human pools agreeing with each other is the ceiling a judge can be held to.
Mean score per model
Does a calibration map help?
Kappa against the human pool the mapping never saw (mean absolute error in brackets), cross-fitted. Features: output length and model family.
Biases, every criterion
Length: partial rank correlation of the score with the output's length, a human pool's score held fixed. Self-preference: how much more the judge favours GPT-family outputs over Cohere's than the humans do, at the same human score; highlighted where the 95% interval leaves out zero.