LLM engineering · Python

Judge Calibration

Before an LLM judge's scores drive decisions, measure it against people. Here GPT-4's 1-to-5 ratings of 1,399 instruction-following outputs are compared with two independent human pools on five criteria: how often they agree, whether the judge favours long answers or its own model family, and whether a calibration map brings its scores in line.

The judge's score against people's

Each cell counts outputs by the judge's score (rows) and the human pool's average of two raters (columns). Agreement is quadratic-weighted kappa: 1 is perfect, 0 is chance. The two human pools agreeing with each other is the ceiling a judge can be held to.

Mean score per model

Does a calibration map help?

Kappa against the human pool the mapping never saw (mean absolute error in brackets), cross-fitted. Features: output length and model family.

Biases, every criterion

Length: partial rank correlation of the score with the output's length, a human pool's score held fixed. Self-preference: how much more the judge favours GPT-family outputs over Cohere's than the humans do, at the same human score; highlighted where the 95% interval leaves out zero.