LLM engineering · Python

LLM Arbitration

When several models review an answer, they disagree, and someone has to decide. Here the critics' verdicts go to a naive Bayes adjudicator that has learned, per critic and task, how often each one passes right and wrong answers, and it returns the probability the answer is right. Everything below is out-of-fold on 4,551 HELM Lite questions, with other models' recorded answers as the critics.

Where to draw the line

The arbiter flags an answer for review when its probability of being right falls below the threshold. Lower the threshold and fewer answers are flagged, more of them truly wrong; raise it and more wrong answers are caught at the cost of more reviews. The dots are the simpler rules it replaces.

What the critics said, and what the arbiter concluded

All six answering models

Precision: flagged answers that were wrong. Recall: wrong answers that were flagged. The arbiter's threshold here is 0.5.