From logs to candidates
What happens to the traffic before a single label is spent.
Three ways to spend 200 labels
Uniform: a random sample. Stratified: spread across clusters. Boosted: stratified, with outliers and low-confidence answers drawn more often. Means over 300 draws each.
Why every case carries a weight
Boosted sampling finds failures by oversampling them, so a plain average over the set understates accuracy. Weighting each case by the inverse of its chance of being drawn (a Hájek estimator) corrects that. Error of the accuracy estimate against the model's true accuracy on all the traffic, in points: