LLM engineering · Python

Eval Dataset Generator

The best evaluation set comes from your own traffic, but nobody can label all of it. This pipeline redacts and deduplicates logged messages, clusters them, and picks which ones to label: spread over the clusters, and tilted toward outliers and low-confidence answers, where the model fails most. Every case carries a weight, so the accuracy measured on the set still estimates accuracy on all the traffic. Below: 3,080 logged banking questions, a budget of 200 labels, 300 draws per strategy.

From logs to candidates

What happens to the traffic before a single label is spent.

Three ways to spend 200 labels

Uniform: a random sample. Stratified: spread across clusters. Boosted: stratified, with outliers and low-confidence answers drawn more often. Means over 300 draws each.

Why every case carries a weight

Boosted sampling finds failures by oversampling them, so a plain average over the set understates accuracy. Weighting each case by the inverse of its chance of being drawn (a Hájek estimator) corrects that. Error of the accuracy estimate against the model's true accuracy on all the traffic, in points: