LLM engineering · Python

Groundtruth

A retrieval evaluation harness for a legal research assistant: 100 questions over 510 real contracts, each answered by a passage lawyers labelled. The recommended configuration puts that passage in the top 10 for 70.0% of questions, against 36.8% for plain BM25, and CI fails any change that costs more than a point.

Every configuration, on the same 100 questions

Recall@10: share of the labelled passage found in the top 10. "At 2,560 words" compares chunk sizes at the same amount of text handed to the model.

ChunksRetrievalRerankerContract expansionRecall@10Recall@20At 2,560 wordsMRR@10ms

Recall@10 by clause type

Five questions per clause type, so each figure moves in steps of about 20 points: read the pattern, not the decimals.

plain BM25recommended

Compare the answers, question by question

The top five passages each configuration returns. Green is the passage lawyers labelled, amber is the right contract but the wrong passage, grey is another contract.

The gate