A retrieval evaluation harness for a legal research assistant: 100 questions over 510 real contracts, each answered by a passage lawyers labelled. The recommended configuration puts that passage in the top 10 for 70.0% of questions, against 36.8% for plain BM25, and CI fails any change that costs more than a point.
Recall@10: share of the labelled passage found in the top 10. "At 2,560 words" compares chunk sizes at the same amount of text handed to the model.
| Chunks | Retrieval | Reranker | Contract expansion | Recall@10 | Recall@20 | At 2,560 words | MRR@10 | ms |
|---|
Five questions per clause type, so each figure moves in steps of about 20 points: read the pattern, not the decimals.
The top five passages each configuration returns. Green is the passage lawyers labelled, amber is the right contract but the wrong passage, grey is another contract.
results/baseline.json.demo/smaller-chunks shrinks chunks to 128 words, a change that sounds harmless. Recall@10 falls from 70.0 to 49.8 and CI goes red.