Field by field
Share of the 200 held-out receipts where the extracted value matches the label (case, spacing and punctuation aside; dates and totals compared as values). The dashed line is the OCR ceiling: how often any run of OCR lines matches the label exactly.
More training receipts
The trained extractor with 100, 200 and all 426 training receipts.
Speed and cost
p95 latency per receipt and compute cost per 1,000 on one CPU core, OCR aside.
Try the parsers
The schema's date and amount parsers, the same rules as slotfill/schema.py, running in your browser. Type a line the way a receipt prints it.