LLM engineering · Python

Warmstart

A cache in front of a bank's AI support assistant, which answers 400,000 questions a month. It answers exact repeats and close paraphrases from the cache, keeps each customer's account answers to that customer, and empties itself when the prompt or the fees change. Below: 10,000 real support questions replayed through it.

Hit rate through the replay

Share of each 250 questions answered from the cache. It starts cold, warms up, drops to nothing when the system prompt changes, and dips when the fee answers are invalidated.

exact repeatparaphrase (semantic)

How wrong can it afford to be?

Each held-out Banking77 question answered from its nearest cached neighbour among the training questions. A hit is wrong when that neighbour wanted a different answer. Move the similarity floor and watch the trade.

similarity floor onlywith the agreement check

Every variant on the same 10,000 questions

Cost and waiting times use the assumptions at the bottom of the page. The last row is the trap: keying on the question alone.

Hit rateWrongLeakedStale$ / 1,000p50p95

Paraphrases it answered from the cache

A sample of semantic hits from the replay, lowest similarity first: the new question, and the earlier one whose answer it got.

Ask it yourself — hit or miss?

The cache running live in your browser over the project's own cached questions, matched with TF-IDF cosine similarity (one of the three embedders measured above). No server, no key. Try a reworded repeat, then move the floor.

The wrong ones

What is measured and what is assumed