textsearch
Inverted index and BM25 in the browser β every score is computed live, not faked.
1 Search
Twelve documents about information retrieval. Tokens are lowercased, stripped of stop words and optionally stemmed; ranked with BM25 (Lucene idf).
2 Explain a term
idf and per-document contribution β the textsearch explain view.
3 Inverted index
Postings for a term: document ids with term frequency and positions.
4 Textbook BM25 idf goes negative
The published idf falls below zero once a term appears in more than half the collection β the score then subtracts. Lucene's 1 + β¦ floors it above zero. Measured on this corpus:
| form | value for document (df=10/12) |
|---|---|
| ln((N β df + 0.5) / (df + 0.5)) | β1.4351 |
| ln(1 + (N β df + 0.5) / (df + 0.5)) | +0.2136 |
The Lucene form is whatβs running above. Set k1=0 to flatten saturation, b=0 to switch length normalisation off β both update live.
5 Corpus
6 Index statistics dashboard
Live counters recomputed from the inverted index whenever the stemmer setting changes.
7 Advanced query syntax
Everything the parser understands β click a row to try it.
term1 AND term2intersection β both terms must appearterm1 OR term2union β either term matchesNOT termcomplement β exclude documents containing term( β¦ )grouping β NOT binds tightest, then AND, then ORphrase: a b cexact adjacency using positional postingsfuzzy: mispelledit-distance fallback when nothing matches exactlyplain termsdefault BM25 bag-of-words rankingimplicit ANDadjacent plain terms in boolean mode intersect8 IDF visualizer
Drag the document-frequency slider: textbook idf dives below zero past df > N/2, Lucene's 1 + ln((N β df + 0.5)/(df + 0.5))-style floor stays positive. Curve plotted against df for this corpus (N = 12).
The tf slider shows the BM25 saturation numerator tfΒ·(k1+1)/(tf + k1Β·(1βb+bΒ·dl/avgdl)) β value displayed live.
9 Porter stemmer steps
Type a word β each suffix rule application is shown in order, matching the stemmer running in search.
10 Typo tolerance
Fuzziness slider: how many edits away a vocabulary term may be. Match quality (1 β d/max(len)) shown per candidate.
11 Search performance
Time the query: postings scanned (sum of posting-list lengths touched), documents scored, wall-clock ms.
12 Relevance feedback
Mark documents relevant to your query β Rocchio-style expansion suggests terms from those docs that are absent from the query.