← All projects
Data & analytics Β· Python
T

textsearch

Inverted index and BM25 in the browser β€” every score is computed live, not faked.

1 Search

Twelve documents about information retrieval. Tokens are lowercased, stripped of stop words and optionally stemmed; ranked with BM25 (Lucene idf).

2 Explain a term

idf and per-document contribution β€” the textsearch explain view.

Press Explain.

3 Inverted index

Postings for a term: document ids with term frequency and positions.

Press Look up.

4 Textbook BM25 idf goes negative

The published idf falls below zero once a term appears in more than half the collection β€” the score then subtracts. Lucene's 1 + … floors it above zero. Measured on this corpus:

formvalue for document (df=10/12)
ln((N βˆ’ df + 0.5) / (df + 0.5))βˆ’1.4351
ln(1 + (N βˆ’ df + 0.5) / (df + 0.5))+0.2136

The Lucene form is what’s running above. Set k1=0 to flatten saturation, b=0 to switch length normalisation off β€” both update live.

5 Corpus

Loading…

6 Index statistics dashboard

Live counters recomputed from the inverted index whenever the stemmer setting changes.

7 Advanced query syntax

Everything the parser understands β€” click a row to try it.

term1 AND term2intersection β€” both terms must appear
term1 OR term2union β€” either term matches
NOT termcomplement β€” exclude documents containing term
( … )grouping β€” NOT binds tightest, then AND, then OR
phrase: a b cexact adjacency using positional postings
fuzzy: mispelledit-distance fallback when nothing matches exactly
plain termsdefault BM25 bag-of-words ranking
implicit ANDadjacent plain terms in boolean mode intersect

8 IDF visualizer

Drag the document-frequency slider: textbook idf dives below zero past df > N/2, Lucene's 1 + ln((N βˆ’ df + 0.5)/(df + 0.5))-style floor stays positive. Curve plotted against df for this corpus (N = 12).

The tf slider shows the BM25 saturation numerator tfΒ·(k1+1)/(tf + k1Β·(1βˆ’b+bΒ·dl/avgdl)) β€” value displayed live.

9 Porter stemmer steps

Type a word β€” each suffix rule application is shown in order, matching the stemmer running in search.

10 Typo tolerance

Fuzziness slider: how many edits away a vocabulary term may be. Match quality (1 βˆ’ d/max(len)) shown per candidate.

quality = 1 βˆ’ distance / max(len(query), len(term)). Ties broken by df (higher wins).

11 Search performance

Time the query: postings scanned (sum of posting-list lengths touched), documents scored, wall-clock ms.

12 Relevance feedback

Mark documents relevant to your query β€” Rocchio-style expansion suggests terms from those docs that are absent from the query.

Select relevant docs, then press Expand query.
115 tests Β· Python 3.10–3.12 Β· standard library only Β· Built by Umer Hashmi