← All projects
Machine learning & AI · Python
:)
sentiment
Naive Bayes written from scratch — the counting, the add-alpha smoothing, the log-space arithmetic — plus a tokeniser built for reviews. Type and watch every word's vote.
01 Classify a review
The trained model (1,500 synthetic reviews, α=1.0 Laplace smoothing) is embedded in this page. Unknown words are ignored rather than zeroing a class; negation is tagged until the clause ends (not_good); repeated letters are squeezed; capitals become a _caps marker.
Live as you type · examples below
02 Decision
positive
—
Per-token contribution to log-odds (positive − negative)
03 How the tokeniser helps
| problem | handled by |
|---|---|
not good ≠ good | negation tagging until clause ends |
loooove, LOVE | letters squeezed, caps → _caps |
:) 👍 | emoticons/emoji survive |
! sentiment, , not | punctuation filtered, not blanket-stripped |
| unseen words | ignored, not allowed to zero a class |
| long documents | log-space sums, no underflow |
04 Results (Python package)
| model | accuracy | f1 |
|---|---|---|
| from scratch (unigrams) | 0.968 | 0.969 |
| from scratch (+ bigrams) | 0.972 | 0.973 |
| from scratch (no negation) | 0.964 | 0.966 |
| sklearn MultinomialNB | 0.968 | 0.969 |
| sklearn TF-IDF + logistic | 0.972 | 0.973 |
The from-scratch model matches sklearn's MultinomialNB to three decimals — it is the same algorithm, written out. Known limitation: contrast after “but” is not weighted; the model adds evidence and leans by volume.