← All projects
Machine learning & AI · Python
:)

sentiment

Naive Bayes written from scratch — the counting, the add-alpha smoothing, the log-space arithmetic — plus a tokeniser built for reviews. Type and watch every word's vote.

01 Classify a review

The trained model (1,500 synthetic reviews, α=1.0 Laplace smoothing) is embedded in this page. Unknown words are ignored rather than zeroing a class; negation is tagged until the clause ends (not_good); repeated letters are squeezed; capitals become a _caps marker.

Live as you type · examples below

02 Decision

positive —

Per-token contribution to log-odds (positive − negative)

    03 How the tokeniser helps

    problemhandled by
    not good ≠ goodnegation tagging until clause ends
    loooove, LOVEletters squeezed, caps → _caps
    :) 👍emoticons/emoji survive
    ! sentiment, , notpunctuation filtered, not blanket-stripped
    unseen wordsignored, not allowed to zero a class
    long documentslog-space sums, no underflow

    04 Results (Python package)

    modelaccuracyf1
    from scratch (unigrams)0.9680.969
    from scratch (+ bigrams)0.9720.973
    from scratch (no negation)0.9640.966
    sklearn MultinomialNB0.9680.969
    sklearn TF-IDF + logistic0.9720.973

    The from-scratch model matches sklearn's MultinomialNB to three decimals — it is the same algorithm, written out. Known limitation: contrast after “but” is not weighted; the model adds evidence and leans by volume.

    18 tests · Python 3.10+ · no runtime dependencies · Built by Umer Hashmi