Skip to content

Naive Bayes Classifiers

A fast probabilistic classifier built on a bold simplifying assumption — and why it still works well for text.

Editorial team 2 min read

Naive Bayes classifies examples using Bayes' theorem, combining how common each class is with how likely the observed features are under each class.

The "Naive" Assumption

It assumes every feature is independent of the others given the class. For text, that means treating each word as unrelated to the next — clearly untrue, but the simplification makes the model extremely fast and it often classifies well anyway.

Variants

  • Multinomial Naive Bayes — word counts, the classic choice for text.
  • Bernoulli Naive Bayes — presence or absence of words.
  • Gaussian Naive Bayes — continuous features assumed to follow a normal distribution.

Strengths

  • Trains and predicts very quickly, even on large vocabularies.
  • Needs relatively little training data.
  • A strong baseline for spam filtering, sentiment and topic classification.

Weaknesses

  • Probability estimates are often poorly calibrated (too confident).
  • Correlated features are effectively double-counted.
  • Usually outperformed by logistic regression or modern language models when there is enough data.

Practical Tip

Use Laplace smoothing (the alpha parameter) so a word never seen with a class during training doesn't force that class's probability to zero.

from sklearn.feature_extraction.text import CountVectorizer
from sklearn.naive_bayes import MultinomialNB
from sklearn.pipeline import make_pipeline
clf = make_pipeline(CountVectorizer(), MultinomialNB()).fit(texts, labels)

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026