Naive Bayes classifies examples using Bayes' theorem, combining how common each class is with how likely the observed features are under each class.
The "Naive" Assumption
It assumes every feature is independent of the others given the class. For text, that means treating each word as unrelated to the next — clearly untrue, but the simplification makes the model extremely fast and it often classifies well anyway.
Variants
- Multinomial Naive Bayes — word counts, the classic choice for text.
- Bernoulli Naive Bayes — presence or absence of words.
- Gaussian Naive Bayes — continuous features assumed to follow a normal distribution.
Strengths
- Trains and predicts very quickly, even on large vocabularies.
- Needs relatively little training data.
- A strong baseline for spam filtering, sentiment and topic classification.
Weaknesses
- Probability estimates are often poorly calibrated (too confident).
- Correlated features are effectively double-counted.
- Usually outperformed by logistic regression or modern language models when there is enough data.
Practical Tip
Use Laplace smoothing (the alpha parameter) so a word never seen with a class during training doesn't force that class's probability to zero.
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.naive_bayes import MultinomialNB
from sklearn.pipeline import make_pipeline
clf = make_pipeline(CountVectorizer(), MultinomialNB()).fit(texts, labels)