Skip to content

Handling Text Data for Classical Machine Learning

Turning text into features with bag-of-words, TF-IDF and n-grams, and when embeddings are a better choice.

Editorial team 1 min read

Before deep learning, text was converted into numerical features for classical models. These methods remain fast, effective baselines.

Preprocessing

  • Lowercasing and removing punctuation.
  • Tokenisation into words.
  • Optionally removing common stop words.
  • Stemming or lemmatisation to group word forms.

Bag of Words

Count how often each word appears in each document. Simple but surprisingly effective.

TF-IDF

Weight words by how frequent they are in a document relative to how common they are across all documents. Distinctive words gain importance.

N-Grams

Include word pairs or triples ("not good", "credit card") to capture some word order.

Models

Logistic regression, linear SVMs and naive Bayes work well on these sparse, high-dimensional features.

Embeddings

Embeddings from pretrained language models capture meaning and synonyms better, and usually perform better with limited labelled data.

When to Use Which

  • Large labelled data, need for speed and interpretability → TF-IDF with a linear model.
  • Nuanced meaning, less data → embeddings or fine-tuned models.

Always compare against the simple baseline.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026