Most machine learning data can be pictured as a table. Understanding its parts is the first step in any project.
Rows Are Examples
Each row is one example: a customer, a transaction, an image, a sentence.
Features Are Inputs
Features are the columns the model uses to make its prediction — age, account balance, pixel values, word counts. Good features are available at prediction time, relevant to the outcome and reliably measured.
The Label Is the Answer
The label (or target) is what the model predicts: whether the customer churned, the sale price, the species of a flower. In supervised learning, every training row needs a label.
Feature Engineering
Raw data rarely arrives in the best form. Feature engineering turns it into more useful inputs, for example:
- extracting the day of the week from a timestamp;
- turning a text category into numeric columns (one-hot encoding);
- computing ratios, such as debt to income;
- aggregating history, such as purchases in the last 30 days.
For tabular data, thoughtful features often improve results more than switching algorithms.
Choosing Labels Carefully
The label defines what the model optimises. "Was hired" is not the same as "would do the job well", and "was arrested" is not the same as "committed a crime". A convenient but flawed label bakes the flaw into every prediction.
Beware Leakage
A feature that is only known after the outcome — such as a cancellation date when predicting cancellations — makes the model look brilliant in testing and useless in practice.