Lesson 1 of 4
The Palmer penguins problem
Frame the task and explore the data before modelling.
10 min 3-question quiz 3 guides to read next
The Palmer penguins dataset records 344 penguins of three species from three islands in Antarctica. Our task: predict the species from the bird's measurements. That is a three-class classification problem.
import pandas as pd
url = "https://raw.githubusercontent.com/allisonhorst/palmerpenguins/main/inst/extdata/penguins.csv"
penguins = pd.read_csv(url)
penguins["species"].value_counts()
penguins.groupby("species")[["bill_length_mm", "flipper_length_mm", "body_mass_g"]].mean()
You will see the classes are not balanced — there are many more Adélie than Chinstrap penguins — and that Gentoo penguins are clearly heavier with longer flippers. Adélie and Chinstrap are harder to tell apart, but their bills differ.
Before modelling, write down:
- What would a useless model score? Always guessing "Adélie" is right about 44% of the time. Any real model must beat that.
- Which features would be available in practice? Measurements, yes. The
yearcolumn tells you when the bird was measured — it shouldn't help predict species, and if it does, that is a warning sign.
Check your understanding
3 questions · pass with 2 correct
Enrol for free to save your progress, unlock every lesson and earn a certificate.
Sign in to enrolFurther reading
Guides that go deeper on this lesson.
-
Baseline Models: Why You Always Need One
A simple baseline turns a metric into a meaningful result. The baselines to use for common problem types.
1 min read
-
Multi-Class and Multi-Label Classification
The difference between choosing one class from many and assigning several labels at once, with methods and metrics for each.
1 min read
-
Exploratory Data Analysis Checklist
A step-by-step checklist for getting to know a new dataset before modelling or reporting.
1 min read