Parquet is a columnar file format designed for analytics. It has become a standard for data lakes and large datasets.
Columnar Storage
CSV stores data row by row. Parquet stores it column by column. Analytical queries usually read a few columns across many rows, so a columnar layout lets tools read only what's needed.
Advantages
- Smaller files: columns of similar values compress very well.
- Faster queries: read only the columns and row groups you need.
- Types preserved: integers, decimals, dates and timestamps keep their types — no guessing.
- Schema included: column names and types travel with the file.
- Wide support: pandas, Polars, DuckDB, Spark, cloud warehouses and many BI tools read it directly.
Disadvantages
- Not human-readable in a text editor.
- Less convenient for small, frequently edited files.
- Appending individual rows is inefficient; it suits write-once, read-many data.
Working With Parquet
import pandas as pd
df = pd.read_parquet("trips.parquet", columns=["pickup_time", "fare_amount"])
df.to_parquet("clean.parquet", index=False)
DuckDB can query Parquet files with SQL directly, without loading them into memory first.
Partitioning
Large datasets are often split into many Parquet files partitioned by a column such as date, so queries can skip irrelevant partitions.
When to Choose It
Use Parquet for analytical datasets beyond a few megabytes, for data shared between tools, and for anything queried repeatedly. Keep CSV for small files meant for people to open and edit.