glitchdata Becomes a Hub for Datasets and Models — Without Storing Any of Them
By glitchdata Team 3 min read
Over the past month glitchdata has changed shape. What started as a learning portal now leads with a hub for datasets and models, organised the way anyone who has used Hugging Face will recognise — with one deliberate difference: we don't keep the files.
One Page per Dataset or Model
Each dataset and model lives at an address made of its owner and its name, such as glitchdata/wine-quality or glitchdata/whisper-tiny. Its page has three tabs. The card explains what it is, where it came from, how to use it, its licence and how to cite it. The data viewer shows the first rows of each CSV file, so you can see the columns before you download anything. And files lists every file with its size and the host it comes from.
You can browse by tag, licence, format or task, sort by downloads, likes or recent updates, and like the ones you find useful. There is also a small read-only JSON API for scripts.
References, Not Copies
Every file on glitchdata is a link to where its publisher hosts it — the UCI repository, NASA, the World Bank, GitHub or Hugging Face. When you download, you are sent straight to that host; we count the download on the way and nothing more. The upload feature that briefly existed during development has been removed from the code entirely, so there is no way for a file to end up stored here.
That keeps the data where its owner maintains it, keeps licences and versions with the source, and means glitchdata never becomes an unofficial mirror of someone else's work.
Real Data From Day One
The hub starts with twelve open datasets — Iris, Palmer penguins, wine quality, Adult census income, OWID CO₂ emissions, MovieLens, a month of NYC yellow taxi trips, Natural Earth boundaries, NASA's GISTEMP record, the live USGS earthquake feed, World Bank population and Capital Bikeshare — and four open models: all-MiniLM-L6-v2, a DistilBERT sentiment classifier, Whisper tiny and ResNet-50. Row counts and file sizes were measured from the files themselves, and each card carries the real licence and citation.
Anyone Can Share — With Review
Signed-in members can add their own datasets and models under their username, or under an organisation they belong to. Every submission is held until an admin approves it, and so is every later change to a live item, so an approved dataset can't have its links quietly swapped. If an admin asks for changes, the owner sees the note on their dashboard.
Why It Matters
Finding trustworthy data is mostly about context: who published it, under what licence, what's actually in the file. A catalogue that documents and previews data — and then points you to the source — gives you that context without adding another copy to keep in sync. See the Help page for a walkthrough.