One Search Box, and a JSON API for Everything Public
By glitchdata Team 2 min read
Three things that should have shipped together finally have: a way to find things, a way to read them as data, and a way for other software to know they exist.
Search
Search covers live datasets and models, published guides, courses and news, from the box in the header. With no filter it shows the best few of each kind with counts; pick a kind and you get a paginated list of just those.
Every word has to match, so adding a word narrows the results — which is what people expect and not what a naive search does. Where a result's summary does not contain what you searched for, the snippet comes from the body, trimmed to the line the match is on.
The API
Everything public is also JSON, at /api. Datasets and models, courses with their lessons, guides with their full text, news posts, topics, and the same search as the site.
There is no key and nothing to agree to: 120 requests a minute, listings answer with data, total and next, and a single item comes back as a bare object. It enforces exactly the same rules as the pages do — drafts, submissions awaiting review and lessons that need enrolment are never returned, signed in or not. A free preview lesson includes its text; the rest do not.
Feeds, Sitemap and Structured Data
News and guides have RSS feeds, linked from every page's head. The sitemap is built from the database on request, so it lists exactly what a signed-out visitor can open and never drifts; there is a readable version for people, too.
Dataset pages carry schema.org data describing each file as a download, which is what Google Dataset Search reads — so a dataset catalogued here can be found by people who have never heard of us.