mrkeyoor.com_
Tue 01 Sept 17:45 UTC
Dataevaluationupdated 28 Aug 2026

awesome-public-datasets review

Awesome Public Datasets is an English-language directory of public data sources, arranged by subject and published as a generated README. It helps researchers and developers find a promising source before they spend time searching individual agencies, archives, and dataset hosts.

+55 / 3dstars / 7d
Verdict

Our 2026-08-27 sandbox did not run commit f74e3e5 because it had no supported language ecosystem or Dockerfile, confirming that Awesome Public Datasets is a reading list rather than software to install. Use it to build a shortlist across unfamiliar subjects, then verify the chosen source's URL, terms, and data quality yourself. Skip it when your workflow needs stable schemas, mirrored files, or a programmatic catalog.

We ran it

Screenshot of awesome-public-datasets (awesomedataworld.slack.com)

Answers from our run

Did you run awesome-public-datasets yourself?

No. GitHub reports no primary language for it, and it carries no manifest our lab installs from, and no Dockerfile, so there was nothing standard to install, build or test. This review is written from the repository's own documentation.

Who should not use awesome-public-datasets?

Teams that need a ready-to-import dataset bundle: commit f74e3e5 was mainly a generated README, with only a Titanic archive stored alongside it.

What are the alternatives to awesome-public-datasets?

Data Package Core Datasets, Google Dataset Search, Kaggle Datasets. Use it to build a shortlist across unfamiliar subjects, then verify the chosen source's URL, terms, and data quality yourself.

Setup4/5No install path; browsing starts immediately
Docs4/5Clear scope and contribution route, sparse entry-level detail
Community4/578,689 stars and current pushes, with a large open queue
Maturity4/5A long-running list whose external links still decay

Discussed on

  1. hnAwesome-public-datasets: A topic-centric list of open datasets5 points
  2. hnAwesomedata/awesome-public-datasets: A topic-centric list of HQ open datasets4 points
  3. hnA topic-centric list of HQ open datasets3 points

Who it’s for

Researchers who need a broad first pass across subjects such as biology, government, finance, climate, and transport.
Developers looking for public data sources to evaluate before building an import pipeline.
Teachers and students who want examples beyond the usual classroom datasets.
Maintainers willing to submit or repair entries through the separate apd-core repository.

Who it’s NOT for

Teams that need a ready-to-import dataset bundle: commit f74e3e5 was mainly a generated README, with only a Titanic archive stored alongside it.
Buyers who require every source to be free and immediately downloadable: the README says some listed datasets are not free, and its status markers include entries that need repair.
Security-sensitive researchers who cannot inspect every outbound link: open issue 526 reports that a Visual Genome URL led to casino spam, while issue 515 reports an inaccessible computer-networks source.
Contributors who expect to edit this repository directly: the README says it is generated by apd-core and tells people not to modify the file by hand.

Setup reality

Our 3-CPU, 8 GB Debian sandbox did not run commit f74e3e5. The checkout declared no supported language ecosystem and had no Dockerfile, so there was no install, build, or test command for the lab to execute. This is a directory to read, not an application to launch.

Browsing the list needs no project credentials, services, or configuration. Access rules belong to each linked dataset, and the README warns that some entries are not free. A useful source still needs its own license, format, login, and availability check.

The main platform wrinkle is editorial: the README is generated from awesomedata/apd-core. Contributors work on YAML there, where the guide documents local Python-based validation, rather than changing this repository's README directly.

The repository is a directory, not a dataset package

commit f74e3e5 contains 3 files in its Git tree: a generated README, an MIT license, and one bundled Titanic CSV archive. The README collects links under subjects that range from agriculture and biology to government, sports, and transport. It is written in English, with a table of contents and short descriptions beside many entries. For someone entering an unfamiliar field, that beats guessing which agency, university lab, or archive owns the useful source.

The listed climate records, genomics collections, transit feeds, and other resources live elsewhere. Cloning the repository therefore gives you the catalog, not a local copy of the cataloged data. That distinction matters before anyone designs a pipeline around it. The single Titanic archive is an exception inside the snapshot, not evidence that the linked collections have been mirrored for stable or offline access.

What happened when we ran it

Our 3-CPU, 8 GB sandbox did not run Awesome Public Datasets. The lab found no supported language ecosystem and no Dockerfile at commit f74e3e5, so it had no install, build, or test command to execute. There is no failed package manager or test log to diagnose. The correct result is simpler: this repository is published reference material rather than runnable software.

The sandbox was a fresh unprivileged Debian container with no secrets. Its 3 CPUs and 8 GB of RAM were available, but there was no application path to exercise. We therefore have no install result, build result, or test result to report, and opening the generated README would not count as a successful run.

The 237,438-byte README favors browsing over filtering

The 237,438-byte README at commit f74e3e5 is useful when you know the domain but not the source. A transport researcher can scan public transit, flight, traffic, and bike-share entries in one place; a biology student can move among genome, protein, microscopy, and taxonomy resources without inventing a dozen search queries. The short descriptions sometimes include a format, date span, or scale, which helps remove obvious mismatches before visiting the host.

The document offers no query language, schema filter, license filter, or machine-oriented catalog endpoint. Search is browser search over one large RST file. That is fine for a human making a shortlist, but awkward for a service that must find all CSV sources with a certain license or geographic scope. Google Dataset Search is better for open-ended retrieval, while Data Package Core Datasets is more useful when consistent packaging matters.

Two open reports show why every destination needs checking

Two 2026 issue reports capture the maintenance burden of any large link collection. Issue 526 says the listed Visual Genome API URL had become casino spam and proposes a different university-hosted address. Issue 515 identifies a computer-networks URL that no longer provides access. These specific reports do not establish that the entire directory is bad. They show why an OK marker cannot replace opening the link yourself.

The status legend distinguishes entries considered healthy from entries marked for repair. That is useful editorial signaling, yet it does not tell you when a source was last checked, what was tested, or whether its terms changed. Before using a dataset, confirm the publisher identity, download path, license, update date, documentation, and a small sample. For sensitive work, check redirects and domains before downloading anything, especially when an issue already describes a hijacked destination.

Contributions go through apd-core and three required fields

The apd-core guide requires 3 YAML fields for a new entry: title, homepage, and category. The README warns contributors not to edit it directly because apd-core generates the file. Proposed entries pass automatic validation and maintainer review. That separation keeps the published index consistent, though it is easy to miss if someone lands on the familiar GitHub edit button.

The stated acceptance policy prefers material that is legally available, useful to a specific domain, directly downloadable, and free of advertising or reputation promotion. The same guide says outdated entries may be removed when no correction arrives. Those are sensible standards. Open broken-link reports show that enforcing them across outside websites remains ongoing work, so contributors who repair an existing source may be as useful as people adding another one.

A 2026 push matters more than the 2023 release tag

GitHub recorded 78,689 stars, 11,815 forks, and 159 combined open issues and pull requests on August 28, 2026. A search of that queue separated 82 open issues from the remaining pull requests. The latest repository push also landed on August 28 and updated the generated README. That activity says the catalog is still being regenerated and discussed, even though the newest tagged release is v20230903 from September 2023.

The queue cuts both ways. Current dataset suggestions and link reports show that people still use the project, while dozens of open items mean a report or submission may wait. The generated README can keep changing without a conventional software release because it is the product. Judge freshness by the push history, the source metadata in apd-core, and the response to concrete link reports, rather than treating the old tag alone as an abandonment signal.

Use it for a shortlist, then inspect the source

Awesome Public Datasets earns a bookmark because it can expose a useful institution or archive you did not know to search for. Its value ends at discovery. commit f74e3e5 offered no executable catalog service, and the 2026 broken-link reports show that outside destinations can decay or become unsafe. Pick two or three plausible entries, verify their owners and terms, download a sample, and record the exact source URL in your own project. If you need stable files or automated filtering, start with a packaged catalog instead.

Alternatives

ProjectWhat it isPick it when
Data Package Core DatasetsA collection of packaged datasets with machine-readable metadata and consistent files.pick this instead when you want data checked into repositories in a repeatable package format rather than a long directory of external sources.
Google Dataset SearchA search engine for dataset pages published across the web.pick this instead when keyword search and coverage matter more than a maintainer-curated GitHub list.
Kaggle DatasetsA hosted dataset catalog with downloads, notebooks, metadata, and user discussion.pick this instead when you want files and analysis tools on the same platform and can accept Kaggle account and licensing constraints.

What people are saying

  1. [velocity-scout] awesomedata/awesome-public-datasets

Sources

  1. Awesome Public Datasets README
  2. Awesome Public Datasets repository facts
  3. apd-core contribution guide
  4. Visual Genome link report
  5. Computer Networks broken-link report
  6. Latest tagged release v20230903

More data reviews

turso · TrackersListCollection · dash · getcontact-cli · awesome-zhuiju-free · iggy · the whole board →