The 145-paper boundary is stricter than an ordinary reading list
This repository asks a specific question: does a reasoning procedure survive a change in wording, symbols, composition, length, difficulty, domain, language, modality, environment, tool, or knowledge? Its survey argues that held-out accuracy alone cannot answer that. An unseen item may still follow a familiar template, while a renamed symbol or one extra inference step can expose that the model learned the surface pattern.
The corpus contains 145 core studies, split into 4 pillars: training, inference, architecture, and analysis. Each paper needs an identifiable reasoning task, reference training exposure, test shift, and evidence of transfer or failure. That filter excludes generic benchmark gains and broad reliability papers where reasoning is only one task. It gives the collection a sharper use than a chronological feed of everything containing the word reasoning.
A 6,012-record screen leaves an audit trail
The draft survey says the maintainers screened 6,012 arXiv records, checked relevant ACL Anthology and OpenReview entries, and read every retained abstract. Full papers were inspected when the abstract did not identify the evaluation split. The final manifest contains 151 sources: 145 core studies plus 4 nearby surveys and 2 contextual comparisons. Exclusion records include reasons, which makes disputed scope choices visible.
That record is the project's strongest feature. A researcher can inspect paper_manifest.tsv, selection.tsv, excluded_papers.tsv, and the larger candidate table instead of trusting a polished diagram. The manifest carries titles, authors, abstracts, taxonomy labels, source URLs, and local PDF paths. Those machine-specific paths reduce portability, but they also reveal that the corpus was managed as working research material rather than assembled only for a README.
What happened when we ran it
Our sandbox installed commit 1438f72 in 12 seconds with Python 3.12 on Debian, 3 CPUs, 8 GB of RAM, no secrets, and no elevated privileges. It added 35 packages and occupied 37 MB. The build completed in 5 seconds. Pip-audit reported 0 known vulnerabilities in that environment.
There was no test script or target, so tests were skipped. The checkout held 16 files, roughly 707 lines of source, and 11.7 MB. It had no CI workflow files, Dockerfile, or tests directory. That is reasonable for a small research corpus, yet it leaves manifest consistency, link health, duplicate detection, and README generation without a visible automated release gate. The successful build does not prove that 145 paper classifications are correct.
Four pillars prevent unlike claims from collapsing together
Training papers change learned parameters through data or feedback. Inference papers keep base weights fixed and alter prompting, search, verification, tools, or memory. Architecture papers ask whether recurrence, positional design, modules, or structured state support transfer. Analysis separates observed behavior from proposed mechanisms and theoretical results. Assigning one primary label stops a hybrid paper from being counted in every favorable category.
The survey also distinguishes model, procedure, and system generalization. A calculator may let a tool-using system solve longer problems without teaching the base model a length-general algorithm. A verifier can improve selection without making the generator stronger. This vocabulary is useful in an engineering review because it forces a claim to name what changed and what stayed fixed. It can prevent a system result from being advertised as a new model capability.
The corpus is current through September 10, not continuously verified
The README says the literature search was updated on September 10, 2026. Of the 145 core papers, 86 are from 2026, 39 from 2025, 14 from 2024, and 6 are earlier foundations. GitHub records the same date for the last push and shows 154 stars with 0 open issues or pull requests. There is no tagged release.
Recency carries a cost. A large share of the corpus is recent arXiv work, and the repository does not claim that every result has passed peer review or independent replication. Paper metadata and links are not quality ratings. The list helps you find relevant experiments and compare their claimed shift boundaries. You still need to read methods, check baselines, and decide whether each test resembles the work you care about.
The missing license and draft citation limit reuse
GitHub reports no repository license, so copying the survey text, scripts, or tabular corpus into another product needs permission or a separate rights analysis. The README supplies a temporary BibTeX entry for citing the living repository, but says the paper citation will arrive when the survey is released. That is a clear sign that the scholarly artifact is still in progress.
Use this collection as an orientation map and an audit seed. The 4-pillar taxonomy makes scattered reasoning claims easier to compare, while the 6,012-record screening trail gives future reviewers somewhere concrete to disagree. Its limits are just as concrete: 0 tests in our run, 0 CI workflows, no license, and no released survey citation. Those facts should stay attached whenever the corpus is reused.

