One Jev call maps page text to federal form and page kind
tax-doc-classifier turns the text from one page into two bounded decisions. Jev chooses among page kinds and federal form candidates, then the library returns the winning form ID, probability distribution, confidence, and a boolean gate. Blank pages are handled without a request. Five complex form families can trigger a second, smaller question to distinguish a parent form from its schedules.
The catalog covers 261 IRS forms and 7 page kinds. The first form question uses 230 options plus a catch-all, while schedules for forms such as 5471 and 8865 are resolved hierarchically. IDs follow IRS Modernized e-File naming, so Form 1040 Schedule A becomes form-1040-schedule-a. That predictable grammar is useful in storage, routing, and downstream rules.
The 0.95 gate is the product, not a promise of correctness
A result is gated only when form confidence reaches 0.95. The README tells callers to act above the gate and use their existing fallback below it. Confidence is the minimum across classification steps, so a shaky schedule decision cannot hide behind a strong parent-form score. That is the right shape for document operations where an uncertain page should reach a person rather than silently enter the wrong queue.
Calibration still belongs to the adopter. The included corpus does not prove behavior on photographed pages, client annotations, damaged scans, new IRS revisions, or a firm's own preprocessing. The README says to calibrate on your pages before choosing a threshold. This library identifies documents. It does not interpret tax law, validate a return, or justify an automatic filing decision.
What happened when we ran it
Our sandbox cloned commit 3e95a77, a 0.2 MB checkout with 19 files and about 698 lines of source. The pnpm install failed with exit code 1 after 18 seconds. Its final output reported ERR_PNPM_IGNORED_BUILDS: scripts for esbuild 0.21.5 and 0.28.2 had been ignored, and pnpm directed the user to run pnpm approve-builds to select allowed dependencies.
The log also said the lockfile passed its supply-chain policies. It does not tell us whether approving those scripts would have produced a successful install, so we did not treat that command as a proven fix. No build or test result followed. Our scan found no CI workflow files, no Dockerfile, and no tests directory. The README's setup section says pnpm install, but does not mention this approval step.
PDF input adds poppler and hosted Jev to the tiny library
The TypeScript package itself has one runtime dependency, but the working pipeline is larger. Node 20 or newer and pnpm install the project. The PDF helpers call poppler's pdftotext and pdfinfo, so those binaries must be on the host. A TypeSafe API key in TYPESAFE_API_KEY powers the only backend implementation shipped. Passing lines directly to classifyPage removes poppler, not the hosted decision call.
The criteria file is generated from the IRS accepted-forms spreadsheet and individual PDFs. Its builder collects printed labels, titles, page counts, parents, selected box labels, and confusing sibling forms. That file must be regenerated when the IRS publishes revisions. The repository includes Apache-2.0 code and a separate data-license document, which is worth checking before redistributing the generated criteria.
The published evaluation supports a narrower claim than the description
The README reports two author-run corpora. On 314 filled pages covering 15 forms, it lists 0 wrong answers and 0 strict errors. On 753 blank-form pages covering 261 forms, it lists 0 wrong answers and 38 strict errors because confidence fell below 0.95. The author says those low-confidence cases were instruction pages, deep corporate-form pages, and schedules confused with parents. We did not rerun either evaluation.
Open issue 3 identifies the resulting wording problem. GitHub's repository description says 100% strict accuracy across 261 IRS forms, while the 261-form table reports a 5.05% strict-error rate. The issue suggests that form coverage and strict page accuracy may have been conflated. Until the maintainer clarifies it, cite the two corpus tables separately and avoid repeating the headline as one result.
Three unsupported page kinds limit real intake coverage
The parser is text-only and English-only. Scans need OCR first, otherwise a page without text is reported as blank. Federal forms can be identified, but state forms only receive the broad state_tax_form kind. The README also says state tax forms, broker or bank statements, and letters are defined as kinds but have not yet been evaluated. A mixed client upload will therefore travel beyond the demonstrated corpus quickly.
Docling or Unstructured can provide a broader ingestion stage, while docTR can supply OCR for image-only pages. tax-doc-classifier can sit after that stage when federal form identity is the question. Keep the original page, extracted text, returned distribution, model version, and review outcome together so a changed criterion file or threshold can be audited later.
Recent fixes help, but there is no release line yet
GitHub showed 497 stars, 2 combined open issues and pull requests, an Apache-2.0 license, and a last push on September 29, 2026. The repository was created September 18 and has no published release. A closed issue had found that the documented package import and GitHub-install paths failed. The current package manifest includes the requested export and prepare script. An open pull request addresses script-path resolution.
This is promising, unusually focused code with honest limits in the README, but it is still early. Our failed 18-second install is an immediate adoption cost, and the unresolved accuracy wording weakens the easiest marketing claim. Fix the pnpm approval path, reproduce the evaluation on your documents, and require human review below a locally chosen gate before this belongs in a tax workflow.

