mrkeyoor.com_
Tue 06 Oct 06:35 UTC
Dataevaluationupdated 06 Oct 2026

tax-doc-classifier review

tax-doc-classifier is a TypeScript library that sends extracted page text to Jev and returns an IRS form ID, page kind, probability distribution, and confidence gate. It helps tax-document pipelines route pages without training or hosting a custom classifier, but it covers federal English text pages rather than tax preparation itself.

Verdict

Our tax-doc-classifier install stopped after 18 seconds on unapproved esbuild scripts, so the repository did not clear the first setup gate despite its small 0.2 MB checkout. The design is worth testing if you process English federal forms and can keep low-confidence pages with a human. Do not rely on the headline accuracy claim or automate tax decisions until your own corpus reproduces the gate.

We ran it

Lab card: what happened when we ran tax-doc-classifierScreenshot of tax-doc-classifier (github.com/kyotofin/tax-doc-classifier)
Install✗ · 18s
Build—
Repo19 files~698 lines of source · 0.2 MB · 0 CI workflows

Answers from our run

Does tax-doc-classifier build from source?

The dependency install failed, and the project has no separate build step. We cloned commit 3e95a77 into a clean Debian container with 3 CPUs and no project-specific setup.

Who should not use tax-doc-classifier?

Scanned-document workflows without OCR: a page with no text layer is reported as blank.

What are the alternatives to tax-doc-classifier?

Docling, Unstructured, docTR. Our tax-doc-classifier install stopped after 18 seconds on unapproved esbuild scripts, so the repository did not clear the first setup gate despite its small 0.

Setup2/5Install failed, and PDF use needs poppler plus an API key
Docs4/5Detailed data, evaluation, IDs, limits, and fallback behavior
Community2/5497 stars and 2 open items in a repository weeks old
Maturity2/5Focused implementation, but install failed and claims need clarity

Who it’s for

Tax software teams already using Jev and processing born-digital federal forms.
Document pipelines that need a form ID and an explicit low-confidence review path.
Developers willing to regenerate criteria when the IRS publishes revised forms.
Teams that can validate the classifier against their own client documents before automation.

Who it’s NOT for

Scanned-document workflows without OCR: a page with no text layer is reported as blank.
State-tax identification: state forms can be routed as a kind but are not identified by form.
Non-English tax documents, which the README lists as unsupported.
Teams without poppler or a TypeSafe API key when using the PDF helpers and shipped backend.
Buyers relying on the repository's 100% strict-accuracy description: open issue 3 notes that the published 261-form corpus has 38 strict errors.
Compliance systems that would act on a model score without review: the README requires calibration on your own pages and a fallback below the gate.

Setup reality

Our sandbox cloned commit 3e95a77 and the pnpm install failed after 18 seconds. The log reported ERR_PNPM_IGNORED_BUILDS because build scripts for esbuild 0.21.5 and 0.28.2 were not approved, and it directed the user to pnpm approve-builds. We did not reach build or tests.

The documented path requires Node 20 or newer, pnpm, poppler's pdftotext and pdfinfo for PDF input, and a TypeSafe API key. Calling classifyPage with text lines avoids poppler, but the shipped jevBackend still needs the hosted service.

The checkout had 19 files, about 698 source lines, and occupied 0.2 MB. Our scan found no CI workflows, Dockerfile, or tests directory. The README's plain pnpm install instructions do not mention the build-script approval that stopped our fresh container.

One Jev call maps page text to federal form and page kind

tax-doc-classifier turns the text from one page into two bounded decisions. Jev chooses among page kinds and federal form candidates, then the library returns the winning form ID, probability distribution, confidence, and a boolean gate. Blank pages are handled without a request. Five complex form families can trigger a second, smaller question to distinguish a parent form from its schedules.

The catalog covers 261 IRS forms and 7 page kinds. The first form question uses 230 options plus a catch-all, while schedules for forms such as 5471 and 8865 are resolved hierarchically. IDs follow IRS Modernized e-File naming, so Form 1040 Schedule A becomes form-1040-schedule-a. That predictable grammar is useful in storage, routing, and downstream rules.

The 0.95 gate is the product, not a promise of correctness

A result is gated only when form confidence reaches 0.95. The README tells callers to act above the gate and use their existing fallback below it. Confidence is the minimum across classification steps, so a shaky schedule decision cannot hide behind a strong parent-form score. That is the right shape for document operations where an uncertain page should reach a person rather than silently enter the wrong queue.

Calibration still belongs to the adopter. The included corpus does not prove behavior on photographed pages, client annotations, damaged scans, new IRS revisions, or a firm's own preprocessing. The README says to calibrate on your pages before choosing a threshold. This library identifies documents. It does not interpret tax law, validate a return, or justify an automatic filing decision.

What happened when we ran it

Our sandbox cloned commit 3e95a77, a 0.2 MB checkout with 19 files and about 698 lines of source. The pnpm install failed with exit code 1 after 18 seconds. Its final output reported ERR_PNPM_IGNORED_BUILDS: scripts for esbuild 0.21.5 and 0.28.2 had been ignored, and pnpm directed the user to run pnpm approve-builds to select allowed dependencies.

The log also said the lockfile passed its supply-chain policies. It does not tell us whether approving those scripts would have produced a successful install, so we did not treat that command as a proven fix. No build or test result followed. Our scan found no CI workflow files, no Dockerfile, and no tests directory. The README's setup section says pnpm install, but does not mention this approval step.

PDF input adds poppler and hosted Jev to the tiny library

The TypeScript package itself has one runtime dependency, but the working pipeline is larger. Node 20 or newer and pnpm install the project. The PDF helpers call poppler's pdftotext and pdfinfo, so those binaries must be on the host. A TypeSafe API key in TYPESAFE_API_KEY powers the only backend implementation shipped. Passing lines directly to classifyPage removes poppler, not the hosted decision call.

The criteria file is generated from the IRS accepted-forms spreadsheet and individual PDFs. Its builder collects printed labels, titles, page counts, parents, selected box labels, and confusing sibling forms. That file must be regenerated when the IRS publishes revisions. The repository includes Apache-2.0 code and a separate data-license document, which is worth checking before redistributing the generated criteria.

The published evaluation supports a narrower claim than the description

The README reports two author-run corpora. On 314 filled pages covering 15 forms, it lists 0 wrong answers and 0 strict errors. On 753 blank-form pages covering 261 forms, it lists 0 wrong answers and 38 strict errors because confidence fell below 0.95. The author says those low-confidence cases were instruction pages, deep corporate-form pages, and schedules confused with parents. We did not rerun either evaluation.

Open issue 3 identifies the resulting wording problem. GitHub's repository description says 100% strict accuracy across 261 IRS forms, while the 261-form table reports a 5.05% strict-error rate. The issue suggests that form coverage and strict page accuracy may have been conflated. Until the maintainer clarifies it, cite the two corpus tables separately and avoid repeating the headline as one result.

Three unsupported page kinds limit real intake coverage

The parser is text-only and English-only. Scans need OCR first, otherwise a page without text is reported as blank. Federal forms can be identified, but state forms only receive the broad state_tax_form kind. The README also says state tax forms, broker or bank statements, and letters are defined as kinds but have not yet been evaluated. A mixed client upload will therefore travel beyond the demonstrated corpus quickly.

Docling or Unstructured can provide a broader ingestion stage, while docTR can supply OCR for image-only pages. tax-doc-classifier can sit after that stage when federal form identity is the question. Keep the original page, extracted text, returned distribution, model version, and review outcome together so a changed criterion file or threshold can be audited later.

Recent fixes help, but there is no release line yet

GitHub showed 497 stars, 2 combined open issues and pull requests, an Apache-2.0 license, and a last push on September 29, 2026. The repository was created September 18 and has no published release. A closed issue had found that the documented package import and GitHub-install paths failed. The current package manifest includes the requested export and prepare script. An open pull request addresses script-path resolution.

This is promising, unusually focused code with honest limits in the README, but it is still early. Our failed 18-second install is an immediate adoption cost, and the unresolved accuracy wording weakens the easiest marketing claim. Fix the pnpm approval path, reproduce the evaluation on your documents, and require human review below a locally chosen gate before this belongs in a tax workflow.

Alternatives

ProjectWhat it isPick it when
Docling gh↗A document conversion toolkit that extracts structured content from many file types.pick this instead when parsing and structure recovery matter more than identifying a specific IRS form.
UnstructuredAn ETL toolkit for partitioning complex documents into clean structured elements.pick this instead when you need a broad ingestion layer before building your own domain classifier.
docTRA document-text recognition library for OCR pipelines.pick this instead when scanned pages lack the text layer this classifier requires.

What people are saying

  1. [velocity-scout] kyotofin/tax-doc-classifier

Sources

  1. tax-doc-classifier README
  2. tax-doc-classifier repository facts
  3. Strict-accuracy wording issue
  4. Package installation issue and fix

More data reviews

awesome-jev · awesome-jev · awesome-jev · lead · prophet · awesome-zhuiju-free · the whole board →