mrkeyoor.com_
Mon 28 Sept 06:38 UTC
Dataevaluationupdated 28 Sept 2026

opendataloader-pdf review

OpenDataLoader PDF turns PDFs into Markdown, HTML, or structured JSON with page positions, and it can write structure tags back into untagged PDFs. Its local path handles ordinary digital files, while an optional local hybrid service adds OCR and model-assisted handling for harder pages.

Verdict

Our commit d523d31 checkout installed 0 npm packages in 10 seconds, then exposed no build or test target, so the clean audit result does not prove the parser itself worked. OpenDataLoader PDF is worth a trial when local PDF structure, coordinates, and Tagged PDF output belong in one pipeline. Do not treat the open-source auto-tagging step as finished PDF/UA compliance, and test your own headings, scans, and extracted images before adopting it.

We ran it

Lab card: what happened when we ran opendataloader-pdfScreenshot of opendataloader-pdf (opendataloader.org)
Install✓ · 10s0 packages · 66 MB
Buildn/ano build script
Testsn/ano test script
Known vulns00 critical · 0 high · 0 moderate · 0 low (npm audit)
Repo335 files~48,515 lines of source · 36.7 MB · 6 CI workflows

Answers from our run

Does opendataloader-pdf build from source?

Dependencies installed in 10 seconds (0 packages), and the project has no separate build step. We cloned commit d523d31 into a clean Debian container with 3 CPUs and no project-specific setup.

Does opendataloader-pdf have tests you can run?

Not through a standard command: the project exposes no test script or target that our harness could run.

Does opendataloader-pdf have known vulnerabilities in its dependencies?

npm audit found none in the dependency tree at the time of our run.

Who should not use opendataloader-pdf?

Teams that need certified PDF/UA output entirely from the open-source package: PDF/UA-1 and PDF/UA-2 export are enterprise features.

What are the alternatives to opendataloader-pdf?

Docling, Marker, Unstructured. Our commit d523d31 checkout installed 0 npm packages in 10 seconds, then exposed no build or test target, so the clean audit result does not prove the parser itself worked.

Setup3/5Quick wrapper install, but Java and hybrid services add work
Docs4/5Modes, SDKs, limits, and enterprise boundaries are stated
Community4/529,386 stars with September 2026 issue and PR activity
Maturity3/5Frequent releases, with open output-correctness reports

Who it’s for

RAG teams that need reading order, element types, and bounding boxes for citations.
Developers who want PDF extraction to stay on their own machine.
Accessibility teams that need an open-source first pass from an untagged file to a Tagged PDF.
Python, Node.js, or Java shops willing to batch conversions around a JVM-backed tool.

Who it’s NOT for

Teams that need certified PDF/UA output entirely from the open-source package: PDF/UA-1 and PDF/UA-2 export are enterprise features.
Pipelines centered on DOCX, XLSX, or PPTX: the capability table says those formats are unsupported, while a broader Hancom integration is described as coming soon.
Windows users whose profile paths contain non-ASCII characters and need the hybrid backend today: open issue 725 reports an HTTP 500 in that case.
Workflows that must preserve embedded images at their native resolution: open issue 562 says extraction rasterizes pages at 144 DPI.
Accessibility teams that cannot inspect heading order: issue 751 documents skipped heading levels in Tagged PDF output.

Setup reality

Our sandbox checkout at commit d523d31 contained 335 files, about 48,515 source lines, and 36.7 MB. The npm install step succeeded in 10 seconds, installed 0 packages, and left 66 MB on disk. There was no build script or test target, so both steps were skipped; npm audit reported 0 known vulnerabilities.

The documented user path is not a typical Node app. Python use requires Python 3.10 or newer plus Java 11 or newer, and Node and Java SDKs are also available. Standard local conversion needs no service or secret. Hybrid work adds a separate local backend and extra model dependencies.

Each Python or Node conversion starts a JVM process, so the README tells users to batch files in one call. OCR, formulas, and image descriptions require hybrid options. Picture description can serialize other requests, and the server needs file-size and image-area limits if untrusted callers can reach it.

One parser covers extraction and a first accessibility pass

OpenDataLoader PDF does two related jobs that usually live in separate tools. It extracts Markdown, HTML, text, annotated PDF, or JSON with page numbers and bounding boxes. It can also add a structure tree to an untagged PDF, producing a Tagged PDF for screen readers. Python, Node.js, and Java SDKs wrap the same core, and the Apache-2.0 license covers the extraction and auto-tagging code.

That pairing is the reason to consider it. A RAG pipeline can keep element coordinates for citations, while an accessibility workflow can reuse the detected headings, lists, tables, and reading order. The 36.7 MB checkout contained about 48,515 source lines across 335 files, so this is a substantial parser rather than a thin API client. Ordinary digital PDFs stay local and do not require a hosted account.

Local mode is simpler than hybrid mode

The default path is aimed at digital PDFs. It uses deterministic layout analysis and exposes structure through several output formats. If the source already has good tags, use_struct_tree reads the author's structure instead. The README warns that poor source tags can produce worse output than heuristic or hybrid parsing, which means the presence of a tag tree is not proof that you should trust it.

Hybrid mode adds a separate local service for OCR, complex or borderless tables, formulas, and picture descriptions. It can route difficult pages through a model while keeping the files on your machine. That privacy story is useful, but the operational shape changes: you install extra Python dependencies, start the backend, and choose client and server flags that must agree. Combining use_struct_tree with hybrid does not combine their results because the structure tree takes precedence.

What happened when we ran it

Our sandbox cloned commit d523d31 with 3 CPUs and 8 GB of RAM. The npm install step succeeded in 10 seconds, installed 0 packages, and left 66 MB on disk. Npm audit found 0 known vulnerabilities. The repository had 6 CI workflow files, no Dockerfile, and no tests directory.

There was no build script or target, so we skipped the build. There was also no test script or target, so we could not run a project test suite. Those are findings about the checked-out repository under the detected npm path. They do not demonstrate conversion quality, OCR accuracy, or Tagged PDF validity, and we will not turn an empty dependency install into a claim that the Java and Python product ran.

That distinction matters here because the documented quick start requires Python 3.10 or newer and Java 11 or newer. The Node package is a wrapper around a JVM-backed conversion path, and the README says each conversion invocation starts a JVM process. Batch several files in one call rather than starting a new process for each PDF. Hybrid mode adds its own server and model downloads beyond the 10-second install our sandbox measured.

Tagged PDF output stops before open-source PDF/UA export

The free path can analyze an untagged file and write headings, paragraphs, lists, tables, and reading order into a Tagged PDF. Full PDF/UA-1 or PDF/UA-2 export and the visual editing studio sit in the enterprise tier. A procurement or compliance team should keep that line visible: a Tagged PDF is useful work, but the project does not promise that the free output completes the whole PDF/UA workflow.

Open issue 751 gives a concrete reason to inspect the result. The reporter found heading sequences that skipped levels and caused a PDF/UA-1 validation rule to fail on 8 documents in a 135-file sample. Issue 755 reports flattened nested lists, headings classified as list items, and separate structural lines merged into paragraphs in version 2.5.11. These user reports concern the same document structure that makes the project attractive. Our sandbox did not reproduce or test either case.

Images and Windows paths have specific failure cases

Image extraction does not always preserve the original pixels. Issue 562 says the current path rasterizes the page at 144 DPI before cropping detected image regions, which can shrink detailed figures and labels. If your retrieval system needs the text around a chart, that may be acceptable. If it archives figures or feeds them into another vision model, compare the extracted dimensions with the embedded source.

Windows hybrid users have another sharp edge. Issue 725 reports an HTTP 500 when a user profile path contains non-ASCII characters, with the log ending in an illegal character-sequence filesystem error. The same report says explicitly selecting a different Docling backend handled the file. That is narrow, reproducible evidence rather than a claim that Windows support as a whole is broken.

September activity is high, and so is the review burden

GitHub showed 29,386 stars, 96 open issues and pull requests, and a last push on September 28, 2026. Release v2.5.11 arrived on September 22 with logging and backend page-batch configuration changes. Recent issues and pull requests cover heading structure, scanned-page routing, table headers, image resolution, and Windows paths. The project is plainly being worked on, while the open queue also shows how many document shapes remain contested.

OpenDataLoader PDF makes the most sense when you can test against a representative PDF set and keep a human or validator in the accessibility loop. Its combination of local extraction, coordinates, and Tagged PDF writing is unusual enough to justify that work. If you mainly ingest office formats, choose a broader converter. If PDF/UA conformance is the deliverable, budget for validation and understand where the enterprise boundary begins.

Alternatives

ProjectWhat it isPick it when
Docling gh↗A document conversion toolkit that covers PDFs and several office and web formats.pick this instead when input-format breadth matters more than writing Tagged PDFs.
MarkerA focused converter for turning PDFs into Markdown and JSON.pick this instead when Markdown conversion is the job and PDF accessibility output is not.
UnstructuredA document ETL toolkit for partitioning many file types into structured elements.pick this instead when a mixed-format ingestion pipeline matters more than PDF-specific tagging.

What people are saying

  1. [velocity-scout] opendataloader-project/opendataloader-pdf

Sources

  1. OpenDataLoader PDF README
  2. OpenDataLoader PDF v2.5.11 release
  3. Issue 751: skipped heading levels in tagged output
  4. Issue 562: extracted image resolution
  5. Issue 725: Windows non-ASCII profile paths

More data reviews

polyledger · timeseries-atlas · data-formulator · toasty · gfwlist · simdjson · the whole board →