mrkeyoor.com_
Sat 03 Oct 15:38 UTC
AI Toolsevaluationupdated 03 Oct 2026

chandra review

Chandra is an OCR model and Python toolkit that turns images and PDFs into Markdown, HTML, or JSON while retaining page layout. It is aimed at documents that ordinary text extraction handles badly, including tables, handwriting, forms, diagrams, and mathematical notation.

Verdict

Our Chandra run used 5,829 MB for 123 packages, then its only test failed at the 120-second limit, so local adoption needs more proof than the 4-second build provides. Trial it if complex page layout is the problem and you already have GPU infrastructure, but score your own documents and examine every output. Skip it when the model license, runtime size, or a clean fresh-container test is a release requirement.

We ran it

Lab card: what happened when we ran chandraScreenshot of chandra (www.datalab.to)
Install✓ · 129s123 packages · 5829 MB
Build✓ · 4s
Tests✗ · 138s0 passed · 1 failed of 1 (pytest)
Known vulns0(pip-audit)
Repo53 files~1,985 lines of source · 10.3 MB · 2 CI workflows · tests dir

Answers from our run

Does chandra build from source?

Dependencies installed in 129 seconds (123 packages), and the build succeeded in 4 seconds. We cloned commit d4f7467 into a clean Debian container with 3 CPUs and no project-specific setup.

Do chandra's tests pass?

Not all of them: 0 of 1 passed and 1 failed when we ran the project's own test command (pytest). Some failures need services or credentials a bare container does not have.

Does chandra have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use chandra?

Commercial users above the model license's $2 million funding or revenue threshold, or services that compete with Datalab's API, unless they buy broader rights.

What are the alternatives to chandra?

olmOCR, MinerU, PaddleOCR. Our Chandra run used 5,829 MB for 123 packages, then its only test failed at the 120-second limit, so local adoption needs more proof than the 4-second build provides.

Setup2/5129-second install, 5,829 MB footprint, and timed-out test
Docs4/5Clear CLI and backends, but deployment detail is still thin
Community4/512,392 stars with September issue and pull request activity
Maturity3/5v0.2.0 has broad OCR scope, while our sole test timed out

Who it’s for

Document-processing teams that need layout blocks, reading order, and extracted images alongside text.
Python developers with NVIDIA GPU access who can run the recommended vLLM path.
Researchers comparing OCR across multilingual, mathematical, tabular, or handwritten documents.
Startups eligible under the model license that want to self-host open weights.

Who it’s NOT for

Commercial users above the model license's $2 million funding or revenue threshold, or services that compete with Datalab's API, unless they buy broader rights.
Small CPU-only deployments expecting a light package: our install used 5,829 MB before production model serving.
Teams that require a clean integration test in a fresh container: the only test failed at the 120-second timeout in our run.
Pipelines that accept .tif files without conversion: open issue 111 says that extension is rejected while .tiff is allowed.
Workflows that cannot review generated text: open issue 116 reports massive gibberish for one specific input combination.

Setup reality

Our sandbox installed commit d4f7467 in 129 seconds, pulling 123 packages and using 5,829 MB. The build succeeded in 4 seconds. The only pytest case failed after the 120-second pytest-timeout limit, leaving 0 passed and 1 failed; pip-audit found 0 known vulnerabilities.

The base package talks to a vLLM server, while local Hugging Face inference needs the hf extra and PyTorch. The recommended vLLM launcher starts Docker and expects NVIDIA GPU selection through configuration. Model weights must also be downloaded.

This is a small repository with 53 files, about 1,985 source lines, and a 10.3 MB checkout, but the runtime is not small. There is no repository Dockerfile even though the vLLM command launches a container, so inspect the packaged launcher and pin its image before production use.

Chandra 2 turns page structure into usable output

Chandra does more than read a line of text from an image. It converts a page into Markdown, HTML, or JSON and retains layout information for sections, tables, forms, images, equations, code, and other page blocks. That makes it relevant when reading order and structure matter as much as the words. The project also extracts images and can describe diagrams or charts, which moves it closer to document reconstruction than plain OCR.

The current Chandra 2 release uses a 4-billion-parameter model. Datalab publishes benchmark tables for 43 and 90 languages, plus results on the olmOCR set, but those figures come from the project. Our 3-CPU, 8 GB sandbox did not measure recognition accuracy, page throughput, or language quality. Buyers should build a private sample containing their worst scans, handwriting, tables, and scripts before treating a public average as a deployment forecast.

The base CLI expects vLLM, while local inference adds PyTorch

The quickest install is pip install chandra-ocr. By default, the CLI sends requests to a vLLM endpoint, and chandra_vllm starts a Docker container configured for NVIDIA GPUs. The alternative hf extra adds Hugging Face inference and PyTorch inside the local Python environment. Input can be one file or a directory; output includes Markdown, HTML, metadata JSON, and any extracted images.

Our source checkout was modest: 53 files, about 1,985 lines of source, and 10.3 MB. Installation changed the picture. The environment pulled 123 packages and occupied 5,829 MB. Model weights and production serving come after that total. A team choosing the Hugging Face path should budget for PyTorch and optional FlashAttention; a vLLM team needs Docker, compatible NVIDIA drivers, GPU capacity, and a pinned serving image.

What happened when we ran it

We cloned and ran commit d4f7467 in a fresh, unprivileged Python 3.12 container. Installation succeeded in 129 seconds. The build also succeeded in 4 seconds, and pip-audit reported 0 known vulnerabilities. Those results show that the package can be installed and built in the lab image. They do not show that Chandra can process a production document within an acceptable time or memory budget.

Pytest ran one integration case, test_inference_image, and it failed when pytest-timeout stopped it after 120 seconds. The whole test step took 138 seconds and finished with 0 passed and 1 failed. The log tail also listed several mtp weights as unexpected and explained that such weights can be ignored across different tasks or architectures, but not when identical architecture is expected. The log does not establish whether those weights caused the timeout, so we will not connect them.

A 4-second build does not validate OCR output

The build result is clean but narrow. Chandra's only test in our run attempted image inference, and that is precisely where the 120-second cap ended the suite. There were no passing unit tests around file ranges, output structure, table markup, or repeated-token handling in the supplied run. A production evaluation needs known-answer documents and checks for empty pages, duplicated text, missing regions, malformed markup, and output that expands far beyond the source.

Open issue 116 gives that caution a concrete shape: the reporter supplied a specific text combination that produced a large block of gibberish. Other open reports describe repeated tokens and no-response behavior. These are user reports, not outcomes from our sandbox. Still, the 1 failed test out of 1 means our run offers no counterweight. Put human review or automated document checks between Chandra and any archive, search index, or customer-facing export.

The model license draws a commercial boundary

The Python code uses Apache-2.0, but the model weights use a modified OpenRAIL-M license. The README says research, personal use, and startups under $2 million in funding or revenue may use the weights for free. It also bars competitive use against Datalab's API under those terms. Larger companies, competing services, and teams needing different rights must arrange a commercial license.

That distinction matters because the useful artifact is the model, not only the 1,985 lines in the repository. Legal review should cover the exact model revision and your product, especially if documents flow through a paid service. Datalab also says its managed platform runs an improved model rather than exactly the open weights, so API results should not be treated as proof of self-hosted output. Compare the route you intend to deploy.

Recent issue work continues after the June code push

GitHub showed 12,392 stars and 61 combined open issues and pull requests. The repository's last push was June 26, 2026, while an open pull request was updated on September 10. Release v0.2.0 arrived on March 18. That is not an abandoned project, but the public code has moved less recently than the issue queue. One open pull request proposes saving per-page layout chunks beside the transcript.

Chandra is worth a controlled trial when tables, forms, handwriting, or page geometry defeat simpler OCR. Keep the trial honest: the 129-second install and 5,829 MB environment were successful, while the one supplied integration test did not finish under its 120-second limit. The deciding evidence should come from your documents on your GPU, checked against expected text and layout. Public benchmark tables cannot make that decision for you.

Alternatives

ProjectWhat it isPick it when
olmOCRAllen AI's toolkit for converting PDFs into clean, ordered text at scale.pick this instead when PDF text extraction and a published processing pipeline matter more than Chandra's layout block output.
MinerU gh↗A document parser for turning PDFs into structured Markdown and JSON.pick this instead when you want a broader PDF parsing pipeline with its own layout and formula tooling.
PaddleOCR gh↗A mature OCR suite covering text recognition, layout analysis, and document parsing.pick this instead when deployment choices and established OCR components matter more than one vision-language model.

What people are saying

  1. [github-trending] datalab-to/chandra

Sources

  1. Chandra repository and README
  2. Chandra OCR 2 v0.2.0 release
  3. Model license
  4. TIF extension issue 111
  5. Gibberish generation issue 116

More ai tools reviews

production-agentic-rag-course · GPT-as-Policy · NeuralScreen · mural · recurrent-looped-tranformer · gpu-time · the whole board →