mrkeyoor.com_
Wed 30 Sept 00:12 UTC
AI Toolsevaluationupdated 26 Aug 2026

modlens review

ModLens gives text-only coding models a way to inspect images by sending them to a separate vision provider and returning structured evidence. It runs as a DeepSeek Harness plugin or a skill for Codex, Claude Code, Pi, and OpenCode, with English documentation and a separate Chinese README.

+41stars / 7d
Verdict

Our modlens run installed 120 packages, built in 5 seconds, and passed 818 of 823 tests with 5 skipped, so its repository mechanics are unusually easy to verify. Use it when a preferred text-only model needs occasional screenshots, charts, or documents and you can name a trusted vision provider. Skip it when images are highly sensitive, native attachment interception is mandatory, or Node 24 abort failover is part of the reliability plan.

We ran it

Lab card: what happened when we ran modlensScreenshot of modlens (liustack.dev)
Install✓ · 12s120 packages · 132 MB
Build✓ · 5s
Tests✓ · 41s818 passed · 0 failed · 5 skipped of 823 (vitest)
Repo154 files~31,941 lines of source · 4.9 MB · 2 CI workflows

Answers from our run

Does modlens build from source?

Dependencies installed in 12 seconds (120 packages), and the build succeeded in 5 seconds. We cloned commit 0cd2b60 into a clean Debian container with 3 CPUs and no project-specific setup.

Do modlens's tests pass?

Yes: 818 of 823 passed when we ran the project's own test command (vitest). Some failures need services or credentials a bare container does not have.

Who should not use modlens?

Security teams that require an OS sandbox around untrusted images: the security guide says subprocess isolation reduces exposure but is not a security boundary.

What are the alternatives to modlens?

Codex, Claude Code, OpenCode. Our modlens run installed 120 packages, built in 5 seconds, and passed 818 of 823 tests with 5 skipped, so its repository mechanics are unusually easy to verify.

Setup4/512-second install and 818 passing tests; a provider is still required
Docs5/5Install, routing, output, security, and troubleshooting are explicit
Community3/53,700 stars and active issues, but one maintainer accepts no PRs
Maturity3/5v3.25.0 is active; Node 24 and Codex attachment bugs remain

Who it’s for

Developers using text-only DeepSeek, GLM, or MiMo models inside supported coding harnesses.
Teams that want OCR, layout, entities, and uncertainty returned as structured JSON.
Users with a Gemini, Anthropic, or OpenAI-compatible vision endpoint already available.
Claude Code and Codex users willing to grant explicit reuse of an existing signed-in CLI.

Who it’s NOT for

Security teams that require an OS sandbox around untrusted images: the security guide says subprocess isolation reduces exposure but is not a security boundary.
Codex users expecting every native pasted attachment to be intercepted: open issue 78 reports a text-only gateway rejecting the image before the skill can run.
Node 24 users relying on provider failover after an abort: open issue 85 reports that an AbortError can crash v3.25.0 instead of trying the next provider.
Contributors who require a normal pull-request route: the README says the single maintainer does not accept pull requests.
Organizations unwilling to send image content to an outside model provider or reuse local agent quotas.

Setup reality

Our sandbox installed 120 pnpm packages in 12 seconds and used 132 MB. The build passed in 5 seconds. Vitest finished in 41 seconds with 818 passed, 0 failed, and 5 skipped out of 823.

Actual image reading needs at least one vision route. Options include a Gemini key, Anthropic key, OpenAI-compatible endpoint, Antigravity CLI login, or explicit reuse of signed-in Codex, Claude Code, OpenCode, Pi, Kimi, or Grok tooling.

DeepSeek Harness has a plugin install path; other hosts use a skill and CLI. Remote-image handling differs by provider, and agent subprocesses are exposure reduction rather than an OS sandbox. Node 24 abort handling and native Codex attachment interception have open reports.

ModLens gives text-only agents a separate pair of eyes

ModLens connects a coding harness to a vision service. A user supplies an image, ModLens asks another engine to inspect it, and the text-only model receives transcription, reading order, entities, relations, and uncertainty. It suits developers who prefer DeepSeek, GLM, or MiMo text models but need to discuss screenshots or scanned text. It is not a new vision model.

The project has two integration styles. DeepSeek Harness gets a plugin with wrapped (modlens vision) model entries and a direct-paste path. Codex, Claude Code, Pi, and OpenCode use a skill plus the ModLens CLI. The README says native vision models are left alone when host metadata confirms their capability. That matters because a bridge should not add another network call when the selected model can already read the image.

Six provider routes still require a trust decision

The README lists 6 built-in provider routes, including Gemini, Anthropic, an OpenAI-compatible endpoint, and agent CLI paths. It can also reuse signed-in local tools after the user grants permission. API routes are tried before slower agent processes, and each attempt is recorded in output metadata. A pinned provider disables fallback, which is the safer choice when policy says exactly where an image may go.

There is no free privacy upgrade hidden in that flexibility. A Gemini, Anthropic, or OpenAI-compatible service receives the image or a URL to it. Reused CLIs can spend an existing subscription quota, which ModLens labels in warnings. For remote URLs, some vendors fetch the URL themselves, so ModLens cannot apply its private-address check, magic-byte validation, or 25 MB cap on those paths. Security review has to follow the selected provider, not the package name alone.

What happened when we ran it

Our sandbox cloned commit 0cd2b60 and installed 120 pnpm packages in 12 seconds. Dependencies occupied 132 MB, and the build passed in 5 seconds. Vitest then completed in 41 seconds with 818 passed, 0 failed, and 5 skipped out of 823. That is a convincing repository check for the Node 22 environment we used.

The checkout contained 154 files, about 31,941 lines of source, and 4.9 MB before dependencies. Our scan found 2 CI workflow files, no Dockerfile, and no top-level tests directory. Those facts describe package health, not OCR accuracy or provider speed. We did not benchmark image recognition, compare vision models, or verify the README's per-provider timing claims, so they do not enter our recommendation.

Private temp files do not make an OS sandbox

Local image recovery has several thoughtful controls. Recovered files use mode 0600 inside a newly created 0700 directory. Project identity is checked against the transcript, and subprocess providers receive a copy in a throwaway working directory rather than access to the original file's neighbors. Claude CLI is restricted to the Read tool. Credentials are scrubbed from errors, warnings, attempt records, and stored cooldown reasons before output.

The security guide also states the boundary plainly. An agent process can still read absolute paths, use the network, and start programs. Antigravity runs with --dangerously-skip-permissions because prompt mode does not work in some environments; its prompt asks it to read only the supplied image. Text inside a screenshot can contain hostile instructions, so that prompt is mitigation rather than isolation. For untrusted images, the guide recommends the direct Gemini API route instead of a local agent process.

Node 24 aborts can bypass the failover chain

Open issue 85 documents a concrete v3.25.0 failure on Node 24. When a provider call is aborted, ModLens tries to replace the message on an AbortError whose property is read-only. The resulting exception exits the CLI instead of recording the failed attempt and moving to another provider. This affects the feature that makes a multi-provider setup attractive, and the report includes a short timeout reproduction. Node 22 passed our 823-test run, so the two findings can coexist.

Codex has a different integration limit. Issue 78 reports that a native attachment sent through a text-only OpenAI-compatible gateway can be rejected by the host before the ModLens skill gets control. Supplying an absolute image path as text works in the reporter's setup. Anyone buying ModLens specifically for paste recovery should test the exact host, model route, and attachment form rather than treating the support matrix as one universal hook.

The output contract favors evidence over fake precision

Version 2 removed pixel coordinates and confidence scores because language models can fabricate both convincingly. The current contract records unreadable details as uncertainty and returns provider-attempt metadata. That is the right bias for agent work: a partial transcription with a named gap is more useful than an invented bounding box. Applications still need to decide whether missing evidence stops a task or merely prompts another read.

ModLens can keep later questions tied to an already processed image, which saves repeated pastes. Key rotation handles authentication, rate-limit, and quota failures, while other provider failures move through the broader chain. These are practical behaviors, but they also create audit work. A production caller should persist which provider answered, which warnings appeared, and whether uncertainty touches the field being acted on.

August releases are fast, and maintenance is concentrated

GitHub listed 3,700 stars, 5 combined issues and pull requests, and a last push on August 25, 2026. Release v3.25.0 was published the same day, while issue activity continued on August 26. Several installation and compatibility reports were closed during the preceding week. The repository is moving quickly rather than sitting idle.

That pace depends on one maintainer. The contribution guide says pull requests are not accepted; users can file issues or fork the MIT-licensed code. This can produce a consistent codebase, though an organization cannot assume its fix will enter upstream through the usual review route. The clean 818-test result makes ModLens worth trying. The open Node 24 and Codex attachment reports should be part of the acceptance test, not footnotes discovered after rollout.

Alternatives

ProjectWhat it isPick it when
Codex gh↗A coding agent with native image input on models that support vision.pick this instead when your chosen model already accepts images and you do not need a bridge for text-only routes.
Claude Code gh↗Anthropic's terminal coding agent with multimodal model support.pick this instead when direct image understanding inside one agent is enough and an Anthropic model is acceptable.
OpenCode gh↗An open coding-agent interface that can use different model providers.pick this instead when you can choose a vision-capable model at the host layer rather than translating images for a text-only model.

What people are saying

  1. [github-trending] liustack/modlens

Sources

  1. ModLens README
  2. ModLens security guide
  3. Node 24 AbortError issue
  4. Codex native attachment issue
  5. ModLens v3.25.0 release

More ai tools reviews

voltagent · InferenceX · Bonsai-demo · qwen-audio-agent · wechat-intelligence-hub · dlss5-visual-enhancer · the whole board →