mrkeyoor.com_
Wed 30 Sept 20:35 UTC
LLM Toolsevaluationupdated 27 Aug 2026

deepseek-harness review

DeepSeek Harness is an open source coding-agent runtime with a local web interface, headless mode, model adapters, tools, sessions, approvals, sandboxing, and MCP client support. Its defining idea is that nearly every part of the agent loop is a replaceable Cordis plugin, so developers can change behavior through composition instead of forking a fixed core.

+4,837stars / 7d
Verdict

Our DeepSeek Harness run passed 14,583 of 14,707 tests, with 10 failures and 114 skips, so the large suite was close but not clean. Use it to study or build a replaceable agent runtime when the Cordis plugin model is the point. For daily coding or production automation, the explicit developer-preview warning and absence of a tagged release are enough reason to choose a more settled agent.

We ran it

Lab card: what happened when we ran deepseek-harnessScreenshot of deepseek-harness (deepseek.com/harness)
Install✓ · 22s936 packages · 1393 MB
Build✓ · 111s
Tests✗ · 829s14583 passed · 10 failed · 114 skipped of 14707 (vitest)
Repo7822 files~623,075 lines of source · 50.8 MB · 18 CI workflows

Answers from our run

Does deepseek-harness build from source?

Dependencies installed in 22 seconds (936 packages), and the build succeeded in 111 seconds. We cloned commit b150a55 into a clean Debian container with 3 CPUs and no project-specific setup.

Do deepseek-harness's tests pass?

Not all of them: 14583 of 14707 passed and 10 failed when we ran the project's own test command (vitest). Some failures need services or credentials a bare container does not have.

Who should not use deepseek-harness?

Production teams requiring stable compatibility: the README labels the project a developer preview and explicitly warns of breaking changes.

What are the alternatives to deepseek-harness?

OpenCode, Codex CLI, AutoGen. Our DeepSeek Harness run passed 14,583 of 14,707 tests, with 10 failures and 114 skips, so the large suite was close but not clean.

Setup2/5One-command trial, but source setup is large and tests failed
Docs5/5Deep architecture, subsystem, tutorial, and post-mortem coverage
Community3/5Very current development, but no issue queue or tagged release
Maturity2/5Explicit developer preview with promised compatibility breaks

Discussed on

  1. hnDeepSeek harness: what doors does config over code open?3 points

Who it’s for

Agent developers who need to replace model, tool, persistence, sandbox, or interface layers independently.
Teams researching an event-driven plugin architecture for long-running coding agents.
DeepSeek API users who want a local web agent with workspace tools and explicit approvals.
Contributors comfortable with a large TypeScript monorepo, generated contracts, bilingual docs, and specialized CI gates.

Who it’s NOT for

Production teams requiring stable compatibility: the README labels the project a developer preview and explicitly warns of breaking changes.
Users wanting a settled release channel: GitHub returned no latest release, so there is no current tagged version to pin from that endpoint.
Resource-constrained contributors: our source install added 936 packages and occupied 1,393 MB, while the test run lasted 829 seconds.
Operators assuming the sandbox covers every system boundary: its docs say sandbox modes govern filesystem effects, while network and process visibility are outside that vocabulary.
Windows users requiring complete filesystem confinement: the sandbox documentation describes the Windows ACL backend as partial enforcement.
Developers who want a simple agent loop rather than learning Cordis contexts, plugin trees, bundles, seams, event waterfalls, and generated Host and Client contracts.

Setup reality

At commit b150a55, pnpm installation succeeded in 22 seconds in our fresh Node 22 sandbox, adding 936 packages and using 1,393 MB on disk. The build succeeded in 111 seconds. Tests ran for 829 seconds and failed with exit code 1: 14,583 passed, 10 failed, and 114 skipped out of 14,707.

The checkout contained 7,822 files, about 623,075 source lines, and occupied 50.8 MB. The final test output shows 4 failed files, 859 passed files, and 9 skipped files. Its last failure location names a SQLite schema test about changed columns and non-strict owned tables, but the provided tail does not show enough detail to assign a cause.

For a user trial, npx @deepseek-ai/dsh web is the documented shortcut. A working session still needs a model credential, a selected workspace, and decisions about tool approval and sandbox mode. Source work needs pnpm, the monorepo build order, Git hook setup, and package-specific checks.

Everything in the runtime enters through plugins

DeepSeek Harness is a coding-agent runtime that happens to ship with a local web application. Its architecture is built on Cordis, where plugins register services, typed events, and reversible effects in a shared context. Model adapters, tools, session storage, approvals, credentials, telemetry, and the agent loop itself enter through that system.

That makes the project attractive to developers who have outgrown a fixed agent. A model provider can be swapped without rewriting the session engine. Filesystem and subprocess backends can move together to another execution environment. A profile composes bundles, project patches, and user patches into the plugin tree that boots. The running tree can be dumped for inspection.

The cost is conceptual weight. Terms such as profiles, bundles, seams, scopes, waterfalls, durable events, Host and Client faces, and generated remote contracts appear because they map to real extension boundaries. Someone who only wants an agent to edit a repository will get little benefit from learning all of them.

What the default experience includes

The quickest documented start is npx @deepseek-ai/dsh web. It launches a local server and opens the browser when appropriate. The user adds a DeepSeek API key in settings, chooses a workspace, and starts a session. The agent can read and edit files, run commands, maintain a plan, and delegate work. Other providers and compatible endpoints are documented separately.

A headless profile supports one-shot automation without the web server. Session activity is written as an append-only event log, and the model's history is derived from that log. The architecture document insists that model-visible input must be reconstructable from durable events, which is a sound basis for replay, transcripts, forks, and debugging.

MCP support is present as a client bridge. External server tools register into the same tool service as built-in capabilities. That fits the plugin thesis cleanly, though every added server still expands what the model can call and may introduce credentials or side effects that need separate review.

What happened when we ran it

We cloned commit b150a55 into an unprivileged Node 22 container with three CPUs, 8 GB of RAM, and no secrets. The monorepo contained 7,822 files, about 623,075 lines of source, and occupied 50.8 MB. It had 18 CI workflow files, no Dockerfile, and no top-level tests directory.

Pnpm installation succeeded in 22 seconds. It added 936 packages and occupied 1,393 MB on disk. The build completed successfully in 111 seconds. That is a credible source path, though the installed footprint is substantial.

Tests ran for 829 seconds and exited with failure. Vitest reported 14,583 passed, 10 failed, and 114 skipped out of 14,707 tests. At the file level, 859 passed, 4 failed, and 9 were skipped out of 872. The provided tail points to a SQLite schema test named rejects changed columns and non-strict owned tables, but it does not include the assertion message needed to explain the failure. We can say the suite was mostly green and still not call it passing.

Safety contracts are detailed but bounded

Approval requests produce one closed outcome, and only allowed-once grants the action. Rejection, cancellation, a missing answerer, or an error all fail closed. Sessions can use an ask policy or a never policy that rejects requests without prompting. Approval events are logged as pairs, which gives later inspection a record of what was requested and decided.

The process sandbox has read-only, workspace-write, and danger-full-access modes. Linux can use bwrap or Landlock, macOS uses Seatbelt, and Windows uses an ACL restricted-token backend. The documentation is precise about the boundary: these modes govern filesystem effects. Network access and process visibility are outside the sandbox vocabulary. Windows and older Landlock setups may report partial enforcement.

That candor matters. A label such as workspace-write can sound broader than it is unless operators read the subsystem contract. Work that needs network isolation or a stronger process boundary should use a container, microVM, or remote execution seam designed for the whole environment.

The repository also publishes incident write-ups. One resolved post-mortem explains how a benign older-Landlock notice plus a child's nonzero exit was misclassified as sandbox failure. The defect did not weaken confinement, according to the report, but it corrupted availability and diagnosis. The fix tightened evidence rules and added deterministic coverage. That level of engineering documentation is rare and useful.

Preview status should decide the adoption

The README uses unambiguous language: this is a developer preview, it is changing quickly, and compatibility-breaking changes will happen. GitHub's latest-release endpoint returned no release, so evaluators do not have a current tagged artifact there to use as a stability boundary. The repository was pushed on August 21, 2026, and GitHub listed 0 open issues and pull requests when fetched on August 27. The project directs feedback to Discussions, so an empty issue queue is not proof that users have no problems. GitHub showed 198,423 stars, but neither stars nor a recent push supplies the version boundary that a production operator needs.

Documentation is already deep. Architecture maps, tutorials, subsystem contracts, extension cookbooks, generated API descriptions, bilingual pages, and post-mortems give contributors far more than a quick start. The development guide also explains its two TypeScript build faces, generated contracts, hooks, CI lanes, and real-API tests.

DeepSeek Harness makes sense for teams whose requirement is a replaceable agent runtime and who can absorb breaking changes. A curious agent-platform engineer will find plenty worth studying. A team seeking a stable daily assistant should wait for tagged releases and a clean repeatable test result, or select a narrower tool whose public interface has settled.

Alternatives

ProjectWhat it isPick it when
OpenCode gh↗A model-flexible coding agent with a mature terminal interface and beta desktop client.pick this instead when using an agent matters more than replacing every part of its runtime.
Codex CLI gh↗OpenAI's terminal coding agent with built-in sandbox and approval controls.pick this instead when you want an opinionated working CLI rather than a plugin research platform.
AutoGen gh↗A framework for building applications with interacting agents and tools.pick this instead when multi-agent application code is the focus rather than a complete coding-agent product.

What people are saying

  1. [velocity-scout] fufankeji/deepseek-harness-studio
  2. [velocity-scout] alchaincyf/deepseek-harness-orange-book
  3. [velocity-scout] anywhere-labs/deepseek-harness-desktop

Sources

  1. DeepSeek Harness README
  2. DeepSeek Harness architecture
  3. Process sandbox documentation
  4. User approval documentation
  5. Landlock classification post-mortem
  6. MCP package documentation

More llm tools reviews

agent-toolkit-for-aws · agent-memory · codex-astra-luna-orchestrator · okf-agent-memory · mlc-llm · awesome-openclaw-skills · the whole board →