mrkeyoor.com_
Tue 01 Sept 17:45 UTC
Dataevaluationupdated 25 Aug 2026

DeepSeek-V4-J-Space-Capability-Realization-Report review

This benchmark report is primarily written in Chinese, and a full English translation is included. It claims to compare DeepSeek V4-Flash-Vision-Exp with and without J-Space V3.7 across selected agent and reasoning benchmarks, but the repository publishes summary tables rather than the harness, task traces, or result files needed to audit those claims.

+2stars / 7d
Verdict

The v2.0 repository ships 2 report files and no executable harness, raw task traces, or machine-readable results, so its score table is not decision-grade evidence. Read it as a set of claims and experiment ideas, then reproduce the relevant subset before changing models or adopting J-Space. The open reproduction dispute makes that verification mandatory, not optional.

We ran it

Screenshot of DeepSeek-V4-J-Space-Capability-Realization-Report (github.com/Tiger3807861189/DeepSeek-V4-J-Space-Capability-Realization-Report)

Answers from our run

Did you run DeepSeek-V4-J-Space-Capability-Realization-Report yourself?

No. GitHub reports no primary language for it, and it carries no manifest our lab installs from, and no Dockerfile, so there was nothing standard to install, build or test. This review is written from the repository's own documentation.

Who should not use DeepSeek-V4-J-Space-Capability-Realization-Report?

Anyone choosing a model or agent stack from reproducible evidence: the v2.0 release lists only Chinese and English report files.

What are the alternatives to DeepSeek-V4-J-Space-Capability-Realization-Report?

OpenAI Evals, Language Model Evaluation Harness, Terminal-Bench. The v2.

Setup1/5Easy to read, but impossible to reproduce from the files provided
Docs2/5Bilingual summary tables without runnable methods or raw evidence
Community2/5Active discussion is dominated by unresolved evidence disputes
Maturity1/5A young report repository with no auditable evaluation artifacts

Discussed on

  1. hnDeepSeek V4 J-Space Capability Realization-Report5 points
  2. hnDeepSeek-V4-Pro-0813 outperforming Fable3 points

Who it’s for

Researchers collecting claims about skill-based changes to agent performance.
DeepSeek users who want hypotheses to test with their own harness.
Benchmark maintainers studying how summarized results can become hard to audit.
Readers comfortable treating every score as author-reported rather than independently verified.

Who it’s NOT for

Anyone choosing a model or agent stack from reproducible evidence: the v2.0 release lists only Chinese and English report files.
Teams that need task-level outputs, run logs, seeds, configurations, or machine-readable results: none are supplied in this repository.
Readers who need an uncontested Terminal-Bench result: issue 13 reports a failed reproduction and questions whether the displayed accuracy can map to integer task counts.
Buyers comparing J-Space across models: issue 4 notes that the other model columns were not tested with the same skill attached.
People who need to adapt or remix the report: release v2.0 applies CC BY-ND 4.0, which prohibits distributing derivatives.

Setup reality

We did not run this repository because the lab found no supported executable ecosystem, reported its language as null, and found no Dockerfile. It is a document repository, so there was no install, build, or test result to measure.

Reading it requires no credentials or service. Reproducing its claims is another matter: the report names DeepSeek Harness, J-Space V3.7, selected benchmark subsets, matching sampling, accuracy, and wall-clock, but it does not include runnable harness code or the recorded task trajectories.

Release v2.0 lists 2 files, README.md and README.en.md. A reproducer must source the model, benchmark datasets, compute, harness, exact configuration, and J-Space companion suite independently.

Two README files carry the entire benchmark claim

The repository is a report, not an evaluation package. Its v2.0 release lists README.md in Chinese and README.en.md in English. Both present an A/B comparison of DeepSeek V4-Flash-Vision-Exp with and without J-Space V3.7. The table spans HLE, Terminal-Bench 2.1, NL2Repo, CyberGym, DeepSWE, Toolathlon-Verified, Agents' Last Exam, and AutomationBench. A second table reports relative wall-clock, token use, and cost changes.

The headline author-reported average rises from 56.99 to 58.61 across the 7 rows where all 6 model columns have values. Reported improvements vary by task: Terminal-Bench moves from 83.9 to 85.5, DeepSWE from 59.3 to 61.8, and AutomationBench from 25.7 to 27.6. These numbers come from the repository author. They are not our lab measurements, and the repository provides no result file from which to recompute them.

The stated method names subsets but omits the run record

The report says the model, environment, and sampling were held constant while only J-Space changed. It names selected subsets, including 20 medium and 10 hard Terminal-Bench 2.1 tasks, plus 34 DeepSWE tasks split across TypeScript, Python, Go, JavaScript, and Rust. Accuracy and wall-clock are the two stated measurements. That is enough to understand the comparison the author intended, but not enough to recreate the exact run.

Missing artifacts include harness code, task identifiers, prompts, seeds, model endpoint settings, J-Space configuration, timestamps, per-task outcomes, token logs, wall-clock logs, failure categories, and hardware details. The README calls the method theoretically reproducible, yet a reader must reconstruct every operational choice from elsewhere. Even successful reconstruction would show a new run, not prove that the published table was calculated from the same inputs.

What happened when we ran it

The lab recorded no supported ecosystem, language null, and no Dockerfile for this repository. Release v2.0 describes only 2 report files. There is no package manifest, executable harness, build target, or test suite for the sandbox to invoke. The resulting lab card cannot confirm or reject any score, speedup, token reduction, or cost figure printed in the report.

That boundary is especially important here because the publication is numerical. A normal code review can inspect an algorithm even when execution is unavailable. This repository offers only the final tables and a short method paragraph. The companion J-Space repository is linked, but a companion skill does not substitute for the task list and output record of this evaluation. No claim in the report should be relabeled as something our sandbox observed.

An open reproduction report directly challenges one row

Issue 13 says its author spent a day reproducing Terminal-Bench 2.1 and obtained 69 successful tasks out of 89, about 77.5 percent. The report's current table shows 83.9 without J-Space and 85.5 with it, while the issue discusses a different displayed figure from an earlier state. The reporter also argues that some percentages do not correspond cleanly to integer task counts for a single run. This is an allegation in an open issue, not an independently established correction.

The responsible response is to preserve both facts: a reproduction challenge exists, and this repository does not contain enough evidence to resolve it. Task-level traces would show the selected subset, pass criteria, retries, and exact denominator. Commit history could explain which table the reporter saw. Without those artifacts, readers cannot distinguish a changed methodology, a changed report, a calculation problem, or a failed reproduction.

Cross-model columns do not isolate the effect of J-Space

The main table places GLM-5.3, Kimi-K3, Opus-4.8, and Fable 5 beside the two DeepSeek columns. Issue 4 points out that these comparisons do not show what happens when the same skill is attached to each other model. That matters because an intervention can improve several models by different amounts. A fair product comparison would either test the same intervention across models or keep the table limited to the controlled DeepSeek A/B result.

The HLE rows need another caveat. The report states that undisclosed HLE values were carried over from DeepSeek V4-Flash-0731. Those are not measurements of the named experimental model in this repository. The average excludes rows without all 6 values, but copied figures still sit beside the reported experiment and can be mistaken for fresh observations. A machine-readable schema should label provenance per cell.

Read it for hypotheses, not procurement

The repository was created on 2026-08-16, last pushed on 2026-08-23, and released as v2.0 that day. GitHub listed 13 open issues and pull requests combined. Several request traces or question reproducibility, while others contain personal attacks that add heat rather than evidence. Recent activity therefore shows attention, but issue volume does not validate the table.

A useful next step is small and concrete: choose one published subset, freeze task IDs and model settings, save every trajectory, and publish a script that recomputes the table from raw rows. Until that exists, OpenAI Evals, the Language Model Evaluation Harness, or Terminal-Bench itself offers a better base for a decision. This report may suggest that structured skills help long tasks, but its 2 files cannot establish how much.

Alternatives

ProjectWhat it isPick it when
OpenAI EvalsAn evaluation framework for defining tests and recording model results in code.pick this instead when you need an executable evaluation rather than a standalone score table.
Language Model Evaluation HarnessA widely used framework for repeatable language-model evaluation across many tasks.pick this instead when standardized task definitions and machine-readable runs matter.
Terminal-BenchThe executable benchmark behind one of the report's most disputed result rows.pick this instead when you want to run terminal-agent tasks and inspect the outcomes yourself.

What people are saying

  1. [velocity-scout] Tiger3807861189/DeepSeek-V4-J-Space-Capability-Realization-Report

Sources

  1. DeepSeek V4 J-Space report
  2. Report v2.0 release
  3. Terminal-Bench reproduction issue 13
  4. Cross-model comparison issue 4
  5. Request for task traces issue 7

More data reviews

turso · TrackersListCollection · dash · getcontact-cli · awesome-zhuiju-free · iggy · the whole board →