Two README files carry the entire benchmark claim
The repository is a report, not an evaluation package. Its v2.0 release lists README.md in Chinese and README.en.md in English. Both present an A/B comparison of DeepSeek V4-Flash-Vision-Exp with and without J-Space V3.7. The table spans HLE, Terminal-Bench 2.1, NL2Repo, CyberGym, DeepSWE, Toolathlon-Verified, Agents' Last Exam, and AutomationBench. A second table reports relative wall-clock, token use, and cost changes.
The headline author-reported average rises from 56.99 to 58.61 across the 7 rows where all 6 model columns have values. Reported improvements vary by task: Terminal-Bench moves from 83.9 to 85.5, DeepSWE from 59.3 to 61.8, and AutomationBench from 25.7 to 27.6. These numbers come from the repository author. They are not our lab measurements, and the repository provides no result file from which to recompute them.
The stated method names subsets but omits the run record
The report says the model, environment, and sampling were held constant while only J-Space changed. It names selected subsets, including 20 medium and 10 hard Terminal-Bench 2.1 tasks, plus 34 DeepSWE tasks split across TypeScript, Python, Go, JavaScript, and Rust. Accuracy and wall-clock are the two stated measurements. That is enough to understand the comparison the author intended, but not enough to recreate the exact run.
Missing artifacts include harness code, task identifiers, prompts, seeds, model endpoint settings, J-Space configuration, timestamps, per-task outcomes, token logs, wall-clock logs, failure categories, and hardware details. The README calls the method theoretically reproducible, yet a reader must reconstruct every operational choice from elsewhere. Even successful reconstruction would show a new run, not prove that the published table was calculated from the same inputs.
What happened when we ran it
The lab recorded no supported ecosystem, language null, and no Dockerfile for this repository. Release v2.0 describes only 2 report files. There is no package manifest, executable harness, build target, or test suite for the sandbox to invoke. The resulting lab card cannot confirm or reject any score, speedup, token reduction, or cost figure printed in the report.
That boundary is especially important here because the publication is numerical. A normal code review can inspect an algorithm even when execution is unavailable. This repository offers only the final tables and a short method paragraph. The companion J-Space repository is linked, but a companion skill does not substitute for the task list and output record of this evaluation. No claim in the report should be relabeled as something our sandbox observed.
An open reproduction report directly challenges one row
Issue 13 says its author spent a day reproducing Terminal-Bench 2.1 and obtained 69 successful tasks out of 89, about 77.5 percent. The report's current table shows 83.9 without J-Space and 85.5 with it, while the issue discusses a different displayed figure from an earlier state. The reporter also argues that some percentages do not correspond cleanly to integer task counts for a single run. This is an allegation in an open issue, not an independently established correction.
The responsible response is to preserve both facts: a reproduction challenge exists, and this repository does not contain enough evidence to resolve it. Task-level traces would show the selected subset, pass criteria, retries, and exact denominator. Commit history could explain which table the reporter saw. Without those artifacts, readers cannot distinguish a changed methodology, a changed report, a calculation problem, or a failed reproduction.
Cross-model columns do not isolate the effect of J-Space
The main table places GLM-5.3, Kimi-K3, Opus-4.8, and Fable 5 beside the two DeepSeek columns. Issue 4 points out that these comparisons do not show what happens when the same skill is attached to each other model. That matters because an intervention can improve several models by different amounts. A fair product comparison would either test the same intervention across models or keep the table limited to the controlled DeepSeek A/B result.
The HLE rows need another caveat. The report states that undisclosed HLE values were carried over from DeepSeek V4-Flash-0731. Those are not measurements of the named experimental model in this repository. The average excludes rows without all 6 values, but copied figures still sit beside the reported experiment and can be mistaken for fresh observations. A machine-readable schema should label provenance per cell.
Read it for hypotheses, not procurement
The repository was created on 2026-08-16, last pushed on 2026-08-23, and released as v2.0 that day. GitHub listed 13 open issues and pull requests combined. Several request traces or question reproducibility, while others contain personal attacks that add heat rather than evidence. Recent activity therefore shows attention, but issue volume does not validate the table.
A useful next step is small and concrete: choose one published subset, freeze task IDs and model settings, save every trajectory, and publish a script that recomputes the table from raw rows. Until that exists, OpenAI Evals, the Language Model Evaluation Harness, or Terminal-Bench itself offers a better base for a decision. This report may suggest that structured skills help long tasks, but its 2 files cannot establish how much.
