mrkeyoor.com_
Fri 02 Oct 14:57 UTC
AI Toolsevaluationupdated 02 Oct 2026

DeepSelect review

DeepSelect is a bilingual English-and-Chinese PyTorch extension for the TopK operation used in DeepSeek Sparse Attention and sampling. It supplies specialized NVIDIA CUDA and Huawei Ascend kernels for a narrow set of bfloat16 and float32 workloads, with full English setup and algorithm notes alongside Chinese material.

Verdict

Our DeepSelect run installed 35 packages and built in 24 seconds combined, then pytest executed zero tests because all 11 items hit collection/setup errors. Trial it only when its narrow TopK contract matches supported hardware and you can run correctness checks on that hardware yourself. The active issue queue contains enough memory, device, and compiler concerns that we would not drop commit bfa4507 into an unattended inference service without that work.

We ran it

Lab card: what happened when we ran DeepSelectScreenshot of DeepSelect (github.com/deepseek-ai/DeepSelect)
Install✓ · 17s35 packages · 37 MB
Build✓ · 7s
Tests✗ · 6s0 passed · 0 failed · 11 errors of 11 (pytest)
Known vulns0(pip-audit)
Repo7049 files~1,020,766 lines of source · 154.4 MB · 0 CI workflows · tests dir

Answers from our run

Does DeepSelect build from source?

Dependencies installed in 17 seconds (35 packages), and the build succeeded in 7 seconds. We cloned commit bfa4507 into a clean Debian container with 3 CPUs and no project-specific setup.

Do DeepSelect's tests pass?

Yes: 0 of 11 passed when we ran the project's own test command (pytest), with 11 collection errors. Some failures need services or credentials a bare container does not have.

Does DeepSelect have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use DeepSelect?

General PyTorch users who need a portable topk: DeepSelect accepts only bfloat16 or float32 inputs, caps topk at 4096, and imposes aligned row strides.

What are the alternatives to DeepSelect?

PyTorch, LiteTopK, CUTLASS. Our DeepSelect run installed 35 packages and built in 24 seconds combined, then pytest executed zero tests because all 11 items hit collection/setup errors.

Setup2/5Build passed, but pytest stopped at 11 collection/setup errors
Docs4/5Bilingual usage, limits, benchmark method, and algorithm notes
Community3/5450 stars and active PRs, but only weeks of public history
Maturity2/5Narrow hardware targets and unresolved correctness reports

Who it’s for

Inference engineers running DeepSeek-style sparse attention on supported NVIDIA Blackwell or Huawei Ascend hardware.
CUDA or Ascend kernel teams that can validate outputs, memory behavior, and compiler compatibility on their exact fleet.
PyTorch services whose TopK workload uses small topk values, especially 512, and can meet the documented stride rules.

Who it’s NOT for

General PyTorch users who need a portable topk: DeepSelect accepts only bfloat16 or float32 inputs, caps topk at 4096, and imposes aligned row strides.
Teams expecting a clean fresh-container test gate: our run executed no tests and stopped with 11 collection/setup errors.
NVIDIA fleets outside the checked-in SM100 and SM103 targets: commit bfa4507 requires NVCC 12.9 or newer for its CUDA build, while SM90 support was still an open pull request.
Multi-GPU or unusual tensor-view deployments that cannot audit unresolved reports: open issues describe wrong-device launch setup and vector accesses beyond accepted tensor boundaries.
Buyers who require tagged artifacts and release notes: the README names v1.0.0, but GitHub had no tags or release record.

Setup reality

Our sandbox installed commit bfa4507 in 17 seconds, adding 35 packages and using 37 MB. The build succeeded in 7 seconds. Pytest failed after 6 seconds: 0 passed, 0 failed, and all 11 collected items ended as collection/setup errors. Pip-audit reported 0 known vulnerabilities.

No hosted credential is required. A real CUDA build needs PyTorch with CUDA, NVCC 12.9 or newer, the CUTLASS submodule, and supported hardware. The Ascend path needs PyTorch, torch_npu, the Ascend toolkit, and its library paths. The package metadata does not declare these requirements for pip to resolve.

The test log explicitly showed ModuleNotFoundError: No module named 'torch' for a bundled CUTLASS CuTeDSL example, then named several more CUTLASS Python tests before the 11-error summary. The log does not establish that every error had the same cause, so we cannot call this a kernel failure or a missing-system-package fix.

TopK values above 4096 are outside the product

DeepSelect replaces one operation rather than serving a whole model. It takes a two-dimensional tensor and returns the largest values or their indices, tuned for the TopK patterns used by DeepSeek Sparse Attention and samplers. The README caps topk at 4096 and says its main tuning target is 512. The bfloat16 path covers the sparse-attention case, while float32 input is available only on CUDA for sampling.

That specialization buys control over details PyTorch users rarely manage themselves. Callers can skip returned values, ask for indices sorted by position, provide per-row end points, add index offsets, or supply an output buffer. These options matter inside an inference engine. They also narrow the safe operating envelope. Inputs need aligned row strides, outputs may be non-contiguous, and Ascend supports only bfloat16 input and int32 indices.

commit bfa4507 builds CUDA only for SM100 and SM103

The checked-in CUDA setup compiles sm_100a and sm_103a targets and rejects NVCC versions older than 12.9. Comments say SM80 and SM90 builds are skipped. An open pull request adds an SM90 path for Hopper GPUs, but that code was not part of the commit we measured. If your fleet is H100 or an older architecture, the default source build is not the documented ready path.

Huawei Ascend arrived in the September 30 commit we tested. That build path imports torch_npu, looks for the Ascend toolkit under /usr/local/Ascend/ascend-toolkit/latest unless configured otherwise, and defaults to the dav-3510 target. DeepSelect detects the backend from the host, though an environment variable can choose CUDA or Ascend. There is no CPU implementation to fall back to.

What happened when we ran it

In our unprivileged Debian container with 3 CPUs and 8 GB of RAM, installation succeeded in 17 seconds. It added 35 packages, occupied 37 MB, and pip-audit reported 0 known vulnerabilities. The build then completed in 7 seconds. The checkout was much larger than the installed dependency layer: 7,049 files, roughly 1,020,766 source lines, and 154.4 MB.

Pytest failed after 6 seconds without executing a test. Its summary was 0 passed, 0 failed, and 11 collection/setup errors out of 11 items. One line explicitly reported ModuleNotFoundError: No module named 'torch' while collecting a bundled CUTLASS CuTeDSL example. The tail also listed several csrc/3rdparty/cutlass/test/python/pycute files and a sharding test, then ended with 11 errors in 0.70s.

The log does not say that all 11 errors share the missing torch cause, and it contains no kernel result. We therefore treat the run as a packaging and test-discovery failure, not proof that CUDA or Ascend math is wrong. It still matters: setup.py has no declared install requirements, and a fresh environment that completed the supplied install and build steps did not reach the project's own tests.

NaN detection can abort the entire kernel

DeepSelect checks for NaN values on every call. With the default abort_when_nan_found=True, the CUDA kernel invokes a trap and aborts. The README also notes a CUDA exception: rows no longer than topk skip that NaN check. Setting the flag to false changes the output to a sentinel pattern instead. This behavior belongs in an integration test because malformed model data can change a request failure into a process-level event.

Memory layout deserves the same attention. The input row stride must match the byte requirement returned by get_stride_requirement(), and unaligned inputs need padding. Caller-owned index buffers have their own alignment rule. Open issue 5 presents source-derived cases where vector accesses can extend past accepted tensor boundaries; its paired pull request was open when checked. Treat custom views and preallocated outputs as review points, not ordinary tensors.

Six open issues and 11 pull requests are active design work

GitHub reported 17 combined open issues and pull requests on October 2, 2026. The API list separated them into 6 issues and 11 pull requests. Reports covered index-offset overflow, launching from the wrong CUDA device, tensor-boundary checks, host setup overhead, GCC 11 compatibility, and attribution to LiteTopK. Several proposed fixes include focused tests, but open pull requests are not behavior in commit bfa4507.

The repository was pushed on September 30, the day Ascend support landed, so it is actively changing. It had 450 stars and an MIT license, but no GitHub tags or release objects even though the README calls the initial code v1.0.0. DeepSelect is best read as young kernel source for teams with the exact hardware and expertise to qualify it. Our 11 collection errors make that qualification your job, not a box already checked upstream.

Alternatives

ProjectWhat it isPick it when
PyTorch gh↗The general tensor framework whose built-in `torch.topk` is the compatibility baseline here.pick this instead when portability, packaged binaries, and broad hardware coverage matter more than a workload-specific TopK kernel.
LiteTopKA CUDA workspace for specialized TopK and sparse-attention kernels on SM100 hardware.pick this instead when its FP8 or FP4 DeepSeek-style paths match your B200 workload and you can run its machine-specific benchmark setup.
CUTLASSNVIDIA's CUDA templates and CuTe tooling for building custom GPU kernels.pick this instead when you need primitives for your own kernel rather than DeepSelect's fixed TopK interface.

What people are saying

  1. [velocity-scout] deepseek-ai/DeepSelect

Sources

  1. DeepSelect README at the tested commit
  2. DeepSelect build configuration
  3. DeepSelect Python interface
  4. DeepSelect algorithm notes
  5. Open tensor-boundary issue
  6. Open wrong-device issue

More ai tools reviews

xialingguo-ip · reelbench-skills · SoL-Pi · flybook · crypto-rag · anything2explainer · the whole board →