TopK values above 4096 are outside the product
DeepSelect replaces one operation rather than serving a whole model. It takes a two-dimensional tensor and returns the largest values or their indices, tuned for the TopK patterns used by DeepSeek Sparse Attention and samplers. The README caps topk at 4096 and says its main tuning target is 512. The bfloat16 path covers the sparse-attention case, while float32 input is available only on CUDA for sampling.
That specialization buys control over details PyTorch users rarely manage themselves. Callers can skip returned values, ask for indices sorted by position, provide per-row end points, add index offsets, or supply an output buffer. These options matter inside an inference engine. They also narrow the safe operating envelope. Inputs need aligned row strides, outputs may be non-contiguous, and Ascend supports only bfloat16 input and int32 indices.
commit bfa4507 builds CUDA only for SM100 and SM103
The checked-in CUDA setup compiles sm_100a and sm_103a targets and rejects NVCC versions older than 12.9. Comments say SM80 and SM90 builds are skipped. An open pull request adds an SM90 path for Hopper GPUs, but that code was not part of the commit we measured. If your fleet is H100 or an older architecture, the default source build is not the documented ready path.
Huawei Ascend arrived in the September 30 commit we tested. That build path imports torch_npu, looks for the Ascend toolkit under /usr/local/Ascend/ascend-toolkit/latest unless configured otherwise, and defaults to the dav-3510 target. DeepSelect detects the backend from the host, though an environment variable can choose CUDA or Ascend. There is no CPU implementation to fall back to.
What happened when we ran it
In our unprivileged Debian container with 3 CPUs and 8 GB of RAM, installation succeeded in 17 seconds. It added 35 packages, occupied 37 MB, and pip-audit reported 0 known vulnerabilities. The build then completed in 7 seconds. The checkout was much larger than the installed dependency layer: 7,049 files, roughly 1,020,766 source lines, and 154.4 MB.
Pytest failed after 6 seconds without executing a test. Its summary was 0 passed, 0 failed, and 11 collection/setup errors out of 11 items. One line explicitly reported ModuleNotFoundError: No module named 'torch' while collecting a bundled CUTLASS CuTeDSL example. The tail also listed several csrc/3rdparty/cutlass/test/python/pycute files and a sharding test, then ended with 11 errors in 0.70s.
The log does not say that all 11 errors share the missing torch cause, and it contains no kernel result. We therefore treat the run as a packaging and test-discovery failure, not proof that CUDA or Ascend math is wrong. It still matters: setup.py has no declared install requirements, and a fresh environment that completed the supplied install and build steps did not reach the project's own tests.
NaN detection can abort the entire kernel
DeepSelect checks for NaN values on every call. With the default abort_when_nan_found=True, the CUDA kernel invokes a trap and aborts. The README also notes a CUDA exception: rows no longer than topk skip that NaN check. Setting the flag to false changes the output to a sentinel pattern instead. This behavior belongs in an integration test because malformed model data can change a request failure into a process-level event.
Memory layout deserves the same attention. The input row stride must match the byte requirement returned by get_stride_requirement(), and unaligned inputs need padding. Caller-owned index buffers have their own alignment rule. Open issue 5 presents source-derived cases where vector accesses can extend past accepted tensor boundaries; its paired pull request was open when checked. Treat custom views and preallocated outputs as review points, not ordinary tensors.
Six open issues and 11 pull requests are active design work
GitHub reported 17 combined open issues and pull requests on October 2, 2026. The API list separated them into 6 issues and 11 pull requests. Reports covered index-offset overflow, launching from the wrong CUDA device, tensor-boundary checks, host setup overhead, GCC 11 compatibility, and attribution to LiteTopK. Several proposed fixes include focused tests, but open pull requests are not behavior in commit bfa4507.
The repository was pushed on September 30, the day Ascend support landed, so it is actively changing. It had 450 stars and an MIT license, but no GitHub tags or release objects even though the README calls the initial code v1.0.0. DeepSelect is best read as young kernel source for teams with the exact hardware and expertise to qualify it. Our 11 collection errors make that qualification your job, not a box already checked upstream.

