DeepGEMM targets recent NVIDIA data-center GPUs
DeepGEMM puts four numeric families, FP8, FP4, BF16, and TF32, behind one CUDA library for dense layers, mixture-of-experts models, and parts of attention. The interfaces include grouped GEMMs and fused Mega MoE operations. Its audience is an engineer who already knows which tensor layout a model produces and wants control over the kernel that consumes it. This is infrastructure below an inference server, not an inference server itself.
The hardware boundary is firm. The README requires an NVIDIA SM90 or SM100 GPU, CUDA Toolkit 12.9 or newer, PyTorch 2.3+, and a compiler whose standard library supports C++20 format. Python 3.8 is the stated minimum. DeepGEMM also expects CUTLASS 4.0 or newer through a Git submodule, so the documented clone is recursive. A laptop GPU, AMD node, or CPU host cannot exercise the intended path.
Callers supply the layouts that the kernels expect
SM90 expects FP32 scale values, while SM100 packs four UE8M0 values into one integer. A dense FP8 call also distinguishes NT, NN, TN, and TT memory layouts, while grouped operations impose rules on which dimensions may vary. The API names reveal how close this code sits to the metal. Those details affect the tensors you prepare before the multiplication begins.
DeepGEMM leaves input transposition and FP8 casting to the caller. It provides utility functions for alignment and scale transformations, but the README warns that those helpers may be slower than folding the work into an earlier kernel. That is a sensible trade for a specialist library. It is also why adopting it means changing a model pipeline, measuring the result on your own shapes, and checking numerical output. Merely replacing a package name will not do the job.
What happened when we ran it
Our sandbox installed commit 057ca59 in 18 seconds, pulling 35 packages and using 37 MB on disk. The build completed in another 4 seconds. The container supplied 3 CPUs, 8 GB of RAM, Python 3.12 on Debian, no credentials, and no elevated privileges. Pip-audit reported 0 known vulnerabilities. These numbers describe repository setup only; we did not measure GPU throughput or kernel accuracy.
Pytest failed after 4 seconds. It reported 0 passed, 0 failed, and 24 collection/setup errors across all 24 test modules it tried to collect. The final lines named six PyCuTe tests under the vendored CUTLASS tree and three DeepJIT tests for Ascend, CUDA, and exceptions, then ended with 24 errors in 1.27s. The supplied log tail does not include the underlying exceptions, so it cannot support a claim about missing drivers, packages, or defects.
The distinction matters because the install and build both passed. A Python environment can accept the package while the broad repository test command still fails before executing one assertion. DeepGEMM has 3 CI workflow files and a tests directory, so testing is present, but our generic fresh-container result was not a usable release gate. A prospective adopter should run the project’s hardware-specific cases on the exact GPU and CUDA combination intended for production.
Runtime compilation moves work into deployment
DeepJIT compiles kernels at runtime with the CUDA 12.9+ toolchain instead of compiling CUDA during package installation. The README exposes controls for the compiler path, C++ standard, cache directory, forced builds, debug output, line information, and PTX or SASS dumps. The default cache lives under the user's home directory, and a cache miss compiles into the first configured cache path.
That design is useful when kernel shapes change, but it gives operators another stateful directory to size and persist. A read-only home, short-lived container, or fleet with mixed CUDA installations needs an explicit cache and compiler policy. Mega MoE raises the bar further: its example uses multiple processes, symmetric memory, NVLink communication, and PyTorch 2.9 or newer for the buffer API. Those requirements belong in capacity planning before any claimed kernel win matters.
Active code does not widen the supported hardware
GitHub showed 8,576 stars and 149 open issues and pull requests on October 6, 2026. Search split that combined figure into 69 issues and 80 pull requests. The last push was September 30, the same date the README announced a separate DeepGEMM Ascend project and more locality work. The latest tagged GitHub release was v2.1.1.post3 from October 15, 2025, but the newer push and pull-request activity show that development continued after that tag.
DeepGEMM earns a trial when its exact kernel families match a production model on Hopper or Blackwell-class infrastructure. The 142.4 MB checkout, 6,880 files, and roughly 852,594 source lines in our snapshot also make clear that this is no tiny teaching sample, despite the focused public API. Start with one real tensor shape, compare outputs against the existing framework path, and keep that path available until the JIT cache and hardware-specific tests are repeatable.

