This stack qualifies one workload at a time
CUDA-for-AMD-Windows joins a CUDA-targeted Windows program to an AMD GPU through ZLUDA and the Windows HIP SDK. ZLUDA presents CUDA-facing DLLs, then routes supported work toward AMD libraries such as rocBLAS, hipBLASLt, rocSPARSE, and rocFFT. The repository adds PowerShell installation, runtime staging, GPU detection, diagnostics, pinned asset manifests, numerical probes, and patches for experimental paths.
The narrow claim is the useful one. The stable public reference uses a Radeon RX 9060 XT (gfx1200), ZLUDA v6-preview.69, HIP SDK 6.4, and LibTorch 2.3.0+cu118. The compatibility table also lists the RX 9070 XT as externally validated and the Radeon 890M as partial. Every other recognized AMD GPU remains an unverified candidate until workload evidence arrives.
A detected AMD GPU proves almost nothing
The documentation separates detection, library loading, safe refusal, functional correctness, and end-to-end application validation. A machine can expose an AMD device through the CUDA API while a later kernel returns the wrong values. The project therefore compares selected output against CPU or native references and records unsupported operations separately from timeouts, process errors, and incorrect tensors. That is the right standard for a translation layer.
Issue 9 shows why this distinction matters. On an RX 7900 XT at commit 4a6fb1b, basic matrix operations and the math attention path passed, while the memory-efficient attention path returned a finite but numerically wrong tensor. Issue 3 documents a Radeon 890M path where convolution hung and another attention route produced wrong output. These are community reports with exact environments, not results from our sandbox.
What happened when we ran it
Our sandbox installed commit 4a6fb1b in 20 seconds. It added 35 packages and used 37 MB on disk. The build succeeded in 8 seconds, and pip-audit reported 0 known vulnerabilities. We ran this in an unprivileged Debian container with Python 3.12, 3 CPUs, 8 GB of RAM, and no secrets.
The repository provided no test script or target, so the test step was skipped. Our checkout contained 81 files, about 6,474 lines of source, and occupied 0.7 MB before installation. It had 4 CI workflow files, no Dockerfile, and no tests directory. That verifies a small, buildable source tree; it does not verify PowerShell execution, an AMD driver, HIP, ZLUDA, or CUDA translation.
A Linux packaging pass is particularly limited here because the actual workflow is Windows x64 on AMD hardware. We did not have that hardware path in the supplied sandbox, and we make no claim about GPU speed, numerical correctness, or application compatibility from the 8-second build. The project's own functional scripts remain mandatory on the target PC.
Installation is short only after the prerequisites exist
The recommended installer detects the GPU, checks HIP, downloads pinned assets, stages the runtime, performs a low-level smoke check, and can run numerical probes when a compatible Python environment is present. The operator first needs a current AMD driver and the correct HIP SDK for the GPU architecture. The Radeon 890M profile, for example, has a different minimum than the historical gfx1200 reference.
Launching an application is a separate step through run-zluda.ps1. doctor.ps1 checks the environment, test-runtime.ps1 checks CUDA-facing libraries, and test-functional.ps1 compares selected work through PyTorch. A runtime check that loads 5 core library groups still does not establish that an application's custom kernels, graph behavior, convolution path, or attention backend is safe.
cuDNN and NVIDIA-specific paths are the hard boundary
The stable Windows HIP path covers the CUDA driver, selected BLAS operations, sparse work, and FFTs on the reference system. It lacks a complete cuDNN equivalent. The experimental v7 branch carries a limited cuDNN v8 to MIOpen bridge and a cuSOLVER proxy, but the docs constrain these to named operations and the reference machine. Neither component turns the stack into general cuDNN or CUDA support.
Applications that depend on NCCL, TensorRT, NVIDIA cubins, custom CUDA extensions, or architecture-specific kernels can fail. Fused attention is a useful warning: a fallback may have the right shape and finite numbers while missing the intended compute path. The project's patched candidate rejects known unsafe cases instead of accepting plausible output. Anyone evaluating another program should adopt the same fail-closed rule.
The repository has evidence, but no project release
GitHub showed 307 stars, 3 open issues, 4 combined issues and pull requests, and a last push on September 30, 2026. The latest-release endpoint returned no release. Instead, the installer pins upstream ZLUDA and HIP-era components through repository manifests, while an experimental patch directory tracks a newer ZLUDA line separately.
Project-owned scripts and documentation use the MIT license, but GitHub reports no single SPDX license for the whole stack. ZLUDA, AMD components, PyTorch, LibTorch, and NVIDIA packages retain their own terms. The third-party notice also excludes recovered local artifacts from the public dependency chain. Teams distributing a prepared runtime need to review every included component rather than assuming one repository license covers it.
This is a specialist bridge for a machine and workload you can afford to disprove. Its best feature is the refusal to call a visible GPU a success. Start with the exact functional probe that resembles your application, compare the numbers, and stop when the evidence level ends.

