mrkeyoor.com_
Sat 03 Oct 06:49 UTC
Dev Toolsevaluationupdated 03 Oct 2026

CUDA-for-AMD-Windows review

CUDA-for-AMD-Windows is a PowerShell-led installer and validation kit for running some Windows CUDA applications on AMD GPUs through ZLUDA and AMD's HIP libraries. It assembles pinned upstream parts, maps selected CUDA-facing libraries to AMD backends, and checks whether a workload truly ran on the GPU with correct output; it is not a complete CUDA implementation.

Verdict

Our checkout installed 35 packages, built in 8 seconds, and offered no test target, so the source package is easy to assemble but did not prove one CUDA workload on AMD hardware. Use this project as a qualification harness for one exact Windows application and GPU, especially near its RX 9060 XT reference path. Do not treat device detection or a passing matrix multiply as permission to run an untested training or inference stack.

We ran it

Lab card: what happened when we ran CUDA-for-AMD-WindowsScreenshot of CUDA-for-AMD-Windows (github.com/Speedstu/CUDA-for-AMD-Windows)
Install✓ · 20s35 packages · 37 MB
Build✓ · 8s
Testsn/ano test script
Known vulns0(pip-audit)
Repo81 files~6,474 lines of source · 0.7 MB · 4 CI workflows

Answers from our run

Does CUDA-for-AMD-Windows build from source?

Dependencies installed in 20 seconds (35 packages), and the build succeeded in 8 seconds. We cloned commit 4a6fb1b into a clean Debian container with 3 CPUs and no project-specific setup.

Does CUDA-for-AMD-Windows have tests you can run?

Not through a standard command: the project exposes no test script or target that our harness could run.

Does CUDA-for-AMD-Windows have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use CUDA-for-AMD-Windows?

Anyone expecting arbitrary CUDA software to work: the README lists NCCL, TensorRT, custom extensions, NVIDIA cubins, and architecture-specific kernels as possible failures.

What are the alternatives to CUDA-for-AMD-Windows?

ZLUDA, TheRock, HIPIFY. Our checkout installed 35 packages, built in 8 seconds, and offered no test target, so the source package is easy to assemble but did not prove one CUDA workload on AMD hardware.

Setup2/5Small source build; real setup needs exact Windows GPU components
Docs5/5Compatibility, evidence levels, failures, and probes are explicit
Community3/5307 stars, 3 open issues, and a September 30 push
Maturity2/5No project release and support remains workload-specific

Who it’s for

Windows developers with a supported AMD GPU and a CUDA-only application worth testing.
Researchers who can compare numerical output against a CPU or native reference before trusting a workload.
Maintainers debugging one CUDA library or kernel path with exact GPU, driver, HIP, and ZLUDA versions.
Radeon RX 9060 XT or RX 9070 XT owners closest to the project's validated configurations.

Who it’s NOT for

Anyone expecting arbitrary CUDA software to work: the README lists NCCL, TensorRT, custom extensions, NVIDIA cubins, and architecture-specific kernels as possible failures.
Users who only know that their machine has an AMD GPU: detection is not functional validation, and most recognized models remain unverified candidates.
Production workloads that cannot tolerate a wrong tensor, GPU hang, or Windows crash during qualification: current compatibility reports document all three classes on partial paths.
Teams requiring a normal versioned release and automated source test target: GitHub has no project release, and our checkout exposed no test script or target.
Linux users or developers who can port source to HIP: native ROCm or HIPIFY avoids this Windows compatibility layer.

Setup reality

Our sandbox installed commit 4a6fb1b in 20 seconds, adding 35 packages and using 37 MB. The build passed in 8 seconds. There was no test script or target, so tests were skipped; pip-audit found 0 known vulnerabilities.

That Debian run checked repository packaging, not the product's GPU path. Real use needs Windows x64, a supported AMD driver, an architecture-appropriate HIP SDK, downloaded ZLUDA assets, and a CUDA-targeted application. Optional functional checks need a compatible Python and PyTorch environment.

The reference configuration is narrow: RX 9060 XT with ZLUDA v6-preview.69 and HIP SDK 6.4. Other GPUs and the v7 patch set have separate evidence levels, failure modes, and version floors.

This stack qualifies one workload at a time

CUDA-for-AMD-Windows joins a CUDA-targeted Windows program to an AMD GPU through ZLUDA and the Windows HIP SDK. ZLUDA presents CUDA-facing DLLs, then routes supported work toward AMD libraries such as rocBLAS, hipBLASLt, rocSPARSE, and rocFFT. The repository adds PowerShell installation, runtime staging, GPU detection, diagnostics, pinned asset manifests, numerical probes, and patches for experimental paths.

The narrow claim is the useful one. The stable public reference uses a Radeon RX 9060 XT (gfx1200), ZLUDA v6-preview.69, HIP SDK 6.4, and LibTorch 2.3.0+cu118. The compatibility table also lists the RX 9070 XT as externally validated and the Radeon 890M as partial. Every other recognized AMD GPU remains an unverified candidate until workload evidence arrives.

A detected AMD GPU proves almost nothing

The documentation separates detection, library loading, safe refusal, functional correctness, and end-to-end application validation. A machine can expose an AMD device through the CUDA API while a later kernel returns the wrong values. The project therefore compares selected output against CPU or native references and records unsupported operations separately from timeouts, process errors, and incorrect tensors. That is the right standard for a translation layer.

Issue 9 shows why this distinction matters. On an RX 7900 XT at commit 4a6fb1b, basic matrix operations and the math attention path passed, while the memory-efficient attention path returned a finite but numerically wrong tensor. Issue 3 documents a Radeon 890M path where convolution hung and another attention route produced wrong output. These are community reports with exact environments, not results from our sandbox.

What happened when we ran it

Our sandbox installed commit 4a6fb1b in 20 seconds. It added 35 packages and used 37 MB on disk. The build succeeded in 8 seconds, and pip-audit reported 0 known vulnerabilities. We ran this in an unprivileged Debian container with Python 3.12, 3 CPUs, 8 GB of RAM, and no secrets.

The repository provided no test script or target, so the test step was skipped. Our checkout contained 81 files, about 6,474 lines of source, and occupied 0.7 MB before installation. It had 4 CI workflow files, no Dockerfile, and no tests directory. That verifies a small, buildable source tree; it does not verify PowerShell execution, an AMD driver, HIP, ZLUDA, or CUDA translation.

A Linux packaging pass is particularly limited here because the actual workflow is Windows x64 on AMD hardware. We did not have that hardware path in the supplied sandbox, and we make no claim about GPU speed, numerical correctness, or application compatibility from the 8-second build. The project's own functional scripts remain mandatory on the target PC.

Installation is short only after the prerequisites exist

The recommended installer detects the GPU, checks HIP, downloads pinned assets, stages the runtime, performs a low-level smoke check, and can run numerical probes when a compatible Python environment is present. The operator first needs a current AMD driver and the correct HIP SDK for the GPU architecture. The Radeon 890M profile, for example, has a different minimum than the historical gfx1200 reference.

Launching an application is a separate step through run-zluda.ps1. doctor.ps1 checks the environment, test-runtime.ps1 checks CUDA-facing libraries, and test-functional.ps1 compares selected work through PyTorch. A runtime check that loads 5 core library groups still does not establish that an application's custom kernels, graph behavior, convolution path, or attention backend is safe.

cuDNN and NVIDIA-specific paths are the hard boundary

The stable Windows HIP path covers the CUDA driver, selected BLAS operations, sparse work, and FFTs on the reference system. It lacks a complete cuDNN equivalent. The experimental v7 branch carries a limited cuDNN v8 to MIOpen bridge and a cuSOLVER proxy, but the docs constrain these to named operations and the reference machine. Neither component turns the stack into general cuDNN or CUDA support.

Applications that depend on NCCL, TensorRT, NVIDIA cubins, custom CUDA extensions, or architecture-specific kernels can fail. Fused attention is a useful warning: a fallback may have the right shape and finite numbers while missing the intended compute path. The project's patched candidate rejects known unsafe cases instead of accepting plausible output. Anyone evaluating another program should adopt the same fail-closed rule.

The repository has evidence, but no project release

GitHub showed 307 stars, 3 open issues, 4 combined issues and pull requests, and a last push on September 30, 2026. The latest-release endpoint returned no release. Instead, the installer pins upstream ZLUDA and HIP-era components through repository manifests, while an experimental patch directory tracks a newer ZLUDA line separately.

Project-owned scripts and documentation use the MIT license, but GitHub reports no single SPDX license for the whole stack. ZLUDA, AMD components, PyTorch, LibTorch, and NVIDIA packages retain their own terms. The third-party notice also excludes recovered local artifacts from the public dependency chain. Teams distributing a prepared runtime need to review every included component rather than assuming one repository license covers it.

This is a specialist bridge for a machine and workload you can afford to disprove. Its best feature is the refusal to call a visible GPU a success. Start with the exact functional probe that resembles your application, compare the numbers, and stop when the evidence level ends.

Alternatives

ProjectWhat it isPick it when
ZLUDAThe upstream CUDA-on-non-NVIDIA translation layer used by this project.pick this instead when you want the upstream runtime directly and can assemble your own AMD stack and tests.
TheRockAMD's build system and distribution project for HIP and ROCm components.pick this instead when your application can use native HIP and you want a current AMD runtime build.
HIPIFYTools for translating CUDA source into portable HIP source.pick this instead when you control the source and can maintain a real CUDA-to-HIP port.

What people are saying

  1. [velocity-scout] Speedstu/CUDA-for-AMD-Windows
  2. [hackernews] CUDA for AMD on Windows

Sources

  1. CUDA for AMD on Windows README
  2. Compatibility matrix
  3. Validation record
  4. Architecture guide
  5. RX 7900 XT compatibility report
  6. Radeon 890M compatibility report

More dev tools reviews

ToolReplay · DuoFold-Android · wutw-public · viserys-agent · birdview · YOINK · the whole board →