GPU backends¶
cuPDLPx supports NVIDIA GPUs through CUDA and AMD GPUs through ROCm/HIP. Both backends use the same algorithm and Python, command-line, and C APIs. The compatibility layer selects the GPU libraries.
Backend matrix¶
| Target | Compiler | Dense algebra | Sparse algebra | Parallel primitives |
|---|---|---|---|---|
| NVIDIA CUDA | NVCC | cuBLAS | cuSPARSE | CUB |
| AMD ROCm | hipcc |
hipBLAS | hipSPARSE | hipCUB / rocPRIM |
With USE_HIP=ON, compatibility headers map CUDA calls to HIP so that the
same .cu sources compile for AMD GPUs.
See SpMV for sparse matrix storage, backend selection, and workspace reuse.
Select an architecture¶
CUDA¶
CMake selects default architectures based on the CUDA version. To reduce binary size, specify the target architecture:
Use an architecture supported by the installed CUDA toolkit and the deployment GPU.
ROCm¶
HIP builds require a target architecture. The default is gfx90a; override it
for other devices:
Platform notes¶
- CI builds both the CUDA and ROCm configurations on Linux.
- CUDA builds are also exercised on Windows.
- The native CLI is disabled automatically for MSVC builds because its current argument parser depends on POSIX headers; the libraries and Python binding are separate build targets.
- ROCm packages installed outside CMake's search path may require an explicit
CMAKE_PREFIX_PATH.
Vendor documentation¶
- NVIDIA: CUDA Toolkit, cuBLAS, cuSPARSE, and CUB.
- AMD: ROCm, hipBLAS, hipSPARSE, and rocPRIM.
- The backend dispatch described here is implemented in
internal/cusparse_compat.handsrc/spmv_backend.cu.