CUDA Graphs¶
cuPDLPx uses CUDA Graphs to reduce kernel launch overhead. A graph records GPU operations and their dependencies, allowing the sequence to be replayed with a single CPU launch. Kernel fusion also reduces memory traffic within each iteration.
Graph capture and replay¶
cuPDLPx evaluates termination and restart criteria every
termination_evaluation_frequency iterations
(200 by default). The iterations between consecutive checks form a
window, executed as follows:
- Execute the first iteration of each window outside the graph. After a restart, its PDHG iterate gives the initial fixed-point error of the new epoch, the baseline for the restart criteria and the active-set step size boost.
- Launch a graph covering iterations 2 through
termination_evaluation_frequencywithin the window. Before its first launch, the graph is captured and instantiated; subsequent windows reuse it. The final iteration also saves the intermediate values needed for the checks. - Compute residual vectors and perform reductions on the GPU. The CPU uses the resulting scalar measures to evaluate termination and restart criteria.
Graph reuse across restarts¶
The iterate arrays and work buffers keep the same device addresses across restarts. cuPDLPx updates the anchor, iteration counter, and step sizes in GPU memory, so the same graph can be reused without recapture.
Check frequency
Increasing termination_evaluation_frequency reduces graph launches,
reductions, and host synchronization, but delays convergence checks,
restarts, and detection of time or iteration limits.
ROCm builds use the corresponding HIP Graph APIs through the GPU backend compatibility layer.