Skip to content

CUDA Graphs

cuPDLPx uses CUDA Graphs to reduce kernel launch overhead. A graph records GPU operations and their dependencies, allowing the sequence to be replayed with a single CPU launch. Kernel fusion also reduces memory traffic within each iteration.

Individual kernel launches compared with a single CUDA Graph launch Individual kernel launches compared with a single CUDA Graph launch
Launch overhead can leave gaps between short GPU kernels; a graph launch submits the recorded sequence together. Adapted from NVIDIA, Effortless CUDA Graphs, GTC 2021.

Graph capture and replay

cuPDLPx evaluates termination and restart criteria every termination_evaluation_frequency iterations (200 by default). The iterations between consecutive checks form a window, executed as follows:

  1. Execute the first iteration of each window outside the graph. After a restart, its PDHG iterate gives the initial fixed-point error of the new epoch, the baseline for the restart criteria and the active-set step size boost.
  2. Launch a graph covering iterations 2 through termination_evaluation_frequency within the window. Before its first launch, the graph is captured and instantiated; subsequent windows reuse it. The final iteration also saves the intermediate values needed for the checks.
  3. Compute residual vectors and perform reductions on the GPU. The CPU uses the resulting scalar measures to evaluate termination and restart criteria.

Graph reuse across restarts

The iterate arrays and work buffers keep the same device addresses across restarts. cuPDLPx updates the anchor, iteration counter, and step sizes in GPU memory, so the same graph can be reused without recapture.

Check frequency

Increasing termination_evaluation_frequency reduces graph launches, reductions, and host synchronization, but delays convergence checks, restarts, and detection of time or iteration limits.

ROCm builds use the corresponding HIP Graph APIs through the GPU backend compatibility layer.