Skip to content

Build against cuOpt 26.10 nightly: multi-GPU PDLP, new options, iteration/node reporting - #20

Open
0x17 wants to merge 19 commits into
mainfrom
multigpu-pdlp-nightly
Open

0x17 wants to merge 19 commits into
mainfrom
multigpu-pdlp-nightly

Conversation

@0x17

@0x17 0x17 commented Oct 2, 2026 •

Copy link
Copy Markdown
Member
  • Build both workflows against a cuOpt 26.10 nightly from the RAPIDS nightly index, currently pinned to libcuopt-cuXX==26.10.0a223.post261002122804 as a stopgap
    • needs --prerelease=allow and --index-strategy unsafe-best-match so RAPIDS packages come from the nightly index and CUDA libraries from pypi.nvidia.com
    • a223 is the last nightly before cuOpt switched to modular component wheels (a224+). Against the modular a226 wheels the link builds, but every cuOptGetIntegerParameter/cuOptSetIntegerParameter call returns CUOPT_INVALID_ARGUMENT, so every solve ends with status 13
    • pins libcuopt-cuXX directly: there is no cuopt-cuXX wheel for a223, and pinning cuopt-cuXX==26.10.0a222 still resolves libcuopt to the modular a226
    • packaging for the modular layout (8edd916) is reverted in 4855ada and kept in history; unpin and reapply it once working modular 26.10 packages from NVIDIA are on PyPI
  • Adapt link and bundles to the split cuOpt libraries inside the libcuopt wheel (both workflows and build-link.sh)
    • link -lcuopt_mathopt, bundle libcuopt/lib64/libcuopt_mathopt.so + libcuopt_client.so
    • copy cuOpt's cuDSS threading layer libcuopt/lib64/libcudss_mtlayer_cuopt.so; do not bundle libcudss_mtlayer_gomp.so
    • libcuopt_client.so needs system libssl.so.3/libcrypto.so.3/libz.so.1 (not bundled)
  • Support multi-GPU PDLP (Add C API support for multiGPU PDLP  NVIDIA/cuopt#1958)
    • optcuopt.def: num_gpus now accepts -1..72, new multigpu_pdlp_partitioner option
    • bug fix: gmscuopt.c creates the problem with cuOptCreateRangedProblem instead of cuOptCreateProblem, since cuOpt's multi-GPU path only reads constraint lower/upper bounds and saw 0 constraints (validation error with presolve 0)
    • skip initial primal/dual solutions for multi-GPU PDLP, cuOpt rejects them
    • README section on usage and limitations
  • Update optcuopt.def to cuOpt 26.10
    • remove 4 options dropped upstream (mip_hyper_heuristic_presolve_time_ratio, mip_hyper_heuristic_presolve_max_time, mip_hyper_diving_min_node_depth, mip_hyper_submip_node_limit_base)
    • new method 4 (primal simplex) and barrier_dual_initial_point 2 (SeDuMi-style), fix changed default of mip_hyper_heuristic_related_vars_time_limit
    • add 18 new options (e.g. concurrent_nnz_cutoff, primal_simplex_pricing, mip_rens, mip_mutation, barrier regularization, Curtis-Reid scaling); sequence_solve left out (Python re-solve cache only)
  • Report iterations and nodes to GAMS via the new solution attributes (cuOptGetSolutionIntAttribute)
    • LP iterations, MIP nodes and simplex iterations; node-limit detection now uses the actual node count
    • guarded by #ifdef, still compiles against 26.08 headers
  • Pass GAMS nodlim to cuOpt node_limit (was ignored before)
  • Bug fix: drop the PDLP reduced-cost workaround in gmscuopt.c, cuOpt now returns correct reduced costs (Correctly return reduced costs for PDLP (stable3) NVIDIA/cuopt#1797)
  • Verified
    • pinned nightly 26.10.0a223 (CUDA 13, single RTX A1000): regression tests 10/10, gamslib test 73/73 match CPLEX
    • PR CI builds and packages CUDA 12 and CUDA 13 on x86_64 and ARM64 with the pin (no solves in CI)
    • earlier against 26.10.0a222: multi-GPU code path (method 1, num_gpus -1, runs even on 1 GPU): 33/35 gamslib LPs match CPLEX within 1e-3; egypt hits the iteration limit (also with single-GPU PDLP), indus89 doesn't converge within 120 s (single-GPU PDLP: optimal in 25 s)
    • iterations/nodes and nodlim checked on trnsport and cube
    • not tested on a machine with several GPUs
  • Not done: cuOptSetLogCallback for a live log, since it only receives lines from the calling thread (a MIP log loses ~75% of its lines)

@0x17 0x17 self-assigned this Oct 2, 2026
@0x17 0x17 changed the title Build against cuOpt nightly and support multi-GPU PDLP Build against cuOpt nightly, add support multi-GPU PDLP, preparations for 26.10 Oct 2, 2026
@0x17 0x17 changed the title Build against cuOpt nightly, add support multi-GPU PDLP, preparations for 26.10 Build against cuOpt 26.10 nightly: multi-GPU PDLP, new options, iteration/node reporting Oct 2, 2026
Comment thread gmscuopt.c
}

status = cuOptCreateProblem(
// Ranged form, since cuOpt's multi-GPU PDLP (without presolve) ignores the row types + RHS

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Huh? please let us know about these types of issues :)
@Bubullzz

@0x17 0x17 Oct 2, 2026 •

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah, you're right. Didn't have the time to properly structure this finding. The issue is now here NVIDIA/cuopt#2042.

@0x17
0x17 marked this pull request as draft October 4, 2026 07:55
@0x17
0x17 marked this pull request as ready for review October 4, 2026 07:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants