[BUILD] CUDA_QUANT_PREPROCESS off by default and Adjust CI#29687
Merged
Conversation
Contributor
There was a problem hiding this comment.
Pull request overview
Updates ONNX Runtime’s CUDA build configuration to make the CUDA quantization preprocessing (weight-packing) module opt-in by default, while ensuring key CUDA CI workflows explicitly enable it and improving workflow flag readability.
Changes:
- Switched
onnxruntime_BUILD_CUDA_QUANT_PREPROCESSCMake default fromONtoOFF. - Updated Linux CUDA GitHub Actions workflows to explicitly enable
onnxruntime_BUILD_CUDA_QUANT_PREPROCESS=ONin CUDA builds. - Reformatted long
extra_build_flagsstrings in CUDA workflows into multi-line folded scalars for maintainability, and added a couple of CI-oriented CMake defines in the no-cuDNN workflow.
Reviewed changes
Copilot reviewed 4 out of 4 changed files in this pull request and generated 1 comment.
| File | Description |
|---|---|
| cmake/CMakeLists.txt | Makes CUDA quant preprocess module opt-in by default via a CMake option default change. |
| .github/workflows/linux_cuda_ci.yml | Enables CUDA quant preprocess in Linux CUDA build/test jobs and improves build flag readability. |
| .github/workflows/linux_cuda_plugin_ci.yml | Enables CUDA quant preprocess in the CUDA EP-as-plugin CI build flags. |
| .github/workflows/linux_cuda_no_cudnn.yml | Reformats build flags and adds CI-specific CMake defines for the no-cuDNN CUDA build. |
apsonawane
approved these changes
Jul 16, 2026
tianleiwu
enabled auto-merge (squash)
July 17, 2026 00:11
This was referenced Jul 17, 2026
tianleiwu
added a commit
that referenced
this pull request
Jul 18, 2026
This cherry-picks the following commits for the release: | Commit ID | PR Number | Commit Title | |-----------|-----------|-------------| | dd32f35 | #29590 | Fix libcudart.so.13 hard dependency in pybind module breaking import on CPU-only Linux | | cc44a4d | #29706 | [CUDA] Fix XQA GroupQueryAttention cudaErrorInvalidValue on Blackwell (sm_120) | | 23a7e9d | #29705 | [CUDA] Do not link nvrtc | | ee93f83 | #29711 | [CUDA] Update cuda arch list for packages of cuda 12.8 | | fea45a3 | #29620 | [CUDA] Add cuDNN-free ArgMax/ArgMin/ReduceSum and fix LogSoftmax on plugin EP | | f05b218 | #29624 | Enable Spectre-mitigated MSVC libs for BinSkim builds | | 1c89b86 | #29687 | [BUILD] CUDA_QUANT_PREPROCESS off by default and Adjust CI | | 41bd391 | #29658 | [CUDA] Fix null allocator passed to plugin EP kernel PrePack | | 405fbea | #28896 | Add Windows ARM64 CUDA plugin package and align CUDA metadata/artifact naming | | 308f24c | #29622 | Enable fpA_intB GEMM in CUDA builds and add configurable options | | 16ebc1d | #29731 | [Build] Use GPU pool to unblock CI temporarily | |5911a3a263| #29748 | Add OrtErrorCode::ORT_DEVICE_RESET | |6217f73ec5 | #29663 | Fix plugin EP allocator deleter lifetime | --------- Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com> Co-authored-by: GitHub Copilot <copilot@example.com> Co-authored-by: Edward Chen <18449977+edgchen1@users.noreply.github.com> Co-authored-by: Yen-Shi Wang <yenshiw@nvidia.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This pull request updates the build and CI configuration for CUDA-related workflows and the main CMake options. The main changes are the addition of new CMake build flags to enable CUDA quantization preprocessing, improved formatting for build flags, and a change in the default for the CUDA quant preprocess build option. These updates improve clarity, make it easier to customize builds, and ensure that the CUDA quant preprocess module is only built when explicitly requested.
Build configuration changes:
--cmake_extra_defines onnxruntime_BUILD_CUDA_QUANT_PREPROCESS=ONflag to the CUDA build jobs in.github/workflows/linux_cuda_ci.ymland.github/workflows/linux_cuda_plugin_ci.yml, enabling the CUDA quantization preprocessing module during CI builds. [1] [2] [3]--cmake_extra_defines onnxruntime_QUICK_BUILD=ONand--cmake_extra_defines onnxruntime_USE_FPA_INTB_GEMM=OFFflags to the CUDA no-cudnn build job for faster builds and to disable a specific GEMM implementation. [1] [2]Formatting and maintainability:
extra_build_flagsstrings in workflow YAML files to use multi-line lists for improved readability and easier maintenance. [1] [2] [3]CMake option default change:
onnxruntime_BUILD_CUDA_QUANT_PREPROCESSCMake option fromONtoOFFincmake/CMakeLists.txt, so the CUDA quantization preprocessing module is only built when explicitly enabled.