You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[SYCL][NewOffloadModel] Support -fno-sycl-rdc with native_cpu and -fsycl-embed-ir #23407
With the new offload model, -fno-sycl-rdc finalizes each translation unit's device code at compile time: a per-TU clang-linker-wrapper --no-sycl-rdc --emit-fatbin-only job produces a single wrapper module (bitcode), which the host compilation links in via -foffload-include-binary.
Two configurations produce additional host objects in clang-linker-wrapper besides that wrapper module, which the compile-step embedding cannot carry. They are currently rejected in the driver (#22833):
error: '-fno-sycl-rdc' is not supported with '-fsycl-targets=native_cpu' when using the new offloading model
error: '-fno-sycl-rdc' is not supported with '-fsycl-embed-ir' when using the new offloading model
Gaps to address
-fsycl-targets=native_cpu: each split kernel module is compiled to a host object and added to the wrapper output (postLinkProcessModule in ClangLinkerWrapper.cpp, Triple.isNativeCPU() branch). The wrapper module only declares the kernel functions in its __sycl_native_cpu_decls table; their definitions live in those objects.
-fsycl-embed-ir with NVPTX/AMDGCN targets: the embedded-IR image is wrapped and compiled to an extra host object (runWrapperAndCompile(..., IsEmbeddedIR=true) in postLinkProcessModule).
Supporting these with -fno-sycl-rdc requires the compile step to produce a single artifact the host compilation can consume, for example by keeping the device output as bitcode and folding it into the same wrapper module. Once supported, the diagnostic in Driver::CreateOffloadingDeviceToolChains and its tests in clang/test/Driver/sycl-no-rdc-new-driver.cpp should be removed.
Independently of this, any -fno-sycl-rdc compile for NVPTX currently hits a driver assertion in NVPTX::FatBinary::ConstructJob (Cuda.cpp, !GpuArch.empty()), so the -fsycl-embed-ir case is not reachable for NVPTX today.
With the new offload model,
-fno-sycl-rdcfinalizes each translation unit's device code at compile time: a per-TUclang-linker-wrapper --no-sycl-rdc --emit-fatbin-onlyjob produces a single wrapper module (bitcode), which the host compilation links in via-foffload-include-binary.Two configurations produce additional host objects in
clang-linker-wrapperbesides that wrapper module, which the compile-step embedding cannot carry. They are currently rejected in the driver (#22833):Gaps to address
-fsycl-targets=native_cpu: each split kernel module is compiled to a host object and added to the wrapper output (postLinkProcessModuleinClangLinkerWrapper.cpp,Triple.isNativeCPU()branch). The wrapper module only declares the kernel functions in its__sycl_native_cpu_declstable; their definitions live in those objects.-fsycl-embed-irwith NVPTX/AMDGCN targets: the embedded-IR image is wrapped and compiled to an extra host object (runWrapperAndCompile(..., IsEmbeddedIR=true)inpostLinkProcessModule).Supporting these with
-fno-sycl-rdcrequires the compile step to produce a single artifact the host compilation can consume, for example by keeping the device output as bitcode and folding it into the same wrapper module. Once supported, the diagnostic inDriver::CreateOffloadingDeviceToolChainsand its tests inclang/test/Driver/sycl-no-rdc-new-driver.cppshould be removed.Related
-fsycl-targets=native_cpu -fno-sycl-rdc -cexited successfully but embedded an empty device binary (no kernels).-fno-sycl-rdccompile for NVPTX currently hits a driver assertion inNVPTX::FatBinary::ConstructJob(Cuda.cpp,!GpuArch.empty()), so the-fsycl-embed-ircase is not reachable for NVPTX today.