Skip to content

[FIX][TIRx][CUDA] Support low-precision, tensor-map and cluster kernels in CUDA-host bundles - #20490

Merged
tqchen merged 1 commit into
apache:mainfrom
spectrometerHBH:fix/cuda-host-cxx-launch
Sep 29, 2026
Merged

tqchen merged 1 commit into
apache:mainfrom
spectrometerHBH:fix/cuda-host-cxx-launch

Conversation

@spectrometerHBH

Copy link
Copy Markdown
Contributor

A CUDA-host bundle (tvm.backend.cuda.export_cuda_host, #20395) compiles device and host code as one NVCC C++ translation unit. Building the native kernels of mlc-ai/TIRx-kernels this way exposed several patterns that the bundle could not compile or launch:

  • Low-precision types in the host pass. The CUDA device header included cuda_fp16.h, cuda_bf16.h, cuda_fp8.h, cuda_fp6.h and cuda_fp4.h (and defined the fp8_e4_t-style aliases) only under defined(__CUDA_ARCH__). NVCC's host pass then fails on kernel signatures that use half or nv_bfloat16. The guards now also admit the host pass (!defined(__CUDA_ARCH__) || __CUDA_ARCH__ >= N); NVRTC and device-only NVCC compilation always define __CUDA_ARCH__ and are unchanged.
  • Tensor-map parameters. The host wrapper bound a T.TensorMap() parameter as CUtensorMap* x = ((void*)...), which C accepts and C++ rejects. It now casts explicitly.
  • Argument types at the launch. Host C code spells some device types differently (bfloat16 is uint16_t* on the host, nv_bfloat16* in the kernel), so the direct <<<>>> launch did not type-check. Launches now go through a small helper that converts each argument to the kernel's own parameter type and calls cudaLaunchKernelEx.
  • Launch attributes. clusterCtaIdx.*, preferredClusterCtaIdx.*, programmatic dependent launch and cooperative launch were rejected. They now become the corresponding cudaLaunchAttributes, following CUDAWrappedFunc in the CUDA runtime module (including cudaFuncAttributeNonPortableClusterSizeAllowed for cluster launches and omitting a unit preferred cluster). Required block dimensions remain unsupported and are still diagnosed.

Testing

On B200 (sm_100a), CUDA 13.2, this branch:

  • tests/python/codegen/test_target_codegen_cuda.py -k cuda_host: 8 passed (4 tests, each under NVCC and NVRTC device compilation). New tests:
    • bfloat16 buffers with a two-CTA cluster launch: runs and checks both the values and each CTA's cluster rank ([0, 1, 0, 1]);
    • a programmatic dependent launch flag: runs and checks the launch attribute;
    • a tensor-map parameter: compiles the bundle with NVCC.
  • tests/python/codegen/test_target_codegen_cuda.py, tests/python/tirx/codegen/test_codegen_cuda.py, tests/python/tirx-transform/test_tir_transform_split_host_device.py, tests/python/tirx-transform/test_tir_transform_lower_tvm_builtin.py: 487 passed, 8 skipped. The remaining 120 failures are all parametrizations of test_ptx_cp_async, which fails while parsing its TVMScript body (prim._OpEQ receives a string), before any code generation; this change does not touch that path.

The same change is proposed for the v0.27.0 release branch in #20489. There it was also verified end to end with mlc-ai/TIRx-kernels (whose kernels currently target the v0.27.0 script APIs): with mlc-ai/TIRx-kernels#3, all 220 correctness configs of its nine single-GPU native kernels pass when every CUDA compile is built with host="cuda_host", bundled with export_cuda_host, compiled by tvm_ffi.cpp.build_inline and loaded with tvm_ffi.load_module.

…ls in CUDA-host bundles

A CUDA-host bundle compiles device and host code as one NVCC C++
translation unit. Several kernel patterns failed there:

- The low-precision headers and type aliases emitted for CUDA device code
  were guarded by __CUDA_ARCH__, so NVCC's host pass could not parse kernel
  signatures that use half, nv_bfloat16 or the fp8/fp6/fp4 types. Expose
  them to the host pass as well; NVRTC and device-only NVCC compilation are
  unchanged.
- Tensor-map parameters were bound through C's implicit void* conversion,
  which C++ rejects. Cast to CUtensorMap* explicitly.
- Host C types spell some device types differently (bfloat16 is uint16_t),
  so direct <<<>>> launches did not type-check. Launch through a helper
  that converts each argument to the kernel's parameter type and calls
  cudaLaunchKernelEx.
- Cluster, preferred-cluster, programmatic dependent and cooperative launch
  parameters were rejected. Emit the matching launch attributes, following
  the CUDA runtime module. Required block dimensions remain unsupported.
@tqchen
tqchen merged commit 2a11297 into apache:main Sep 29, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants