Skip to content

[FIX][TIRx][CUDA] Reevaluate conditional wait predicates after each load - #20441

Merged
MasterJH5574 merged 1 commit into
apache:v0.27.0from
jinhongyii:fix/cuda-wait-until-predicate
Sep 25, 2026
Merged

MasterJH5574 merged 1 commit into
apache:v0.27.0from
jinhongyii:fix/cuda-wait-until-predicate

Conversation

@jinhongyii

Copy link
Copy Markdown
Contributor

Fix an infinite loop in CUDA codegen for T.cuda.wait_until predicates containing T.if_then_else. Previously, conditional temporaries were evaluated before the polling macro, so the predicate could keep testing a stale value after the destination was
reloaded.

Keep statements generated for the predicate inside its repeated evaluation, preserving lazy branch semantics. Also correct the API documentation to describe load-before-test behavior.

Add eight GPU regression cases covering signed/unsigned 32/64-bit words, with and without backoff. Tests run in isolated subprocesses with timeouts to contain hangs.

Targets v0.27.0. Validated on TVM 0.27.0rc1 with CUDA 13.2 and B200:

  • Stock rc1 reproduces the regression as a timeout.
  • Wait codegen tests: 25 passed.
  • Broader CUDA codegen tests: 202 passed, 4 environmental skips.
  • Downstream tirx-tools wait tests: 52 passed.
  • Compute Sanitizer synccheck and memcheck: 0 errors.
  • Scoped pre-commit checks and git diff --check: passed.

@jinhongyii
jinhongyii marked this pull request as draft September 24, 2026 22:40
@jinhongyii
jinhongyii marked this pull request as ready for review September 24, 2026 23:40
@jinhongyii

Copy link
Copy Markdown
Contributor Author

cc: @tqchen

@MasterJH5574
MasterJH5574 force-pushed the fix/cuda-wait-until-predicate branch from 3c87b8f to c3073d9 Compare September 25, 2026 01:12
Keep statements emitted while printing wait_until predicates inside the repeated macro expression. Preserve lazy conditional evaluation and leave ordinary predicates as direct expressions. Add bounded-process GPU regressions and correct the load-first API documentation.

[FIX][TIRx][CUDA] Make lazy macro arguments explicit

Record zero-based lazy_args in cuda_func_call attributes and preserve their
evaluation sites in CUDA codegen. Let wait_until mark its predicate instead
of relying on a wait-specific printer branch. Preserve TVMScript round-trip
and cover per-argument, repeated, nested conditional evaluation on GPU.

[TEST][TIRx][CUDA] Remove redundant wait predicate dtype variants

Keep one int32 regression for each backoff macro form. The previous dtype
matrix used identical small positive values and did not exercise distinct
signedness or width boundaries.
@MasterJH5574
MasterJH5574 force-pushed the fix/cuda-wait-until-predicate branch from c3073d9 to 6fed867 Compare September 25, 2026 01:12
@MasterJH5574
MasterJH5574 merged commit 6fed867 into apache:v0.27.0 Sep 25, 2026
0 of 5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants