[FIX][TIRx][CUDA] Reevaluate conditional wait predicates after each load - #20441
Merged
MasterJH5574 merged 1 commit intoSep 25, 2026
Merged
Conversation
jinhongyii
marked this pull request as draft
September 24, 2026 22:40
jinhongyii
marked this pull request as ready for review
September 24, 2026 23:40
Contributor
Author
|
cc: @tqchen |
tqchen
approved these changes
Sep 25, 2026
MasterJH5574
force-pushed
the
fix/cuda-wait-until-predicate
branch
from
September 25, 2026 01:12
3c87b8f to
c3073d9
Compare
Keep statements emitted while printing wait_until predicates inside the repeated macro expression. Preserve lazy conditional evaluation and leave ordinary predicates as direct expressions. Add bounded-process GPU regressions and correct the load-first API documentation. [FIX][TIRx][CUDA] Make lazy macro arguments explicit Record zero-based lazy_args in cuda_func_call attributes and preserve their evaluation sites in CUDA codegen. Let wait_until mark its predicate instead of relying on a wait-specific printer branch. Preserve TVMScript round-trip and cover per-argument, repeated, nested conditional evaluation on GPU. [TEST][TIRx][CUDA] Remove redundant wait predicate dtype variants Keep one int32 regression for each backoff macro form. The previous dtype matrix used identical small positive values and did not exercise distinct signedness or width boundaries.
MasterJH5574
force-pushed
the
fix/cuda-wait-until-predicate
branch
from
September 25, 2026 01:12
c3073d9 to
6fed867
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fix an infinite loop in CUDA codegen for
T.cuda.wait_untilpredicates containingT.if_then_else. Previously, conditional temporaries were evaluated before the polling macro, so the predicate could keep testing a stale value after the destination wasreloaded.
Keep statements generated for the predicate inside its repeated evaluation, preserving lazy branch semantics. Also correct the API documentation to describe load-before-test behavior.
Add eight GPU regression cases covering signed/unsigned 32/64-bit words, with and without backoff. Tests run in isolated subprocesses with timeouts to contain hangs.
Targets
v0.27.0. Validated on TVM 0.27.0rc1 with CUDA 13.2 and B200:git diff --check: passed.