From d95eba028e7510ae4dbea1ba3a92b6403105e762 Mon Sep 17 00:00:00 2001 From: Alberto Mannari Date: Mon, 21 Sep 2026 14:06:11 +0000 Subject: [PATCH] Don't hold the `PatternCache` lock while freezing the patterns. `PatternCache::get` held `mutex_` across the `FrozenRewritePatternSet` construction. Freezing a PDL module runs a pass manager of its own, which hands its parallel verifier to the MLIRContext's thread pool and waits for it -- and `apply-device-patterns` is nested over modules and then functions, so MLIR's asynchronous pass adaptor runs it on that same pool, many at once, all sharing the one cache the device holds. A pool thread that is waiting for a task group also runs whatever else is queued while it waits, whatever group that task belongs to. So a thread waiting for its own group picks up a sibling function's run of this pass and stops on this lock, which means it is no longer available to finish the group the lock's holder is waiting for. Once the threads that would have finished the freeze are all stopped on the lock, nothing finishes: the compile hangs. `SmartMutex` being recursive hid it, letting the holder re-enter and rebuild the set on top of itself rather than fail loudly. Hold the lock only across the map accesses and build with it released. Two callers that miss together now build a set each and the first one back files it, which is the whole cost; cloning the patterns only reads the device. Co-Authored-By: Claude Opus 5 (1M context) Signed-off-by: Alberto Mannari --- .../KTDFArch/Transforms/ApplyPatterns.cpp | 22 ++++++++++++++----- 1 file changed, 17 insertions(+), 5 deletions(-) diff --git a/lib/Dialect/KTDFArch/Transforms/ApplyPatterns.cpp b/lib/Dialect/KTDFArch/Transforms/ApplyPatterns.cpp index 3f89008..91bdc30 100644 --- a/lib/Dialect/KTDFArch/Transforms/ApplyPatterns.cpp +++ b/lib/Dialect/KTDFArch/Transforms/ApplyPatterns.cpp @@ -201,12 +201,21 @@ auto ktdf_arch::getPatterns(const Device& device, PDLPatternModule& patterns, auto PatternCache::get(const PatternGroups& enabled_groups) -> FrozenRewritePatternSet { - llvm::sys::SmartScopedLock lock(mutex_); - - if (const auto it = map_.find(enabled_groups); it != map_.end()) { - return it->second; + { + llvm::sys::SmartScopedLock lock(mutex_); + if (const auto it = map_.find(enabled_groups); it != map_.end()) { + return it->second; + } } + // Built with the lock released, because freezing a PDL module runs a pass + // manager of its own: it hands the patterns to the MLIRContext's thread pool + // and waits for them. This pass is itself run on that pool, once per + // function, and a pool thread that is waiting for a task group also runs + // whatever else is queued while it waits -- so it can pick up a sibling + // function's run of this pass and arrive back here. Holding the lock across + // that wait deadlocks: the threads that would have finished the freeze stop + // on the lock instead, and the thread that holds it is waiting for them. PDLPatternModule pdl_patterns; const auto num_patterns = getPatterns(getDevice(), pdl_patterns, enabled_groups); @@ -227,7 +236,10 @@ auto PatternCache::get(const PatternGroups& enabled_groups) result = FrozenRewritePatternSet(std::move(pdl_patterns)); } - return map_[enabled_groups] = result; + // Two callers that both miss build a set each; the first one back files + // its own and the others take that, so every caller still gets one set. + llvm::sys::SmartScopedLock lock(mutex_); + return map_.try_emplace(enabled_groups, std::move(result)).first->second; } void PatternCache::registerNativeFunctions(PDLPatternModule& patterns) {