Skip to content

[Runtime] Clear strides when the paged KV cache copies one row of aux data - #20491

Merged
tlopex merged 1 commit into
apache:mainfrom
akaashrp:fix/kv-cache-aux-copy-strides
Sep 30, 2026
Merged

tlopex merged 1 commit into
apache:mainfrom
akaashrp:fix/kv-cache-aux-copy-strides

Conversation

@akaashrp

Copy link
Copy Markdown
Contributor

PlainPagedKVCacheAuxDataManager fills the length info and the compact copy positions one row at a time into a {rows, n} view. CopyVecDataToArray narrows the destination to one dimension for each row but keeps the strides of the two-dimensional view, so CopyDataFromTo rejects it as non-contiguous once n is above one. With sliding window support enabled that is any forward pass over more than one sequence, on every device that uses the plain manager (Metal, Vulkan, OpenCL, WebGPU). MLC LLM hits it when serving concurrent Gemma 4 requests on Metal.

  1. Clear the strides together with the shape in CopyVecDataToArray, as the cached manager already does
  2. Add a test that runs two sequences through a cache with sliding window support on CUDA and Metal

…data

PlainPagedKVCacheAuxDataManager writes the length info and the compact
copy positions one row at a time into a {rows, n} view. CopyVecDataToArray
narrowed the destination to one dimension but kept the strides of the
view, so the contiguity check in CopyDataFromTo failed as soon as n was
above one. That is every forward pass with more than one sequence on a
cache with sliding window support, on any device other than CUDA and CPU.

Clear the strides with the shape, as the cached manager already does.
@akaashrp
akaashrp requested a review from tlopex September 29, 2026 19:53
@tlopex
tlopex merged commit 4d3f3da into apache:main Sep 30, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants