Skip to content

perf: run k2 CTC loss on the GPU instead of the CPU - #77

Open
lumpidu wants to merge 1 commit into
Stylish-TTS:mainfrom
lumpidu:perf/k2_ctc_on_gpu
Open

perf: run k2 CTC loss on the GPU instead of the CPU#77
lumpidu wants to merge 1 commit into
Stylish-TTS:mainfrom
lumpidu:perf/k2_ctc_on_gpu

Conversation

@lumpidu

@lumpidu lumpidu commented May 29, 2026

Copy link
Copy Markdown

CTCLossWithLabelPriors.to() left k2_device hard-set to "cpu", so every step moved the full log_probs tensor from the GPU to the CPU and ran the k2 CTC graph there.
That transfer stalls the step and leaves the GPU idle for most of the alignment stage.

Select the device based on what k2 was built with: use the model device when k2 has CUDA support, and fall back to the CPU otherwise. Operation to() runs during setup of every stage, including stages that never touch the alignment loss, so the k2 import is guarded

A missing k2 keeps the harmless cpu default and only the alignment stage, which anyways needs k2, fails later. The supervision segments stay on the CPU at both DenseFsaVec call sites because k2 requires that argument there regardless of where log_probs lives.

Computed loss is identical to the CPU path, only placement changes.

CTCLossWithLabelPriors.to() left k2_device hard-set to "cpu", so every step
moved the full log_probs tensor from the GPU to the CPU and ran the k2 CTC
graph there. That transfer stalls the step and leaves the GPU idle for most
of the alignment stage.

Select the device based on what k2 was built with: use the model device
when k2 has CUDA support, and fall back to the CPU otherwise. to() runs
during setup of every stage, including stages that never touch the
alignment loss, so the k2 import is guarded; missing k2 keeps the harmless
cpu default and only the alignment stage, which genuinely needs k2, fails
later. The supervision segments stay on the CPU at both DenseFsaVec call
sites because k2 requires that argument there regardless of where log_probs
lives.

The computed loss is identical to the CPU path; only the placement changes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant