Captured before GPU pipeline rollout. Re-measure after COMPOSE_ENABLE_GPU=true deploy.
| Setting | Value |
|---|---|
| VM | e2-standard-4, Frankfurt (europe-west3-b) |
| GPU | None |
| STT | STT_DEVICE=cpu, whisper-medium |
| NMT | CPU torch, num_beams=5 |
| TTS | OMNIVOICE_DEVICE=cpu, OMNIVOICE_NUM_STEP=32 |
| Stage | Wall-clock | % of total |
|---|---|---|
| STT | ~50s | ~3% |
| NMT | ~50s | ~3% |
| TTS | ~28 min | ~93% |
| Merge | negligible | <1% |
| Total | ~30 min | 100% |
| Check | Target |
|---|---|
| End-to-end 20s fullDubbing | < 10 min (T4) |
| TTS share of total | < 20% |
dablja_stage_duration_seconds p95 |
Compare in Grafana dablja-pipeline |
histogram_quantile(0.95, sum(rate(dablja_stage_duration_seconds_bucket[5m])) by (le, stage))
LogQL per job:
{service=~"stt|nmt|tts"} |= "YOUR_JOB_ID"
TTS per-segment timing (after code change):
{service="tts"} | json | job_id="YOUR_JOB_ID"