Skip to content

Latest commit

 

History

History
49 lines (36 loc) · 1.16 KB

File metadata and controls

49 lines (36 loc) · 1.16 KB

Inference benchmark baseline (pre-GPU)

Captured before GPU pipeline rollout. Re-measure after COMPOSE_ENABLE_GPU=true deploy.

Environment (CPU baseline)

Setting Value
VM e2-standard-4, Frankfurt (europe-west3-b)
GPU None
STT STT_DEVICE=cpu, whisper-medium
NMT CPU torch, num_beams=5
TTS OMNIVOICE_DEVICE=cpu, OMNIVOICE_NUM_STEP=32

Sample: 20s video, fullDubbing

Stage Wall-clock % of total
STT ~50s ~3%
NMT ~50s ~3%
TTS ~28 min ~93%
Merge negligible <1%
Total ~30 min 100%

Post-GPU targets (acceptance)

Check Target
End-to-end 20s fullDubbing < 10 min (T4)
TTS share of total < 20%
dablja_stage_duration_seconds p95 Compare in Grafana dablja-pipeline

How to re-measure

histogram_quantile(0.95, sum(rate(dablja_stage_duration_seconds_bucket[5m])) by (le, stage))

LogQL per job:

{service=~"stt|nmt|tts"} |= "YOUR_JOB_ID"

TTS per-segment timing (after code change):

{service="tts"} | json | job_id="YOUR_JOB_ID"