This document is a step-by-step plan to enable NVIDIA GPU acceleration for the DabljaAR AI pipeline on bare metal, local Docker, and GCP (Terraform) hosts.
It covers host drivers, Docker / Compose overlays, runtime environment variables, Terraform compute settings, validation, and troubleshooting.
| Service | Model | GPU benefit | CPU fallback |
|---|---|---|---|
| stt-service | faster-whisper (Whisper) | High — 20s audio can drop from minutes to seconds | Yes (STT_DEVICE=cpu) |
| tts-service | OmniVoice | High — synthesis is the slowest stage in full dubbing | Yes (OMNIVOICE_DEVICE=cpu) |
| nmt-service | NLLB-200 | Moderate — auto-uses CUDA when torch sees a GPU | Yes (CPU torch in default image) |
| media-service | FFmpeg mux only | Optional — NVENC/NVDEC via Dockerfile.gpu |
Yes (CPU FFmpeg) |
Services that do not need GPU: backend, orchestrator, postgres, rabbitmq, caddy.
- Whisper STT runs on CUDA inside the container.
- OmniVoice TTS runs on CUDA inside the container.
- NMT uses GPU when a CUDA PyTorch wheel is present.
- Docker Compose passes GPU devices into worker containers.
- GCP VM is provisioned with a GPU, NVIDIA drivers, and NVIDIA Container Toolkit.
Understanding what already exists avoids duplicate work.
| Path | Purpose |
|---|---|
| stt-service/Dockerfile | CPU torch via libs/docker-torch/install_cpu.sh |
| stt-service/Dockerfile.gpu | nvidia/cuda:12.3.2-runtime base + cu118 torch (device probe) + faster-whisper/CTranslate2 |
| tts-service/Dockerfile | CPU image (OmniVoice) |
| tts-service/Dockerfile.gpu | cu118 torch 2.4.0+cu118 + torchaudio 2.4.0+cu118 + verify_gpu_torch.py |
| nmt-service/Dockerfile | CPU torch via libs/docker-torch/install_cpu.sh |
| nmt-service/Dockerfile.gpu | cu118 torch 2.6.0+cu118 + verify_gpu_torch.py (pip-bundled CUDA 11 libs) |
| media-service/Dockerfile.gpu | CUDA runtime base for FFmpeg NVENC (optional) |
| File | Status |
|---|---|
| docker-compose.gpu.yml | Dev overlay — STT + TTS only; uses gpus: all |
| docker-compose.microservices.prod.yml | Production microservices — defaults to CPU |
docker-compose.microservices.gpu.yml |
Production GPU overlay — merged when COMPOSE_ENABLE_GPU=true |
| Resource | GPU support |
|---|---|
| infra/terraform/modules/compute/main.tf | guest_accelerator block when gpu_count > 0 |
| infra/terraform/scripts/vm-bootstrap.sh.tftpl | Installs NVIDIA Container Toolkit when nvidia-smi exists |
| Default boot image | common-cu128-ubuntu-2204-nvidia-570 (Deep Learning VM, drivers preinstalled) |
docker-compose.gpu.ymlstill sets legacySILMA_DEVICE=cuda— TTS now usesOMNIVOICE_DEVICE(see §6).- No production GPU overlay for
docker-compose.microservices.prod.yml. - NMT GPU Dockerfile —
nmt-service/Dockerfile.gpuexists; uses pip-bundled cu118 libs (not CUDA 12 runtime base). - Single-GPU contention — STT, NMT, and TTS share one GPU unless you scale out or serialize stages (orchestrator already runs stages sequentially per job).
flowchart TB
subgraph host [GPU Host]
Driver[NVIDIA Driver]
NCT[NVIDIA Container Toolkit]
Docker[Docker Engine]
end
subgraph compose [Docker Compose stack]
STT[stt-service\nDockerfile.gpu]
NMT[nmt-service\nCPU or GPU image]
TTS[tts-service\nDockerfile.gpu]
Media[media-service\noptional GPU]
end
Driver --> NCT
NCT --> Docker
Docker -->|"gpus: all"| STT
Docker -->|"gpus: all"| NMT
Docker -->|"gpus: all"| TTS
Docker --> Media
Principle: GPU wheels are baked into *.gpu Dockerfiles; Compose assigns devices with gpus: all (or a device list); runtime env vars (STT_DEVICE, OMNIVOICE_DEVICE) select the compute device inside the container.
Complete these before building GPU images or running Compose.
- NVIDIA GPU with Compute Capability ≥ 6.1 (Pascal+) for
cu118wheels used in STT/TTS GPU Dockerfiles. - Host driver new enough for CUDA 11.8 userland: ≥ 450.80 (Linux).
Verify on the host:
nvidia-smiExpected: driver version, GPU name, memory summary, no errors.
Required so Docker can pass /dev/nvidia* into containers.
Ubuntu / Debian (manual install):
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey \
| sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -fsSL https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list \
| sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#' \
| sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt-get update
sudo apt-get install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart dockerGCP VM: The Terraform bootstrap script runs the same steps automatically when nvidia-smi is present (vm-bootstrap.sh.tftpl).
GPU support requires Compose v2 with the Compose GPU integration:
docker compose version # v2.x required
docker info | grep -i nvidiadocker run --rm --gpus all nvidia/cuda:12.3.2-base-ubuntu22.04 nvidia-smiIf this fails, fix host driver / toolkit before continuing.
From the repo root, merge the base stack with the GPU overlay:
docker compose \
-f docker-compose.yml \
-f docker-compose.gpu.yml \
up -d --build stt-service tts-serviceFor local dev with MinIO and full stack:
docker compose \
-f docker-compose.yml \
-f docker-compose.gpu.yml \
up -d --buildValidate merged config:
docker compose -f docker-compose.yml -f docker-compose.gpu.yml config \
| grep -A3 'stt-service:' | head -20Confirm gpus and dockerfile: Dockerfile.gpu appear for STT/TTS.
The existing overlay should be extended for OmniVoice and NMT. Target content:
# docker-compose.gpu.yml — recommended shape
services:
stt-service:
build:
dockerfile: Dockerfile.gpu
additional_contexts:
shared: ./libs/dablja-worker
docker_torch: ./libs/docker-torch
environment:
STT_DEVICE: cuda
STT_COMPUTE_TYPE: float16
gpus: all
tts-service:
build:
dockerfile: Dockerfile.gpu
additional_contexts:
shared: ./libs/dablja-worker
docker_torch: ./libs/docker-torch
docker_onnx: ./libs/docker-onnx
environment:
OMNIVOICE_DEVICE: cuda
OMNIVOICE_DTYPE: float16
PREWARM_TTS_MODEL: "true"
gpus: all
nmt-service:
# Optional until nmt-service/Dockerfile.gpu exists — see §7.1
gpus: allReplace SILMA_DEVICE with OMNIVOICE_DEVICE when editing the live file.
Production uses docker-compose.microservices.prod.yml. Create docker-compose.microservices.gpu.yml as a thin overlay:
# docker-compose.microservices.gpu.yml
# Usage:
# docker compose --env-file .env.production \
# -f docker-compose.microservices.prod.yml \
# -f docker-compose.microservices.gpu.yml \
# up -d --build
services:
stt-service:
build:
dockerfile: Dockerfile.gpu
environment:
STT_DEVICE: cuda
STT_COMPUTE_TYPE: float16
gpus: all
tts-service:
build:
dockerfile: Dockerfile.gpu
environment:
OMNIVOICE_DEVICE: cuda
OMNIVOICE_DTYPE: float16
gpus: all
nmt-service:
gpus: allDeploy command:
docker compose --env-file .env.production \
-f docker-compose.microservices.prod.yml \
-f docker-compose.microservices.gpu.yml \
up -d --buildAlso set GPU variables in .env.production (§6) so rebuilds stay consistent.
If the host has multiple GPUs, replace gpus: all with explicit IDs:
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ["0"]
capabilities: [gpu]Or use NVIDIA_VISIBLE_DEVICES=0 in environment for a single service.
Add to .env, .env.production, or Secret Manager env-production.
Default stack: demo-aligned dubbing is enabled (DUBBING_TUNE_ENABLED=true, PIPELINE_SEGMENTS_MODE=tts_focused). Whisper defaults to medium in compose files. Set GROQ_API_KEY for NMT length adjustment (SYNC).
STT_DEVICE=cuda
STT_COMPUTE_TYPE=float16
STT_MODEL_SIZE=medium
PREWARM_STT_MODEL=true| Variable | GPU value | Notes |
|---|---|---|
STT_DEVICE |
cuda or auto |
auto picks CUDA when torch reports it |
STT_COMPUTE_TYPE |
float16 |
Use int8_float32 on older GPUs if unstable |
STT_MODEL_SIZE |
medium / large-v3 |
Larger models need more VRAM |
OMNIVOICE_DEVICE=cuda
OMNIVOICE_DTYPE=float16
OMNIVOICE_NUM_STEP=32
PREWARM_TTS_MODEL=true| Variable | GPU value | Notes |
|---|---|---|
OMNIVOICE_DEVICE |
cuda or auto |
Replaces legacy SILMA_DEVICE |
OMNIVOICE_DTYPE |
float16 |
Use float32 if generation is unstable |
OMNIVOICE_NUM_STEP |
32 |
Lower for speed, higher for quality |
No dedicated device env var — nmt-service/app/model.py uses cuda:0 when torch.cuda.is_available(). GPU requires a CUDA torch wheel inside the image (§7.1).
Keep model weights on the persistent volume:
HF_HOME=/model-cache/hf
HUGGINGFACE_HUB_CACHE=/model-cache/hf/hub
TORCH_HOME=/model-cache/torchSTT uses nvidia/cuda:12.3.2-runtime because CTranslate2 (faster-whisper) requires system libcublas.so.12.
NMT and TTS use python:3.12-slim runtime with cu118 PyTorch wheels — CUDA libs come from pip nvidia-* packages bundled with torch, not the system linker path.
GPU images install CUDA PyTorch before requirements-base.txt so pip does not pull mismatched torch/torchaudio from PyPI. TTS pins both torch and torchaudio from the cu118 index.
Build manually:
docker build -f stt-service/Dockerfile.gpu \
--build-context shared=libs/dablja-worker \
-t dabljaar/stt-service:gpu stt-service
docker build -f nmt-service/Dockerfile.gpu \
--build-context shared=libs/dablja-worker \
--build-context docker_torch=libs/docker-torch \
-t dabljaar/nmt-service:gpu nmt-service
docker build -f tts-service/Dockerfile.gpu \
--build-context shared=libs/dablja-worker \
--build-context docker_torch=libs/docker-torch \
-t dabljaar/tts-service:gpu tts-serviceVerify inside image:
docker run --rm --gpus all dabljaar/nmt-service:gpu \
python -c "import torch; print(torch.__version__, torch.cuda.is_available())"
docker run --rm --gpus all dabljaar/tts-service:gpu \
python -c "import torch, torchaudio; print(torch.__version__, torchaudio.__version__, torch.cuda.is_available())"nmt-service/Dockerfile.gpu is wired in both GPU compose overlays with gpus: all. Build-time verify_gpu_torch.py checks the cu118 wheel and pip nvidia.cublas libs.
NMT auto-selects cuda:0 when torch.cuda.is_available() — no separate device env var.
media-service/Dockerfile.gpu uses nvidia/cuda:12.3.2-runtime-ubuntu22.04 for hardware encode/decode during dubbing merge. Add to a GPU overlay only if merge latency is a bottleneck — ML inference is not the bottleneck here.
Run after deploy on any environment.
nvidia-smi
docker run --rm --gpus all nvidia/cuda:12.3.2-base-ubuntu22.04 nvidia-smidocker exec dabljaar_stt_service python -c \
"import torch; print('STT cuda:', torch.cuda.is_available())"
docker exec dabljaar_tts_service python -c \
"import torch; print('TTS cuda:', torch.cuda.is_available())"
docker exec dabljaar_nmt_service python -c \
"import torch; print('NMT cuda:', torch.cuda.is_available())"curl -s http://localhost:8001/health/model | jq .
curl -s http://localhost:8005/health/model | jq .
curl -s http://localhost:8002/health/model | jq .Expect device fields reporting cuda / cuda:0.
- Upload a short video (~30s) with
output_type=fullDubbingortranslationAndTTS. - Watch worker logs:
docker logs -f dabljaar_stt_service
docker logs -f dabljaar_tts_service- Compare wall-clock STT time in logs (
Transcription done | ... time=) — GPU should be dramatically faster than CPU (e.g. 532s → tens of seconds for 20s audio).
watch -n1 nvidia-smiGCP GPUs are zone-scoped. Recommended: g2-standard-8 with embedded L4 (gpu_count = 0) in europe-west3-a/b/c. Legacy T4 on N1 requires europe-west3-b and gpu_count = 1.
Confirm quota and availability:
gcloud compute accelerator-types list --filter="zone:europe-west3-b"| Profile | machine_type |
gpu_count |
GPU |
|---|---|---|---|
| G2 + L4 (recommended) | g2-standard-8 |
0 |
Embedded L4 (~24 GB) |
| N1 + T4 | n1-standard-8 |
1 |
nvidia-tesla-t4 (16 GB) |
Set in terraform.tfvars:
region = "europe-west3"
zone = "europe-west3-b"
machine_type = "g2-standard-8"
gpu_count = 0 # required for G2 — L4 is embedded
enable_spot = false # GPU spot can preempt — bad for long TTS jobs
boot_disk_image = "projects/deeplearning-platform-release/global/images/family/common-cu129-ubuntu-2204-nvidia-580"
boot_disk_size = 100
data_disk_size = 500 # model cache + Docker layers
models_bucket_suffix = "15fcda3e" # prod: pin existing GCS bucket
models_bucket_location = "EUROPE-WEST2" # if bucket predates compute region change
startup_script_enabled = trueScheduling note: GPU VMs use on_host_maintenance = TERMINATE (compute/main.tf) — host maintenance stops the VM; plan for restarts.
cd infra/terraform
terraform init
terraform plan -var-file=terraform.tfvars
terraform apply -var-file=terraform.tfvarsRequest G2 / L4 quota if apply fails with quota errors (G2 profile). For N1 + T4, request NVIDIA T4 GPUs:
gcloud compute project-info describe --project=YOUR_PROJECT_ID
# GCP Console → IAM & Admin → Quotas → NVIDIA L4 GPUs (G2) or NVIDIA T4 GPUs (N1)INSTANCE="$(terraform output -raw instance_name)"
ZONE="$(terraform output -raw instance_zone)"
# GPU visible on host
gcloud compute ssh "$INSTANCE" --zone="$ZONE" --command='nvidia-smi'
# Bootstrap completed (Docker + toolkit)
gcloud compute ssh "$INSTANCE" --zone="$ZONE" \
--command='sudo test -f /var/lib/vm-bootstrap.done && docker compose version'
# Toolkit wired into Docker
gcloud compute ssh "$INSTANCE" --zone="$ZONE" \
--command='docker run --rm --gpus all nvidia/cuda:12.3.2-base-ubuntu22.04 nvidia-smi'Terraform output shortcut:
terraform output check_gpu_statusIn Secret Manager env-production (or repo .env.production before upload), set GPU variables from §6:
STT_DEVICE=cuda
STT_COMPUTE_TYPE=float16
OMNIVOICE_DEVICE=cuda
OMNIVOICE_DTYPE=float16Re-apply Terraform if the secret is managed by Terraform, then refresh on VM:
gcloud compute ssh "$INSTANCE" --zone="$ZONE" \
--command='sudo google_metadata_script_runner startup'SSH to the VM (or use GitHub Actions deploy workflow) and run:
cd /opt/dabljaar/web # or your deploy path
docker compose --env-file .env.production \
-f docker-compose.microservices.prod.yml \
-f docker-compose.microservices.gpu.yml \
up -d --buildEnsure the deploy workflow references the GPU overlay when gpu_count > 0 (update .github/workflows/deploy-gcp.yml if it only uses the CPU prod file today).
To run without GPU (dev/staging cost savings):
gpu_count = 0Keep the Deep Learning boot image or switch to a standard Ubuntu image; workers use CPU Dockerfiles from prod compose only.
Rough planning for one shared GPU running stages sequentially (orchestrator default):
| Component | Approx VRAM (medium/float16) |
|---|---|
| Whisper medium | 2–4 GB |
| NLLB-600M | 2–3 GB |
| OmniVoice | 4–8+ GB (model dependent) |
T4 (16 GB) fits all three sequentially but not comfortably in parallel. For concurrent multi-user load, use:
- Larger GPU (A100), or
- Separate GPU VMs per worker type, or
- Horizontal scaling with one GPU per replica.
Low-VRAM mitigations (also disables demo tuning if DUBBING_TUNE_ENABLED=false):
DUBBING_TUNE_ENABLED=false
PIPELINE_SEGMENTS_MODE=single
STT_MODEL_SIZE=small
STT_COMPUTE_TYPE=int8
OMNIVOICE_NUM_STEP=24
OMNIVOICE_DTYPE=float16| Symptom | Likely cause | Fix |
|---|---|---|
could not select device driver "" with capabilities: [[gpu]] |
NVIDIA Container Toolkit not installed / Docker not restarted | §4.2 |
torch.cuda.is_available() → False inside container |
CPU image used, or no gpus: in Compose |
Use Dockerfile.gpu + overlay |
CUDA error: no kernel image |
GPU too old for wheel SM target | Use Pascal+ or CPU torch |
| OOM during TTS | OmniVoice + other models exceed VRAM | §10 mitigations; ensure stages are not parallel on one GPU |
nvidia-smi works on host but not in container |
nvidia-ctk runtime configure not run |
Re-run toolkit configure; restart Docker |
| GCP apply fails on GPU | Quota or zone availability | Change zone or request quota |
| VM terminates on maintenance | Expected for GPU instances | Use enable_spot = false; monitor uptime |
| STT still slow after “GPU enable” | STT_DEVICE=cpu in env |
Set STT_DEVICE=cuda and rebuild |
libcublas.so.12 is not found during STT transcribe |
CTranslate2 needs CUDA 12 runtime; slim GPU image lacks it | Rebuild stt-service with Dockerfile.gpu (CUDA 12 runtime base); run validate-gpu-deploy.sh --strict |
| TTS crashes or odd CUDA errors after GPU enable | torch cu118 paired with torchaudio from default PyPI (version skew) |
Rebuild TTS GPU image; expect both 2.4.0+cu118; check version_match: True in validate script |
NMT/TTS torch.cuda.is_available() True but job fails |
Pip CUDA libs not loadable at kernel launch | Run validate-gpu-deploy.sh --strict (CUDA tensor alloc); check worker logs |
| TTS ignores GPU | OMNIVOICE_DEVICE=cpu or CPU image |
Set OMNIVOICE_DEVICE=cuda; use Dockerfile.gpu |
Debug a single service:
docker compose ... up -d --build stt-service
docker logs -f dabljaar_stt_service
docker exec -it dabljaar_stt_service python3.12 -c "
import ctypes
ctypes.CDLL('libcublas.so.12')
from app.model import WhisperModelManager
m = WhisperModelManager()
print('device:', m.device, 'compute:', m.compute_type)
"Use this ordered checklist when rolling out GPU in a new environment.
- GPU hardware or GCP
gpu_count >= 1 -
nvidia-smiOK on host - NVIDIA Container Toolkit installed
-
docker run --gpus all nvidia/cuda:12.3.2-base-ubuntu22.04 nvidia-smiOK
- Update docker-compose.gpu.yml:
OMNIVOICE_DEVICE, STT env, NMTgpus - Create
docker-compose.microservices.gpu.ymlfor production - Add
nmt-service/Dockerfile.gpu - TTS GPU image pins matched
torch+torchaudiocu118 wheels - (Optional) Wire
media-service/Dockerfile.gpufor NVENC
- Set §6 variables in
.env.production/ Secret Manager - Confirm
ai_model_cachevolume has enough disk for HF + torch caches
- Set
machine_type(g2-standard-8+gpu_count=0recommended), disk sizes, andmodels_bucket_suffixin prod interraform.tfvars - Confirm GPU quota in target zone
-
terraform apply - Verify bootstrap +
nvidia-smion VM
- Build with GPU overlay
- Run §8 validation
- Run one full dubbing job; confirm STT/TTS log timings and
nvidia-smiusage
- Update GitHub Actions deploy to pass GPU compose overlay when appropriate
- Document rollback: remove GPU overlay and set
STT_DEVICE=cpu,OMNIVOICE_DEVICE=cpu
- docker_setup.md — general Compose usage and overlay examples
- runbook.md — operational commands (includes legacy GPU env names)
- microservices_lld.md — pipeline architecture
- infra/terraform/README.md — full GCP bootstrap order
- ai_inference_optimization_plan.md — CPU vs GPU performance context
To revert to CPU without removing GPU hardware:
docker compose --env-file .env.production \
-f docker-compose.microservices.prod.yml \
up -d --buildEnsure env:
STT_DEVICE=cpu
OMNIVOICE_DEVICE=cpuCPU images (Dockerfile without .gpu) will be used when the GPU overlay is not merged.