This repo demonstrates an observability-only GCP Batch setup where:
- a team workload container issues Gemini-like HTTPS requests
- a mock Vertex-like Cloud Run service returns deterministic
200and429responses - a platform governance container runs
archodexalongside the workload - Cloud Logging provides the proof that the platform layer saw the workload behavior
In a successful Batch run, the job has one task with two runnables on the same VM. The governance runnable starts as a background container with host-level visibility, then the team workload runs as the foreground container and exits normally.
The workload then issues 60 Gemini-like requests with concurrency 12 and no
retries. The mock service returns 4 successful responses and 56
429 Too Many Requests responses. The governance container sees the responses and emits logs, e.g.:
{
"message": "archodex_http_exchange",
"httpRequest": {
"requestMethod": "POST",
"requestUrl": "https://mock-vertex-xxxx.a.run.app/v1/projects/demo/locations/us-west1/publishers/google/models/gemini-demo:generateContent",
"status": 429
}
}Those governance logs are the proof point of the POC: the platform-owned container can
see the workload's 429 behavior from the Batch VM without changing the
workload image or code.
mock-vertex/ Cloud Run mock service
team-workload/ Python workload container
governance/ Governance wrapper image and Archodex stub
batch/ Batch job template and submission helpers
scripts/ Build and deploy helpers
docs/ High-level and low-level design docs
- Docker with
buildx - Python 3.12+ for local testing
gcloudinstalled and configured against the demo project (see GCP Setup)
The demo runs in an already-bootstrapped project (archodex-batch-demo-20260422).
You're joining that project, not creating one — the steps below verify your
access and confirm the project is configured. If a check fails, ask the project
owner to add you or fix the project-level config; you shouldn't need to create
or modify any resources yourself.
export PROJECT_ID=archodex-batch-demo-20260422
gcloud auth login
gcloud auth application-default login
gcloud config set project "${PROJECT_ID}"Sanity check:
gcloud auth list
gcloud config get-value projectgcloud services list --enabled \
--filter="config.name:(batch.googleapis.com OR run.googleapis.com OR artifactregistry.googleapis.com OR cloudbuild.googleapis.com OR compute.googleapis.com OR logging.googleapis.com)"All six should be listed.
gcloud artifacts repositories list --project="${PROJECT_ID}"Expected: archodex-demo in us-west1.
gcloud projects get-iam-policy "${PROJECT_ID}" \
--flatten="bindings[].members" \
--filter="bindings.members:serviceAccount:batch-runtime@${PROJECT_ID}.iam.gserviceaccount.com" \
--format="value(bindings.role)"Expected (all three):
roles/artifactregistry.readerroles/batch.agentReporterroles/logging.logWriter
PROJECT_NUMBER="$(gcloud projects describe "${PROJECT_ID}" --format='value(projectNumber)')"
gcloud projects get-iam-policy "${PROJECT_ID}" \
--flatten="bindings[].members" \
--filter="bindings.members:serviceAccount:${PROJECT_NUMBER}@cloudbuild.gserviceaccount.com OR bindings.members:serviceAccount:${PROJECT_NUMBER}-compute@developer.gserviceaccount.com" \
--format="table(bindings.members,bindings.role)"At least one of those two service accounts should have roles/cloudbuild.builds.builder.
gcloud auth configure-docker us-west1-docker.pkg.dev --quietThis is a one-time per-machine setup. It's required for ./scripts/build-and-push.sh to push images to Artifact Registry; without it the push step fails with an auth error.
The governance container bundles the archodex-agent binary (eBPF-based network observer, source in the sibling archodex-agent repository). The binaries are gitignored and not produced by any script in this repo, so you must build them and copy them into governance/bin/<arch>/ before docker build.
# From a clone of the archodex-agent repo (sibling directory by convention):
cd ../archodex-agent
cargo build --release --target=aarch64-unknown-linux-gnu
mkdir -p ../archodex-backpressure-demo/governance/bin/arm
cp target/aarch64-unknown-linux-gnu/release/archodex-agent \
../archodex-backpressure-demo/governance/bin/arm/archodex-agent
cargo build --release --target=x86_64-unknown-linux-gnu
mkdir -p ../archodex-backpressure-demo/governance/bin/x86
cp target/x86_64-unknown-linux-gnu/release/archodex-agent \
../archodex-backpressure-demo/governance/bin/x86/archodex-agentThe binary target must match the Batch VM architecture selected by ARCH (arm by default, or x86 for the x86 Batch path). The same arm binary works for local Docker iteration on Apple Silicon.
The x86 Batch path uses a custom Batch-compatible COS 125 VM image:
projects/archodex-batch-demo-20260422/global/images/batch-cos-125-archodex-bmc-20260508-200903
That image is derived from cos-cloud COS 125 and includes the Batch managed-container support files required to run Batch container runnables. It is separate from the governance container image; rebuild and push governance:x86 whenever you want to test a new local x86 archodex-agent binary.
Install the Python dependencies used by the tests:
python3 -m pip install -r requirements-dev.txt \
-r mock-vertex/requirements.txt \
-r team-workload/requirements.txtRun the component test suites separately so their app packages do not collide:
PYTHONPATH=mock-vertex python3 -m pytest mock-vertex/tests
PYTHONPATH=team-workload python3 -m pytest team-workload/tests
python3 -m pytest governance/testsFor faster image and runtime debugging, you can run the demo locally on Docker Desktop without waiting for GCP Batch.
This is a useful approximation for:
- governance image and binary startup
- local mock-service behavior
- workload request generation
- local container wiring between governance and workload
It is not a perfect substitute for Batch host-level observability on GCP.
Build local images:
./scripts/build-local-images.shRun the local harness against a local mock service:
./scripts/run-local-batch.shThat script:
- starts
mock-vertexlocally ifTARGET_URLis not set - runs
governancein privileged mode - runs
team-workloadin the governance container's network and PID namespaces - prints governance and mock logs at the end
To point the local harness at Cloud Run instead of the local mock:
export REGION=us-west1
export MOCK_BASE_URL="$(gcloud run services describe mock-vertex --region="${REGION}" --format='value(status.url)')"
export TARGET_URL="${MOCK_BASE_URL}/v1/projects/demo/locations/us-west1/publishers/google/models/gemini-demo:generateContent"
./scripts/run-local-batch.shUseful overrides:
export TOTAL_REQUESTS=20
export CONCURRENCY=4
export KEEP_CONTAINERS=1Build and push all demo images:
export PROJECT_ID="your-project"
export REGION="us-west1"
export REPO="archodex-demo"
./scripts/build-and-push.shThe build script now publishes only the Batch-side images:
team-workloadgovernance
It publishes arm images by default:
ARCH=arm ./scripts/build-and-push.shTo build x86 images:
ARCH=x86 ./scripts/build-and-push.shThe ARCH switch selects the Docker platform, image tag, and governance agent binary path. Images are tagged as :arm or :x86.
Deploy the mock Vertex service:
./scripts/deploy-mock-vertex.shThat script uses Cloud Build to build mock-vertex remotely, pushes it to
Artifact Registry, and then deploys it to Cloud Run with the Invoker IAM check
disabled so the service is public without needing an allUsers IAM binding.
Render and submit a Batch job:
export MOCK_BASE_URL="$(gcloud run services describe mock-vertex --region="${REGION}" --format='value(status.url)')"
export MOCK_VERTEX_URL="${MOCK_BASE_URL}/v1/projects/demo/locations/us-west1/publishers/google/models/gemini-demo:generateContent"
export ARCH=arm
./batch/render-job.sh
./batch/submit.shUse ARCH=x86 to submit an x86 Batch job. The render script derives the matching machine type, VM image, and image tags unless you explicitly override them with BATCH_MACHINE_TYPE, BATCH_VM_IMAGE_URI, GOVERNANCE_IMAGE, or TEAM_WORKLOAD_IMAGE.
The Batch governance runnable uses the reduced Archodex capability set:
--init --pid=host --cap-add=BPF --cap-add=PERFMON --cap-add=SYS_PTRACE --cap-add=SYS_ADMIN
It does not require --privileged, /proc or /sys bind mounts, or seccomp/AppArmor overrides.