STITCH = Site Training, Integration, Transfer, and Collaborative Harmonization.
A Python package for federated AI/ML training across CLIF consortium sites, where the sites coordinate through a shared storage object (Azure Blob) instead of a live server.
Scope — STITCH is a communication + file-sync + logging layer, not an ML framework. It does the mechanism: where files live, polling, orchestration, integrity, the end marker, retries, and syncing files between a local mirror and the blob. It does not do policy: building, training, or aggregating models — those are the lead's own
train/aggregatescripts. Crucially, STITCH never reads a model into memory or deserializes it — it dumps two files per pass (an opaquemodel.<ext>whose format is the user's choice, plus a fixedmetrics.json) from the site to storage and back, and returns a status code. The user's own code writes and reads the model file at the paths STITCH provides.
Classic federated learning uses a parameter server: every site opens a network connection to a central coordinator. In a hospital consortium that means inbound firewall rules, a hosted service, certificates, and a long IT/security approval cycle at every participating institution.
CLIF-STITCH avoids all of that:
- There is no direct link — sites never connect to each other or to a coordinator.
- The only thing every site does is outbound HTTPS to one Azure Blob container (which IT already permits).
- The blob is a passive mailbox. Sites write their model objects and poll/read others'. Coordination is just files appearing in the container.
- Models are served just-in-time: a site pulls the current global model only when it is about to train, and pushes its result back. Nothing is held open.
The entire behaviour is driven by a single config.yaml — that file is the
communication protocol. Sites share the same config; the only thing that differs per site
is its own identity and its secret token.
Lead Site Azure Blob (rendezvous) Other Sites
(e.g. GMU) ┌──────────────────────────────┐ (UCMC, RUSH, NU, RUMC)
│ │ global/ <- lead writes │ │
│ ── write ──► │ sites/<ID>/ <- each site │ ◄── read ── │
│ ◄─ read ── │ orchestration.json (the plan) │ ── write ─► │
│ └──────────────────────────────┘ │
└────────────────── no site-to-site link ───────────────────-─┘
Roles: one lead site (aggregator + participant) and N local sites (participants). On the whiteboard GMU is the lead; UCMC, RUSH, NU, RUMC are local sites.
Pass 0 (init)
Lead : build initial model -> write global/pass_000/{model,metrics}
All sites : pull global model (JIT)
Pass k = 1..N ("Fed Pass k")
Each site : train locally on its CLIF data [red = local training]
: push sites/<ID>/pass_k/{model, metrics.json} (metrics written last)
Lead : wait for ALL expected sites to report completed (block policy)
: USER aggregate() (e.g. FedAvg, weighted by n_samples) [blue = fed avg]
: write global/pass_k/{model, metrics.json}
: advance orchestration.json to pass k+1
All sites : poll global -> pull new model -> repeat
The drawing shows FedAvg, but STITCH doesn't implement it — the lead's own aggregate()
does. Every arrow is a file sync (local↔blob), never a socket, and the lead is itself a
participant (its own sites/GMU/pass_k/ holds a model too, as on the whiteboard).
★ Insight ─────────────────────────────────────
Because there's no server to hold state, the run state lives in the blob as a small
orchestration.json (lead-written) plus per-site events/. Actors advance by polling those,
not by being pushed to. This is the classic trade-off of blackboard systems: simpler &
approval-free infra, at the cost of poll latency and needing explicit markers so the lead knows
a pass is complete — and it is exactly what makes any actor freely resumable.
─────────────────────────────────────────────────
Every (role, pass) is a folder holding exactly two files: an opaque model.<ext> (the
extension comes from model.format in config) and a fixed-schema metrics.json. STITCH keeps
a local mirror whose paths are identical to the blob, so syncing is a dumb copy with no
path translation:
./stitch_store/ # local mirror root (storage.local_dir in config)
└── <run_id>/ ... # SAME tree as below — STITCH syncs the two directions
The blob container holds the identical tree (this is the on-the-wire protocol):
<container>/ # e.g. "fedrun"
└── <run_id>/ # one federated experiment, e.g. mortality_v1_2026q2
├── config.snapshot.yaml # frozen copy of the shared config for this run
│
├── global/ # "Lead CPK" on the whiteboard — only the lead writes here
│ ├── orchestration.json # THE INSTRUCTION SET: lead writes, everyone reads first
│ │ # {current_pass, phase, directive, site_status, ...} (§7)
│ ├── pass_000/ # Model init
│ │ ├── model.safetensors # opaque model (format from config)
│ │ └── metrics.json # FIXED schema (the completion marker)
│ ├── pass_001/ # aggregated global (Fed avg Pass 1)
│ │ ├── model.safetensors
│ │ └── metrics.json
│ ├── ... # pass_NNN/ up to num_passes
│ └── DONE.json # END MARKER: run finished -> every site stops on sight
│ # {final_pass, final_model, reason, final_metrics}
│
└── sites/ # each site owns exactly one subfolder
├── GMU/ # lead also trains -> has its own local models
│ ├── pass_001/
│ │ ├── model.safetensors
│ │ └── metrics.json
│ └── events/ # INSERTION FOLDER: append-only JSON reports (§7)
│ ├── pass_001.started.json
│ ├── pass_001.completed.json # "report back": sha256 + n_samples + metrics
│ └── heartbeat.json # refreshed every heartbeat_minutes (liveness)
├── UCMC/
│ ├── pass_001/ # "Pass1" on the whiteboard
│ │ ├── model.safetensors
│ │ └── metrics.json
│ └── events/ ...
├── RUSH/ NU/ RUMC/ ...
The local mirror (./stitch_store/<run_id>/) holds this exact tree; STITCH syncs the two.
Rules:
- Local mirror == blob. The user's code only ever touches the local tree (at paths
STITCH hands it);
push_*copies a local pass folder up to the blob,pull_*copies a blob pass folder down. Same paths both sides → no mapping, and the local cache is what makes re-runs idempotent and resumable (already_synced). - A site writes only inside
sites/<its own id>/— its 2 files per pass plus append-only JSONs under its ownevents/. The lead is the only writer ofglobal/(models +orchestration.json+DONE.json). - A site reads
global/(the instruction set + the current model); only the lead reads othersites/*/(its aggregation script + the event rollup). metrics.jsonis written last and is the completion marker + integrity record + weight. STITCH computesmodel_sha256by hashing the file bytes — it does not open the model. Its fixed schema (onlyn_samples+metricscome from the user; STITCH fills the rest):
{
"schema_version": 1,
"pass_id": 1,
"site_id": "NU",
"model_file": "model.safetensors",
"model_format": "safetensors",
"model_sha256": "ab12…", // integrity of the paired model file
"n_samples": 1284, // REQUIRED — the aggregation weight
"metrics": { "auc": 0.86, "loss": 0.31 }, // free-form, user's values
"created_at": "2026-06-24T18:03:00Z"
}There are two config shapes, but they are not independent: a site config is a
projection of the lead config (the shared protocol: block, copied verbatim) plus that
site's identity and scoped token. A human authors exactly one file — the lead config.
The lead's CLI derives the blob layout and one ready-to-run config per site.
★ Insight ─────────────────────────────────────
Because we chose per-site scoped SAS (§5), a single shared file is impossible — the
token must differ per site. The lead must mint N tokens regardless, so generating a full
config around each token is nearly free and removes all manual site-side editing. The
shared protocol: block is what guarantees lead and sites speak the same protocol even
though their files differ.
─────────────────────────────────────────────────
stitch_version: "0.1"
role: lead
protocol: # <-- shared core; copied VERBATIM into every site config
run_id: "mortality_v1_2026q2" # unique per experiment; also the top folder name
num_passes: 5 # number of "Fed Pass" passes
policy: # how the lead orchestrates each pass
require: all_expected # wait for EVERY expected site before aggregating (block)
pass_timeout_minutes: 120 # how long before a non-reporting site is flagged stalled
heartbeat_minutes: 3 # how often a busy site refreshes its heartbeat.json
model:
format: safetensors # EXTENSION ONLY — STITCH never opens the file
# (pkl | joblib | npz | safetensors | pt | <any ext>)
storage:
backend: azure_blob # or `localfs` to develop/run the whole loop OFFLINE
account_url: "https://clifstitch.blob.core.windows.net" # azure_blob only
container: "fedrun" # azure_blob only
local_dir: "./stitch_store" # local mirror root; paths identical to the blob
# blob_dir: "./blob" # localfs ONLY: a directory standing in for the container
# (no Azure, no secrets — the offline path used by tests)
layout: # path templates -> must match across all sites
global_dir: "global"
sites_dir: "sites"
pass_dir: "pass_{pass:03d}" # nested per-pass folder
model_name: "model" # + .<model.format> -> model.safetensors
metrics_name: "metrics.json" # fixed-schema marker, written last
events_dir: "events" # per-site append-only insertion folder
orchestration_name: "orchestration.json" # lead-written instruction set
end_marker_name: "DONE.json"
transfer: # how STITCH syncs + recovers from failures
max_retries: 5 # retries PER failed file op before returning `failed`
retry_backoff_seconds: 10 # backoff between retries
poll_seconds: 30 # how often to poll for a new pass / all-expected
security:
whitelist: ["model.*", "*.json"] # ONLY model files + control/event JSONs may sync
forbid_patterns: ["*.parquet", "*.csv", "*.json.gz"] # data formats can NEVER leave
max_file_mb: 512
verify_sha256: true
sites: # the roster the lead names (lead is included)
- id: GMU
lead: true # the lead also trains locally, like any participant
- id: UCMC
- id: RUSH
- id: NU
- id: RUMC
minting: # how the CLI mints per-site scoped SAS (LEAD-ONLY secret)
auth: account_key # account_key | user_delegation
account_key_env: "CLIF_STITCH_ACCOUNT_KEY" # never copied into any site config
sas_expiry_days: 30
embed_sas: true # write the token INTO each site file (one-artifact handoff)
# set false to instead emit sas_token_env + deliver token separatelystitch_version: "0.1"
role: local
generated_by: "clif-stitch gen-configs" # provenance; CLI also stamps generated_at
protocol: # IDENTICAL copy of the lead's protocol: block
run_id: "mortality_v1_2026q2"
# ... (verbatim) ...
site:
site_id: "NU" # which sites/<id>/ folder this machine owns
sas_token: "sv=...&sp=racwl&sig=..." # LIMITED: read/add/create/write/list, NO deleteThe lead's own runnable file (stitch.lead-run.yaml) is generated the same way with
role: lead and a full-access token embedded as sas_token (sp=racwdl — also delete).
The account key never leaves the minting machine (env only) and is written into no
config; only short-lived SAS are distributed. (Set embed_sas: false to instead emit a
sas_token_env name per actor and deliver the token out of band.)
clif-stitch is added to a uv project (uv add clif-stitch) and run via uv run
(see SCAFFOLD.md §1). init scaffolds into the current folder (uv/git/npm convention),
so the verb for setting up the blob run is launch.
# Anyone, in a uv project — scaffold model-code skeleton IN PLACE (see SCAFFOLD.md)
uv run clif-stitch init # writes project/, stitch.lead.yaml, data/, .gitignore here
# (additive: never overwrites uv's pyproject.toml)
# Lead, one-time setup -------------------------------------------------------
uv run clif-stitch launch --config stitch.lead.yaml
# writes global/orchestration.json (pass 0, phase=launching) and
# config.snapshot.yaml (the protocol: block, NO secrets),
# validates blob access. (Azure "folders" are virtual prefixes —
# they materialize on first write, so no explicit mkdir is needed.)
uv run clif-stitch gen-configs --config stitch.lead.yaml --out ./dist/
# for each site in `sites:` -> mint scoped SAS + write ./dist/stitch.site.<id>.yaml
# also writes ./dist/stitch.lead-run.yaml for the lead itself.
# Lead distributes ./dist/stitch.site.<id>.yaml to each site over a SECURE channel,
# then runs the federation:
uv run clif-stitch run --config ./dist/stitch.lead-run.yaml # aggregate loop + lead's local training
# Each local site (git clone + uv sync, then receive its file) ---------------
uv run clif-stitch run --config stitch.site.<id>.yaml # local loop; RESUMES if re-run
# Helpers --------------------------------------------------------------------
uv run clif-stitch status --config <any> # read orchestration.json: pass, phase, per-site state + heartbeat
uv run clif-stitch resume --config <any> # alias for `run` (auto-resume); --restart to start over
uv run clif-stitch verify --config <any> # check SAS validity + that whitelist/layout line up
uv run clif-stitch reset --config stitch.lead.yaml # (lead) re-init a run folder
run infers the role from role: in the config — there is no --role flag to get wrong.
So: humans write one lead file; gen-configs fans out to N ready-to-run site files;
sites just run.
★ Insight ─────────────────────────────────────
The cardinal rule of clinical federated learning: patient data never leaves the site —
only model parameters do. Every security control below exists to make a coding mistake or
a malicious file unable to violate that rule, since the blob is a shared surface.
─────────────────────────────────────────────────
-
Whitelist enforcement.
clif-stitchrefuses to upload or download any object whose name doesn't matchsecurity.whitelist, rejectsforbid_patterns(parquet/csv = data), and rejects files larger thanmax_file_mb. A bug that tries to push a cohort parquet is blocked before it touches the network. -
Two-tier SAS tokens (no Azure CLI). The lead mints both tiers from the account key in Python (
generate_container_sas) and embeds each in the matching config:- lead — a full container token,
sp=racwdl(read/add/create/write/delete/list); - site — a limited token,
sp=racwl(read/add/create/write/list, no delete), so a site cannot destroy models or the global lane, and tokens expire (sas_expiry_days).
The account key never leaves the minting machine (env only) — it is written into no config. Honest caveat: a vanilla Blob container SAS limits permissions, not path — it cannot confine a site to only its own
sites/<id>/prefix. Single-writer-of-global/and each-site-writes-only-its-own-events/are therefore enforced as a protocol convention (the code only writes within its lane), not by the token on a flat container. True per-folder write isolation needs ADLS Gen2 (hierarchical-namespace directory SAS) or Azure RBAC + managed identity — deferred per §8. - lead — a full container token,
-
Integrity check. Each
metrics.jsoncarries the paired model'smodel_sha256. After a download STITCH re-hashes the bytes and compares before declaringdownloaded. -
STITCH never executes the model. It only copies and hashes bytes — it never imports a ML framework, never unpickles, never loads a model into memory. So a malicious
model.pklcannot run code inside STITCH; the risk lives entirely in the user's loader. The user controls that risk viamodel.formatin config — choosingsafetensors/npzmeans even their own loader never executes code. (Whitelist still blocks anything butmodel.*+metrics.json+ control files, so data files can never sync.) -
Token expiry. SAS tokens expire; config keeps tokens out of source control and reads them from env / secret store so rotation doesn't require editing shared files.
clif-stitch/ # published so `uv add clif-stitch` works
├── README.md
├── pyproject.toml # defines the `clif-stitch` console entry point (cli:main)
├── config.example.yaml
├── docs/
│ └── DESIGN.md # this document
├── src/
│ └── clif_stitch/
│ ├── __init__.py # exports StitchClient, Status
│ ├── config.py # pydantic: ProtocolConfig + LeadConfig + SiteConfig
│ ├── cli.py # subcommands: init, launch, gen-configs, run, resume, status, verify, reset
│ ├── client.py # StitchClient: pull/push + path helpers (the user-facing API)
│ ├── status.py # the Status enum: uploaded/downloaded/no_new_pass/
│ │ # failed/run_complete/already_synced
│ ├── sync.py # copy a pass folder local<->blob, with retry (max_retries)
│ │ # + sha256 verify. NEVER opens model.<ext>.
│ ├── resume.py # reconciler: compute resume point from mirror + events
│ │
│ ├── scaffold.py # `clif-stitch init` — copy templates/ into the CURRENT folder
│ ├── templates/ # the project skeleton stamped out by `init` (see SCAFFOLD.md)
│ │ └── project/ # train.py / aggregate.py / data.py stubs, stitch.lead.yaml
│ │
│ ├── provision/ # LEAD-only: turn the roster into a live run
│ │ ├── launch.py # create orchestration.json + config.snapshot on the blob
│ │ ├── gen_configs.py # project lead config -> per-site config files
│ │ └── sas.py # mint scoped SAS per site (account_key/user_delegation)
│ │
│ ├── storage/ # pluggable storage backend (bytes in/out only)
│ │ ├── base.py # StorageBackend ABC: put/get/list/exists
│ │ ├── azure_blob.py # Azure SAS implementation
│ │ └── localfs.py # fake backend: run the whole loop offline, no Azure
│ │
│ ├── protocol/ # the "wire" rules
│ │ ├── layout.py # path templates from config -> keys (local == blob)
│ │ ├── orchestration.py # the instruction set: read/write + advance (lead-only writer)
│ │ ├── events.py # emit/read append-only events + heartbeat (insertion folder)
│ │ └── whitelist.py # enforce whitelist / forbid / size limits
│ │
│ └── roles/
│ ├── lead.py # reconcile: skip-if-published, wait-all, aggregate, advance, DONE
│ └── local.py # reconcile: resume point, pull global, train, push, emit events
└── tests/
├── test_config_projection.py # site config == lead protocol block + (id, sas)
├── test_whitelist.py
├── test_sync_retry.py # failure retries up to max_retries then -> failed
├── test_status.py # no_new_pass / already_synced / run_complete branches
├── test_resume.py # killed mid-pass / post-train / lead-crash all resume correctly
└── test_roundtrip_localfs.py # run the whole loop against the local-FS fake backend
Note there is no aggregation/ or model.py — STITCH has no FedAvg and no model class.
Aggregation and the model live in the user's project (see SCAFFOLD.md).
★ Insight ─────────────────────────────────────
StorageBackend as an ABC is the key seam: a LocalFsBackend lets the entire loop run on one
laptop (no Azure, no secrets) for tests and demos, while AzureBlobBackend is the production
path. sync.py deals only in byte streams + paths, so neither backend ever needs to know what
a model is.
─────────────────────────────────────────────────
The training and aggregation are user-supplied hooks — clif-stitch tells the site code
where to write its model file and where to read the global from, moves those files, and
returns a Status. STITCH owns coordination & transport; the site owns the model, its loader,
and its data.
Every actor is an idempotent reconciler: it compares desired state (the lead's
orchestration.json) against observed state (its local mirror + its events/), and does
only the missing work. Re-running clif-stitch run after a crash or a closed terminal is
therefore always safe and always converges — that is the resume mechanism.
{
"schema_version": 1, "run_id": "mortality_v1_2026q2", "num_passes": 5,
"current_pass": 2, "phase": "training", // launching|training|aggregating|complete|failed
"directive": { "pass": 2, "from_global": "global/pass_001",
"expected": ["GMU","UCMC","RUSH","NU"] }, // ALL must report (block policy)
"site_status": { // lead's rollup of events -> one-GET `status`
"UCMC": {"pass":2,"state":"completed","heartbeat_age_s":5},
"NU": {"pass":2,"state":"stalled","heartbeat_age_s":840}
},
"updated_at": "..."
}Append-only JSONs, idempotent keys (a re-run overwrites the same file, never duplicates):
pass_{NNN}.started.json, pass_{NNN}.completed.json (the "report back" — sha256 + n_samples
- metrics),
pass_{NNN}.failed.json, andheartbeat.json(refreshed everyheartbeat_minutes). The append-only design means concurrent site writes never collide and the folder replays to reconstruct progress on resume.
LOCAL reconciler (run/resume): LEAD reconciler (run/resume):
read orchestration.json (+ DONE?) read orchestration.json (+ DONE?)
k = first pass I haven't completed k = first pass with no global published
loop: loop:
pull_global(k-1) train + push_local(k) (it's a participant)
run_complete -> stop wait until ALL `expected` reported completed
if model_path(k) exists -> just push (block; stale heartbeat -> flag stalled)
else emit started; train; heartbeat; pull_sites(k); USER aggregate -> global/pass_k
push_local(k); emit completed advance orchestration.json -> pass k+1
k += 1 write global/DONE.json -> phase=complete
- Resume granularity = the pass boundary. A completed pass is the durable checkpoint;
STITCH does not checkpoint inside a single
traincall. On resumepull_global(k-1)fetches the current pass's global, so work picks up where it left off — never from pass 0. - Block-until-all. The lead aggregates only once every
expectedsite has acompletedevent. A down site halts that pass; its stale heartbeat is whatstatussurfaces so an operator knows which site to revive. Revival = that operator re-runsclif-stitch run(no actor is remotely controllable — outbound only). The lead is revived the same way; while it's down, sites still finish their pass and wait. - Termination. When the lead finishes (
Npasses or the aggregator signals stop) it writesglobal/DONE.json; everypull_globalchecks it first, returnsrun_complete, and stops. failed(a transfer that exhaustedmax_retries) aborts the current pass; re-running resumes (already_syncedfor whatever already uploaded).- Where to look on a crash. Canonical state is the local mirror + each site's
events/(exactly whatstatusreads); the verbose stdout/traceback of the user'strain/aggregateis teed toruns/<run_id>/(local-only, never synced). See FLOW §7.
config.py—ProtocolConfig/LeadConfig/SiteConfig+stitch.lead.example.yaml.status.py(the enum) +StorageBackendABC +LocalFsBackend(run the loop offline first).sync.py— copy a pass folder local↔backend withmax_retries+ sha256, returning aStatus;client.pypath helpers (model_path/metrics_path/global_model_path).protocol/—layout,orchestration(lead-only writer),events(append-only + heartbeat),whitelist; theDONE.jsonend marker.resume.py— compute the resume point from the local mirror +events/; makerunresume-by-default (--restart/--from-passoverrides).provision/—launch(orchestration.json + snapshot) andgen-configs. SAS minting stubbed onLocalFsBackendfirst, thenAzureBlobBackend+provision/sas.pyreal scoped tokens.roles/local.py,roles/lead.pyreconcilers (calling USERtrain/aggregatehooks),cli.py(incl.status+resume),scaffold.py(clif-stitch init).- Example: a tiny mortality model trained via
clifpy, driven by a generated lead config, demoing the full multi-pass loop + a mid-run kill/resume onLocalFsBackend.
Defer: additional model.format adapters in the scaffold, user_delegation SAS, config
delivery via secure portal, parallel multi-file sync, partial-quorum (require: min_sites) policy.
- Unit: whitelist rejects
*.parquet/ oversized files and anything butmodel.*+metrics.json+ control files; a failed transfer retries up tomax_retriesthen returnsfailed;pull_globalreturnsno_new_passwhen absent,already_syncedwhen the local mirror already has it,run_completewhenDONE.jsonexists;gen-configsproduces onerole: localfile per site whoseprotocol:block is byte-identical to the lead's (no account key); sha256 mismatch on download is rejected. - Integration (no cloud): run lead + 3 local roles as separate processes against a shared
temp dir via
LocalFsBackend; assert the run reachesphase=complete. Termination: writeglobal/DONE.jsonearly and assert every local role's nextpull_globalreturnsrun_completeand exits. - Resume (no cloud): (a) kill a local role mid-
train→ re-run retrains pass k fromglobal/pass_{k-1}, not 0; (b) kill it after train but before push → re-run pushes the existingmodel_path(k), no retrain; (c) kill the lead after publishingglobal/pass_k→ re-run skips to k+1; (d) hold one site down → lead blocks at that pass andstatusshows itstalledvia stale heartbeat; on the site's re-run the pass completes and the run advances; (e)--restartignores the mirror/events and starts at pass 0. - End-to-end (cloud): point two sites at a real Azure container with scoped SAS tokens;
confirm each can write only its own folder (models + its
events/), both readglobal/, retries survive a forced network blip, and a full run completes.