Skip to content

Latest commit

 

History

History
516 lines (442 loc) · 30.6 KB

File metadata and controls

516 lines (442 loc) · 30.6 KB

CLIF-STITCH — Design

STITCH = Site Training, Integration, Transfer, and Collaborative Harmonization.

A Python package for federated AI/ML training across CLIF consortium sites, where the sites coordinate through a shared storage object (Azure Blob) instead of a live server.

Scope — STITCH is a communication + file-sync + logging layer, not an ML framework. It does the mechanism: where files live, polling, orchestration, integrity, the end marker, retries, and syncing files between a local mirror and the blob. It does not do policy: building, training, or aggregating models — those are the lead's own train / aggregate scripts. Crucially, STITCH never reads a model into memory or deserializes it — it dumps two files per pass (an opaque model.<ext> whose format is the user's choice, plus a fixed metrics.json) from the site to storage and back, and returns a status code. The user's own code writes and reads the model file at the paths STITCH provides.


1. Context — why a storage object instead of a server

Classic federated learning uses a parameter server: every site opens a network connection to a central coordinator. In a hospital consortium that means inbound firewall rules, a hosted service, certificates, and a long IT/security approval cycle at every participating institution.

CLIF-STITCH avoids all of that:

  • There is no direct link — sites never connect to each other or to a coordinator.
  • The only thing every site does is outbound HTTPS to one Azure Blob container (which IT already permits).
  • The blob is a passive mailbox. Sites write their model objects and poll/read others'. Coordination is just files appearing in the container.
  • Models are served just-in-time: a site pulls the current global model only when it is about to train, and pushes its result back. Nothing is held open.

The entire behaviour is driven by a single config.yaml — that file is the communication protocol. Sites share the same config; the only thing that differs per site is its own identity and its secret token.

   Lead Site                Azure Blob (rendezvous)              Other Sites
  (e.g. GMU)          ┌──────────────────────────────┐      (UCMC, RUSH, NU, RUMC)
      │               │  global/   <- lead writes     │              │
      │  ── write ──► │  sites/<ID>/ <- each site      │ ◄── read ──  │
      │  ◄─ read  ──  │  orchestration.json (the plan) │  ── write ─► │
      │               └──────────────────────────────┘              │
      └────────────────── no site-to-site link ───────────────────-─┘

2. The federated cycle (decoded from the whiteboard)

Roles: one lead site (aggregator + participant) and N local sites (participants). On the whiteboard GMU is the lead; UCMC, RUSH, NU, RUMC are local sites.

Pass 0  (init)
  Lead         : build initial model        -> write global/pass_000/{model,metrics}
  All sites    : pull global model (JIT)

Pass k = 1..N   ("Fed Pass k")
  Each site    : train locally on its CLIF data           [red = local training]
               : push  sites/<ID>/pass_k/{model, metrics.json}   (metrics written last)
  Lead         : wait for ALL expected sites to report completed (block policy)
               : USER aggregate() (e.g. FedAvg, weighted by n_samples) [blue = fed avg]
               : write global/pass_k/{model, metrics.json}
               : advance orchestration.json to pass k+1
  All sites    : poll global -> pull new model -> repeat

The drawing shows FedAvg, but STITCH doesn't implement it — the lead's own aggregate() does. Every arrow is a file sync (local↔blob), never a socket, and the lead is itself a participant (its own sites/GMU/pass_k/ holds a model too, as on the whiteboard).

★ Insight ───────────────────────────────────── Because there's no server to hold state, the run state lives in the blob as a small orchestration.json (lead-written) plus per-site events/. Actors advance by polling those, not by being pushed to. This is the classic trade-off of blackboard systems: simpler & approval-free infra, at the cost of poll latency and needing explicit markers so the lead knows a pass is complete — and it is exactly what makes any actor freely resumable. ─────────────────────────────────────────────────


3. Folder structure — one tree, mirrored locally and on the blob

Every (role, pass) is a folder holding exactly two files: an opaque model.<ext> (the extension comes from model.format in config) and a fixed-schema metrics.json. STITCH keeps a local mirror whose paths are identical to the blob, so syncing is a dumb copy with no path translation:

./stitch_store/                           # local mirror root (storage.local_dir in config)
└── <run_id>/  ...                         # SAME tree as below — STITCH syncs the two directions

The blob container holds the identical tree (this is the on-the-wire protocol):

<container>/                              # e.g. "fedrun"
└── <run_id>/                             # one federated experiment, e.g. mortality_v1_2026q2
    ├── config.snapshot.yaml              # frozen copy of the shared config for this run
    │
    ├── global/                           # "Lead CPK" on the whiteboard — only the lead writes here
    │   ├── orchestration.json            # THE INSTRUCTION SET: lead writes, everyone reads first
    │   │                                 #   {current_pass, phase, directive, site_status, ...} (§7)
    │   ├── pass_000/                     # Model init
    │   │   ├── model.safetensors         #   opaque model (format from config)
    │   │   └── metrics.json              #   FIXED schema (the completion marker)
    │   ├── pass_001/                     # aggregated global (Fed avg Pass 1)
    │   │   ├── model.safetensors
    │   │   └── metrics.json
    │   ├── ...                           # pass_NNN/  up to num_passes
    │   └── DONE.json                     # END MARKER: run finished -> every site stops on sight
    │                                     #   {final_pass, final_model, reason, final_metrics}
    │
    └── sites/                            # each site owns exactly one subfolder
        ├── GMU/                          # lead also trains -> has its own local models
        │   ├── pass_001/
        │   │   ├── model.safetensors
        │   │   └── metrics.json
        │   └── events/                   # INSERTION FOLDER: append-only JSON reports (§7)
        │       ├── pass_001.started.json
        │       ├── pass_001.completed.json   # "report back": sha256 + n_samples + metrics
        │       └── heartbeat.json            # refreshed every heartbeat_minutes (liveness)
        ├── UCMC/
        │   ├── pass_001/                 # "Pass1" on the whiteboard
        │   │   ├── model.safetensors
        │   │   └── metrics.json
        │   └── events/  ...
        ├── RUSH/  NU/  RUMC/  ...

The local mirror (./stitch_store/<run_id>/) holds this exact tree; STITCH syncs the two.

Rules:

  • Local mirror == blob. The user's code only ever touches the local tree (at paths STITCH hands it); push_* copies a local pass folder up to the blob, pull_* copies a blob pass folder down. Same paths both sides → no mapping, and the local cache is what makes re-runs idempotent and resumable (already_synced).
  • A site writes only inside sites/<its own id>/ — its 2 files per pass plus append-only JSONs under its own events/. The lead is the only writer of global/ (models + orchestration.json + DONE.json).
  • A site reads global/ (the instruction set + the current model); only the lead reads other sites/*/ (its aggregation script + the event rollup).
  • metrics.json is written last and is the completion marker + integrity record + weight. STITCH computes model_sha256 by hashing the file bytes — it does not open the model. Its fixed schema (only n_samples + metrics come from the user; STITCH fills the rest):
{
  "schema_version": 1,
  "pass_id": 1,
  "site_id": "NU",
  "model_file": "model.safetensors",
  "model_format": "safetensors",
  "model_sha256": "ab12…",          // integrity of the paired model file
  "n_samples": 1284,                // REQUIRED — the aggregation weight
  "metrics": { "auc": 0.86, "loss": 0.31 },   // free-form, user's values
  "created_at": "2026-06-24T18:03:00Z"
}

4. Configuration & provisioning — lead generates every config

There are two config shapes, but they are not independent: a site config is a projection of the lead config (the shared protocol: block, copied verbatim) plus that site's identity and scoped token. A human authors exactly one file — the lead config. The lead's CLI derives the blob layout and one ready-to-run config per site.

★ Insight ───────────────────────────────────── Because we chose per-site scoped SAS (§5), a single shared file is impossible — the token must differ per site. The lead must mint N tokens regardless, so generating a full config around each token is nearly free and removes all manual site-side editing. The shared protocol: block is what guarantees lead and sites speak the same protocol even though their files differ. ─────────────────────────────────────────────────

4a. Lead config — stitch.lead.yaml (the ONE human-authored file)

stitch_version: "0.1"
role: lead

protocol:                           # <-- shared core; copied VERBATIM into every site config
  run_id: "mortality_v1_2026q2"     # unique per experiment; also the top folder name
  num_passes: 5                     # number of "Fed Pass" passes
  policy:                           # how the lead orchestrates each pass
    require: all_expected           # wait for EVERY expected site before aggregating (block)
    pass_timeout_minutes: 120       # how long before a non-reporting site is flagged stalled
    heartbeat_minutes: 3            # how often a busy site refreshes its heartbeat.json
  model:
    format: safetensors             # EXTENSION ONLY — STITCH never opens the file
                                    #   (pkl | joblib | npz | safetensors | pt | <any ext>)
  storage:
    backend: azure_blob             # or `localfs` to develop/run the whole loop OFFLINE
    account_url: "https://clifstitch.blob.core.windows.net"   # azure_blob only
    container: "fedrun"                                        # azure_blob only
    local_dir: "./stitch_store"     # local mirror root; paths identical to the blob
    # blob_dir: "./blob"            # localfs ONLY: a directory standing in for the container
                                    #   (no Azure, no secrets — the offline path used by tests)
  layout:                           # path templates -> must match across all sites
    global_dir: "global"
    sites_dir:  "sites"
    pass_dir:   "pass_{pass:03d}"   # nested per-pass folder
    model_name: "model"             # + .<model.format> -> model.safetensors
    metrics_name: "metrics.json"    # fixed-schema marker, written last
    events_dir: "events"            # per-site append-only insertion folder
    orchestration_name: "orchestration.json"   # lead-written instruction set
    end_marker_name: "DONE.json"
  transfer:                         # how STITCH syncs + recovers from failures
    max_retries: 5                  # retries PER failed file op before returning `failed`
    retry_backoff_seconds: 10       # backoff between retries
    poll_seconds: 30                # how often to poll for a new pass / all-expected
  security:
    whitelist: ["model.*", "*.json"]   # ONLY model files + control/event JSONs may sync
    forbid_patterns: ["*.parquet", "*.csv", "*.json.gz"]   # data formats can NEVER leave
    max_file_mb: 512
    verify_sha256: true

sites:                              # the roster the lead names (lead is included)
  - id: GMU
    lead: true                      # the lead also trains locally, like any participant
  - id: UCMC
  - id: RUSH
  - id: NU
  - id: RUMC

minting:                            # how the CLI mints per-site scoped SAS (LEAD-ONLY secret)
  auth: account_key                 # account_key | user_delegation
  account_key_env: "CLIF_STITCH_ACCOUNT_KEY"   # never copied into any site config
  sas_expiry_days: 30
  embed_sas: true                   # write the token INTO each site file (one-artifact handoff)
                                    # set false to instead emit sas_token_env + deliver token separately

4b. Generated site config — stitch.site.<id>.yaml (machine-written, handed to the site)

stitch_version: "0.1"
role: local
generated_by: "clif-stitch gen-configs"   # provenance; CLI also stamps generated_at

protocol:                           # IDENTICAL copy of the lead's protocol: block
  run_id: "mortality_v1_2026q2"
  # ... (verbatim) ...

site:
  site_id: "NU"                     # which sites/<id>/ folder this machine owns
  sas_token: "sv=...&sp=racwl&sig=..."   # LIMITED: read/add/create/write/list, NO delete

The lead's own runnable file (stitch.lead-run.yaml) is generated the same way with role: lead and a full-access token embedded as sas_token (sp=racwdl — also delete). The account key never leaves the minting machine (env only) and is written into no config; only short-lived SAS are distributed. (Set embed_sas: false to instead emit a sas_token_env name per actor and deliver the token out of band.)

4c. CLI — init, launch, gen-configs, run

clif-stitch is added to a uv project (uv add clif-stitch) and run via uv run (see SCAFFOLD.md §1). init scaffolds into the current folder (uv/git/npm convention), so the verb for setting up the blob run is launch.

# Anyone, in a uv project — scaffold model-code skeleton IN PLACE (see SCAFFOLD.md)
uv run clif-stitch init             # writes project/, stitch.lead.yaml, data/, .gitignore here
                                    # (additive: never overwrites uv's pyproject.toml)

# Lead, one-time setup -------------------------------------------------------
uv run clif-stitch launch        --config stitch.lead.yaml
    # writes global/orchestration.json (pass 0, phase=launching) and
    #        config.snapshot.yaml (the protocol: block, NO secrets),
    #        validates blob access. (Azure "folders" are virtual prefixes —
    #        they materialize on first write, so no explicit mkdir is needed.)

uv run clif-stitch gen-configs   --config stitch.lead.yaml --out ./dist/
    # for each site in `sites:`  -> mint scoped SAS + write ./dist/stitch.site.<id>.yaml
    # also writes ./dist/stitch.lead-run.yaml for the lead itself.

# Lead distributes ./dist/stitch.site.<id>.yaml to each site over a SECURE channel,
# then runs the federation:
uv run clif-stitch run --config ./dist/stitch.lead-run.yaml   # aggregate loop + lead's local training

# Each local site (git clone + uv sync, then receive its file) ---------------
uv run clif-stitch run --config stitch.site.<id>.yaml         # local loop; RESUMES if re-run

# Helpers --------------------------------------------------------------------
uv run clif-stitch status --config <any>   # read orchestration.json: pass, phase, per-site state + heartbeat
uv run clif-stitch resume --config <any>   # alias for `run` (auto-resume); --restart to start over
uv run clif-stitch verify --config <any>   # check SAS validity + that whitelist/layout line up
uv run clif-stitch reset  --config stitch.lead.yaml   # (lead) re-init a run folder

run infers the role from role: in the config — there is no --role flag to get wrong. So: humans write one lead file; gen-configs fans out to N ready-to-run site files; sites just run.


5. Security model (this is clinical data — non-negotiable)

★ Insight ───────────────────────────────────── The cardinal rule of clinical federated learning: patient data never leaves the site — only model parameters do. Every security control below exists to make a coding mistake or a malicious file unable to violate that rule, since the blob is a shared surface. ─────────────────────────────────────────────────

  • Whitelist enforcement. clif-stitch refuses to upload or download any object whose name doesn't match security.whitelist, rejects forbid_patterns (parquet/csv = data), and rejects files larger than max_file_mb. A bug that tries to push a cohort parquet is blocked before it touches the network.

  • Two-tier SAS tokens (no Azure CLI). The lead mints both tiers from the account key in Python (generate_container_sas) and embeds each in the matching config:

    • lead — a full container token, sp=racwdl (read/add/create/write/delete/list);
    • site — a limited token, sp=racwl (read/add/create/write/list, no delete), so a site cannot destroy models or the global lane, and tokens expire (sas_expiry_days).

    The account key never leaves the minting machine (env only) — it is written into no config. Honest caveat: a vanilla Blob container SAS limits permissions, not path — it cannot confine a site to only its own sites/<id>/ prefix. Single-writer-of-global/ and each-site-writes-only-its-own-events/ are therefore enforced as a protocol convention (the code only writes within its lane), not by the token on a flat container. True per-folder write isolation needs ADLS Gen2 (hierarchical-namespace directory SAS) or Azure RBAC + managed identity — deferred per §8.

  • Integrity check. Each metrics.json carries the paired model's model_sha256. After a download STITCH re-hashes the bytes and compares before declaring downloaded.

  • STITCH never executes the model. It only copies and hashes bytes — it never imports a ML framework, never unpickles, never loads a model into memory. So a malicious model.pkl cannot run code inside STITCH; the risk lives entirely in the user's loader. The user controls that risk via model.format in config — choosing safetensors/npz means even their own loader never executes code. (Whitelist still blocks anything but model.* + metrics.json + control files, so data files can never sync.)

  • Token expiry. SAS tokens expire; config keeps tokens out of source control and reads them from env / secret store so rotation doesn't require editing shared files.


6. Python package structure

clif-stitch/                          # published so `uv add clif-stitch` works
├── README.md
├── pyproject.toml                    # defines the `clif-stitch` console entry point (cli:main)
├── config.example.yaml
├── docs/
│   └── DESIGN.md                     # this document
├── src/
│   └── clif_stitch/
│       ├── __init__.py               # exports StitchClient, Status
│       ├── config.py                 # pydantic: ProtocolConfig + LeadConfig + SiteConfig
│       ├── cli.py                    # subcommands: init, launch, gen-configs, run, resume, status, verify, reset
│       ├── client.py                 # StitchClient: pull/push + path helpers (the user-facing API)
│       ├── status.py                 # the Status enum: uploaded/downloaded/no_new_pass/
│       │                             #   failed/run_complete/already_synced
│       ├── sync.py                   # copy a pass folder local<->blob, with retry (max_retries)
│       │                             #   + sha256 verify. NEVER opens model.<ext>.
│       ├── resume.py                 # reconciler: compute resume point from mirror + events
│       │
│       ├── scaffold.py               # `clif-stitch init` — copy templates/ into the CURRENT folder
│       ├── templates/                # the project skeleton stamped out by `init` (see SCAFFOLD.md)
│       │   └── project/              #   train.py / aggregate.py / data.py stubs, stitch.lead.yaml
│       │
│       ├── provision/                # LEAD-only: turn the roster into a live run
│       │   ├── launch.py             # create orchestration.json + config.snapshot on the blob
│       │   ├── gen_configs.py        # project lead config -> per-site config files
│       │   └── sas.py                # mint scoped SAS per site (account_key/user_delegation)
│       │
│       ├── storage/                  # pluggable storage backend (bytes in/out only)
│       │   ├── base.py               # StorageBackend ABC: put/get/list/exists
│       │   ├── azure_blob.py         # Azure SAS implementation
│       │   └── localfs.py            # fake backend: run the whole loop offline, no Azure
│       │
│       ├── protocol/                 # the "wire" rules
│       │   ├── layout.py             # path templates from config -> keys (local == blob)
│       │   ├── orchestration.py      # the instruction set: read/write + advance (lead-only writer)
│       │   ├── events.py             # emit/read append-only events + heartbeat (insertion folder)
│       │   └── whitelist.py          # enforce whitelist / forbid / size limits
│       │
│       └── roles/
│           ├── lead.py               # reconcile: skip-if-published, wait-all, aggregate, advance, DONE
│           └── local.py              # reconcile: resume point, pull global, train, push, emit events
└── tests/
    ├── test_config_projection.py     # site config == lead protocol block + (id, sas)
    ├── test_whitelist.py
    ├── test_sync_retry.py            # failure retries up to max_retries then -> failed
    ├── test_status.py                # no_new_pass / already_synced / run_complete branches
    ├── test_resume.py                # killed mid-pass / post-train / lead-crash all resume correctly
    └── test_roundtrip_localfs.py     # run the whole loop against the local-FS fake backend

Note there is no aggregation/ or model.py — STITCH has no FedAvg and no model class. Aggregation and the model live in the user's project (see SCAFFOLD.md).

★ Insight ───────────────────────────────────── StorageBackend as an ABC is the key seam: a LocalFsBackend lets the entire loop run on one laptop (no Azure, no secrets) for tests and demos, while AzureBlobBackend is the production path. sync.py deals only in byte streams + paths, so neither backend ever needs to know what a model is. ─────────────────────────────────────────────────

The training and aggregation are user-supplied hooks — clif-stitch tells the site code where to write its model file and where to read the global from, moves those files, and returns a Status. STITCH owns coordination & transport; the site owns the model, its loader, and its data.


7. Orchestration, lifecycle & resume

Every actor is an idempotent reconciler: it compares desired state (the lead's orchestration.json) against observed state (its local mirror + its events/), and does only the missing work. Re-running clif-stitch run after a crash or a closed terminal is therefore always safe and always converges — that is the resume mechanism.

The instruction set — global/orchestration.json (lead writes, everyone reads first)

{
  "schema_version": 1, "run_id": "mortality_v1_2026q2", "num_passes": 5,
  "current_pass": 2, "phase": "training",   // launching|training|aggregating|complete|failed
  "directive": { "pass": 2, "from_global": "global/pass_001",
                 "expected": ["GMU","UCMC","RUSH","NU"] },   // ALL must report (block policy)
  "site_status": {                          // lead's rollup of events -> one-GET `status`
    "UCMC": {"pass":2,"state":"completed","heartbeat_age_s":5},
    "NU":   {"pass":2,"state":"stalled","heartbeat_age_s":840}
  },
  "updated_at": "..."
}

The insertion folder — sites/<id>/events/ (site appends, lead reads all)

Append-only JSONs, idempotent keys (a re-run overwrites the same file, never duplicates): pass_{NNN}.started.json, pass_{NNN}.completed.json (the "report back" — sha256 + n_samples

  • metrics), pass_{NNN}.failed.json, and heartbeat.json (refreshed every heartbeat_minutes). The append-only design means concurrent site writes never collide and the folder replays to reconstruct progress on resume.
LOCAL reconciler (run/resume):              LEAD reconciler (run/resume):
  read orchestration.json (+ DONE?)           read orchestration.json (+ DONE?)
  k = first pass I haven't completed          k = first pass with no global published
  loop:                                       loop:
    pull_global(k-1)                            train + push_local(k) (it's a participant)
      run_complete -> stop                      wait until ALL `expected` reported completed
    if model_path(k) exists -> just push          (block; stale heartbeat -> flag stalled)
    else emit started; train; heartbeat;        pull_sites(k); USER aggregate -> global/pass_k
         push_local(k); emit completed          advance orchestration.json -> pass k+1
    k += 1                                     write global/DONE.json -> phase=complete
  • Resume granularity = the pass boundary. A completed pass is the durable checkpoint; STITCH does not checkpoint inside a single train call. On resume pull_global(k-1) fetches the current pass's global, so work picks up where it left off — never from pass 0.
  • Block-until-all. The lead aggregates only once every expected site has a completed event. A down site halts that pass; its stale heartbeat is what status surfaces so an operator knows which site to revive. Revival = that operator re-runs clif-stitch run (no actor is remotely controllable — outbound only). The lead is revived the same way; while it's down, sites still finish their pass and wait.
  • Termination. When the lead finishes (N passes or the aggregator signals stop) it writes global/DONE.json; every pull_global checks it first, returns run_complete, and stops.
  • failed (a transfer that exhausted max_retries) aborts the current pass; re-running resumes (already_synced for whatever already uploaded).
  • Where to look on a crash. Canonical state is the local mirror + each site's events/ (exactly what status reads); the verbose stdout/traceback of the user's train/aggregate is teed to runs/<run_id>/ (local-only, never synced). See FLOW §7.

8. MVP scope (suggested first cut)

  1. config.py — ProtocolConfig / LeadConfig / SiteConfig + stitch.lead.example.yaml.
  2. status.py (the enum) + StorageBackend ABC + LocalFsBackend (run the loop offline first).
  3. sync.py — copy a pass folder local↔backend with max_retries + sha256, returning a Status; client.py path helpers (model_path / metrics_path / global_model_path).
  4. protocol/ — layout, orchestration (lead-only writer), events (append-only + heartbeat), whitelist; the DONE.json end marker.
  5. resume.py — compute the resume point from the local mirror + events/; make run resume-by-default (--restart / --from-pass overrides).
  6. provision/ — launch (orchestration.json + snapshot) and gen-configs. SAS minting stubbed on LocalFsBackend first, then AzureBlobBackend + provision/sas.py real scoped tokens.
  7. roles/local.py, roles/lead.py reconcilers (calling USER train/aggregate hooks), cli.py (incl. status + resume), scaffold.py (clif-stitch init).
  8. Example: a tiny mortality model trained via clifpy, driven by a generated lead config, demoing the full multi-pass loop + a mid-run kill/resume on LocalFsBackend.

Defer: additional model.format adapters in the scaffold, user_delegation SAS, config delivery via secure portal, parallel multi-file sync, partial-quorum (require: min_sites) policy.


9. Verification

  • Unit: whitelist rejects *.parquet / oversized files and anything but model.* + metrics.json + control files; a failed transfer retries up to max_retries then returns failed; pull_global returns no_new_pass when absent, already_synced when the local mirror already has it, run_complete when DONE.json exists; gen-configs produces one role: local file per site whose protocol: block is byte-identical to the lead's (no account key); sha256 mismatch on download is rejected.
  • Integration (no cloud): run lead + 3 local roles as separate processes against a shared temp dir via LocalFsBackend; assert the run reaches phase=complete. Termination: write global/DONE.json early and assert every local role's next pull_global returns run_complete and exits.
  • Resume (no cloud): (a) kill a local role mid-train → re-run retrains pass k from global/pass_{k-1}, not 0; (b) kill it after train but before push → re-run pushes the existing model_path(k), no retrain; (c) kill the lead after publishing global/pass_k → re-run skips to k+1; (d) hold one site down → lead blocks at that pass and status shows it stalled via stale heartbeat; on the site's re-run the pass completes and the run advances; (e) --restart ignores the mirror/events and starts at pass 0.
  • End-to-end (cloud): point two sites at a real Azure container with scoped SAS tokens; confirm each can write only its own folder (models + its events/), both read global/, retries survive a forced network blip, and a full run completes.