Skip to content

v1.8.x MLFlow registry and live GPU metrics - #309

Merged
dferguson992 merged 21 commits into
awslabs:mainfrom
dferguson992:main
Oct 2, 2026
Merged

dferguson992 merged 21 commits into
awslabs:mainfrom
dferguson992:main

Conversation

@dferguson992

@dferguson992 dferguson992 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

v1.8.0 delivers six major themes:

  1. Plain EKS deployment target + serve-layer plugin system — a new eks target that works on any Kubernetes cluster without the HyperPod inference operator, backed by a self-describing per-engine plugin manifest interface
  2. Deploy-time EJS rendering — Kubernetes manifests re-rendered from EJS source at every do/deploy, not frozen at mcc generate time
  3. MLflow model family + dataset registry — MLflow as the system of record for model lineage and adapter sub-models, and the primary read source for dataset tracking (the S3 sidecar remains the durable write record); SageMaker MPG as the deployment/governance target
  4. Live benchmark metrics — plugin-aware background poller — Phase 2 metrics polled throughout the benchmark run, not just at the end; poller configuration driven by serve-layer manifests
  5. Adapter and LoRA cohort — unified do/adapter across targets, HF search in adapter/draft MCPs, LoRA enabled by default on all vLLM deployments
  6. Engine + predictor plugin expansion and template consolidation — a llama.cpp serving engine on the BL105 plugin interface, the HTTP predictor stack reworked into the same self-describing plugin shape, Triton config.pbtxt derived from a single source, dead-template cleanup, and the removal of the dormant marketplace path from the docs

Note on versioning: the v1.8.0 git tag was cut after Themes 1–5. Theme 6 landed on main afterward and is folded into this same release — the published v1.8.0 corresponds to this full document, not the intermediate tag.


Theme 1: Plain EKS Target + Serve-layer Plugin System

BL103 — do/deploy --target eks

Added a new eks deployment target that deploys a standard Kubernetes Deployment + Service + ConfigMap without the HyperPod inference operator or SageMaker endpoint registration. Works on any conformant EKS cluster, including HyperPod EKS clusters where the inference operator is not needed.

do/benchmark, do/test, do/clean, do/optimize, and do/draft all route correctly for the eks target. do/optimize explicitly N/A-exits (no SageMaker endpoint to optimize against). do/draft engine guards allow vLLM/SGLang and now permit tgi/triton with a warning rather than a hard error. 15 new property tests; 4,687 passing.

BL111 — EJS re-render at deploy time for eks target

BL098 (v1.7) applied deploy-time EJS rendering to hyperpod-eks. BL111 extends the same contract to the eks target from day one so it never has the "baked at generate time" class of bug.

The generated eks project ships both eks/*.yaml.ejs (the EJS source, unrendered) and eks/*.yaml (the frozen envsubst fallback). A new do/lib/render-eks-manifests.cjs helper detects node+ejs availability at deploy time; when present, renders from .ejs source with full EJS semantics; otherwise falls back to envsubst over the frozen .yaml with a clear warning. 14 new tests covering all 4 design properties.

BL105 — Serve-layer plugin interface: serve.d/<engine>/manifest.json

Restructured serve.d/ from flat engine wrappers (serve.d/vllm.ejs) to per-engine plugin directories (serve.d/vllm/manifest.json + serve.d/vllm/vllm.ejs). Each engine's manifest is a governed JSON artifact that declares:

  • engine name and env_var_prefix
  • supported_algorithms for speculative decoding
  • algorithm_map (MLCC hyphenated name → engine-specific enum)
  • metrics_endpoint ({path, port, format}) for Phase 2 poller discovery
  • hot_reload boolean
  • Optional dimension_map for benchmark dimension → config key derivation

do/draft algorithm validation now reads supported_algorithms from the active engine's manifest instead of hardcoded case statements. do/deploy.d/hyperpod-eks sources the env-var prefix from the manifest. .optimize_engine.py retires _DIMENSION_CONFIG_KEY_BY_TARGET (resolving the # TODO BL105 comment) and derives config keys from env_var_prefix + dimension_map[dim]. scripts/validate-serve-manifests.js validates all manifests in CI. Manifests are copied to .mlcc/serve.d/ at mcc generate time.

BL107 — SGLang serve-layer plugin (second reference implementation)

Completed the SGLang plugin as the second reference implementation of the BL105 interface. Verified the manifest covers all five SGLang algorithms (eagle3, eagle2, eagle, draft-model → STANDALONE, mtp; notably excludes ngram which SGLang doesn't support). The sglang.ejs wrapper's env-var prefix and deploy mapping are now fully manifest-driven. Fixed the --help text which previously mislabeled SGLang's supported algorithm set.


Theme 2: Deploy-time EJS Rendering

BL111

Covered above under Theme 1.


Theme 3: MLflow Model Family + Dataset Registry

BL109 (spike) — Model family architecture

Conducted a thorough architecture spike researching how base models, fine-tuned flavors, LoRA adapters, and draft models fit into a unified MLflow-first family hierarchy. Key findings incorporated throughout the wave:

  • MLflow has no native model→model derivation edge; the family relationship is a convention layer (tags + params)
  • Critical naming constraint: forward slashes in registered model names silently break MLflow loading (Issue #8801). Use -- as the delimiter: meta-llama--Llama-3.1-8B__adapter__legal-lora
  • parent_run_id is for HPO run trees, not base→derivative lineage; use mlcc.base_model_run_id param instead
  • MetaDataset + mlflow.log_input() is the right primitive for dataset tracking (works on any backend; genai.datasets requires SQL backend)
  • SageMaker MPG is the deployment/governance target, bridged via AutoModelRegistrationEnabled and CustomerMetadataProperties

Design report: workspace/deep_research/model_family_architecture_spike/model_family_architecture_spike.md.

BL-FAM-01/02 — MLflow model family foundation

New shared helper templates/do/lib/python/mlcc_mlflow.py:

  • sanitize_name(hf_id) — replaces / with --, strips spaces; guards register() against the slash bug
  • family_tags(base_id, artifact_type) — mlcc.family, mlcc.artifact_type, mlcc.managed_by
  • family_params(base_id, **kwargs) — mlcc.base_model_id, optional mlcc.base_model_run_id, mlcc.adapter_type, mlcc.training_technique, mlcc.draft_algorithm
  • search_family(base_id) — queries MLflow by mlcc.family tag, exhausts pagination
  • register(...) — sanitize guard + create registered model/version + attach params/aliases
  • log_dataset(source, name, context, meta) — logs MetaDataset as run input, idempotent by (name, digest); carries provenance in source so --list/resolve can reconstruct full entries from MLflow alone
  • _mlflow_configured() — network-free predicate (checks MLFLOW_TRACKING_URI/MLFLOW_TRACKING_SERVER_ARN env + bootstrap config)

30 tests (19 unit + 11 Hypothesis property). MLflow interactions are exercised through mocks/injected clients; the module is structured so importing it never requires MLflow installed.

BL110 — do/register dataset: MLflow-backed registry

Integrated do/register dataset with mlflow.data + mlflow.log_input() + MetaDataset. The two paths are asymmetric by design:

  • Write (do/register dataset): the S3 sidecar (_dataset.json) remains the durable source of truth; when MLflow is configured the dataset is additionally logged as a MetaDataset run input. An MLflow failure here is non-fatal — the persisted S3 registration still stands.
  • Read (--list, name resolution): when MLflow is configured it is the primary source — --list reads MLflow run inputs (list_dataset_inputs) and do/tune --dataset <name> resolves via MLflow (resolve_dataset_by_name). A configured-but-unreachable store is a hard error on these read paths rather than a silent fall-through. The S3 sidecar is the read source only when MLflow is not configured.

The round-trip (log → list → resolve) is covered by the BL110 test suite (MLflow mocked).

BL056 — Adapters as sub-models in MLflow + MPG lineage

do/tune/do/register adapter now writes adapters as <base_id>__adapter__<name> registered models in MLflow, stamped with family tags/params (mlcc.family, mlcc.artifact_type=adapter, mlcc.base_model_run_id). Base run resolution happens before MLflow registration; BASE_RUN_NOT_FOUND exits cleanly before any registration side-effects. SageMaker MPG adapter versions carry mlcc.family + mlcc.base_model_id in CustomerMetadataProperties. do/register adapter --group-by-family provides the MLflow-backed family-grouped listing. 19 adapter-family tests (14 unit + 5 Hypothesis property).


Theme 4: Live Benchmark Metrics — Plugin-aware Background Poller

BL108 — Background Phase 2 metrics poller

Replaced the post-run point-in-time /metrics scrape with a background poller that samples throughout the benchmark run. Architecture is plugin-aware: the poller reads metrics_endpoint from the active engine's serve-layer manifest (BL105); if absent, Phase 2 is silently skipped.

  • templates/do/lib/python/phase2_poller.py — sample_once, aggregate, run_poller, should_spawn, CLI. The vLLM registry in gpu_metrics.py tracks the spec_decode_num_accepted_tokens / spec_decode_num_draft_tokens counter pair; collect_engine_metrics derives the spec_decode_acceptance_rate output (accepted ÷ draft) from them.
  • do/benchmark — two new lifecycle functions: _phase2_start_poller() (port-forward + poller spawn) and _phase2_stop_poller() (join poller → stop port-forward). Spawned at BENCHMARK_RUN_START, stopped at BENCHMARK_RUN_END. trap ... EXIT guarantees cleanup on abnormal exit.
  • Poller writes directly to benchmarks/.last_gpu_metrics.json (fixed location, no temp file lifecycle). Callsites updated to reference this fixed sidecar.
  • Constraint honored: .benchmark_writer.py was not modified.
  • 20/20 tests passing (15 unit + 5 Hypothesis property). Full suite runs in <1s.

Theme 5: Adapter and LoRA Cohort

BL102 — mcc bootstrap update --module <name> scoped redeploy

Added --module <name> flag to mcc bootstrap update. Validates the name against provisioned modules, narrows the CDK deploy to that module only, merges outputs rather than overwriting the full profile on save. Useful when one module's stack is stuck in UPDATE_ROLLBACK_FAILED — the stuck stack can be bypassed by deploying only the module that needs updating. 7 new unit tests.

BL112 — do/adapter unified across targets

Unified do/adapter so add/remove/list/update work on hyperpod-eks via the vLLM LoRA HTTP API (POST /v1/load_lora_adapter, POST /v1/unload_lora_adapter, GET /v1/models), using the same direct-pod port-forward pattern as do/test and do/benchmark. The confusing split between add (SMAI only) and --load-lora (HyperPod only) is gone. Sourcing verbs (from-hub, from-tune, from-train, from-registry) remain SMAI IC-specific with a clear error and guidance on hyperpod-eks. --load-lora/--unload-lora are deprecated with a stderr warning that redirects to the unified verbs. --profile now correctly threaded to all AWS CLI calls in the hyperpod-eks path. 25 unit + 4 property tests.

BL113 — adapter-picker MCP server: HuggingFace adapter search

New bundled MCP server servers/adapter-picker/ with three tools:

  • search_hf_adapters(base_model, task) — searches HF Hub for PEFT/LoRA adapters; exact-match compatibility on base_model_name_or_path
  • get_adapter_metadata(hf_id) — full metadata with adapter type classification (DoRA → QLoRA → LoRA → unknown)
  • recommend_adapter(base_model, task) — top-ranked compatible adapter with all_options

Registered in config/mcp.json and config_loader.py defaults. 11 tests (4 property + 7 example). Existing S3 loading paths untouched.

BL114 — draft-model-picker MCP: S3 model loading

Extended draft-model-picker with S3 support for enterprise environments where draft heads are stored internally. New get_draft_from_s3(s3_uri) tool validates the S3 path contains a valid draft head. do/draft set now allows s3:// URIs (previously hard-rejected). draft-models.json catalog and schema updated with "source": "s3" entries and an optional s3_uri field. 26 tests.

BL115 — LoRA enabled by default on all vLLM targets

ENABLE_LORA=true is now the default in do/config for all vLLM deployment configs. The schema's three LoRA parameters (enableLora, maxLoras, maxLoraRank) now list all supported deployment targets (managed-inference, hyperpod-eks, async-inference, batch-transform, eks) in appliesTo.deploymentTargets. The eks ConfigMap template emits VLLM_ENABLE_LORA: "true" gated on HP_LORA_ENABLED. do/deploy emits a warning when both ENABLE_LORA=true and HP_SPECULATIVE_ALGORITHM are set. ENABLE_LORA=false opt-out preserved. 23 tests. Serve wrapper has no hardcoded --max-loras 4; values flow through the existing VLLM_* forwarding mechanism (30/64 from schema).

BL116 — Benchmark run name (petname)

Every do/benchmark invocation now generates a human-readable run identifier in the form {adjective-noun}_{workload}_{max-concurrency} (e.g. coral-hawk_sample_8). Stored as BENCHMARK_RUN_NAME in do/config, passed to .benchmark_writer.py via --run-name, written to Athena as a run_name column. --list shows a RUN NAME column; --set-baseline accepts the petname as a lookup key alongside the timestamp job name. Old runs show —; new runs show the petname. Schema migration: mcc bootstrap update --module benchmark.

Workload names hyphenated (same session): multi_turn_chat → multi-turn-chat, rag_document_qa → rag-document-qa, etc. across catalog, templates, tests, docs, and CLI.


Theme 6: Engine + Predictor Plugin Expansion and Template Consolidation

This theme landed after the intermediate v1.8.0 tag (see the versioning note in the Summary). It extends the BL105 plugin interface to a second class of engine, generalizes the HTTP predictor stack into the same self-describing shape, pushes more per-target/per-engine knowledge into single declarative sources, and clears dead template code.

llama.cpp serving engine (BL105 plugin interface)

Added templates/code/serve.d/llama-cpp/ as a serve-layer plugin: a manifest.json (env_var_prefix: SM_LLAMA_CPP_, metrics_endpoint {/metrics, 8080, prometheus}, dimension_map, and engine_features for gpu_layers / threads / flash_attn) plus the llama-cpp.ejs wrapper. The deployment config transformers-llama-cpp resolves to { architecture: transformers, backend: llama-cpp }. The engine ships both a SageMaker CPU DLC image (the default, N_GPU_LAYERS=0) and a CUDA image, catalogued in model-servers.json. 9 unit tests (test/unit/llama-cpp-serve-plugin.test.js).

Because llama.cpp is the first engine to ship a CPU-default image, the base-image catalog's two llama-cpp entries use distinct framework_version values (1.0.0-cpu / 1.0.0-cuda) so both survive the version-keyed framework registry rather than one silently overwriting the other; their created dates order the CPU (default) image first.

Predictor framework plugin system

Reworked the HTTP predictor handlers (sklearn / xgboost / tensorflow) from a single branching model_handler.py into per-framework plugin directories under templates/code/predictors.d/<framework>/ (each a manifest.json + handler.py), read by a new src/lib/predictor-manifest-reader.js against templates/code/predictors.d/manifest.schema.json. The monolithic templates/code/model_handler.py and its dead test/test_model_handler.py sglang branch were removed. 10 conformance tests (test/unit/predictor-framework-conformance.test.js); 6 serve-engine registration-drift tests (test/unit/serve-engine-registration-drift.test.js); the deployment-config-resolver suite updated to 19 tests.

Triton config.pbtxt derivation

templates/triton/config.pbtxt is now derived from a single source rather than carrying hand-maintained per-backend duplication, guarded by 8 conformance tests (test/unit/triton-config-pbtxt-conformance.test.js) that fail loudly if the derivation and its consumers drift.

Template cleanup (Tier 1) and marketplace deprecation

Removed dead sibling-template code: the unreachable framework === 'sglang' branch in test/test_model_handler.py and the unused diffusors/start_server.sh (diffusors boots via code/serve; the diffusors arm now unlinks the bulk-copied TensorRT-LLM code/start_server.sh). The deferred Tier 2–4 cleanup items are specced under bl-v19-code-dir-cleanup. The dormant marketplace deployment path — already hard-refused at the generator entry — is now explicitly marked removed in the user-facing docs (configuration.md, deployments.md, index.md); full catalog/resolver removal is tracked separately under bl-marketplace-removal.


Testing

  • Python unit tests: 1,009+ passing across all new modules
  • JS suites: all new property and unit suites passing, including the Theme 6 additions (llama.cpp plugin, predictor-framework conformance, serve-engine registration-drift, Triton config.pbtxt conformance) and the full property suite
  • MLflow: family/dataset/adapter flows covered via mocks and injected clients (no live-store dependency in the gate)
  • End-to-end: eks target generation + deploy-time render verified; project generation re-verified across http, a transformers engine, a Triton backend, and diffusors after the Theme 6 template changes; poller full suite <1s

Upgrade Notes

Athena schema migration required

v1.8.0 adds one new Athena column (run_name). Run after updating:

mcc bootstrap update --module benchmark

npm link recommended for template development

cd /path/to/ml-container-creator && npm link

mcc regenerate now preserves runtime-written config vars automatically via RUNTIME_OWNED_VARS in regenerate-command-handler.js — no more manual re-entry after regeneration.

serve.d/ directory structure changed

serve.d/vllm.ejs and serve.d/sglang.ejs moved to serve.d/vllm/vllm.ejs and serve.d/sglang/sglang.ejs. EJS include paths updated accordingly. The **/serve.d/** ignore glob already excluded wrappers from generated output — no change to generated projects. Theme 6 adds further plugin directories under the same convention (serve.d/llama-cpp/, and the predictor stack under code/predictors.d/<framework>/).

New transformers-llama-cpp deployment config

A llama.cpp serving option is now selectable (--deployment-config=transformers-llama-cpp). It serves GGUF models via llama-server (OpenAI-compatible) and defaults to the SageMaker CPU DLC image; set SM_LLAMA_CPP_N_GPU_LAYERS=-1 on the CUDA image for full GPU offload.

marketplace path marked removed in docs

The marketplace deployment config (already hard-refused at the generator entry point in prior releases) is now explicitly documented as removed. Both --deployment-config=marketplace and the marketplace:// model-name prefix are refused; build and deploy your own image with a HuggingFace id, s3:// artifact, or registry:// package instead.

@github-actions github-actions Bot added documentation Improvements or additions to documentation generator tests ci dependencies labels Sep 28, 2026
@dferguson992
dferguson992 merged commit ba50442 into awslabs:main Oct 2, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci dependencies documentation Improvements or additions to documentation generator tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants