v1.8.x MLFlow registry and live GPU metrics - #309
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
v1.8.0 delivers six major themes:
ekstarget that works on any Kubernetes cluster without the HyperPod inference operator, backed by a self-describing per-engine plugin manifest interfacedo/deploy, not frozen atmcc generatetimedo/adapteracross targets, HF search in adapter/draft MCPs, LoRA enabled by default on all vLLM deploymentsconfig.pbtxtderived from a single source, dead-template cleanup, and the removal of the dormant marketplace path from the docsTheme 1: Plain EKS Target + Serve-layer Plugin System
BL103 —
do/deploy --target eksAdded a new
eksdeployment target that deploys a standard Kubernetes Deployment + Service + ConfigMap without the HyperPod inference operator or SageMaker endpoint registration. Works on any conformant EKS cluster, including HyperPod EKS clusters where the inference operator is not needed.do/benchmark,do/test,do/clean,do/optimize, anddo/draftall route correctly for theekstarget.do/optimizeexplicitly N/A-exits (no SageMaker endpoint to optimize against).do/draftengine guards allow vLLM/SGLang and now permit tgi/triton with a warning rather than a hard error. 15 new property tests; 4,687 passing.BL111 — EJS re-render at deploy time for eks target
BL098 (v1.7) applied deploy-time EJS rendering to
hyperpod-eks. BL111 extends the same contract to theekstarget from day one so it never has the "baked at generate time" class of bug.The generated eks project ships both
eks/*.yaml.ejs(the EJS source, unrendered) andeks/*.yaml(the frozen envsubst fallback). A newdo/lib/render-eks-manifests.cjshelper detectsnode+ejsavailability at deploy time; when present, renders from.ejssource with full EJS semantics; otherwise falls back to envsubst over the frozen.yamlwith a clear warning. 14 new tests covering all 4 design properties.BL105 — Serve-layer plugin interface:
serve.d/<engine>/manifest.jsonRestructured
serve.d/from flat engine wrappers (serve.d/vllm.ejs) to per-engine plugin directories (serve.d/vllm/manifest.json+serve.d/vllm/vllm.ejs). Each engine's manifest is a governed JSON artifact that declares:enginename andenv_var_prefixsupported_algorithmsfor speculative decodingalgorithm_map(MLCC hyphenated name → engine-specific enum)metrics_endpoint({path, port, format}) for Phase 2 poller discoveryhot_reloadbooleandimension_mapfor benchmark dimension → config key derivationdo/draftalgorithm validation now readssupported_algorithmsfrom the active engine's manifest instead of hardcoded case statements.do/deploy.d/hyperpod-ekssources the env-var prefix from the manifest..optimize_engine.pyretires_DIMENSION_CONFIG_KEY_BY_TARGET(resolving the# TODO BL105comment) and derives config keys fromenv_var_prefix + dimension_map[dim].scripts/validate-serve-manifests.jsvalidates all manifests in CI. Manifests are copied to.mlcc/serve.d/atmcc generatetime.BL107 — SGLang serve-layer plugin (second reference implementation)
Completed the SGLang plugin as the second reference implementation of the BL105 interface. Verified the manifest covers all five SGLang algorithms (
eagle3,eagle2,eagle,draft-model→STANDALONE,mtp; notably excludesngramwhich SGLang doesn't support). Thesglang.ejswrapper's env-var prefix and deploy mapping are now fully manifest-driven. Fixed the--helptext which previously mislabeled SGLang's supported algorithm set.Theme 2: Deploy-time EJS Rendering
BL111
Covered above under Theme 1.
Theme 3: MLflow Model Family + Dataset Registry
BL109 (spike) — Model family architecture
Conducted a thorough architecture spike researching how base models, fine-tuned flavors, LoRA adapters, and draft models fit into a unified MLflow-first family hierarchy. Key findings incorporated throughout the wave:
--as the delimiter:meta-llama--Llama-3.1-8B__adapter__legal-loraparent_run_idis for HPO run trees, not base→derivative lineage; usemlcc.base_model_run_idparam insteadMetaDataset+mlflow.log_input()is the right primitive for dataset tracking (works on any backend;genai.datasetsrequires SQL backend)AutoModelRegistrationEnabledandCustomerMetadataPropertiesDesign report:
workspace/deep_research/model_family_architecture_spike/model_family_architecture_spike.md.BL-FAM-01/02 — MLflow model family foundation
New shared helper
templates/do/lib/python/mlcc_mlflow.py:sanitize_name(hf_id)— replaces/with--, strips spaces; guardsregister()against the slash bugfamily_tags(base_id, artifact_type)—mlcc.family,mlcc.artifact_type,mlcc.managed_byfamily_params(base_id, **kwargs)—mlcc.base_model_id, optionalmlcc.base_model_run_id,mlcc.adapter_type,mlcc.training_technique,mlcc.draft_algorithmsearch_family(base_id)— queries MLflow bymlcc.familytag, exhausts paginationregister(...)— sanitize guard + create registered model/version + attach params/aliaseslog_dataset(source, name, context, meta)— logsMetaDatasetas run input, idempotent by(name, digest); carries provenance in source so--list/resolve can reconstruct full entries from MLflow alone_mlflow_configured()— network-free predicate (checksMLFLOW_TRACKING_URI/MLFLOW_TRACKING_SERVER_ARNenv + bootstrap config)30 tests (19 unit + 11 Hypothesis property). MLflow interactions are exercised through mocks/injected clients; the module is structured so importing it never requires MLflow installed.
BL110 —
do/register dataset: MLflow-backed registryIntegrated
do/register datasetwithmlflow.data+mlflow.log_input()+MetaDataset. The two paths are asymmetric by design:do/register dataset): the S3 sidecar (_dataset.json) remains the durable source of truth; when MLflow is configured the dataset is additionally logged as aMetaDatasetrun input. An MLflow failure here is non-fatal — the persisted S3 registration still stands.--list, name resolution): when MLflow is configured it is the primary source —--listreads MLflow run inputs (list_dataset_inputs) anddo/tune --dataset <name>resolves via MLflow (resolve_dataset_by_name). A configured-but-unreachable store is a hard error on these read paths rather than a silent fall-through. The S3 sidecar is the read source only when MLflow is not configured.The round-trip (log → list → resolve) is covered by the BL110 test suite (MLflow mocked).
BL056 — Adapters as sub-models in MLflow + MPG lineage
do/tune/do/register adapternow writes adapters as<base_id>__adapter__<name>registered models in MLflow, stamped with family tags/params (mlcc.family,mlcc.artifact_type=adapter,mlcc.base_model_run_id). Base run resolution happens before MLflow registration;BASE_RUN_NOT_FOUNDexits cleanly before any registration side-effects. SageMaker MPG adapter versions carrymlcc.family+mlcc.base_model_idinCustomerMetadataProperties.do/register adapter --group-by-familyprovides the MLflow-backed family-grouped listing. 19 adapter-family tests (14 unit + 5 Hypothesis property).Theme 4: Live Benchmark Metrics — Plugin-aware Background Poller
BL108 — Background Phase 2 metrics poller
Replaced the post-run point-in-time
/metricsscrape with a background poller that samples throughout the benchmark run. Architecture is plugin-aware: the poller readsmetrics_endpointfrom the active engine's serve-layer manifest (BL105); if absent, Phase 2 is silently skipped.templates/do/lib/python/phase2_poller.py—sample_once,aggregate,run_poller,should_spawn, CLI. The vLLM registry ingpu_metrics.pytracks thespec_decode_num_accepted_tokens/spec_decode_num_draft_tokenscounter pair;collect_engine_metricsderives thespec_decode_acceptance_rateoutput (accepted ÷ draft) from them.do/benchmark— two new lifecycle functions:_phase2_start_poller()(port-forward + poller spawn) and_phase2_stop_poller()(join poller → stop port-forward). Spawned atBENCHMARK_RUN_START, stopped atBENCHMARK_RUN_END.trap ... EXITguarantees cleanup on abnormal exit.benchmarks/.last_gpu_metrics.json(fixed location, no temp file lifecycle). Callsites updated to reference this fixed sidecar..benchmark_writer.pywas not modified.Theme 5: Adapter and LoRA Cohort
BL102 —
mcc bootstrap update --module <name>scoped redeployAdded
--module <name>flag tomcc bootstrap update. Validates the name against provisioned modules, narrows the CDK deploy to that module only, merges outputs rather than overwriting the full profile on save. Useful when one module's stack is stuck inUPDATE_ROLLBACK_FAILED— the stuck stack can be bypassed by deploying only the module that needs updating. 7 new unit tests.BL112 —
do/adapterunified across targetsUnified
do/adaptersoadd/remove/list/updatework on hyperpod-eks via the vLLM LoRA HTTP API (POST /v1/load_lora_adapter,POST /v1/unload_lora_adapter,GET /v1/models), using the same direct-pod port-forward pattern asdo/testanddo/benchmark. The confusing split betweenadd(SMAI only) and--load-lora(HyperPod only) is gone. Sourcing verbs (from-hub,from-tune,from-train,from-registry) remain SMAI IC-specific with a clear error and guidance on hyperpod-eks.--load-lora/--unload-loraare deprecated with a stderr warning that redirects to the unified verbs.--profilenow correctly threaded to all AWS CLI calls in the hyperpod-eks path. 25 unit + 4 property tests.BL113 —
adapter-pickerMCP server: HuggingFace adapter searchNew bundled MCP server
servers/adapter-picker/with three tools:search_hf_adapters(base_model, task)— searches HF Hub for PEFT/LoRA adapters; exact-match compatibility onbase_model_name_or_pathget_adapter_metadata(hf_id)— full metadata with adapter type classification (DoRA → QLoRA → LoRA →unknown)recommend_adapter(base_model, task)— top-ranked compatible adapter withall_optionsRegistered in
config/mcp.jsonandconfig_loader.pydefaults. 11 tests (4 property + 7 example). Existing S3 loading paths untouched.BL114 —
draft-model-pickerMCP: S3 model loadingExtended
draft-model-pickerwith S3 support for enterprise environments where draft heads are stored internally. Newget_draft_from_s3(s3_uri)tool validates the S3 path contains a valid draft head.do/draft setnow allowss3://URIs (previously hard-rejected).draft-models.jsoncatalog and schema updated with"source": "s3"entries and an optionals3_urifield. 26 tests.BL115 — LoRA enabled by default on all vLLM targets
ENABLE_LORA=trueis now the default indo/configfor all vLLM deployment configs. The schema's three LoRA parameters (enableLora,maxLoras,maxLoraRank) now list all supported deployment targets (managed-inference,hyperpod-eks,async-inference,batch-transform,eks) inappliesTo.deploymentTargets. TheeksConfigMap template emitsVLLM_ENABLE_LORA: "true"gated onHP_LORA_ENABLED.do/deployemits a warning when bothENABLE_LORA=trueandHP_SPECULATIVE_ALGORITHMare set.ENABLE_LORA=falseopt-out preserved. 23 tests. Serve wrapper has no hardcoded--max-loras 4; values flow through the existingVLLM_*forwarding mechanism (30/64 from schema).BL116 — Benchmark run name (petname)
Every
do/benchmarkinvocation now generates a human-readable run identifier in the form{adjective-noun}_{workload}_{max-concurrency}(e.g.coral-hawk_sample_8). Stored asBENCHMARK_RUN_NAMEindo/config, passed to.benchmark_writer.pyvia--run-name, written to Athena as arun_namecolumn.--listshows aRUN NAMEcolumn;--set-baselineaccepts the petname as a lookup key alongside the timestamp job name. Old runs show—; new runs show the petname. Schema migration:mcc bootstrap update --module benchmark.Workload names hyphenated (same session):
multi_turn_chat→multi-turn-chat,rag_document_qa→rag-document-qa, etc. across catalog, templates, tests, docs, and CLI.Theme 6: Engine + Predictor Plugin Expansion and Template Consolidation
This theme landed after the intermediate
v1.8.0tag (see the versioning note in the Summary). It extends the BL105 plugin interface to a second class of engine, generalizes the HTTP predictor stack into the same self-describing shape, pushes more per-target/per-engine knowledge into single declarative sources, and clears dead template code.llama.cpp serving engine (BL105 plugin interface)
Added
templates/code/serve.d/llama-cpp/as a serve-layer plugin: amanifest.json(env_var_prefix: SM_LLAMA_CPP_,metrics_endpoint {/metrics, 8080, prometheus},dimension_map, andengine_featuresforgpu_layers/threads/flash_attn) plus thellama-cpp.ejswrapper. The deployment configtransformers-llama-cppresolves to{ architecture: transformers, backend: llama-cpp }. The engine ships both a SageMaker CPU DLC image (the default,N_GPU_LAYERS=0) and a CUDA image, catalogued inmodel-servers.json. 9 unit tests (test/unit/llama-cpp-serve-plugin.test.js).Because llama.cpp is the first engine to ship a CPU-default image, the base-image catalog's two llama-cpp entries use distinct
framework_versionvalues (1.0.0-cpu/1.0.0-cuda) so both survive the version-keyed framework registry rather than one silently overwriting the other; theircreateddates order the CPU (default) image first.Predictor framework plugin system
Reworked the HTTP predictor handlers (sklearn / xgboost / tensorflow) from a single branching
model_handler.pyinto per-framework plugin directories undertemplates/code/predictors.d/<framework>/(each amanifest.json+handler.py), read by a newsrc/lib/predictor-manifest-reader.jsagainsttemplates/code/predictors.d/manifest.schema.json. The monolithictemplates/code/model_handler.pyand its deadtest/test_model_handler.pysglang branch were removed. 10 conformance tests (test/unit/predictor-framework-conformance.test.js); 6 serve-engine registration-drift tests (test/unit/serve-engine-registration-drift.test.js); the deployment-config-resolver suite updated to 19 tests.Triton
config.pbtxtderivationtemplates/triton/config.pbtxtis now derived from a single source rather than carrying hand-maintained per-backend duplication, guarded by 8 conformance tests (test/unit/triton-config-pbtxt-conformance.test.js) that fail loudly if the derivation and its consumers drift.Template cleanup (Tier 1) and marketplace deprecation
Removed dead sibling-template code: the unreachable
framework === 'sglang'branch intest/test_model_handler.pyand the unuseddiffusors/start_server.sh(diffusors boots viacode/serve; the diffusors arm now unlinks the bulk-copied TensorRT-LLMcode/start_server.sh). The deferred Tier 2–4 cleanup items are specced underbl-v19-code-dir-cleanup. The dormantmarketplacedeployment path — already hard-refused at the generator entry — is now explicitly marked removed in the user-facing docs (configuration.md,deployments.md,index.md); full catalog/resolver removal is tracked separately underbl-marketplace-removal.Testing
config.pbtxtconformance) and the full property suiteUpgrade Notes
Athena schema migration required
v1.8.0 adds one new Athena column (
run_name). Run after updating:npm linkrecommended for template developmentmcc regeneratenow preserves runtime-written config vars automatically viaRUNTIME_OWNED_VARSinregenerate-command-handler.js— no more manual re-entry after regeneration.serve.d/ directory structure changed
serve.d/vllm.ejsandserve.d/sglang.ejsmoved toserve.d/vllm/vllm.ejsandserve.d/sglang/sglang.ejs. EJS include paths updated accordingly. The**/serve.d/**ignore glob already excluded wrappers from generated output — no change to generated projects. Theme 6 adds further plugin directories under the same convention (serve.d/llama-cpp/, and the predictor stack undercode/predictors.d/<framework>/).New
transformers-llama-cppdeployment configA llama.cpp serving option is now selectable (
--deployment-config=transformers-llama-cpp). It serves GGUF models viallama-server(OpenAI-compatible) and defaults to the SageMaker CPU DLC image; setSM_LLAMA_CPP_N_GPU_LAYERS=-1on the CUDA image for full GPU offload.marketplace path marked removed in docs
The
marketplacedeployment config (already hard-refused at the generator entry point in prior releases) is now explicitly documented as removed. Both--deployment-config=marketplaceand themarketplace://model-name prefix are refused; build and deploy your own image with a HuggingFace id,s3://artifact, orregistry://package instead.