Skip to content

Latest commit

 

History

History
323 lines (258 loc) · 12.1 KB

File metadata and controls

323 lines (258 loc) · 12.1 KB

Platform extension and custom deployment guide

This guide is for an evaluator or developer running UMSH on a host that is not described by the checked-in five-platform deployment. It covers two related cases:

  1. deploying an existing Intel or CUDA backend at a different path or host;
  2. adding a genuinely new platform policy implementation.

Do not edit an archived campaign to make a new host look like a paper host. A new device is an extension result until its hardware, binding, implementation, and workload coordinates are shown to match the paper experiment.

Choose the starting point

Use the closest backend guide first:

  • Intel guide for a supported Intel CPU with an integrated Intel GPU and OpenCL USM;
  • NVIDIA guide for the native ARM64 CUDA backend on Orin or Spark.

The current native compile-time platforms are RAPTOR, ARROW, LUNAR, METEOR, ORIN, and SPARK. UMSH_TARGET_PLATFORM=AUTO detects these on the build host. The CUDA backend is currently selected only for ARM CPU plus CUDA; this is not a generic discrete-NVIDIA or cross-compilation path.

Prepare a native checkout

Every campaign host needs its own Git checkout and build directory. Check out the same commit on every machine and build on the target machine; do not copy a build directory from a different CPU, GPU, driver, or CUDA/OpenCL runtime.

From each repository root:

git rev-parse HEAD
git status --short

cmake -S . -B build-release \
  -DCMAKE_BUILD_TYPE=Release \
  -DUMSH_BUILD_EXAMPLES=ON \
  -DUMSH_BUILD_KERNEL_MODULE=OFF \
  -DUMSH_TARGET_PLATFORM=AUTO
cmake --build build-release --parallel
ctest --test-dir build-release --output-on-failure

The userspace-only build is the portable starting point. Kernel modules are machine prerequisites, not something the campaign runner installs. Build and load an optional module only after the stock path works, using the safety procedure in the backend guide.

Confirm what CMake selected:

grep -E '^UMSH_(BACKEND|RESOLVED_TARGET_PLATFORM):' \
  build-release/CMakeCache.txt
build-release/examples/custom/policy_probe/umsh_policy_probe

An UNKNOWN compile-time platform leaves the backend's low-level APIs available but disables the Table 2 default selector. Do not force the name of a different product to obtain a policy default. Add a real platform definition instead.

Create a deployment manifest

The checked-in platforms.json contains exact author-machine SSH aliases and paths. Keep it as the reference deployment. Put evaluator-specific paths in an ignored copy so the source checkout can remain clean:

mkdir -p build
cp tools/cross_platform/platforms.json build/ae-platforms.json

Edit build/ae-platforms.json and add or replace a platform object. A local Intel entry starts with this shape:

{
  "ae-intel": {
    "transport": "local",
    "backend": "intel",
    "repo_root": "/absolute/path/to/umsh",
    "build_root": "build-release",
    "remote_work_root": "/tmp/umsh-cross-platform",
    "expected_hardware": {
      "platform": "describe the actual CPU and GPU"
    },
    "cases": [
      "intel-correctness-smoke",
      "intel-correctness-4k",
      "intel-correctness-64k",
      "intel-correctness-1m",
      "intel-perf-bind",
      "intel-perf-cpu-read",
      "intel-perf-cpu-write",
      "intel-perf-gpu-read",
      "intel-perf-gpu-write",
      {
        "id": "gpu-contention",
        "suite": "performance",
        "state": "unsupported",
        "profiles": ["full"],
        "result_scope": "UNSUPPORTED",
        "paper_claim": "performance.gpu-contention",
        "reason": "not_registered_until_the_new_platform_mechanism_is_verified",
        "command": []
      },
      "intel-perf-ssd"
    ]
  }
}

This is a conservative platform-object fragment, not a complete manifest. Insert it under the existing top-level platforms object. The contention case starts unsupported because Raptor and an unknown Intel GPU cannot inherit the Lunar/Arrow mechanism by name. Replace it with the intel-perf-contention catalog reference only after the implementation and final mechanism have been verified on the new platform.

For CUDA, use orin-new or spark only as a structural reference. Do not copy platform overrides verbatim. In particular, reset Orin's SSD ready override to the shared cuda-perf-ssd catalog reference, which is external_setup, and re-enable it only after the new path passes the storage procedure below. Re-evaluate every platform-specific unsupported or patched-UVM row as well. For SSH transport, set host and an explicit ssh_options list, and make repo_root the exact remote checkout. The runner does not search alternate paths or deploy/build source.

Copy the closest existing platform rather than silently omitting cases. If a case does not apply, replace that case reference with an explicit object whose state and reason describe the missing capability.

Declare support honestly

Each resolved case has one of three pre-execution states:

Manifest state Use it when
ready the command is non-privileged and all required artifact contracts are declared
unsupported no implementation is registered for this backend/platform coordinate
external_setup the experiment needs privileged mutation, an explicitly selected storage path, exclusive hardware use, or another out-of-band procedure

Use command: [] for unsupported and external_setup, with a concrete reason. Do not turn an unsupported policy into the nearest available instruction. Do not call a raw CUDA cache operator canonical Bypass, or call an uncached CPU binding a DMA bypass implementation.

A runtime requires.loaded_modules gate is appropriate when a non-privileged command is valid only after an operator has preloaded a helper module. If the module is absent, the campaign records SKIP_PRECONDITION rather than running a reduced matrix under the same case name.

Validate before execution

The --manifest option precedes the subcommand:

python3 tools/cross_platform/orchestrate.py \
  --manifest build/ae-platforms.json validate

python3 tools/cross_platform/orchestrate.py \
  --manifest build/ae-platforms.json \
  plan --platform ae-intel --profile smoke

python3 tools/cross_platform/orchestrate.py \
  --manifest build/ae-platforms.json \
  inventory --platform ae-intel --note purpose=ae-preflight --execute

plan and run without --execute make no SSH connection and create no files. Review every expanded command, source path, build path, case state, and reason before executing. Inventory is read-only; missing optional tools remain visible in its raw output.

Bring up correctness first

Run the smoke profile before enabling optional mappings or performance setup:

commit=$(git rev-parse HEAD)
python3 tools/cross_platform/orchestrate.py \
  --manifest build/ae-platforms.json \
  run --platform ae-intel \
  --suite correctness --profile smoke \
  --expected-commit "$commit" --require-clean --execute

Then run the full canonical profile:

python3 tools/cross_platform/orchestrate.py \
  --manifest build/ae-platforms.json \
  run --platform ae-intel \
  --suite correctness --profile full \
  --expected-commit "$commit" --require-clean --execute

The smoke profile uses 4 KiB, one seed, and two rounds. The full profile uses 4 KiB, 64 KiB, and 1 MiB, three fixed seeds, and 20 rounds. Unsupported rows must remain in the matrix. Interpret PASS_IMMEDIATE, PASS_CONVERGED, FAIL_STABLE, skips, and inconclusive results using the correctness matrix reference.

Only after stock correctness is understood should optional UC mappings, patched drivers, raw instruction characterization, or larger retry budgets be introduced. Start a new campaign whenever the source, driver, module, binding, or protocol changes.

Add performance cases

Before a full performance run, follow the backend guide to configure THP and HugeTLB, preload any declared helper modules, select the power/performance mode, and record the original host state. Then inspect and execute the performance suite:

python3 tools/cross_platform/orchestrate.py \
  --manifest build/ae-platforms.json \
  plan --platform ae-intel --suite performance --profile full

python3 tools/cross_platform/orchestrate.py \
  --manifest build/ae-platforms.json \
  run --platform ae-intel \
  --suite performance --profile full \
  --expected-commit "$commit" --require-clean --execute

The orchestrator has no individual-case filter. Use the direct executable commands in intel.md or nvidia.md for a focused diagnostic. Such a manual CSV is not equivalent to a campaign archive unless the same inventory, source identity, timing metadata, and artifact checks are preserved.

Every ready performance case must declare:

  • a non-empty performance_metadata object with direction, result scope, timing boundary, units, and binding/policy classification sources;
  • required artifact paths and real columns emitted by the binary;
  • row count or bounds, unique coordinates, numeric constraints, allowed classification values, and evidence-invalid values;
  • required JSON sidecars when the CSV cannot prove a coordinate by itself.

Never synthesize a missing CSV field in the harness. Put stable case-level context in a sidecar and keep per-row facts in the binary output.

SSD extension cases

Keep a shared SSD catalog case external_setup until a platform-specific file path has been selected and verified. A strict Figure 7 override must retain:

--temp_file={build}/ssd-load-data/umsh_ssd_test.dat
--storage_metadata={work}/storage-metadata.json
--require_ssd=true
--reuse_existing=true
--cache_flush_method=fadvise

Copy the u285h or orin-new override only when the file resolves to the intended physical NVMe. The binary must rediscover the mount and device, pass an aligned 512-byte O_DIRECT write/read/compare probe without fallback, and reject virtual block layers. The campaign artifact gate must additionally reject a sparse data file and validate the complete 48-row matrix. SATA, USB, RAID, multi-device, or diagnostic --require_ssd=false results need a separately named extension scope; do not weaken the paper case contract.

Add a genuinely new platform

If AUTO cannot identify a platform that needs new Table 2 defaults, complete all of these steps:

  1. add an accepted UMSH_TARGET_PLATFORM value and truthful native detector;
  2. define the binding requirements and policy plan without fallback;
  3. register only implementations that exist on that backend;
  4. add compile-time positive and negative tests;
  5. add runtime binding, policy-resolution, and persistent-kernel evidence;
  6. register the new campaign platform with unsupported cells still visible.

Source names, capability bits, or a passing kernel-completion copy are not enough to register a canonical implementation. Retain final device ISA or driver evidence for mechanisms whose semantics depend on compilation or PTE configuration.

Archive and report

Campaign output defaults to build/experiment-results/cross-platform/<run-id>/. Preserve the complete directory, including:

campaign.json
cases.csv
manifest.snapshot.json
SHA256SUMS
<platform>/inventory/
<platform>/cases/<ordinal>-<case>/

RECORDED means that a command exited successfully and its artifact contract passed. It does not mean that an absolute paper number was reproduced. Compare absolute performance only when hardware, workload, binding, implementation, timing boundary, and power/clock state match. Otherwise report the new device as a platform characterization and compare qualitative behavior.

Run the local orchestration checks after changing a checked-in manifest or artifact contract:

python3 -m py_compile tools/cross_platform/orchestrate.py
python3 -m unittest discover -s tools/cross_platform/tests -v
python3 tools/cross_platform/orchestrate.py validate
python3 tools/cross_platform/orchestrate.py run --profile all
python3 tools/cross_platform/orchestrate.py inventory

The final two commands are dry runs unless --execute is present.