Skip to content

Latest commit

 

History

History
742 lines (590 loc) · 28.1 KB

File metadata and controls

742 lines (590 loc) · 28.1 KB

Intel artifact-evaluation guide

This guide covers native deployment, correctness validation, and performance experiments for the Intel backend. The shortest useful evaluation path is:

  1. build the userspace targets with kernel modules disabled;
  2. run the build checks and the stock-runtime binding probe;
  3. run the canonical persistent-kernel correctness smoke matrix; and
  4. use the campaign orchestrator for any result that will be retained as AE evidence.

Privileged kernel modules and page-system changes are not part of the default build. They are needed only for the corresponding optional performance rows or for a complete preconfigured campaign.

Supported Intel platforms

The native detector currently recognizes these Intel platform families:

Platform family Build identity Checked-in campaign ID Driver in the checked-in deployment Status
Lunar Lake LUNAR / LNL ultra258v xe Paper platform class
Arrow Lake ARROW / ARL u285h i915 Paper platform class
Raptor Lake class RAPTOR / RPL tianx i915 Control platform with restricted GPU policy support
Meteor Lake METEOR / MTL none not fixed Source-supported, but not covered by the checked-in campaign

Configuration is native-only. Do not copy a build directory between machines: the selected platform macros come from the build host, and the correctness binary also checks the live Intel GPU identity before using an architecture-specific implementation.

Install userspace prerequisites

The Intel backend requires CMake 3.24 or later, a C++23 compiler, Python 3, Ninja, Intel OpenCL development headers, and an Intel OpenCL runtime exposing cl_intel_unified_shared_memory.

On Ubuntu or Debian:

sudo apt update
sudo apt install build-essential cmake ninja-build git python3 \
  ocl-icd-opencl-dev opencl-clhpp-headers intel-opencl-icd clinfo

Confirm that the intended integrated GPU is visible and that the required extension is reported:

cmake --version
c++ --version
clinfo -l
clinfo | grep -m1 cl_intel_unified_shared_memory

Verify that the reported CMake is at least 3.24 and that the compiler supports the required C++23 library features. Some distribution releases package an older CMake even though the command is available.

The correctness test rejects non-empty IGC_*, VISA_OPTIONS, NEO_OCL_*, and NEO_Inject* overrides because they can change final GPU code while leaving the requested policy label unchanged. NEOReadDebugKeys=0 is accepted, but an unset environment is clearer. Inspect the environment before collecting evidence:

env | grep -E '^(IGC_|VISA_OPTIONS=|NEOReadDebugKeys=|NEO_OCL_|NEO_Inject)' || true

Build the userspace targets

Kernel modules are deliberately disabled in the default AE build. The empty GITHUB_MIRROR value selects the original Abseil download URL rather than the mirror configured by the repository preset.

cmake -S . -B build-release -G Ninja \
  -DCMAKE_BUILD_TYPE=Release \
  -DUMSH_TARGET_PLATFORM=AUTO \
  -DUMSH_BUILD_EXAMPLES=ON \
  -DUMSH_BUILD_KERNEL_MODULE=OFF \
  -DGITHUB_MIRROR=
cmake --build build-release --parallel

If network access is unavailable and a verified Abseil source tree is already present, configure with -DFETCHCONTENT_SOURCE_DIR_ABSL=/absolute/path/to/abseil instead. Do not point this option at an unverified or version-incompatible source tree.

Check the detected backend and target in the configure output or cache:

grep -E '^UMSH_(BACKEND|RESOLVED_TARGET_PLATFORM):' \
  build-release/CMakeCache.txt

The expected backend is Intel; the resolved platform must match the host. An unknown Intel CPU does not silently select another platform's policy.

Installation is not required for the AE workflow. A staging install is useful only when validating the install layout; see the Intel policy-transfer reference for that procedure.

Sanity checks

Run the host-independent library and artifact-tool tests first:

ctest --test-dir build-release --output-on-failure
python3 -m unittest discover -s tools/cross_platform/tests -v
python3 tools/cross_platform/orchestrate.py validate

These commands check compile-time defaults, storage-report logic, and the campaign implementation. They are not CPU/GPU data-transfer correctness tests.

Print the implementation manifest compiled into the Intel correctness binary without opening an OpenCL device:

build-release/examples/correctness_tests/intel/policy_transfer/correctness_intel_policy_transfer \
  --list_policies

Then run one real stock-runtime allocation and binding probe:

build-release/examples/custom/policy_probe/umsh_policy_probe

Require BINDING status=PASS, then inspect the fields on that line separately. For the stock Intel path, same_backing=1, the expected Intel userptr backing, and the expected verification method must all be present. The probe's PASS decision checks bind and runtime-coherency resolution; it does not itself gate same_backing or the verification field. The probe validates deployment and capability resolution, not persistent-kernel data visibility.

Correctness tests

For the public control API on Arrow/Lunar, build and run:

cmake --build build-release --target correctness_intel_control --parallel
timeout 120s build-release/examples/correctness_tests/intel/control/\
correctness_intel_control 10000

This checks wait/set with the existing intra-kernel default read/write operations. Raptor is unsupported. See the control-plane guide for the experimental Intel protocol contract and expected output. Arrow's inter-kernel CPU -> GPU default is W_F^WB -> R_I^WB; the control test uses the separate intra-kernel profile.

For data-policy characterization, use the canonical persistent-kernel policy matrix. It observes the consumer while the GPU kernel is still active, before a completion boundary can implicitly publish or invalidate candidate data.

Standalone smoke matrix

This command matches the campaign's 4 KiB, one-seed, two-round smoke profile:

python3 examples/correctness_tests/run_policy_matrix.py \
  --backend intel \
  --matrix-kind canonical \
  --binary "$PWD/build-release/examples/correctness_tests/intel/policy_transfer/correctness_intel_policy_transfer" \
  --output "$PWD/build/ae-intel/canonical-smoke-4k.csv" \
  --bytes 4096 \
  --rounds 2 \
  --seeds 0x6a09e667f3bcc909 \
  --case-timeout-seconds 120 \
  --execute

The runner creates a sibling canonical-smoke-4k.logs/ directory containing the raw output for every executed cell. It refuses to replace an existing CSV unless --force is passed; prefer a new output name for a new observation.

The complete profile repeats the matrix at 4 KiB, 64 KiB, and 1 MiB with these three seeds and 20 rounds:

0x6a09e667f3bcc909
0xbb67ae8584caa73b
0x3c6ef372fe94f82b

Use the campaign command below for the complete profile. It creates separate, validated artifacts for all three sizes.

Interpret correctness outcomes

The matrix deliberately includes positive paths, stale-data paths, unsupported coordinates, and optional backings. Do not apply an ordinary "every row must pass" unit-test rule.

Outcome Meaning
PASS_IMMEDIATE The first adaptive attempt observed the exact payload without a post-issue retry
PASS_CONVERGED A later attempt or post-issue retry observed the exact payload
FAIL_STABLE All configured budgets completed with a stable stale, torn, or corrupt signature
INCONCLUSIVE_* The bounded protocol did not establish either an exact observation or a stable mismatch
SKIP_BACKING The requested allocation or GPU import was unavailable
SKIP_UNSUPPORTED The implementation or its runtime prerequisite was unavailable
SKIP_CONTROL_UNSUPPORTED The independent persistent control protocol did not pass its live gate
ERROR, ERROR_HUNG, or runner contract error Infrastructure failure; not a policy result

The matrix runner treats a stable mismatch and a declared skip as observations, continues to later cells, and preserves their raw logs. Consult the matrix contract and the Intel protocol details before interpreting an individual cell.

Checked-in campaign manifest

The campaign orchestrator inventories source and machine state, executes only manifest cases marked ready, validates their declared artifacts, and writes a hash-covered local archive. It never builds the repository, invokes sudo, loads a module, or changes a kernel parameter.

The checked-in platform entries are deployment records with literal internal paths and SSH aliases:

ID Transport Repository path expected by the manifest
ultra258v local /home/yjr/xpu_sync/umsh
u285h ssh u285h /home/yjr/work/xpu_sync/umsh-campaign
tianx ssh tianx /home/yjr/work/xpu_sync/umsh-campaign

These paths are not a fallback search list. They are directly usable only on the preconfigured artifact hosts. An evaluator deploying to a different host, checkout path, or SSH alias must follow the platform-extension guide and validate a separate manifest entry before execution. Do not relabel a new machine as one of the three checked-in systems merely to reuse its commands.

Inspect a campaign without connecting to a remote host or creating files:

python3 tools/cross_platform/orchestrate.py validate
python3 tools/cross_platform/orchestrate.py plan \
  --platform u285h --suite correctness --profile smoke
python3 tools/cross_platform/orchestrate.py plan \
  --platform u285h --suite performance --profile full

Run correctness through the campaign

Run from the controller checkout described by the selected manifest entry. A clean source tree is recommended because it makes the recorded source identity easy to audit:

git status --short
ae_commit=$(git rev-parse HEAD)

python3 tools/cross_platform/orchestrate.py run \
  --platform u285h \
  --suite correctness \
  --profile smoke \
  --expected-commit "$ae_commit" \
  --require-clean \
  --execute

python3 tools/cross_platform/orchestrate.py run \
  --platform u285h \
  --suite correctness \
  --profile full \
  --expected-commit "$ae_commit" \
  --require-clean \
  --execute

git status --short must be empty when --require-clean is used. Without that flag, the orchestrator still records the dirty status and a content hash; it does not make a dirty run anonymous.

Replace u285h with ultra258v or tianx only when the corresponding manifest deployment is actually in use. Repeated --platform options select multiple hosts.

Performance prerequisites

Most Intel performance binaries configure either Transparent Huge Pages, explicit HugeTLB pages, or both. They run a page-system preflight before opening their result CSV, so a failed preflight cannot be mistaken for a partial benchmark.

Also record the selected CPU governor, available clock controls, power mode, thermal state, and any platform-specific firmware setting used for the run. The generic orchestrator inventory does not capture every Intel frequency or power control. Keep that operator record, with its own hashes, beside the immutable campaign archive rather than editing files inside the archive after SHA256SUMS has been generated.

Benchmark THP madvise required 2 MiB HugeTLB pool Optional module requirement in the full manifest
bind yes 2048 logical peak; 3072 recommended operational pool none
CPU read yes 512 uc_mem for UC rows; wbinvd_mod for coarse rows
CPU write yes 512 uc_mem for UC rows; wbinvd_mod for coarse rows
GPU read/write yes none none
GPU contention read no none none
SSD load yes none none

The 3072-page bind recommendation is measured release headroom for the current i915 matrix, not a portable promise. The logical peak is 2048 pages, but that exact pool passed preflight and later encountered a transient shortage on the validated Raptor host. Verify the actual free pool after every reservation.

Configure THP for the experiment window

Inspect the active global and 2 MiB policies:

cat /sys/kernel/mm/transparent_hugepage/enabled
cat /sys/kernel/mm/transparent_hugepage/defrag
cat /sys/kernel/mm/transparent_hugepage/hpage_pmd_size
test ! -e /sys/kernel/mm/transparent_hugepage/hugepages-2048kB/enabled || \
  cat /sys/kernel/mm/transparent_hugepage/hugepages-2048kB/enabled

The required global enabled and defrag modes are madvise, and hpage_pmd_size must be 2097152. Save the current selections before changing them:

ae_old_thp_enabled=$(sed -n 's/.*\[\([^]]*\)\].*/\1/p' \
  /sys/kernel/mm/transparent_hugepage/enabled)
ae_old_thp_defrag=$(sed -n 's/.*\[\([^]]*\)\].*/\1/p' \
  /sys/kernel/mm/transparent_hugepage/defrag)
test -n "$ae_old_thp_enabled"
test -n "$ae_old_thp_defrag"

printf 'saved THP enabled=%s defrag=%s\n' \
  "$ae_old_thp_enabled" "$ae_old_thp_defrag"

echo madvise | sudo tee /sys/kernel/mm/transparent_hugepage/enabled
echo madvise | sudo tee /sys/kernel/mm/transparent_hugepage/defrag

The preflight rejects always because it can promote Default mappings and invalidate the Default/HugePage/THP comparison. A successful policy preflight does not prove that every individual THP VMA was promoted; strict page-backing claims require separate smaps or equivalent evidence.

Reserve HugeTLB pages

Confirm that the default HugeTLB size is 2 MiB and save the existing pool:

grep -E 'HugePages_(Total|Free|Rsvd)|Hugepagesize|Hugetlb' /proc/meminfo
ae_old_nr_hugepages=$(sysctl -n vm.nr_hugepages)
printf 'saved vm.nr_hugepages=%s\n' "$ae_old_nr_hugepages"

Keep the three saved values in the same operator shell through restoration. If the campaign may outlive that shell, copy the printed values into the campaign operator record and verify them before reassigning the variables. Never guess the previous settings during restoration.

For CPU read or CPU write alone, reserve at least 512 pages. For the complete performance profile, including bind, use the observed 3072-page pool if the host has sufficient memory:

sudo sysctl -w vm.nr_hugepages=3072
grep -E 'HugePages_(Total|Free|Rsvd)|Hugepagesize|Hugetlb' /proc/meminfo

HugePages_Free - HugePages_Rsvd must meet the selected benchmark's requirement. Fragmentation may prevent a late reservation even when the host has enough total RAM; do not proceed based only on the requested sysctl value.

Run the performance campaign

Prepare THP, HugeTLB, and any selected optional modules before starting the non-privileged orchestrator. Then run:

ae_commit=$(git rev-parse HEAD)

python3 tools/cross_platform/orchestrate.py run \
  --platform u285h \
  --suite performance \
  --profile full \
  --expected-commit "$ae_commit" \
  --require-clean \
  --execute

The manifest runs these Intel cases:

Case Complete artifact contract
bind exactly 93 rows
CPU read exactly 51 rows with UC and coarse options enabled
CPU write exactly 51 rows with UC and coarse options enabled
GPU read 26 to 52 rows, depending on registered platform operations
GPU write 26 to 52 rows, depending on registered platform operations
GPU contention read exactly 8 rows; unsupported on Raptor
SSD load exactly 48 rows when enabled; external setup otherwise

A missing required helper module makes the CPU performance case SKIP_PRECONDITION; the orchestrator does not silently remove the optional rows and accept a smaller CSV as the full case.

Run individual benchmarks

Direct binary execution is useful for bring-up and diagnosis. It does not create the campaign inventory, source snapshot, performance metadata, artifact validation, or top-level checksums. Use the orchestrator for retained AE evidence.

Choose a new output directory for each observation:

ae_out="$PWD/build/ae-intel-manual"
mkdir -p "$ae_out"

Bind

build-release/examples/benchmarks/intel/bind/benchmark_intel_bind \
  --output="$ae_out/bind.csv" \
  --log_dir="$ae_out/bind-logs" \
  --alsologtostderr=true

CPU read and write

The primary module-free CPU WB/Fine matrices are:

build-release/examples/benchmarks/intel/cpu_read/benchmark_intel_cpu_read \
  --output="$ae_out/cpu-read-primary.csv" \
  --log_dir="$ae_out/cpu-read-primary-logs" \
  --alsologtostderr=true

build-release/examples/benchmarks/intel/cpu_write/benchmark_intel_cpu_write \
  --output="$ae_out/cpu-write-primary.csv" \
  --log_dir="$ae_out/cpu-write-primary-logs" \
  --alsologtostderr=true

After the optional modules have been independently verified and loaded, append --include_uncached=true --include_coarse=true to obtain the 51-row CPU matrices used by the full manifest.

GPU read and write

build-release/examples/benchmarks/intel/gpu_read/benchmark_intel_gpu_read \
  --maintenance_scope=block \
  --output="$ae_out/gpu-read.csv" \
  --log_dir="$ae_out/gpu-read-logs" \
  --alsologtostderr=true

build-release/examples/benchmarks/intel/gpu_write/benchmark_intel_gpu_write \
  --maintenance_scope=block \
  --output="$ae_out/gpu-write.csv" \
  --log_dir="$ae_out/gpu-write-logs" \
  --alsologtostderr=true

GPU contention

build-release/examples/benchmarks/intel/gpu_contention_read/benchmark_intel_gpu_contention_read \
  --output="$ae_out/gpu-contention-read.csv" \
  --log_dir="$ae_out/gpu-contention-read-logs" \
  --alsologtostderr=true

The bind, CPU read/write, and GPU read/write binaries accept --core_id=N. Rows fail when the requested affinity cannot be established. The GPU read and write binaries also accept --max_buffer_bytes=N for a bounded diagnostic run after a driver hang. A bounded run is incomplete and must not be used to synthesize omitted coordinates.

SSD load

The checked-in manifest enables the Intel SSD case as ready on u285h. Its reference path was observed on a Samsung PCIe 5.0 NVMe, but the per-run storage sidecar is authoritative. Lunar and Raptor retain explicit external_setup rows; enabling the benchmark there produces new-platform characterization rather than a Figure 7 reproduction.

The benchmark performs an untimed storage preflight and writes a JSON sidecar containing the resolved file, mount, filesystem, block stack, physical NVMe identity, firmware, PCI BDF and negotiated link, capacity, aligned O_DIRECT write/read/compare probe, and final allocation information. Its in-process eligibility gate rejects an ineligible device, mount, virtual block layer, or direct-I/O path. The stricter test_file.fully_allocated=true requirement is enforced by the campaign artifact validator, not by the benchmark's exit code alone.

The checked-in campaign command can be inspected with:

python3 tools/cross_platform/orchestrate.py plan \
  --platform u285h --suite performance --profile full

For a standalone full-shape run on the verified SSD filesystem:

ae_out="$PWD/build/ae-intel-manual"
mkdir -p "$ae_out/ssd-load"

build-release/examples/benchmarks/intel/ssd_load/benchmark_intel_ssd_load \
  --temp_file="$PWD/build-release/ssd-load-data/umsh_ssd_test.dat" \
  --storage_metadata="$ae_out/ssd-load/storage-metadata.json" \
  --require_ssd=true \
  --reuse_existing=true \
  --cache_flush_method=fadvise \
  --output="$ae_out/ssd-load/results.csv" \
  --log_dir="$ae_out/ssd-load/logs" \
  --alsologtostderr=true

The full experiment creates a fully allocated, benchmark-owned 1 GiB file when the path is absent, or reuses an exact-size existing file, and emits 16 sizes times 3 methods for exactly 48 rows. An existing file is never silently truncated or extended and qualifies as evidence only when the final sidecar also proves it is fully allocated.

A standalone exit of zero is therefore insufficient evidence when an existing file was reused: inspect the final sidecar and require it to report a regular, fully allocated 1 GiB file. Prefer the campaign path for retained evidence, because its JSON and CSV contracts reject a sparse file, a missing coordinate, or an unverified row as ERROR_ARTIFACT.

For a launch-only diagnostic, use a separate 512-byte file and append --max_size=512 --iterations_override=1. That diagnostic is not a Figure 7 artifact and must not share its file with the full run.

The method mapping is:

CSV type Interpretation
Direct Figure 7 Write-Bypass candidate only when O_DIRECT, no-fallback, same-backing, payload, and storage-sidecar gates all pass
Mmap Paper Mmap+Copy control
Copy Additional buffered-read plus OpenCL-copy diagnostic; not the paper control line

POSIX_FADV_DONTNEED is advisory. The metadata proves that fadvise was requested, not that every cache level was physically cold. drop_caches is a root-only, host-wide diagnostic mode and is deliberately excluded from the ready campaign. Never use --require_ssd=false for Figure 7 evidence.

Evidence layout and acceptance

The default local archive is:

build/experiment-results/cross-platform/<run-id>/
  campaign.json
  cases.csv
  manifest.snapshot.json
  SHA256SUMS
  <platform>/
    inventory/
    cases/<ordinal>-<case>/
      case.json
      performance-metadata.json
      stdout.txt
      stderr.txt
      artifacts/

Remote programs write to a unique directory below /tmp/umsh-cross-platform/<run-id>/; the orchestrator fetches declared artifacts but does not delete the remote directory. An existing local run ID is an error and is never overwritten.

RECORDED means that the command exited successfully and all declared artifact contracts passed. It does not by itself mean that a paper claim was reproduced. Performance interpretation must use the CSV scope and implementation columns together with performance-metadata.json and inventory. Intel GPU paper-policy comparisons must select both policy_scope=CANONICAL and result_scope=SHARED_BINDING. The executable archive and validator contract is defined by platforms.json and orchestrate.py.

Optional Intel helper modules

The userspace build, stock binding probe, WB correctness cells, GPU benchmarks, contention benchmark, and SSD benchmark do not require an out-of-tree kernel module.

The optional helpers enable these additional rows:

Module Additional coverage Device
uc_mem CPU Uncached binding rows /dev/uc_mem
wbinvd_mod coarse CPU flush/invalidate rows /dev/wbinvd_dev

Both module implementations are included under external/. Configure a separate module build against the exact running-kernel tree, build only the two module targets, and verify their identities before loading:

Install a compiler, GNU Make, kmod, psmisc, and the build tree for the exact running kernel. For a Debian/Ubuntu distribution kernel, the headers can usually be installed with sudo apt install linux-headers-"$(uname -r)" kmod psmisc. For the custom Arrow kernel, retain its configured build tree and pass it through KDIR; do not expect that distribution package name to exist.

ae_kernel_build=${KDIR:-/lib/modules/$(uname -r)/build}
test -f "$ae_kernel_build/Makefile"
test "$(make -s -C "$ae_kernel_build" kernelrelease)" = "$(uname -r)"

cmake -S . -B build-policy-modules -G Ninja \
  -DUMSH_BUILD_EXAMPLES=OFF \
  -DBUILD_TESTING=OFF \
  -DUMSH_BUILD_KERNEL_MODULE=ON \
  -DKDIR="$ae_kernel_build"
cmake --build build-policy-modules \
  --target uc_mem wbinvd_mod --parallel

ae_uc_ko=$PWD/build-policy-modules/external/module_uc_mem/uc_mem.ko
ae_wbinvd_ko=$PWD/build-policy-modules/external/module_wbinvd/wbinvd_mod.ko
modinfo "$ae_uc_ko"
modinfo "$ae_wbinvd_ko"
test "$(modinfo -F vermagic "$ae_uc_ko" | awk '{print $1}')" = \
  "$(uname -r)"
test "$(modinfo -F vermagic "$ae_wbinvd_ko" | awk '{print $1}')" = \
  "$(uname -r)"
sha256sum "$ae_uc_ko" "$ae_wbinvd_ko"

The convenience insertion targets now refuse to replace a loaded instance. For a controlled experiment, still use ordinary insertion after the identity checks above:

test ! -e /sys/module/uc_mem
test ! -e /sys/module/wbinvd_mod
sudo insmod "$ae_uc_ko"
sudo insmod "$ae_wbinvd_ko"
sudo udevadm settle

test -c /dev/uc_mem
test -c /dev/wbinvd_dev
lsmod | grep -E '^(uc_mem|wbinvd_mod)[[:space:]]'

Secure Boot may reject unsigned modules. Do not disable a machine's security policy merely to turn an optional row into a result.

wbinvd_mod exposes a world-writable device whose writes invalidate CPU caches on every core. Load it only on an isolated test system for the coarse-policy window. uc_mem can also remain referenced while a VMA or GPU/GUP pin is being released. Before unloading either module, wait for benchmark teardown, require zero module reference counts, and verify that no process holds either device:

for _ in {1..60}; do
  ae_wbinvd_refcount=$(cat /sys/module/wbinvd_mod/refcnt)
  ae_uc_refcount=$(cat /sys/module/uc_mem/refcnt)
  test "$ae_wbinvd_refcount" = 0 && test "$ae_uc_refcount" = 0 && break
  sleep 1
done
test "$ae_wbinvd_refcount" = 0
test "$ae_uc_refcount" = 0
if sudo fuser /dev/wbinvd_dev; then
  echo "/dev/wbinvd_dev is still open; refusing to unload" >&2
  exit 1
fi
if sudo fuser /dev/uc_mem; then
  echo "/dev/uc_mem is still open; refusing to unload" >&2
  exit 1
fi

sudo rmmod wbinvd_mod
sudo rmmod uc_mem

Never force-unload either module. Retain module hashes and the relevant kernel log window with any result that used them. The detailed implementations and additional safety checks are documented in the uc_mem reference and wbinvd_mod reference.

Restore page-system settings

After all benchmark processes have exited, wait for HugeTLB reservations to drain before shrinking the pool:

for _ in {1..60}; do
  ae_hugepages_rsvd=$(awk '/^HugePages_Rsvd:/ {print $2}' /proc/meminfo)
  test "$ae_hugepages_rsvd" = 0 && break
  sleep 1
done
test "$ae_hugepages_rsvd" = 0

sudo sysctl -w vm.nr_hugepages="$ae_old_nr_hugepages"
test "$(sysctl -n vm.nr_hugepages)" = "$ae_old_nr_hugepages"

echo "$ae_old_thp_enabled" | \
  sudo tee /sys/kernel/mm/transparent_hugepage/enabled
echo "$ae_old_thp_defrag" | \
  sudo tee /sys/kernel/mm/transparent_hugepage/defrag

grep -E 'HugePages_(Total|Free|Rsvd|Surp)' /proc/meminfo
cat /sys/kernel/mm/transparent_hugepage/enabled
cat /sys/kernel/mm/transparent_hugepage/defrag

These settings affect only the current boot unless the operator separately configured persistence. Start a new campaign after any module, THP, HugeTLB, driver, compiler, or source change; never edit an old archive to describe the new state.

Intel-specific limitations

  • CPU access on an Uncached mapping remains canonical Read-Alone or Write-Alone. It is not the missing CPU DMA Bypass implementation.
  • Intel GPU Base rows use device-local IntelDeviceUsm; CC and policy rows use imported userptr. A Base/CC ratio therefore changes backing and cannot reproduce a same-backing CC-overhead bar.
  • --maintenance_scope=block adds explicitly experimental GPU maintenance rows. The current single-work-item correctness protocol cannot establish Block participation semantics.
  • Intel contention rows use IntelDeviceUsm and have result_scope=DEVICE_LOCAL_REFERENCE; they characterize the contention mechanism, not end-to-end canonical shared-memory Read-Bypass.
  • Raptor exposes only ordinary GPU candidate access. Bypass, flush, invalidate, and contention paths remain unsupported, and a failed live control gate is not a candidate-data PASS or FAIL.
  • Meteor Lake source support is not evidence of a completed five-platform campaign. It must be onboarded and reported as a distinct platform.
  • A THP request plus a successful system preflight is not proof that every VMA was promoted to a 2 MiB mapping.
  • The archived Arrow GPU performance campaign has no complete 1 GiB matrix; a bounded 256 MiB diagnostic must not be used to fill the missing rows.
  • Source-level Intel policy names and successful OpenCL compilation are not automatic final-ISA attestation when the IGC version changes.