▶ Watch the 53-second Turbofit overview
A local-model provider for Hermes Agent that fits itself around the way you use your computer.
Turbofit detects the machine's physical accelerator topology, selects a local main/auxiliary model ladder, downloads the required artifacts through Turbohaul Manager, and exposes stable OpenAI-compatible model IDs to Hermes Agent.
When another program needs VRAM, Turbofit contracts one step at a time: it can remove a dedicated auxiliary residency, route auxiliary work through the main model, reduce context, or move to a smaller local model. After memory remains available long enough, it heals back toward the selected ceiling.
Current scope is local-only. Turbofit does not select, configure, or fall back to API models. API orchestration is recorded under Later development, not presented as a current feature.
- Hermes provider: one local OpenAI-compatible endpoint with stable
auto,active:main, andactive:auxIDs. - Hardware-aware auto selection: canonical profiles from 8 GB through 300+ GB; no 48 GB runtime special case.
- Manual profile selection: pin the adaptive ceiling to a compatible hardware profile.
- Managed model acquisition: first activation pulls missing GGUF artifacts from pinned Hugging Face commits, verifies SHA-256, deduplicates shared blobs, installs Turbohaul manifests, and verifies the final tags before inference.
- Adaptive local scaling: contraction dwell, expansion dwell, hysteresis, cooldown, rollback, and flap quarantine.
- External workload priority: Turbofit never kills or signals games, editors, renderers, or other GPU consumers.
- Verified publication: a new route is published only after its local model rung loads and passes verification.
- Bounded auxiliary lifecycle:
active:auxforwards SSE frames immediately, propagates client disconnect cancellation through Turbohaul to llama.cpp, disables hidden-thinking by default, and caps each generation at 4,096 tokens unless explicitly configured otherwise. - Portable configuration: Turbofiles contain no credentials, machine-local paths, mutable process state, or embedded model binaries.
The graphic shows the full product ladder: hardware tiers, context targets, auxiliary choices, pressure response, and the model families moving through promotion. The table below distinguishes the local ladders available now from higher-context and broader-model combinations still awaiting evidence.
Turbofit ships local-only profiles for these physical classes:
| Class | Canonical topology | Current local ladder |
|---|---|---|
| 8 GB | 1x8 |
Bonsai 27B 1-bit, 64K shared main/aux floor |
| 16 GB | 1x16 |
Bonsai 27B 1-bit: 262K → 128K → 64K |
| 24 GB | 1x24 |
GRM 2.6 Plus 128K → Bonsai 262K → 128K → 64K |
| 48 GB | 2x24 |
GRM 2.6 Plus 262K shared main/aux → GRM 128K → Bonsai floors |
| 64 GB | 2x32 |
same verified dual-model ladder with additional headroom |
| 96 GB | 4x24 |
same verified dual-model ladder; unused cards remain available to other work |
| 200 GB | 2x100 |
same verified dual-model ladder while larger candidates are promoted |
| 300+ GB | 3x100 |
same verified dual-model ladder; unused cards remain available to other work |
The current dual-24 GB production ceiling routes both active:main and active:aux to the same GRM 2.6 Plus 262K runtime. Stage-v1 passed with 100% quality, 100% context retrieval, 3.823 effective end-to-end tok/s, and 12,369 / 15,437 MiB peak GPU use (sha256:2dd320671f7f891ede49a22820754fa657335581ed1272a3444fe45af9426223). A separate live 256-token probe reported 55.885 llama-server decode tok/s; effective throughput includes gateway and manager orchestration.
Profiles are selected from physical capacity, not transient free VRAM. Larger cards can satisfy a smaller per-card envelope when card count and topology shape remain compatible; 1x48 is still not treated as 2x24.
The Bonsai floor was measured on the current benchmark host. The 8/16/64/96/200/300 class mappings are portable recommendations, not claims of completed benchmarking on every accelerator family. Activation remains fail-closed if Turbohaul cannot load or verify a selected rung.
# Show local profiles, rungs, and compatibility with this machine
scripts/turbofit-runtime list
# Let Turbofit choose from physical hardware
scripts/turbofit-runtime set auto
# Or select a compatible local profile explicitly
scripts/turbofit-runtime set hardware-16gb
# Inspect the persisted selection
scripts/turbofit-runtime status
# Run one controller reconciliation
scripts/turbofit-controller --once
# Optional persistent user service
scripts/install-controller-service --startA new selection starts at its smallest local floor—not at an API fallback. Before that floor is published, the controller:
- Resolves every model tag required by the rung.
- Checks Turbohaul's installed tags and content digests.
- Pulls each missing Hugging Face artifact from an exact commit.
- Requires the downloaded SHA-256 to match the acquisition catalog.
- Reuses a verified blob when several model tags share it.
- Installs the context/runtime manifest for each tag.
- Loads and verifies the selected local rung.
- Atomically publishes the stable routes.
Acquisition recipes live in runtime-profiles/acquisitions.json. Model lifecycle authority remains in Turbohaul Manager; Turbofit does not create a second model store.
Run the Turbofit gateway at http://127.0.0.1:8091, then configure one provider:
custom_providers:
- name: turbofit
base_url: http://127.0.0.1:8091/v1
api_key: not-needed
api_mode: chat_completions
models:
auto: {}
active:main: {}
active:aux: {}
model:
provider: custom:turbofit
default: autoStable model IDs:
| ID | Meaning |
|---|---|
auto |
current selected main route |
active:main |
current main residency |
active:aux |
dedicated auxiliary when present, otherwise shared main |
The IDs stay constant while the controller changes the backing local model and context.
hermes plugins install SouthpawIN/turbofit --enableRestart Hermes, then open the setup screen:
hermes dashboardSelect Turbofit in the dashboard. The plugin can scan/select a compatible
hardware profile, register the custom:turbofit endpoint, set auto as the
primary model, add or remove Turbofit from the canonical
fallback_providers chain, and publish both provider and dashboard endpoints
privately with Tailscale Serve. Tailnet publishing defaults to separate HTTPS
ports and never exposes a public Funnel route. It also registers /turbofit,
turbofit_status, and turbofit_configure for CLI and gateway sessions.
Linux/WSL2 NVIDIA launches use CUDA. Apple Silicon is detected as one unified
Metal device; the portable 8/16/24 GB profiles compile Docker-only Bonsai
recipes to native llama-server processes with Metal enabled. Every generated
launch path includes --jinja for tool-call templates.
Development should happen from a Git checkout, not the installed plugin directory.
Current requirements:
- Python 3.11+
- Hermes Agent
- Turbohaul Manager v0.7
- PyYAML for YAML Turbofiles
- a supported local runtime/accelerator backend
- network access to the pinned Hugging Face artifacts on first acquisition
CUDA on Linux/WSL2 and Metal on Apple Silicon are implemented. NVIDIA remains the primary measured backend; Metal launch compilation is portable but still requires host-specific benchmark promotion before claiming equivalent performance. AMD/Intel backends remain discovery-only.
A representative ladder is:
local main + dedicated local auxiliary
→ local main shared with auxiliary work
→ smaller local model/context
→ minimum local floor
Contraction occurs only after a sustained deficit. Healing occurs one rung at a time only after the configured margin, dwell, hysteresis, cooldown, and flap controls pass.
A transition:
- Blocks new auxiliary admission when leaving dedicated mode.
- Drains active auxiliary streams.
- Requests clean unload through Turbohaul.
- Activates or acquires the target local rung.
- Verifies the target.
- Atomically publishes routes.
- Restores and verifies the previous state on failure.
At the minimum local floor, Turbofit does not route to an API model. If no lower local rung fits, it holds the floor and reports the capacity condition.
| Path | Purpose |
|---|---|
runtime-profiles/*gb.yaml |
production hardware profiles |
runtime-profiles/acquisitions.json |
pinned sources, hashes, and Turbohaul tag recipes |
runtime-profiles/runtime-resolutions.json |
rung-to-model-tag resolution |
runtime-profiles/rung-requirements.json |
per-card VRAM requirements |
references/model-catalog.json |
requested model variants and capabilities |
references/configuration-matrix.json |
generated main × auxiliary × context candidate space |
benchmarks/suite.yaml |
promotion gates |
references/results/ |
measured machine-readable evidence |
The matrix contains the requested 12 main variants × 4 auxiliary modes × 4 contexts: 192 research configurations. Every row now compiles to a concrete, --jinja-enabled launch recipe; FP16, quantized, vision-projector, and DSpark artifacts resolve independently. A compilable row is not automatically a production recommendation. Production promotion requires artifact, runtime, performance, quality, and pressure/self-heal evidence.
The hybrid system-RAM catalog separately defines executable 64K, 128K, 262K, and 1M candidates for Laguna, MiniMax M3, and GLM 5.2. Higher-context rows are explicitly configured-unmeasured, use CPU/system-RAM policies (including MoE expert offload where supported), and cannot be promoted until real evidence exists.
The current source list includes:
- GLM/GRM 5.2 2.788 bpw
- MiniMax M3 GGUF
- Laguna S 2.1
- Laguna S 2.1 GGUF
- GRM 2.6 Plus GGUF
- Carwin MoE Nano GGUF
- Ternary Bonsai 27B GGUF
- Bonsai 27B GGUF
Candidate ranking follows:
quality
→ reach 128K context
→ reach 20 tok/s
→ reach 262K context
→ reach 100 tok/s
→ reach 1M context
→ maximize speed
Measured claims remain attached to their exact artifact, runtime flags, context, host fingerprint, throughput, TTFT, and per-card VRAM evidence.
runtime-profiles/hybrid-models.json defines dual-24 GB GPU + system-RAM placements for Laguna S 2.1 Q4_K_M, MiniMax M3 MXFP4_MOE, and GLM 5.2 2.788 bpw. Every artifact is bound to an immutable Hugging Face revision, required SHA-256 identity, and exact file size. A benchmark pass validates only that exact artifact, runtime, flags, context, and host class; it does not automatically add the candidate to the production adaptation ladder.
| Candidate / exact placement | Stage-v1 quality | Context retrieval | Effective output | Peak GPU MiB | Peak server RSS | Evidence |
|---|---|---|---|---|---|---|
| Laguna S 2.1 Q4_K_M · Poolside Laguna runtime · 24 GPU layers · 64K | 100% | 100% | 2.667 tok/s | 19,597 / 18,270 | 38,403 MiB | sha256:a332e44a601c129b90262c877ad1f62e0d8fd54780dba5f566ee010d51225ec5 |
| MiniMax M3 MXFP4_MOE · PR 24523 runtime · 40 GPU layers + 56 CPU-MoE layers · 64K | 100% | 100% | 0.622 tok/s | 7,983 / 23,168 | 226,678 MiB | sha256:2edf1976c39a661de1b8aa49fcd726602d7763923ecb3760723f7990e33ecabc |
| GLM 5.2 2.788 bpw · ik_llama.cpp DSA + CPU-MoE · layer split · 64K | 100% | 100% | 0.973 tok/s | 11,734 / 11,468 | 234,926 MiB | sha256:02ab108a18e2a264c61472f7c98da0ee10bb70be4acf87436ec7066bed51c49e |
The first MiniMax launch proved that 99 GPU layers left no room for the 64K KV cache; the measured configuration uses 40. GLM uses layer split because ik_llama.cpp explicitly disables the DSA indexer under graph/tensor-parallel attention split.
The configuration checker reports both static hardware_fits and current launch_ready, so a machine is not called ready while another resident model still occupies required VRAM:
scripts/turbofit-hybrid-config list
scripts/turbofit-hybrid-config check glm-5-2-2-788bpw dual-24gb-64kThe evidence-first benchmark stage records raw responses, measured token usage, exact-answer quality checks, passkey context retrieval, effective end-to-end output throughput, host RAM, per-GPU VRAM, and an evidence SHA-256:
scripts/turbofit-benchmark-stage \
--candidate <model-id> \
--configuration dual-24gb-64k \
--base-url http://127.0.0.1:<port>/v1 \
--model <served-model> \
--output references/results/<run>.json \
--disable-thinking# Unit, integration, schema, profile, link, and simulated adaptation checks
scripts/release-check
# Adds live NVML, Turbohaul, stable-route, controlled-pressure,
# external-process survival, contraction, and healing checks
scripts/release-check --realA simulated pass is not represented as real pressure evidence. The latest machine-readable acceptance record is stored at references/results/adaptive-runtime-acceptance.json.
- External GPU processes are never terminated or signaled.
- Physical capacity selects the profile; transient availability selects the rung.
- All model lifecycle operations go through Turbohaul.
- Downloads use pinned revisions and required SHA-256 values.
- Missing or mismatched artifacts fail closed.
- New routes are not published before local verification.
- Credentials and machine-local paths never enter portable profiles.
- Research candidates never become production recommendations automatically.
The following items remain outside the verified release boundary:
- host-specific Metal performance evidence and AMD/Intel runtime backends
- promotion evidence for every one of the 192 compiled model configurations
- Mixture-of-Agents presets
- pricing-aware opt-in API routing
- richer daemon/service management outside systemd
Model-source intelligence is refreshed daily by .github/workflows/model-intel.yml; discovery never bypasses the benchmark promotion gates.
MIT License. Copyright (c) 2026 sovthpaw (SouthpawIN). See LICENSE.
src/turbofit_runtime/ schemas, selection, acquisition, policy, controller
runtime-profiles/ production profiles and runtime catalogs
benchmarks/ promotion suite
research/ candidate-only discovery
scripts/turbofit-runtime list/set/status selection CLI
scripts/turbofit-controller adaptive local controller
scripts/turbofit-gateway.py stable OpenAI-compatible gateway
references/results/ measured evidence
assets/turbofit-hero.png README image
tests/ unit and integration gates



