Skip to content
Merged
Show file tree
Hide file tree
Changes from 8 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -209,8 +209,10 @@ transport maps to the contract 503.
| Sampler | top-p 0.9, repeat penalty 1.1, lookback window 8192 (generated tokens only) | [desktop/main/backend/inference/penalties.ts](desktop/main/backend/inference/penalties.ts) (`PENALTY_FULL_CONTEXT_TOKENS`) |
| History | at most 12 turns carried into the prompt | `MAX_HISTORY_TURNS` |

Vulkan stays off (reserved). The memory governor can force the runtime
profile to `fast` (below); the wizard's RAM gate uses the same estimate
The compute backend is decided by an out-of-process probe (issue #155): a working
Vulkan device is used when one is found and CPU inference is the fallback;
`inference.vulkan` can pin either choice. The memory governor can force the
runtime profile to `fast` (below); the wizard's RAM gate uses the same estimate
family (file size + 1 GiB KV-cache + 1 GiB overhead,
[desktop/main/first-run/ram-gate.ts](desktop/main/first-run/ram-gate.ts)).

Expand Down
30 changes: 29 additions & 1 deletion CHANGELOG.md

Large diffs are not rendered by default.

6 changes: 4 additions & 2 deletions INSTALL.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,8 @@ Complete installation guide for the Document Q&A Assistant, including standard P
- **Python**: 3.11 or higher

### Optional Components
- **NVIDIA GPU**: Not required for GGUF backend (CPU-only inference)
- **NVIDIA GPU**: not required. The GGUF backend uses a working Vulkan device when one is
present and falls back to CPU inference otherwise

## Standard Installation

Expand Down Expand Up @@ -262,7 +263,8 @@ Not required. Application runs as a standard executable.

### System Requirements for Power Users

The application runs CPU-only using GGUF models. No GPU or NPU acceleration is required.
The application runs GGUF models on a GPU when a probe finds a working Vulkan device, and on
the CPU otherwise. GPU acceleration is optional; no particular GPU or NPU is required.

## Post-Installation

Expand Down
15 changes: 9 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,15 +86,16 @@ The desktop app runs GGUF models via node-llama-cpp (Node main-process backend,
- **Quality profile (default)**: Gemma 4 E2B-it (Q4_K_M GGUF per [ADR-0002](docs/adr/0002-llm-profiles.md); ~2.9 GB nominal per PACKAGING.md; 2,620,370,976 bytes measured per bench/RESULTS.md) — bundled
- **Fast profile**: lfm2.5-vl-450m (Q4_K_M GGUF per ADR-0002) — bundled
- Profile selection: automatic free-RAM gate or in-app choice; `TRAININGAPP_DESKTOP_INFERENCE_PROFILE` (`quality` / `fast` / `auto`) is the desktop env override. (`RAG_GGUF_PATH` / `--gguf-path` select a custom GGUF on the legacy Python harness only.)
- No GPU required
- No particular GPU required: a working Vulkan device is used when one is found, and
CPU inference is the fallback
- No network access required (unless you turn on the external model or update checks)
- Measured decode throughput and first-token latency per model/profile: see [bench/RESULTS.md](bench/RESULTS.md) (issue #52 benchmark harness)

### Hardware Requirements
#### Minimum (Intel 11th Gen i5, 16GB RAM)
- Windows 11 (64-bit)
- Intel Core i5 11th generation or newer (or equivalent AMD Ryzen 5000+)
- Intel integrated graphics (present on all 11th gen+ Intel CPUs) — no discrete GPU required
- Intel integrated graphics (present on all 11th gen+ Intel CPUs); a discrete card is not needed
- 16GB RAM
- ~6.4 GB free storage for models + app (measured installed footprint 6,378,451,601 bytes; staged model resources 4,111,872,009 bytes; see bench/RESULTS.md)
- **Performance**: measured numbers per model and quantization are recorded in [bench/RESULTS.md](bench/RESULTS.md)
Expand All @@ -109,7 +110,7 @@ The desktop app runs GGUF models via node-llama-cpp (Node main-process backend,
#### High-Performance (Intel 13th Gen i9, 64GB RAM)
- High-end CPU (Intel Core i9 or AMD Ryzen 9)
- 64GB RAM
- **Performance**: measured CPU-only GGUF numbers are recorded in [bench/RESULTS.md](bench/RESULTS.md)
- **Performance**: measured CPU and GPU numbers, and which machine each was measured on, are recorded in [bench/RESULTS.md](bench/RESULTS.md)

> **Pending**: the offline/low-RAM reference-laptop validation matrix (issue #86) —
> reference-i5 rows are not yet measured; no reference-hardware numbers are claimed here.
Expand Down Expand Up @@ -298,7 +299,9 @@ The desktop app runs GGUF models via node-llama-cpp (Node main-process backend,
verification of the bundled tree → knowledge-pack activation → license notices →
complete. Every gate names its failure reason; setup can be re-run from Settings.

No Python, no GPU, and no network access are required (unless you turn on the external model or update checks).
No Python and no network access are required (unless you turn on the external model or update
checks). GPU acceleration is used when the machine has one that works; otherwise inference
runs on the CPU.

### Building the desktop app from source

Expand Down Expand Up @@ -679,7 +682,7 @@ Mirrors [ARCHITECTURE.md](ARCHITECTURE.md) (the authoritative map):
- Cross-encoder rerank (ettin-reranker-32m-v1) with a calibrated relevance floor

**LLM Interface**
- GGUF via node-llama-cpp (desktop, CPU-only, fully offline)
- GGUF via node-llama-cpp (desktop; GPU when a probe finds one working, CPU otherwise, fully offline)
- GGUF via wllama WASM (browser, CPU/SIMD, fully offline)

**RAG Engine**
Expand Down Expand Up @@ -1152,4 +1155,4 @@ Legacy Python harness only (CI conformance surface):
---
**Version**: 2.3.0
**Last Updated**: 2026-09-29 (v3 documentation refresh, issue #89)
**Hardware**: CPU-only optimized for Intel 11th gen i5 and above (16GB RAM minimum)
**Hardware**: optimized for Intel 11th gen i5 and above (16GB RAM minimum); GPU acceleration is opportunistic and CPU inference is the floor
18 changes: 18 additions & 0 deletions api_server.py
Original file line number Diff line number Diff line change
Expand Up @@ -771,6 +771,24 @@ async def get_status_models(auth: dict = Security(require_auth())):
)


@app.post("/settings/inference/gpu-test")
async def test_gpu_inference(auth: dict = Security(require_auth())):
"""
GPU capability re-probe (issue #155).

The probe runs as a child process against node-llama-cpp's Vulkan backend
on the Electron/Node desktop surface only
(desktop/main/backend/inference/gpu-probe.ts); this Python host has no
llama.cpp GPU backend and never wires a probe, so the route exists to keep
the shared contract (contracts/api.openapi.yaml) consistent across backends
and always answers the documented unwired 503 - the path is known, never
404-absent.
"""
raise HTTPException(
status_code=503, detail="GPU probing is not wired on this host"
)


# --- C7 (issue #74): knowledge pack lifecycle -------------------------------
# Same wire shapes as the Node backend (desktop/main/backend/server.ts).
# C8 (issue #75): zip extraction is delegated to the shared pack_extract
Expand Down
33 changes: 33 additions & 0 deletions bench/RESULTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -152,6 +152,39 @@ see the provenance rule 4 above) on the real staged tree, machine-tagged:
| devstation | startup integrity gate latency (streaming sha256 of the full staged tree) | 2,377 ms | packaged-mode pass, 0 failures, 2026-09-23 |
| reference-i5 | all E1 size/latency rows | PENDING | operator runs the same commands (E3/#86 owns the matrix) |

### GPU device matrix (issue #155)

Which compute backend the desktop backend actually selects on each machine. Issue #155 turned
GPU acceleration from a hard-disabled constant into an out-of-process probe plus a CPU
fallback, so a row here is a statement about what the probe concludes there.

Recording discipline for this table (the same one the `reference-i5` E1 rows above already
follow): every row is either a figure measured on the machine it names, or an explicit `PENDING`
naming what is unmeasured. No tok/s number appears here unless it was measured on the machine
named. The issue #155 comments carry prefill/decode tok/s figures from an external benchmark
harness; this table does not restate them as its own evidence, because this revision did not
re-measure them.

| machine | device | backend selected | GPU offload | notes |
|---|---|---|---|---|
| devstation | Intel Arc Pro B50 (discrete), driver 32.0.101.8805 | vulkan | supportsGpuOffloading=true | measured: node-llama-cpp 3.20.0, same process reported gpu=false for the shipped CPU-only option and gpu=vulkan for `{gpu:{type:auto,exclude:[cuda]}}` |
| devstation | Intel Arc Pro B50 (discrete), driver 32.0.101.8805 | cpu | gpu=false, supportsGpuOffloading=false | measured: the same host and the same node-llama-cpp 3.20.0 process under the shipped `{gpu:false}` option - the comparison the GPU row above is measured against |
| reference-i5 | 12th-gen Core i5 mobile, Intel Iris Xe integrated | PENDING | PENDING | no integrated-GPU host was available; prefill and decode tok/s and the probe's own verdict are unmeasured, and this is the floor-spec machine the speed bar for the quality tier depends on |
| reference-amd-igpu | AMD integrated, RDNA | PENDING | PENDING | no AMD host was available; backend selection, probe verdict and throughput are unmeasured |
| reference-amd-dgpu | AMD discrete, RDNA | PENDING | PENDING | no AMD host was available; backend selection, probe verdict and throughput are unmeasured |
| reference-nvidia | NVIDIA discrete | PENDING | PENDING | no NVIDIA host was available; note the CUDA backend is deliberately not shipped, so this host is expected to resolve to Vulkan-or-CPU rather than CUDA |

One upstream llama.cpp defect has no runtime mitigation in the pinned library and is covered
only by the probe plus its CPU fallback, not fixed here: **#27638** (device loss at
`ubatch >= 2048`; `ubatch` does not appear anywhere in the installed node-llama-cpp 3.20.0
`dist/` except one code comment). Verified by searching the installed package, not assumed.

**#29054** (deterministic hang on a q8_0 KV cache) is a different case and is **not** claimed as
unmitigated: node-llama-cpp 3.20.0 does expose `experimentalKvCacheKeyType` /
`experimentalKvCacheValueType` on `LlamaContextOptions`, and **both already default to F16**, so
the f16 mitigation the upstream report asks for is in force by default. The application does not
surface that experimental override to users, deliberately.

**Known limit (recorded, never hidden): the single-file NSIS target cannot
embed the real-weights payload.** `makensis.exe` aborts with `File: failed
creating mmap of …-x64.nsis.7z` because the app archive is 4,177,827,908
Expand Down
72 changes: 72 additions & 0 deletions contracts/api.openapi.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -434,6 +434,32 @@ paths:
"503":
description: Model status is not wired on this host

/settings/inference/gpu-test:
post:
tags: [settings]
summary: Re-run the GPU capability probe (issue #155)
description: >-
Runs the out-of-process GPU probe again and adopts the result, so a
machine whose driver or hardware changed can recover from a stale
negative verdict. Takes no request body and persists nothing itself; the
host adopts and persists the returned verdict.

"No usable GPU" is an ANSWER, not an error, so a failed probe is a 200
with `ok: false` and a populated `reason` - the same shape
`POST /settings/external/test` uses. A 503 means the host has no probe
wired at all (a known route that cannot be served, never a 404).
Transport-token-guarded like every route.
operationId: testGpu
responses:
"200":
description: The probe ran; the verdict says whether a GPU is usable.
content:
application/json:
schema:
$ref: "#/components/schemas/GpuProbeResponse"
"503":
description: GPU probing is not wired on this host

/packs:
get:
tags: [packs]
Expand Down Expand Up @@ -572,6 +598,24 @@ components:
in: header
name: X-API-Key
schemas:
GpuProbeResponse:
type: object
description: >-
The result of a GPU capability probe run (issue #155). Same shape as the
`gpu` member of StatusModelsResponse.
required: [backend, ok, reason]
properties:
backend:
type: string
enum: [vulkan, cpu]
ok:
type: boolean
reason:
type: string
device:
type: string
nullable: true

StatusModelsResponse:
type: object
required: [engine, profile, models]
Expand All @@ -585,6 +629,34 @@ components:
GGUFs never block chat.
profile:
type: string
gpu:
type: object
description: >-
The GPU decision (issue #155), as the engine actually made it.
Present for the local llama.cpp backend and absent when an external
endpoint generates, because then no local compute backend is
involved. `backend` is what runs; `ok` is whether a GPU was usable;
`reason` is ALWAYS non-empty so a CPU fallback can be explained
rather than guessed at.
required: [backend, ok, reason]
properties:
backend:
type: string
enum: [vulkan, cpu]
description: >-
The resolved compute backend. CPU means either the probe found no
usable device, the operator pinned CPU, or nothing has been
probed yet.
ok:
type: boolean
description: True only when a GPU backend loaded AND produced a sane generation.
reason:
type: string
description: Human-readable explanation; never empty.
device:
type: string
nullable: true
description: Adapter identity the backend reported, when it reported one.
resident:
type: object
description: >-
Expand Down
22 changes: 16 additions & 6 deletions desktop/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -126,8 +126,14 @@ profile auto-selection.
`lfm2.5-vl-450m/model.gguf` — both relative to the model dir (assumption A2
pending ADR-0002 #56). Embeddings remain the stub until B5 (#63).
- **Threads**: `min(cores, 8)` by default — explicitly NOT the browser WASM
4-cap (`web_ui/src/lib/llm/wllama-service.ts`). `inference.vulkan` is
reserved, default `false` (llama.cpp #17389).
4-cap (`web_ui/src/lib/llm/wllama-service.ts`).
- **Compute backend**: `inference.vulkan` takes `'auto'` (the default, follow
the probe), `true` (force the GPU) or `false` (force CPU). The probe runs in
a separate OS process, so a driver fault cannot take the app down; its
verdict and a human-readable reason persist in `gpu-probe.json` beside the
other profile sidecars and are reported by `GET /status/models`. A GPU load
that fails on the automatic path retries once on CPU and records why; an
explicitly forced GPU surfaces the error instead of degrading silently.
- **Model location**: `TRAININGAPP_INFERENCE_MODEL_DIR` env (or dev-server
`--model-dir`) -> `<userData>/models` (Electron injects the path) ->
`~/.trainingapp/models` headless fallback. A missing model makes `/ask` and
Expand All @@ -138,15 +144,19 @@ profile auto-selection.
in-flight generation ends); a client disconnect stops emission via a 20ms
cancellation poll + abort signal — well inside the 200ms budget.
- **Settings**: `inference.profile` | `inference.profileThresholdGb` |
`inference.threads` (1..64) | `inference.vulkan` via PUT /settings; the
`inference.threads` (1..64) | `inference.vulkan` ('auto'|true|false) via
PUT /settings; the
rag_* keys still round-trip (owned by the composed stub until B5/B6).
Headless env equivalents: `TRAININGAPP_DESKTOP_INFERENCE_PROFILE`
(quality|fast|auto; invalid values fall back to auto) and
`TRAININGAPP_DESKTOP_INFERENCE_THREADS` (same 1..64 integer gate as the
settings key; invalid values fall back to the min(cores, 8) default).
- **Packaged installs**: the installer does NOT yet unpack node-llama-cpp's
native addon from the asar archive — packaged inference lands with #84
(E1). Dev runs and the headless dev-server are unaffected.
- **Packaged installs**: `electron-builder.yml` pins
`@node-llama-cpp/win-x64-vulkan` and `@node-llama-cpp/win-x64` in
`asarUnpack`, so the compute backends are always on disk rather than
depending on electron-builder's implicit native-module heuristic. The CUDA
backend packages are excluded from the installer and no code path selects
them. Dev runs and the headless dev-server are unaffected.
- **History**: contract-supplied history is capped to the last 12 turns
before seeding the model, so an oversized array cannot overflow the 8192
context (the browser client already caps at 6).
Expand Down
17 changes: 17 additions & 0 deletions desktop/electron-builder.yml
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,14 @@ directories:
files:
- dist/**/*
- package.json
# issue #155: the CUDA backend packages are excluded from the installer.
# Nothing can select them: the GPU decision resolves to 'vulkan' or 'cpu' only,
# and the auto probe passes exclude: ['cuda'] because these packages are not
# shipped. They measured ~510 MB unpacked, so carrying them is pure weight.
# The exclusion is declared here rather than in npm config so `npm ci` still
# installs everything the dev tree and the vulkan path need.
- '!**/node_modules/@node-llama-cpp/win-x64-cuda/**'
- '!**/node_modules/@node-llama-cpp/win-x64-cuda-ext/**'
# The desktop:build script copies web_ui/dist into renderer/ before packaging
# (stage-installer-resources.mjs excludes the weight files duplicated into
# installer-resources); the main process serves it from <resourcesPath>/web_ui
Expand Down Expand Up @@ -56,3 +64,12 @@ asarUnpack:
- '**/node_modules/pdfjs-dist/**'
- '**/node_modules/sqlite-vec*/**'
- '**/node_modules/better-sqlite3/**'
# issue #155: the llama.cpp compute backends are pinned explicitly for the
# same reason sqlite-vec is above. They used to reach disk only through
# electron-builder's implicit native-module heuristic; the Vulkan binary is
# now LOADED AT RUNTIME by the GPU probe, so relying on a heuristic that
# electron-builder may change would silently turn every packaged install
# back into CPU-only. The CPU backend is pinned alongside it because the
# probe's CPU fallback path must work even where no GPU binary exists.
- '**/node_modules/@node-llama-cpp/win-x64-vulkan/**'
- '**/node_modules/@node-llama-cpp/win-x64/**'
Loading
Loading