Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions .github/workflows/desktop-build.yml
Original file line number Diff line number Diff line change
Expand Up @@ -55,6 +55,14 @@ jobs:
npm --prefix packtool ci --no-audit --no-fund
npm --prefix packtool run build

# issue #155 review PRR-006: the GPU probe's child module is a COMPILED
# artifact (a spawned node cannot load the .ts source), so its contract
# tests need dist/. Without this the child had ZERO coverage in any CI job:
# `dist/` is git-excluded, so the probe-child tests skipped and every
# behaviour in gpu-probe-child.ts could be deleted with CI green.
- name: Compile the desktop main process
run: npm --prefix desktop run compile

- name: Acceptance tests (vitest, stubbed electron)
run: npm --prefix desktop test

Expand Down
6 changes: 4 additions & 2 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -209,8 +209,10 @@ transport maps to the contract 503.
| Sampler | top-p 0.9, repeat penalty 1.1, lookback window 8192 (generated tokens only) | [desktop/main/backend/inference/penalties.ts](desktop/main/backend/inference/penalties.ts) (`PENALTY_FULL_CONTEXT_TOKENS`) |
| History | at most 12 turns carried into the prompt | `MAX_HISTORY_TURNS` |

Vulkan stays off (reserved). The memory governor can force the runtime
profile to `fast` (below); the wizard's RAM gate uses the same estimate
The compute backend is decided by an out-of-process probe (issue #155): a working
Vulkan device is used when one is found and CPU inference is the fallback;
`inference.vulkan` can pin either choice. The memory governor can force the
runtime profile to `fast` (below); the wizard's RAM gate uses the same estimate
family (file size + 1 GiB KV-cache + 1 GiB overhead,
[desktop/main/first-run/ram-gate.ts](desktop/main/first-run/ram-gate.ts)).

Expand Down
30 changes: 29 additions & 1 deletion CHANGELOG.md

Large diffs are not rendered by default.

6 changes: 4 additions & 2 deletions INSTALL.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,8 @@ Complete installation guide for the Document Q&A Assistant, including standard P
- **Python**: 3.11 or higher

### Optional Components
- **NVIDIA GPU**: Not required for GGUF backend (CPU-only inference)
- **NVIDIA GPU**: not required. The GGUF backend uses a working Vulkan device when one is
present and falls back to CPU inference otherwise

## Standard Installation

Expand Down Expand Up @@ -262,7 +263,8 @@ Not required. Application runs as a standard executable.

### System Requirements for Power Users

The application runs CPU-only using GGUF models. No GPU or NPU acceleration is required.
The application runs GGUF models on a GPU when a probe finds a working Vulkan device, and on
the CPU otherwise. GPU acceleration is optional; no particular GPU or NPU is required.

## Post-Installation

Expand Down
15 changes: 9 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -86,15 +86,16 @@ The desktop app runs GGUF models via node-llama-cpp (Node main-process backend,
- **Quality profile (default)**: Gemma 4 E2B-it (Q4_K_M GGUF per [ADR-0002](docs/adr/0002-llm-profiles.md); ~2.9 GB nominal per PACKAGING.md; 2,620,370,976 bytes measured per bench/RESULTS.md) — bundled
- **Fast profile**: lfm2.5-vl-450m (Q4_K_M GGUF per ADR-0002) — bundled
- Profile selection: automatic free-RAM gate or in-app choice; `TRAININGAPP_DESKTOP_INFERENCE_PROFILE` (`quality` / `fast` / `auto`) is the desktop env override. (`RAG_GGUF_PATH` / `--gguf-path` select a custom GGUF on the legacy Python harness only.)
- No GPU required
- No particular GPU required: a working Vulkan device is used when one is found, and
CPU inference is the fallback
- No network access required (unless you turn on the external model or update checks)
- Measured decode throughput and first-token latency per model/profile: see [bench/RESULTS.md](bench/RESULTS.md) (issue #52 benchmark harness)

### Hardware Requirements
#### Minimum (Intel 11th Gen i5, 16GB RAM)
- Windows 11 (64-bit)
- Intel Core i5 11th generation or newer (or equivalent AMD Ryzen 5000+)
- Intel integrated graphics (present on all 11th gen+ Intel CPUs) — no discrete GPU required
- Intel integrated graphics (present on all 11th gen+ Intel CPUs); a discrete card is not needed
- 16GB RAM
- ~6.4 GB free storage for models + app (measured installed footprint 6,378,451,601 bytes; staged model resources 4,111,872,009 bytes; see bench/RESULTS.md)
- **Performance**: measured numbers per model and quantization are recorded in [bench/RESULTS.md](bench/RESULTS.md)
Expand All @@ -109,7 +110,7 @@ The desktop app runs GGUF models via node-llama-cpp (Node main-process backend,
#### High-Performance (Intel 13th Gen i9, 64GB RAM)
- High-end CPU (Intel Core i9 or AMD Ryzen 9)
- 64GB RAM
- **Performance**: measured CPU-only GGUF numbers are recorded in [bench/RESULTS.md](bench/RESULTS.md)
- **Performance**: measured CPU and GPU numbers, and which machine each was measured on, are recorded in [bench/RESULTS.md](bench/RESULTS.md)

> **Pending**: the offline/low-RAM reference-laptop validation matrix (issue #86) —
> reference-i5 rows are not yet measured; no reference-hardware numbers are claimed here.
Expand Down Expand Up @@ -298,7 +299,9 @@ The desktop app runs GGUF models via node-llama-cpp (Node main-process backend,
verification of the bundled tree → knowledge-pack activation → license notices →
complete. Every gate names its failure reason; setup can be re-run from Settings.

No Python, no GPU, and no network access are required (unless you turn on the external model or update checks).
No Python and no network access are required (unless you turn on the external model or update
checks). GPU acceleration is used when the machine has one that works; otherwise inference
runs on the CPU.

### Building the desktop app from source

Expand Down Expand Up @@ -679,7 +682,7 @@ Mirrors [ARCHITECTURE.md](ARCHITECTURE.md) (the authoritative map):
- Cross-encoder rerank (ettin-reranker-32m-v1) with a calibrated relevance floor

**LLM Interface**
- GGUF via node-llama-cpp (desktop, CPU-only, fully offline)
- GGUF via node-llama-cpp (desktop; GPU when a probe finds one working, CPU otherwise, fully offline)
- GGUF via wllama WASM (browser, CPU/SIMD, fully offline)

**RAG Engine**
Expand Down Expand Up @@ -1152,4 +1155,4 @@ Legacy Python harness only (CI conformance surface):
---
**Version**: 2.3.0
**Last Updated**: 2026-09-29 (v3 documentation refresh, issue #89)
**Hardware**: CPU-only optimized for Intel 11th gen i5 and above (16GB RAM minimum)
**Hardware**: optimized for Intel 11th gen i5 and above (16GB RAM minimum); GPU acceleration is opportunistic and CPU inference is the floor
18 changes: 18 additions & 0 deletions api_server.py
Original file line number Diff line number Diff line change
Expand Up @@ -771,6 +771,24 @@ async def get_status_models(auth: dict = Security(require_auth())):
)


@app.post("/settings/inference/gpu-test")
async def test_gpu_inference(auth: dict = Security(require_auth())):
"""
GPU capability re-probe (issue #155).

The probe runs as a child process against node-llama-cpp's Vulkan backend
on the Electron/Node desktop surface only
(desktop/main/backend/inference/gpu-probe.ts); this Python host has no
llama.cpp GPU backend and never wires a probe, so the route exists to keep
the shared contract (contracts/api.openapi.yaml) consistent across backends
and always answers the documented unwired 503 - the path is known, never
404-absent.
"""
raise HTTPException(
status_code=503, detail="GPU probing is not wired on this host"
)


# --- C7 (issue #74): knowledge pack lifecycle -------------------------------
# Same wire shapes as the Node backend (desktop/main/backend/server.ts).
# C8 (issue #75): zip extraction is delegated to the shared pack_extract
Expand Down
88 changes: 88 additions & 0 deletions bench/RESULTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,6 +63,53 @@ different machine.
| wllama | PENDING |
| onnxruntime | PENDING |

### reference-amd-igpu

Registered by PR #159 review PRR-038 so the GPU device matrix can name a machine that
does not exist yet. `bench/append_results.py` refuses any row whose `machine` tag is
absent from this registry. No AMD host has been measured; every field stays PENDING.
Fill these in only on that hardware, using the same commands as `devstation` above.

NOTE: each tag needs its OWN `### ` heading. `append_results.py` takes the whole
heading line as one literal tag, so a combined "a / b / c" heading would register as
the single string `a / b / c` and leave all three individual tags unregistered.

| field | value |
|---|---|
| CPU model | PENDING |
| RAM | PENDING |
| OS build | PENDING |
| GPU + driver | PENDING |
| node-llama-cpp (the engine issue #155 probes) | PENDING |
| Backend the probe selects | PENDING |

### reference-amd-dgpu

As `reference-amd-igpu` above, for a discrete AMD GPU.

| field | value |
|---|---|
| CPU model | PENDING |
| RAM | PENDING |
| OS build | PENDING |
| GPU + driver | PENDING |
| node-llama-cpp (the engine issue #155 probes) | PENDING |
| Backend the probe selects | PENDING |

### reference-nvidia

As `reference-amd-igpu` above, for an NVIDIA GPU. Note the CUDA backend is
deliberately not shipped, so this host is expected to resolve to Vulkan-or-CPU.

| field | value |
|---|---|
| CPU model | PENDING |
| RAM | PENDING |
| OS build | PENDING |
| GPU + driver | PENDING |
| node-llama-cpp (the engine issue #155 probes) | PENDING |
| Backend the probe selects | PENDING |

## Native llama.cpp CPU results

One row per (model x quant x threads x prompt length); decode tok/s over the
Expand Down Expand Up @@ -144,6 +191,7 @@ see the provenance rule 4 above) on the real staged tree, machine-tagged:
| devstation | models/reranker (ettin-reranker-32m-v1 q8 + root tokenizers) | 39,611,408 | rerank-worker dtype q8; AutoTokenizer loads from the model ROOT |
| devstation | models/llm-quality (gemma-4-e2b-it Q4_K_M + mmproj) | 3,606,025,056 | ADR-0002; mmproj has no native consumer at E1 |
| devstation | models/llm-fast (lfm2.5-vl-450m Q4_K_M + mmproj) | 332,128,736 | ADR-0002 |
| devstation | backend packages EXCLUDED from the installer: @node-llama-cpp/win-x64-cuda (170,658,131) + win-x64-cuda-ext (362,957,501) | 533,615,632 | issue #155: unreachable - the probe passes `exclude: ['cuda']`. Measured `du -sb desktop/node_modules/@node-llama-cpp/win-x64-cuda{,-ext}`, i.e. 508.90 MiB / 533.6 MB decimal. The "~510 MB" quoted in `desktop/electron-builder.yml` is the MiB figure. PR #159 review PRR-037 |
| devstation | packs (bundled-docs + training fixtures) | 1,695 | contracts/fixtures/packs layout fixtures |
| devstation | docs (licenses.md) | 6,240 | the first-run licensing seam |
| devstation | staged resources total | 4,111,872,009 | 3.83 GiB (17 model files after the review round added the reranker root tokenizers) |
Expand All @@ -152,6 +200,46 @@ see the provenance rule 4 above) on the real staged tree, machine-tagged:
| devstation | startup integrity gate latency (streaming sha256 of the full staged tree) | 2,377 ms | packaged-mode pass, 0 failures, 2026-09-23 |
| reference-i5 | all E1 size/latency rows | PENDING | operator runs the same commands (E3/#86 owns the matrix) |

### GPU device matrix (issue #155)

Which compute backend the desktop backend actually selects on each machine. Issue #155 turned
GPU acceleration from a hard-disabled constant into an out-of-process probe plus a CPU
fallback, so a row here is a statement about what the probe concludes there.

Recording discipline for this table (the same one the `reference-i5` E1 rows above already
follow): every row is either a figure measured on the machine it names, or an explicit `PENDING`
naming what is unmeasured. No tok/s number appears here unless it was measured on the machine
named. The issue #155 comments carry prefill/decode tok/s figures from an external benchmark
harness; this table does not restate them as its own evidence, because this revision did not
re-measure them.

| machine | device | backend selected | GPU offload | notes |
|---|---|---|---|---|
| devstation | Intel Arc Pro B50 (discrete), driver 32.0.101.8805 | vulkan | supportsGpuOffloading=true | measured: node-llama-cpp 3.20.0, same process reported gpu=false for the shipped CPU-only option and gpu=vulkan for `{gpu:{type:auto,exclude:[cuda]}}`. Reproduce: `node -e "import('node-llama-cpp').then(n=>n.getLlama({gpu:{type:'auto',exclude:['cuda']},build:'never'}).then(l=>console.log(l.gpu,l.supportsGpuOffloading)))"` |
| devstation | Intel Arc Pro B50 (discrete), driver 32.0.101.8805 | cpu | gpu=false, supportsGpuOffloading=false | measured: the same host and the same node-llama-cpp 3.20.0 process under the shipped `{gpu:false}` option - the comparison the GPU row above is measured against. Reproduce: `node -e "import('node-llama-cpp').then(n=>n.getLlama({gpu:false,build:'never'}).then(l=>console.log(l.gpu,l.supportsGpuOffloading)))"` |
| reference-i5 | 12th-gen Core i5 mobile, Intel Iris Xe integrated | PENDING | PENDING | no integrated-GPU host was available; prefill and decode tok/s and the probe's own verdict are unmeasured, and this is the floor-spec machine the speed bar for the quality tier depends on |
| reference-amd-igpu | AMD integrated, RDNA | PENDING | PENDING | no AMD host was available; backend selection, probe verdict and throughput are unmeasured |
| reference-amd-dgpu | AMD discrete, RDNA | PENDING | PENDING | no AMD host was available; backend selection, probe verdict and throughput are unmeasured |
| reference-nvidia | NVIDIA discrete | PENDING | PENDING | no NVIDIA host was available; note the CUDA backend is deliberately not shipped, so this host is expected to resolve to Vulkan-or-CPU rather than CUDA |

PR #159 review PRR-038: the three `reference-amd-*` / `reference-nvidia` tags above are
not in this file's machine registry, and `bench/append_results.py` raises `SystemExit`
on any unregistered tag - so a future measured row for them would be rejected. They
are registered here with the all-PENDING shape the registry already uses for
`reference-i5` (which likewise has `PENDING` cells), so the tag is usable the day
that hardware appears.

One upstream llama.cpp defect has no runtime mitigation in the pinned library and is covered
only by the probe plus its CPU fallback, not fixed here: **#27638** (device loss at
`ubatch >= 2048`; `ubatch` does not appear anywhere in the installed node-llama-cpp 3.20.0
`dist/` except one code comment). Verified by searching the installed package, not assumed.

**#29054** (deterministic hang on a q8_0 KV cache) is a different case and is **not** claimed as
unmitigated: node-llama-cpp 3.20.0 does expose `experimentalKvCacheKeyType` /
`experimentalKvCacheValueType` on `LlamaContextOptions`, and **both already default to F16**, so
the f16 mitigation the upstream report asks for is in force by default. The application does not
surface that experimental override to users, deliberately.

**Known limit (recorded, never hidden): the single-file NSIS target cannot
embed the real-weights payload.** `makensis.exe` aborts with `File: failed
creating mmap of …-x64.nsis.7z` because the app archive is 4,177,827,908
Expand Down
72 changes: 72 additions & 0 deletions contracts/api.openapi.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -434,6 +434,32 @@ paths:
"503":
description: Model status is not wired on this host

/settings/inference/gpu-test:
post:
tags: [settings]
summary: Re-run the GPU capability probe (issue #155)
description: >-
Runs the out-of-process GPU probe again and adopts the result, so a
machine whose driver or hardware changed can recover from a stale
negative verdict. Takes no request body and persists nothing itself; the
host adopts and persists the returned verdict.

"No usable GPU" is an ANSWER, not an error, so a failed probe is a 200
with `ok: false` and a populated `reason` - the same shape
`POST /settings/external/test` uses. A 503 means the host has no probe
wired at all (a known route that cannot be served, never a 404).
Transport-token-guarded like every route.
operationId: testGpu
responses:
"200":
description: The probe ran; the verdict says whether a GPU is usable.
content:
application/json:
schema:
$ref: "#/components/schemas/GpuProbeResponse"
"503":
description: GPU probing is not wired on this host

/packs:
get:
tags: [packs]
Expand Down Expand Up @@ -572,6 +598,24 @@ components:
in: header
name: X-API-Key
schemas:
GpuProbeResponse:
type: object
description: >-
The result of a GPU capability probe run (issue #155). Same shape as the
`gpu` member of StatusModelsResponse.
required: [backend, ok, reason]
properties:
backend:
type: string
enum: [vulkan, cpu]
ok:
type: boolean
reason:
type: string
device:
type: string
nullable: true

StatusModelsResponse:
type: object
required: [engine, profile, models]
Expand All @@ -585,6 +629,34 @@ components:
GGUFs never block chat.
profile:
type: string
gpu:
type: object
description: >-
The GPU decision (issue #155), as the engine actually made it.
Present for the local llama.cpp backend and absent when an external
endpoint generates, because then no local compute backend is
involved. `backend` is what runs; `ok` is whether a GPU was usable;
`reason` is ALWAYS non-empty so a CPU fallback can be explained
rather than guessed at.
required: [backend, ok, reason]
properties:
backend:
type: string
enum: [vulkan, cpu]
description: >-
The resolved compute backend. CPU means either the probe found no
usable device, the operator pinned CPU, or nothing has been
probed yet.
ok:
type: boolean
description: True only when a GPU backend loaded AND produced a sane generation.
reason:
type: string
description: Human-readable explanation; never empty.
device:
type: string
nullable: true
description: Adapter identity the backend reported, when it reported one.
resident:
type: object
description: >-
Expand Down
Loading
Loading