A fully offline RAG-based document question answering system optimized for Windows PCs. Features semantic search, hybrid retrieval, and CPU-based LLM inference with GGUF models. Nothing leaves the machine unless you turn on the external model or update checks (both off by default).
The shipped delivery options share the same offline RAG capabilities:
- Desktop app (primary) β an Electron installer with a first-run wizard, a Node main-process backend (ADR-0003), bundled models (Quality/Fast profiles per ADR-0002), and knowledge packs β fully offline after install.
- HTML5 web app (
web_ui/) β a fully self-contained, STIG-scannable archive that runs entirely in the browser with no runtime downloads (the same build is the desktop app's renderer).
A legacy Python harness (api_server.py, pip install) survives only as the CI
contract-conformance surface β see the scope notes at the top of USAGE.md, INSTALL.md,
and CONFIGURATION.md.
- Knowledge packs β installable document/training packs built and verified with
packtool, managed in-app (install / supersede / rollback / remove); see the pack authoring guide and the training-pack refresh runbook. - Learn panel with Open-in-training deep links β answers cite the training slides that teach them, and jump straight into the embedded Storyline player.
- First-run wizard β hardware detection, profile selection, sha256 integrity verification of the bundled tree, pack activation, and license notices.
- Signed, opt-in update channel β Ed25519-signed pack updates and detect-and-notify app updates, off until you switch them on (ADR-0010, docs/updates.md).
- Quality/Fast inference profiles β
gemma-4-e2b-itQ4_K_M vslfm2.5-vl-450mQ4_K_M with a free-RAM auto gate (ADR-0002).
See ARCHITECTURE.md for the full system map.
The browser app is a complete, offline RAG client. See PACKAGING.md for the build/bundle steps.
- Fully offline, packaged models β embeddings (arctic-embed-m ONNX), ONNX Runtime WASM, and the
browser LLM are served same-origin from
public/models/; nothing is fetched from a CDN or the HuggingFace Hub at runtime. A readiness gate reports "models ready vs missing". - Browser LLM engine: wllama β llama.cpp WASM, CPU/SIMD, no WebGPU, the default and most robust on i5/Iris Xe. A hardware-capability panel detects WebGPU/threads/memory. The WebLLM (WebGPU) engine remains selectable code but ships no weights in the offline manifest, so it is not part of the air-gapped configuration.
- Multimodal β attach a screenshot in chat and ask about it (wllama + Gemma 4 E2B-it mmproj), offline.
- Chat UX β streaming with interactive source citations, regenerate, conversation export (Markdown/JSON), and Fast/Balanced/Quality RAG presets.
- Self-contained archive β
npm run build:offlineproduces a validatedweb_ui/dist/the Electron desktop app (or any root static host) serves with the COOP/COEP headers wllama needs. - Optional external model (both apps, off by default) β Settings β Model & connection connects to any OpenAI- or Anthropic-compatible server on this computer, on your network, or in the cloud (with an API key). Retrieval stays local; see the "External model" section of CONFIGURATION.md and ADR-0011.
- Application Shell: Navigation rail with Chat, Documents, Settings pages and responsive flexbox layout
- Theme System: Dark/light mode toggle with system preference detection and localStorage persistence
- Design Token Foundation (Phase 1): Comprehensive CSS custom property system on 8px grid with Inter font, status color tokens (info/warning/success), and radius tokens (sm/md/lg)
- Toast Notifications: Non-blocking toast system with success/error/info variants, entrance and exit fades, and a 5 second auto-dismiss that pauses on hover and focus (only the dismiss button closes a toast early), at most 5 toasts at once (the oldest is dropped), identical messages not duplicated (the repeat restarts the timer), and a viewport portaled to
document.body - Keyboard Shortcuts: Ctrl+Enter (send), Ctrl+L (clear chat), Ctrl+, (open settings) with input/textarea focus guard
- Testing Framework: vitest configured with @testing-library/react and jsdom environment
- Centered Transcript Layout: Message list now centered with 768px max-width for comfortable reading
- Rich Empty State: Hero heading "How can I help with your documents?" with 3 clickable suggested prompt cards
- Suggested Prompts: Click any prompt card to immediately send that question to the chat
- Assistant Message Styling: Full-width prose layout (no bubble background/radius) for improved readability
- User Message Styling: 75% width bubbles aligned right, maintaining visual distinction
- Action Row Copy Button: Copy button relocated below message content in a dedicated action row
- Composer Redesign: Raised card input with 20px radius, enhanced focus feedback (border color + shadow), elevation shadow, and 12px radius buttons
- Offline-First Design: No internet required after initial setup
- Multi-format Support: PDF, DOCX, PPTX, TXT, MD documents
- Hybrid Retrieval: keyword (FTS5/BM25) + vector search fused with Reciprocal Rank Fusion (RRF, k=60 on every surface)
- Window Expansion: Automatically fetches adjacent context chunks
- Smart Chunking: Paragraph and sentence boundary aware
- Cross-Encoder Reranking: ettin-reranker (ModernBERT) for precise ranking
The desktop app runs GGUF models via node-llama-cpp (Node main-process backend, ADR-0003); the browser app uses the same GGUF weights through wllama (llama.cpp WASM) β fully offline on both:
- Quality profile (default): Gemma 4 E2B-it (Q4_K_M GGUF per ADR-0002; ~2.9 GB nominal per PACKAGING.md; 2,620,370,976 bytes measured per bench/RESULTS.md) β bundled
- Fast profile: lfm2.5-vl-450m (Q4_K_M GGUF per ADR-0002) β bundled
- Profile selection: automatic free-RAM gate or in-app choice;
TRAININGAPP_DESKTOP_INFERENCE_PROFILE(quality/fast/auto) is the desktop env override. (RAG_GGUF_PATH/--gguf-pathselect a custom GGUF on the legacy Python harness only.) - No particular GPU required: a working Vulkan device is used when one is found, and CPU inference is the fallback
- No network access required (unless you turn on the external model or update checks)
- Measured decode throughput and first-token latency per model/profile: see bench/RESULTS.md (issue #52 benchmark harness)
- Windows 11 (64-bit)
- Intel Core i5 11th generation or newer (or equivalent AMD Ryzen 5000+)
- Intel integrated graphics (present on all 11th gen+ Intel CPUs); a discrete card is not needed
- 16GB RAM
- ~6.4 GB free storage for models + app (measured installed footprint 6,378,451,601 bytes; staged model resources 4,111,872,009 bytes; see bench/RESULTS.md)
- Performance: measured numbers per model and quantization are recorded in bench/RESULTS.md
- Intel Core i7 12th generation or newer (or equivalent AMD Ryzen 7000+)
- Intel Iris Xe integrated graphics or discrete GPU
- 32GB RAM
- SSD for vector database
- Performance: measured numbers per model and quantization are recorded in bench/RESULTS.md
- High-end CPU (Intel Core i9 or AMD Ryzen 9)
- 64GB RAM
- Performance: measured CPU and GPU numbers, and which machine each was measured on, are recorded in bench/RESULTS.md
Pending: the offline/low-RAM reference-laptop validation matrix (issue #86) β reference-i5 rows are not yet measured; no reference-hardware numbers are claimed here.
The phases below describe the 2026
web_uioverhaul and are preserved for history; the current feature set is described above and in ARCHITECTURE.md.
- Streaming Chat Interface: Full-featured chat page (
ChatPage.tsx) with real-time token streaming display using RAF-batched updates viaTokenStreamManager - Role-Based Message Bubbles: Distinct styling for user, assistant, and system messages with relative timestamps ("2m ago", "just now")
- Inline Markdown Renderer: react-markdown + remark-gfm based renderer supporting CommonMark + GFM (tables, strikethrough, task lists, autolinks, nested emphasis), fenced code blocks with language chips and per-block copy, and a URL allowlist (allows http/https/mailto/tel; rejects javascript:, data:, and scheme-less/relative URLs)
- Source Citation Pills: Expandable/collapsible source pills with filename truncation, full path reveal on click, and one-click copy-to-clipboard
- Inference Mode Toggle: Status indicator (green/yellow/red) for browser-local vs API mode with server connectivity check against
/auth/statusendpoint - Streaming Cursor Animation: Blinking cursor (
@keyframes blink) appended to assistant messages during streaming for visual feedback - Copy Message: Hover-to-reveal copy button on user and assistant bubbles with 1.5s "Copied!" feedback
- Streaming Indicator: Bouncing dots animation (setInterval-based, 3 dots cycling at 200ms) shown below messages during generation
- Operation Cancellation: Cancel button stops
TokenStreamManager, clears pending mock timers, and marks streaming messages complete
Historical (Phase 3): the browser API-server mode and its server URL described here were removed by settings-wiring-honesty, and with them the browser-side "Server not connected" header warning;
apimode now exists only in the desktop app (where that warning still reports the desktop backend). See CONFIGURATION.md, App Settings.
- Dual-Mode Context:
InferenceModeContext(InferenceModeContext.tsx) managesbrowser-localvsapimode via React context - localStorage Persistence: Mode preference and server URL stored under
inference-modekey; survives page refresh - Server Connectivity Check:
checkServerConnectivity()pings/auth/statuswith 5s timeout, handles abort for rapid toggles, updatesisServerConnectedandmodeErrorstate - Model Loading Progress:
modelLoadingProgress(0β100) displayed in blocking overlay when browser-local model is initializing - API Mode Warning: "Server not connected" warning shown in header when API mode is active but server is unreachable
- Dexie.js Integration: IndexedDB-based conversation persistence via
DocQADatabaseclass (db/index.ts) and CRUD operations (db/conversations.ts) with pagination support - Sidebar Navigation: Responsive 260px sidebar (
Sidebar.tsx) with collapsible state, showing conversation history - Conversation Context Menu: Right-click to delete or rename conversations (
SidebarConversationItem.tsx) - Controlled ChatPage: Refactored with
messages,onMessagesChange, andonSaveConversationprops for explicit state management - App Wiring:
useConversationshook connects AppLayout and ChatPage for automatic conversation loading and saving - Simplified Header: Compact padding, right-aligned controls, removed title text
- Elevation Tokens: A shadow hierarchy and surface colors (since replaced by the Lumen
--shadow-*and--bg-*tokens) for consistent depth - Relative Timestamps:
relativeTime.tsutility formats conversation timestamps as "2m ago", "Yesterday", etc.
- Expandable Citations: Click to expand source pills showing full filename, page number, and content preview
- One-Click Copy: Copy button on each pill copies citation text to clipboard
- Hover Preview: Hover shows truncated source preview with tooltip for full content
- Phase Attribution: Pills labeled with "Phase 3" or "Phase 4" indicating extraction source
- CTkTooltip Class: Non-blocking hover tooltips with 500ms delay for all settings fields
- Contextual Help: Each RAG configuration field has descriptive hint text explaining its purpose
- Dark Theme Tooltips: Tooltips use dark background (#3a3a4e) with white text for consistent visibility
- Browser-Side Extraction: All document processing happens locally in the browser with no server uploads
- Multi-Format Support: PDF, DOCX, XLSX, PPTX, TXT, and MD files via dedicated extractors
- Extractor Factory:
ExtractorFactoryselects the appropriate extractor based on MIME type - Semantic Chunking: Faithful Python port with paragraph/sentence boundary awareness, configurable overlap, page mapping, and SHA256 content IDs
- IndexedDB Storage: Documents, chunks, and metadata persisted locally via
document-store.ts - Documents Page: Full-featured
/documentspage with drag-and-drop upload, file processing pipeline, and document list with status tracking - DropZone Component: Drag-and-drop or click-to-browse file input with visual feedback and progress indication
- DocumentList Component: Paginated document list showing name, type, size, status, and date with delete functionality
| Format | Extractor | Library |
|---|---|---|
pdf-extractor.ts |
pdfjs-dist | |
| DOCX | docx-extractor.ts |
mammoth |
| XLSX | xlsx-extractor.ts |
xlsx |
| PPTX | pptx-extractor.ts |
jszip + xml parsing |
| TXT/MD | txt-extractor.ts |
Native text processing |
pdfjs-dist^4.4.168mammoth^1.8.0xlsx^0.18.5jszip^3.10.1
- Real-time UI Updates: Font size slider now applies to all widgets immediately when saved
- Debug Mode: Toggle debug-level logging for troubleshooting
- Log File Persistence: Customizable log file path with automatic persistence
- Auto-Reconfiguration: RAG settings (chunk size, n_results, etc.) trigger engine reinitialization when changed
- Thread-Safe RAG Engine: Full serialization via
asyncio.to_thread()wrapping for blocking endpoints - ChromaDB Locking:
RLockfor vector store operations preventing concurrent access corruption - BM25 Index Threadsafety: Incremental add operations protected by RLock for safe concurrent document ingestion
- Lazy LLM Initialization: On-demand LLM loading reduces memory footprint for CLI/API modes
- Cancellation Propagation:
cancellation_eventpassed through query processing for responsive long-operation termination - Memory Budget Checks: Pre-ingestion memory validation prevents OOM errors on large document sets
- QueryTransformer Singleton: Shared transformer instance across requests with thread-safe initialization
- Cross-Encoder threadsafety:
__new__pattern ensures single instance with RLock for concurrent reranking - Neighborhood Expansion: Increased k from 3 to 5 chunks for better context coverage in streaming mode
- Embedding Batch Normalization: Consistent batch sizes for predictable memory usage during ingestion
- Transformers.js Embeddings: Browser-side embedding generation using snowflake-arctic-embed-m-v1.5 ONNX model for offline use
- HNSW Vector Index:
EdgeVecRust/WASM-based HNSW index with native IndexedDB persistence for semantic search - FlexSearch Keyword Index: Full-text keyword search with resolution-based scoring for BM25-style matching
- Reciprocal Rank Fusion: Ported RRF algorithm for hybrid retrieval combining semantic and keyword results
- Cross-Encoder Reranking: ettin-reranker-32m-v1 (ModernBERT) reranker with conditional activation (skipped on low-memory devices)
- Memory-Aware Model Selection: Device memory detection with tier-based configuration (low/medium/high memory tiers)
- Browser LLM engines: wllama (llama.cpp WASM, CPU/SIMD, the default) drives the
bundled Gemma 4 E2B-it GGUF + mmproj weights β
web_ui/public/models/manifest.jsonships the wllama runtime and the Gemma files, nothing is fetched at runtime. The Phase-6 WebLLM (WebGPU) engine (@mlc-ai/web-llm, retired model era) remains selectable code but ships no weights in the offline manifest, so it is not part of the air-gapped configuration (see PACKAGING.md) - Model Download Manager: Progress tracking with speed/ETA calculation, cancellation support, and storage quota error handling
- ModelDownloadProgress UI: Accessible progress bar with ARIA attributes, download speed, ETA countdown, and cancel button
- Model Readiness Gate: Pre-flight checks for WebGPU availability, memory sufficiency (2GB minimum), and OPFS cache status; guides users to the wllama engine or an external model server (Settings β Model & connection) when requirements aren't met
- RAG Orchestrator: Full retrieval pipeline connecting embeddingβvector searchβkeyword searchβRRF fusionβrerankingβLLM generation; emits typed
RAGEventstream for UI progress - WebGPU Watchdog: Context loss detection via
GPUDevice.lostpromise/event monitoring;createRecoveryHandlerautomatically re-initializes the service after loss
- Thinking Indicator: Animated "Thinking..." with dots while LLM generates responses
- Smart Regeneration: "Regenerate" button replaces the last assistant message instead of creating duplicates
- Feedback System: Working thumbs up/down buttons that persist to database
- Conversation Context Menu: Right-click options to delete or rename conversations
- Time Display: Relative timestamps in sidebar (e.g., "2 min ago", "Yesterday")
Historical (Phase 7): the Inference Mode toggle (browser-local vs API server), the Server Configuration / Server URL controls, the Model Selection dropdown and the stored
serverUrlpreference listed here no longer exist; the current Settings page, including the External model region, is described in CONFIGURATION.md, App Settings.
- Dedicated Settings Page (
SettingsPage.tsx): Full-featured settings UI with 6 sections:- Inference Mode: Toggle between browser-local (WebGPU) and API server modes with real-time state sync
- Server Configuration: Server URL input with connection test button and status indicators
- Model Selection: Dropdown for AI model choice with cache status, download progress, and cancel support
- Appearance: Theme selector (light/dark/system) with immediate UI application
- Storage: Memory budget display, memory pressure status, and two-click cache clear with confirmation
- About: Version info and app description
- IndexedDB Persistence: User preferences (theme, preferredModel, serverUrl) stored in IndexedDB with automatic load/save
- InferenceModeProvider at Root: Provider moved to
App.tsxroot level for shared state across all pages (Chat, Documents, Settings) - Cross-Browser Compatibility (
browser-compat.ts): Detection for Chrome/Edge 113+ (full WebGPU) and Firefox 112 or newer (supported; WebGPU is experimental there, so in-browser model features may be degraded); classifies the browser by name and version so the app can show the unsupported-browser notice. Firefox is verified in CI by the Playwright browser e2e suite (it runs under both Chromium and Firefox). Safari and every other WebKit browser (all iOS/iPadOS browsers included, FxiOS among them) are classified unsupported (as is Firefox below 112), and the browser app shows a dismissible "This browser isn't supported" notice at load (the course player's start-failure message names Safari as unsupported too) - Reusable UI Components:
ErrorBoundary.tsx: Class-based error boundary catching render errors with retry functionalityui/EmptyState: Contextual empty state primitive from the Lumen component library (src/ui/); the earlierLoadingSkeletonandcomponents/EmptyState.tsxwere removed (Lumen phases 7-8)
- Dual-Mode Streaming:
ChatPagenow connects toRAGOrchestratorfor browser-local inference (WebGPU) andSSEStreamConsumerfor API server streaming, with seamless mode switching - DocumentsPage Search Wiring: Document search now uses the full search pipeline (vector-index + keyword-index + RRF fusion)
- Service Initialization Hook (
useServiceInitialization.ts): Sequential service initialization with proper cleanup on unmount; manages embedding service, vector index, and keyword index lifecycle - Loading Overlay: Service initialization state surfaced via blocking overlay in
App.tsxduring startup - Production Build Fixes: edgevec WASM snippet stub plugin for Vite; pdfjs worker initialization fix for production
- Enter Key Submission: Press Enter to submit questions (no need to click "Ask" button)
- Escape Key: Clears input field or cancels active operations
- Ctrl+Enter: Alternative shortcut for submitting questions
- Ctrl+L: Quick clear chat shortcut
- Ctrl+,: Open settings dialog shortcut
- Inline Typing Indicator: "Thinking..." indicator appears in chat area while processing (replaces status bar overwrite)
- Clear Chat Confirmation: Clear button requires a second click within 3 seconds to prevent accidental deletion
- Settings Switch Labels: CTkSwitch widgets now display descriptive text labels ("Enable Hybrid Search", "Enable Reranking")
| Component | File | Description |
|---|---|---|
ChatPage.tsx |
src/pages/ |
Primary chat page with streaming, message state, send/cancel/clear |
ChatMessageList.tsx |
src/components/ |
Centered transcript (768px max-width), rich empty state with suggested prompts (Phase 2) |
ChatMessageBubble.tsx |
src/components/ |
Role-based messages: assistant full-width prose, user 75% bubbles, action-row copy (Phase 2) |
ChatInput.tsx |
src/components/ |
Raised card composer with focus feedback, elevation shadow, 20px radius (Phase 2) |
MarkdownRenderer.tsx |
src/components/ |
react-markdown + remark-gfm renderer (CommonMark + GFM, URL allowlist) |
SourceCitation.tsx |
src/components/ |
Expandable citation pills with copy-to-clipboard |
InferenceModeToggle.tsx |
src/components/ |
Status dot (green/yellow/red) for browser-local vs API mode |
StreamingIndicator.tsx |
src/components/ |
Bouncing dots animation during generation |
DropZone.tsx |
src/components/ |
Drag-and-drop file upload with progress indication |
DocumentList.tsx |
src/components/ |
Paginated document list with status tracking |
ModelDownloadProgress.tsx |
src/components/ |
Accessible progress bar for model download |
ErrorBoundary.tsx |
src/components/ |
Error boundary with retry functionality |
EmptyState |
src/ui/ |
Contextual empty state (Lumen primitive; replaces the removed LoadingSkeleton and components/EmptyState.tsx) |
Sidebar.tsx |
src/components/ |
Responsive 260px sidebar with conversation history (Phase 3) |
SidebarConversationItem.tsx |
src/components/ |
Conversation list item with context menu (Phase 3) |
- Download the installer β unsigned NSIS x64 build (see the repo Releases; the measured installed footprint is ~6.4 GB, staged model resources 4,111,872,009 bytes β bench/RESULTS.md). Models, knowledge packs, and license docs are bundled; nothing downloads at runtime.
- Run the installer and launch the app.
- Complete the first-run wizard: hardware detection β profile selection (Quality/Fast, with an automatic free-RAM recommendation) β sha256 integrity verification of the bundled tree β knowledge-pack activation β license notices β complete. Every gate names its failure reason; setup can be re-run from Settings.
No Python and no network access are required (unless you turn on the external model or update checks). GPU acceleration is used when the machine has one that works; otherwise inference runs on the CPU.
cd desktop
npm install
npm run desktop:build # builds the web_ui renderer, stages models/packs, compiles,
# generates the sha256 manifest, runs electron-builder (NSIS x64)
npm run desktop:dev # vite dev server + Electron, for development
npm test # desktop test suitesModel weights must be staged first β see PACKAGING.md and desktop/README.md for the staging and packaging details.
cd web_ui
npm install
npm run dev # Development server
npm run build:offline # Self-contained offline archive (see PACKAGING.md)
npm run typecheck # TypeScript validation
npm test # Run tests with vitestThe Python stack below is not the shipped product β it survives as the CI
contract-conformance surface for the frozen API (contracts/tests/run_conformance.py).
It is retained here for maintainers running that suite.
- Windows 10 or later
- Python 3.10+
- pip package manager
-
Clone or download the repository
cd doc_qa_app
-
Install dependencies
pip install -r requirements.txt -
Models
GGUF Model (Required for LLM inference)
# The harness uses a local GGUF via RAG_GGUF_PATH (e.g. the ADR-0002 # gemma-4-e2b-it Q4_K_M model.gguf); any GGUF format model works # From Hugging Face: https://huggingface.co/models?search=gguf
Embedding Model (Required for search)
# Snowflake/snowflake-arctic-embed-m-v1.5 is packaged for offline use # Can be manually downloaded if needed for offline installation
-
Run the harness
GUI Mode (default):
python main.py
CLI Mode:
python main.py --cliAPI Server:
python main.py --api --port 8080
-
Download the offline installer bundle
- Includes Python embeddable, wheels, and model files
-
Extract the bundle
- Unzip to a directory on your machine
-
Install
- Run the provided installer or execute
main.py
- Run the provided installer or execute
-
No internet required after installation
Desktop app (main override seams; more are documented in docs/electron-mode.md):
| Variable | Description | Default |
|---|---|---|
TRAININGAPP_DESKTOP_INFERENCE_PROFILE |
Force the inference profile (quality / fast / auto) |
auto (free-RAM gate) |
TRAININGAPP_DESKTOP_BACKEND_MODE |
Backend host selection (node / sidecar) |
node (ADR-0003) |
TRAININGAPP_DESKTOP_DEV_ORIGINS |
Extra dev origins allowed by the loopback guard (unpackaged builds only) | - |
TRAININGAPP_DESKTOP_FREE_RAM_BYTES |
Override free RAM for the wizard's RAM gate (dev/test seam) | real reading |
Legacy Python harness only (api_server.py / main.py):
| Variable | Description | Default |
|---|---|---|
RAG_DB_PATH |
Vector database location | ./doc_qa_db |
RAG_GGUF_PATH |
Path to GGUF model file | - |
RAG_CHUNK_SIZE |
Document chunk size (words) | 512 |
RAG_N_RESULTS |
Context chunks to retrieve | 3 |
RAG_MAX_TOKENS |
Max response tokens | 1024 |
RAG_TEMPERATURE |
LLM temperature | 0.3 |
API_PORT |
API server port | 8080 |
The shipped desktop app has no user-facing auth: its backend binds loopback only, on an OS-assigned random port, and every request carries a per-launch 256-bit
X-Desktop-Tokenβ see docs/security/desktop.md and docs/electron-mode.md. TheENABLE_AUTH/API_KEYmaterial below applies to the legacy Python harness only.
Set both environment variables to enable authentication:
| Variable | Description | Example |
|---|---|---|
ENABLE_AUTH |
Enable authentication (any value enables) | true |
API_KEY |
Secret API key for authentication | your-secure-api-key |
export ENABLE_AUTH=true
export API_KEY="your-secure-api-key"
python main.py --api --port 8080$env:ENABLE_AUTH=$true
$env:API_KEY="your-secure-api-key"
python main.py --api --port 8080All API requests require authentication headers:
- API Key:
X-API-Key: <your-api-key> - JWT Bearer Token:
Authorization: Bearer <jwt-token>
import requests
import os
# Configure authentication
os.environ["ENABLE_AUTH"] = "true"
os.environ["API_KEY"] = "your-secure-api-key"
# Make authenticated request
headers = {
"X-API-Key": os.environ["API_KEY"]
}
response = requests.post("http://localhost:8080/ask", json={
"question": "What are the main findings?",
"n_results": 3
}, headers=headers)
print(response.json())- Always use HTTPS in production
- Rotate API keys regularly
- Store API keys in environment variables, never in code
- See USAGE.md for complete authentication documentation
Backend Selection (legacy Python harness):
The harness uses GGUF models only via llama-cpp-python.
If RAG_GGUF_PATH is set, that model is used. Otherwise, it defaults to the bundled Gemma 4
artifact. The desktop app instead loads its bundled ADR-0002 profile models through
node-llama-cpp (see the LLM Backend section above).
The GUI/CLI/API flows below are the legacy Python harness. The shipped desktop app exposes ingestion, chat, knowledge packs, and training through its UI; its API is the same frozen contract, served on a per-launch loopback address.
GUI Mode:
- Click "Ingest" button
- Select document folder (folder-based ingestion)
- Wait for processing to complete
Note: GUI supports folder-based batch ingestion. For single-file upload, use API or CLI mode.
CLI Mode:
# Ingest all documents in a directory
python main.py --ingest "C:\Documents\reports"
# Ingest a single file
python main.py --ingest "C:\Documents\report.pdf"API Mode:
import requests
# Ingest entire directory
response = requests.post("http://localhost:8080/ingest", json={
"directory": "C:/Documents/reports"
})
print(response.json())
# Upload and ingest single file
with open("C:/Documents/report.pdf", "rb") as f:
response = requests.post(
"http://localhost:8080/ingest/file",
files={"file": ("report.pdf", f, "application/pdf")}
)
print(response.json())GUI Mode:
- Type your question in the input field
- Press Enter or click "Ask"
- View the answer with source citations
CLI Mode:
# Single question
python main.py --query "What are the main findings?"
# Interactive mode
python main.py --cliAPI Mode:
import requests
response = requests.post("http://localhost:8080/ask", json={
"question": "What are the main findings?",
"n_results": 3
})
print(response.json())Combines BM25 keyword search with vector semantic search using RRF fusion:
- BM25: Fast keyword matching
- Vector: Semantic understanding
- RRF Fusion: Combines both for optimal results
Automatically fetches adjacent chunks around retrieved results:
- Configurable window size (default: 1 chunk)
- Ensures context continuity
- Improves answer quality for multi-part questions
ettin-reranker-32m-v1 (ModernBERT, enabled by default on the desktop backend):
- Ranks retrieved chunks by relevance after initial retrieval
- Higher accuracy than pure hybrid search
- Lightweight β ~40 MB staged total (39,611,408 bytes: q8 ONNX + root tokenizers, per bench/RESULTS.md) β optimized for minimum-spec hardware
- Can be tuned via
TRAININGAPP_RETRIEVAL_RERANK(desktop) / the Settings dialog (legacy harness)
Keyword-based query expansion (disabled by default):
- Extracts key terms from questions to improve retrieval
- Note: The LLM-based step-back transformation is not wired (latency cost too high for minimum-spec hardware)
GUI settings dialog and CLI options below are the legacy Python harness; the desktop app is configured in-app (Settings) plus the
TRAININGAPP_*seams listed under Installation. Harness env vars are documented in CONFIGURATION.md.
LLM Settings:
- GGUF Model Path: Path to
.ggufmodel file
RAG Settings:
- Chunk Size: Number of words per chunk
- Results to Retrieve: Number of chunks for context
- Max Tokens: Maximum response length
- Temperature: Response creativity (0.0-1.0)
Advanced Settings:
- Hybrid Search: Enable/disable BM25+Vector search
- Window Expansion: Number of adjacent chunks to fetch
- Cross-Encoder Reranking: Enable/disable reranking
python main.py [OPTIONS]
Options:
--api Run API server
--cli Run in interactive CLI mode
--ingest PATH Ingest documents from directory
--query QUESTION Ask a question
--db-path PATH Path to vector database (default: ./doc_qa_db)
--model-path PATH Path to GGUF model file (legacy alias for --gguf-path)
--gguf-path PATH GGUF model path
--port PORT API server port (default: 8080)
--chunk-size SIZE Chunk size in words (default: 512)
--chunk-overlap N Chunk overlap in words (default: 50)Mirrors ARCHITECTURE.md (the authoritative map):
+--------------------------- Windows desktop app (desktop/) ---------------------------+
| |
| Electron renderer (web_ui build) Electron main process |
| +--------------------------------+ +--------------------------------------+ |
| | app://index.html | IPC | desktop/main/index.ts | |
| | React pages (web_ui/src/pages) |<------>| lockdown, integrity gate, first-run | |
| | desktopApi bridge (preload) | | wizard, signed update checker | |
| +--------+-----------------------+ +------------------+-------------------+ |
| | HTTP 127.0.0.1:<random port> + X-Desktop-Token (per-launch) |
| v |
| +----------------------------------------------------------------------------------+ |
| | Node backend host (desktop/main/backend) β ADR-0003 | |
| | LlamaEngine (node-llama-cpp, Quality/Fast profiles) - ingest pipeline | |
| | hybrid retrieval (vec0 KNN + FTS5 + RRF k=60 + ettin rerank) - pack manager | |
| | learn assembler - memory governor | |
| +------------------------------------+---------------------------------------------+ |
| v |
| <userData>/profiles/default/store.sqlite |
| better-sqlite3 + sqlite-vec vec0 KNN + FTS5 (contracts/store.schema.sql) |
+---------------------------------------------------------------------------------------+
Plain browser (web_ui/, no Electron): wllama WASM LLM + ONNX embeddings +
IndexedDB/EdgeVec/FlexSearch in-page; knowledge packs in OPFS, courses on an
isolated player origin (ADR-0012).
Renderer (Electron)
- The web_ui React build served from the
app://protocol with a strict CSP - Chat / Documents (knowledge packs) / Training (embedded Storyline player) / Settings pages
Node backend (Electron main process)
- 18 contract routes behind the loopback guard (random port + per-launch
X-Desktop-Token) - GGUF inference via node-llama-cpp with Quality/Fast profiles (ADR-0002)
- Pack lifecycle: install / supersede / rollback / remove (packtool-built packs)
Vector Store
- SQLite + sqlite-vec
vec0KNN and an FTS5 mirror (contracts/store.schema.sql, ADR-0005) - Reciprocal Rank Fusion (RRF, k=60) for hybrid results β same constant on every surface
- Cross-encoder rerank (ettin-reranker-32m-v1) with a calibrated relevance floor
LLM Interface
- GGUF via node-llama-cpp (desktop; GPU when a probe finds one working, CPU otherwise, fully offline)
- GGUF via wllama WASM (browser, CPU/SIMD, fully offline)
RAG Engine
- Query processing and routing
- Hybrid search orchestration
- Context assembly and answer generation with
groundingprovenance - Source citation and Learn-panel deep-link tracking
The pip-based entries below diagnose the legacy Python harness. Desktop-app issues surface through the first-run wizard's named gates and the startup integrity check (failures block backend start and name path/expected/actual β reinstall if the bundled tree fails verification).
Solution 1: GGUF Model Not Found
# Desktop: both profile models are bundled and integrity-checked at startup:
# <resources>/models/llm-quality/gemma-4-e2b-it/model.gguf (Q4_K_M, ADR-0002)
# <resources>/models/llm-fast/lfm2.5-vl-450m/model.gguf (Q4_K_M, ADR-0002)
# A missing model means a broken install β re-run the installer.
# Legacy harness: point RAG_GGUF_PATH at a local GGUF file; custom GGUF models
# can be downloaded from https://huggingface.co/models?search=ggufSolution 2: Wrong Model Path
- Check Settings dialog for correct path
- Use "Browse" button to select model file
pip install chromadb --break-system-packagespip install sentence-transformers# CPU-only build (recommended)
pip install llama-cpp-python
# With CUDA support (if you have NVIDIA GPU)
pip install llama-cpp-python --extra-index-url https://abetlen.github.io/llama-cpp-python/whl/cu121- Nothing downloads: the embedding model (bge-small-en-v1.5 on desktop, snowflake-arctic-embed-m in the browser), the reranker, and both LLM profiles are bundled
- Desktop first run verifies the sha256 integrity manifest over the staged model tree (2,377 ms measured, streaming; bench/RESULTS.md)
- Legacy harness only: BM25 index is built on first ingestion
Solution 1: Reduce chunk size
python main.py --chunk-size 128Solution 2: Increase chunk overlap
python main.py --chunk-size 256 --chunk-overlap 100Solution 3: Reduce number of results
$env:RAG_N_RESULTS=2Check BM25 is enabled:
# In API, check config
from rag_engine import create_engine_from_env
engine = create_engine_from_env()
print(engine.config.hybrid_search) # Should be TrueVerify both backends loaded:
# Check vector store stats
stats = engine.vector_store.get_stats()
print(f"Embedding model: {stats['embedding_model']}")
print(f"BM25 index: {'Ready' if engine.vector_store.bm25_index else 'Not built'}")| Endpoint | Method | Description |
|---|---|---|
/health |
GET | Health check (liveness probe) |
/stats |
GET | Engine statistics |
/ask |
POST | Ask a question (non-streaming) |
/ask/stream |
POST | Ask a question with SSE streaming |
/search |
POST | Search documents |
/ingest |
POST | Ingest directory |
/ingest/file |
POST | Upload and ingest single file |
/ingest/batch |
POST | Batch upload and ingest (up to 20 files) |
/documents |
GET | List documents |
/documents |
DELETE | Clear all documents |
/settings |
GET | Get current RAG settings |
/settings |
PUT | Update RAG settings |
/auth/status |
GET | Authentication status |
/auth/token |
POST | Obtain JWT token |
/telemetry/memory |
GET | Memory telemetry snapshot and downgrade state |
/status/models |
GET | Per-profile inference model presence (first-run gate) |
import requests
import json
# Configure the engine (legacy Python harness)
os.environ["RAG_GGUF_PATH"] = "path/to/gemma-4-e2b-it/model.gguf" # Q4_K_M, ADR-0002
# Start API server in another terminal
# python main.py --api --port 8080
# Ask a question
response = requests.post("http://localhost:8080/ask", json={
"question": "What are the main findings?",
"n_results": 3
})
result = response.json()
print(f"Answer: {result['answer']}")
print(f"Sources: {result['sources']}")
print(f"Grounding: {result['grounding']}") # "grounded" or "general" (C5, issue #72)
print(f"Inference time: {result['inference_time']:.2f}s")import requests
# Ask with streaming response
with requests.post(
"http://localhost:8080/ask/stream",
json={"question": "What are the main findings?", "n_results": 3},
headers={"Authorization": "Bearer <token>"},
stream=True
) as response:
for line in response.iter_lines():
if line.startswith("data: "):
data = json.loads(line[6:])
if "token" in data:
print(data["token"], end="", flush=True)
elif data.get("done"):
print(f"\n\nSources: {data['sources']}")
print(f"Inference time: {data['inference_time']:.2f}s")import requests
# Upload multiple files at once (up to 20)
files = [
("files", ("report1.pdf", open("report1.pdf", "rb"), "application/pdf")),
("files", ("report2.docx", open("report2.docx", "rb"), "application/vnd.openxmlformats-officedocument.wordprocessingml.document")),
("files", ("notes.txt", open("notes.txt", "rb"), "text/plain")),
]
response = requests.post(
"http://localhost:8080/ingest/batch",
files=files,
headers={"Authorization": "Bearer <token>"}
)
result = response.json()
print(f"Total: {result['total_files']}, Succeeded: {result['successful']}, Failed: {result['failed']}")
for r in result["results"]:
status = "β" if r["success"] else "β"
print(f" {status} {r['filename']}: {r.get('error', r.get('chunks_added', 0))} chunks")import requests
# Get current settings
response = requests.get(
"http://localhost:8080/settings",
headers={"Authorization": "Bearer <token>"}
)
settings = response.json()
print(f"Chunk size: {settings['chunk_size']}, Overlap: {settings['chunk_overlap']}")
# Update settings (partial update supported)
response = requests.put(
"http://localhost:8080/settings",
json={"rag_temperature": 0.7, "rag_chunk_size": 768},
headers={"Authorization": "Bearer <token>"}
)
updated = response.json()
print(f"New temperature: {updated['temperature']}, chunk size: {updated['chunk_size']}")The shipped desktop binary is the Electron NSIS installer (see "Building the desktop app from source" above). The PyInstaller/Inno Setup flow below built the retired Python desktop product and is retained for history.
pip install pyinstallerpython build.pyThe executable will be created in dist/DocumentQA.exe.
To create an offline installer:
# Prepare installer files
python scripts/build_installer.py
# Manually download:
# 1. GGUF model to build_installer/models/
# 2. Embedding model to build_installer/embeddings/
# 3. Python embeddable to python_embeddable/
# Run Inno Setup
iscc build_installer/setup.issThis creates an offline installer with all dependencies and models included.
The browser-based interface (web_ui/) is one of the two shipped surfaces β it is also
the desktop app's renderer. This section documents its development flow.
- Vite 6 + React 18 + TypeScript 5
- Pure CSS design token system (no Tailwind)
- vitest + @testing-library/react for testing
The web UI is styled only by the Lumen design tokens in web_ui/src/styles/lumen-tokens.css: colors (--bg-*, --text-*, --accent*, status), --type-* typography, --space-* spacing, --r-* radii and --shadow-1/2/3, each with a dark-theme override. Design rules and contrast guarantees: docs/design/design-language.md.
The earlier --color-*, --spacing-*, --radius-*, --font-family, --font-size-*, --line-height-* and --shadow-sm/md/lg tokens (59 names, frozen in web_ui/src/styles/retired-tokens.ts) were retired in Lumen phase 8 (legacy to Lumen map: web_ui/src/styles/token-remap.ts).
Font: Inter (self-hosted via @fontsource/inter, weights 400/500/600/700) as --font-sans
Dark mode overrides via [data-theme="dark"] attribute on <html>.
cd web_ui
npm install
npm run dev # Development server
npm run build # Production build
npm run typecheck # TypeScript validation
npm test # Run tests with vitestThe web UI includes a typed API client (src/lib/api/) for all backend endpoints:
| File | Description |
|---|---|
client.ts |
ApiClient class with methods for all endpoints |
streaming.ts |
SSEStreamConsumer for POST-based SSE streaming |
auth.ts |
Token storage with Safari private mode fallback |
types.ts |
TypeScript interfaces matching FastAPI models |
index.ts |
Barrel export and default client instance |
The web UI includes browser-side document processing with no server uploads:
import { ExtractorFactory } from './lib/processing/extractor-factory';
import { TextChunker } from './lib/processing/text-chunker';
import { DocumentStore } from './lib/storage/document-store';
// Extract text from uploaded file
const extractor = ExtractorFactory.getExtractor(file);
const extraction = await extractor.extract(file);
// Chunk with semantic boundaries
const chunker = new TextChunker({ chunkSize: 512, overlap: 50 });
const chunks = chunker.chunk(extraction.text, extraction.metadata);
// Store in IndexedDB
const store = new DocumentStore();
await store.saveDocument({
id: crypto.randomUUID(),
name: file.name,
type: file.type,
size: file.size,
chunks,
createdAt: new Date()
});
// List all documents
const docs = await store.loadDocuments();
console.log(`Loaded ${docs.length} documents`);| File | Description |
|---|---|
src/types/chat.ts |
Shared ChatMessage, MessageRole, and ChatState types |
src/lib/streaming/TokenStreamManager.ts |
RAF-batched token delivery, unified callbacks for SSE/WebLLM, cancellation support |
src/lib/inference/InferenceModeContext.tsx |
React context for browser-local/api mode with localStorage persistence |
Usage:
import { apiClient, SSEStreamConsumer, login } from './lib/api';
// Ask a question
const answer = await apiClient.ask("What are the main findings?");
// Stream tokens with SSE
const stream = new SSEStreamConsumer('/ask/stream', { question: "Tell me more" });
stream.onToken(token => appendToAnswer(token));
stream.onDone(data => showSources(data.sources));
stream.start();
// Batch upload
const batch = await apiClient.uploadBatch([file1, file2, file3]);
console.log(`Uploaded ${batch.successful}/${batch.total_files} files`);
// Settings
const settings = await apiClient.getSettings();
await apiClient.updateSettings({ rag_temperature: 0.8 });The ML spike page validates Transformers.js, EdgeVec, and FlexSearch on target hardware.
Test Categories:
- Transformers.js: Hugging Face transformers running in browser (feature-extraction pipeline)
- EdgeVec: HNSW-based vector similarity search (edgevec npm package)
- FlexSearch: Full-text search indexing (flexsearch npm package)
Results show pass/fail/skip status, duration, and memory delta for each library.
The web UI implements a complete browser-side search pipeline:
Query β Embeddings (Transformers.js) β HNSW (EdgeVec) β RRF Fusion β Reranker (optional)
β
Keyword Index (FlexSearch) βββββββββββββββββββββββββββββββ
| Component | File | Description |
|---|---|---|
| Embedding Service | src/lib/embeddings/embedding-service.ts |
Transformers.js pipeline with snowflake-arctic-embed-m-v1.5 ONNX (768-dim, q8) |
| Memory-Aware Selection | src/lib/embeddings/memory-aware.ts |
Device memory detection, tier-based model configuration |
| Vector Index | src/lib/search/vector-index.ts |
EdgeVec HNSW index with IndexedDB persistence |
| Keyword Index | src/lib/search/keyword-index.ts |
FlexSearch with resolution-based scoring |
| RRF Fusion | src/lib/search/rrf-fusion.ts |
Reciprocal Rank Fusion for hybrid results |
| Reranker | src/lib/search/reranker.ts |
Cross-encoder reranker (ettin-reranker-32m-v1, ModernBERT) |
| Types | src/types/embedding.ts |
EmbeddingDocument, EmbeddingResult interfaces |
| Types | src/types/search.ts |
SearchResult, HybridSearchResult interfaces |
Dependencies added: @huggingface/transformers ^3.0.0, edgevec ^0.6.0, flexsearch ^0.8.0
trainingapp/
βββ desktop/ # Electron desktop app (the shipped product)
β βββ main/ # main process: backend host, security/, first-run/,
β β # update-checker.ts, app:// protocol
β βββ preload/ # contextBridge: desktopApi token bridge
β βββ renderer/ # build-time staging of web_ui/dist (gitignored)
β βββ scripts/ # resource stager, sha256 manifest generator,
β β # pack builders, packaged smoke test
β βββ src/__tests__/ # desktop vitest suites (frozen acceptance specs)
βββ web_ui/ # HTML5 web app β browser surface AND desktop renderer
β βββ src/ # React app: pages/, components/, lib/ (api, llm, rag, ...)
β βββ public/models/ # manifest.json + packaged model weights (gitignored)
β βββ scripts/ # prepare-models.mjs, validate-build.mjs
βββ packtool/ # Knowledge Pack build/verify CLI (Node): build/, storyline/, links/
βββ contracts/ # frozen contracts + conformance suite + fixtures
β βββ api.openapi.yaml # the frozen HTTP API (both backends)
β βββ store.schema.sql # SQLite store schema (v3)
β βββ pack.schema.json # Knowledge Pack manifest schema
β βββ pack-feed.schema.json # signed update-feed schema (ADR-0010)
βββ docs/ # adr/ (ADR-0001..0010), security/, pack/training/update
β βββ archive/pre-v3/ # retired pre-v3 planning/audit/release docs (indexed)
βββ eval/ # tier-0 eval harness (questions.jsonl, corpus, runner)
βββ bench/ # measured performance results (RESULTS.md)
βββ scripts/ # repo scripts (CI path classifier, export_seed_chunks.py)
βββ tests/ # Python (legacy harness) test suites
βββ .github/workflows/ # CI: test, conformance, desktop-build, web-ui, ...
βββ api_server.py # legacy Python harness entry (CI conformance surface),
β # with config.py / rag_engine.py / vector_store.py / ...
βββ README.md # This file
- Offline-Only: No data leaves your machine
- No Cloud Services: All processing is local
- Model Bundling: Models are stored locally
- Opt-in Updates Only: zero update-related network calls until you enable the signed channel (ADR-0010)
MIT License - See LICENSE for details.
Bundled model weights carry their own licenses (LLM, embedding, reranker) β see docs/licenses.md for the per-model review.
- Fork the repository
- Create a feature branch
- Make your changes
- Add tests if applicable
- Submit a pull request
Desktop stack (v3):
- node-llama-cpp - GGUF inference (node bindings)
- better-sqlite3 - SQLite store
- sqlite-vec - sqlite vector search extension
- pdfjs-dist - PDF processing (Apache-2.0)
- @huggingface/transformers - In-browser ML models (Apache-2.0)
- edgevec - In-browser vector database (MIT OR Apache-2.0)
- flexsearch - Full-text search (Apache-2.0)
- mammoth - DOCX processing (BSD-2-Clause)
- xlsx - XLSX processing (Apache-2.0)
- jszip - ZIP handling (MIT OR GPL-3.0-or-later)
Legacy Python harness only (CI conformance surface):
- ChromaDB - Vector database
- Sentence Transformers - Embedding models
- llama-cpp-python - GGUF inference
- PyMuPDF - PDF processing
- CustomTkinter - GUI toolkit of the retired desktop product
- @mlc-ai/web-llm - Optional WebGPU browser engine (weights not bundled)
Version: 2.3.0 Last Updated: 2026-09-29 (v3 documentation refresh, issue #89) Hardware: optimized for Intel 11th gen i5 and above (16GB RAM minimum); GPU acceleration is opportunistic and CPU inference is the floor