Skip to content

EricFinland/hearth

Repository files navigation

hearth

build NixOS Ollama Sandboxed Audited License release

Local LLMs and autonomous agents, sandboxed by default. Every run audited. The whole OS reproducible from one flake.

📖 Documentation  ·  🚀 Quickstart  ·  🧠 Architecture


Most people run local agents with full system privileges and no record of what they did. hearth flips that. Agents are contained at the operating-system level, every run records its tokens, cost, latency, and errors to a local database, and the entire system is defined in one flake.nix you can rebuild identically and roll back in a single command.

It is not a custom kernel or a remastered distro. It is a declarative NixOS system you nixos-rebuild switch into existence.

Status: v1.5, stable. The whole stack runs on real hardware: sandboxed agents, the audit log, the reproducible flake, the web cockpit, an OpenAI-compatible API, a local knowledge base, a standing-missions scheduler, and the self-improvement loop (which only ever produces reviewable, gated branches, never auto-changing a running system). Since 1.0, per-run containment (tool and host allowlists, with hosts enforced at the kernel), honeyfile tripwires, a flight recorder with a replay viewer, and a governor (a daily spend breaker plus push alerts) have landed on top. Local model quality is the honest ceiling. See the CHANGELOG.

What makes it different

🛡️ Sandboxed by default Agents run as ephemeral, isolated systemd units. No host secrets, no writes outside their own workspace, no privilege escalation.
🧾 Every run audited Tokens, cost, latency, and errors land in local SQLite. One command shows the last 20 runs. A failed run still leaves a trail.
♻️ Reproducible from boot One flake builds the whole OS. Atomic, bootloader-level rollback. Two builds from the same lock are identical.
🧠 Local and private Ollama on your own GPU, agents that actually use tools, a web command center. Zero cloud, nothing leaves the box.

Architecture

flowchart LR
  Dev["💻 Your laptop<br/>edit · git push"] --> GH["GitHub"]
  GH -->|"nixos-rebuild --flake"| Host

  subgraph Host["🔥 hearth host"]
    direction TB
    LLM["Ollama + CUDA"]
    AG["Sandboxed agents"]
    DB[("SQLite audit")]
    MAP["Web command center"]
    LLM --> AG
    AG --> DB
    AG --> MAP
  end
Loading

See it run

$ hearth-status
● ollama       active (running)   llama3.2:3b, mistral:7b
● tailscale    connected
● recent runs  3 in the last hour

$ hearth-runs
AGENT   MODEL          TOKENS   LATENCY   COST
demo    llama3.2:3b      142     0.9s     $0.00
build   qwen2.5-coder    2.1k    14s      $0.00
chat    mistral:7b       430     3.2s     $0.00

How a run stays contained

sequenceDiagram
  actor You
  participant Agent as Agent (sandboxed)
  participant Model as Local model
  participant Tools as Tools (workspace only)
  participant Audit as Audit log

  You->>Agent: goal
  loop until done
    Agent->>Model: think
    Model-->>Agent: tool call
    Agent->>Tools: run · no escape · secrets by name only
    Tools-->>Agent: result
  end
  Agent->>Audit: tokens · cost · latency
  Agent-->>You: result + receipt
Loading

Quickstart

git clone https://github.com/EricFinland/hearth
cd hearth

nix flake check               # validate the whole system
bash scripts/build-image.sh   # build a bootable image

Full install paths (existing NixOS host, fresh VM, or a Linux primer) live in the docs:

Use it

# Point any OpenAI client at your local models (audited):
curl http://your-hearth:8770/v1/chat/completions \
  -d '{"model":"llama3.2:3b","messages":[{"role":"user","content":"hello"}],"stream":true}'

# Check the install is healthy:
hearth-doctor

# Watch activity in Grafana, etc:
curl http://your-hearth:8770/metrics
The full feature set
  • Declarative NixOS system. The entire OS is one flake; nixos-rebuild switch applies changes atomically.
  • Ollama on boot with a declarative model manifest pulled on activation, CUDA-accelerated.
  • Tool-using agent loop (hearth-loop): a model gets a goal and tools (run commands, read and write files, HTTP), runs in a per-run workspace, and is audited.
  • Least-privilege sandbox with a written threat model: ProtectSystem=strict, ProtectHome, NoNewPrivileges, empty capabilities, a syscall filter, and per-run private temp.
  • Per-run audit log in SQLite, queryable with hearth-runs.
  • Web command center: chat with a local model and launch sandboxed agents from the browser, with permission modes, an approvals queue, and a kill switch.
  • OpenAI-compatible API (/v1/chat/completions with real token streaming, /v1/models): any OpenAI client uses your local models, every call audited.
  • Knowledge base (RAG): ingest docs or a whole repo (index_dir), semantic retrieval via local embeddings with lexical fallback, auto-recalled into agent context.
  • Standing missions: a scheduler that runs missions on a cadence (the works-while-you-sleep layer).
  • Self-improvement loop: an always-on daemon proposes changes to hearth's own config, validates them with nix flake check, compounds and learns, and produces reviewable branches with one-click promote-to-live and an auto-rollback watchdog.
  • OS-level egress enforcement: when a run declares its allowed hosts, hearth programs per-run nftables walls that drop everything else at the kernel, so a run cannot slip the tool-layer allowlist by shelling out to curl.
  • Flight recorder and replay viewer: every run records a structured per-step event stream; a scrubber replays it step by step, with each tool call's args, output, duration, and permission verdict inspectable.
  • Run diff: POST /diff runs the same prompt against two local models and returns tokens, latency, and output side by side, so "which model for this?" is a live test, not a guess.
  • Spend circuit breaker: a hard daily token budget across all runs; at the cap, running agents halt gracefully and new runs refuse to start until local midnight.
  • Unified alerting: one fan-out pushes every alert to Telegram and ntfy, firing on errors, tripwire trips, and budget breaches so the box can reach your phone.
  • Declarative missions: scheduled cron-as-flake missions you define in the flake, rendered read-only and merged with cockpit-created ones, each launch carrying its capability manifest and egress allowlist.
  • Observability: a Prometheus /metrics endpoint, a usage-over-time stats view, and hearth-doctor for a one-command health check.
  • Agent credentials by name: keys are substituted at request time via systemd credentials, so the model never sees the secret value.
  • More agent tools: grep, multi-file edit, directory tree, web-to-knowledge, and more.
  • Optional KDE Plasma desktop, a Tailscale mesh with a tight firewall, secrets via sops-nix, and a boot dashboard.

Contributing & security

Contributions are welcome, see CONTRIBUTING.md for the build and self-test workflow. Found a security issue? Please follow SECURITY.md rather than opening a public issue.

First-boot note: the config ships a default console password for the very first local login (SSH is key-only). Change it immediately with passwd. See SECURITY.md.


Built by Eric Catalano  ·  MIT licensed  ·  Docs  ·  Contribute  ·  Security

About

Local LLMs and autonomous agents on NixOS: sandboxed by default, every run audited, the whole OS reproducible from one flake.

Topics

Resources

License

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Packages

 
 
 

Contributors