Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 18 additions & 7 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,10 +13,12 @@ abstracts away the low-level Firecracker socket API. There's also a low-level

```
knaller/
vm.go High-level API: Run(), List(), VM type
vm.go High-level API: Run(), List(), VM type, AdoptVM(), Kill()
vm_direct.go RunDirect() — per-VM kernel netns mode (Kubernetes-friendly)
config.go Config struct with defaults and validation
network.go Network config derivation + pasta namespace setup script
disk.go Per-VM rootfs copy management + host DNS detection
snapshot.go CreateSnapshot, CreateSnapshotRaw, LoadSnapshot helpers
Containerfile_guest Guest rootfs container definition (Ubuntu + sshd + systemd)
Makefile Build targets: build, test, create-guest
firecracker/
Expand All @@ -41,11 +43,16 @@ knaller/
stdin is not connected. Use SSH to interact with the guest. The guest IP is
printed on start and available via `knaller list`.

- **Rootless networking via pasta.** Each VM runs inside a pasta network namespace
(from the passt project). pasta creates a user+network namespace with a TAP device
and provides L2↔L4 translation to the host — all without root privileges. Inside
the namespace, a second TAP device is created for Firecracker's guest NIC using
`ip tuntap` (works because we have CAP_NET_ADMIN within the namespace).
- **Two networking modes.** `Run` puts the VM in a pasta-managed user+network
namespace (rootless, no privileges). `RunDirect` is **not rootless** — it
puts the VM in a per-VM kernel network namespace and `nsenter`'s firecracker
into it. The supervisor needs CAP_NET_ADMIN + CAP_SYS_ADMIN + CAP_NET_RAW in
the host netns (i.e. root or root-equivalent), and also write access to
`/sys/fs/cgroup` if `EscapeCgroupSlice` is set. The trade-off pays for itself
inside Kubernetes pods where pasta's user namespace breaks `KVM_CREATE_VM`.
RunDirect manages two nft tables — `knaller_box_nat` per-netns and
`knaller_host` shared — for in/out NAT and a default-deny egress filter that
lets guests reach the internet but blocks the host's RFC1918 neighbours.

- **Per-VM rootfs copies.** Each VM gets its own copy of the base rootfs at
`~/.local/share/knaller/vms/<name>/rootfs.ext4`, using `cp --reflink=auto`
Expand All @@ -58,7 +65,11 @@ knaller/

- **One Firecracker process per VM.** Firecracker is not a daemon — each process is
exactly one VM with one API socket. Knaller starts a new Firecracker process for
each `Run()` call and manages its lifecycle.
each `Run()` call and manages its lifecycle. `AdoptVM(name, socketPath, pid)`
re-attaches to a process the current binary did not start, for supervisors that
outlive their VMs (e.g. across container restarts). Pair with
`Config.EscapeCgroupSlice` to move firecracker into a host-level cgroupv2 slice
on launch so it survives the supervisor's container being killed.

- **Cleanup is explicit.** Call `vm.Cleanup()` after `vm.Wait()` returns. This removes
the API socket and rootfs copy. Network namespace cleanup is automatic when the
Expand Down
118 changes: 115 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,13 +7,19 @@ knaller ([/ˈknalɐ/](https://de.wiktionary.org/wiki/Knaller)) — a Go library
- start/stop/pause/resume/snapshot microvms
- set limits on CPU, memory, network bandwidth, disk bandwidth and disk IOPS
- start new microvm from an existing snapshot
- rootless operation (user space networking with passt)
- two networking modes:
- **rootless** (default): user-space networking via [pasta](https://passt.top/) — no privileges required
- **direct**: per-VM kernel network namespace — needed where pasta breaks KVM (e.g. Kubernetes pods that already run in a user namespace)
- adopt running VMs across supervisor restarts (`AdoptVM`) — useful for daemons that outlive their VMs
- raw-disk mode: hand a pre-attached block device to firecracker (NBD, LVM, etc.) and skip the rootfs copy
- raw snapshots: pause/snapshot/resume a VM with a `whilePaused` hook for callers managing the disk lifecycle out of band

## Requirements

- Linux with KVM
- [Firecracker binary](https://github.com/firecracker-microvm/firecracker/releases)
- [pasta](https://passt.top/) for rootless user space networking
- For rootless mode: [pasta](https://passt.top/)
- For direct mode (**not rootless**): `iproute2`, `nftables`, `util-linux` (`nsenter`) on `PATH`. The supervisor needs `CAP_NET_ADMIN` + `CAP_NET_RAW` + `CAP_SYS_ADMIN` in the host network namespace — in practice that means running as root, or as a Kubernetes pod with `hostNetwork: true` plus a privileged security context. `e2fsprogs` is also required if you use `Config.RootFSSize`. If you set `Config.EscapeCgroupSlice`, the supervisor must additionally be able to write under `/sys/fs/cgroup`.
- Podman for building the guest rootfs image

## Install
Expand Down Expand Up @@ -131,4 +137,110 @@ func main() {
// Block until VM exits
vm.Wait()
}
```
```

## Direct networking mode

`knaller.RunDirect` is a drop-in replacement for `knaller.Run` that puts the
VM inside a per-VM **kernel** network namespace instead of pasta's user+network
namespace. Use it where pasta isn't viable — most commonly when your supervisor
already runs in a user namespace (Kubernetes pods, rootless containers), since
KVM's `KVM_CREATE_VM` ioctl returns `EPERM` from inside a user namespace.

What direct mode sets up per VM:

- a kernel netns named `kn-<hash>` containing the firecracker TAP device
- a veth pair (`vh-<hash>` host-side / `vg-<hash>` guest-side) with a /30 in
`172.20.0.0/16`, plumbed on both sides
- an in-netns `knaller_box_nat` nft table that DNATs port 22 (and any
`Config.Ports`) to the guest IP, and SNATs new outbound flows to the
veth-guest IP so siblings on the host see distinct sources
- a host-side `knaller_host` nft table that DNATs `host:<sshPort>` (and
forwarded ports) to the per-VM veth-guest IP, masquerades outbound flows so
the upstream NIC sees the host's IP, and rejects guest→RFC1918 reachability
by default (DNS to `169.254.169.253` is allowed; everything else in
`10/8`, `192.168/16`, `100.64/10`, `224/4`, `169.254/16` and the knaller
veth supernet itself is rejected)
- the firecracker process is `nsenter`'d into the netns, so `/proc/<pid>/cmdline`
shows `firecracker` (not `nsenter`) and discovery/adoption can match on the
command line

Required capabilities and binaries:

- **Direct mode is not rootless.** The supervisor needs `CAP_NET_ADMIN`
(for `ip link`, `nft`, sysctls) plus `CAP_SYS_ADMIN` (to create the
kernel netns and `nsenter` into it) plus `CAP_NET_RAW`, all in the host
network namespace. In Kubernetes that's `hostNetwork: true` plus a
privileged security context (or the explicit capability set); on a bare
host it means running as root or granting the equivalent file
capabilities to the binary.
- `ip` (iproute2), `nft` (nftables), `nsenter` (util-linux) on `PATH`.
- If you set `Config.EscapeCgroupSlice`, the supervisor must also be able
to write into `/sys/fs/cgroup` (i.e. the host cgroupv2 hierarchy must be
mounted RW and visible to the process).

```go
vm, err := knaller.RunDirect(ctx, &knaller.Config{
Name: "myvm",
Kernel: "/path/to/vmlinux",
RootFS: "/path/to/rootfs.ext4",
CPUs: 2,
Memory: 2048,
})
```

### Raw-disk mode

Set `Config.RawDiskPath` to a block device or file the caller manages out of
band (e.g. an NBD device backed by a content-addressed cache, or an LVM
logical volume). knaller does **not** copy, truncate, or resize the device,
and `Cleanup()` leaves it alone — the caller owns the disk lifecycle. This
also disables the per-VM rootfs copy, so VM start time is bounded by the
firecracker handshake instead of by `cp` + `resize2fs`.

```go
vm, err := knaller.RunDirect(ctx, &knaller.Config{
Name: "myvm",
Kernel: "/path/to/vmlinux",
RawDiskPath: "/dev/nbd0",
CPUs: 2,
Memory: 2048,
})
```

### Adopting a running VM

A long-running supervisor that gets restarted (e.g. a Kubernetes DaemonSet)
can re-attach to VMs from its previous lifetime with `AdoptVM`. Persist
`vm.Name`, `vm.SocketPath`, and `vm.PID` somewhere durable; on restart, call:

```go
vm, err := knaller.AdoptVM(name, socketPath, pid)
```

`AdoptVM` verifies the firecracker is alive (`kill -0 pid`) and that its
API socket still answers (`GetInfo` with a 2 s timeout). The returned `*VM`
has `cmd == nil`; `Wait` switches to polling `/proc/<pid>` and `Kill` falls
back to `syscall.Kill(pid, SIGKILL)`. Pair this with
`Config.EscapeCgroupSlice` to move the firecracker process into a host-level
cgroupv2 slice on launch so it survives the supervisor's container being
restarted.

### Raw snapshots

`CreateSnapshotRaw` is like `CreateSnapshot` but skips the rootfs copy and
the drive-path patching, so it composes with `RawDiskPath`. It returns
timing for the paused-window so callers can attribute pause-tail latency,
and accepts a `whilePaused` callback that runs after the firecracker
state+memory dump is written but before the VM is resumed — useful for
flushing a dirty queue or copying an external manifest into `snapDir`.

```go
res, err := knaller.CreateSnapshotRaw(ctx, "myvm", os.Stderr, func(snapDir string) error {
return copyManifestInto(snapDir)
})
```

On restore (`RunDirect` with `SnapshotID` + `RawDiskPath`), `LoadSnapshot`
is followed by `PatchDrive` so the new `RawDiskPath` replaces whatever
device path was baked into the state file.
49 changes: 43 additions & 6 deletions config.go
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,40 @@ type Config struct {
PastaBin string // path to pasta binary (default: "pasta")
Stdout io.Writer // serial console log output (default: io.Discard)
Stderr io.Writer // firecracker process stderr (default: io.Discard)

// RootFSSize, when > 0 and larger than the source rootfs, expands the
// per-VM rootfs to this byte size after the cp+reflink (truncate to
// size, then e2fsck -fp + resize2fs). Sparse — actual host disk
// consumption is what the guest writes. Requires e2fsprogs on PATH.
// Ignored when RawDiskPath is set (caller owns the device).
RootFSSize int64

// RawDiskPath, when non-empty, bypasses the per-VM rootfs copy and
// hands the named block device or file to firecracker as the rootfs
// drive. Use this when the caller manages the disk lifecycle out of
// band (e.g. an NBD device backed by a content-addressed cache, or a
// raw image on shared storage). Knaller does not touch its contents
// — no copy, no truncate, no resize — and Cleanup() leaves it alone.
// Snapshot restore with RawDiskPath set will PatchDrive the
// post-LoadSnapshot drive path to point here.
RawDiskPath string

// Netns, when non-empty, overrides the per-VM kernel network namespace
// name that RunDirect would otherwise derive from cfg.Name. Adoption
// uses this to pin the netns identity captured by an external state
// store, so the new process attaches to the same namespace the
// previous lifetime created.
Netns string

// EscapeCgroupSlice, when non-empty, names a cgroupv2 slice (e.g.
// "knaller-vms.slice") that the firecracker process is moved into
// immediately after spawn. Used in container-managed environments
// (Kubernetes, systemd-nspawn) so the VM survives a restart of the
// supervising container's own cgroup. Requires the host's cgroupv2
// hierarchy to be visible at /sys/fs/cgroup. The slice is created
// on demand. Errors are non-fatal — the VM still starts; it just
// shares the parent process's lifetime.
EscapeCgroupSlice string
}

// setDefaults fills in zero-value fields with sensible defaults.
Expand Down Expand Up @@ -64,7 +98,8 @@ func (c *Config) setDefaults() {

// validate checks that all required fields are set and valid. When restoring
// from a snapshot (SnapshotID is set), kernel/rootfs/cpus/memory come from the
// snapshot and are not validated here.
// snapshot and are not validated here. When the caller manages the rootfs out
// of band (RawDiskPath is set), the RootFS path is similarly skipped.
func (c *Config) validate() error {
if c.SnapshotID != "" {
return nil
Expand All @@ -75,11 +110,13 @@ func (c *Config) validate() error {
if _, err := os.Stat(c.Kernel); err != nil {
return fmt.Errorf("kernel: %w", err)
}
if c.RootFS == "" {
return errors.New("rootfs path is required")
}
if _, err := os.Stat(c.RootFS); err != nil {
return fmt.Errorf("rootfs: %w", err)
if c.RawDiskPath == "" {
if c.RootFS == "" {
return errors.New("rootfs path is required")
}
if _, err := os.Stat(c.RootFS); err != nil {
return fmt.Errorf("rootfs: %w", err)
}
}
if c.CPUs <= 0 {
return errors.New("cpus must be > 0")
Expand Down
32 changes: 32 additions & 0 deletions config_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -140,6 +140,38 @@ func TestConfigValidateSnapshotSkipsKernelRootfs(t *testing.T) {
}
}

func TestConfigValidateRawDiskPathSkipsRootFS(t *testing.T) {
dir := t.TempDir()
kernel := filepath.Join(dir, "vmlinux")
os.WriteFile(kernel, []byte("fake"), 0o644)

// RawDiskPath is set; RootFS is intentionally empty + nonexistent.
cfg := &Config{Kernel: kernel, RawDiskPath: "/dev/nbd0"}
cfg.setDefaults()
if err := cfg.validate(); err != nil {
t.Fatalf("expected no error with RawDiskPath set, got: %v", err)
}
}

func TestConfigDefaultsNewFieldsZero(t *testing.T) {
// New direct-mode fields default to their zero value — knaller does not
// turn on cgroup escape, raw-disk, or netns override unless asked.
cfg := &Config{}
cfg.setDefaults()
if cfg.RootFSSize != 0 {
t.Errorf("RootFSSize = %d, want 0", cfg.RootFSSize)
}
if cfg.RawDiskPath != "" {
t.Errorf("RawDiskPath = %q, want empty", cfg.RawDiskPath)
}
if cfg.Netns != "" {
t.Errorf("Netns = %q, want empty", cfg.Netns)
}
if cfg.EscapeCgroupSlice != "" {
t.Errorf("EscapeCgroupSlice = %q, want empty", cfg.EscapeCgroupSlice)
}
}

func TestRandomName(t *testing.T) {
name1 := randomName()
name2 := randomName()
Expand Down
28 changes: 27 additions & 1 deletion disk.go
Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,11 @@ func vmDataDir(name string) string {
// has its own writable filesystem. Uses cp --reflink=auto to get copy-on-write
// behavior on filesystems that support it (btrfs, xfs), which makes the copy
// nearly instant and only uses disk space for blocks that the VM actually changes.
func prepareDisk(name, baseRootFS string) (string, error) {
//
// If rootFSSize > 0 and larger than the source image, the copy is grown
// (truncate, then e2fsck -fp + resize2fs) so the guest sees the expanded
// filesystem. Requires e2fsprogs (e2fsck, resize2fs) on PATH.
func prepareDisk(name, baseRootFS string, rootFSSize int64) (string, error) {
dir := vmDataDir(name)
if err := os.MkdirAll(dir, 0o755); err != nil {
return "", fmt.Errorf("create vm dir: %w", err)
Expand All @@ -32,6 +36,28 @@ func prepareDisk(name, baseRootFS string) (string, error) {
if out, err := cmd.CombinedOutput(); err != nil {
return "", fmt.Errorf("copy rootfs: %s: %w", out, err)
}
if rootFSSize > 0 {
st, err := os.Stat(dst)
if err != nil {
return "", fmt.Errorf("stat rootfs: %w", err)
}
if st.Size() < rootFSSize {
if err := os.Truncate(dst, rootFSSize); err != nil {
return "", fmt.Errorf("truncate rootfs to %d: %w", rootFSSize, err)
}
// e2fsck returns 1 for "errors fixed", which is fine on a
// fresh copy. Only escalate on >=4 (unfixed errors), per
// e2fsck(8) exit codes.
if out, err := exec.Command("e2fsck", "-fp", dst).CombinedOutput(); err != nil {
if exitErr, ok := err.(*exec.ExitError); !ok || exitErr.ExitCode() >= 4 {
return "", fmt.Errorf("e2fsck rootfs: %s: %w", out, err)
}
}
if out, err := exec.Command("resize2fs", dst).CombinedOutput(); err != nil {
return "", fmt.Errorf("resize2fs rootfs: %s: %w", out, err)
}
}
}
return dst, nil
}

Expand Down
32 changes: 30 additions & 2 deletions disk_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ func TestPrepareDisk(t *testing.T) {
defer os.Setenv("HOME", origHome)

name := "test-vm"
diskPath, err := prepareDisk(name, baseRootFS)
diskPath, err := prepareDisk(name, baseRootFS, 0)
if err != nil {
t.Fatal(err)
}
Expand All @@ -42,6 +42,34 @@ func TestPrepareDisk(t *testing.T) {
}
}

func TestPrepareDiskRootFSSizeNoGrowthBelowCurrent(t *testing.T) {
// When rootFSSize <= current size, prepareDisk must not invoke
// e2fsck/resize2fs (which would fail on our non-ext4 fake content).
dir := t.TempDir()
baseRootFS := filepath.Join(dir, "rootfs.ext4")
content := []byte("rootfs content larger than the requested grow target")
if err := os.WriteFile(baseRootFS, content, 0o644); err != nil {
t.Fatal(err)
}

origHome := os.Getenv("HOME")
os.Setenv("HOME", dir)
defer os.Setenv("HOME", origHome)

// rootFSSize is smaller than current → growth path must be skipped.
diskPath, err := prepareDisk("vm-no-grow", baseRootFS, 4)
if err != nil {
t.Fatalf("prepareDisk with sub-current rootFSSize: %v", err)
}
st, err := os.Stat(diskPath)
if err != nil {
t.Fatal(err)
}
if st.Size() != int64(len(content)) {
t.Errorf("size = %d, want %d (no growth expected)", st.Size(), len(content))
}
}

func TestRemoveDisk(t *testing.T) {
dir := t.TempDir()
baseRootFS := filepath.Join(dir, "rootfs.ext4")
Expand All @@ -52,7 +80,7 @@ func TestRemoveDisk(t *testing.T) {
defer os.Setenv("HOME", origHome)

name := "test-rm-vm"
diskPath, err := prepareDisk(name, baseRootFS)
diskPath, err := prepareDisk(name, baseRootFS, 0)
if err != nil {
t.Fatal(err)
}
Expand Down
Loading
Loading