A lightweight, production-quality simulation of a Kubernetes-inspired scheduler. It manages nodes, pods, scheduling, failover, load balancing, resource accounting, and health monitoring entirely in memory — no Docker, no Kubernetes, no VMs, no databases.
Written in Go with only the standard library.
- Goroutines give cheap concurrency for the scheduler loop and heartbeat monitor.
- Single static binary; near-zero runtime overhead; trivial memory footprint.
- Strong typing and simple data structures keep per-object cost small.
go build -o bin/ksim ./cmd/sim
./bin/ksim # interactive shell
./bin/ksim <<'EOF' # or scripted via stdin
demo
exit
EOFRun the built-in example simulation (2 nodes, 100 random pods, node failure, recovery):
./bin/ksim --log-level=error
ksim> demo| Requirement | Target | Measured |
|---|---|---|
| Idle memory | < 100 MB | 7.6 MB (100 nodes, 5000 pods) |
| 100 nodes / 5000 pods | few seconds | ~19 ms |
| Avg scheduling latency | low | ~5–100 µs |
- Nodes — register/remove/list/describe, capacity, allocation, health, schedulable flag, heartbeats.
- Pods — create/delete/list with CPU+memory requests, priority, lifecycle
statuses (
Pending,Running,Succeeded,Failed,Evicted). - Scheduler — event-driven, O(n) filter + score (least CPU, least memory, balanced utilization, pod count). Runs only when a pod is added, a node is added/recovered, or a pod is deleted — never polls.
- Failover — mark a node unhealthy (explicitly or via heartbeat timeout),
evict its pods to
Pending, and reschedule automatically with no data loss. - Load balancing — scoring spreads work across nodes by utilization.
- Resource accounting — every bind consumes capacity; scheduling is rejected when resources are insufficient.
- Heartbeats — a single timer-driven goroutine simulates heartbeats and detects timeouts (silence a node to watch it fail).
- Events — ring buffer capped at the latest 1000 events.
- Metrics — cluster CPU/memory usage, per-node utilization, pending/running/ failed pods, average scheduling latency.
- Logging — tiny leveled structured logger (INFO/WARN/ERROR).
.
├── cmd/sim/ # entrypoint: flags, wiring, main loop
├── internal/
│ ├── model/ # Node, Pod, Quantity, statuses, event types
│ ├── registry/ # generic in-memory id-keyed registries
│ ├── scheduler/ # pure O(n) filter + scoring algorithm
│ ├── cluster/ # control plane: state, scheduler loop, failover, metrics
│ ├── event/ # fixed-capacity ring buffer
│ ├── heartbeat/ # heartbeat simulation + timeout detection
│ ├── logging/ # leveled structured logger
│ └── cli/ # interactive command interface
├── docs/ # architecture, algorithm, resource model, workflows
└── scripts/ # runnable example scripts
- Architecture — design, diagram, module responsibilities, concurrency model
- Scheduling algorithm — filtering, scoring, complexity
- Resource model — units, accounting, rejection
- Failover workflow — failure, eviction, rescheduling, recovery
- CLI reference — every command
- Example simulations — demo transcript and manual walkthrough
go build ./... # compile
go vet ./... # static checks
go test ./... # unit tests
go test -race ./... # race detector
go test -bench=. ./internal/cluster/ # scheduling benchmarks- Structs over interfaces; value types for small objects.
- No reflection, no recursion, no global locks, no polling loops.
- Two background goroutines total (scheduler loop + heartbeat monitor); all other work is request-driven from the CLI.
- Single
sync.RWMutexon the cluster state; the scheduler loop is the sole background writer of scheduling state. - Event-driven triggers via a buffered channel; enqueue is never done while holding the cluster mutex, so the system cannot deadlock under load.