Skip to content

Latest commit

 

History

History
165 lines (133 loc) · 4.99 KB

File metadata and controls

165 lines (133 loc) · 4.99 KB

Example simulations

1. Built-in demo: 100 pods, node failure, recovery

The demo command performs the full scenario described in the requirements: two nodes, 100 pods with random resource requests, scheduling decisions, load balancing, node failure, automatic rescheduling, and the final cluster state.

./bin/ksim --log-level=error
ksim> demo

Observed behavior (seed 42):

=== Example simulation (seed 42) ===
Adding node node-a (4 CPU, 8GiB) and node-b (8 CPU, 16GiB) ...
Creating 100 pods with random resource requests ...

--- Scheduling decisions (first 25) ---
ID       NAME     CPU(m)  MEM(MiB)  STATUS   NODE
pod-1    app-001  85      107       Running  node-a
pod-10   app-010  79      75        Running  node-b
pod-11   app-011  53      234       Running  node-b
... (all 100 pods Running) ...

--- Cluster state after scheduling ---
Nodes:          2 (healthy: 2)
CPU:            8.22 / 12 cores (68.5%)
Memory:         12669Mi / 24576Mi (51.6%)
Pods:           pending=0 running=100
Node utilization:
ID      CPU    MEM    PODS  HEALTH
node-a  68.2%  54.2%  34    healthy
node-b  68.7%  50.2%  66    healthy

Load balancing: node-a and node-b converge to the same CPU utilization (68.2% vs 68.7%) even though node-b has twice the capacity and therefore twice the pods (34 vs 66). This is the utilization-based score equalizing the two.

=== Simulating failure of node-a ===
--- Cluster state after failover ---
Nodes:          2 (healthy: 1)
Pods:           pending=2 running=98
ID      CPU     MEM    PODS  HEALTH
node-a  0.0%    0.0%   0     UNHEALTHY
node-b  100.0%  76.4%  98    healthy

--- Recent events ---
SEQ  TIME      KIND            OBJECT  MESSAGE
...  NodeFailed    node-a  node failed
...  PodRescheduled  pod-52  scheduled on node-b
...  PodRescheduled  pod-54  scheduled on node-b
...  PodRescheduled  pod-91  scheduled on node-b

All 34 pods from node-a were evicted and rescheduled onto node-b; the 2 that could not fit (node-b hit 100% CPU) remain Pending. Resources on node-a are fully released (0.0%). PodRescheduled events record every move.

=== Recovering node-a ===
Nodes:          2 (healthy: 2)
Pods:           pending=0 running=100
ID      CPU     MEM    PODS  HEALTH
node-a  5.6%    1.9%   2     healthy
node-b  100.0%  76.4%  98    healthy

After recovery the two pending pods are scheduled onto the empty node-a. Pods already running on node-b are not migrated back (no unnecessary migrations).

2. Manual walkthrough

./bin/ksim --log-level=error
ksim> add-node node-a node-a-host 4 8192
ksim> add-node node-b node-b-host 8 16384
ksim> create-pod web-1 250 512
ksim> create-pod web-2 500 1024
ksim> create-pod worker-1 750 2048

ksim> list-pods
ID     NAME      CPU(m)  MEM(MiB)  PRIORITY  STATUS   NODE
pod-1  web-1     250     512       0         Running  node-a
pod-2  web-2     500     1024      0         Running  node-b
pod-3  worker-1  750     2048      0         Running  node-a

ksim> cluster-status
Nodes:          2 (healthy: 2)
CPU:            1.50 / 12 cores (12.5%)
Memory:         3584Mi / 24576Mi (14.6%)
Pods:           pending=0 running=3
Node utilization:
ID      CPU    MEM    PODS  HEALTH
node-a  25.0%  31.2%  2     healthy
node-b  6.2%   6.2%   1     healthy

ksim> simulate-failure node-a
node node-a marked unhealthy; hosted pods evicted and rescheduled

ksim> list-pods
ID     NAME      CPU(m)  MEM(MiB)  PRIORITY  STATUS   NODE
pod-1  web-1     250     512       0         Running  node-b
pod-2  web-2     500     1024      0         Running  node-b
pod-3  worker-1  750     2048      0         Running  node-b

ksim> recover-node node-a
node node-a recovered

3. Heartbeat timeout (silence → fail → recover)

./bin/ksim --log-level=warn --heartbeat-interval=1s --heartbeat-timeout=3s
ksim> add-node node-a host-a 4 8192
ksim> add-node node-b host-b 4 8192
ksim> create-pod p1 500 1024
ksim> create-pod p2 500 1024
ksim> silence-node node-a
node node-a silenced; it will stop sending heartbeats
# ...about 3 seconds later the monitor acts...
WARN node heartbeat timeout, marking unhealthy node=node-a
WARN node failed, evicting pods node=node-a evicted=1

ksim> cluster-status
Nodes:          2 (healthy: 1)
Pods:           pending=0 running=2
ID      CPU    MEM    PODS  HEALTH
node-a  0.0%   0.0%   0     UNHEALTHY
node-b  25.0%  25.0%  2     healthy

ksim> recover-node node-a
node node-a recovered

4. Scale test

Create 100 nodes and 5000 pods with a piped script and observe latency, throughput, and memory:

# generate the script, then
/usr/bin/time -v ./bin/ksim --log-level=warn < scale.txt > /dev/null

Measured on modest hardware: ~19 ms wall time, ~7.6 MB peak RSS, all 5000 pods scheduled. The cluster-status summary in that run reports an average scheduling latency of ~5 µs.

5. Reproducibility

The demo uses a fixed seed (42 by default; override with -seed). The scheduler is deterministic: ties break by node registration order, so the same script always produces the same placements.