diff --git a/infra/docs/04-arc-and-caching.md b/infra/docs/04-arc-and-caching.md index 504ebff46..6c20776e0 100644 --- a/infra/docs/04-arc-and-caching.md +++ b/infra/docs/04-arc-and-caching.md @@ -201,12 +201,15 @@ template: mountPath: /opt/bazel-repo-cache subPath: bazel-repo-cache resources: + # Burstable on purpose: requests bin-pack 8 runners onto the + # 32-vCPU / 125 GiB host; limits are each Kata VM's hotplug + # ceiling. Keep sum(memory limits) under host RAM. requests: - cpu: "8" - memory: "24Gi" + cpu: "3" + memory: "10Gi" limits: cpu: "8" - memory: "24Gi" + memory: "14Gi" volumes: - name: runner-cache persistentVolumeClaim: @@ -259,14 +262,17 @@ Field by field: mounts to the `arc-runners/runner-cache` PVC. `ReadWriteOnce` is enough on this single-node k3s host; use a RWX-capable storage class before spreading runners across nodes. -- **`resources`** - requests `8` CPU / `24Gi`, limits `8` CPU / `24Gi`. Kata reads - these and sizes the guest accordingly: the VM now boots at the same - guaranteed floor (`default_vcpus: 2`, `default_memory: 4096`) and only - hotplugs beyond that toward the limits, with `default_maxvcpus: 0` allowing up - to all host CPUs. Effectively the **requests are the boot-time VM size** and - the **limits are the hotplug ceiling**. See [02-kata-runtime.md](02-kata-runtime.md) - for the runtime knobs and [`infra/tune-kata-runtime.sh`](../tune-kata-runtime.sh) - for the SSH-driven patch helper. +- **`resources`** - requests `3` CPU / `10Gi`, limits `8` CPU / `14Gi` + (burstable; see the `maxRunners` bullet above). Kata sizes the guest from + these: every VM boots at the fixed floor from the runtime config + (`default_vcpus: 2`, `default_memory: 4096` — deliberately at or below the + pod request so boot stays cheap) and hotplugs beyond it toward the pod + **limits**, with `default_maxvcpus: 0` allowing up to all host CPUs. + Effectively the **boot shape is a fixed floor**, the **requests are the + scheduler's bin-packing unit**, and the **limits are the hotplug ceiling**. + See [02-kata-runtime.md](02-kata-runtime.md) for the runtime knobs and + [`infra/tune-kata-runtime.sh`](../tune-kata-runtime.sh) for the SSH-driven + patch helper. --- diff --git a/infra/tune-kata-runtime.sh b/infra/tune-kata-runtime.sh index 2766d7c24..9eda821e2 100755 --- a/infra/tune-kata-runtime.sh +++ b/infra/tune-kata-runtime.sh @@ -1,6 +1,11 @@ #!/usr/bin/env bash -# Patch the live Kata QEMU config to match the runner pod's guaranteed boot -# shape, raise guest and host open-file limits, and enlarge the virtiofsd worker pool. +# Patch the live Kata QEMU config: set the guest BOOT floor (default_vcpus / +# default_memory), raise guest and host open-file limits, and enlarge the +# virtiofsd worker pool. Runner pods are burstable (see reload-runner.sh): +# the boot floor stays deliberately at or below the pod's CPU/memory REQUEST +# so VMs boot small and fast, and Kata hotplugs each guest toward the pod +# LIMIT under load. Keep BOOT_VCPUS/BOOT_MEMORY_MIB <= the request when +# retuning either side. # The script runs over SSH, which keeps the desired values version-controlled. # It then smoke-tests a new kata-qemu pod. #