docs(infra): reconciled runner sizing docs with burstable scale set

- Updated the arc-omp-values example block and resources bullet to the live
  3cpu/10Gi request + 8cpu/14Gi limit shape instead of the old guaranteed
  8cpu/24Gi sizing.
- Documented the Kata boot floor vs request vs hotplug-limit relationship in
  tune-kata-runtime.sh so boot defaults are kept at or below pod requests.
This commit is contained in:
can1357
2026-07-30 08:16:19 +02:00
parent 38846a7e96
commit 17ea2d65fa
2 changed files with 24 additions and 13 deletions
+17 -11
View File
@@ -201,12 +201,15 @@ template:
mountPath: /opt/bazel-repo-cache mountPath: /opt/bazel-repo-cache
subPath: bazel-repo-cache subPath: bazel-repo-cache
resources: resources:
# Burstable on purpose: requests bin-pack 8 runners onto the
# 32-vCPU / 125 GiB host; limits are each Kata VM's hotplug
# ceiling. Keep sum(memory limits) under host RAM.
requests: requests:
cpu: "8" cpu: "3"
memory: "24Gi" memory: "10Gi"
limits: limits:
cpu: "8" cpu: "8"
memory: "24Gi" memory: "14Gi"
volumes: volumes:
- name: runner-cache - name: runner-cache
persistentVolumeClaim: persistentVolumeClaim:
@@ -259,14 +262,17 @@ Field by field:
mounts to the `arc-runners/runner-cache` PVC. `ReadWriteOnce` is enough on this mounts to the `arc-runners/runner-cache` PVC. `ReadWriteOnce` is enough on this
single-node k3s host; use a RWX-capable storage class before spreading runners single-node k3s host; use a RWX-capable storage class before spreading runners
across nodes. across nodes.
- **`resources`** - requests `8` CPU / `24Gi`, limits `8` CPU / `24Gi`. Kata reads - **`resources`** - requests `3` CPU / `10Gi`, limits `8` CPU / `14Gi`
these and sizes the guest accordingly: the VM now boots at the same (burstable; see the `maxRunners` bullet above). Kata sizes the guest from
guaranteed floor (`default_vcpus: 2`, `default_memory: 4096`) and only these: every VM boots at the fixed floor from the runtime config
hotplugs beyond that toward the limits, with `default_maxvcpus: 0` allowing up (`default_vcpus: 2`, `default_memory: 4096` — deliberately at or below the
to all host CPUs. Effectively the **requests are the boot-time VM size** and pod request so boot stays cheap) and hotplugs beyond it toward the pod
the **limits are the hotplug ceiling**. See [02-kata-runtime.md](02-kata-runtime.md) **limits**, with `default_maxvcpus: 0` allowing up to all host CPUs.
for the runtime knobs and [`infra/tune-kata-runtime.sh`](../tune-kata-runtime.sh) Effectively the **boot shape is a fixed floor**, the **requests are the
for the SSH-driven patch helper. scheduler's bin-packing unit**, and the **limits are the hotplug ceiling**.
See [02-kata-runtime.md](02-kata-runtime.md) for the runtime knobs and
[`infra/tune-kata-runtime.sh`](../tune-kata-runtime.sh) for the SSH-driven
patch helper.
--- ---
+7 -2
View File
@@ -1,6 +1,11 @@
#!/usr/bin/env bash #!/usr/bin/env bash
# Patch the live Kata QEMU config to match the runner pod's guaranteed boot # Patch the live Kata QEMU config: set the guest BOOT floor (default_vcpus /
# shape, raise guest and host open-file limits, and enlarge the virtiofsd worker pool. # default_memory), raise guest and host open-file limits, and enlarge the
# virtiofsd worker pool. Runner pods are burstable (see reload-runner.sh):
# the boot floor stays deliberately at or below the pod's CPU/memory REQUEST
# so VMs boot small and fast, and Kata hotplugs each guest toward the pod
# LIMIT under load. Keep BOOT_VCPUS/BOOT_MEMORY_MIB <= the request when
# retuning either side.
# The script runs over SSH, which keeps the desired values version-controlled. # The script runs over SSH, which keeps the desired values version-controlled.
# It then smoke-tests a new kata-qemu pod. # It then smoke-tests a new kata-qemu pod.
# #