feat(infra): added kata-based ARC CI docs and reload tooling for k3s
- Added k3s host and cluster setup docs, including pinned install, module checks, and readiness commands. - Added kata-runtime setup docs with containerd integration, RuntimeClass, and guest-kernel validation steps. - Added ARC ARC/runner-and-cache deployment docs covering helm values, secrets, and pod network policy. - Added reload-runner automation and a preloaded runner image Dockerfile with pinned toolchain and rollout checks.
This commit is contained in:
@@ -0,0 +1,448 @@
|
||||
# 01 — Host preparation & single-node k3s cluster
|
||||
|
||||
This guide takes a fresh Linux host from bare OS to a **working single-node [k3s](https://k3s.io) cluster** that is ready to run Kata-isolated CI runners. It covers host prerequisites (hardware virtualization, kernel modules, time sync, the nginx port constraint), the exact k3s install, kubeconfig setup, cluster networking (Flannel CNI, CIDRs, CoreDNS), and pod-to-internet egress via host firewalld NAT.
|
||||
|
||||
Read [README.md](README.md) first for the overall architecture and the placeholder/redaction table. When this guide is done, continue with **[02-kata-runtime.md](02-kata-runtime.md)** to install the Kata Containers runtime and register the `kata-qemu` RuntimeClass.
|
||||
|
||||
All configs below are copied from the live reference host and then redacted. Substitute the placeholders from the [README redaction table](README.md#redaction--placeholders) (notably `<CI_HOST>`, `<PUBLIC_IP>`, `<EXT_IFACE>`, `<TAILNET_IP>`) with your own values. CIDRs (`10.42.0.0/16`, `10.43.0.0/16`), the CoreDNS IP (`10.43.0.10`), and version numbers are kept as-is.
|
||||
|
||||
The reference host is a bare-metal **CentOS Stream 10** box, 32 vCPU / 125 GiB RAM, AMD CPU, running **k3s `v1.35.5+k3s1`** (bundled containerd `v2.2.3-k3s1`). Commands are shown for RHEL-family (`dnf` / `firewalld`); adapt package and firewall commands for your distro.
|
||||
|
||||
---
|
||||
|
||||
## 1. Host prerequisites
|
||||
|
||||
### 1.1 Hardware virtualization / KVM
|
||||
|
||||
Every CI job boots its own QEMU/KVM microVM, so the host **must** expose working KVM. On bare metal this means VT-x (Intel) or AMD-V (AMD) enabled in firmware; on a VM you need working *nested* virtualization.
|
||||
|
||||
Check the CPU virtualization flag (`vmx` = Intel, `svm` = AMD) and that the KVM device and modules are present:
|
||||
|
||||
```bash
|
||||
# CPU supports virtualization? (non-zero count = yes)
|
||||
grep -E -c '(vmx|svm)' /proc/cpuinfo
|
||||
|
||||
# which flavor
|
||||
grep -E -om1 '(vmx|svm)' /proc/cpuinfo # reference host prints: svm (AMD)
|
||||
lscpu | grep -i virtualization
|
||||
|
||||
# /dev/kvm must exist and be accessible
|
||||
ls -l /dev/kvm # crw-rw-rw-. 1 root kvm 10, 232 ... /dev/kvm
|
||||
|
||||
# KVM kernel modules loaded
|
||||
lsmod | grep -E '^kvm' # kvm_amd ... kvm (or kvm_intel on Intel)
|
||||
```
|
||||
|
||||
On the reference host this yields:
|
||||
|
||||
```
|
||||
$ ls -l /dev/kvm
|
||||
crw-rw-rw-. 1 root kvm 10, 232 /dev/kvm
|
||||
|
||||
$ lsmod | grep -E '^kvm'
|
||||
kvm_amd 237568 99
|
||||
kvm 1470464 78 kvm_amd
|
||||
```
|
||||
|
||||
The module loads automatically when the CPU flag is present; if `/dev/kvm` is missing, load it explicitly and persist it:
|
||||
|
||||
```bash
|
||||
modprobe kvm_amd # or: modprobe kvm_intel
|
||||
echo kvm_amd > /etc/modules-load.d/kvm.conf
|
||||
```
|
||||
|
||||
On Debian/Ubuntu you can instead run `kvm-ok` (from the `cpu-checker` package); on RHEL-family the checks above are the equivalent.
|
||||
|
||||
> Kata also needs the `vhost_vsock` and `vhost_net` modules for its agent vsock channel and VM networking. Those are part of the Kata runtime setup and are covered in [02-kata-runtime.md](02-kata-runtime.md); the KVM availability above is the only virtualization prerequisite for this guide.
|
||||
|
||||
### 1.2 Kernel modules and sysctls for k3s networking
|
||||
|
||||
k3s needs the `br_netfilter` and `overlay` modules and a couple of sysctls so that bridged pod traffic is seen by iptables and so the host can route/NAT pod traffic. The k3s systemd unit loads the modules on start (`ExecStartPre=-/sbin/modprobe br_netfilter` / `overlay`) and the installer sets the sysctls, but set them explicitly so they survive reboots and are correct before install:
|
||||
|
||||
```bash
|
||||
cat >/etc/modules-load.d/k3s.conf <<'EOF'
|
||||
br_netfilter
|
||||
overlay
|
||||
EOF
|
||||
modprobe br_netfilter overlay
|
||||
|
||||
cat >/etc/sysctl.d/90-k3s.conf <<'EOF'
|
||||
net.ipv4.ip_forward = 1
|
||||
net.bridge.bridge-nf-call-iptables = 1
|
||||
EOF
|
||||
sysctl --system
|
||||
```
|
||||
|
||||
Verify (these are the live values on the reference host):
|
||||
|
||||
```
|
||||
$ sysctl net.ipv4.ip_forward net.bridge.bridge-nf-call-iptables
|
||||
net.ipv4.ip_forward = 1
|
||||
net.bridge.bridge-nf-call-iptables = 1
|
||||
|
||||
$ lsmod | grep -E 'br_netfilter|overlay'
|
||||
br_netfilter 36864 0
|
||||
bridge 409600 1 br_netfilter
|
||||
overlay 229376 49
|
||||
```
|
||||
|
||||
`net.ipv4.ip_forward = 1` is what lets the host route (and NAT — see [section 5](#5-cluster-networking-cni-cidrs--nat-egress)) pod traffic out to the internet.
|
||||
|
||||
### 1.3 SELinux
|
||||
|
||||
The reference host runs SELinux in **Permissive** mode:
|
||||
|
||||
```
|
||||
$ getenforce
|
||||
Permissive
|
||||
```
|
||||
|
||||
The k3s installer installs an SELinux policy (`k3s-selinux`) when SELinux is Enforcing on RHEL-family hosts, so Enforcing also works; Permissive is used here to keep the Kata/QEMU + virtio-fs path unencumbered during bring-up. Pick one consistently — if you run Enforcing, make sure `container-selinux` / `k3s-selinux` are installed (the k3s installer pulls them).
|
||||
|
||||
### 1.4 Base packages
|
||||
|
||||
The k3s install script needs only `curl`; everything else (its own containerd, CNI, kubectl) is bundled. Make sure the host has current packages and the basics:
|
||||
|
||||
```bash
|
||||
dnf -y update
|
||||
dnf -y install curl tar iptables
|
||||
```
|
||||
|
||||
> Do **not** pre-install a separate containerd/Docker for k3s to use — k3s ships and manages its own containerd v2. (A separate Docker install can coexist for *building* the runner image; that is covered in [03-runner-image.md](03-runner-image.md).)
|
||||
|
||||
### 1.5 Time synchronization
|
||||
|
||||
Clock skew breaks TLS to the Kubernetes API and to GitHub. Keep an NTP client running. The reference host uses `chrony`:
|
||||
|
||||
```bash
|
||||
dnf -y install chrony
|
||||
systemctl enable --now chronyd
|
||||
timedatectl # "System clock synchronized: yes", "NTP service: active"
|
||||
```
|
||||
|
||||
### 1.6 The nginx port constraint (why Traefik and servicelb are disabled)
|
||||
|
||||
The reference host **also runs nginx**, which owns ports 80 and 443:
|
||||
|
||||
```
|
||||
$ systemctl is-active nginx
|
||||
active
|
||||
|
||||
$ ss -tlnp | grep -E ':80 |:443 '
|
||||
LISTEN 0 511 0.0.0.0:80 0.0.0.0:* users:(("nginx",...))
|
||||
LISTEN 0 511 0.0.0.0:443 0.0.0.0:* users:(("nginx",...))
|
||||
```
|
||||
|
||||
A default k3s install would deploy **Traefik** (an ingress controller that wants :80/:443) and **servicelb** (the Klipper load-balancer, which binds `LoadBalancer` service ports directly on the host). Both would collide with nginx. We therefore disable both at install time (next section). k3s' own API server listens on **:6443**, which does not conflict with nginx, so the cluster is fully functional without those two add-ons.
|
||||
|
||||
> Note on swap: the reference host has swap enabled (`/dev/md1`, 16 GiB) and k3s runs fine with it. If you prefer the upstream-Kubernetes convention of swap-off, disabling it is also supported — it is not required here.
|
||||
|
||||
---
|
||||
|
||||
## 2. Install k3s
|
||||
|
||||
Install k3s as a single-node server, pinning the version and disabling Traefik and servicelb. This is the exact configuration baked into the reference host's systemd unit:
|
||||
|
||||
```bash
|
||||
curl -sfL https://get.k3s.io | \
|
||||
INSTALL_K3S_VERSION=v1.35.5+k3s1 \
|
||||
INSTALL_K3S_EXEC="server --disable=traefik --disable=servicelb" \
|
||||
sh -
|
||||
```
|
||||
|
||||
What each piece does:
|
||||
|
||||
| Token | Meaning |
|
||||
| --- | --- |
|
||||
| `INSTALL_K3S_VERSION=v1.35.5+k3s1` | Pin the exact k3s release (reproducible installs; omit to track the stable channel). |
|
||||
| `server` | Run this node as a **control-plane + worker** (single-node cluster — it both schedules and runs pods). |
|
||||
| `--disable=traefik` | Do **not** deploy the bundled Traefik ingress controller, so nothing tries to bind host :80/:443 (owned by nginx — see [1.6](#16-the-nginx-port-constraint-why-traefik-and-servicelb-are-disabled)). |
|
||||
| `--disable=servicelb` | Do **not** deploy Klipper servicelb, so `LoadBalancer` services do not bind host ports. Runners need no inbound `LoadBalancer`; the ARC listener reaches GitHub via **outbound** long-poll. |
|
||||
|
||||
Everything else is left at k3s defaults *on purpose* — those defaults are what the rest of this doc set relies on:
|
||||
|
||||
| Default (not overridden) | Value | Why we keep it |
|
||||
| --- | --- | --- |
|
||||
| CNI | **Flannel**, VXLAN backend | Simple single-node overlay; see [section 5](#5-cluster-networking-cni-cidrs--nat-egress). |
|
||||
| `--cluster-cidr` (pod network) | `10.42.0.0/16` | Pod IP range. |
|
||||
| `--service-cidr` (service network) | `10.43.0.0/16` | ClusterIP range. |
|
||||
| `--cluster-dns` (CoreDNS) | `10.43.0.10` | In-cluster DNS resolver. |
|
||||
| Container runtime | bundled **containerd v2** | Kata is wired into *this* containerd in [02-kata-runtime.md](02-kata-runtime.md). |
|
||||
|
||||
Because we do not pass `--node-ip` / `--flannel-iface`, k3s auto-detects the host's primary interface and uses its address as the node IP (the public IPv4 on the reference host). If your host has multiple NICs, set `--node-ip` / `--flannel-iface` explicitly.
|
||||
|
||||
The installer writes the systemd unit `/etc/systemd/system/k3s.service`. On the reference host its `ExecStart` is exactly:
|
||||
|
||||
```ini
|
||||
ExecStartPre=-/sbin/modprobe br_netfilter
|
||||
ExecStartPre=-/sbin/modprobe overlay
|
||||
ExecStart=/usr/local/bin/k3s \
|
||||
server \
|
||||
'--disable=traefik' \
|
||||
'--disable=servicelb' \
|
||||
```
|
||||
|
||||
There is **no** `/etc/rancher/k3s/config.yaml` on the host — the two `--disable` flags above are the *only* customization; everything else is the default set listed above.
|
||||
|
||||
Enable and check the service:
|
||||
|
||||
```bash
|
||||
systemctl enable --now k3s
|
||||
systemctl status k3s --no-pager
|
||||
journalctl -u k3s -f # follow startup logs until the node is Ready
|
||||
```
|
||||
|
||||
Confirm the version:
|
||||
|
||||
```
|
||||
$ k3s --version
|
||||
k3s version v1.35.5+k3s1 (6a4781ad)
|
||||
go version go1.25.9
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. kubeconfig
|
||||
|
||||
k3s writes an admin kubeconfig to `/etc/rancher/k3s/k3s.yaml`. Point `kubectl` at it:
|
||||
|
||||
```bash
|
||||
export KUBECONFIG=/etc/rancher/k3s/k3s.yaml
|
||||
# persist for future shells:
|
||||
echo 'export KUBECONFIG=/etc/rancher/k3s/k3s.yaml' >> ~/.bashrc
|
||||
```
|
||||
|
||||
The file targets the local API server over loopback (TLS material redacted):
|
||||
|
||||
```yaml
|
||||
apiVersion: v1
|
||||
clusters:
|
||||
- cluster:
|
||||
certificate-authority-data: <REDACTED>
|
||||
server: https://127.0.0.1:6443
|
||||
name: default
|
||||
contexts:
|
||||
- context:
|
||||
cluster: default
|
||||
user: default
|
||||
name: default
|
||||
current-context: default
|
||||
kind: Config
|
||||
users:
|
||||
- name: default
|
||||
user:
|
||||
client-certificate-data: <REDACTED>
|
||||
client-key-data: <REDACTED>
|
||||
```
|
||||
|
||||
> This kubeconfig embeds cluster-admin credentials. Treat the file as a secret (`chmod 600`, root-only). To administer the cluster from another machine, copy the file and replace `127.0.0.1` with the host's reachable address — on the reference host that is done over **Tailscale** (`tailscale0`), so the API server is never exposed on the public interface. Do not commit this file.
|
||||
|
||||
Quick check that the client can talk to the server:
|
||||
|
||||
```bash
|
||||
kubectl version # client + server versions
|
||||
kubectl cluster-info
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Verify the node and system pods
|
||||
|
||||
Wait until the node reports `Ready`:
|
||||
|
||||
```bash
|
||||
kubectl get nodes -o wide
|
||||
```
|
||||
|
||||
Reference output (redacted — `<CI_HOST>` is the hostname, `<PUBLIC_IP>` the auto-detected node IP):
|
||||
|
||||
```
|
||||
NAME STATUS ROLES AGE VERSION INTERNAL-IP EXTERNAL-IP OS-IMAGE KERNEL-VERSION CONTAINER-RUNTIME
|
||||
<CI_HOST> Ready control-plane,master ... v1.35.5+k3s1 <PUBLIC_IP> <none> CentOS Stream 10 (Coughlan) 7.0.10-1.el10.elrepo.x86_64 containerd://2.2.3-k3s1
|
||||
```
|
||||
|
||||
Then confirm the core system pods are running:
|
||||
|
||||
```bash
|
||||
kubectl get pods -A
|
||||
```
|
||||
|
||||
You should see the k3s base set in `kube-system` (these are what remain after disabling Traefik and servicelb):
|
||||
|
||||
```
|
||||
NAMESPACE NAME READY STATUS RESTARTS AGE
|
||||
kube-system coredns-<hash> 1/1 Running 0 ...
|
||||
kube-system local-path-provisioner-<hash> 1/1 Running 0 ...
|
||||
kube-system metrics-server-<hash> 1/1 Running 0 ...
|
||||
```
|
||||
|
||||
`local-path-provisioner` is the default storage class (used later for the RustFS cache PVC in [04-arc-and-caching.md](04-arc-and-caching.md)); `coredns` is cluster DNS; `metrics-server` backs `kubectl top`. There is intentionally **no** `traefik` or `svclb-*` pod.
|
||||
|
||||
---
|
||||
|
||||
## 5. Cluster networking (CNI, CIDRs & NAT egress)
|
||||
|
||||
### 5.1 Flannel CNI
|
||||
|
||||
k3s installs Flannel and writes its CNI config to `/var/lib/rancher/k3s/agent/etc/cni/net.d/10-flannel.conflist`. This file is generated by k3s — copied here verbatim (no secrets; nothing to redact):
|
||||
|
||||
```json
|
||||
{
|
||||
"name":"cbr0",
|
||||
"cniVersion":"1.0.0",
|
||||
"plugins":[
|
||||
{
|
||||
"type":"flannel",
|
||||
"delegate":{
|
||||
"hairpinMode":true,
|
||||
"forceAddress":true,
|
||||
"isDefaultGateway":true
|
||||
}
|
||||
},
|
||||
{
|
||||
"type":"portmap",
|
||||
"capabilities":{
|
||||
"portMappings":true
|
||||
}
|
||||
},
|
||||
{
|
||||
"type":"bandwidth",
|
||||
"capabilities":{
|
||||
"bandwidth":true
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
- `flannel` delegates to the `bridge` plugin (`cbr0`), with `isDefaultGateway` so the pod's default route points at the node — this is the path pod egress takes to reach the host's NAT.
|
||||
- `portmap` and `bandwidth` are standard chained plugins (host-port mapping and per-pod bandwidth shaping).
|
||||
- The backend is Flannel's default **VXLAN** (we did not override it at install). On a single node, pod-to-pod traffic stays on the local bridge.
|
||||
|
||||
### 5.2 Address ranges
|
||||
|
||||
The cluster uses the k3s defaults — keep these as-is (they are referenced throughout the doc set):
|
||||
|
||||
| Range | CIDR | Notes |
|
||||
| --- | --- | --- |
|
||||
| Pod network (cluster-cidr) | `10.42.0.0/16` | Single node carves a `/24` from this: `kubectl get node -o jsonpath='{.items[0].spec.podCIDR}'` → `10.42.0.0/24`. |
|
||||
| Service network (service-cidr) | `10.43.0.0/16` | ClusterIP services. |
|
||||
| CoreDNS service IP | `10.43.0.10` | Cluster DNS resolver (`kube-dns` Service). |
|
||||
|
||||
Confirm CoreDNS:
|
||||
|
||||
```
|
||||
$ kubectl -n kube-system get svc kube-dns
|
||||
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE
|
||||
kube-dns ClusterIP 10.43.0.10 <none> 53/UDP,53/TCP,9153/TCP ...
|
||||
```
|
||||
|
||||
### 5.3 Pod-to-internet egress via host firewalld NAT
|
||||
|
||||
Runner microVMs need outbound internet (to reach GitHub, fetch toolchains, etc.) but the host must **not** expose the cluster on its public interface. The path is:
|
||||
|
||||
```
|
||||
pod (10.42.0.0/16) --default route--> cbr0/flannel --> host routing --> firewalld masquerade (SNAT) --> <EXT_IFACE> --> internet (as <PUBLIC_IP>)
|
||||
```
|
||||
|
||||
This relies on `net.ipv4.ip_forward = 1` (set in [1.2](#12-kernel-modules-and-sysctls-for-k3s-networking)) plus firewalld **masquerade** (SNAT) and **forwarding** on the public zone. On the reference host the external interface lives in the default `public` zone with masquerade enabled:
|
||||
|
||||
```
|
||||
$ firewall-cmd --list-all
|
||||
public (default, active)
|
||||
target: default
|
||||
icmp-block-inversion: no
|
||||
interfaces: <EXT_IFACE>
|
||||
sources:
|
||||
services: cockpit dhcpv6-client http https ssh
|
||||
ports:
|
||||
protocols:
|
||||
forward: yes
|
||||
masquerade: yes
|
||||
forward-ports:
|
||||
source-ports:
|
||||
icmp-blocks:
|
||||
rich rules:
|
||||
```
|
||||
|
||||
The two lines that make pod egress work are **`masquerade: yes`** (SNAT pod source IPs to `<PUBLIC_IP>` on the way out) and **`forward: yes`** (allow routing between interfaces). The `services` list (`ssh`, `http`, `https`, `cockpit`, `dhcpv6-client`) is the host's own inbound allow-list and is unrelated to pod egress — note that `http`/`https` here are for the **host nginx**, not k3s.
|
||||
|
||||
The pod and service CIDRs (plus the Tailscale admin interface) are placed in the **trusted** zone so intra-cluster and admin traffic is accepted without per-rule firewalling:
|
||||
|
||||
```
|
||||
$ firewall-cmd --zone=trusted --list-all
|
||||
trusted (active)
|
||||
target: ACCEPT
|
||||
interfaces: tailscale0
|
||||
sources: 10.42.0.0/16 10.43.0.0/16
|
||||
forward: yes
|
||||
masquerade: no
|
||||
...
|
||||
```
|
||||
|
||||
```
|
||||
$ firewall-cmd --get-active-zones
|
||||
public (default)
|
||||
interfaces: <EXT_IFACE>
|
||||
trusted
|
||||
interfaces: tailscale0
|
||||
sources: 10.42.0.0/16 10.43.0.0/16
|
||||
docker # br-* bridges from the separate Docker stack (unrelated to k3s)
|
||||
interfaces: ...
|
||||
```
|
||||
|
||||
To reproduce this NAT setup on a fresh host (replace `<EXT_IFACE>` with your public NIC, e.g. `eth0`):
|
||||
|
||||
```bash
|
||||
# external interface in the public zone with NAT + forwarding
|
||||
firewall-cmd --permanent --zone=public --change-interface=<EXT_IFACE>
|
||||
firewall-cmd --permanent --zone=public --add-masquerade
|
||||
firewall-cmd --permanent --zone=public --add-forward # firewalld >= 0.9
|
||||
|
||||
# trust intra-cluster traffic (pod + service CIDRs) and the Tailscale admin iface
|
||||
firewall-cmd --permanent --zone=trusted --add-source=10.42.0.0/16
|
||||
firewall-cmd --permanent --zone=trusted --add-source=10.43.0.0/16
|
||||
firewall-cmd --permanent --zone=trusted --change-interface=tailscale0
|
||||
|
||||
firewall-cmd --reload
|
||||
```
|
||||
|
||||
Verify masquerade and forwarding are live:
|
||||
|
||||
```
|
||||
$ firewall-cmd --query-masquerade
|
||||
yes
|
||||
$ firewall-cmd --query-forward
|
||||
yes
|
||||
```
|
||||
|
||||
> This is host-level NAT only. A second, finer-grained layer — the `runner-egress-lockdown` Kubernetes **NetworkPolicy** — restricts *which* destinations runner pods may reach (it blocks `<PUBLIC_IP>`, the tailnet `100.64.0.0/10`, RFC-1918 ranges, etc., while allowing the public internet and cluster DNS). That policy is part of the runner setup and is documented in [04-arc-and-caching.md](04-arc-and-caching.md).
|
||||
|
||||
---
|
||||
|
||||
## 6. Verification: DNS & egress smoke test
|
||||
|
||||
The system pods being `Running` (section 4) already prove in-cluster networking. To prove **DNS resolution** and **pod-to-internet egress** end to end, run a throwaway pod (delete it afterward):
|
||||
|
||||
```bash
|
||||
# in-cluster DNS: resolve the kubernetes Service via CoreDNS (10.43.0.10)
|
||||
kubectl run dns-test --image=busybox:1.36 --restart=Never --rm -it -- \
|
||||
nslookup kubernetes.default.svc.cluster.local
|
||||
|
||||
# external DNS + egress: resolve and reach the internet through host NAT
|
||||
kubectl run egress-test --image=busybox:1.36 --restart=Never --rm -it -- \
|
||||
sh -c 'nslookup github.com && wget -qO- https://api.github.com/zen'
|
||||
```
|
||||
|
||||
Expected: the first command resolves to a `10.43.x.x` ClusterIP; the second resolves a public name and prints a line of text fetched from the internet (proving SNAT/masquerade works). If DNS fails, recheck CoreDNS (`kubectl -n kube-system get pods`); if egress fails, recheck `ip_forward`, `masquerade`, and `forward` from [section 5.3](#53-pod-to-internet-egress-via-host-firewalld-nat).
|
||||
|
||||
> On the live reference host, the ARC scale-set **listener** pod (in `arc-systems`) is itself continuous proof of working egress: it stays `Running` only because it can reach GitHub outbound through this exact NAT path.
|
||||
|
||||
---
|
||||
|
||||
## Next steps
|
||||
|
||||
The host now runs a healthy single-node k3s cluster with working networking and NAT egress, and Traefik/servicelb disabled so nginx keeps :80/:443. Continue with:
|
||||
|
||||
- **[02-kata-runtime.md](02-kata-runtime.md)** — install Kata Containers, wire it into k3s' bundled containerd, and register the `kata-qemu` RuntimeClass.
|
||||
- [README.md](README.md) — architecture overview, full component map, and the placeholder/redaction reference.
|
||||
@@ -0,0 +1,470 @@
|
||||
# 02 — Kata Containers runtime for k3s
|
||||
|
||||
This guide installs **Kata Containers 3.31.0** and wires it into the k3s-bundled
|
||||
containerd as a named runtime, `kata-qemu`, so that any pod carrying
|
||||
`runtimeClassName: kata-qemu` boots inside its own QEMU/KVM microVM with a
|
||||
**separate guest kernel** from the host.
|
||||
|
||||
- **Previous:** [`01-host-and-cluster.md`](./01-host-and-cluster.md) — host prep, k3s install, networking. You need a working single-node k3s, KVM enabled (`/dev/kvm` present), and nested-virt off (this is bare metal).
|
||||
- **Next:** [`03-runner-image.md`](./03-runner-image.md) — the preloaded GitHub Actions runner image that runs inside these microVMs.
|
||||
|
||||
Why a microVM per pod: the CI jobs run untrusted code (third-party deps, fork
|
||||
PRs). A `runc` container shares the host kernel; a Kata pod gets its own guest
|
||||
kernel and a hardware-virtualization boundary (Intel VT-x / AMD-V), so a kernel
|
||||
exploit inside a job does not reach the host. On `<CI_HOST>` the host runs the
|
||||
CentOS Stream 10 kernel `7.0.10-1.el10.elrepo.x86_64`, while every Kata guest
|
||||
runs `6.18.28-194` — the verification in the last section turns that gap into a
|
||||
one-line proof.
|
||||
|
||||
All host paths and commands below were taken from the live host (read-only) and
|
||||
redacted per the doc-set redaction map. Run them on **your** reproduction host;
|
||||
on the live host only the read-only inspections (`kata-runtime check`,
|
||||
`kata-runtime env`, `kubectl get …`) are safe.
|
||||
|
||||
---
|
||||
|
||||
## Step 1 — Install the Kata static release into `/opt/kata`
|
||||
|
||||
Kata ships a self-contained **static tarball**: a pinned QEMU, the guest kernel,
|
||||
the guest rootfs image, virtiofsd, the runtime, and the containerd shim, all
|
||||
rooted at `/opt/kata`. Nothing links against host libraries, so it coexists
|
||||
cleanly with the host's own QEMU/libvirt and survives OS upgrades.
|
||||
|
||||
```bash
|
||||
# amd64 host; pin the exact version so the kernel/image/QEMU triple is reproducible.
|
||||
KATA_VER=3.31.0
|
||||
curl -fsSL -o kata-static.tar.xz \
|
||||
"https://github.com/kata-containers/kata-containers/releases/download/${KATA_VER}/kata-static-${KATA_VER}-amd64.tar.xz"
|
||||
|
||||
# The archive is rooted at ./opt/kata, so extracting at / lands everything in /opt/kata.
|
||||
sudo tar -xf kata-static.tar.xz -C /
|
||||
```
|
||||
|
||||
Put the **shim** and the **CLI** on `PATH`. The shim symlink is what containerd
|
||||
resolves at launch time (see Step 2); the `kata-runtime` symlink is for
|
||||
host-side inspection and `kata-runtime check`/`env`:
|
||||
|
||||
```bash
|
||||
sudo ln -sf /opt/kata/bin/containerd-shim-kata-v2 /usr/local/bin/containerd-shim-kata-v2
|
||||
sudo ln -sf /opt/kata/bin/kata-runtime /usr/local/bin/kata-runtime
|
||||
```
|
||||
|
||||
On the live host both symlinks are in place:
|
||||
|
||||
```text
|
||||
/usr/local/bin/containerd-shim-kata-v2 -> /opt/kata/bin/containerd-shim-kata-v2
|
||||
/usr/local/bin/kata-runtime -> /opt/kata/bin/kata-runtime
|
||||
```
|
||||
|
||||
### Inspect the install layout
|
||||
|
||||
```text
|
||||
/opt/kata
|
||||
├── VERSION # "3.31.0"
|
||||
├── versions.yaml # pinned component versions (QEMU, kernel, rootfs)
|
||||
├── bin/ # qemu-system-x86_64, containerd-shim-kata-v2, kata-runtime, ...
|
||||
├── libexec/ # virtiofsd
|
||||
└── share/
|
||||
├── defaults/kata-containers/
|
||||
│ ├── configuration-qemu.toml # the active config for the kata-qemu runtime
|
||||
│ └── configuration.toml -> configuration-qemu.toml
|
||||
└── kata-containers/
|
||||
├── vmlinux.container -> vmlinux-6.18.28-194 # guest kernel
|
||||
└── kata-containers.img -> kata-ubuntu-noble.image # guest rootfs
|
||||
```
|
||||
|
||||
`bin/` ships several hypervisors (`qemu-system-x86_64`, `cloud-hypervisor`,
|
||||
`firecracker`, `jailer`) and the QEMU confidential-computing variants
|
||||
(`-snp-experimental`, `-tdx-experimental`); this setup uses plain
|
||||
`qemu-system-x86_64`. `share/defaults/kata-containers/` also carries
|
||||
`configuration-clh.toml`, `configuration-fc.toml`, etc. — one per hypervisor.
|
||||
We only use `configuration-qemu.toml`.
|
||||
|
||||
### Confirm version and capability
|
||||
|
||||
```console
|
||||
$ /opt/kata/bin/kata-runtime --version
|
||||
kata-runtime : 3.31.0
|
||||
commit : ddb8a5de89891f12e1ce0013eb066a330b2988b9
|
||||
OCI specs: 1.2.1
|
||||
```
|
||||
|
||||
The component pins live in `/opt/kata/versions.yaml`. The two that matter for
|
||||
the guest are the QEMU and kernel versions:
|
||||
|
||||
```yaml
|
||||
assets:
|
||||
hypervisor:
|
||||
qemu:
|
||||
version: "v10.2.1"
|
||||
kernel:
|
||||
version: "v6.18.28"
|
||||
```
|
||||
|
||||
`kata-runtime check` confirms the host can actually start a microVM (KVM
|
||||
present, CPU virtualization usable). Run it as a host-side smoke test before
|
||||
touching k3s:
|
||||
|
||||
```console
|
||||
$ /opt/kata/bin/kata-runtime check
|
||||
level=warning msg="Not running network checks as super user" arch=amd64 ...
|
||||
System is capable of running Kata Containers
|
||||
System can currently create Kata Containers
|
||||
```
|
||||
|
||||
`kata-runtime env` cross-checks the resolved kernel/image/hypervisor and the
|
||||
host kernel — note the guest kernel (`6.18.28-194`) versus the host kernel
|
||||
(`7.0.10`):
|
||||
|
||||
```console
|
||||
$ /opt/kata/bin/kata-runtime env | grep -iE 'Kernel|Image|MachineType|Path|Hypervisor'
|
||||
[Kernel]
|
||||
Path = "/opt/kata/share/kata-containers/vmlinux-6.18.28-194"
|
||||
[Image]
|
||||
Path = "/opt/kata/share/kata-containers/kata-ubuntu-noble.image"
|
||||
[Hypervisor]
|
||||
MachineType = "q35"
|
||||
Version = "QEMU emulator version 10.2.1 (kata-static) ..."
|
||||
Path = "/opt/kata/bin/qemu-system-x86_64"
|
||||
[Host]
|
||||
Kernel = "7.0.10-1.el10.elrepo.x86_64"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Step 2 — Register `kata-qemu` with k3s's containerd
|
||||
|
||||
k3s embeds its **own** containerd (v2 here) — not the host's. On every start it
|
||||
**regenerates** `/var/lib/rancher/k3s/agent/etc/containerd/config.toml` from a
|
||||
template; the header says so:
|
||||
|
||||
```toml
|
||||
# File generated by k3s. DO NOT EDIT. Use config-v3.toml.tmpl instead.
|
||||
version = 3
|
||||
imports = ["/var/lib/rancher/k3s/agent/etc/containerd/config-v3.toml.d/*.toml"]
|
||||
root = "/var/lib/rancher/k3s/agent/containerd"
|
||||
state = "/run/k3s/containerd"
|
||||
...
|
||||
[plugins.'io.containerd.cri.v1.runtime'.containerd.runtimes.runc]
|
||||
runtime_type = "io.containerd.runc.v2"
|
||||
```
|
||||
|
||||
Editing `config.toml` directly is pointless — k3s overwrites it on the next
|
||||
restart. The durable hook is the generated `imports` line: k3s merges any
|
||||
`*.toml` under `config-v3.toml.d/` into the final config. That directory is the
|
||||
**supported drop-in path** and survives k3s upgrades.
|
||||
|
||||
> k3s picks the drop-in directory from the containerd config schema version. With
|
||||
> the generated `version = 3` config the path is `config-v3.toml.d/`. (On older
|
||||
> k3s that emitted a v2 config it was `config.toml.d/`.) Match whatever your
|
||||
> generated `config.toml` declares.
|
||||
|
||||
Create the drop-in. This is the real file from the host, verbatim:
|
||||
|
||||
```toml
|
||||
# /var/lib/rancher/k3s/agent/etc/containerd/config-v3.toml.d/kata.toml
|
||||
# Kata Containers (QEMU/KVM microVM) runtime for k3s containerd.
|
||||
# Added out-of-band; merged via the generated config's `imports`.
|
||||
[plugins.'io.containerd.cri.v1.runtime'.containerd.runtimes.kata-qemu]
|
||||
runtime_type = "io.containerd.kata.v2"
|
||||
runtime_path = "/opt/kata/bin/containerd-shim-kata-v2"
|
||||
[plugins.'io.containerd.cri.v1.runtime'.containerd.runtimes.kata-qemu.options]
|
||||
ConfigPath = "/opt/kata/share/defaults/kata-containers/configuration-qemu.toml"
|
||||
```
|
||||
|
||||
What each line does:
|
||||
|
||||
- The table key `…runtimes.kata-qemu` defines a CRI runtime **named**
|
||||
`kata-qemu`. A RuntimeClass whose `handler` is `kata-qemu` (Step 3) selects
|
||||
exactly this entry.
|
||||
- `runtime_type = "io.containerd.kata.v2"` tells containerd this is a v2 (shim)
|
||||
runtime. By the v2 naming convention containerd would look for
|
||||
`containerd-shim-kata-v2` on `PATH` — which is why we symlinked it in Step 1.
|
||||
- `runtime_path = "/opt/kata/bin/containerd-shim-kata-v2"` pins the **exact**
|
||||
shim binary regardless of `PATH`, so the runtime is unambiguous even if the
|
||||
symlink is missing or another shim shadows it.
|
||||
- `options.ConfigPath` points the shim at the hypervisor configuration analyzed
|
||||
in Step 4. This is how a single shim binary can back multiple runtimes (e.g. a
|
||||
second `kata-clh` runtime pointing at `configuration-clh.toml`).
|
||||
|
||||
Restart k3s so the bundled containerd reloads and merges the drop-in:
|
||||
|
||||
```bash
|
||||
sudo systemctl restart k3s
|
||||
```
|
||||
|
||||
Verify containerd now knows the runtime (CRI reports the registered runtime
|
||||
handlers):
|
||||
|
||||
```bash
|
||||
sudo k3s crictl info | grep -A2 '"kata-qemu"'
|
||||
```
|
||||
|
||||
The `runc` default is untouched: pods without a `runtimeClassName` keep running
|
||||
as ordinary host-kernel containers. Only pods that opt into the RuntimeClass
|
||||
below get a microVM.
|
||||
|
||||
---
|
||||
|
||||
## Step 3 — Create the `kata-qemu` RuntimeClass
|
||||
|
||||
A Kubernetes [RuntimeClass](https://kubernetes.io/docs/concepts/containers/runtime-class/)
|
||||
maps a friendly name a pod can request to the containerd runtime handler
|
||||
registered in Step 2. The `handler` value **must** equal the runtime name in the
|
||||
drop-in (`kata-qemu`).
|
||||
|
||||
```yaml
|
||||
# kata-qemu-runtimeclass.yaml
|
||||
apiVersion: node.k8s.io/v1
|
||||
kind: RuntimeClass
|
||||
metadata:
|
||||
name: kata-qemu
|
||||
handler: kata-qemu
|
||||
```
|
||||
|
||||
```bash
|
||||
kubectl apply -f kata-qemu-runtimeclass.yaml
|
||||
```
|
||||
|
||||
Confirm it exists (live host, read-only):
|
||||
|
||||
```console
|
||||
$ kubectl get runtimeclass kata-qemu -o yaml
|
||||
apiVersion: node.k8s.io/v1
|
||||
handler: kata-qemu
|
||||
kind: RuntimeClass
|
||||
metadata:
|
||||
annotations:
|
||||
kubectl.kubernetes.io/last-applied-configuration: |
|
||||
{"apiVersion":"node.k8s.io/v1","handler":"kata-qemu","kind":"RuntimeClass","metadata":{"annotations":{},"name":"kata-qemu"}}
|
||||
creationTimestamp: "2026-06-14T19:18:16Z"
|
||||
name: kata-qemu
|
||||
resourceVersion: "580"
|
||||
uid: 50a199ac-3936-4ae8-9deb-be389ce9f042
|
||||
```
|
||||
|
||||
From here, any pod with `spec.runtimeClassName: kata-qemu` is scheduled onto the
|
||||
Kata shim. The ARC runner pod template sets exactly this — see
|
||||
[`04-arc-and-caching.md`](./04-arc-and-caching.md).
|
||||
|
||||
---
|
||||
|
||||
## Step 4 — Kata hypervisor configuration (`configuration-qemu.toml`)
|
||||
|
||||
The shim reads `/opt/kata/share/defaults/kata-containers/configuration-qemu.toml`
|
||||
(via `ConfigPath`). The file is long and mostly defaults; below are the **active
|
||||
settings that shape this deployment**, copied from the host with the noise
|
||||
stripped. Each is explained afterward.
|
||||
|
||||
```toml
|
||||
[hypervisor.qemu]
|
||||
path = "/opt/kata/bin/qemu-system-x86_64"
|
||||
kernel = "/opt/kata/share/kata-containers/vmlinux.container"
|
||||
image = "/opt/kata/share/kata-containers/kata-containers.img"
|
||||
machine_type = "q35"
|
||||
rootfs_type = "ext4"
|
||||
cpu_features = "pmu=off"
|
||||
kernel_params = "cgroup_no_v1=all systemd.unified_cgroup_hierarchy=1"
|
||||
|
||||
default_vcpus = 1
|
||||
default_maxvcpus = 0
|
||||
default_memory = 2048
|
||||
default_maxmemory = 0
|
||||
memory_slots = 10
|
||||
|
||||
shared_fs = "virtio-fs"
|
||||
virtio_fs_daemon = "/opt/kata/libexec/virtiofsd"
|
||||
virtio_fs_cache = "auto"
|
||||
virtio_fs_extra_args = ["--thread-pool-size=1", "--announce-submounts"]
|
||||
|
||||
disable_block_device_use = true
|
||||
block_device_driver = "virtio-scsi"
|
||||
block_device_aio = "io_uring"
|
||||
|
||||
[factory]
|
||||
enable_template = false
|
||||
vm_cache_number = 0
|
||||
|
||||
[runtime]
|
||||
internetworking_model = "tcfilter"
|
||||
emptydir_mode = "shared-fs"
|
||||
static_sandbox_resource_mgmt = false
|
||||
sandbox_cgroup_only = false
|
||||
```
|
||||
|
||||
### Boot artifacts: `path` / `kernel` / `image`
|
||||
|
||||
`path` is the bundled QEMU 10.2.1; `kernel` is the guest kernel
|
||||
(`vmlinux.container -> vmlinux-6.18.28-194`); `image` is the guest rootfs
|
||||
(`kata-containers.img -> kata-ubuntu-noble.image`, an Ubuntu Noble rootfs with
|
||||
`kata-agent` baked in as PID 1's manager). These three are the entire guest —
|
||||
none of them is the host kernel, which is the whole point. `machine_type =
|
||||
"q35"` is the modern PCIe QEMU machine (needed for PCIe hotplug);
|
||||
`cpu_features = "pmu=off"` disables the virtual perf-monitoring unit (avoids
|
||||
spurious PMU passthrough issues). `kernel_params` forces cgroup v2-only in the
|
||||
guest, matching a modern systemd userspace.
|
||||
|
||||
### vCPU / memory sizing — hotplug from pod requests/limits
|
||||
|
||||
This is the most important block to understand for a CI runner.
|
||||
|
||||
- `default_vcpus = 1` and `default_memory = 2048` (MiB) are the **boot-time**
|
||||
size. Every microVM starts tiny: 1 vCPU, 2 GiB.
|
||||
- `default_maxvcpus = 0` means "no fixed ceiling — use the host's physical CPU
|
||||
count" (32 on this box). `default_maxmemory = 0` likewise means "host total
|
||||
RAM". `memory_slots = 10` is the number of ACPI DIMM hotplug slots, i.e. how
|
||||
many memory-grow operations the guest can accept.
|
||||
- `static_sandbox_resource_mgmt = false` enables **dynamic** sizing: Kata reads
|
||||
the pod's container CPU/memory **limits** that the kubelet/CRI hands the shim
|
||||
and **hotplugs** vCPUs and memory after boot to match. (Set it to `true` and
|
||||
the VM stays fixed at the defaults regardless of pod limits.) So a runner pod
|
||||
requesting `2 CPU / 4Gi` with limits `6 CPU / 12Gi` boots at 1 vCPU/2 GiB and
|
||||
grows toward 6 vCPU / 12 GiB. If a pod sets no limits, the VM stays at the
|
||||
defaults.
|
||||
|
||||
The practical knobs: raise `default_vcpus`/`default_memory` only if jobs need a
|
||||
big VM *immediately* at boot (hotplug has latency); otherwise leave them small
|
||||
and let the pod's `resources.limits` drive the size. Sizing the runner pod is
|
||||
covered in [`04-arc-and-caching.md`](./04-arc-and-caching.md).
|
||||
|
||||
### `shared_fs = "virtio-fs"` — sharing the container rootfs into the VM
|
||||
|
||||
The container's rootfs is prepared on the **host** by containerd's overlayfs
|
||||
snapshotter (see below). Rather than repackaging it as a virtual disk, Kata runs
|
||||
**virtiofsd** (`/opt/kata/libexec/virtiofsd`) on the host to export that
|
||||
directory over **virtio-fs**, and the guest mounts it as the container root.
|
||||
This is why `disable_block_device_use = true`: the rootfs travels in over the
|
||||
shared filesystem, not as a block device. Benefits: no image-to-block
|
||||
conversion, near-instant rootfs availability, and host/guest can both see the
|
||||
files. `virtio_fs_cache = "auto"` and `--thread-pool-size=1` are conservative
|
||||
caching/concurrency defaults; `--announce-submounts` makes nested mounts visible
|
||||
to the guest.
|
||||
|
||||
`emptydir_mode = "shared-fs"` extends the same mechanism to Kubernetes
|
||||
`emptyDir` volumes — they are shared into the guest over virtio-fs instead of
|
||||
being block devices.
|
||||
|
||||
### `block_device_driver = "virtio-scsi"` — for the volumes that *are* blocks
|
||||
|
||||
Even with `disable_block_device_use = true` for the rootfs, any genuine block
|
||||
volume (e.g. a `local-path` PVC presented as a device) is attached over a
|
||||
**virtio-scsi** controller, with `block_device_aio = "io_uring"` for efficient
|
||||
async I/O. virtio-scsi (vs virtio-blk) supports more disks per controller and
|
||||
hotplug, which matters when volumes attach after boot.
|
||||
|
||||
### vsock agent channel
|
||||
|
||||
The shim on the host talks to `kata-agent` inside the guest over a
|
||||
**VIRTIO-VSOCK** channel — a host↔guest socket transport that needs no guest IP
|
||||
or network. All container lifecycle operations (create/start/exec/IO/metrics)
|
||||
are ttRPC calls over that vsock link. In Kata 3.x vsock is the default and only
|
||||
agent transport, so there is no `use_vsock` toggle to set; the config instead
|
||||
shows `use_legacy_serial = false`, confirming the guest console/agent path is on
|
||||
the modern virtio channel rather than a legacy serial port. The upshot: the
|
||||
agent control plane is isolated from the pod's data-plane networking entirely.
|
||||
|
||||
### Networking into the guest
|
||||
|
||||
`internetworking_model = "tcfilter"` is how the pod's CNI veth reaches the VM:
|
||||
Kata creates a TAP device for the guest NIC and installs a TC (traffic-control)
|
||||
filter that mirrors packets between the CNI-provided veth and the TAP. The pod
|
||||
keeps the IP Flannel assigned it; the VM transparently sits behind it. Egress
|
||||
restrictions are enforced one layer up by a NetworkPolicy — see
|
||||
[`04-arc-and-caching.md`](./04-arc-and-caching.md).
|
||||
|
||||
### `factory.enable_template = false` — a fresh VM per job, deliberately
|
||||
|
||||
Kata's **VM factory/templating** can pre-create a paused "template" VM and fork
|
||||
new microVMs from it via copy-on-write memory, shaving boot time. It is **off**
|
||||
here (`enable_template = false`, `vm_cache_number = 0`) on purpose: CI jobs must
|
||||
be mutually isolated and reproducible, so each job gets a **pristine VM built
|
||||
from scratch** with no memory state inherited from a previous job. The boot cost
|
||||
(a second or two) is an acceptable price for clean isolation, and the preloaded
|
||||
runner image ([`03-runner-image.md`](./03-runner-image.md)) is what removes the
|
||||
*real* per-job cost (dependency installs), not VM templating.
|
||||
|
||||
### Interaction with the overlayfs snapshotter
|
||||
|
||||
k3s's containerd uses the default **overlayfs** snapshotter. For a `runc` pod
|
||||
that overlay mount *is* the container root. For a Kata pod, containerd still
|
||||
builds the same overlayfs rootfs on the host, but because `shared_fs =
|
||||
"virtio-fs"` it is **exported into the guest by virtiofsd** rather than used
|
||||
directly. So the two cooperate cleanly: the snapshotter assembles image layers
|
||||
on the host (image pulls, layer caching, dedup all work normally), and virtio-fs
|
||||
projects the result into the microVM. No special snapshotter (devmapper /
|
||||
blockfile) is needed — that would only be required if you wanted the rootfs
|
||||
delivered as a block device instead of a shared filesystem.
|
||||
|
||||
---
|
||||
|
||||
## Step 5 — Verify microVM isolation
|
||||
|
||||
The defining test: a pod under `kata-qemu` must report a **different kernel**
|
||||
than the host. Run a throwaway pod with `--rm` so nothing is left behind. On
|
||||
your reproduction host:
|
||||
|
||||
```bash
|
||||
# Inside the microVM (Kata): guest kernel.
|
||||
kubectl run kata-smoke --rm -it --restart=Never \
|
||||
--image=busybox \
|
||||
--overrides='{"spec":{"runtimeClassName":"kata-qemu"}}' \
|
||||
-- uname -r
|
||||
```
|
||||
|
||||
Expected — the **guest** kernel:
|
||||
|
||||
```text
|
||||
6.18.28-194
|
||||
```
|
||||
|
||||
Compare with the host kernel:
|
||||
|
||||
```console
|
||||
$ uname -r
|
||||
7.0.10-1.el10.elrepo.x86_64
|
||||
```
|
||||
|
||||
Different kernel string == the workload is genuinely inside a separate guest
|
||||
kernel, not a namespaced host process. As a control, the same pod **without**
|
||||
the RuntimeClass runs on `runc` and prints the **host** kernel
|
||||
(`7.0.10-1.el10.elrepo.x86_64`) — proving the difference comes from Kata, not
|
||||
the image.
|
||||
|
||||
For a closer look at the VM's resources (confirming the hotplug sizing from Step
|
||||
4), use an image with more tools:
|
||||
|
||||
```bash
|
||||
kubectl run kata-smoke --rm -it --restart=Never \
|
||||
--image=ubuntu:24.04 \
|
||||
--overrides='{"spec":{"runtimeClassName":"kata-qemu"}}' \
|
||||
-- bash -lc 'uname -r; nproc; grep MemTotal /proc/meminfo'
|
||||
```
|
||||
|
||||
This prints the guest kernel, the hotplugged vCPU count, and guest RAM (MiB) —
|
||||
which track the pod's `resources.limits`, not the host's 32c/125G.
|
||||
|
||||
> **On the live host, do not run throwaway pods.** It is production CI. Use the
|
||||
> read-only host checks instead: `kata-runtime check` and `kata-runtime env`
|
||||
> (Step 1) prove the runtime is healthy without scheduling anything. The
|
||||
> equivalent VM-boot proof for the actual runner image is the `preload-verify`
|
||||
> recipe documented in [`03-runner-image.md`](./03-runner-image.md).
|
||||
|
||||
---
|
||||
|
||||
## Recap
|
||||
|
||||
1. Static tarball extracted to `/opt/kata` (QEMU + guest kernel + rootfs +
|
||||
virtiofsd + shim), shim and CLI symlinked onto `PATH`.
|
||||
2. Drop-in `config-v3.toml.d/kata.toml` registers the `kata-qemu` runtime
|
||||
(`io.containerd.kata.v2`, pinned `runtime_path`, `ConfigPath`), merged via
|
||||
k3s's generated `imports`; picked up on `systemctl restart k3s`.
|
||||
3. `RuntimeClass/kata-qemu` (`handler: kata-qemu`) lets pods opt in.
|
||||
4. `configuration-qemu.toml` boots a small QEMU q35 VM (1 vCPU / 2 GiB) that
|
||||
hotplugs up to the pod's limits, shares the container rootfs in over
|
||||
virtio-fs, talks to the agent over vsock, and builds a fresh VM per job
|
||||
(templating off).
|
||||
5. A `kata-qemu` pod reports guest kernel `6.18.28-194` vs host
|
||||
`7.0.10-1.el10.elrepo.x86_64` — isolation confirmed.
|
||||
|
||||
Continue to [`03-runner-image.md`](./03-runner-image.md) to build the runner
|
||||
image that boots inside these microVMs.
|
||||
@@ -0,0 +1,432 @@
|
||||
# 03 - Preloaded runner image
|
||||
|
||||
Each ephemeral CI job boots a fresh Kata microVM and runs inside a single
|
||||
runner container. To keep that cold microVM **warm** (no per-job dependency
|
||||
fetches), the container does not use the stock GitHub Actions runner image
|
||||
directly: it uses a locally built image that bakes every dependency CI installs
|
||||
on every job into the snapshot. This document covers building that image,
|
||||
importing it into k3s containerd, pointing ARC at it, and verifying it.
|
||||
|
||||
Navigation: previous - [02-kata-runtime.md](./02-kata-runtime.md) (Kata +
|
||||
containerd + RuntimeClass) | next - [04-arc-and-caching.md](./04-arc-and-caching.md)
|
||||
(ARC runners, RustFS cache, egress policy).
|
||||
|
||||
> The image built here is referenced by the ARC runner pod template
|
||||
> (`template.spec.containers[0].image` + `imagePullPolicy: IfNotPresent`),
|
||||
> documented in [04-arc-and-caching.md](./04-arc-and-caching.md). The pinned
|
||||
> `runtimeClassName: kata-qemu` on that same template comes from
|
||||
> [02-kata-runtime.md](./02-kata-runtime.md).
|
||||
|
||||
All host commands below run on `<CI_HOST>` (the single k3s node) as root.
|
||||
|
||||
---
|
||||
|
||||
## 1. Why a preloaded image
|
||||
|
||||
The base `ghcr.io/actions/actions-runner:latest` is a clean Ubuntu 24.04 runner.
|
||||
On a normal (GitHub-hosted-style) runner, the CI workflow installs its system
|
||||
dependencies at the start of every job: the cairo/pango native stack for canvas
|
||||
builds, `fd`/`ripgrep`/`imagemagick`, `bun`, and a pinned Rust nightly with the
|
||||
cross targets. Inside a Kata microVM that is destroyed after a single job, paying
|
||||
that apt/bun/rustup cost on **every** job is pure latency - the microVM starts
|
||||
cold each time.
|
||||
|
||||
The preloaded image moves that work to build time. Every ephemeral runner then
|
||||
starts with the toolchain already present: apt deps are not re-fetched, `bun` and
|
||||
`cargo`/`rustc` are on `PATH`, and the pinned Rust toolchain is already the
|
||||
default so target/component installs in CI become no-ops.
|
||||
|
||||
### Stay in sync with `setup-system-deps`
|
||||
|
||||
The apt set baked into the image **must** match the repo's
|
||||
`.github/actions/setup-system-deps` composite action. That action is the
|
||||
self-healing counterpart: it probes for the baked tools and skips the apt
|
||||
round-trip when they are present (preloaded image), but installs the exact same
|
||||
set on a stock runner so CI still works anywhere. Its detection probes are:
|
||||
|
||||
```bash
|
||||
command -v fd && command -v rg && command -v magick \
|
||||
&& pkg-config --exists cairo pango
|
||||
```
|
||||
|
||||
If you add a dependency that the action probes for, add it in **both** the
|
||||
Dockerfile apt line and `setup-system-deps`. If they drift, either the action
|
||||
re-installs deps the image already has (slow) or CI breaks on a tool the image
|
||||
forgot to bake.
|
||||
|
||||
---
|
||||
|
||||
## 2. The Dockerfile
|
||||
|
||||
Build context lives at `/root/omp-kata-runner-image/`. The Dockerfile below is
|
||||
reproduced verbatim (it contains no secrets or redactable host identifiers; the
|
||||
`ARG` pins, the full apt set, and the toolchain steps are real).
|
||||
|
||||
> The canonical copy of this Dockerfile is version-controlled at
|
||||
> [`infra/runner.Dockerfile`](../runner.Dockerfile);
|
||||
> the `/root/omp-kata-runner-image/` copy is overwritten from it by the
|
||||
> repo-driven reload script (see section 3 below).
|
||||
|
||||
```dockerfile
|
||||
# syntax=docker/dockerfile:1
|
||||
# Preloaded omp-kata runner image.
|
||||
#
|
||||
# Stock GitHub Actions runner (Ubuntu 24.04) with the dependencies CI installs
|
||||
# on every job baked in, so each ephemeral Kata microVM boots with them already
|
||||
# present instead of re-fetching them per job:
|
||||
# - APT system deps (canvas/cairo stack + fd/ripgrep/imagemagick) + fd/magick shims
|
||||
# - GitHub CLI (gh) — present on GitHub-hosted runners; the coding-agent github
|
||||
# tool and release workflows expect it
|
||||
# - C/build toolchain the native + canvas builds need
|
||||
# - bun (system-wide, on PATH)
|
||||
# - rust nightly toolchain (pinned) + clippy/rustfmt + linux-arm64/windows-msvc targets
|
||||
#
|
||||
# Rebuild + reimport (see /root/omp-kata-runner.md) after bumping the ARGs below
|
||||
# or the apt set. Keep the apt set in sync with .github/actions/setup-system-deps.
|
||||
FROM ghcr.io/actions/actions-runner:latest
|
||||
|
||||
ARG RUST_NIGHTLY=nightly-2026-04-29
|
||||
ARG BUN_VERSION=1.3.14
|
||||
|
||||
USER root
|
||||
ENV DEBIAN_FRONTEND=noninteractive
|
||||
|
||||
# Mirrors the "Install system deps" block in .github/workflows/ci.yml plus the
|
||||
# C/build toolchain (native + canvas builds) and the GitHub CLI. The gh apt repo
|
||||
# is added first so `gh` installs in the same apt transaction.
|
||||
RUN curl -fsSL https://cli.github.com/packages/githubcli-archive-keyring.gpg -o /usr/share/keyrings/githubcli-archive-keyring.gpg \
|
||||
&& chmod go+r /usr/share/keyrings/githubcli-archive-keyring.gpg \
|
||||
&& echo "deb [arch=$(dpkg --print-architecture) signed-by=/usr/share/keyrings/githubcli-archive-keyring.gpg] https://cli.github.com/packages stable main" > /etc/apt/sources.list.d/github-cli.list \
|
||||
&& apt-get update \
|
||||
&& apt-get install -y \
|
||||
build-essential pkg-config curl ca-certificates git unzip xz-utils zstd gh \
|
||||
libcairo2-dev libpango1.0-dev libjpeg-dev libgif-dev librsvg2-dev \
|
||||
fd-find ripgrep imagemagick \
|
||||
&& ln -sf "$(command -v fdfind)" /usr/local/bin/fd \
|
||||
&& ln -sf /usr/bin/convert /usr/local/bin/magick \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# bun, system-wide (BUN_INSTALL/bin == /usr/local/bin, already on PATH).
|
||||
ENV BUN_INSTALL=/usr/local
|
||||
RUN curl -fsSL https://bun.sh/install | bash -s "bun-v${BUN_VERSION}" \
|
||||
&& bun --version
|
||||
|
||||
# rust toolchain for the runner user; rustup default == pinned nightly so
|
||||
# dtolnay/rust-toolchain@nightly and target/component adds are no-ops in CI.
|
||||
USER runner
|
||||
ENV RUSTUP_HOME=/home/runner/.rustup \
|
||||
CARGO_HOME=/home/runner/.cargo \
|
||||
PATH=/home/runner/.cargo/bin:/usr/local/bin:${PATH}
|
||||
RUN curl --proto '=https' --tlsv1.2 -fsSL https://sh.rustup.rs \
|
||||
| sh -s -- -y --default-toolchain "${RUST_NIGHTLY}" --profile minimal \
|
||||
&& rustup component add clippy rustfmt \
|
||||
&& rustup target add aarch64-unknown-linux-gnu x86_64-pc-windows-msvc \
|
||||
&& cargo --version && rustc --version
|
||||
```
|
||||
|
||||
### Stage-by-stage annotation
|
||||
|
||||
**`# syntax=docker/dockerfile:1` + `FROM ghcr.io/actions/actions-runner:latest`.**
|
||||
BuildKit frontend pin, then the stock Actions runner base (Ubuntu 24.04). The
|
||||
base already ships the runner agent and its `/home/runner/run.sh` entrypoint, the
|
||||
non-root `runner` user, and `sudo`. Everything below layers onto that base; the
|
||||
runner agent itself is never modified.
|
||||
|
||||
**`ARG RUST_NIGHTLY` / `ARG BUN_VERSION`.** The two version knobs you bump. They
|
||||
are build args so you can also override them ad hoc with
|
||||
`docker build --build-arg RUST_NIGHTLY=... --build-arg BUN_VERSION=...` without
|
||||
editing the file. `RUST_NIGHTLY` must match what the repo's
|
||||
`dtolnay/rust-toolchain@nightly` step expects so the toolchain install in CI is a
|
||||
no-op (see step 4 below).
|
||||
|
||||
**`USER root` + `ENV DEBIAN_FRONTEND=noninteractive`.** Switch to root for the
|
||||
apt and bun system installs; `noninteractive` suppresses debconf/tzdata prompts
|
||||
during `apt-get install`.
|
||||
|
||||
**The apt `RUN` block.** This is the set that must mirror `setup-system-deps`.
|
||||
In order:
|
||||
- The first three lines add the **GitHub CLI apt repository** (keyring + signed
|
||||
source list) *before* `apt-get update`, so `gh` resolves and installs in the
|
||||
same apt transaction as everything else. `gh` is present on GitHub-hosted
|
||||
runners and is expected by the release workflows and the coding-agent `github`
|
||||
tool.
|
||||
- `apt-get install` pulls three groups:
|
||||
- **build toolchain / utilities:** `build-essential pkg-config curl
|
||||
ca-certificates git unzip xz-utils zstd gh`. `build-essential` + `pkg-config`
|
||||
are needed by the native and canvas builds; `zstd` is the codec the bun and
|
||||
sccache cache tarballs use (see [04-arc-and-caching.md](./04-arc-and-caching.md)).
|
||||
- **canvas / cairo native stack:** `libcairo2-dev libpango1.0-dev libjpeg-dev
|
||||
libgif-dev librsvg2-dev` - the `-dev` headers the canvas/rsvg native modules
|
||||
compile against.
|
||||
- **CLI tools:** `fd-find ripgrep imagemagick`, used by the agent and tests.
|
||||
- **The two shims** normalize Debian's binary names to what callers expect:
|
||||
Debian ships `fd` as `fdfind`, so `ln -sf "$(command -v fdfind)"
|
||||
/usr/local/bin/fd` exposes it as `fd`; ImageMagick installs `convert`, so
|
||||
`ln -sf /usr/bin/convert /usr/local/bin/magick` exposes the v7-style `magick`
|
||||
name. These two shims are exactly what `setup-system-deps` recreates on a stock
|
||||
runner.
|
||||
- `rm -rf /var/lib/apt/lists/*` drops the apt index to keep the layer smaller.
|
||||
|
||||
**bun (`ENV BUN_INSTALL=/usr/local` + install `RUN`).** Setting
|
||||
`BUN_INSTALL=/usr/local` makes the official installer drop the binary at
|
||||
`/usr/local/bin/bun`, which is already on `PATH` for every user - so bun is
|
||||
**system-wide** with no per-user shell init. The version is pinned via
|
||||
`bun-v${BUN_VERSION}`, and `bun --version` fails the build if the install is
|
||||
broken.
|
||||
|
||||
**Rust toolchain (`USER runner` + rustup `RUN`).** The toolchain is installed as
|
||||
the **`runner` user** - the UID jobs execute as - so cargo/rustc are owned by and
|
||||
visible to the job without sudo. `RUSTUP_HOME`/`CARGO_HOME` are pinned under
|
||||
`/home/runner`, and `~/.cargo/bin` is prepended to `PATH`. rustup installs the
|
||||
pinned nightly as the **default toolchain** (`--profile minimal`), then adds the
|
||||
`clippy` and `rustfmt` components and the `aarch64-unknown-linux-gnu`
|
||||
(Linux arm64) and `x86_64-pc-windows-msvc` (Windows cross) targets. Because the
|
||||
default toolchain already *is* the pinned nightly with these components/targets,
|
||||
the corresponding `rustup` steps in CI become no-ops - the warm-start payoff.
|
||||
`cargo --version && rustc --version` is the final build-time sanity check.
|
||||
|
||||
---
|
||||
|
||||
## 3. Build, import, and roll out (`reload.sh`)
|
||||
|
||||
`/root/omp-kata-runner-image/reload.sh` does the whole cycle: build, in-image
|
||||
smoke test, import into k3s containerd, point ARC at the new tag, and verify the
|
||||
rollout. It is idempotent and cache-fast on an unchanged rebuild. Reproduced
|
||||
verbatim (no secrets; the `/root` and kubeconfig paths are the real host paths):
|
||||
|
||||
```bash
|
||||
#!/usr/bin/env bash
|
||||
# Rebuild the preloaded omp-kata runner image, import it into k3s containerd,
|
||||
# point the ARC runner scale set at it, and roll it out. Idempotent: safe to
|
||||
# re-run after editing ./Dockerfile. Docker layer cache makes an unchanged
|
||||
# rebuild near-instant.
|
||||
#
|
||||
# ./reload.sh # build tag omp-kata-runner:YYYY-MM-DD-HHMMSS
|
||||
# ./reload.sh 2026-06-20 # build tag omp-kata-runner:2026-06-20
|
||||
# ./reload.sh foo:bar # build an explicit repo:tag
|
||||
set -euo pipefail
|
||||
export KUBECONFIG=/etc/rancher/k3s/k3s.yaml
|
||||
cd "$(dirname "$0")"
|
||||
|
||||
arg="${1:-$(date +%Y-%m-%d-%H%M%S)}"
|
||||
case "$arg" in *:*) IMAGE="$arg";; *) IMAGE="omp-kata-runner:$arg";; esac
|
||||
|
||||
echo "==> [1/5] building $IMAGE"
|
||||
DOCKER_BUILDKIT=1 docker build -t "$IMAGE" -t omp-kata-runner:preloaded .
|
||||
|
||||
echo "==> [2/5] verifying baked tools"
|
||||
docker run --rm --entrypoint bash "$IMAGE" -lc '
|
||||
set -e
|
||||
for b in gh fd rg magick bun cargo rustc pkg-config zstd; do
|
||||
command -v "$b" >/dev/null || { echo "MISSING: $b"; exit 1; }
|
||||
done
|
||||
echo "tools OK | $(bun --version) | $(rustc --version) | gh $(gh --version | head -1 | cut -d" " -f3)"
|
||||
'
|
||||
|
||||
echo "==> [3/5] importing into k3s containerd (k8s.io namespace)"
|
||||
docker save "$IMAGE" | k3s ctr -n k8s.io images import --platform linux/amd64 -
|
||||
|
||||
echo "==> [4/5] pointing ARC runner scale set at $IMAGE"
|
||||
sed -i "s#image: omp-kata-runner:.*#image: $IMAGE#" /root/arc-omp-values.yaml
|
||||
helm upgrade omp-kata --namespace arc-runners --version 0.14.2 \
|
||||
-f /root/arc-omp-values.yaml \
|
||||
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set >/dev/null
|
||||
|
||||
echo "==> [5/5] verifying rollout"
|
||||
live="$(kubectl get autoscalingrunnerset omp-kata -n arc-runners -o jsonpath='{.spec.template.spec.containers[0].image}')"
|
||||
echo "ARC runner image is now: $live"
|
||||
[ "$live" = "$IMAGE" ] && echo "OK: reloaded $IMAGE" || { echo "MISMATCH: expected $IMAGE"; exit 1; }
|
||||
```
|
||||
|
||||
Run it with no argument for an auto-dated tag:
|
||||
|
||||
```bash
|
||||
cd /root/omp-kata-runner-image
|
||||
./reload.sh
|
||||
```
|
||||
|
||||
### What each step does
|
||||
|
||||
**Preamble.** `set -euo pipefail` aborts on the first error; `KUBECONFIG` points
|
||||
at the k3s admin config; `cd` into the build context. The tag is resolved from
|
||||
`$1`: no arg gives a timestamped `omp-kata-runner:YYYY-MM-DD-HHMMSS`; an argument
|
||||
containing a colon (`foo:bar`) is used as an explicit `repo:tag`; anything else
|
||||
is treated as a tag suffix on `omp-kata-runner:`.
|
||||
|
||||
**[1/5] build.** `DOCKER_BUILDKIT=1 docker build` tags the result twice: the
|
||||
immutable `$IMAGE` (dated) and the moving `omp-kata-runner:preloaded` alias.
|
||||
BuildKit + the docker layer cache make an unchanged rebuild near-instant.
|
||||
|
||||
**[2/5] verify baked tools.** Runs the freshly built image with a bash entrypoint
|
||||
and asserts every expected binary is on `PATH`
|
||||
(`gh fd rg magick bun cargo rustc pkg-config zstd`), failing the whole script if
|
||||
any is missing, then prints the bun / rustc / gh versions. This catches a broken
|
||||
apt set, missing shim, or bad toolchain pin **before** anything touches the
|
||||
cluster.
|
||||
|
||||
**[3/5] import into k3s containerd.**
|
||||
`docker save "$IMAGE" | k3s ctr -n k8s.io images import --platform linux/amd64 -`
|
||||
streams the image tarball straight from the docker daemon into k3s's **own**
|
||||
containerd, in the `k8s.io` namespace - the namespace the kubelet pulls from.
|
||||
`--platform linux/amd64` matches the host architecture. Note `docker save` is
|
||||
given only the dated `$IMAGE`, so **only the dated tag is imported**; the
|
||||
`:preloaded` alias stays a docker-local convenience and is never imported or
|
||||
referenced by ARC.
|
||||
|
||||
**[4/5] point ARC at the new tag.** `sed -i` rewrites the single
|
||||
`image: omp-kata-runner:...` line in `/root/arc-omp-values.yaml` to the new tag,
|
||||
then `helm upgrade` re-renders the runner scale set with the chart pinned to
|
||||
`0.14.2`. (That values file is the runner pod template, documented in
|
||||
[04-arc-and-caching.md](./04-arc-and-caching.md).)
|
||||
|
||||
**[5/5] verify rollout.** Reads the image back off the live
|
||||
`autoscalingrunnerset` via `kubectl ... jsonpath` and asserts it equals `$IMAGE`,
|
||||
printing `OK` or exiting non-zero with `MISMATCH`. New ephemeral runner pods
|
||||
created after this point boot from the new image; in-flight jobs finish on the
|
||||
old one (the scale set is scale-to-zero, so this drains quickly).
|
||||
|
||||
### Running it from the repo (over SSH)
|
||||
|
||||
You do not have to keep `reload.sh` on the host. The repo ships the version-
|
||||
controlled Dockerfile plus an SSH-driven wrapper that performs the same build,
|
||||
import, and rollout remotely, so the whole image lifecycle is managed from a
|
||||
checkout:
|
||||
|
||||
- [`infra/runner.Dockerfile`](../runner.Dockerfile) - the image definition (source of truth).
|
||||
- [`infra/reload-runner.sh`](../reload-runner.sh) - copies that Dockerfile to the host, then runs the build / verify / import / `helm upgrade` / verify steps over SSH.
|
||||
|
||||
The host is never hardcoded; point it at your node with `CI_HOST`:
|
||||
|
||||
```bash
|
||||
CI_HOST=<CI_HOST> ./infra/reload-runner.sh # dated tag
|
||||
CI_HOST=<CI_HOST> ./infra/reload-runner.sh 2026-06-20 # explicit tag
|
||||
```
|
||||
|
||||
It honors the same defaults as the host script (remote build dir, ARC values
|
||||
path, release name, namespace, chart version), each overridable via the
|
||||
environment variables documented in the script header.
|
||||
|
||||
---
|
||||
|
||||
## 4. Why import into k3s containerd instead of using a registry
|
||||
|
||||
This is a single-node cluster, and the **only** consumer of the runner image is
|
||||
the kubelet/containerd on that same node. A registry would add a service to run,
|
||||
secure, and authenticate against, for zero benefit. Instead:
|
||||
|
||||
- `docker save | k3s ctr -n k8s.io images import` places the image directly into
|
||||
the containerd instance k3s schedules from. containerd normalizes the short
|
||||
reference `omp-kata-runner:<tag>` to `docker.io/library/omp-kata-runner:<tag>`
|
||||
in its store (verified: the imported tags appear under that prefix).
|
||||
- The ARC pod template sets `imagePullPolicy: IfNotPresent`. Because the image is
|
||||
already present locally, the kubelet **uses the local copy and never attempts a
|
||||
pull** - no registry, no pull credentials, no registry egress (which the
|
||||
runner egress lockdown in [04-arc-and-caching.md](./04-arc-and-caching.md)
|
||||
would block anyway).
|
||||
|
||||
Trade-off: the image must be (re-)imported on every node that schedules runners.
|
||||
Here that is exactly one node, so a re-roll is simply rebuild + re-import, and the
|
||||
next job's microVM starts cold but with warm dependencies from the local store.
|
||||
|
||||
---
|
||||
|
||||
## 5. Tag conventions
|
||||
|
||||
| Tag | Mutability | Imported into containerd? | Referenced by ARC? | Purpose |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| `omp-kata-runner:YYYY-MM-DD-HHMMSS` | immutable | yes | yes | the build of record; what runners actually boot |
|
||||
| `omp-kata-runner:preloaded` | moving | no | no | docker-local alias to the most recent build |
|
||||
|
||||
- The default `reload.sh` tag is timestamped (`date +%Y-%m-%d-%H%M%S`). You can
|
||||
also pass a date-only tag (`./reload.sh 2026-06-20`) or an explicit `repo:tag`.
|
||||
- ARC always pins the **immutable dated tag**, never `:preloaded`. That keeps a
|
||||
rollout reproducible and makes rollback trivial: `sed` the image line back to
|
||||
the prior dated tag (it is still in the local store) and `helm upgrade`.
|
||||
- Live example at time of writing: the scale set references
|
||||
`omp-kata-runner:2026-06-15-002621`, held in containerd as
|
||||
`docker.io/library/omp-kata-runner:2026-06-15-002621`.
|
||||
|
||||
---
|
||||
|
||||
## 6. Bumping bun / Rust / the apt set and re-rolling
|
||||
|
||||
1. **bun:** edit `ARG BUN_VERSION=` in the Dockerfile (or pass
|
||||
`--build-arg BUN_VERSION=...`).
|
||||
2. **Rust:** edit `ARG RUST_NIGHTLY=` to the new pinned nightly. Keep it equal to
|
||||
what the repo's `dtolnay/rust-toolchain@nightly` step resolves, so the CI
|
||||
toolchain install stays a no-op.
|
||||
3. **apt set:** edit the `apt-get install` line. You **must** mirror the change in
|
||||
`.github/actions/setup-system-deps` (and, if you add a tool the action probes
|
||||
for, in its detection block - currently `fd`, `rg`, `magick`,
|
||||
`pkg-config --exists cairo pango`).
|
||||
4. Re-roll:
|
||||
|
||||
```bash
|
||||
cd /root/omp-kata-runner-image
|
||||
./reload.sh
|
||||
```
|
||||
|
||||
The rebuild is cache-fast for unchanged layers, the baked-tools check guards
|
||||
the change, and the import/helm-upgrade/verify steps roll the new tag out.
|
||||
The next job's microVM boots warm with the updated toolchain.
|
||||
|
||||
---
|
||||
|
||||
## 7. Verification
|
||||
|
||||
### Baked-tools check (any tag)
|
||||
|
||||
`reload.sh` step 2 already runs this on every build. To re-check an existing tag
|
||||
standalone:
|
||||
|
||||
```bash
|
||||
docker run --rm --entrypoint bash omp-kata-runner:preloaded -lc '
|
||||
set -e
|
||||
for b in gh fd rg magick bun cargo rustc pkg-config zstd; do
|
||||
command -v "$b" >/dev/null || { echo "MISSING: $b"; exit 1; }
|
||||
done
|
||||
echo "tools OK | $(bun --version) | $(rustc --version)"
|
||||
'
|
||||
```
|
||||
|
||||
Confirm the live ARC reference and that the tag exists in the k3s image store
|
||||
(both read-only):
|
||||
|
||||
```bash
|
||||
export KUBECONFIG=/etc/rancher/k3s/k3s.yaml
|
||||
kubectl get autoscalingrunnerset omp-kata -n arc-runners \
|
||||
-o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
|
||||
k3s ctr -n k8s.io images ls | grep omp-kata-runner
|
||||
```
|
||||
|
||||
### Kata microVM boot check
|
||||
|
||||
The baked-tools check above runs the image under plain docker; it does **not**
|
||||
prove the image boots inside a Kata QEMU/KVM microVM. To verify that, launch a
|
||||
throwaway pod with `runtimeClassName: kata-qemu` (the RuntimeClass set up in
|
||||
[02-kata-runtime.md](./02-kata-runtime.md)) and confirm both the deps and the
|
||||
**guest** kernel:
|
||||
|
||||
```bash
|
||||
export KUBECONFIG=/etc/rancher/k3s/k3s.yaml
|
||||
kubectl run preload-verify -n arc-runners --restart=Never \
|
||||
--image=omp-kata-runner:2026-06-15-002621 \
|
||||
--overrides='{"spec":{"runtimeClassName":"kata-qemu"}}' \
|
||||
--command -- bash -lc 'uname -r; bun --version; rustc --version; magick -version | head -1'
|
||||
|
||||
kubectl logs preload-verify -n arc-runners
|
||||
kubectl delete pod preload-verify -n arc-runners
|
||||
```
|
||||
|
||||
Use the tag that is currently live (or any imported tag). `uname -r` should show
|
||||
the Kata **guest** kernel (a 6.x `vmlinux.container` build), not the host kernel -
|
||||
confirming the image really booted in its own microVM - and the `bun`/`rustc`/
|
||||
`magick` lines confirm the baked toolchain is present inside the VM. This `kubectl
|
||||
run` is the only step here that creates a cluster object; delete the pod
|
||||
afterward as shown.
|
||||
|
||||
---
|
||||
|
||||
Continue to [04-arc-and-caching.md](./04-arc-and-caching.md) for how ARC
|
||||
references this image in the runner pod template, wires in the RustFS shared
|
||||
cache, and locks down runner egress.
|
||||
@@ -0,0 +1,649 @@
|
||||
# 04 - ARC runners, shared cache, and egress policy
|
||||
|
||||
This is the last setup step. By now the node runs k3s with the `kata-qemu`
|
||||
RuntimeClass ([02-kata-runtime.md](02-kata-runtime.md)) and the preloaded runner
|
||||
image has been imported into the cluster containerd ([03-runner-image.md](03-runner-image.md)).
|
||||
Here we install **actions-runner-controller (ARC)**, register an ephemeral
|
||||
**scale set** whose pods each boot inside their own Kata microVM, stand up the
|
||||
in-cluster **RustFS (S3)** shared cache, and lock down runner egress with a
|
||||
NetworkPolicy. See [README.md](README.md) for the architecture overview.
|
||||
|
||||
Everything below is read against the live cluster; set the kubeconfig once:
|
||||
|
||||
```bash
|
||||
export KUBECONFIG=/etc/rancher/k3s/k3s.yaml
|
||||
```
|
||||
|
||||
ARC's `gha-runner-scale-set` flavour has three moving parts:
|
||||
|
||||
- **Controller** (`arc` release, ns `arc-systems`) - watches `AutoscalingRunnerSet`
|
||||
custom resources and reconciles them.
|
||||
- **Listener** (one pod per scale set, ns `arc-systems`) - long-polls the GitHub
|
||||
Actions service for jobs targeting the scale set's `runs-on` label.
|
||||
- **Scale set** (`omp-kata` release, ns `arc-runners`) - the `AutoscalingRunnerSet`
|
||||
plus the pod template; the controller turns assigned jobs into ephemeral runner
|
||||
pods here.
|
||||
|
||||
---
|
||||
|
||||
## 1. GitHub App and the `arc-github` secret
|
||||
|
||||
The listener authenticates to GitHub. The durable option is a **GitHub App**
|
||||
(no expiring user token, scoped to exactly the repos you install it on).
|
||||
|
||||
1. Create the App at **GitHub - Settings - Developer settings - GitHub Apps - New GitHub App**.
|
||||
- **Repository permissions**: `Administration: Read and write` (register/remove
|
||||
self-hosted runners) and `Metadata: Read-only` (granted automatically).
|
||||
- No webhook is needed for the scale-set flavour; uncheck **Active** under Webhook.
|
||||
- Generate and download a **private key** (`.pem`).
|
||||
2. **Install** the App on the target repo or org (App page - **Install App** -
|
||||
pick `<OWNER>/<REPO>` or "All repositories"). Note the **App ID** and the
|
||||
**Installation ID** (the trailing number in the install settings URL,
|
||||
`.../installations/<id>`).
|
||||
3. Create the secret in the runners namespace. The three key names below are
|
||||
exactly what the chart reads:
|
||||
|
||||
```bash
|
||||
kubectl create namespace arc-runners
|
||||
|
||||
kubectl -n arc-runners create secret generic arc-github \
|
||||
--from-literal=github_app_id=<GITHUB_APP_ID> \
|
||||
--from-literal=github_app_installation_id=<GITHUB_APP_INSTALLATION_ID> \
|
||||
--from-literal=github_app_private_key=<GITHUB_APP_PRIVATE_KEY>
|
||||
```
|
||||
|
||||
`<GITHUB_APP_PRIVATE_KEY>` is the full PEM body (use `--from-file=github_app_private_key=key.pem`
|
||||
to avoid shell-quoting the multi-line value).
|
||||
|
||||
Verify the live secret carries those three keys (names only - never print values):
|
||||
|
||||
```bash
|
||||
kubectl -n arc-runners get secret arc-github \
|
||||
-o go-template='{{range $k,$v := .data}}{{$k}}{{"\n"}}{{end}}'
|
||||
# github_app_id
|
||||
# github_app_installation_id
|
||||
# github_app_private_key
|
||||
```
|
||||
|
||||
**Token alternative.** ARC also accepts a single-key secret with a classic PAT
|
||||
(scope `repo`) or a fine-grained PAT (`Administration: RW` + `Metadata: R`):
|
||||
|
||||
```bash
|
||||
kubectl -n arc-runners create secret generic arc-github \
|
||||
--from-literal=github_token=<GITHUB_PAT>
|
||||
```
|
||||
|
||||
The App is preferred: it does not expire, it is scoped per-installation, and one
|
||||
installation covers every repo you grant it (useful for [adding another repo](#7-operate)).
|
||||
Whichever you choose, the `githubConfigSecret` value in step 3's chart points at
|
||||
this secret by name.
|
||||
|
||||
---
|
||||
|
||||
## 2. Install ARC (controller + scale set)
|
||||
|
||||
ARC ships as OCI Helm charts; no `helm repo add` is required. Both the controller
|
||||
and the scale set are pinned to the same chart version, **0.14.2** (matches the
|
||||
live `helm list -A`).
|
||||
|
||||
**Controller** (installed with chart defaults - `helm get values arc` is empty):
|
||||
|
||||
```bash
|
||||
helm install arc \
|
||||
--namespace arc-systems --create-namespace \
|
||||
--version 0.14.2 \
|
||||
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set-controller
|
||||
```
|
||||
|
||||
**Scale set** (`omp-kata`), using the values file from step 3:
|
||||
|
||||
```bash
|
||||
helm install omp-kata \
|
||||
--namespace arc-runners --create-namespace \
|
||||
--version 0.14.2 \
|
||||
-f arc-omp-values.yaml \
|
||||
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set
|
||||
```
|
||||
|
||||
Confirm both releases and the running controller image:
|
||||
|
||||
```bash
|
||||
helm list -A
|
||||
# arc arc-systems deployed gha-runner-scale-set-controller-0.14.2 0.14.2
|
||||
# omp-kata arc-runners deployed gha-runner-scale-set-0.14.2 0.14.2
|
||||
|
||||
kubectl -n arc-systems get deploy arc-gha-rs-controller \
|
||||
-o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
|
||||
# ghcr.io/actions/gha-runner-scale-set-controller:0.14.2
|
||||
```
|
||||
|
||||
Within a few seconds the controller spawns the listener in `arc-systems`:
|
||||
|
||||
```bash
|
||||
kubectl -n arc-systems get pods
|
||||
# arc-gha-rs-controller-xxxxxxxxxx-xxxxx 1/1 Running
|
||||
# omp-kata-<hash>-listener 1/1 Running
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Scale-set values (`arc-omp-values.yaml`)
|
||||
|
||||
This is the live `arc-omp-values.yaml` verbatim, with only the repo owner/name in
|
||||
`githubConfigUrl` redacted:
|
||||
|
||||
```yaml
|
||||
githubConfigUrl: "https://github.com/<OWNER>/<REPO>"
|
||||
githubConfigSecret: arc-github
|
||||
runnerScaleSetName: omp-kata
|
||||
minRunners: 0
|
||||
maxRunners: 10
|
||||
# none: each job runs inside the runner container, which itself lives in a Kata microVM
|
||||
containerMode:
|
||||
type: ""
|
||||
template:
|
||||
spec:
|
||||
runtimeClassName: kata-qemu # <-- every runner pod boots its own KVM microVM
|
||||
containers:
|
||||
- name: runner
|
||||
# Preloaded image: stock ghcr.io/actions/actions-runner + CI deps baked in
|
||||
# (apt cairo/pango/jpeg/gif/rsvg stack, fd/ripgrep/imagemagick, bun, rust
|
||||
# nightly + clippy/rustfmt + arm64/msvc targets). Built + imported locally;
|
||||
# see /root/omp-kata-runner-image/. IfNotPresent uses the local image.
|
||||
image: omp-kata-runner:2026-06-15-002621
|
||||
imagePullPolicy: IfNotPresent
|
||||
command: ["/home/runner/run.sh"]
|
||||
# Shared sccache backend (in-cluster RustFS S3). Exposes SCCACHE_BUCKET/
|
||||
# ENDPOINT/REGION/USE_SSL + AWS creds to every job; CI flips RUSTC_WRAPPER
|
||||
# on for rust builds only. GitHub-hosted runners lack this env and keep the
|
||||
# GHA cache backend. See /root/sccache-rustfs/.
|
||||
envFrom:
|
||||
- secretRef:
|
||||
name: sccache-s3
|
||||
resources:
|
||||
requests:
|
||||
cpu: "2"
|
||||
memory: "4Gi"
|
||||
limits:
|
||||
cpu: "6"
|
||||
memory: "12Gi"
|
||||
```
|
||||
|
||||
Field by field:
|
||||
|
||||
- **`githubConfigUrl`** - the repo (or org) the scale set serves. Jobs reach it
|
||||
with `runs-on: omp-kata`.
|
||||
- **`githubConfigSecret: arc-github`** - the auth secret from [step 1](#1-github-app-and-the-arc-github-secret).
|
||||
- **`runnerScaleSetName: omp-kata`** - the runner label. This is the string that
|
||||
goes in a workflow's `runs-on:`.
|
||||
- **`minRunners: 0` / `maxRunners: 10`** - **scale-to-zero**. With no queued jobs
|
||||
there are zero runner pods (and zero microVMs) consuming the node; the listener
|
||||
scales up to ten concurrent runners on demand. (The older ops notes capped this
|
||||
at 3; the live value is 10.)
|
||||
- **`containerMode.type: ""`** - **none**. The default chart offers `dind`
|
||||
(Docker-in-Docker sidecar) or `kubernetes` mode for job-container isolation;
|
||||
both are unnecessary here because the *whole runner pod* is already isolated in
|
||||
a microVM. The job runs directly in the runner container - no privileged dind
|
||||
sidecar, no extra attack surface.
|
||||
- **`template.spec.runtimeClassName: kata-qemu`** - the critical line. It binds
|
||||
the pod to the Kata QEMU runtime ([02-kata-runtime.md](02-kata-runtime.md)), so
|
||||
every runner boots its own KVM microVM with a guest kernel distinct from the host.
|
||||
- **`image` / `imagePullPolicy: IfNotPresent`** - the locally built, dependency-baked
|
||||
runner image ([03-runner-image.md](03-runner-image.md)). `IfNotPresent` uses the
|
||||
copy already imported into cluster containerd; there is no registry. Bump the tag
|
||||
here when you rebuild the image (see [Operate](#7-operate)).
|
||||
- **`command: ["/home/runner/run.sh"]`** - the stock actions-runner entrypoint;
|
||||
overridden explicitly because the custom image keeps the upstream layout.
|
||||
- **`envFrom.secretRef.name: sccache-s3`** - injects the shared-cache S3
|
||||
configuration into every job's environment ([step 5](#5-shared-cache-rustfs-s3)).
|
||||
- **`resources`** - requests `2` CPU / `4Gi`, limits `6` CPU / `12Gi`. Kata reads
|
||||
these and sizes the guest accordingly: the VM boots at the runtime defaults
|
||||
(`default_vcpus: 1`, `default_memory: 2048`) and **hotplugs** vCPUs and RAM to
|
||||
cover the pod's containers, with `default_maxvcpus: 0` allowing up to all host
|
||||
CPUs. Effectively the **requests are the guaranteed VM size** and the **limits
|
||||
are the hotplug ceiling**. See [02-kata-runtime.md](02-kata-runtime.md) for the
|
||||
hotplug mechanics.
|
||||
|
||||
---
|
||||
|
||||
## 4. Job lifecycle and the no-permission ServiceAccount
|
||||
|
||||
One job runs in one fresh microVM that is destroyed afterward:
|
||||
|
||||
1. The **listener** (ns `arc-systems`) long-polls the GitHub Actions service for
|
||||
jobs whose `runs-on` matches `omp-kata`.
|
||||
2. When jobs are assigned, the controller reconciles the `AutoscalingRunnerSet`
|
||||
and creates an **`EphemeralRunnerSet`** sized to the demand (bounded by
|
||||
`minRunners`/`maxRunners`).
|
||||
3. Each replica becomes an **ephemeral runner pod** registered **just-in-time
|
||||
(JIT)** with GitHub - a per-runner registration secret is minted, not a
|
||||
long-lived token.
|
||||
4. Because the pod's `runtimeClassName` is `kata-qemu`, it **boots a microVM**,
|
||||
pulls the one assigned job, runs it, and exits.
|
||||
5. ARC **deletes the pod** (and its microVM); a clean VM is created for the next
|
||||
job. There is no VM templating - state never leaks between jobs.
|
||||
|
||||
Observe the chain live:
|
||||
|
||||
```bash
|
||||
kubectl -n arc-runners get autoscalingrunnerset omp-kata
|
||||
kubectl -n arc-runners get ephemeralrunnerset
|
||||
kubectl -n arc-runners get pods -o wide # one pod per in-flight job; empty when idle
|
||||
```
|
||||
|
||||
**No-permission ServiceAccount.** The scale-set chart runs every runner pod under
|
||||
a ServiceAccount with no RBAC bindings:
|
||||
|
||||
```bash
|
||||
kubectl -n arc-runners get sa
|
||||
# default
|
||||
# omp-kata-gha-rs-no-permission
|
||||
```
|
||||
|
||||
Job code therefore has no Kubernetes API rights - it cannot read secrets, list
|
||||
pods, or touch the cluster, even though it executes inside the cluster. Combined
|
||||
with microVM isolation and the egress policy ([step 6](#6-runner-egress-lockdown)),
|
||||
a compromised job is boxed into a throwaway VM with no cluster reach.
|
||||
|
||||
---
|
||||
|
||||
## 5. Shared cache (RustFS S3)
|
||||
|
||||
GitHub's hosted cache backend is only reachable over the node's NAT egress, so on
|
||||
a busy matrix (many concurrent jobs) it becomes the bottleneck. Instead an
|
||||
**S3-compatible object store, RustFS, runs inside the cluster** and serves the
|
||||
cache at LAN speed over `rustfs.sccache.svc.cluster.local:9000`.
|
||||
|
||||
### 5a. Deploy RustFS
|
||||
|
||||
The store lives in its own `sccache` namespace: a `local-path` PVC for durability,
|
||||
a single-replica Deployment, and a ClusterIP Service. This is `rustfs.yaml`
|
||||
verbatim (no secrets inline - credentials come from a separate secret):
|
||||
|
||||
```yaml
|
||||
apiVersion: v1
|
||||
kind: Namespace
|
||||
metadata:
|
||||
name: sccache
|
||||
labels:
|
||||
kubernetes.io/metadata.name: sccache
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: rustfs-data
|
||||
namespace: sccache
|
||||
spec:
|
||||
accessModes: [ReadWriteOnce]
|
||||
storageClassName: local-path
|
||||
resources:
|
||||
requests:
|
||||
storage: 100Gi
|
||||
---
|
||||
apiVersion: apps/v1
|
||||
kind: Deployment
|
||||
metadata:
|
||||
name: rustfs
|
||||
namespace: sccache
|
||||
labels: { app: rustfs }
|
||||
spec:
|
||||
replicas: 1
|
||||
strategy: { type: Recreate }
|
||||
selector:
|
||||
matchLabels: { app: rustfs }
|
||||
template:
|
||||
metadata:
|
||||
labels: { app: rustfs }
|
||||
spec:
|
||||
containers:
|
||||
- name: rustfs
|
||||
image: rustfs/rustfs:latest
|
||||
imagePullPolicy: IfNotPresent
|
||||
env:
|
||||
- name: RUSTFS_ACCESS_KEY
|
||||
valueFrom: { secretKeyRef: { name: rustfs-creds, key: RUSTFS_ACCESS_KEY } }
|
||||
- name: RUSTFS_SECRET_KEY
|
||||
valueFrom: { secretKeyRef: { name: rustfs-creds, key: RUSTFS_SECRET_KEY } }
|
||||
- name: RUSTFS_VOLUMES
|
||||
value: "/data"
|
||||
- name: RUSTFS_ADDRESS
|
||||
value: ":9000"
|
||||
- name: RUSTFS_CONSOLE_ENABLE
|
||||
value: "false"
|
||||
- name: RUSTFS_OBS_LOG_DIRECTORY
|
||||
value: "/logs"
|
||||
ports:
|
||||
- { name: s3, containerPort: 9000 }
|
||||
volumeMounts:
|
||||
- { name: data, mountPath: /data }
|
||||
- { name: logs, mountPath: /logs }
|
||||
readinessProbe:
|
||||
tcpSocket: { port: 9000 }
|
||||
initialDelaySeconds: 5
|
||||
periodSeconds: 5
|
||||
livenessProbe:
|
||||
tcpSocket: { port: 9000 }
|
||||
initialDelaySeconds: 15
|
||||
periodSeconds: 20
|
||||
resources:
|
||||
requests: { cpu: "200m", memory: "256Mi" }
|
||||
limits: { cpu: "2", memory: "2Gi" }
|
||||
volumes:
|
||||
- name: data
|
||||
persistentVolumeClaim: { claimName: rustfs-data }
|
||||
- name: logs
|
||||
emptyDir: {}
|
||||
---
|
||||
apiVersion: v1
|
||||
kind: Service
|
||||
metadata:
|
||||
name: rustfs
|
||||
namespace: sccache
|
||||
spec:
|
||||
selector: { app: rustfs }
|
||||
ports:
|
||||
- { name: s3, port: 9000, targetPort: 9000, protocol: TCP }
|
||||
```
|
||||
|
||||
Notes:
|
||||
|
||||
- **`strategy: Recreate`** with a single replica and an RWO `local-path` PVC: the
|
||||
data is node-local and only one pod ever mounts it.
|
||||
- **`RUSTFS_CONSOLE_ENABLE: "false"`** - only the S3 API on `:9000` is exposed;
|
||||
no admin console.
|
||||
- The `kubernetes.io/metadata.name: sccache` namespace label is what the egress
|
||||
NetworkPolicy's `namespaceSelector` matches ([step 6](#6-runner-egress-lockdown)).
|
||||
|
||||
The RustFS pod credentials come from a two-key secret in the `sccache` namespace
|
||||
(values are the object-store root credentials - use placeholders):
|
||||
|
||||
```bash
|
||||
kubectl -n sccache create secret generic rustfs-creds \
|
||||
--from-literal=RUSTFS_ACCESS_KEY=<S3_ACCESS_KEY> \
|
||||
--from-literal=RUSTFS_SECRET_KEY=<S3_SECRET_KEY>
|
||||
```
|
||||
|
||||
Apply and verify:
|
||||
|
||||
```bash
|
||||
kubectl apply -f rustfs.yaml
|
||||
kubectl -n sccache get deploy,svc,pvc
|
||||
# deployment.apps/rustfs 1/1
|
||||
# service/rustfs ClusterIP 10.43.x.x 9000/TCP
|
||||
# persistentvolumeclaim/rustfs-data Bound 100Gi local-path
|
||||
```
|
||||
|
||||
Create the `sccache` bucket once (any S3 client - e.g. the `aws` CLI or `mc`
|
||||
pointed at the endpoint with the root creds): `mb s3://sccache`.
|
||||
|
||||
### 5b. The `sccache-s3` secret (injected into every runner)
|
||||
|
||||
Every runner pod gets the cache configuration via `envFrom` ([step 3](#3-scale-set-values-arc-omp-valuesyaml)).
|
||||
The secret lives in `arc-runners` (the runners' namespace) and has six keys:
|
||||
|
||||
```bash
|
||||
kubectl -n arc-runners get secret sccache-s3 \
|
||||
-o go-template='{{range $k,$v := .data}}{{$k}}{{"\n"}}{{end}}'
|
||||
# AWS_ACCESS_KEY_ID
|
||||
# AWS_SECRET_ACCESS_KEY
|
||||
# SCCACHE_BUCKET
|
||||
# SCCACHE_ENDPOINT
|
||||
# SCCACHE_REGION
|
||||
# SCCACHE_S3_USE_SSL
|
||||
```
|
||||
|
||||
Recreate it (the two credential values must equal the `rustfs-creds` above; the
|
||||
rest are non-sensitive cluster-local config):
|
||||
|
||||
```bash
|
||||
kubectl -n arc-runners create secret generic sccache-s3 \
|
||||
--from-literal=AWS_ACCESS_KEY_ID=<S3_ACCESS_KEY> \
|
||||
--from-literal=AWS_SECRET_ACCESS_KEY=<S3_SECRET_KEY> \
|
||||
--from-literal=SCCACHE_BUCKET=sccache \
|
||||
--from-literal=SCCACHE_ENDPOINT=rustfs.sccache.svc.cluster.local:9000 \
|
||||
--from-literal=SCCACHE_REGION=us-east-1 \
|
||||
--from-literal=SCCACHE_S3_USE_SSL=false
|
||||
```
|
||||
|
||||
- `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` - the **only sensitive entries**;
|
||||
RustFS's root credentials (S3 SigV4 auth).
|
||||
- `SCCACHE_BUCKET: sccache` - bucket name.
|
||||
- `SCCACHE_ENDPOINT` - the in-cluster Service DNS + port.
|
||||
- `SCCACHE_REGION: us-east-1` - arbitrary region label SigV4 requires.
|
||||
- `SCCACHE_S3_USE_SSL: false` - the endpoint is plain HTTP on the cluster network.
|
||||
|
||||
### 5c. The two consumers
|
||||
|
||||
The presence of `$SCCACHE_BUCKET` in the environment is the repo's single signal
|
||||
for "am I on the self-hosted infra?". Both consumers branch on it and fall back
|
||||
to GitHub-hosted cache backends off-infra (GitHub-hosted macOS/arm runners never
|
||||
get the secret and so cannot reach the private RustFS).
|
||||
|
||||
**(a) sccache for Rust** - [`.github/actions/build-native`](../../.github/actions/build-native/action.yml).
|
||||
It installs `sccache`, then sets `RUSTC_WRAPPER=sccache` and `CARGO_INCREMENTAL=0`
|
||||
(sccache silently no-ops with incremental enabled). The backend is conditional:
|
||||
|
||||
- `$SCCACHE_BUCKET` set - sccache reads `SCCACHE_BUCKET/ENDPOINT/REGION` and the
|
||||
AWS creds straight from the inherited pod env and uses the **shared S3 (RustFS)**.
|
||||
- otherwise - it exports `SCCACHE_GHA_ENABLED=true` and uses the **GitHub Actions
|
||||
cache**.
|
||||
|
||||
`Swatinem/rust-cache` still caches `target/` on top; sccache fills the gaps when
|
||||
`target/` is cold.
|
||||
|
||||
**(b) bun dependency cache** - [`.github/actions/bun-install`](../../.github/actions/bun-install/action.yml),
|
||||
a composite action wrapping `bun install --frozen-lockfile`. A "Detect cache
|
||||
backend" step checks `$SCCACHE_BUCKET` + `$AWS_ACCESS_KEY_ID`:
|
||||
|
||||
- on-infra - it runs `rustfs-cache.sh restore` before install and `... save` after;
|
||||
- off-infra - it uses stock `actions/cache@v4` for the bun store.
|
||||
|
||||
`rustfs-cache.sh` talks to RustFS directly with `curl --aws-sigv4` (S3 SigV4), no
|
||||
SDK. It keys two objects per lockfile under the `bun-cache/` prefix of the
|
||||
`sccache` bucket, derived from `sha256(bun.lock)`:
|
||||
|
||||
- `store-<os>-<lockhash>` - the bun global package store (`~/.bun/install/cache`),
|
||||
plus a rolling `store-<os>-latest` alias so a changed lockfile still warm-starts
|
||||
from the previous store and `bun install` fetches only the delta.
|
||||
- `nm-<os>-<lockhash>` - the installed `node_modules` trees (repo root, every
|
||||
`packages/*`, and `python/robomp/web`).
|
||||
|
||||
On **restore**, a `node_modules` hit short-circuits everything - the subsequent
|
||||
`bun install --frozen-lockfile` is a no-op, so the store is neither fetched nor
|
||||
saved. Archives are multi-threaded **zstd** (`.tzst`, baked into the runner image)
|
||||
with a **gzip** (`.tgz`) fallback so the action still works on an older image; the
|
||||
suffix records the codec and restore only inflates what the host can decompress.
|
||||
On **save**, it writes the store/`node_modules` objects only when this exact
|
||||
lockfile has none yet (`s3_exists` check), avoiding redundant uploads.
|
||||
|
||||
This caching is why RustFS sits inside the egress allow-list on `tcp/9000`
|
||||
([step 6](#6-runner-egress-lockdown)).
|
||||
|
||||
---
|
||||
|
||||
## 6. Runner egress lockdown
|
||||
|
||||
Runner pods reach the public internet (GitHub, package registries, crates.io,
|
||||
npm) but must **not** reach the host's own services, the LAN, the tailnet, or
|
||||
arbitrary cluster workloads. A single NetworkPolicy in `arc-runners` enforces
|
||||
this. Because the pod template sets no special labels, the policy uses
|
||||
`podSelector: {}` to cover **every** pod in the namespace.
|
||||
|
||||
> k3s ships a built-in NetworkPolicy controller (kube-router based) that enforces
|
||||
> policies even though the CNI is Flannel - so this policy actually takes effect.
|
||||
> Do not start k3s with `--disable-network-policy` ([01-host-and-cluster.md](01-host-and-cluster.md)),
|
||||
> or the lockdown silently becomes a no-op.
|
||||
|
||||
Live spec (captured with `kubectl get networkpolicy -n arc-runners runner-egress-lockdown -o yaml`;
|
||||
server-managed metadata omitted, host public IP redacted):
|
||||
|
||||
```yaml
|
||||
apiVersion: networking.k8s.io/v1
|
||||
kind: NetworkPolicy
|
||||
metadata:
|
||||
name: runner-egress-lockdown
|
||||
namespace: arc-runners
|
||||
spec:
|
||||
podSelector: {}
|
||||
policyTypes:
|
||||
- Ingress
|
||||
- Egress
|
||||
egress:
|
||||
# 1. Cluster DNS only (CoreDNS + kube-system).
|
||||
- to:
|
||||
- ipBlock:
|
||||
cidr: 10.43.0.10/32
|
||||
- namespaceSelector:
|
||||
matchLabels:
|
||||
kubernetes.io/metadata.name: kube-system
|
||||
ports:
|
||||
- port: 53
|
||||
protocol: UDP
|
||||
- port: 53
|
||||
protocol: TCP
|
||||
# 2. Public internet, MINUS all private/infra ranges and the host's own public IP.
|
||||
- to:
|
||||
- ipBlock:
|
||||
cidr: 0.0.0.0/0
|
||||
except:
|
||||
- 10.0.0.0/8
|
||||
- 172.16.0.0/12
|
||||
- 192.168.0.0/16
|
||||
- 169.254.0.0/16
|
||||
- 100.64.0.0/10
|
||||
- <PUBLIC_IP>/32
|
||||
# 3. RustFS shared cache (S3) over the cluster network.
|
||||
- to:
|
||||
- ipBlock:
|
||||
cidr: 10.43.0.0/16
|
||||
- namespaceSelector:
|
||||
matchLabels:
|
||||
kubernetes.io/metadata.name: sccache
|
||||
ports:
|
||||
- port: 9000
|
||||
protocol: TCP
|
||||
```
|
||||
|
||||
The allow-list, rule by rule:
|
||||
|
||||
- **Rule 1 - DNS.** UDP/TCP 53 to CoreDNS (`10.43.0.10/32`) and the `kube-system`
|
||||
namespace. Without this, name resolution breaks and rule 2 is useless.
|
||||
- **Rule 2 - public internet only.** `0.0.0.0/0` with an `except` list that
|
||||
carves out every range a job has no business reaching: RFC1918 private space
|
||||
(`10/8`, `172.16/12`, `192.168/16`), link-local (`169.254/16`), the CGNAT range
|
||||
used by the **tailnet** (`100.64.0.0/10`), and the **host's own public IP**
|
||||
(`<PUBLIC_IP>/32`). Note `10.0.0.0/8` covers the pod CIDR (`10.42.0.0/16`) and
|
||||
service CIDR (`10.43.0.0/16`), so this rule alone gives a job **zero** in-cluster
|
||||
reach - rules 1 and 3 punch the only two holes the job legitimately needs.
|
||||
- **Rule 3 - RustFS cache.** TCP 9000 to the service CIDR (`10.43.0.0/16`) and the
|
||||
`sccache` namespace - the shared cache from [step 5](#5-shared-cache-rustfs-s3).
|
||||
- **Ingress.** `policyTypes` lists `Ingress` but no ingress rule is defined, which
|
||||
is a **default-deny**: nothing can open a connection *into* a runner pod.
|
||||
|
||||
Egress that survives rule 2 leaves the node via the host's firewalld masquerade
|
||||
(SNAT to the public IP) over the default interface - see
|
||||
[01-host-and-cluster.md](01-host-and-cluster.md) for the host firewall side.
|
||||
|
||||
### Security model
|
||||
|
||||
- **Kernel isolation.** Each job runs in a Kata microVM with its own guest kernel
|
||||
(6.x), separate from the host kernel (7.0.x) - a kernel exploit hits a throwaway
|
||||
VM, not the host. See [02-kata-runtime.md](02-kata-runtime.md).
|
||||
- **No cluster rights.** Jobs run under `omp-kata-gha-rs-no-permission` with no
|
||||
RBAC ([step 4](#4-job-lifecycle-and-the-no-permission-serviceaccount)).
|
||||
- **Constrained network.** The policy above blocks the host, LAN, tailnet, and
|
||||
arbitrary cluster pods; only DNS, the public internet, and RustFS are reachable.
|
||||
- **Ephemeral.** One job per VM, destroyed afterward - no state, secret, or
|
||||
artifact survives into the next job.
|
||||
- **Public-repo recommendation.** For a public repo, require approval for fork
|
||||
PRs so untrusted code cannot auto-run on the infra: **repo - Settings - Actions
|
||||
- General - Fork pull request workflows from outside collaborators - Require
|
||||
approval for all outside collaborators**.
|
||||
|
||||
---
|
||||
|
||||
## 7. Operate
|
||||
|
||||
```bash
|
||||
export KUBECONFIG=/etc/rancher/k3s/k3s.yaml
|
||||
```
|
||||
|
||||
**Status / scale**
|
||||
|
||||
```bash
|
||||
kubectl -n arc-runners get autoscalingrunnerset omp-kata # min/max/current runners
|
||||
kubectl -n arc-runners get ephemeralrunnerset # desired vs current replicas
|
||||
kubectl -n arc-runners get pods -o wide # live runner VMs (empty when idle)
|
||||
```
|
||||
|
||||
**Logs**
|
||||
|
||||
```bash
|
||||
# Listener (job dispatch / scaling decisions)
|
||||
kubectl -n arc-systems logs -l app.kubernetes.io/component=runner-scale-set-listener -f
|
||||
# Controller (reconciliation)
|
||||
kubectl -n arc-systems logs deploy/arc-gha-rs-controller -f
|
||||
# A specific runner / its job
|
||||
kubectl -n arc-runners logs <runner-pod>
|
||||
```
|
||||
|
||||
**Verify the cache is being used.** A warm job logs `bun cache: ... HIT` and
|
||||
`sccache backend: shared S3 (sccache @ rustfs.sccache.svc.cluster.local:9000)` in
|
||||
its step output. To inspect objects directly, point any S3 client at the endpoint
|
||||
(with the RustFS root creds) and list `s3://sccache/bun-cache/`.
|
||||
|
||||
**Resize a job's VM** - edit the `resources` block in `arc-omp-values.yaml`
|
||||
([step 3](#3-scale-set-values-arc-omp-valuesyaml); requests = guaranteed VM size,
|
||||
limits = hotplug ceiling) and roll out:
|
||||
|
||||
```bash
|
||||
helm upgrade omp-kata \
|
||||
--namespace arc-runners --version 0.14.2 \
|
||||
-f arc-omp-values.yaml \
|
||||
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set
|
||||
```
|
||||
|
||||
**Change scale-to-zero bounds** - edit `minRunners` / `maxRunners` in the same
|
||||
file and `helm upgrade` as above. (Keep `maxRunners` within the node's CPU/RAM
|
||||
budget: each runner can hotplug up to its `limits`.)
|
||||
|
||||
**Update the runner image** - bump `template.spec.containers[0].image` to the new
|
||||
tag, then `helm upgrade` as above; confirm with:
|
||||
|
||||
```bash
|
||||
kubectl -n arc-runners get autoscalingrunnerset omp-kata \
|
||||
-o jsonpath='{.spec.template.spec.containers[0].image}{"\n"}'
|
||||
```
|
||||
|
||||
See [03-runner-image.md](03-runner-image.md) for building and importing the image.
|
||||
|
||||
**Add another repo.** Because the GitHub App installation can cover multiple repos,
|
||||
reuse the same `arc-github` secret and install a second scale set with its own
|
||||
`githubConfigUrl`, `runnerScaleSetName` (the new `runs-on:` label), and release
|
||||
name:
|
||||
|
||||
```bash
|
||||
helm install <release> \
|
||||
--namespace arc-runners --version 0.14.2 \
|
||||
--set githubConfigUrl=https://github.com/<OWNER>/<OTHER_REPO> \
|
||||
--set githubConfigSecret=arc-github \
|
||||
--set runnerScaleSetName=<other-repo>-kata \
|
||||
-f arc-omp-values.yaml \
|
||||
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set
|
||||
```
|
||||
|
||||
Jobs in the other repo then target `runs-on: <other-repo>-kata`. (On this host a
|
||||
convenience wrapper, `omp-add-repo-runner <OWNER>/<REPO> [label]`, performs exactly
|
||||
this install.)
|
||||
|
||||
**Uninstall** (leaves k3s/Kata in place):
|
||||
|
||||
```bash
|
||||
helm uninstall omp-kata -n arc-runners
|
||||
helm uninstall arc -n arc-systems
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
**Previous:** [03-runner-image.md](03-runner-image.md) - the preloaded runner image.
|
||||
**Overview:** [README.md](README.md) - architecture and the full doc set.
|
||||
@@ -0,0 +1,110 @@
|
||||
# Self-hosted Kata CI
|
||||
|
||||
This is a self-hosted GitHub Actions setup where **every CI job runs inside its own throwaway Kata Containers QEMU/KVM microVM**. A single bare-metal Linux host runs a one-node [k3s](https://k3s.io) cluster; [actions-runner-controller (ARC)](https://github.com/actions/actions-runner-controller) watches GitHub for queued jobs and, for each one, creates a just-in-time ephemeral runner pod that boots a fresh microVM (its own guest kernel, isolated from the host), runs exactly one job, and is then destroyed. Runners share an in-cluster **RustFS** (S3-compatible) object store for `sccache` and Bun dependency caching, and egress the public internet through host NAT under a restrictive NetworkPolicy. The result is hardware-isolated, scale-to-zero CI on hardware you control.
|
||||
|
||||
These docs are written as a **from-scratch setup guide**: read this overview first, then follow the numbered guides in order to reproduce the system on your own host.
|
||||
|
||||
## Architecture
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
GH["GitHub Actions<br/>repo <OWNER>/<REPO> + GitHub App"]
|
||||
|
||||
subgraph HOST["CI host (<CI_HOST>) — CentOS Stream 10, KVM/bare-metal"]
|
||||
NAT["firewalld masquerade<br/>egress NAT → public IP <PUBLIC_IP>"]
|
||||
subgraph K3S["single-node k3s v1.35.5 (own containerd v2)"]
|
||||
subgraph SYS["ns: arc-systems"]
|
||||
LIS["Runner scale-set listener<br/>long-polls GitHub"]
|
||||
CTRL["ARC controller 0.14.2<br/>scales EphemeralRunnerSet"]
|
||||
end
|
||||
subgraph RUN["ns: arc-runners"]
|
||||
NP["NetworkPolicy<br/>runner-egress-lockdown"]
|
||||
POD["Ephemeral runner pod (JIT)<br/>runtimeClassName: kata-qemu"]
|
||||
subgraph VM["Kata QEMU/KVM microVM — separate guest kernel"]
|
||||
JOB["actions/runner + one job's steps"]
|
||||
end
|
||||
end
|
||||
subgraph CACHE["ns: sccache"]
|
||||
RUSTFS["RustFS (S3)<br/>svc rustfs:9000 · PVC rustfs-data 100Gi"]
|
||||
end
|
||||
SEC["Secret sccache-s3<br/>S3 creds + endpoint"]
|
||||
end
|
||||
end
|
||||
|
||||
GH <-->|"long-poll / JIT registration"| LIS
|
||||
LIS --> CTRL
|
||||
CTRL -->|"creates 1 pod per job"| POD
|
||||
POD --> VM
|
||||
SEC -.->|"envFrom"| POD
|
||||
NP -.->|"filters egress"| POD
|
||||
JOB -->|"sccache + Bun cache (S3 SigV4)"| RUSTFS
|
||||
POD -->|"allowed egress"| NAT
|
||||
NAT -->|"checkout / API / internet"| GH
|
||||
```
|
||||
|
||||
Key properties baked into this design:
|
||||
|
||||
- **One job = one VM.** Runner pods are ephemeral and JIT-registered; there is no VM templating or pooling, so a job never inherits state from a previous job.
|
||||
- **Scale-to-zero.** `minRunners: 0` / `maxRunners: 10` — when no jobs are queued, zero runner pods (and zero microVMs) exist.
|
||||
- **Host-kernel isolation.** Jobs see the microVM's guest kernel, not the host kernel, so a kernel exploit in a job does not reach the host.
|
||||
- **No external registry.** The runner image is built on the host and imported straight into k3s' containerd.
|
||||
- **Shared, in-cluster cache.** `sccache` and the Bun dependency cache both target RustFS over the cluster network; nothing cache-related leaves the host.
|
||||
|
||||
## End-to-end job lifecycle
|
||||
|
||||
1. A workflow job targeting the self-hosted label (`runs-on:`) is **queued** on GitHub.
|
||||
2. The **scale-set listener** in `arc-systems` is long-polling the GitHub Actions service and receives the job-assignment message.
|
||||
3. The listener signals demand to the **ARC controller**, which scales the **EphemeralRunnerSet** up by one.
|
||||
4. The controller creates a single **JIT-registered ephemeral runner pod** in `arc-runners`, with `runtimeClassName: kata-qemu` and the `sccache-s3` secret injected via `envFrom`.
|
||||
5. containerd hands the pod to the Kata shim, which **boots a fresh QEMU/KVM microVM** (own guest kernel; the container rootfs is shared in over virtio-fs). No templating — every job gets a clean VM.
|
||||
6. The runner agent inside the microVM **registers just-in-time and picks up exactly one job**. Steps run isolated from the host, using RustFS over S3 for `sccache`/Bun caching and NAT egress for the public internet, all constrained by the `runner-egress-lockdown` NetworkPolicy.
|
||||
7. The job finishes; the ephemeral runner **deregisters and the pod (and its microVM) is destroyed** — never reused.
|
||||
8. When no jobs remain queued, the EphemeralRunnerSet **scales back to zero**, leaving no idle runners or VMs.
|
||||
|
||||
## Component map (bill of materials)
|
||||
|
||||
| Component | What it is | Version | Documented in |
|
||||
| --- | --- | --- | --- |
|
||||
| Host + k3s cluster | Bare-metal CentOS Stream 10 node running single-node k3s (own containerd v2, Flannel CNI; Traefik + servicelb disabled so host nginx keeps :80/:443); firewalld provides NAT egress | k3s `v1.35.5+k3s1` | [01-host-and-cluster.md](01-host-and-cluster.md) |
|
||||
| Kata Containers runtime | QEMU/KVM microVM runtime: containerd drop-in registering `kata-qemu` + the `kata-qemu` RuntimeClass | Kata `3.31.0` | [02-kata-runtime.md](02-kata-runtime.md) |
|
||||
| Preloaded runner image | Custom `actions/runner` image (build toolchain, Bun, Rust nightly + cross targets, native-build deps) built on the host and imported into k3s containerd — no registry | local dated tag | [03-runner-image.md](03-runner-image.md) |
|
||||
| ARC (runner scale set) | actions-runner-controller, `gha-runner-scale-set` flavor: controller in `arc-systems`, one scale set + listener, GitHub App auth | ARC `0.14.2` | [04-arc-and-caching.md](04-arc-and-caching.md) |
|
||||
| RustFS shared cache | In-cluster S3-compatible store (`svc rustfs:9000`, 100Gi PVC) backing `sccache` and the Bun cache, plus the `sccache-s3` secret and the egress NetworkPolicy | in-cluster service | [04-arc-and-caching.md](04-arc-and-caching.md) |
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Before starting, you need:
|
||||
|
||||
- **A Linux host with hardware virtualization.** Intel VT-x or AMD-V enabled, KVM available (`/dev/kvm` present and accessible). Bare metal is simplest; on a VM you need working nested virtualization. The reference host is 32 vCPU / 125 GiB RAM — size to roughly `maxRunners × per-job resources` plus cluster overhead.
|
||||
- **Root (or full sudo)** on that host: you will install k3s, Kata, kernel modules, firewall rules, and a container image.
|
||||
- **A public-ish egress path.** The host must reach GitHub; runner microVMs NAT out through the host's public IP. No inbound ports are required for the runners (the listener uses outbound long-poll).
|
||||
- **A GitHub repository** to attach runners to, and a **GitHub App** (recommended) or PAT installed on it with permissions to manage self-hosted runners. You will record the App ID, installation ID, and private key as a Kubernetes secret.
|
||||
- **CLI tooling on the host:** `kubectl` and `helm` (k3s bundles a kubectl), plus Docker/buildkit for building the runner image (see [03-runner-image.md](03-runner-image.md)).
|
||||
|
||||
## Redaction & placeholders
|
||||
|
||||
The configs in this doc set are copied verbatim from the live host and then redacted. Wherever you see one of these tokens, substitute your own value:
|
||||
|
||||
| Placeholder | Substitute with |
|
||||
| --- | --- |
|
||||
| `<CI_HOST>` | Your CI host's hostname / SSH target |
|
||||
| `<PUBLIC_IP>` | The host's public IPv4 address |
|
||||
| `<TAILNET_IP>` | Your Tailscale/tailnet admin IP(s) (the generic CGNAT range `100.64.0.0/10` is kept as-is) |
|
||||
| `<OWNER>/<REPO>` | Your GitHub repository owner and name |
|
||||
| `<GITHUB_APP_ID>` | Your GitHub App ID |
|
||||
| `<GITHUB_APP_INSTALLATION_ID>` | Your GitHub App installation ID |
|
||||
| `<GITHUB_APP_PRIVATE_KEY>` | Your GitHub App private key (PEM) |
|
||||
| `<S3_ACCESS_KEY>` | RustFS/S3 access key ID |
|
||||
| `<S3_SECRET_KEY>` | RustFS/S3 secret access key |
|
||||
| `<PLACEHOLDER>` | Any other password/key/token (named in context where it appears) |
|
||||
|
||||
Secret **values** never appear in these docs — only key names and placeholders. The following are intentionally **kept as-is** because they are not sensitive and are needed to follow along: the pod CIDR `10.42.0.0/16`, the service CIDR `10.43.0.0/16`, the CoreDNS service IP `10.43.0.10`, the bucket name `sccache`, in-cluster service DNS names and ports, and all version numbers.
|
||||
|
||||
## Recommended setup order
|
||||
|
||||
Work through the numbered guides in order — each builds on the previous:
|
||||
|
||||
1. **[01-host-and-cluster.md](01-host-and-cluster.md)** — Host prep (KVM, firewall/NAT) and the single-node k3s install, networking, and CNI.
|
||||
2. **[02-kata-runtime.md](02-kata-runtime.md)** — Install Kata Containers, wire it into k3s' containerd, and register the `kata-qemu` RuntimeClass.
|
||||
3. **[03-runner-image.md](03-runner-image.md)** — Build the preloaded runner image and import it into k3s containerd.
|
||||
4. **[04-arc-and-caching.md](04-arc-and-caching.md)** — Install ARC and the runner scale set, deploy the RustFS shared cache, wire up the `sccache-s3` secret, and apply the egress NetworkPolicy.
|
||||
Executable
+81
@@ -0,0 +1,81 @@
|
||||
#!/usr/bin/env bash
|
||||
# Build + roll the preloaded omp-kata runner image onto the self-hosted CI host,
|
||||
# driven over SSH from this repo. The Dockerfile next to this script is the
|
||||
# source of truth: it is copied to the host, built there, imported into k3s
|
||||
# containerd, and the ARC runner scale set is pointed at the new tag and rolled.
|
||||
#
|
||||
# The host is intentionally NOT hardcoded (this repo is public). Set CI_HOST to
|
||||
# your ssh target; the remaining knobs default to the reference deployment.
|
||||
#
|
||||
# Usage:
|
||||
# CI_HOST=my-ci-host ./infra/reload-runner.sh # tag: omp-kata-runner:YYYY-MM-DD-HHMMSS
|
||||
# CI_HOST=my-ci-host ./infra/reload-runner.sh 2026-06-20 # tag: omp-kata-runner:2026-06-20
|
||||
# CI_HOST=my-ci-host ./infra/reload-runner.sh my/repo:tag # explicit repo:tag
|
||||
#
|
||||
# Env knobs (defaults match the reference deployment):
|
||||
# CI_HOST ssh target of the CI host (required)
|
||||
# REMOTE_CTX remote build dir for the Dockerfile [/root/omp-kata-runner-image]
|
||||
# ARC_VALUES remote ARC scale-set helm values file [/root/arc-omp-values.yaml]
|
||||
# ARC_RELEASE helm release name of the runner scale set [omp-kata]
|
||||
# ARC_NAMESPACE namespace the runner scale set lives in [arc-runners]
|
||||
# ARC_CHART_VERSION gha-runner-scale-set chart version [0.14.2]
|
||||
# KUBECONFIG_REMOTE kubeconfig path on the host [/etc/rancher/k3s/k3s.yaml]
|
||||
set -euo pipefail
|
||||
|
||||
: "${CI_HOST:?set CI_HOST to the ssh target of your CI host, e.g. CI_HOST=my-ci-host}"
|
||||
REMOTE_CTX="${REMOTE_CTX:-/root/omp-kata-runner-image}"
|
||||
ARC_VALUES="${ARC_VALUES:-/root/arc-omp-values.yaml}"
|
||||
ARC_RELEASE="${ARC_RELEASE:-omp-kata}"
|
||||
ARC_NAMESPACE="${ARC_NAMESPACE:-arc-runners}"
|
||||
ARC_CHART_VERSION="${ARC_CHART_VERSION:-0.14.2}"
|
||||
KUBECONFIG_REMOTE="${KUBECONFIG_REMOTE:-/etc/rancher/k3s/k3s.yaml}"
|
||||
|
||||
arg="${1:-$(date +%Y-%m-%d-%H%M%S)}"
|
||||
case "$arg" in *:*) IMAGE="$arg";; *) IMAGE="omp-kata-runner:$arg";; esac
|
||||
|
||||
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
[ -f "$here/runner.Dockerfile" ] || { echo "no runner.Dockerfile next to $0" >&2; exit 1; }
|
||||
|
||||
echo "==> [0/5] copying Dockerfile to ${CI_HOST}:${REMOTE_CTX}"
|
||||
ssh "$CI_HOST" "mkdir -p '$REMOTE_CTX'"
|
||||
scp -q "$here/runner.Dockerfile" "${CI_HOST}:${REMOTE_CTX}/Dockerfile"
|
||||
|
||||
# All build/import/rollout steps run on the host. Config is passed as positional
|
||||
# args (no secrets, no spaces) so it survives the ssh command-string re-parse
|
||||
# regardless of the host's login shell.
|
||||
ssh "$CI_HOST" bash -s -- \
|
||||
"$IMAGE" "$REMOTE_CTX" "$ARC_VALUES" "$ARC_RELEASE" "$ARC_NAMESPACE" "$ARC_CHART_VERSION" "$KUBECONFIG_REMOTE" <<'REMOTE'
|
||||
set -euo pipefail
|
||||
IMAGE="$1"; REMOTE_CTX="$2"; ARC_VALUES="$3"; ARC_RELEASE="$4"; ARC_NAMESPACE="$5"; ARC_CHART_VERSION="$6"
|
||||
export KUBECONFIG="$7"
|
||||
cd "$REMOTE_CTX"
|
||||
|
||||
echo "==> [1/5] building $IMAGE"
|
||||
DOCKER_BUILDKIT=1 docker build -t "$IMAGE" -t omp-kata-runner:preloaded .
|
||||
|
||||
echo "==> [2/5] verifying baked tools"
|
||||
docker run --rm --entrypoint bash "$IMAGE" -lc '
|
||||
set -e
|
||||
for b in gh fd rg magick bun cargo rustc pkg-config zstd; do
|
||||
command -v "$b" >/dev/null || { echo "MISSING: $b"; exit 1; }
|
||||
done
|
||||
echo "tools OK | $(bun --version) | $(rustc --version) | gh $(gh --version | head -1 | cut -d" " -f3)"
|
||||
'
|
||||
|
||||
echo "==> [3/5] importing into k3s containerd (k8s.io namespace)"
|
||||
docker save "$IMAGE" | k3s ctr -n k8s.io images import --platform linux/amd64 -
|
||||
|
||||
echo "==> [4/5] pointing ARC runner scale set at $IMAGE"
|
||||
sed -i "s#image: omp-kata-runner:.*#image: $IMAGE#" "$ARC_VALUES"
|
||||
helm upgrade "$ARC_RELEASE" --namespace "$ARC_NAMESPACE" --version "$ARC_CHART_VERSION" \
|
||||
-f "$ARC_VALUES" \
|
||||
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set >/dev/null
|
||||
|
||||
echo "==> [5/5] verifying rollout"
|
||||
live="$(kubectl get autoscalingrunnerset "$ARC_RELEASE" -n "$ARC_NAMESPACE" \
|
||||
-o jsonpath='{.spec.template.spec.containers[0].image}')"
|
||||
echo "ARC runner image is now: $live"
|
||||
[ "$live" = "$IMAGE" ] && echo "OK: reloaded $IMAGE" || { echo "MISMATCH: expected $IMAGE"; exit 1; }
|
||||
REMOTE
|
||||
|
||||
echo "OK: $IMAGE built on ${CI_HOST}, imported into k3s, and rolled out to ARC."
|
||||
@@ -0,0 +1,54 @@
|
||||
# syntax=docker/dockerfile:1
|
||||
# Preloaded omp-kata runner image.
|
||||
#
|
||||
# Stock GitHub Actions runner (Ubuntu 24.04) with the dependencies CI installs
|
||||
# on every job baked in, so each ephemeral Kata microVM boots with them already
|
||||
# present instead of re-fetching them per job:
|
||||
# - APT system deps (canvas/cairo stack + fd/ripgrep/imagemagick) + fd/magick shims
|
||||
# - GitHub CLI (gh) — present on GitHub-hosted runners; the coding-agent github
|
||||
# tool and release workflows expect it
|
||||
# - C/build toolchain the native + canvas builds need
|
||||
# - bun (system-wide, on PATH)
|
||||
# - rust nightly toolchain (pinned) + clippy/rustfmt + linux-arm64/windows-msvc targets
|
||||
#
|
||||
# Rebuild + reimport (see /root/omp-kata-runner.md) after bumping the ARGs below
|
||||
# or the apt set. Keep the apt set in sync with .github/actions/setup-system-deps.
|
||||
FROM ghcr.io/actions/actions-runner:latest
|
||||
|
||||
ARG RUST_NIGHTLY=nightly-2026-04-29
|
||||
ARG BUN_VERSION=1.3.14
|
||||
|
||||
USER root
|
||||
ENV DEBIAN_FRONTEND=noninteractive
|
||||
|
||||
# Mirrors the "Install system deps" block in .github/workflows/ci.yml plus the
|
||||
# C/build toolchain (native + canvas builds) and the GitHub CLI. The gh apt repo
|
||||
# is added first so `gh` installs in the same apt transaction.
|
||||
RUN curl -fsSL https://cli.github.com/packages/githubcli-archive-keyring.gpg -o /usr/share/keyrings/githubcli-archive-keyring.gpg \
|
||||
&& chmod go+r /usr/share/keyrings/githubcli-archive-keyring.gpg \
|
||||
&& echo "deb [arch=$(dpkg --print-architecture) signed-by=/usr/share/keyrings/githubcli-archive-keyring.gpg] https://cli.github.com/packages stable main" > /etc/apt/sources.list.d/github-cli.list \
|
||||
&& apt-get update \
|
||||
&& apt-get install -y \
|
||||
build-essential pkg-config curl ca-certificates git unzip xz-utils zstd gh \
|
||||
libcairo2-dev libpango1.0-dev libjpeg-dev libgif-dev librsvg2-dev \
|
||||
fd-find ripgrep imagemagick \
|
||||
&& ln -sf "$(command -v fdfind)" /usr/local/bin/fd \
|
||||
&& ln -sf /usr/bin/convert /usr/local/bin/magick \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# bun, system-wide (BUN_INSTALL/bin == /usr/local/bin, already on PATH).
|
||||
ENV BUN_INSTALL=/usr/local
|
||||
RUN curl -fsSL https://bun.sh/install | bash -s "bun-v${BUN_VERSION}" \
|
||||
&& bun --version
|
||||
|
||||
# rust toolchain for the runner user; rustup default == pinned nightly so
|
||||
# dtolnay/rust-toolchain@nightly and target/component adds are no-ops in CI.
|
||||
USER runner
|
||||
ENV RUSTUP_HOME=/home/runner/.rustup \
|
||||
CARGO_HOME=/home/runner/.cargo \
|
||||
PATH=/home/runner/.cargo/bin:/usr/local/bin:${PATH}
|
||||
RUN curl --proto '=https' --tlsv1.2 -fsSL https://sh.rustup.rs \
|
||||
| sh -s -- -y --default-toolchain "${RUST_NIGHTLY}" --profile minimal \
|
||||
&& rustup component add clippy rustfmt \
|
||||
&& rustup target add aarch64-unknown-linux-gnu x86_64-pc-windows-msvc \
|
||||
&& cargo --version && rustc --version
|
||||
Reference in New Issue
Block a user