FR
live

Kubernetes schedules cgroup v1 removal in v1.38 and blocks the kubelet on cgroup v1 nodes from v1.35

Kubernetes has deprecated cgroup v1 and plans its full removal in v1.38; since v1.35, the kubelet refuses to start on a node still running cgroup v1. If your nodes are still on cgroup v1, migrate them to cgroup v2 before upgrading, or the cluster will fail to start.

A row of identical server drawers in a dark rack, one drawer pulled open and empty.

October 6, 2026. Kubernetes published a migration guide to cgroup v2, authored by Paco Xu of DaoCloud, and the message is unambiguous: cgroup v1 support is deprecated, the kubelet refuses to start on a cgroup v1 node by default since v1.35, and full removal is scheduled for v1.38. Why it matters: plenty of nodes still run cgroup v1 without anyone realizing, and the first upgrade to v1.35 or later will end in a kubelet startup failure.

What is deprecated, and on what schedule

The timeline is precise. cgroup v2 has been stable in Kubernetes since v1.25, released in 2022, and cgroup v1 moved into maintenance mode with v1.31. The deprecation is now real: starting with v1.35, the failCgroupV1 option defaults to true, so the kubelet will not start on a node still running cgroup v1. An administrator can temporarily set failCgroupV1: false in the kubelet configuration, but removal will follow the Kubernetes deprecation policy.

The removal work is tracked in KEP-5573 (“Remove cgroup v1 support”). The fallback is scheduled for removal in v1.38. In practice: if you are still on a release older than v1.35, migrate every Linux node to cgroup v2 before upgrading, or plan the failCgroupV1: false override. If you are already on v1.35 or later, confirm that every node runs cgroup v2 — or deliberately keep the override. For kubeadm-managed clusters, v1.35 also tightens the early check: the SystemVerification preflight returns an error during kubeadm init, kubeadm join, and kubeadm upgrade when it detects cgroup v1 with a kubelet at v1.35 or later.

Why cgroup v1 no longer cuts it

The guide lists the concrete limits of the old interface, which drive the migration beyond the deprecation itself.

The sneakiest one is about memory: the kubelet treats active_file memory as not reclaimable. Under I/O-intensive workloads, a large page cache can therefore trigger memory pressure and evict pods even though the memory is actually reclaimable. This is bug kubernetes/kubernetes#43916, and migrating to cgroup v2 does not change that calculation on its own: the documented workaround is to set equal memory requests and limits for containers doing heavy I/O.

More importantly, cgroup v2 unlocks recent features that cgroup v1 cannot provide. Memory QoS, introduced as alpha in v1.22 and still alpha in v1.36, relies on the cgroup v2 memory controller: memory.high provides throttling, while memory.min and memory.low give hard and soft protection under tiered reservation. Per-container OOM handling also depends on v2: by default singleProcessOOMKill is false, and the kubelet sets memory.oom.group to kill every process in a container together instead of leaving a partially functioning one. Finally, cgroup v2 is the only interface that officially supports delegation to less-privileged containers — the foundation of rootless containers.

Structurally, the core difference is the unified hierarchy. cgroup v1 stacks separate hierarchies per controller, which complicates delegation and makes limit consistency harder to audit; cgroup v2 unifies them into a single hierarchy that is simpler to reason about and secure. That is the property both rootless containers and the memory controller lean on.

The failure mode is silent until it is not. A node still on cgroup v1 will keep running fine through v1.34, then fail the kubelet on the upgrade to v1.35 — and because the kubelet does not start, the node never rejoins the cluster, taking its pods down with it. The preflight check in kubeadm catches this early, but only if you run kubeadm upgrade rather than replacing the node image in place. That is why the guide’s advice is to verify cgroup2fs on every node before the version bump, not after.

The genealogy of cgroup v2 in the kernel

The move to v2 is also a kernel version story, and the guide lays out the timeline. When Kubernetes was announced in 2014, only cgroup v1 existed. cgroup v2 appeared in Linux 4.5, released in 2016, with the io, memory, and pids controllers. The cpu controller arrived in 4.15, and PSI (Pressure Stall Information) from 4.20. The project advises against using cgroup v2 with a kernel older than 5.2, for lack of cgroup-level task-freezer support, and documents 5.8 as the minimum — the version that added the root cgroup’s cpu.stat file. The memory.high livelock fix used by Memory QoS is only present from 5.9.

That timeline explains why so many nodes remain on cgroup v1: distributions adopted v2 by default gradually, and a server installed years ago can happily run v1 with nothing signaling it. The stat -fc %T /sys/fs/cgroup/ command is exactly the way to clear up the ambiguity in a second.

What the migration requires

The requirements are concrete and checkable. You need a Kubernetes version with v2 support (stable since v1.25), a Linux kernel 5.8 or later (5.9 recommended for Memory QoS), and a compatible container runtime: containerd v1.4 minimum, or v2.0 for automatic cgroup-driver detection; CRI-O v1.20 minimum. The kubelet and the runtime must use the same cgroup driver.

The guide recommends the systemd driver when the cluster is kubeadm-managed, since kubeadm runs the kubelet as a systemd service. Since v1.34, automatic detection of the runtime’s cgroup driver through the CRI (KEP-4033) is stable, provided the runtime implements the RuntimeConfig RPC (containerd v2.0+ or CRI-O v1.28+).

Two details deserve a note. The first is CPU weight conversion: cgroup v1 uses cpu.shares, cgroup v2 uses cpu.weight, and newer runtimes apply a more faithful non-linear conversion — available in crun v1.23 and runc v1.3.2. After upgrading a runtime, monitoring tools that predict exact cpu.weight values may need updating. The second is PSI (Pressure Stall Information), which exposes CPU, memory, and I/O contention per node, pod, and container: it requires cgroup v2, a kernel 4.20 or later, and CONFIG_PSI=y. The kubelet now exposes it by default (KubeletPSI is stable).

Checking after the migration

To verify a node, one command is enough: stat -fc %T /sys/fs/cgroup/ must return cgroup2fs. To list Kubernetes-related systemd units, systemctl list-units "kube*" --type=slice followed by systemd-cgls /sys/fs/cgroup/* shows the effective hierarchies. When tracing a pod to its cgroup, the guide calls out a classic trap: with the systemd driver, the value info.runtimeSpec.linux.cgroupsPath is a systemd unit path (slice:runtime:id), not a directory under /sys/fs/cgroup — do not treat it as a filesystem path.

The full flow to confirm a limit is actually applied fits in a few commands:

bash
# 1. Is the node actually on cgroup v2?
stat -fc %T /sys/fs/cgroup/          # must print "cgroup2fs"

# 2. Find a specific container's cgroup (systemd driver)
CONTAINER_ID=$(crictl ps \
  --label io.kubernetes.pod.namespace=<ns> \
  --label io.kubernetes.pod.name=<pod> \
  --name <container> -q | head -n1)
PID=$(crictl inspect "$CONTAINER_ID" | jq -r '.info.pid')
CGROUP="/sys/fs/cgroup$(awk -F: '$1=="0"{print $3}' /proc/$PID/cgroup)"

# 3. Read the effective values
cat "$CGROUP/cpu.weight"    # CPU request: shares converted to weight
cat "$CGROUP/cpu.max"       # CPU limit
cat "$CGROUP/memory.max"    # memory limit

The reading is direct: if cpu.weight or memory.max do not match what the manifest declares, it is either a runtime conversion to revisit or a pod mid-resize whose spec and status are out of sync.

For teams running mixed-fleet clusters, the migration is best staged: convert node pools one at a time, watch the kubelet logs for cgroup warnings after each step, and keep the failCgroupV1: false override only as a short-term bridge while stragglers are rebuilt.

Verdict

If your nodes are still on cgroup v1, do not upgrade until the migration is done: v1.35 fails the kubelet by default, and v1.38 removes the fallback. The path is mechanical — kernel 5.8+, an up-to-date runtime, the systemd driver, then the stat -fc %T check. If you are already on cgroup v2, use the opportunity to enable the features v2 unlocks, starting with Memory QoS in tiered reservation if your workloads need fine-grained memory protection. Either way, treat the deprecation as a migration deadline, not a warning you can postpone indefinitely.

References

The cyber brief, every Tuesday

The flaws that matter and the patches to apply, in a ten-minute read.

No spam. One-click unsubscribe.
read next

On the same topic

GitHub stacked pull requests go GA with 9% more merged code

GitHub announced general availability for stacked pull requests, which split a large change into small, independently reviewed pull requests that merge together at the end. If your branches keep blocking each other, GitHub’s numbers — 9% more merged code and 5% faster time-to-merge — justify adopting them now.

← Back to the feed

Type at least two characters.

↑ ↓ navigate ↵ open esc dismiss