FR
live

Kubernetes 1.37 enables HPA scale-to-zero by default and shields etcd from apiserver overload

On August 26, 2026, Kubernetes shipped version 1.37 "Garhwal", with 67 enhancements and 16 graduates to stable. The HorizontalPodAutoscaler can now idle a workload down to zero pods by default, and the apiserver no longer floods etcd while its watch cache warms up.

A row of identical supermarket checkout lanes in a closed, dark store, all dark except one with a single amber standby light still lit.

August 26, 2026. Kubernetes shipped version 1.37, named “Garhwal” after a Himalayan region of Uttarakhand, India. The release tally is unusual: 67 enhancements — 16 graduating to Stable, 23 to Beta, 27 entering Alpha, and one deprecation — with no headline feature in sight. Why it matters: this is a consolidation release, and that is exactly what production teams have been asking for. Two changes land directly on cost and control-plane resilience: the HorizontalPodAutoscaler can now idle a workload down to zero pods by default, and the kube-apiserver stops flooding etcd while its watch cache warms up.

The HPA learns to stop: scale-to-zero in beta, on by default

The most profitable feature in 1.37 is quiet. First introduced in v1.16, HorizontalPodAutoscaler scale-to-zero moves to Beta and — more importantly — becomes enabled by default. For workloads driven by object or external metrics, spec.minReplicas: 0 is all it takes for the HPA to scale replicas down to zero when the queue is empty, then restore them when demand returns.

The boundary is drawn up front: scaling to zero based on CPU or memory is not supported, because those metrics only exist while pods are running. The real target is elsewhere — queue consumers, batch jobs, GPU workloads — where idleness is measurable through an external metric that does not depend on live pods.

The mechanism is clean. While the HPA holds a workload at zero, it exposes a ScaledToZero condition set to True in its status. The controller uses that condition to tell apart a workload it scaled to zero itself — and will bring back when the metric returns — from a deployment a human set to zero manually. When the workload comes back, the condition flips to False with the reason NotScaledToZero.

The manifest is a few lines. The decisive part is combining minReplicas: 0 with an external metric that does not depend on live pods:

yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: file-processor
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: file-processor
  minReplicas: 0
  maxReplicas: 10
  metrics:
    - type: External
      external:
        metric:
          name: queue_depth
        target:
          type: AverageValue
          averageValue: "1"

Here the HPA reads an external metric (queue_depth) that lives in the message broker, not in the pods: it stays measurable even when the deployment is at zero. Once the queue averages more than one message, the controller scales the deployment back up; when it falls to zero, it closes it down again. For a Kafka consumer or an SQS queue that only runs intermittently, the saving is immediate, with no change to the business logic.

The end of a nine-year oddity: the metrics.k8s.io API goes stable

The other notable promotion is the metrics.k8s.io API, which finally reaches Stable after nearly nine years in beta. It provides the standard way to fetch CPU and memory usage for pods and nodes, and it powers two daily drivers: the kubectl top command and the HorizontalPodAutoscaler.

The graduation follows the project’s stated goal of avoiding permanent Beta APIs. Now that a v1 exists, future releases will move to it; v1beta1 stays usable throughout the transition, in line with the deprecation policy, so adopting the Stable API breaks no existing workflow. For an SRE, the practical effect is small but real: one less dependency that can shift underneath you during an upgrade.

The symbol matters as much as the mechanics. An API left nine years in beta had become a standing reminder that the releases’ promised stability did not reach the surfaces operators touch every day. By freezing it, Kubernetes closes one of its oldest API debts, and sends the same signal as the KYAML graduation: the daily tools stop being pre-releases.

The apiserver no longer crushes etcd on startup

The most consequential change for large clusters is silent. Kubernetes finished the work on resilient watch cache initialization: the ResilientWatchCacheInitialization feature gate reached Stable in v1.34, and 1.37 locks on the remaining gate, WatchCacheInitializationPostStartHook, now enabled permanently.

What changes is how the apiserver behaves at startup and during recovery. Previously, warming the cache triggered a burst of list and watch requests against etcd; requests piled up while the cache filled. Now the apiserver delegates bounded requests and rejects the overflow with HTTP 429, instead of letting the load crush etcd or exhaust API Priority and Fairness capacity.

The guidance for operators is explicit: clients — including in-house controllers and operators — must handle 429 Too Many Requests gracefully, honoring the Retry-After header and implementing exponential backoff. That is the price of the apiserver’s new self-defense: it only works if your clients do not turn the protection into a cascading failure.

The textbook case is a control-plane restart after an outage. Before, a kube-apiserver coming back could, by replaying thousands of list requests while its cache warmed, drag etcd down again just as the cluster was trying to stabilize. Now recovery is bounded: the apiserver comes back gradually, and well-behaved clients retry instead of hammering.

Admission policies now survive an etcd outage

1.37 also advances manifest-based admission control to Beta. Admission webhooks and CEL-based policies can now be loaded from files on disk, through the staticManifestsDir field of the AdmissionConfiguration, instead of living only inside the Kubernetes API.

The operational benefit is direct: policies loaded this way are enforced from apiserver startup, keep working while etcd is unavailable, and can protect the API-based admission resources themselves from modification. It is a step toward guardrails that no longer depend on the datastore being up to hold.

This matters most in air-gapped or degraded environments, where the control plane must keep enforcing policy even while its backing store recovers.

What to check before you upgrade

Two practical points dominate for anyone planning the upgrade. First, containerd 2.0 is now the minimum: Kubernetes 1.35 was the last release to support containerd 1.x, and 1.37 removes the kubelet flags tied to that end of life. The runtime migration needs to be settled before the upgrade, not after.

Second, the quiet stabilizers of daily life: KYAML goes Stable (kubectl get -o kyaml is now stable), the SELinuxMount and SELinuxChangePolicy features graduate to Stable, and Pod-level checkpoint and restore enters Alpha through new CRI RPCs (CheckpointPod, RestorePod), provided your runtime implements them.

One quiet stabilizer deserves a second look: KYAML. With kubectl get -o kyaml now stable, you get a safer, less ambiguous rendering of manifests without changing any of your files — every KYAML document is valid YAML, so your existing pipelines keep working. It is a small thing, but it is the kind of small thing that removes entire classes of formatting bugs from the daily kubectl workflow.

Verdict

If you run event-driven workloads — queue consumers, batch jobs, GPU tasks — enable scale-to-zero as soon as you land on 1.37: it is on by default, and spec.minReplicas: 0 converts idleness directly into savings without touching your CPU/memory metrics. If you operate a real-size cluster, the priority is elsewhere: confirm your operators tolerate 429s with exponential backoff, because the etcd protection is only worth as much as your clients that will not turn it into a stampede. And in every case, settle the containerd 2.0 migration before upgrading — it is the one prerequisite that does not forgive being handled after the fact.

References

The cyber brief, every Tuesday

The flaws that matter and the patches to apply, in a ten-minute read.

No spam. One-click unsubscribe.
read next

On the same topic

Docker Cloud Sandboxes move long-horizon agent work from the laptop to the cloud with one command

On September 24, 2026, Docker extended its Sandboxes to the cloud: the same microVM environment, running on Docker-managed compute, with a single command to move a project between a laptop and the cloud. Teams handing multi-hour tasks to coding agents no longer have to choose between a laptop that sleeps and isolation they would have to rebuild.

OpenTelemetry and Prometheus finally converge, and the 2026 numbers confirm it

A 2026 survey shows interoperability between OpenTelemetry and Prometheus has improved markedly: the ease-of-use score climbed from 3.1 to 3.6, and the share who find them hard to combine fell from 29% to 10%. For an SRE team still on the fence, now is the time to consolidate on the OTel Collector without abandoning Prometheus.

← Back to the feed

Type at least two characters.

↑ ↓ navigate ↵ open esc dismiss