Homa cuts datacenter short-message latency thirteenfold versus TCP
On October 1, 2026, John Ousterhout, Stanford professor emeritus, made the case for Homa, a message-based transport protocol whose p99 latency on short messages drops to 92 microseconds against 1.2 milliseconds for TCP over a 100 Gbps link. Evaluate it on the datacenter links where every millisecond idles a GPU, but price the adoption cost before replacing anything.
October 1, 2026. John Ousterhout, a Stanford professor emeritus and the creator of Tcl/Tk as well as the Raft consensus algorithm, is publicly arguing for Homa, a transport protocol designed to replace TCP inside datacenters. 92 microseconds. That is Homa’s 99th-percentile latency on short messages, against 1.2 milliseconds for TCP. March 2026. The protocol was backported to Red Hat Enterprise Linux 8 and 9.5. Why it matters: TCP carries nearly the entire web, but it was built for byte streams, not for the bursts of short messages that drive AI — and in a datacenter, a millisecond lost is a GPU sitting idle.
What Homa changes relative to TCP
TCP is a byte-stream protocol: messages are serialized into one continuous flow, with no differentiation and no priority. A receiver cannot tell a short message from a long one, and congestion control belongs to the sender, which must guess the right pace from the arrival of acknowledgments. That machinery, designed by Vint Cerf and his peers to tame heterogeneous networks, becomes a bottleneck the moment latency is the metric.
Homa flips the roles. It is a message-based protocol, close to a remote procedure call (RPC): each message’s length is explicit. The receiver drives congestion control — the first packet received announces how much data is coming, and the receiver then schedules when each packet should be sent. That inversion makes a SRPT (shortest-remaining-processing-time) algorithm possible: short messages overtake long ones instead of queueing behind them in an undifferentiated stream.
The result is measurable at both extremes. Over a 100 Gbps network at 80% utilization, Homa’s p99 for short messages falls to 92 microseconds, thirteen times below TCP’s 1.2 milliseconds. Even on the longest messages, Homa is twice as fast, according to Ousterhout.
Why now: GPUs do not wait
The timing of this pitch is no accident. The labs training large models push weight gradients, model weights, KV-cache entries and checkpoints across the network — big transfers that must now share bandwidth with short bursts from agents and control tasks such as metadata coordination and cache lookups.
“For these workloads, what really matters is latency,” Ousterhout says. Even a millisecond idles an expensive GPU. That sentence turns an academic protocol debate into a billing problem: when a cluster of several thousand GPUs loses a fraction of a second at every synchronization point, the cost lands in money, not abstractions.
Homa is not the first to route around TCP for that reason. The database world uses DPDK to bypass the network stack and speed up queries. Storage moved to NVMe-oF over RDMA or Fibre Channel. The web adopted QUIC, the basis of HTTP/3, to escape TCP’s head-of-line blocking. AWS ships its Scalable Reliable Datagram (SRD). Homa sits in that lineage, but with a broader ambition: to replace TCP as the datacenter’s default transport.
A deliberately gradual deployment
Ousterhout’s most pragmatic argument fits in one sentence: Homa does not require ripping anything out. You compile the module from its GitHub repository, install it in the Linux kernels of clients and servers, with no reboot. “Homa works side by side with TCP, so you can gradually move applications over,” he writes — adding that running Homa even makes the remaining TCP applications faster.
The protocol did not appear last week. It began as a PhD dissertation by Behnam Montazeri, published in 2019 and now a staff engineer at Google. Ousterhout, now retired from teaching, has made spreading Homa his “life’s mission.” He is drafting an IETF standardization document and working through upstreaming the protocol into the Linux kernel; the backport to RHEL 8 and 9.5 in March 2026 shows the effort has left the lab. He is also helping one large financial-services company with a prototype.
The skeptics have a point
Not everyone shares the enthusiasm. Network architect Ivan Pepelnjak published a scathing position paper in 2023 against Homa, disputing how Ousterhout characterizes TCP’s performance and dismissing Homa as “a solution looking for a problem.” The debate is unresolved, and that is healthy: replacing a protocol proven over four decades demands a high burden of proof.
The reality is more nuanced than the pitch. TCP remains the champion of the web and the cloud, and nothing suggests it will vanish from the public internet. Homa’s real territory is the datacenter’s internal fabric — where the operator controls both ends of the link and can deploy a kernel module without negotiating with thousands of peers. That is precisely the perimeter infrastructure teams already own.
How to evaluate Homa before you trust it
The pitch is strong, but the only benchmark that matters is your own. Run Homa side by side with TCP on the exact links that feed your GPU cluster, and measure the p99 latency of your real message mix — not a synthetic workload. Watch two things in particular. First, Homa’s SRPT scheduling favors short messages by design; that is the point, but it means a flood of tiny control messages can starve long transfers if you do not bound their share. Second, the kernel module is still pre-mainline: a protocol under active IETF drafting and Linux upstreaming will change, so pin the version you validated and re-test on every upgrade. The thirteenfold p99 gain is measured at 100 Gbps and 80% utilization — a number worth reproducing, not importing.
Verdict
If you operate a GPU cluster for AI training or inference, Homa deserves a benchmark on your internal links: a thirteenfold gain on short-message p99 translates directly into recovered GPU time, and side-by-side operation with TCP limits the adoption risk. If your datacenter is mostly web traffic or large transfers, the priority lies elsewhere — QUIC, DPDK or NVMe-oF already address their own bottlenecks, and Homa will not move your metrics. Either way, do not replace TCP on the strength of a chart: measure your own workloads’ p99, on your own topology, before letting an experimental protocol carry production traffic. The protocol has a strong argument — it still has to earn the right to prove it in your environment.