Stanford Professor Emeritus Makes the Case for Homa to Replace TCP in AI Datacenters
John Ousterhout is pursuing IETF standardization and Linux kernel adoption for a message-based networking protocol designed to curb latency in cloud and AI clusters.

Transmission Control Protocol (TCP), the foundational networking standard underlying the modern internet and cloud computing infrastructure, is ill-matched to the latency demands of emerging artificial intelligence workloads, according to John Ousterhout, professor emeritus of computer science at Stanford University. Speaking at the AI Engineer World’s Fair, Ousterhout argued that datacenter networks require a clean-slate transport protocol, advocating for an alternative called Homa in comments reported by The Register (https://www.theregister.com/networks/2026/10/01/stanford-prof-is-beating-the-drum-for-a-new-protocol-to-replace-tcp/5300629).
Developed originally by Vint Cerf and his collaborators to regulate packet delivery across disparate networks, TCP structures data into continuous, unprioritized byte streams. Under TCP architecture, senders manage congestion control by adjusting transmission rates based on delivery acknowledgments, requiring the transmitting node to estimate network capacity with limited insight into the receiver's queue. While adequate for general internet traffic, that serialization model introduces latency variance when short, time-sensitive control bursts must compete with massive bulk data transfers in datacenters.
Homa departs from this model by using a message-based structure with explicitly defined message lengths, akin to remote procedure calls. First detailed in a 2019 Stanford doctoral dissertation by Behnam Montazeri, now a staff engineer at Google, the protocol shifts congestion management to the receiver. The receiving node uses information in the initial packet to explicitly schedule transmission times using a shortest-remaining-processing-time (SRPT) algorithm, prioritizing smaller payloads over larger transfers.
According to test figures cited by Ousterhout, this receiver-managed scheduling reduces tail latency significantly. On a 100 Gbps network running at 80 percent utilization, Homa achieved a 99th percentile (p99) latency of 92 microseconds for short messages, compared to 1.2 milliseconds for TCP—a thirteenfold reduction. Latency for large messages improved by a factor of two. Deploying the protocol does not require hardware alterations or system reboots; administrators compile Homa from source and insert the module into the Linux kernel, allowing it to run concurrently alongside TCP traffic while easing pressure on remaining TCP workloads.
The push for lower transport latency coincides with infrastructure demands from frontier AI labs training and serving large language models. These clusters transfer large volumes of weight gradients, model checkpoints, and key-value cache entries across network fabrics that must simultaneously accommodate brief control messages, metadata lookups, and agent calls. In these environments, millisecond-level transmission delays can leave expensive graphics processing units idling while waiting for synchronization.
Following his retirement from teaching, Ousterhout has prioritized Homa's broader ecosystem adoption. He is drafting an Internet Engineering Task Force (IETF) standardization document and working to upstream the protocol into the mainline Linux kernel. Homa was backported to Red Hat Enterprise Linux versions 8 and 9.5 in March, and Ousterhout is currently collaborating with an unnamed financial services company on a prototype deployment.
The protocol faces established alternatives and skepticism from within the networking community. In 2023, network architect Ivan Pepelnjak published a critical position paper questioning Ousterhout's comparative performance characterizations of TCP and framing Homa as a niche solution. Datacenter operators have also turned to other latency-reduction mechanisms, including the Data Plane Development Kit (DPDK) for databases, NVMe-oF and specialized RDMA fabrics for storage, Amazon Web Services' Scalable Reliable Datagram, and Google's QUIC protocol for web delivery.
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.



