There's a video surfacing on Homa, a networking approach that could replace TCP for AI…
By AI Update World · 2026-10-04

TCP, or Transmission Control Protocol, has been the backbone of reliable data communication since the 1970s. It was designed with a fundamental assumption: networks are lossy and unreliable, so every packet of data needs to be tracked, verified, and retransmitted if lost. This worked brilliantly for email, web browsing, and file transfers, where occasional latency was acceptable. But TCP carries overhead. It requires handshakes before data flows, acknowledgments after receipt, and automatic slowdowns when congestion is detected. For decades, this tradeoff made sense. For AI training clusters, it increasingly does not.
Modern AI clusters are a different beast entirely. Thousands of GPUs sit in the same data center, connected by networks specifically engineered to be fast and reliable. These are not the public internet. The physics of the problem is different: latency matters enormously because training loops iterate millions of times, and packet loss is rare because the hardware is under direct control. When an GPU finishes computing one layer and needs to shuffle data to the next GPU in line, waiting for TCP's safety mechanisms feels like wearing a seatbelt while driving across a parking lot. The overhead is real and compounds across millions of operations.
This is where alternative transport approaches enter the picture. Rather than assuming the network might fail and building in extensive recovery, some researchers ask: what if we design specifically for the conditions we actually have? For cluster communication, you might optimize for low latency over absolute reliability, or allow applications to handle retransmission themselves rather than having the protocol do it invisibly. You might prioritize throughput by letting congestion control happen at the application level rather than deep in the network stack. Different workloads need different guarantees, and TCP offers a one-size-fits-all solution that feels wasteful when you know more about your specific environment.
The history of specialized transport protocols goes back further than many realize. Research communities have been experimenting with alternatives to TCP since at least the 1990s. UDP, or User Datagram Protocol, is simpler and faster but gives up reliability entirely. Various academic projects and industry efforts have explored custom protocols for video streaming, gaming, scientific computing, and cloud workloads. Some succeeded in specific niches. Most were constrained by the fact that TCP was already everywhere and changing infrastructure is costly and risky.
AI clusters represent a new opportunity for this conversation because the infrastructure is being built from scratch, investment in it is enormous, and the performance gains from better networking ripple upward into training speed and hardware utilization. When you're running training jobs that cost millions of dollars and take weeks, even small percentage improvements in data movement efficiency become worth signific