When The East-West Problem Leaves The Building

The next AI bottleneck isn’t compute—it’s connectivity across sites

In this segment, Ben Edmond explores the importance of connecting the back end of AI data centers across sites and why the commercial challenge of that process may be even harder than the engineering one. In illuminating the friction of the current market structure, he showcases the need for a shared, authoritative way to name what is being bought and sold—and why Connectbase is in prime position to support the constraint of coordinating AI at scale.

Inside every large AI training cluster lies a problem that has nothing to do with how fast the chips are. The bottleneck is the network between them—GPU talking to GPU, exchanging gradients hundreds of times a second in tight synchronized lockstep. Training is a bulk-synchronous operation: Every GPU reaches a barrier and waits for the slowest one before the next step begins. A single congested link or lagging node idles tens of thousands of GPUs at once. The most expensive idle asset in the industry is a GPU waiting on the network.

That reality has defined the past two years of networking innovation. The battle is over the fabric, not the silicon. NVLink binds 72 Blackwell GPUs into a single 130 TB/s domain, so tightly coupled the rack behaves as one giant GPU. Ultra Ethernet, InfiniBand, and optical circuit switching all compete over the layers above it. The entire discipline exists to keep GPUs from waiting on each other—and until recently, it was a problem solved inside one building.

This story is part of a paid subscription. Please subscribe for immediate access.

Subscribe Now
Already a member? Login here
Already a member? Log in here