When one datacenter is no longer enough
Training the largest AI models is no longer simply a question of adding more GPUs. As clusters grow, power availability is becoming a hard physical constraint, forcing infrastructure teams to consider how a single training workload can operate across multiple data centres. That creates a very different networking problem. Traditional datacenter interconnect was designed to move traffic between sites. AI training demands something more exacting: huge, synchronous flows, minimal packet loss and tightly coordinated communication between GPUs. If one part of the cluster stalls, the impact can ripple across the entire job. In this interview, The Register’s Tim Phillips speaks with Rakesh Chopra, SVP, Silicon and Systems Architecture, Cisco Fellow, about what Cisco calls the “Scale-Across” imperative and why it believes distributed AI infrastructure requires a new approach to routing. Rakesh explains how the demands of AI are changing the role of the network, from simple transport to an integral part of the compute system. …
You're reading a preview. The full article is published by The Register on their website.
Read the full story on The Register

