Growing pains: how distributed AI training changes the network between datacenters
Large-scale AI training has already escaped the confines of a single datacenter. Google said Gemini was trained synchronously across clusters in multiple locations; Microsoft has connected AI data centres in Wisconsin and Georgia into what it describes as one distributed AI supercomputer; AWS has connected AI compute clusters across wide areas to allow Anthropic to build Claude models; Meta has built high-capacity datacenter interconnects to support model training; and CoreWeave and Google Cloud recently announced cross-cloud training with a private interconnect, with Azure likely to follow later in the year. This growing geographic spread of model training is partly due to the limits of single datacenters or campus clusters, which can be strained as they try to meet the compute and power demands of new models. Cisco estimates that training models today can require clusters with tens of thousands of GPUs. By 2030, the largest individual frontier training runs could draw 4-16GW of power, according to researchers at Epoch AI. …
You're reading a preview. The full article is published by The Register on their website.
Read the full story on The Register
