{"id":17263,"date":"2026-08-07T16:39:47","date_gmt":"2026-08-07T16:39:47","guid":{"rendered":"https:\/\/dmsretail.com\/RetailNews\/scaling-the-future-why-ethernet-is-the-backbone-of-ai-supercomputing\/"},"modified":"2026-08-07T16:39:47","modified_gmt":"2026-08-07T16:39:47","slug":"scaling-the-future-why-ethernet-is-the-backbone-of-ai-supercomputing","status":"publish","type":"post","link":"https:\/\/dmsretail.com\/RetailNews\/scaling-the-future-why-ethernet-is-the-backbone-of-ai-supercomputing\/","title":{"rendered":"Scaling the future: Why Ethernet is the backbone of AI Supercomputing"},"content":{"rendered":"<p> <p><a href=\"https:\/\/dmsretail.com\/online-workshops-list\/\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-496\" src=\"https:\/\/dmsretail.com\/RetailNews\/wp-content\/uploads\/2022\/05\/RETAIL-ONLINE-TRAINING-728-X-90.png\" alt=\"Retail Online Training\" width=\"729\" height=\"91\" srcset=\"https:\/\/dmsretail.com\/RetailNews\/wp-content\/uploads\/2022\/05\/RETAIL-ONLINE-TRAINING-728-X-90.png 729w, https:\/\/dmsretail.com\/RetailNews\/wp-content\/uploads\/2022\/05\/RETAIL-ONLINE-TRAINING-728-X-90-300x37.png 300w\" sizes=\"auto, (max-width: 729px) 100vw, 729px\" \/><\/a><\/p><br \/>\n<\/p>\n<div>\n<p>The rapid evolution of artificial intelligence is fundamentally changing how we architect data centers. As AI models grow more complex, the industry is shifting focus from individual server performance to the data center\u2019s interconnected fabric. Two factors are driving this shift: expanding training clusters and inference workloads that now demand cluster-level performance.<\/p>\n<p>For training, frontier models require large numbers of GPUs, and cluster sizes now exceed the capacity of a single data hall. Clusters span multiple data centers connected by wide-area networks, and the infrastructure must scale to support hundreds of thousands of GPUs across broad geographic regions.<\/p>\n<p>Inference is also transforming the infrastructure. Frontier models, even at FP4 precision, now surpass the capacity of a single GPU. The push for faster token serving is increasing demand for larger inference clusters, matching the same coordinated, high-performance networking as training clusters.<\/p>\n<p>Taken together, these changes make the network more than a connectivity layer. The network is becoming the system-level fabric that determines how much of the AI infrastructure can be used, how quickly jobs complete, and how predictably inference can be served.<\/p>\n<h2>The network is now the system<\/h2>\n<p>A few years ago, GPU compute power was the primary bottleneck for AI model training. As distributed training has scaled, that constraint has shifted decisively from compute to network \u2014 GPU communication now determines overall cluster efficiency. For instance, Meta\u2019s production data shows that in large-scale Deep Neural Network training runs, network overhead accounts for up to 60% of total training iteration time \u2014 a share that increases with cluster size.<\/p>\n<p>This is why we think about the next phase of AI networking as a continuum. Scale-up connects accelerators inside a server or rack, where proprietary technologies such as NVLink and emerging approaches such as UALink, have focused on extremely low latency and high bandwidth. Scale-out connects racks and pods into larger training clusters, where InfiniBand has historically been a common choice for high-performance fabrics. Scale-across connects clusters, storage, front-end networks, and data centers, where Ethernet is already the operational foundation.<\/p>\n<p>At scale, for training and inference alike, the network matters as much as the compute itself. The question is no longer whether AI needs specialized networking behavior. It does. The real question is whether we deliver that behavior through a patchwork of proprietary fabrics, or through one common Ethernet foundation that can grow across the whole continuum.<\/p>\n<h2>Why Ethernet becomes the common foundation<\/h2>\n<p>Proprietary networking solutions have long dominated high-performance computing, but they introduce vendor lock-in and limit scalability across diverse hardware. InfiniBand still has a role in plenty of AI deployments, but the direction of the industry isn\u2019t in question \u2014 Ethernet is becoming the predominant networking technology for AI infrastructure. Embracing Ethernet puts you on the right operating model from day one: open, interoperable, and built to scale across many domains.<\/p>\n<p>Cisco is championing an \u201cEthernet-first\u201d strategy for AI for three core reasons:<\/p>\n<ul>\n<li><strong>Open Standards and Interoperability<\/strong>: Ethernet enables organizations to integrate components from multiple vendors. This flexibility is essential for future-proofing data centers as AI hardware evolves.<\/li>\n<li><strong>Unmatched Scalability<\/strong>: InfiniBand\u2019s proprietary fabric management struggles above ~tens of thousands of GPUs, requiring complex workarounds as clusters grow. Ethernet has no such ceiling \u2014 hyperscalers have already leveraged decades of mature switching architecture and standards-based tooling to validate Ethernet-based clusters at hundreds of thousands of GPUs across multiple data centers.<\/li>\n<li><strong>Investment Protection \u2014 With a Learning Curve<\/strong>: Ethernet builds on familiar infrastructure \u2014 existing switching platforms, management tooling, and a broad engineering talent pool. That foundation matters. But AI fabric operations is not a straight extension of enterprise networking. RoCEv2 and RDMA introduce new failure modes; congestion management (PFC, ECN, buffer tuning) requires careful calibration to avoid GPU stalls; and telemetry at hundred-thousand-GPU scale demands purpose-built tooling. Skills transfer partially, not fully. The advantage over InfiniBand is a more open, composable operational model.<\/li>\n<\/ul>\n<p>That operating model matters because no two AI environments look alike. Training wants ultra-low latency and predictable collective communication. Inference wants QoS that accounts for load, location, and cost. A multi-site deployment wants fault tolerance, tenant isolation, and deterministic telemetry stretched across a much bigger failure domain. Ethernet gives you one foundation that can flex to all those requirements \u2014 instead of stitching together a separate technology island for each one.<\/p>\n<h2>What Ethernet must deliver for AI<\/h2>\n<p>To earn its place as the common AI fabric, Ethernet must handle what makes AI traffic different. This traffic is synchronized, bursty, and expensive to stall. Fall behind on the network, and GPUs sit idle. Let congestion spread, and job completion times stretch out. Take too long to heal a failure, and large jobs lose efficiency.<\/p>\n<p>First up: intelligent load balancing. AI fabrics must spread traffic across many paths without sacrificing single-flow performance, keeping pace with modern NIC bandwidth and putting the whole topology to work. Weighted adaptive routing, multipath transport, source-routed and path-aware forwarding \u2014 these all serve the same goal: react to hotspots fast, without introducing instability.<\/p>\n<p>Second: congestion control and reliable delivery. That means fast congestion detection, precise notification, and recovery that doesn\u2019t throw away useful work. Packet trimming, local link repair, selective retransmission, ordered and unordered retransmission, header optimization \u2014 none of these are standalone features. They\u2019re all doing the same job: keeping AI traffic moving when the fabric is under pressure.<\/p>\n<p>Third: isolation and service assurance. AI clusters increasingly run multiple tenants and multiple jobs side by side, and a fault or noisy neighbor in one must never degrade another\u2019s performance. Delivering that guarantee without heavy per-job configuration \u2014 especially as workloads move off InfiniBand \u2014 is what separates a fabric that merely connects GPUs from one that can be trusted to run production AI at scale.<\/p>\n<p>This is exactly where standards like UEC, ESUN, and Multipath Reliable Connection (MRC) earn their keep. They\u2019re defining how Ethernet picks up the AI-specific behavior it needs \u2014 congestion control, multipath operation, reliable transport, path awareness, telemetry, interoperability \u2014 without giving up the openness that made Ethernet the right choice to begin with.<\/p>\n<h2>Ethernet plus P4 programmability: The multiplying factor<\/h2>\n<p>In AI, networking standards are evolving rapidly. New protocols such as UEC Transport and MRC are being developed to address challenges in AI and ML traffic, including congestion control, efficient use of fabric bandwidth, packet ordering, and telemetry.<\/p>\n<p>New standards such as these often require capabilities in networking that can only be met in the new ASIC generation which is typically available eighteen months later at best.<\/p>\n<p>Historically, this assumption made sense. ASICs are built to a fixed specification, and once set, changes are not possible. If a standard was not included in the original design, it cannot be supported by the chip.<\/p>\n<p>AI is challenging this model.<\/p>\n<p>AI workload requirements are evolving at an unprecedented pace. UEC and MRC are not minor updates; each introduces significant new capabilities required at the switching ASIC level. These changes are arriving faster than traditional silicon development cycles can support.<\/p>\n<p>This presents a significant challenge for customers building infrastructure today. Delaying an AI buildout to wait for new hardware is not feasible. The cost of delay, including lost training runs, reduced competitiveness, and idle capital, is substantial.<\/p>\n<h2>Cisco\u2019s Silicon One was designed to address this challenge.<\/h2>\n<p>Since Silicon One is programmable in P4: it is not limited to the initial set of applications envisioned when the ASIC was designed. P4 enables engineers and customers to define packet processing in software, separating network logic from physical hardware. When a new standard emerges, such as a revised UEC congestion response or new MRC capabilities, we can deliver these updates in software on existing hardware, often within weeks or months rather than waiting for the next product cycle.<\/p>\n<p>That\u2019s the multiplying factor. Standards set the direction for the ecosystem, but P4 programmability decides how fast customers see the benefit on real infrastructure. It also means customer-specific behavior \u2014 scheduler-aware policy, topology-specific routing, tenant isolation \u2014 doesn\u2019t have to wait on a fixed-function silicon roadmap.<\/p>\n<h2>Where Cisco Silicon One fits in<\/h2>\n<p>Cisco Silicon One sits right at the intersection of high-performance Ethernet, emerging AI networking standards, and P4 programmability. That\u2019s not a coincidence \u2014 AI networks need both performance and adaptability at once: performance to keep GPUs fed, adaptability to keep up with standards and customer requirements that are still very much in motion.<\/p>\n<p>We have demonstrated this capability multiple times across real, production-relevant features:<\/p>\n<ul>\n<li><strong>Packet Trimming:<\/strong> Rather than dropping packets outright during congestion events, packet trimming preserves the header while discarding the payload, allowing receivers to selectively request retransmission of only the missing data. This significantly reduces unnecessary full-flow retransmissions and improves throughput under load\u2014delivered on existing Silicon One hardware through a P4 software update, with no silicon changes required.<\/li>\n<li><strong>Full MRC Support:<\/strong> Multipath Reliable Connection introduces a comprehensive suite of load balancing and congestion control mechanisms purpose-built for AI and ML traffic patterns. Because Silicon One is P4-programmable, we were able to implement the complete MRC capability set\u2014including its multipath load balancing and congestion response algorithms\u2014without waiting for a new ASIC generation.<\/li>\n<li><strong>Weighted Adaptive Routing:<\/strong> AI workloads generate highly bursty, asymmetric traffic that can rapidly create hotspots across a fabric. Weighted Adaptive Routing dynamically distributes flows across available paths based on real-time congestion metrics, assigning weights to steer traffic away from congested links and maximize fabric utilization. Delivering this capability on existing hardware requires only a P4 software update.<\/li>\n<li><strong>Multi-tenant and Multi-job Isolation:<\/strong> Most of the AI clusters, with the exception of foundational model training, support multiple tenants and multiple jobs within each tenant. Enforcing tenant- and job-level isolation policies to prevent cross-communication is a critical service that the network operator must provide. As customers migrate from InfiniBand to Ethernet, supporting an efficient solution that minimizes configuration and network churn whenever a tenant and a job are scheduled onto the cluster becomes a key differentiator.<\/li>\n<\/ul>\n<p>MRC is a good illustration of why Cisco\u2019s SRv6 investment pays off here. Its switch-side requirements \u2014 SRv6 uSID forwarding, packet trimming, deterministic path-pinned telemetry \u2014 line up with capabilities we\u2019ve already built through SRv6 and programmable Silicon One forwarding. And because that forwarding behavior is programmable, both these capabilities and customer-specific extensions can keep evolving hardware you\u2019ve already deployed, as the spec matures.<\/p>\n<p>This is not a theoretical advantage; it is the difference between telling a customer \u201cwe support that today\u201d and \u201cwe\u2019ll have silicon for that in 12 to 18 months.\u201d In AI infrastructure, this distinction is critical.<\/p>\n<p>The broader point is that programmability is essential. Given the rapid evolution of AI networking standards, it is the only viable architectural approach. Continuing to build inflexible ASICs to a fixed specification and relying on market stability is increasingly difficult to justify as new protocols are introduced.<\/p>\n<h2>The path forward<\/h2>\n<p>The future of AI depends not only on server silicon but also on the fabric connecting those servers. As we enter the era of large, multi-rack clusters, the industry needs a robust, flexible networking foundation.<\/p>\n<p>That foundation comes down to a single, open building block \u2014 Ethernet \u2014 flexible enough to address three distinct scaling challenges at once:<\/p>\n<ul>\n<li><strong>Scale-up, ultra-optimized:<\/strong> within the rack, Ethernet must match the raw, low-latency performance of dedicated scale-up fabrics between GPUs.<\/li>\n<li><strong>Scale-out, performant and reliable:<\/strong> across racks and pods, it must sustain full throughput and reliable delivery as training clusters scale out to tens of thousands of GPUs.<\/li>\n<li><strong>Scale-across, fault-tolerant and QoS-aware:<\/strong> across data centers and geographies, it must preserve job isolation and predictable performance as thousands of GPUs training clusters \u2014 and increasingly, inference clusters \u2014 span the wide area network.<\/li>\n<\/ul>\n<p>As Ethernet evolves, it solves for all three \u2014 without giving up the open, standards-based ecosystem that makes it the right long-term choice for AI infrastructure.<\/p>\n<p>Cisco is committed to delivering this foundation. By prioritizing open standards, high-performance silicon, and intelligent automation, we ensure tomorrow\u2019s infrastructure can support today\u2019s breakthroughs.<\/p>\n<p>To be clear, this isn\u2019t Ethernet instead of innovation. It\u2019s Ethernet as the open foundation innovation builds on \u2014 multiplied by P4 programmability and delivered in platforms like Cisco Silicon One \u2014 so AI networks can evolve just as fast as the workloads riding on them.<\/p>\n<blockquote>\n<\/blockquote>\n<p>Additional resources:<\/p>\n<\/p><\/div>\n<p><p><a href=\"https:\/\/dmsretail.com\/online-workshops-list\/\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-496\" src=\"https:\/\/dmsretail.com\/RetailNews\/wp-content\/uploads\/2022\/05\/RETAIL-ONLINE-TRAINING-728-X-90.png\" alt=\"Retail Online Training\" width=\"729\" height=\"91\" srcset=\"https:\/\/dmsretail.com\/RetailNews\/wp-content\/uploads\/2022\/05\/RETAIL-ONLINE-TRAINING-728-X-90.png 729w, https:\/\/dmsretail.com\/RetailNews\/wp-content\/uploads\/2022\/05\/RETAIL-ONLINE-TRAINING-728-X-90-300x37.png 300w\" sizes=\"auto, (max-width: 729px) 100vw, 729px\" \/><\/a><\/p><br \/><\/p>\n","protected":false},"excerpt":{"rendered":"<p>The rapid evolution of artificial intelligence is fundamentally changing how we architect data centers. As AI models grow more complex, the industry is shifting focus [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":17264,"comment_status":"","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5],"tags":[],"class_list":["post-17263","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-technology"],"_links":{"self":[{"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/posts\/17263","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/comments?post=17263"}],"version-history":[{"count":0,"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/posts\/17263\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/media\/17264"}],"wp:attachment":[{"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/media?parent=17263"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/categories?post=17263"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/dmsretail.com\/RetailNews\/wp-json\/wp\/v2\/tags?post=17263"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}