Back to News Feed
NVIDIA Blog7d agoJesse Clayton

How XPUs Meet a World-Class AI Factory

In the modern era of artificial intelligence, the true measure of success is not found in the theoretical peak performance of a single chip, but in the relentless, continuous output of an AI factory. For hyperscalers and AI-native enterprises, the economics of intelligence are dictated by a rigorous set of metrics: tokens generated per second, energy efficiency measured in tokens per watt, the total cost per token, and the unwavering reliability of system uptime.

To achieve this, infrastructure must be conceptualized and constructed as a cohesive, high-performance factory rather than a fragmented collection of individual accelerators. Companies currently developing custom XPUs (accelerator processing units) face a daunting reality: the design of the silicon is only a fraction of the challenge. The true hurdle lies in the surrounding ecosystem—scale-up and scale-out networking, rack-scale architecture, production-grade software, and a resilient supply chain.

The Complexity of the AI Factory

Building an AI factory from the ground up is an immensely expensive and complex undertaking. For many, this complexity acts as a fundamental barrier to market entry. By attempting to reinvent the entire stack, companies often lose precious time and resources.

The solution lies in a hybrid approach: combining the unique innovation of custom XPUs with the proven, mature infrastructure of industry leaders. This is the core philosophy behind NVLink Fusion, a program designed to connect custom XPUs to NVIDIA’s world-class AI infrastructure. By leveraging this established foundation, developers can focus their engineering talent on the specific innovations that matter most, while mitigating the risks associated with building out a full-scale data center architecture.

Unlocking Performance Through Scale-Up

Modern AI workloads—ranging from trillion-parameter models to complex mixture-of-experts architectures and agentic AI—demand a scale-up fabric that can keep pace with the compute. If the networking layer fails to match the speed of the silicon, utilization rates plummet and the cost per token skyrockets.

A superior scale-up networking solution must excel across three critical pillars:

  • Delivered Performance: This encompasses end-to-end network throughput, integrated in-network compute capabilities, and a mature software stack that ensures seamless operation.
  • Factory Resiliency: The infrastructure must support continuous health monitoring, real-time telemetry, and component-level serviceability, allowing for maintenance without interrupting the factory’s output.
  • Platform Maturity: By utilizing a proven technology stack with a track record of large-scale deployments, operators can significantly reduce the operational risks associated with new, unproven architectures.

The NVLink Advantage

NVLink Fusion integrates XPUs directly into the NVIDIA NVLink scale-up domain. The sixth-generation NVLink technology provides high-bandwidth, low-latency connectivity across a 72-XPU domain. Compared to off-the-shelf Ethernet alternatives, this architecture delivers end-to-end latency that is 3x lower and a packet rate that is 10x higher.

Furthermore, the integration of NVIDIA NVLink-C2C allows XPUs to connect to NVIDIA Vera CPUs or other ecosystem-compatible processors. This interface offers up to 6x the energy efficiency of traditional PCIe connections, effectively removing the bottlenecks between control and compute that often hinder agentic AI systems.

A Proven Ecosystem for Rapid Deployment

Teams developing custom XPUs often underestimate the sheer scale of the integration effort required to move from a lab prototype to a data center deployment. This process involves sourcing high-speed interfaces, designing compute and switch trays, validating rack-level power and cooling, and managing a complex web of suppliers.

NVLink Fusion provides a standardized platform that allows teams to bypass these hurdles. By utilizing the NVIDIA MGX rack-scale architecture, developers gain access to a supply chain already optimized for high-performance systems.

"NVLink Fusion gives customers the ability to choose the CPU architecture, the performance level, the software capabilities that best meet their needs for the workloads that they care about," says Tim Wilson, vice president and general manager of data center silicon engineering at Intel.

Manufacturing partners are already leaning into this ecosystem. With the Vera Rubin NVL72 architecture, for example, manufacturers are achieving near-total automation in system assembly, an investment that can be directly leveraged by any XPU utilizing NVLink Fusion.

Managing Risk Through Infrastructure Standardization

AI factory planning is a long-lead process that cannot wait for the finalization of silicon. Facility design, cooling requirements, and power procurement must be established years in advance. Locking a data center into a single chip architecture creates significant schedule risk.

NVLink Fusion addresses this by providing a unified, flexible architecture. Because XPU-based systems and standard GPU-based systems share the same rack footprints, cooling, and power delivery systems, operators can move forward with facility construction while deferring the final decision on their silicon mix.

"The value of the NVLink Fusion program is that customers can deploy their rack-level solution with the NVIDIA GPU, and then they can decouple the development of their XPU and put it at a different pace," notes Vince Hu, corporate senior vice president and general manager of the data center and computing business group at MediaTek.

Designed and Validated as a Factory

Factory buildouts are unforgiving; mistakes in the planning phase can lead to catastrophic rework costs. To prevent this, NVLink Fusion aligns with the NVIDIA DSX reference architecture, which co-designs buildings, power, cooling, and compute as a single entity.

Through the NVIDIA Omniverse DSX AI Factory Blueprint, partners can utilize digital twins to model their facilities and technology stacks before a single piece of hardware is installed. At the rack level, this focus on efficiency continues: reference compute trays are designed for 100% liquid cooling, eliminating fans and cables, and allowing for the removal of individual trays without disrupting the operation of the entire rack.

The Software Foundation

Hardware is only as effective as the software that orchestrates it. The NVLink Fusion ecosystem is supported by a robust software suite, including:

  • NVIDIA NCCL: For high-performance distributed workloads.
  • NVIDIA Dynamo and NIXL: For advanced disaggregation.
  • NVIDIA Mission Control: For comprehensive cluster management, telemetry, and debugging.

By integrating these tools, operators can manage a heterogeneous mix of AI infrastructure as a single, coordinated system. NVLink Fusion represents a pivotal shift in the industry, enabling hyperscalers and AI-native companies to build unified, semi-custom AI factories that combine the specialized strengths of multiple builders into a powerful, world-class platform that no single company could achieve in isolation.