Why AI Factories Need Great Interconnect IP, Not Just More Compute
AI infrastructure is often evaluated based on compute: how many accelerators, how much total performance, how fast models can train.
That works up to a point.
At scale, these systems behave fundamentally differently from traditional data center architectures. They are built for sustained, tightly coordinated workloads, where large numbers of accelerators operate as a single system rather than independent resources.
In this environment, adding compute does not guarantee higher performance.
What matters is whether that compute can be kept utilized—and that depends on how efficiently data can be delivered, exchanged, and synchronized across the system.
As AI factories have scaled, the limiting factor has shifted. Performance is no longer defined by peak compute capability, but by communication efficiency.
The constraint moves from compute to interconnect.
AI Factories Are Fundamentally Different Systems
AI factories are not just larger data centers, they operate very differently. Instead of loosely coupled workloads, they run large training jobs across tightly connected systems. Thousands of accelerators operate together, continuously exchanging data and staying synchronized over long periods.
This reflects a broader shift in data center architecture, from general-purpose infrastructure to systems built specifically for sustained AI workloads.
This fundamentally needs a new approach to interconnect, as covered in our latest white paper, "Building AI Factories with IP Solutions".
In this environment:
- Compute resources are interdependent
- Data exchange is continuous
- Synchronization is intrinsic to the workload
Under these conditions, efficiency is determined by the weakest link in the data path. If one part of the system cannot keep up, the entire system slows down.

Compute Doesn't Define Performance
It is easy to assume that faster accelerators will improve performance. But accelerators only create value when they are fully utilized, and that depends on how efficiently data reaches them.
- Bandwidth: Accelerators consume data at extremely high rates. If bandwidth is insufficient, compute stalls regardless of peak capability.
- Latency: Distributed training requires frequent coordination. Latency in collective operations directly reduces throughput.
- Topology: At scale, communication paths become more complex. Poor topology increases hop count, introduces contention, and amplifies latency.
As AI workloads have scaled, the industry has introduced new interconnect approaches, both within nodes and across systems, to address these pressures.
However, higher link speeds alone do not remove the constraint. Efficiency depends on how the interconnect behaves under sustained load.

These Constraints Originate in IP
Interconnect behavior is not defined at deployment. It is determined in the underlying IP.
Key characteristics are set by:
- SerDes design, which determines signaling efficiency, reach, and power
- Protocol implementation, which governs overhead and effective bandwidth
- Interface IP, which defines buffering, flow control, and latency
- Integration across chiplets, packages, and systems
These are architectural decisions made early in the design process. Once established, they constrain what the system can achieve.
As AI factories scale, the impact of these decisions becomes more visible. Tradeoffs that were previously localized begin to affect overall system behavior.
Explore Cadence IP
Why "Good Enough" Interconnect Fails at Scale
At smaller system sizes, an interconnect that meets basic requirements is often sufficient. That is not the case in AI factories.
As systems scale:
- Small inefficiencies increase synchronization costs
- Variability across links creates imbalance
- Local bottlenecks propagate across the system
A design that looks acceptable in isolation can become the primary limiter of system performance. Peak metrics don't tell you much here. What matters is how the system performs under sustained workloads.
Compute, Memory, and Interconnect Must Be Co-Optimized
Compute, memory, and interconnect are tightly coupled. Improving one without the others does not improve overall performance:
- More compute without enough bandwidth reduces utilization
- Faster links without lower latency do not improve coordination
- Local optimizations can introduce system-level inefficiencies
AI factory design requires a system-level approach, where these elements are developed together. Optimizing one dimension in isolation creates imbalance—and lost performance.

Where the Real Constraint Sits
AI factories make the limitation clear. Performance is not defined by how fast individual components are. It is defined by how efficiently the system moves and manages data. That is governed by the interconnect.
Because interconnect behavior is set at the IP level, achieving performance at scale depends on getting interconnect IP right from the start.
Scaling AI systems is not just a matter of adding more compute. It is a question of whether the system can sustain efficient data movement as it grows.
Read the "Building AI Factories With IP Solutions" white paper now.
Related Semiconductor IP
- Highly scalable performance for classic and generative on-device and edge AI solutions
- MIPI CSI-2 TX Controller
- 1.6T Ultra Ethernet Controller
- HBM4E PHY and controller
- HBM3 PHY
Related Blogs
- From GPUs to Memory Pools: Why AI Needs Compute Express Link (CXL)
- Why Hardware Monitoring Needs Infrastructure, Not Just Sensors
- MIPI: Not Just Mobile Any More
- AIoT. Why it's not just another tech buzzword.