At GTC 2026, Nvidia introduced three new systems simultaneously: the Groq LPX inference rack, the Vera ETL256 CPU rack, and the STX storage reference architecture. These launches extend Nvidia's portfolio beyond its traditional GPU compute core into low-latency inference, CPU orchestration, and storage layers, signaling a systematic redefinition of AI infrastructure boundaries.
The introduction of these three systems further strengthens Nvidia's dominance in the AI infrastructure supply chain, driving greater market concentration around its ecosystem while expanding its competitive position.
Groq LPX: the speed play
The Groq LPX system has attracted the most market attention. The product was developed in less than four months after Nvidia secured Groq-related IP licensing and brought in its core team in a deal reportedly worth around US$20 billion. The LP30 chip used in the system is manufactured using Samsung's 4nm process**,** which allows Nvidia to bypass capacity constraints associated with TSMC's advanced N3 process while avoiding reliance on high-bandwidth memory (HBM), which remains in tight supply. As a result, the solution offers incremental capacity advantages that are difficult for competitors to replicate.
Architecturally, the LPX rack deeply integrates Groq's LP30 chips with Nvidia GPUs and introduces Attention-FFN Disaggregation (AFD) technology. By separating attention computation and feed-forward network (FFN) workloads across GPUs and LPUs, the system significantly reduces inference latency, providing a new optimization path for highly interactive large model applications.
According to semiconductor research firm SemiAnalysis, while LPUs may not offer clear cost advantages for large-scale token processing, they deliver substantial value in latency-sensitive applications, making them a critical component within disaggregated architectures.
Vera and STX: dense compute meets storage
Meanwhile, the Vera ETL256 CPU rack integrates 256 CPUs into a single liquid-cooled rack, enabling full interconnectivity within the rack via copper cabling topology. This directly addresses the increasingly prominent CPU supply bottleneck as AI workloads scale.
The STX storage reference architecture further extends Nvidia's influence from compute and networking into the storage layer, completing its full-stack AI infrastructure strategy.
Nvidia also named major storage vendors — including DDN, Dell, Hewlett Packard Enterprise (HPE), IBM, NetApp, Supermicro, and VAST Data — as supporters of the STX standard, reinforcing its strategy of leveraging ecosystem partnerships to strengthen its influence over industry standards.
From GPU supplier to platform provider
A recent report from SemiAnalysis noted that the launch of these three systems sends a clear message: Nvidia is transforming from a GPU supplier into a full-stack AI infrastructure platform provider. Its expansion now spans inference optimization, CPU density scaling, and storage orchestration, potentially reshaping competition across the AI hardware supply chain. Its collaboration with Groq follows an IP licensing and talent integration model rather than a traditional acquisition, enabling rapid access to key technologies and teams while accelerating product commercialization.
From a technical perspective, the LP30 chip adopts a monolithic die design and features 500MB of on-chip SRAM, delivering 1.2 PFLOPS of performance at FP8 precision. This marks a significant improvement over Groq's first-generation LPU (230MB SRAM, 750 TFLOPS INT8), driven largely by the transition from GlobalFoundries' GF16 process to Samsung's SF4 node.
Under the AFD architecture, GPUs handle attention workloads requiring dynamic access to KV cache and fully leverage HBM resources, while LPUs execute statically scheduled FFN computations, maximizing low-latency advantages. The two are connected via all-to-all communication for token distribution and aggregation, using a pipelined mechanism to minimize communication latency and further enhance overall inference efficiency.
Article edited by Jerry Chen