NVIDIA H20 NVL16 vs NVIDIA N1 16SM Comparison
NVIDIA H20 NVL16
N1 16SM
Analysis: NVIDIA H20 NVL16 vs NVIDIA N1 16SM
Where Each One Wins
The recorded data places the NVIDIA H20 NVL16 and the NVIDIA N1 16SM in very different roles within a server environment. The H20 NVL16 is a Hopper-generation accelerator built around the GH100 chip, designed for massive parallel compute. The N1 16SM is a Blackwell 2.0 IGP (integrated graphics processor) built around the GB20B chip, intended for a different class of workload. Neither part has recorded benchmark scores in the database, so the win split is determined entirely by their architectural and specification profiles.
The H20 NVL16 wins on raw compute throughput. It delivers 39.54 TFLOPS of FP32 performance and 79.07 TFLOPS of FP16 performance with a 2:1 ratio. The N1 16SM delivers 9.609 TFLOPS of FP32 and 9.609 TFLOPS of FP16 with a 1:1 ratio. For any workload that scales with floating-point operations, the H20 NVL16 is the clear winner, offering roughly 4.1 times the FP32 throughput and 8.2 times the FP16 throughput. The H20 NVL16 also wins decisively on memory bandwidth: 4.03 TB/s from 96 GB of HBM3 on a 6144-bit bus, versus 273.2 GB/s from 128 GB of LPDDR5X on a 256-bit bus. That is a 14.7 times bandwidth advantage, which matters for memory-bound inference or training loops.
The N1 16SM wins on memory capacity and on integration. It carries 128 GB of LPDDR5X, which is 32 GB more than the H20 NVL16's 96 GB. For models or datasets that exceed 96 GB, the N1 16SM can hold them entirely in local memory. The N1 16SM also has a much higher boost clock: 2346 MHz versus 1980 MHz for the H20 NVL16. Its base clock is lower at 741 MHz versus 1830 MHz, but the boost behavior favors the N1 16SM. The N1 16SM includes 16 RT cores, which the H20 NVL16 lacks entirely (listed as null). The N1 16SM also has a display output (1x HDMI), while the H20 NVL16 has no display outputs. The N1 16SM is an IGP with no power connectors, meaning it draws power from the host platform, while the H20 NVL16 is an SXM module with a 400 W TDP and an 800 W suggested PSU.
The H20 NVL16 wins on texture rate (617.8 GTexel/s versus 300.3 GTexel/s) and on the number of shading units (9984 versus 2048), TMUs (312 versus 128), and tensor cores (312 versus 64). The N1 16SM wins on pixel rate (56.30 GPixel/s versus 47.52 GPixel/s) despite having far fewer ROPs at the same count (24 each). The H20 NVL16 has a larger die (814 mm² versus 382 mm²) and more transistors (80,000 million versus unknown for the N1 16SM), with a transistor density of 98.3M per mm² for the H20 NVL16.
The Verdict
The data indicates a straightforward choice based on workload type. For compute-heavy server tasks, particularly those that exploit FP16 tensor operations or require enormous memory bandwidth, the H20 NVL16 is the only sensible pick. Its 79.07 TFLOPS FP16 throughput and 4.03 TB/s bandwidth are in a different class from the N1 16SM's 9.609 TFLOPS and 273.2 GB/s. The H20 NVL16 also has 312 tensor cores versus 64, making it far more capable for matrix math. Anyone running large-scale training or high-throughput inference should choose the H20 NVL16.
For workloads that fit within 128 GB but not within 96 GB, or for systems where a discrete SXM module is not feasible, the N1 16SM is the choice. Its integrated design, lack of power connectors, and single HDMI output suggest it is meant for embedded or edge deployments where the host CPU and GPU share a package. The N1 16SM's higher boost clock (2346 MHz) and its 16 RT cores give it advantages in latency-sensitive or graphics-adjacent tasks, despite its lower raw throughput. The N1 16SM also has a newer architecture (Blackwell 2.0 versus Hopper) and a later release date (2026 versus 2025), which may matter for software feature support.
The H20 NVL16 is a server accelerator with no display outputs, a 400 W TDP, and an 800 W suggested PSU. The N1 16SM is an IGP with a single HDMI output, no power connectors, and an unknown TDP. These are not competing products in the same socket or price class; they serve different physical and electrical footprints. The verdict is that the H20 NVL16 wins for compute density and bandwidth, while the N1 16SM wins for capacity and integration simplicity.
Head-to-Head Benchmarks
The database contains no recorded head-to-head benchmark scores for these two parts, so the comparison relies on the specification-level measurements that are present. The largest single advantage for the H20 NVL16 is memory bandwidth: 4.03 TB/s versus 273.2 GB/s, a 14.7 times difference. This is the kind of gap that dominates any workload that streams data through the GPU, such as large batch inference or embedding lookups. The H20 NVL16 also leads in FP16 compute by a factor of 8.2 (79.07 TFLOPS versus 9.609 TFLOPS), which is critical for transformer models that use mixed precision. Its FP32 lead is 4.1 times (39.54 versus 9.609), still substantial for traditional HPC codes.
The N1 16SM counters with a higher boost clock: 2346 MHz versus 1980 MHz, a 18.5% advantage. This does not translate into higher throughput because the N1 16SM has only 2048 shading units versus 9984, but it does indicate better per-core efficiency at peak frequency. The N1 16SM also has a higher pixel rate: 56.30 GPixel/s versus 47.52 GPixel/s, a 18.5% advantage, despite having the same 24 ROPs. This suggests the N1 16SM's ROPs are clocked higher and may be more efficient per unit.
Texture rate goes the other way: the H20 NVL16 delivers 617.8 GTexel/s versus 300.3 GTexel/s, a 2.1 times advantage, because it has 312 TMUs versus 128. The H20 NVL16's transistor count of 80,000 million versus unknown for the N1 16SM, and its die size of 814 mm² versus 382 mm², indicate a much larger and more complex chip. The H20 NVL16 uses HBM3 memory on a 6144-bit bus, while the N1 16SM uses LPDDR5X on a 256-bit bus. The memory clock differs: 1313 MHz (5.3 Gbps effective) for the H20 NVL16 versus 1067 MHz (8.5 Gbps effective) for the N1 16SM. The effective data rate per pin is higher on the N1 16SM, but the total bandwidth is far lower due to the narrower bus.
In tensor core count, the H20 NVL16 leads 312 to 64, a 4.9 times advantage. In RT cores, the N1 16SM leads 16 to zero (null for the H20 NVL16). In shading units, the H20 NVL16 leads 9984 to 2048, a 4.9 times advantage. The H20 NVL16's texture rate of 617.8 GTexel/s exceeds the N1 16SM's 300.3 GTexel/s by 2.1 times, while the N1 16SM's pixel rate of 56.30 GPixel/s exceeds the H20 NVL16's 47.52 GPixel/s by 18.5%. These are the measurable wins in each direction.
FAQ
Q: Which GPU has more memory bandwidth?
A: The NVIDIA H20 NVL16 has 4.03 TB/s from 96 GB of HBM3 on a 6144-bit bus. The NVIDIA N1 16SM has 273.2 GB/s from 128 GB of LPDDR5X on a 256-bit bus. The H20 NVL16 leads by a factor of 14.7.
Q: Does the N1 16SM support ray tracing?
A: Yes. The NVIDIA N1 16SM includes 16 RT cores. The NVIDIA H20 NVL16 lists no RT cores in its specification, so it does not have this capability.
Q: Which GPU has a higher boost clock?
A: The NVIDIA N1 16SM has a boost clock of 2346 MHz, while the NVIDIA H20 NVL16 has a boost clock of 1980 MHz. The N1 16SM's boost clock is 18.5% higher.
Q: What are the memory types and sizes?
A: The H20 NVL16 uses 96 GB of HBM3 with a 6144-bit bus. The N1 16SM uses 128 GB of LPDDR5X with a 256-bit bus. The N1 16SM has 32 GB more capacity, but the H20 NVL16 has far more bandwidth.
Q: Which GPU has more tensor cores?
A: The NVIDIA H20 NVL16 has 312 tensor cores, while the NVIDIA N1 16SM has 64. The H20 NVL16 leads by a factor of 4.9.
Q: Are both GPUs compatible with PCIe 5.0?
A: Yes. Both the NVIDIA H20 NVL16 and the NVIDIA N1 16SM use a PCIe 5.0 x16 bus interface.
Architecture Differences
The two processors come from different NVIDIA architectures. The H20 NVL16 uses the Hopper architecture with the GH100 chip, part of the Server Hopper (Hxx) generation. The N1 16SM uses the Blackwell 2.0 architecture with the GB20B chip, part of the Blackwell IGP (N1x) generation. Both are manufactured on a 5 nm process at TSMC, but the H20 NVL16 has a die size of 814 mm² and 80,000 million transistors, while the N1 16SM has a die size of 382 mm² and an unknown transistor count. The H20 NVL16's transistor density is 98.3M per mm².
The H20 NVL16 is built as an SXM module, a discrete accelerator with a 400 W TDP and an 800 W suggested PSU. It has no display outputs and no power connectors listed. The N1 16SM is an IGP with no power connectors, a single HDMI output, and an unknown TDP. The H20 NVL16 uses HBM3 memory, while the N1 16SM uses LPDDR5X. The H20 NVL16 has 9984 shading units, 312 TMUs, and 24 ROPs, with 312 tensor cores and no RT cores. The N1 16SM has 2048 shading units, 128 TMUs, and 24 ROPs, with 64 tensor cores and 16 RT cores.
The H20 NVL16's FP16 performance is listed as 79.07 TFLOPS with a 2:1 ratio, meaning it doubles FP32 throughput when using FP16. The N1 16SM's FP16 performance is 9.609 TFLOPS with a 1:1 ratio, meaning no throughput gain from FP16. This architectural difference is significant for AI workloads that rely on FP16 or mixed precision. The H20 NVL16's memory clock is 1313 MHz with 5.3 Gbps effective data rate, while the N1 16SM's memory clock is 1067 MHz with 8.5 Gbps effective. The N1 16SM's higher effective data rate per pin reflects its LPDDR5X design, but the total bandwidth is far lower.
The release dates differ: the H20 NVL16 was released in September 2025, while the N1 16SM was released in May 2026. The H20 NVL16 has a predecessor (Server Ada) and a successor (Server Blackwell), while the N1 16SM has neither listed. Both parts are marked as Active in production status.
Specification Differences
The following specifications differ between the two parts, based only on the recorded fields:
- Chip: GH100 for the H20 NVL16, GB20B for the N1 16SM.
- Architecture: Hopper for the H20 NVL16, Blackwell 2.0 for the N1 16SM.
- Generation: Server Hopper (Hxx) for the H20 NVL16, Blackwell IGP (N1x) for the N1 16SM.
- Transistors: 80,000 million for the H20 NVL16, unknown for the N1 16SM.
- Die size: 814 mm² for the H20 NVL16, 382 mm² for the N1 16SM.
- Transistor density: 98.3M / mm² for the H20 NVL16, null for the N1 16SM.
- Base clock: 1830 MHz for the H20 NVL16, 741 MHz for the N1 16SM.
- Boost clock: 1980 MHz for the H20 NVL16, 2346 MHz for the N1 16SM.
- Memory clock: 1313 MHz (5.3 Gbps effective) for the H20 NVL16, 1067 MHz (8.5 Gbps effective) for the N1 16SM.
- Memory size: 96 GB for the H20 NVL16, 128 GB for the N1 16SM.
- Memory type: HBM3 for the H20 NVL16, LPDDR5X for the N1 16SM.
- Memory bus width: 6144 bit for the H20 NVL16, 256 bit for the N1 16SM.
- Memory bandwidth: 4.03 TB/s for the H20 NVL16, 273.2 GB/s for the N1 16SM.
- Shading units: 9984 for the H20 NVL16, 2048 for the N1 16SM.
- TMUs: 312 for the H20 NVL16, 128 for the N1 16SM.
- RT cores: null for the H20 NVL16, 16 for the N1 16SM.
- Tensor cores: 312 for the H20 NVL16, 64 for the N1 16SM.
- Pixel rate: 47.52 GPixel/s for the H20 NVL16, 56.30 GPixel/s for the N1 16SM.
- Texture rate: 617.8 GTexel/s for the H20 NVL16, 300.3 GTexel/s for the N1 16SM.
- FP32 performance: 39.54 TFLOPS for the H20 NVL16, 9.609 TFLOPS for the N1 16SM.
- FP16 performance: 79.07 TFLOPS (2:1) for the H20 NVL16, 9.609 TFLOPS (1:1) for the N1 16SM.
- TDP: 400 W for the H20 NVL16, unknown for the N1 16SM.
- Slot width: SXM Module for the H20 NVL16, IGP for the N1 16SM.
- Power connectors: null for the H20 NVL16, None for the N1 16SM.
- Suggested PSU: 800 W for the H20 NVL16, null for the N1 16SM.
- Display outputs: No outputs for the H20 NVL16, 1x HDMI for the N1 16SM.
- Release date: 2025-09-01 for the H20 NVL16, 2026-05-31 for the N1 16SM.
- Predecessor: Server Ada for the H20 NVL16, null for the N1 16SM.
- Successor: Server Blackwell for the H20 NVL16, null for the N1 16SM.
- ROPs: 24 for both, no difference.
- Manufacturer, foundry, process node, bus interface, APIs, production status, launch MSRP: identical or null for both parts.