NVIDIA H20 NVL16 vs NVIDIA H200 NVL Comparison

NVIDIA
GEFORCE

NVIDIA H20 NVL16

CORE STATE GH100
VRAM 96 GB
CLOCK SPEED 1980 MHz
TDP 400 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

H200 NVL

CORE STATE GH100
VRAM 141 GB
CLOCK SPEED 1785 MHz
TDP 600 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2024

PERFORMANCE BENCHMARKS

geekbench_opencl
N/A
334,891

Analysis: NVIDIA H20 NVL16 vs NVIDIA H200 NVL

FAQ

Q: What is the core specification difference between the NVIDIA H20 NVL16 and the NVIDIA H200 NVL?

A: The H20 NVL16 uses 9,984 shading units, 312 tensor cores, and 96 GB of HBM3 memory with 4.03 TB/s bandwidth. The H200 NVL uses 16,896 shading units, 528 tensor cores, and 141 GB of HBM3e memory with 4.89 TB/s bandwidth.

Q: Which GPU has the higher boost clock?

A: The H20 NVL16 has a boost clock of 1980 MHz, while the H200 NVL has a boost clock of 1785 MHz.

Q: What is the power draw difference between the two cards?

A: The H20 NVL16 has a TDP of 400 W with a suggested PSU of 800 W, while the H200 NVL has a TDP of 600 W with a suggested PSU of 1000 W.

Q: How does the memory type differ between the two?

A: The H20 NVL16 uses HBM3 memory, while the H200 NVL uses HBM3e memory. The H200 NVL also has more memory capacity at 141 GB versus 96 GB.

Q: What are the physical dimensions and interface requirements?

A: The H20 NVL16 is an SXM Module, while the H200 NVL is a dual-slot card measuring 267 mm in length and 111 mm in height. Both use PCIe 5.0 x16, but the H200 NVL requires an 8-pin EPS power connector.

Q: What is the benchmark performance of the H200 NVL?

A: The H200 NVL records a Geekbench OpenCL score of 334,891, placing it in the 100th percentile of all GPUs in the database.

The Verdict

The data clearly separates these two Hopper-generation server accelerators by capability tier. The H200 NVL is the higher-performing part across every compute metric, delivering 60.32 TFLOPS FP32 versus 39.54 TFLOPS for the H20 NVL16. Its 16896 shading units and 528 tensor cores represent a 69% and 69% increase over the H20's 9984 shading units and 312 tensor cores, respectively.

The H20 NVL16 does hold advantages in clock speed and power efficiency. Its 1980 MHz boost clock is 10.9% higher than the H200's 1785 MHz boost, and its 400 W TDP is one-third lower than the H200's 600 W TDP. For deployments constrained by power delivery or thermal limits, the H20 NVL16 offers a lower-power path into Hopper architecture.

However, benchmark evidence is decisive. The H200 NVL's Geekbench OpenCL score of 334,891 places it in the 100th percentile of all GPUs, while the H20 NVL16's percentile rank stands at 50 with no recorded benchmark score. The H200 NVL outperforms the AMD Instinct MI300X by 5.3%, trails the NVIDIA B200 by 3.1%, and surpasses the NVIDIA L40S by 13.2%. The H20 NVL16 has no nearest rivals recorded in the database.

Organizations with memory-heavy workloads should favor the H200 NVL. Its 141 GB HBM3e capacity and 4.89 TB/s bandwidth exceed the H20's 96 GB HBM3 and 4.03 TB/s by substantial margins. The H200 NVL also has a longer production history, with a release date in November 2024 compared to the H20's September 2025 release.

The H20 NVL16 is the choice only when power constraints or clock-speed sensitivity dominate the selection criteria. Its 400 W TDP and 1980 MHz boost are the only metrics where it leads. For raw throughput, memory capacity, and sustained compute performance, the H200 NVL is the superior accelerator per the recorded data.

Head-to-Head Benchmarks

The database contains no direct head-to-head benchmark entries for these two GPUs, and the H20 NVL16 has no benchmark scores recorded. The H200 NVL, however, has a single Geekbench OpenCL score of 334,891 that provides a reference point against its nearest rivals.

The H200 NVL sits 3.1% behind the NVIDIA B200, which scores 345,482. This places the H200 NVL within striking distance of NVIDIA's next-generation Blackwell server part, despite being from the earlier Hopper generation. Against the AMD Instinct MI300X, the H200 NVL leads by 5.3% with a score of 317,994 for the AMD part. The H200 NVL also outperforms the NVIDIA L40S by 13.2%, with the L40S scoring 295,763.

The largest gap in the H200 NVL's rivalry set is the 9.4% deficit to the NVIDIA B300 SXM6 AC, which scores 369,831. This indicates that while the H200 NVL is a top-tier performer in the 100th percentile, it does not sit at the absolute apex of the database's server GPU rankings.

For the H20 NVL16, the absence of benchmark data means its 50th percentile rank is based on specifications rather than measured performance. Its 39.54 TFLOPS FP32 throughput is 34% lower than the H200 NVL's 60.32 TFLOPS, and its 79.07 TFLOPS FP16 is also 34% lower than the H200's 120.6 TFLOPS. These compute deltas are consistent with the 69% difference in shading unit counts, indicating that the H200 NVL's architecture scales compute output nearly linearly with its larger execution resource pool.

The texture rate difference further illustrates the performance gap. The H200 NVL achieves 942.5 GTexel/s versus 617.8 GTexel/s for the H20 NVL16, a 52.5% advantage. Pixel rates are closer, with the H20 NVL16 at 47.52 GPixel/s versus 42.84 GPixel/s for the H200 NVL, a 10.9% edge for the H20 that stems from its higher clock speed.

Specification Differences

The two GPUs share foundational specifications but diverge significantly in execution resources and memory. Both use the GH100 chip, built on TSMC's 5 nm process with 80,000 million transistors and an 814 mm² die size, yielding a transistor density of 98.3 million transistors per square millimeter.

Compute Resources: The H200 NVL carries 16,896 shading units, 528 TMUs, and 528 tensor cores. The H20 NVL16 has 9,984 shading units, 312 TMUs, and 312 tensor cores. Both have 24 ROPs. This gives the H200 NVL a 69% advantage in each of these execution resource categories.

Clock Speeds: The H20 NVL16 has a base clock of 1830 MHz and a boost clock of 1980 MHz. The H200 NVL has a base clock of 1365 MHz and a boost clock of 1785 MHz. The H20 NVL16's boost clock is 10.9% higher.

Memory: The H20 NVL16 uses 96 GB of HBM3 with a 6144-bit bus and 4.03 TB/s bandwidth, running at 1313 MHz (5.3 Gbps effective). The H200 NVL uses 141 GB of HBM3e with the same 6144-bit bus but 4.89 TB/s bandwidth, running at 1593 MHz (6.4 Gbps effective). The H200 NVL offers 46.9% more capacity and 21.3% more bandwidth.

Performance Rates: The H20 NVL16 delivers 47.52 GPixel/s pixel rate and 617.8 GTexel/s texture rate. The H200 NVL delivers 42.84 GPixel/s and 942.5 GTexel/s, respectively. FP32 compute is 39.54 TFLOPS for the H20 versus 60.32 TFLOPS for the H200. FP16 compute is 79.07 TFLOPS (2:1) versus 120.6 TFLOPS (2:1).

Power and Physical: The H20 NVL16 has a 400 W TDP, requires an 800 W suggested PSU, and uses an SXM Module form factor. The H200 NVL has a 600 W TDP, requires a 1000 W suggested PSU, uses a dual-slot form factor with an 8-pin EPS power connector, and measures 267 mm by 111 mm.

APIs: Both cards report N/A for DirectX, OpenGL, and Vulkan support, consistent with their server-oriented design and lack of display outputs.

Release and Status: Both are Active in production. The H200 NVL was released in November 2024, while the H20 NVL16 was released in September 2025. Both list Server Ada as their predecessor and Server Blackwell as their successor.

Architecture Differences

Both GPUs are built on the Hopper architecture using the GH100 chip, so their architectural foundations are identical. The differences emerge in how the chip is configured and paired with memory.

Process and Die: Both use TSMC's 5 nm process with an identical 814 mm² die and 80,000 million transistors. There is no manufacturing node advantage for either part.

Execution Configuration: The H200 NVL activates a larger portion of the GH100 die. Its 16,896 shading units and 528 tensor cores represent a fuller implementation of the chip's compute array. The H20 NVL16 uses 9,984 shading units and 312 tensor cores, effectively a reduced configuration that trades compute density for lower power draw and higher clock speeds.

Memory Architecture: The H20 NVL16 pairs the GH100 with HBM3, while the H200 NVL uses HBM3e. Both maintain a 6144-bit memory bus, but the HBM3e modules in the H200 NVL operate at a higher effective data rate of 6.4 Gbps versus 5.3 Gbps, and the H200 NVL carries 141 GB versus 96 GB. This combination gives the H200 NVL both more capacity and higher bandwidth.

Clock Strategy: The H20 NVL16's higher base and boost clocks (1830/1980 MHz versus 1365/1785 MHz) indicate a binning strategy that prioritizes frequency over core count. This yields a higher pixel rate despite fewer ROPs at the same count. The H200 NVL's lower clocks are offset by its 69% larger execution resource pool, which dominates throughput metrics.

Feature Parity: Both cards have no display outputs, no supported graphics APIs, and no ray tracing cores listed. They are compute-only accelerators with identical bus interfaces (PCIe 5.0 x16) and no power connector listed for the H20 NVL16 (the H200 NVL uses an 8-pin EPS).

The architectural relationship is clear: the H20 NVL16 is a clock-optimized, power-constrained variant of the same GH100 silicon, while the H200 NVL is a full-resource, memory-maximized configuration. The 400 W TDP difference (600 W versus 400 W) reflects this divergence in execution resource activation and memory subsystem requirements.

Where Each One Wins

The H200 NVL wins in every compute throughput category. Its FP32 output of 60.32 TFLOPS exceeds the H20 NVL16's 39.54 TFLOPS by 52.5%. FP16 follows the same pattern at 120.6 TFLOPS versus 79.07 TFLOPS. Texture rate favors the H200 at 942.5 GTexel/s versus 617.8 GTexel/s. Its 16,896 shading units and 528 tensor cores provide the execution resources for large-scale parallel workloads, and its 141 GB HBM3e memory with 4.89 TB/s bandwidth supports datasets that would exceed the H20's 96 GB capacity.

The H200 NVL is the clear choice for memory-bound applications. Large language model inference, multi-tenant inference serving, and datasets exceeding 96 GB require the H200 NVL's additional 45 GB of capacity. Its 4.89 TB/s bandwidth also reduces memory access bottlenecks relative to the H20's 4.03 TB/s.

The H20 NVL16 wins where power and clock speed matter. Its 400 W TDP is 33% lower than the H200 NVL's 600 W, and its 1980 MHz boost clock is 10.9% higher. These attributes make it suitable for power-constrained server chassis or facilities with limited power delivery per slot. The H20 NVL16's 47.52 GPixel/s pixel rate also exceeds the H200 NVL's 42.84 GPixel/s, a minor win in pixel throughput.

The H20 NVL16 also wins on release recency. Its September 2025 release date is nearly a year after the H200 NVL's November 2024 launch, which may matter for procurement cycles seeking the newest production status.

Benchmark standing is entirely one-sided. The H200 NVL holds a Geekbench OpenCL score of 334,891 and sits in the 100th percentile of all GPUs, with rival comparisons showing a 5.3% lead over the AMD Instinct MI300X and a 13.2% lead over the NVIDIA L40S. The H20 NVL16 has no recorded benchmark score and sits at the 50th percentile. For any workload where measured performance is the deciding factor, the H200 NVL is the only choice supported by the data.

Workload split: Choose the H200 NVL for AI training, large-batch inference, scientific computing, and any task requiring more than 96 GB of memory. Choose the H20 NVL16 for power-limited deployments, clock-sensitive workloads, or environments where the 400 W TDP and SXM module form factor align with existing infrastructure. The H200 NVL's 600 W TDP and dual-slot physical design require more power headroom and chassis space.

DETAILED SPECIFICATIONS

SPECIFICATION
H20 NVL16
H200 NVL
Core Specs
Shading Units
9,984
16,896 +69.2%
Shaders
9,984
16,896 +69.2%
TMUs
312
528 +69.2%
ROPs
24
24 0.0%
SM Count
78
132 +69.2%
Clocks
Base Clock
1830 MHz
1365 MHz
Boost Clock
1980 MHz
1785 MHz
Memory Clock
1313 MHz 5.3 Gbps effective
1593 MHz 6.4 Gbps effective
Memory
Memory Size
96 GB
141 GB
VRAM (MB)
98,304
144,384 +46.9%
Memory Type
HBM3
HBM3e
Memory Bus
6144 bit
6144 bit
Bandwidth
4.03 TB/s
4.89 TB/s
Cache
L1 Cache
256 KB (per SM)
256 KB (per SM)
L2 Cache
60 MB
50 MB
Performance
Pixel Rate
47.52 GPixel/s
42.84 GPixel/s
Texture Rate
617.8 GTexel/s
942.5 GTexel/s
FP32 (TFLOPS)
39.54 TFLOPS
60.32 TFLOPS
FP64 (TFLOPS)
19.77 TFLOPS (1:2)
30.16 TFLOPS (1:2)
FP16 (TFLOPS)
79.07 TFLOPS (2:1)
120.6 TFLOPS (2:1)
AI/RT
Tensor Cores
312
528 +69.2%
Power
TDP
400 W
600 W
TDP (W)
400
600 +50.0%
Suggested PSU
800 W
1000 W
Power Connectors
—
8-pin EPS
Architecture
Architecture
Hopper
Hopper
GPU Name
GH100
GH100
Generation
Server Hopper (Hxx)
Server Hopper (Hxx)
Process Size
5 nm
5 nm
Transistors
80,000 million
80,000 million
Die Size
814 mm²
814 mm²
Foundry
TSMC
TSMC
Density
98.3M / mm²
98.3M / mm²
API Support
OpenCL
3.0
3.0
CUDA
9.0
9.0
Physical
Slot Width
SXM Module
Dual-slot
Length
—
267 mm 10.5 inches
Height
—
111 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Production
Active
Active
Predecessor
Server Ada
Server Ada
Successor
Server Blackwell
Server Blackwell
View H20 NVL16 Details View H200 NVL Details