NVIDIA H20 vs NVIDIA H200 NVL Comparison

NVIDIA
GEFORCE

NVIDIA H20

CORE STATE GH100
VRAM 96 GB
CLOCK SPEED 1980 MHz
TDP 500 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2024
VS
NVIDIA
GEFORCE

H200 NVL

CORE STATE GH100
VRAM 141 GB
CLOCK SPEED 1785 MHz
TDP 600 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2024

PERFORMANCE BENCHMARKS

geekbench_opencl
N/A
334,891

Analysis: NVIDIA H20 vs NVIDIA H200 NVL

Head-to-Head Benchmarks

The database records no direct head-to-head benchmark comparisons between the NVIDIA H20 and the NVIDIA H200 NVL. The H20 has no benchmark entries, while the H200 NVL carries a single Geekbench OpenCL score of 334,891. This places the H200 NVL at the 100th percentile among all GPUs in the database, a top-tier position. The H20, by contrast, sits at the 50th percentile with an average benchmark score of zero, meaning no performance data has been captured for it.

Without direct comparisons, the nearest rivals of the H200 NVL provide context for its recorded score. The H200 NVL trails the NVIDIA B300 SXM6 AC by 9.4%, as that part scores 369,831. It also sits 3.1% behind the NVIDIA B200, which scores 345,482. Moving down the list, the H200 NVL leads the AMD Instinct MI300X by 5.3%, as that competitor scores 317,994, and it is 13.2% ahead of the NVIDIA L40S, which scores 295,763. These deltas show the H200 NVL operating in a dense performance band, with the B200 and B300 above it and the MI300X and L40S below it.

The H20's lack of benchmark data means its relative standing cannot be quantified from the database. Its percentile of 50 is a default positioning rather than a measured result. The H200 NVL's percentile of 100 reflects its recorded OpenCL score, placing it at the top of the database's distribution. For any workload represented by Geekbench OpenCL, the H200 NVL is the stronger part by the only available metric.

FAQ

Q: What is the recorded benchmark score for each GPU?

A: The H200 NVL has a Geekbench OpenCL score of 334,891. The H20 has no recorded benchmark score, with an average benchmark score of zero.

Q: How does the H200 NVL compare to its nearest rivals?

A: The H200 NVL is 3.1% behind the NVIDIA B200, 9.4% behind the NVIDIA B300 SXM6 AC, 5.3% ahead of the AMD Instinct MI300X, and 13.2% ahead of the NVIDIA L40S.

Q: Which GPU has a higher percentile ranking?

A: The H200 NVL is at the 100th percentile among all GPUs in the database. The H20 is at the 50th percentile.

Q: Do both GPUs use the same chip and architecture?

A: Yes, both use the GH100 chip and the Hopper architecture, manufactured by NVIDIA on a 5 nm process at TSMC.

Q: What are the memory capacities of the two GPUs?

A: The H20 has 96 GB of HBM3 memory. The H200 NVL has 141 GB of HBM3e memory.

Q: What is the form factor difference between the two?

A: The H20 is an SXM Module, while the H200 NVL is a dual-slot card with a length of 267 mm (10.5 inches) and a height of 111 mm (4.4 inches), using an 8-pin EPS power connector.

The Verdict

The data supports selecting the H200 NVL for any compute role where the Geekbench OpenCL score is representative. It is the only one of the two with a measured performance result, and that result ranks at the 100th percentile. The H20 has no benchmark data, so its actual performance cannot be assessed from the database.

The H200 NVL also leads on raw specifications that matter for dense compute workloads. It delivers 60.32 TFLOPS of FP32 throughput versus 39.54 TFLOPS for the H20, and 120.6 TFLOPS of FP16 (2:1) versus 79.07 TFLOPS. Its memory subsystem is larger and faster: 141 GB of HBM3e with 4.89 TB/s of bandwidth, compared to 96 GB of HBM3 at 4.03 TB/s. The H200 NVL also carries more shading units (16,896 versus 9,984) and more tensor cores (528 versus 312).

The H20 does hold advantages in certain clock and rate figures. Its base clock is 1830 MHz versus 1365 MHz, and its boost clock is 1980 MHz versus 1785 MHz. It also has a higher pixel rate at 47.52 GPixel/s versus 42.84 GPixel/s. However, these do not offset the H200 NVL's larger compute and memory resources in the absence of any H20 benchmark result.

For a builder choosing between these two Hopper parts, the H200 NVL is the defensible pick based on recorded data. The H20 cannot be recommended on performance grounds because no measurements exist for it. The H200 NVL's lower clock speeds do not compensate for its substantial lead in core count, TFLOPS, and memory bandwidth.

Specification Differences

The two GPUs share their foundation: both use the GH100 chip, Hopper architecture, and a 5 nm TSMC process with 80,000 million transistors on an 814 mm² die. Both have a 6144-bit memory bus and no display outputs. Both use PCIe 5.0 x16 and have 24 ROPs.

The H20 runs at a base clock of 1830 MHz and a boost clock of 1980 MHz, with memory at 1313 MHz (5.3 Gbps effective). The H200 NVL runs at a base clock of 1365 MHz and a boost clock of 1785 MHz, with memory at 1593 MHz (6.4 Gbps effective). The H200 NVL's lower core clocks are offset by its higher memory clock.

Compute resources differ sharply. The H20 has 9,984 shading units, 312 TMUs, and 312 tensor cores. The H200 NVL has 16,896 shading units, 528 TMUs, and 528 tensor cores. Pixel rates are close: 47.52 GPixel/s for the H20, 42.84 GPixel/s for the H200 NVL. Texture rates favor the H200 NVL at 942.5 GTexel/s versus 617.8 GTexel/s.

Memory capacity and type also differ. The H20 uses 96 GB of HBM3 with 4.03 TB/s bandwidth. The H200 NVL uses 141 GB of HBM3e with 4.89 TB/s bandwidth. The H20 is an SXM Module with a 500 W TDP and a suggested PSU of 900 W. The H200 NVL is a dual-slot card with an 8-pin EPS connector, a 600 W TDP, and a suggested PSU of 1000 W.

Architecture Differences

Both GPUs are built on the Hopper architecture and use the GH100 chip, so the fundamental architecture is identical. They share the same transistor count, die size, and process node. The differences come from how the chip is configured and clocked.

The H200 NVL activates more of the GH100's resources. It has 16,896 shading units and 528 tensor cores, which is a substantially higher utilization of the chip's compute array compared to the H20's 9,984 shading units and 312 tensor cores. The H200 NVL also pairs this with faster HBM3e memory, delivering 4.89 TB/s over the same 6144-bit bus, whereas the H20 uses HBM3 at 4.03 TB/s.

Clocking strategy diverges as well. The H20 runs higher core clocks, with a 1980 MHz boost, likely to compensate for its reduced core count. The H200 NVL runs at a 1785 MHz boost, but its larger core count and faster memory give it higher aggregate throughput. The FP32 and FP16 figures confirm this: 60.32 TFLOPS and 120.6 TFLOPS for the H200 NVL, versus 39.54 TFLOPS and 79.07 TFLOPS for the H20.

The H200 NVL's memory advantage extends beyond bandwidth. Its 141 GB capacity is 47% larger than the H20's 96 GB, which matters for models that exceed the smaller pool. The H20's higher pixel rate (47.52 GPixel/s versus 42.84 GPixel/s) reflects its higher core clock, but this is a minor difference in the context of server compute workloads, where tensor throughput and memory dominate.

Where Each One Wins

The H200 NVL wins on every recorded performance metric and on the majority of the specification sheet. Its Geekbench OpenCL score of 334,891 is the only measured benchmark between the two, and it ranks at the 100th percentile. It leads in FP32 compute by 20.78 TFLOPS, in FP16 compute by 41.53 TFLOPS, in texture rate by 324.7 GTexel/s, and in memory bandwidth by 0.86 TB/s. Its 141 GB memory capacity is also larger, and it has more shading units and tensor cores.

The H20 wins on clock speeds, with a 465 MHz higher base clock and a 195 MHz higher boost clock. It also has a higher pixel rate by 4.68 GPixel/s. It uses less power at 500 W versus 600 W, and it occupies an SXM form factor, which is a different physical mounting standard than the H200 NVL's dual-slot card. Its 96 GB memory capacity is smaller, but for workloads that fit within that pool, the higher core clocks could provide an advantage in latency-sensitive tasks.

For dense compute, large model inference, or any workload that scales with tensor cores and memory capacity, the H200 NVL is the clear choice. For workloads that are sensitive to core clock speed and fit within a smaller memory footprint, the H20's higher clocks and lower power draw may be preferable. The H20's lack of benchmark data means those theoretical advantages cannot be confirmed with measured results, but the specifications support that use case.

DETAILED SPECIFICATIONS

SPECIFICATION
H20
H200 NVL
Core Specs
Shading Units
9,984
16,896 +69.2%
Shaders
9,984
16,896 +69.2%
TMUs
312
528 +69.2%
ROPs
24
24 0.0%
SM Count
78
132 +69.2%
Clocks
Base Clock
1830 MHz
1365 MHz
Boost Clock
1980 MHz
1785 MHz
Memory Clock
1313 MHz 5.3 Gbps effective
1593 MHz 6.4 Gbps effective
Memory
Memory Size
96 GB
141 GB
VRAM (MB)
98,304
144,384 +46.9%
Memory Type
HBM3
HBM3e
Memory Bus
6144 bit
6144 bit
Bandwidth
4.03 TB/s
4.89 TB/s
Cache
L1 Cache
256 KB (per SM)
256 KB (per SM)
L2 Cache
60 MB
50 MB
Performance
Pixel Rate
47.52 GPixel/s
42.84 GPixel/s
Texture Rate
617.8 GTexel/s
942.5 GTexel/s
FP32 (TFLOPS)
39.54 TFLOPS
60.32 TFLOPS
FP64 (TFLOPS)
19.77 TFLOPS (1:2)
30.16 TFLOPS (1:2)
FP16 (TFLOPS)
79.07 TFLOPS (2:1)
120.6 TFLOPS (2:1)
AI/RT
Tensor Cores
312
528 +69.2%
Power
TDP
500 W
600 W
TDP (W)
500
600 +20.0%
Suggested PSU
900 W
1000 W
Power Connectors
8-pin EPS
Architecture
Architecture
Hopper
Hopper
GPU Name
GH100
GH100
Generation
Server Hopper (Hxx)
Server Hopper (Hxx)
Process Size
5 nm
5 nm
Transistors
80,000 million
80,000 million
Die Size
814 mm²
814 mm²
Foundry
TSMC
TSMC
Density
98.3M / mm²
98.3M / mm²
API Support
OpenCL
3.0
3.0
CUDA
9.0
9.0
Physical
Slot Width
SXM Module
Dual-slot
Length
267 mm 10.5 inches
Height
111 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Production
Active
Active
Predecessor
Server Ada
Server Ada
Successor
Server Blackwell
Server Blackwell
View H20 Details View H200 NVL Details