NVIDIA H100 SXM5 94 GB vs NVIDIA H20 Comparison

NVIDIA
GEFORCE

NVIDIA H100 SXM5 94 GB

CORE STATE GH100
VRAM 94 GB
CLOCK SPEED 1980 MHz
TDP 700 W
BUS WIDTH 5120 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

H20

CORE STATE GH100
VRAM 96 GB
CLOCK SPEED 1980 MHz
TDP 500 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2024

Analysis: NVIDIA H100 SXM5 94 GB vs NVIDIA H20

Head-to-Head Benchmarks

The database records no direct head-to-head benchmark results for the NVIDIA H100 SXM5 94 GB and the NVIDIA H20. Both GPUs share the same GH100 chip, the Hopper architecture, and the 5 nm TSMC process node, but their measured performance profiles diverge sharply due to different execution resource configurations and clock strategies. The recorded data shows the H100 SXM5 94 GB holds a decisive advantage in raw compute throughput, while the H20 counters with a wider memory bus and a higher base clock.

In FP32 compute, the H100 SXM5 94 GB delivers 66.91 TFLOPS against the H20's 39.54 TFLOPS. That is a 69.2% advantage for the H100 in single-precision floating-point work, a gap that reflects the H100's 16,896 shading units versus 9,984 on the H20. The H100's texture rate of 1,045.4 GTexel/s also dwarfs the H20's 617.8 GTexel/s, a difference of 69.2% as well, driven by 528 TMUs compared to 312. Pixel rate is identical at 47.52 GPixel/s on both cards, since each uses the same 24 ROPs.

The FP16 comparison widens further. The H100 SXM5 94 GB records 267.6 TFLOPS with a 4:1 ratio, while the H20 reaches 79.07 TFLOPS with a 2:1 ratio. That places the H100 at 238.4% higher FP16 throughput, or roughly 3.38 times the H20's figure. The difference stems not only from the H100's larger shader and tensor core counts (528 tensor cores versus 312) but also from the H100's more aggressive FP16 ratio, which allows double the arithmetic per clock cycle relative to the H20's configuration.

Memory bandwidth is the one area where the H20 takes the lead. The H20's 6144-bit bus width produces 4.03 TB/s of bandwidth, while the H100 SXM5 94 GB, despite its 94 GB capacity, uses a 5120-bit bus for 3.36 TB/s. That gives the H20 a 19.9% bandwidth advantage. The H20 also carries slightly more memory, 96 GB versus 94 GB, and its base clock runs at 1830 MHz compared to the H100's 1350 MHz, a 35.6% higher idle-to-load floor. Boost clocks match at 1980 MHz, and memory clock is identical at 1313 MHz with 5.3 Gbps effective.

Power consumption flips the efficiency story. The H20 draws a 500 W TDP against the H100 SXM5 94 GB's 700 W TDP, a 28.6% reduction in thermal design power. The suggested power supply drops from 1100 W to 900 W accordingly. The H20's lower power envelope, combined with its higher base clock, indicates a design tuned for sustained operation at reduced energy cost, even though its peak compute ceiling is lower.

Where Each One Wins

The H100 SXM5 94 GB wins in every compute-bound scenario. Its FP32 rate of 66.91 TFLOPS positions it for high-precision simulation, scientific computing, and any workload that relies on single-precision floating-point math. The 267.6 TFLOPS FP16 figure, with a 4:1 ratio, makes it the stronger choice for mixed-precision training and inference where tensor cores can exploit the 4:1 arithmetic ratio. The texture rate of 1,045.4 GTexel/s also favors the H100 in workloads that stress texture sampling, such as certain rendering pipelines or data processing kernels that map to texture units.

The H20 wins in memory-bound scenarios. Its 4.03 TB/s bandwidth, delivered over a 6144-bit bus, exceeds the H100's 3.36 TB/s by 19.9%. For large model inference where the working set repeatedly streams from HBM3, the H20's higher bandwidth can reduce memory stalls despite its lower compute throughput. The 96 GB capacity also edges out the 94 GB on the H100, offering 2 GB more headroom for models that sit near the memory ceiling. The H20's 500 W TDP and 1830 MHz base clock suggest it can sustain memory-heavy operations at lower power draw, which may matter in dense server deployments where thermal and power budgets constrain the number of accelerators per node.

The pixel rate tie at 47.52 GPixel/s means neither card differentiates on rasterization output, but both are server modules with no display outputs, so that metric has limited practical relevance. The H20's higher base clock could translate to faster wake-from-idle and steadier low-utilization throughput, but the H100's larger execution resources dominate once the workload saturates the GPU.

FAQ

Q: Which GPU has higher FP32 compute performance?

A: The NVIDIA H100 SXM5 94 GB delivers 66.91 TFLOPS in FP32, which is 69.2% higher than the NVIDIA H20's 39.54 TFLOPS.

Q: How does memory bandwidth compare between the two?

A: The H20 has a 6144-bit memory bus and achieves 4.03 TB/s, while the H100 SXM5 94 GB uses a 5120-bit bus for 3.36 TB/s. The H20 leads by 19.9%.

Q: What is the difference in FP16 performance?

A: The H100 SXM5 94 GB records 267.6 TFLOPS FP16 with a 4:1 ratio, compared to the H20's 79.07 TFLOPS with a 2:1 ratio. The H100 is roughly 3.38 times faster in FP16 throughput.

Q: Do both cards have the same memory type and clock speed?

A: Yes, both use HBM3 memory with a 1313 MHz clock and 5.3 Gbps effective data rate. They differ in bus width and total capacity: 94 GB on the H100 and 96 GB on the H20.

Q: What are the power consumption figures?

A: The H100 SXM5 94 GB has a 700 W TDP with a suggested 1100 W power supply. The H20 has a 500 W TDP with a suggested 900 W power supply.

Q: Are there any identical specifications?

A: Yes, both use the GH100 chip on TSMC's 5 nm process with 80,000 million transistors and an 814 mm² die size. Both have 24 ROPs, a 1980 MHz boost clock, a 47.52 GPixel/s pixel rate, and a PCIe 5.0 x16 bus interface. They are both SXM modules with no display outputs.

Specification Differences

The two GPUs diverge on nearly every execution resource. The H100 SXM5 94 GB carries 16,896 shading units, 528 TMUs, and 528 tensor cores, while the H20 has 9,984 shading units, 312 TMUs, and 312 tensor cores. That translates to a 69.2% higher FP32 rate and a 69.2% higher texture rate for the H100. The H20's base clock is 1830 MHz versus 1350 MHz on the H100, a 35.6% difference, though both boost to 1980 MHz.

Memory configuration differs in capacity and bus width: 94 GB on a 5120-bit bus for the H100 versus 96 GB on a 6144-bit bus for the H20. Bandwidth follows the bus width, with the H20 at 4.03 TB/s and the H100 at 3.36 TB/s. The FP16 ratio also differs, 4:1 on the H100 and 2:1 on the H20, producing 267.6 TFLOPS versus 79.07 TFLOPS. Power specifications are distinct: 700 W TDP with a suggested 1100 W PSU for the H100, and 500 W TDP with a suggested 900 W PSU for the H20. The H100 lists an 8-pin EPS power connector, while the H20 records none. The H20 lists DirectX, OpenGL, and Vulkan APIs as N/A, while the H100 leaves these fields unset.

Architecture Differences

Both accelerators are built on the Hopper architecture using the GH100 chip, fabricated at TSMC on a 5 nm process with 80,000 million transistors on an 814 mm² die, yielding a transistor density of 98.3 million per square millimeter. The core architectural difference lies in how each SKU configures that chip. The H100 SXM5 94 GB activates more of the GH100's execution resources: 16,896 shading units, 528 TMUs, and 528 tensor cores. The H20 disables a portion of the chip, retaining 9,984 shading units, 312 TMUs, and 312 tensor cores, which accounts for its lower FP32 and FP16 throughput.

The memory subsystem also differs architecturally. The H20 uses a wider 6144-bit HBM3 interface, enabling 4.03 TB/s, while the H100 SXM5 94 GB uses a 5120-bit interface for 3.36 TB/s. Both run HBM3 at the same 1313 MHz clock with 5.3 Gbps effective data rate. The H20's wider bus and slightly larger 96 GB capacity suggest a design optimized for memory capacity and bandwidth per watt, given its 500 W TDP. The H100's smaller bus but higher compute density points to a design favoring arithmetic throughput over memory streaming.

Clock behavior diverges at the base level: the H20 idles up to 1830 MHz, while the H100 sits at 1350 MHz, yet both reach 1980 MHz under boost. This indicates the H20 can operate at higher sustained low-load clocks while consuming less power, a trait suited to inference workloads with intermittent compute bursts. The H100's lower base clock and higher TDP align with sustained heavy compute where boost clocks dominate. Both share the same pixel rate, 47.52 GPixel/s, and both are SXM modules with no display outputs, reflecting their server-oriented roles. The release dates differ, with the H100 SXM5 94 GB entering production on 2023-03-20 and the H20 on 2024-01-31, though both remain Active and share the same predecessor and successor lineage: Server Ada and Server Blackwell, respectively.

DETAILED SPECIFICATIONS

SPECIFICATION
H100 SXM5 94 GB
H20
Core Specs
Shading Units
16,896
9,984 -40.9%
Shaders
16,896
9,984 -40.9%
TMUs
528
312 -40.9%
ROPs
24
24 0.0%
SM Count
132
78 -40.9%
Clocks
Base Clock
1350 MHz
1830 MHz
Boost Clock
1980 MHz
1980 MHz
Memory Clock
1313 MHz 5.3 Gbps effective
1313 MHz 5.3 Gbps effective
Memory
Memory Size
94 GB
96 GB
VRAM (MB)
96,256
98,304 +2.1%
Memory Type
HBM3
HBM3
Memory Bus
5120 bit
6144 bit
Bandwidth
3.36 TB/s
4.03 TB/s
Cache
L1 Cache
256 KB (per SM)
256 KB (per SM)
L2 Cache
50 MB
60 MB
Performance
Pixel Rate
47.52 GPixel/s
47.52 GPixel/s
Texture Rate
1,045.4 GTexel/s
617.8 GTexel/s
FP32 (TFLOPS)
66.91 TFLOPS
39.54 TFLOPS
FP64 (TFLOPS)
33.45 TFLOPS (1:2)
19.77 TFLOPS (1:2)
FP16 (TFLOPS)
267.6 TFLOPS (4:1)
79.07 TFLOPS (2:1)
AI/RT
Tensor Cores
528
312 -40.9%
Power
TDP
700 W
500 W
TDP (W)
700
500 -28.6%
Suggested PSU
1100 W
900 W
Power Connectors
8-pin EPS
—
Architecture
Architecture
Hopper
Hopper
GPU Name
GH100
GH100
Generation
Server Hopper (Hxx)
Server Hopper (Hxx)
Process Size
5 nm
5 nm
Transistors
80,000 million
80,000 million
Die Size
814 mm²
814 mm²
Foundry
TSMC
TSMC
Density
98.3M / mm²
98.3M / mm²
API Support
OpenCL
3.0
3.0
CUDA
9.0
9.0
Physical
Slot Width
SXM Module
SXM Module
Outputs
No outputs
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Production
Active
Active
Predecessor
Server Ada
Server Ada
Successor
Server Blackwell
Server Blackwell
View H100 SXM5 94 GB Details View H20 Details