NVIDIA B200 SXM6 vs NVIDIA H20 NVL16 Comparison

NVIDIA
GEFORCE

NVIDIA B200 SXM6

CORE STATE GB100
VRAM 180 GB
CLOCK SPEED 1830 MHz
TDP 1000 W
BUS WIDTH 8192 bit
ARCHITECTURE Blackwell
nm
PROCESS 5 nm
LAUNCH DATE 2024
VS
NVIDIA
GEFORCE

H20 NVL16

CORE STATE GH100
VRAM 96 GB
CLOCK SPEED 1980 MHz
TDP 400 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2025

Analysis: NVIDIA B200 SXM6 vs NVIDIA H20 NVL16

FAQ

Q: What are the core architectural differences between the NVIDIA B200 SXM6 and the NVIDIA H20 NVL16?

A: The B200 SXM6 uses the GB100 chip on the Blackwell architecture, built on a 5 nm process at TSMC with 208,000 million transistors on a 1628 mm² die. The H20 NVL16 uses the GH100 chip on the Hopper architecture, also on a 5 nm TSMC process, but with 80,000 million transistors on an 814 mm² die.

Q: How do the memory subsystems compare between the two accelerators?

A: The B200 SXM6 has 180 GB of HBM3e memory on an 8192-bit bus, delivering 8.19 TB/s bandwidth. The H20 NVL16 has 96 GB of HBM3 memory on a 6144-bit bus, delivering 4.03 TB/s bandwidth. That means the B200 has roughly double the memory capacity and over twice the bandwidth.

Q: What are the clock speed differences?

A: The B200 SXM6 runs a base clock of 120 MHz and a boost clock of 1830 MHz. The H20 NVL16 runs a base clock of 1830 MHz and a boost clock of 1980 MHz. Despite the B200's much lower base clock, its boost clock is only 150 MHz lower than the H20's.

Q: Which card has higher FP32 throughput?

A: The B200 SXM6 delivers 69.34 TFLOPS of FP32 performance, while the H20 NVL16 delivers 39.54 TFLOPS. The B200 is about 75% ahead in raw single-precision compute.

Q: What about FP16 performance?

A: The B200 SXM6 delivers 69.34 TFLOPS of FP16 with a 1:1 ratio. The H20 NVL16 delivers 79.07 TFLOPS of FP16 with a 2:1 ratio. In FP16, the H20 is actually ahead by roughly 14%, although the ratio difference means the B200's FP16 is full-rate while the H20's is half-rate relative to its FP32.

Q: What are the power requirements?

A: The B200 SXM6 has a TDP of 1000 W with a suggested PSU of 1400 W. The H20 NVL16 has a TDP of 400 W with a suggested PSU of 800 W.

The Verdict

The data shows two very different accelerators despite both being NVIDIA server modules. The B200 SXM6 is the higher-capability part in most raw compute and memory metrics, with double the memory size, more than double the bandwidth, and nearly double the FP32 throughput. Its transistor count is 208,000 million versus 80,000 million for the H20, and its die size is 1628 mm² versus 814 mm². The B200 also has more shading units (18944 vs 9984), more TMUs (592 vs 312), and more tensor cores (592 vs 312).

The H20 NVL16, however, wins in FP16 throughput and pixel rate. Its FP16 of 79.07 TFLOPS beats the B200's 69.34 TFLOPS, and its pixel rate of 47.52 GPixel/s beats the B200's 43.92 GPixel/s. The H20 also runs at higher clocks, with a boost of 1980 MHz versus 1830 MHz, and uses far less power at 400 W TDP versus 1000 W TDP.

For a workload dominated by FP32 compute or massive memory capacity and bandwidth, the B200 SXM6 is the clear choice. For FP16-heavy workloads where power draw matters, the H20 NVL16 delivers more FP16 per watt. The B200's launch MSRP is 34,999 USD, while the H20 has no recorded launch MSRP in the database. The B200's higher transistor density (127.8M / mm² vs 98.3M / mm²) also indicates a more advanced design implementation.

Head-to-Head Benchmarks

The database contains no recorded benchmark scores for either accelerator, so direct performance measurements are unavailable. However, the specification data provides a basis for comparing theoretical peak throughput.

In FP32 compute, the B200 SXM6 delivers 69.34 TFLOPS against the H20 NVL16's 39.54 TFLOPS. That is a 29.80 TFLOPS advantage for the B200, representing roughly a 75% lead. This is the largest single-metric gap in the comparison.

In texture rate, the B200's 1,083.4 GTexel/s more than doubles the H20's 617.8 GTexel/s. The B200's TMU count of 592 versus 312 directly explains this advantage.

In memory bandwidth, the B200's 8.19 TB/s is more than double the H20's 4.03 TB/s. The B200's 8192-bit bus width versus 6144-bit, combined with faster HBM3e memory at 2000 MHz (8 Gbps effective) versus HBM3 at 1313 MHz (5.3 Gbps effective), produces this gap.

The H20 NVL16 wins in FP16 throughput at 79.07 TFLOPS versus 69.34 TFLOPS. That is a 9.73 TFLOPS advantage, roughly 14% higher. The H20 achieves this with a 2:1 FP16 ratio, effectively doubling its FP32 rate, while the B200 runs FP16 at a 1:1 ratio, matching its FP32 rate.

The H20 also wins in pixel rate at 47.52 GPixel/s versus 43.92 GPixel/s. Both parts have 24 ROPs, but the H20's higher boost clock of 1980 MHz versus 1830 MHz gives it the edge in pixel throughput.

Specification Differences

| Field | NVIDIA B200 SXM6 | NVIDIA H20 NVL16 |

|---|---|---|

| Chip | GB100 | GH100 |

| Architecture | Blackwell | Hopper |

| Generation | Server Blackwell (Bxx) | Server Hopper (Hxx) |

| Process Node | 5 nm | 5 nm |

| Transistors | 208,000 million | 80,000 million |

| Die Size | 1628 mm² | 814 mm² |

| Transistor Density | 127.8M / mm² | 98.3M / mm² |

| Base Clock | 120 MHz | 1830 MHz |

| Boost Clock | 1830 MHz | 1980 MHz |

| Memory Size | 180 GB | 96 GB |

| Memory Type | HBM3e | HBM3 |

| Memory Bus Width | 8192 bit | 6144 bit |

| Memory Bandwidth | 8.19 TB/s | 4.03 TB/s |

| Memory Clock | 2000 MHz 8 Gbps effective | 1313 MHz 5.3 Gbps effective |

| Shading Units | 18944 | 9984 |

| TMUs | 592 | 312 |

| Tensor Cores | 592 | 312 |

| Pixel Rate | 43.92 GPixel/s | 47.52 GPixel/s |

| Texture Rate | 1,083.4 GTexel/s | 617.8 GTexel/s |

| FP32 | 69.34 TFLOPS | 39.54 TFLOPS |

| FP16 | 69.34 TFLOPS (1:1) | 79.07 TFLOPS (2:1) |

| TDP | 1000 W | 400 W |

| Suggested PSU | 1400 W | 800 W |

| Bus Interface | PCIe 6.0 x16 | PCIe 5.0 x16 |

| Release Date | 2024-10-31 | 2025-09-01 |

Architecture Differences

The B200 SXM6 and H20 NVL16 represent two distinct NVIDIA server generations. The B200 is built on Blackwell, the successor to Hopper, and its predecessor is listed as Server Hopper. The H20 is built on Hopper, with its predecessor listed as Server Ada and its successor as Server Blackwell. This places them on opposite sides of a generational boundary.

The B200 uses the GB100 chip, fabricated on a 5 nm TSMC process with 208,000 million transistors. The die measures 1628 mm², giving a transistor density of 127.8M per mm². The H20 uses the GH100 chip, also on a 5 nm TSMC process, but with 80,000 million transistors on an 814 mm² die, yielding a density of 98.3M per mm². The B200 has more than double the transistor count and a 2.6x higher density.

Memory architecture differs substantially. The B200 uses HBM3e with 180 GB capacity, while the H20 uses HBM3 with 96 GB capacity. The B200's bus width is 8192 bit versus 6144 bit, and its memory clock is 2000 MHz (8 Gbps effective) versus 1313 MHz (5.3 Gbps effective). The resulting bandwidth gap is 8.19 TB/s versus 4.03 TB/s.

The compute layouts differ in scale. The B200 has 18944 shading units, 592 TMUs, and 592 tensor cores. The H20 has 9984 shading units, 312 TMUs, and 312 tensor cores. Both have 24 ROPs. The B200's FP32 rate is 69.34 TFLOPS, while the H20's is 39.54 TFLOPS. In FP16, the B200 runs at 1:1 ratio delivering 69.34 TFLOPS, while the H20 runs at 2:1 ratio delivering 79.07 TFLOPS.

The B200 uses PCIe 6.0 x16, while the H20 uses PCIe 5.0 x16. Both are SXM modules with no display outputs. The B200's base clock of 120 MHz is unusually low, likely reflecting a power management strategy given its 1000 W TDP, while the H20's base clock of 1830 MHz is much higher despite its 400 W TDP.

Where Each One Wins

The B200 SXM6 wins in scenarios requiring maximum FP32 compute. Its 69.34 TFLOPS versus 39.54 TFLOPS makes it the stronger choice for workloads that rely on single-precision floating-point operations, such as certain scientific simulations and data processing tasks. The B200 also dominates in memory-heavy applications. With 180 GB versus 96 GB and 8.19 TB/s versus 4.03 TB/s, it can hold larger datasets and move them faster, which matters for large-scale model training and inference where memory capacity and bandwidth are bottlenecks.

The B200's texture rate of 1,083.4 GTexel/s versus 617.8 GTexel/s gives it a clear advantage in texture-bound operations, though neither part has display outputs, so this is relevant only for compute tasks that exercise texture units. Its PCIe 6.0 x16 interface also provides a newer interconnect compared to the H20's PCIe 5.0 x16.

The H20 NVL16 wins in FP16 throughput, delivering 79.07 TFLOPS versus 69.34 TFLOPS. Workloads that heavily use FP16 arithmetic, such as certain AI training and inference paths with half-precision formats, will see higher peak throughput on the H20. Its 2:1 FP16 ratio means it doubles its FP32 rate for half-precision work, while the B200's 1:1 ratio keeps FP16 at the same rate as FP32.

The H20 also wins in pixel rate at 47.52 GPixel/s versus 43.92 GPixel/s, although with no display outputs this advantage is limited to compute workloads that involve rasterization-style operations. The H20's higher boost clock of 1980 MHz versus 1830 MHz contributes to this win and also helps in any clock-sensitive compute tasks.

The H20's power efficiency is a significant differentiator. At 400 W TDP versus 1000 W, the H20 consumes 600 W less. For FP16 work, the H20 delivers more throughput per watt by a wide margin. The H20's suggested PSU of 800 W versus 1400 W also reflects lower system-level power requirements. Deployments with power constraints or dense multi-GPU configurations would favor the H20.

The B200's launch MSRP is 34,999 USD, while the H20 has no recorded launch MSRP in the database. The B200's release date is 2024-10-31, and the H20's is 2025-09-01. Both are listed as Active production status. The B200's successor is Server Rubin, while the H20's successor is Server Blackwell, which is the generation the B200 belongs to, indicating the H20 is one generation behind in NVIDIA's server roadmap.

DETAILED SPECIFICATIONS

SPECIFICATION
B200 SXM6
H20 NVL16
Core Specs
Shading Units
18,944
9,984 -47.3%
Shaders
18,944
9,984 -47.3%
TMUs
592
312 -47.3%
ROPs
24
24 0.0%
SM Count
148
78 -47.3%
Clocks
Base Clock
120 MHz
1830 MHz
Boost Clock
1830 MHz
1980 MHz
Memory Clock
2000 MHz 8 Gbps effective
1313 MHz 5.3 Gbps effective
Memory
Memory Size
180 GB
96 GB
VRAM (MB)
184,320
98,304 -46.7%
Memory Type
HBM3e
HBM3
Memory Bus
8192 bit
6144 bit
Bandwidth
8.19 TB/s
4.03 TB/s
Cache
L1 Cache
256 KB (per SM)
256 KB (per SM)
L2 Cache
126 MB
60 MB
Performance
Pixel Rate
43.92 GPixel/s
47.52 GPixel/s
Texture Rate
1,083.4 GTexel/s
617.8 GTexel/s
FP32 (TFLOPS)
69.34 TFLOPS
39.54 TFLOPS
FP64 (TFLOPS)
34.67 TFLOPS (1:2)
19.77 TFLOPS (1:2)
FP16 (TFLOPS)
69.34 TFLOPS (1:1)
79.07 TFLOPS (2:1)
AI/RT
Tensor Cores
592
312 -47.3%
Power
TDP
1000 W
400 W
TDP (W)
1,000
400 -60.0%
Suggested PSU
1400 W
800 W
Architecture
Architecture
Blackwell
Hopper
GPU Name
GB100
GH100
Generation
Server Blackwell (Bxx)
Server Hopper (Hxx)
Process Size
5 nm
5 nm
Transistors
208,000 million
80,000 million
Die Size
1628 mm²
814 mm²
Foundry
TSMC
TSMC
Density
127.8M / mm²
98.3M / mm²
API Support
OpenCL
3.0
3.0
CUDA
10.0
9.0
Physical
Slot Width
SXM Module
SXM Module
Outputs
No outputs
No outputs
Bus Interface
PCIe 6.0 x16
PCIe 5.0 x16
Other
Launch Price
34,999 USD
—
Production
Active
Active
Predecessor
Server Hopper
Server Ada
Successor
Server Rubin
Server Blackwell
View B200 SXM6 Details View H20 NVL16 Details