NVIDIA GB10 vs NVIDIA H20 NVL16 Comparison

NVIDIA
GEFORCE

NVIDIA GB10

CORE STATE GB20B
VRAM 128 GB
CLOCK SPEED 2418 MHz
TDP 140 W
BUS WIDTH 256 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

H20 NVL16

CORE STATE GH100
VRAM 96 GB
CLOCK SPEED 1980 MHz
TDP 400 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2025

PERFORMANCE BENCHMARKS

geekbench_opencl
120,137
N/A
geekbench_vulkan
114,648
N/A

Analysis: NVIDIA GB10 vs NVIDIA H20 NVL16

Where Each One Wins

The recorded data splits these two NVIDIA server accelerators along clear architectural and workload lines. The GB10, built on the Blackwell 2.0 architecture, is the only one with benchmark scores in the database. It posts 120,137 points in Geekbench OpenCL and 114,648 points in Geekbench Vulkan. Its average benchmark score sits at 117,393, placing it in the 95th percentile of all GPUs tracked. The H20 NVL16, a Hopper generation part, has no recorded benchmark entries, an average score of zero, and sits in the 50th percentile. That absence of measured data is itself a meaningful distinction: the GB10 delivers verifiable compute results, while the H20 NVL16 has none in the database.

The GB10 also wins on practical integration. It uses an IGP slot width, requires no power connectors, and lists a suggested PSU of 300 W. The H20 NVL16 uses an SXM Module slot width, carries a 400 W TDP, and lists a suggested PSU of 800 W. For systems where physical footprint and power delivery are constrained, the GB10 has the clear advantage. Its dimensions are recorded at 150 mm by 51 mm by 150 mm, while the H20 NVL16 has no recorded dimensions. The GB10 also includes a display output, one HDMI port, whereas the H20 NVL16 has no outputs at all. That makes the GB10 the only one of the two that can drive a monitor directly.

In raw shading throughput, the H20 NVL16 wins. It has 9,984 shading units against the GB10's 6,144, and its FP32 rate of 39.54 TFLOPS exceeds the GB10's 29.71 TFLOPS. The H20 NVL16 also doubles its FP16 throughput to 79.07 TFLOPS with a 2:1 ratio, while the GB10 holds FP16 at 29.71 TFLOPS with a 1:1 ratio. For workloads that scale with shader count or FP16 math, the H20 NVL16 is the stronger part on paper. The GB10 counters with a higher boost clock, 2418 MHz versus 1980 MHz, and a narrower gap between base and boost, but the H20's larger shader array carries the throughput advantage.

Memory capacity splits the two differently. The GB10 has 128 GB of LPDDR5X on a 256-bit bus, delivering 273.2 GB/s. The H20 NVL16 has 96 GB of HBM3 on a 6144-bit bus, delivering 4.03 TB/s. The H20's bandwidth is roughly 14.7 times higher per the recorded figures, which matters for large matrix operations or data movement. The GB10's larger capacity, 32 GB more, suits models that need to reside in memory rather than stream through it. The H20 also uses 80,000 million transistors on an 814 mm² die, while the GB10's transistor count is unknown and its die size is 382 mm². Both are fabricated at TSMC on a 5 nm node.

Architecture Differences

The GB10 uses the GB20B chip under the Blackwell 2.0 architecture, part of the Server Blackwell (Bxx) generation. The H20 NVL16 uses the GH100 chip under the Hopper architecture, part of the Server Hopper (Hxx) generation. This is a generational split: the GB10's predecessor is listed as Server Hopper and its successor as Server Rubin, while the H20 NVL16's predecessor is Server Ada and its successor is Server Blackwell. The GB10 is the newer part, releasing on 2025-10-14, while the H20 NVL16 released on 2025-09-01. Both are currently marked Active in production status.

The H20 NVL16 has no ray tracing cores recorded, while the GB10 has 48. The H20 NVL16 has 312 tensor cores against the GB10's 384. Texture mapping units differ as well: 312 on the H20 versus 384 on the GB10. Raster operation units are 24 on the H20 versus 48 on the GB10. The GB10 posts a pixel rate of 116.1 GPixel/s and a texture rate of 928.5 GTexel/s, both higher than the H20's 47.52 GPixel/s and 617.8 GTexel/s. The H20 compensates with more shading units, 9,984 versus 6,144, and a higher base clock of 1830 MHz versus 1665 MHz.

Memory architecture diverges sharply. The GB10 uses LPDDR5X with an 8.5 Gbps effective data rate, while the H20 uses HBM3 at 5.3 Gbps effective. The bus widths reflect the packaging: 256 bit for the GB10, 6144 bit for the H20. The H20's bandwidth advantage, 4.03 TB/s versus 273.2 GB/s, is the largest single architectural gap between the two. The GB10's 128 GB capacity exceeds the H20's 96 GB, but the H20's memory technology is designed for sustained high throughput rather than capacity.

Both parts share the PCIe 5.0 x16 bus interface. Neither supports DirectX, OpenGL, or Vulkan APIs per the database, which marks all three as N/A for both. The GB10 lists a 140 W TDP and a 300 W suggested PSU, while the H20 lists a 400 W TDP and an 800 W suggested PSU. The GB10 has a launch MSRP of 3,999 USD; the H20 has no recorded launch MSRP.

FAQ

Q: Which part has higher FP32 throughput?

A: The H20 NVL16. It records 39.54 TFLOPS FP32, while the GB10 records 29.71 TFLOPS.

Q: Does the GB10 support ray tracing?

A: Yes. The GB10 has 48 ray tracing cores. The H20 NVL16 has no ray tracing cores recorded.

Q: Which part has more memory bandwidth?

A: The H20 NVL16. It delivers 4.03 TB/s from HBM3 over a 6144-bit bus, versus the GB10's 273.2 GB/s from LPDDR5X over a 256-bit bus.

Q: Which part has a higher benchmark percentile ranking?

A: The GB10. It ranks in the 95th percentile of all GPUs, while the H20 NVL16 ranks in the 50th percentile.

Q: Are there any benchmark scores for the H20 NVL16?

A: No. The database lists no benchmark entries for the H20 NVL16, and its average benchmark score is zero.

Q: What is the transistor count difference?

A: The H20 NVL16 uses 80,000 million transistors on an 814 mm² die. The GB10's transistor count is unknown, and its die size is 382 mm².

Specification Differences

The two parts differ on nearly every measured specification. The GB10 uses the GB20B chip, while the H20 uses the GH100. Architecture is Blackwell 2.0 versus Hopper. Generations are Server Blackwell versus Server Hopper. The GB10's die is 382 mm²; the H20's is 814 mm². Transistor density is recorded only for the H20 at 98.3M per mm²; the GB10 has no transistor count or density listed.

Clock speeds differ: the GB10's base is 1665 MHz, its boost is 2418 MHz, and its memory runs at 1067 MHz with 8.5 Gbps effective. The H20's base is 1830 MHz, its boost is 1980 MHz, and its memory runs at 1313 MHz with 5.3 Gbps effective. The GB10 boosts higher, while the H20 has a higher base clock.

Memory configuration: the GB10 has 128 GB LPDDR5X on a 256-bit bus at 273.2 GB/s. The H20 has 96 GB HBM3 on a 6144-bit bus at 4.03 TB/s. The GB10 has 6,144 shading units, 384 TMUs, 48 ROPs, 48 RT cores, and 384 tensor cores. The H20 has 9,984 shading units, 312 TMUs, 24 ROPs, no RT cores, and 312 tensor cores.

Pixel and texture rates: the GB10 posts 116.1 GPixel/s and 928.5 GTexel/s. The H20 posts 47.52 GPixel/s and 617.8 GTexel/s. FP32 is 29.71 TFLOPS for the GB10 and 39.54 TFLOPS for the H20. FP16 is 29.71 TFLOPS (1:1) for the GB10 and 79.07 TFLOPS (2:1) for the H20.

Power and physical specs: the GB10 has a 140 W TDP, IGP slot width, no power connectors, and a 300 W suggested PSU. The H20 has a 400 W TDP, SXM Module slot width, no recorded power connectors, and an 800 W suggested PSU. The GB10 has one HDMI output; the H20 has no outputs. The GB10 has recorded dimensions of 150 mm by 51 mm by 150 mm; the H20 has no dimensions listed. The GB10 has a launch MSRP of 3,999 USD; the H20 has none.

Head-to-Head Benchmarks

The database contains no head-to-head benchmark entries between the GB10 and the H20 NVL16. The GB10 has two recorded benchmark scores: 120,137 in Geekbench OpenCL and 114,648 in Geekbench Vulkan. These produce an average of 117,393, placing it in the 95th percentile. The H20 NVL16 has zero benchmark scores and an average of zero, placing it in the 50th percentile.

The GB10's nearest rivals in the database provide context for its measured performance. The NVIDIA RTX 4000 SFF Ada Generation averages 117,088, which is 0.3% behind the GB10. The AMD Radeon PRO W7700 averages 118,976, which is 1.3% ahead of the GB10. The NVIDIA Tesla V100 SXM2 16 GB averages 114,395, sitting 2.6% behind. The NVIDIA RTX A5500 Mobile averages 113,944, trailing by 3.0%. The GB10's measured scores place it in a narrow competitive band around the 117,000 to 119,000 average range, with the closest rival, the RTX 4000 SFF Ada Generation, within a fraction of a percent.

The H20 NVL16 has no nearest rivals listed, no benchmark scores, and no average score. Its 50th percentile ranking, which is the midpoint of the database's GPU distribution, does not correspond to any recorded performance data. The only quantitative claims that can be made about the H20 come from its specification sheet: higher FP32, much higher FP16, and dramatically higher memory bandwidth.

The specification comparison indicates the H20 NVL16 should win throughput-bound tasks. Its 79.07 TFLOPS FP16 is over 2.6 times the GB10's 29.71 TFLOPS. Its memory bandwidth of 4.03 TB/s is roughly 14.7 times the GB10's 273.2 GB/s. Those gaps are large enough to dominate memory-heavy or tensor-heavy workloads. The GB10 counters with higher boost clocks, more TMUs, more ROPs, more RT cores, more tensor cores, and higher pixel and texture rates. Its pixel rate of 116.1 GPixel/s is about 2.4 times the H20's 47.52 GPixel/s, and its texture rate of 928.5 GTexel/s is about 1.5 times the H20's 617.8 GTexel/s.

The GB10 also has more memory capacity, 128 GB versus 96 GB, which is a 33% advantage in the recorded figures. That capacity, combined with the IGP form factor and lower power draw, positions the GB10 for dense integration. The H20's SXM module form factor, higher TDP, and HBM3 packaging target a different deployment scenario, one where raw bandwidth and FP16 throughput outweigh physical simplicity.

Neither part has any recorded API support for DirectX, OpenGL, or Vulkan, confirming both are compute-oriented server accelerators rather than graphics products. The GB10's single HDMI output is the only display capability between the two. The absence of head-to-head benchmark data means the performance relationship between these two parts rests entirely on the specification sheet, and the specification sheet shows two designs optimized for opposite ends of the server GPU spectrum.

DETAILED SPECIFICATIONS

SPECIFICATION
GB10
H20 NVL16
Core Specs
Shading Units
6,144
9,984 +62.5%
Shaders
6,144
9,984 +62.5%
TMUs
384
312 -18.8%
ROPs
48
24 -50.0%
SM Count
48
78 +62.5%
Clocks
Base Clock
1665 MHz
1830 MHz
Boost Clock
2418 MHz
1980 MHz
Memory Clock
1067 MHz 8.5 Gbps effective
1313 MHz 5.3 Gbps effective
Memory
Memory Size
128 GB
96 GB
VRAM (MB)
131,072
98,304 -25.0%
Memory Type
LPDDR5X
HBM3
Memory Bus
256 bit
6144 bit
Bandwidth
273.2 GB/s
4.03 TB/s
Cache
L1 Cache
128 KB (per SM)
256 KB (per SM)
L2 Cache
50 MB
60 MB
Performance
Pixel Rate
116.1 GPixel/s
47.52 GPixel/s
Texture Rate
928.5 GTexel/s
617.8 GTexel/s
FP32 (TFLOPS)
29.71 TFLOPS
39.54 TFLOPS
FP64 (TFLOPS)
464.3 GFLOPS (1:64)
19.77 TFLOPS (1:2)
FP16 (TFLOPS)
29.71 TFLOPS (1:1)
79.07 TFLOPS (2:1)
AI/RT
RT Cores
48
Tensor Cores
384
312 -18.8%
Power
TDP
140 W
400 W
TDP (W)
140
400 +185.7%
Suggested PSU
300 W
800 W
Power Connectors
None
Architecture
Architecture
Blackwell 2.0
Hopper
GPU Name
GB20B
GH100
Generation
Server Blackwell (Bxx)
Server Hopper (Hxx)
Process Size
5 nm
5 nm
Transistors
unknown
80,000 million
Die Size
382 mm²
814 mm²
Foundry
TSMC
TSMC
Density
98.3M / mm²
API Support
OpenCL
3.0
3.0
CUDA
12.1
9.0
Physical
Slot Width
IGP
SXM Module
Length
150 mm 5.9 inches
Height
51 mm 2 inches
Outputs
1x HDMI
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Launch Price
3,999 USD
Production
Active
Active
Predecessor
Server Hopper
Server Ada
Successor
Server Rubin
Server Blackwell
View GB10 Details View H20 NVL16 Details