NVIDIA H20 vs NVIDIA Rubin GPU Comparison

NVIDIA
GEFORCE

NVIDIA H20

CORE STATE GH100
VRAM 96 GB
CLOCK SPEED 1980 MHz
TDP 500 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2024
VS
NVIDIA
GEFORCE

Rubin GPU

CORE STATE GR100
VRAM 288 GB
CLOCK SPEED 2267 MHz
TDP 2300 W
BUS WIDTH 16384 bit
ARCHITECTURE Rubin
nm
PROCESS 3 nm
LAUNCH DATE 2026

Analysis: NVIDIA H20 vs NVIDIA Rubin GPU

Head-to-Head Benchmarks

The database contains no recorded benchmark scores for either the NVIDIA H20 or the NVIDIA Rubin GPU. Both parts show an average benchmark score of zero, and the head-to-head benchmark table is empty. This means there is no measured performance data to compare directly, and the wins counter for each side sits at zero. Without benchmark results, the only quantitative performance indicators available are the theoretical compute rates, memory specifications, and clock speeds listed in the database.

The FP32 compute rate tells a clear story. The H20 delivers 39.54 TFLOPS of single-precision performance, while the Rubin GPU reaches 130.0 TFLOPS. That puts the Rubin at roughly 3.3 times the FP32 throughput of the H20. In FP16, the gap follows the same pattern: the H20 produces 79.07 TFLOPS, and the Rubin GPU reaches 260.0 TFLOPS, again a 3.3 times advantage. These are raw shader and tensor core throughput figures, not application-level benchmarks, but they indicate the Rubin GPU has a substantially larger compute pipeline.

The texture rate difference is even more pronounced. The H20 manages 617.8 GTexel/s, while the Rubin GPU records 2,031.2 GTexel/s. That is approximately 3.3 times higher, consistent with the TMU count scaling. Pixel rate is the one area where the two parts are close: the H20 outputs 47.52 GPixel/s, and the Rubin GPU outputs 54.41 GPixel/s, a modest 14% advantage for the Rubin. Both parts use 24 ROPs, so the pixel throughput difference comes entirely from the higher boost clock on the Rubin GPU.

Memory bandwidth separates the two far more than pixel rate. The H20 has 4.03 TB/s of bandwidth across a 6144-bit bus using HBM3. The Rubin GPU has 22.1 TB/s across a 16384-bit bus using HBM4. That is a 5.5 times bandwidth advantage for the Rubin GPU, driven by both a wider bus and faster memory. The effective memory clock on the Rubin is 10.8 Gbps, compared to 5.3 Gbps on the H20.

Clock speeds behave differently between the two. The H20 has a base clock of 1830 MHz and a boost clock of 1980 MHz. The Rubin GPU has a much lower base clock of 700 MHz but a higher boost clock of 2267 MHz. The large base clock gap suggests the Rubin GPU is designed to idle or run at very low clocks under light load, then scale up aggressively under compute load. The H20 runs at a more conventional base-to-boost spread.

Neither part has any display outputs, and both are SXM modules with no API support for DirectX, OpenGL, or Vulkan. These are server accelerators, not graphics cards, so the absence of benchmark scores in the database likely reflects the fact that standard GPU benchmarks are not applicable to this class of hardware.

Architecture Differences

The two accelerators come from different architecture generations and use different manufacturing processes. The H20 is built on the Hopper architecture, specifically the GH100 chip, and is part of the Server Hopper (Hxx) generation. The Rubin GPU uses the Rubin architecture with the GR100 chip, part of the Server Rubin (Rxx) generation. The H20's predecessor is listed as Server Ada, and its successor is Server Blackwell. The Rubin GPU's predecessor is Server Blackwell, and it has no successor listed.

The manufacturing process differs significantly. The H20 uses a 5 nm process at TSMC, while the Rubin GPU uses a 3 nm process at the same foundry. Transistor counts scale accordingly: the H20 has 80,000 million transistors on a die size of 814 mm², giving a transistor density of 98.3 million per mm². The Rubin GPU has 336,000 million transistors on a die size of 1456 mm², giving a density of 230.8 million per mm². That means the Rubin GPU packs more than four times the transistor count into less than twice the die area.

The compute resources differ at every level. The H20 has 9984 shading units, 312 TMUs, 24 ROPs, and 312 tensor cores. The Rubin GPU has 28672 shading units, 896 TMUs, 24 ROPs, and 896 tensor cores. The shading unit count is 2.9 times higher on the Rubin, the TMU count is 2.9 times higher, and the tensor core count is 2.9 times higher. The ROP count is identical at 24.

Memory architecture is a major differentiator. The H20 uses 96 GB of HBM3 on a 6144-bit bus, while the Rubin GPU uses 288 GB of HBM4 on a 16384-bit bus. The Rubin GPU has three times the memory capacity and roughly 2.7 times the bus width. The memory clock also differs: the H20 runs at 1313 MHz with 5.3 Gbps effective data rate, while the Rubin runs at 2695 MHz with 10.8 Gbps effective.

The bus interface has changed between generations. The H20 uses PCIe 5.0 x16, while the Rubin GPU uses PCIe 6.0 x16. Power requirements are vastly different: the H20 has a TDP of 500 W with a suggested PSU of 900 W, while the Rubin GPU has a TDP of 2300 W with a suggested PSU of 2700 W. Both are SXM modules with no power connector details listed.

Release timing also differs. The H20 was released on 2024-01-31, while the Rubin GPU is dated 2025-12-31. Both are listed as Active in production status.

The Verdict

The data shows two accelerators at very different points in their respective architectures. The H20 is a Hopper-generation part with a 500 W TDP, 96 GB of HBM3, and 39.54 TFLOPS of FP32 compute. The Rubin GPU is a next-generation Rubin-architecture part with a 2300 W TDP, 288 GB of HBM4, and 130.0 TFLOPS of FP32 compute.

For workloads that depend on memory capacity and bandwidth, the Rubin GPU is clearly ahead. It has three times the memory capacity, 5.5 times the bandwidth, and a wider bus. For compute-heavy tasks, the Rubin GPU offers 3.3 times the FP32 and FP16 throughput. The Rubin GPU also has 2.9 times the shading units, TMUs, and tensor cores compared to the H20.

However, the H20 has advantages that matter in certain deployment contexts. It consumes 500 W versus 2300 W, which means it requires far less power delivery and cooling infrastructure. The H20 also has a higher base clock of 1830 MHz versus 700 MHz, which could translate to better performance in workloads that do not scale with core count or that are latency-sensitive at low utilization. The H20's smaller die size of 814 mm² versus 1456 mm² also means more dies per wafer, though the database does not include wafer yield data.

The pixel rate difference is small: 47.52 GPixel/s for the H20 versus 54.41 GPixel/s for the Rubin GPU, a 14% gap. This is the only metric where the two parts are close, and it reflects the identical ROP count of 24 on both.

The choice between these two parts depends on whether the workload can utilize the Rubin GPU's massive compute and memory resources, and whether the power and cooling budget can support a 2300 W part. The H20 is a lower-power option that still delivers substantial compute, but it cannot match the Rubin GPU on raw throughput, memory size, or bandwidth. The data favors the Rubin GPU for maximum performance, but favors the H20 for power-constrained environments.

FAQ

Q: Which GPU has higher FP32 compute performance?

A: The NVIDIA Rubin GPU delivers 130.0 TFLOPS of FP32 compute, which is approximately 3.3 times the 39.54 TFLOPS of the NVIDIA H20.

Q: How much memory does each GPU have and what type?

A: The H20 has 96 GB of HBM3, while the Rubin GPU has 288 GB of HBM4. The Rubin GPU has three times the capacity and uses a newer memory type.

Q: What is the memory bandwidth difference between the two?

A: The H20 has 4.03 TB/s of bandwidth, while the Rubin GPU has 22.1 TB/s, a 5.5 times advantage for the Rubin GPU.

Q: Are these GPUs suitable for graphics workloads?

A: No. Both parts have no display outputs and list N/A for DirectX, OpenGL, and Vulkan support. They are server accelerators, not graphics cards.

Q: What is the power consumption of each GPU?

A: The H20 has a TDP of 500 W with a suggested PSU of 900 W. The Rubin GPU has a TDP of 2300 W with a suggested PSU of 2700 W.

Q: Which architecture does each GPU use?

A: The H20 uses the Hopper architecture with the GH100 chip. The Rubin GPU uses the Rubin architecture with the GR100 chip.

Q: What is the transistor count difference?

A: The H20 has 80,000 million transistors on an 814 mm² die. The Rubin GPU has 336,000 million transistors on a 1456 mm² die, over four times the transistor count.

Where Each One Wins

The NVIDIA H20 wins in power efficiency and low-load responsiveness. Its 500 W TDP is far below the Rubin GPU's 2300 W, and its base clock of 1830 MHz is much higher than the Rubin's 700 MHz. For workloads that are bursty, latency-sensitive, or that cannot justify the power delivery and cooling for a 2300 W part, the H20 is the more practical choice. Its pixel rate of 47.52 GPixel/s is also close to the Rubin's, so graphics-like output tasks see only a small penalty.

The NVIDIA Rubin GPU wins in every raw throughput category except pixel rate. It has 3.3 times the FP32 and FP16 compute, 2.9 times the shading units, 2.9 times the TMUs, 2.9 times the tensor cores, 3 times the memory capacity, 5.5 times the memory bandwidth, and a newer PCIe 6.0 x16 interface versus PCIe 5.0 x16 on the H20. The Rubin GPU also has a higher boost clock of 2267 MHz versus 1980 MHz.

For memory-bound workloads such as large model inference or training with massive datasets, the Rubin GPU's 288 GB of HBM4 and 22.1 TB/s bandwidth are decisive. For compute-bound workloads that can scale across 28672 shading units, the Rubin GPU's 130.0 TFLOPS FP32 is the clear winner. The H20 remains viable for smaller models or for deployments where the 500 W TDP fits existing infrastructure. The data does not support recommending the H20 for peak performance, but it does support recommending it for power-constrained or low-utilization scenarios.

DETAILED SPECIFICATIONS

SPECIFICATION
H20
Rubin GPU
Core Specs
Shading Units
9,984
28,672 +187.2%
Shaders
9,984
28,672 +187.2%
TMUs
312
896 +187.2%
ROPs
24
24 0.0%
SM Count
78
224 +187.2%
Clocks
Base Clock
1830 MHz
700 MHz
Boost Clock
1980 MHz
2267 MHz
Memory Clock
1313 MHz 5.3 Gbps effective
2695 MHz 10.8 Gbps effective
Memory
Memory Size
96 GB
288 GB
VRAM (MB)
98,304
294,912 +200.0%
Memory Type
HBM3
HBM4
Memory Bus
6144 bit
16384 bit
Bandwidth
4.03 TB/s
22.1 TB/s
Cache
L1 Cache
256 KB (per SM)
256 KB (per SM)
L2 Cache
60 MB
128 MB
Performance
Pixel Rate
47.52 GPixel/s
54.41 GPixel/s
Texture Rate
617.8 GTexel/s
2,031.2 GTexel/s
FP32 (TFLOPS)
39.54 TFLOPS
130.0 TFLOPS
FP64 (TFLOPS)
19.77 TFLOPS (1:2)
32.50 TFLOPS (1:4)
FP16 (TFLOPS)
79.07 TFLOPS (2:1)
260.0 TFLOPS (2:1)
AI/RT
Tensor Cores
312
896 +187.2%
Power
TDP
500 W
2300 W
TDP (W)
500
2,300 +360.0%
Suggested PSU
900 W
2700 W
Architecture
Architecture
Hopper
Rubin
GPU Name
GH100
GR100
Generation
Server Hopper (Hxx)
Server Rubin (Rxx)
Process Size
5 nm
3 nm
Transistors
80,000 million
336,000 million
Die Size
814 mm²
1456 mm²
Foundry
TSMC
TSMC
Density
98.3M / mm²
230.8M / mm²
API Support
OpenCL
3.0
3.0
CUDA
9.0
10.7
Physical
Slot Width
SXM Module
SXM Module
Outputs
No outputs
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 6.0 x16
Other
Production
Active
Active
Predecessor
Server Ada
Server Blackwell
Successor
Server Blackwell
View H20 Details View Rubin GPU Details