NVIDIA H100 SXM5 94 GB vs NVIDIA Rubin GPU Comparison

NVIDIA
GEFORCE

NVIDIA H100 SXM5 94 GB

CORE STATE GH100
VRAM 94 GB
CLOCK SPEED 1980 MHz
TDP 700 W
BUS WIDTH 5120 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

Rubin GPU

CORE STATE GR100
VRAM 288 GB
CLOCK SPEED 2267 MHz
TDP 2300 W
BUS WIDTH 16384 bit
ARCHITECTURE Rubin
nm
PROCESS 3 nm
LAUNCH DATE 2026

Analysis: NVIDIA H100 SXM5 94 GB vs NVIDIA Rubin GPU

Head-to-Head Benchmarks

The recorded database contains no direct head-to-head benchmark results for the NVIDIA H100 SXM5 94 GB and the NVIDIA Rubin GPU. Both entries show an empty benchmark array, zero average benchmark scores, and zero recorded wins for either part. The percentile versus all GPUs is identical at 50 for both products, indicating that neither has a measured performance position in the global database hierarchy.

The absence of benchmark data means the comparison must rely entirely on the architectural and specification differences recorded in the database. The Rubin GPU carries a massive theoretical compute advantage on paper: its FP32 throughput is listed at 130.0 TFLOPS, which is 94.3% higher than the H100's 66.91 TFLOPS. In FP16 compute, the two are much closer, with the H100 at 267.6 TFLOPS (4:1) and the Rubin at 260.0 TFLOPS (2:1), a difference of only 2.9% in favor of the H100 when using the stated ratios.

Texture rate shows a clear separation. The Rubin GPU reaches 2,031.2 GTexel/s, while the H100 delivers 1,045.4 GTexel/s. That places the Rubin approximately 94.3% ahead in raw texture fill capability. Pixel rate is closer: 54.41 GPixel/s for Rubin versus 47.52 GPixel/s for H100, a 14.5% advantage for the newer part.

Memory bandwidth is where the Rubin GPU separates itself most decisively. The database records 22.1 TB/s for the Rubin GPU against 3.36 TB/s for the H100. That is a 6.6x difference in favor of Rubin. Memory capacity follows the same pattern: 288 GB of HBM4 versus 94 GB of HBM3, a 3.1x increase. The bus width expands from 5,120 bit to 16,384 bit, exactly 3.2x wider.

Clock behavior differs notably between the two. The H100 has a base clock of 1350 MHz and a boost of 1980 MHz. The Rubin GPU lists a much lower base of 700 MHz but a higher boost of 2267 MHz. The boost delta is 14.5% in Rubin's favor, while the base clock delta is 48.1% in the H100's favor. This suggests the Rubin part relies more heavily on boost behavior to reach its rated performance.

Where Each One Wins

Based on the recorded specification data, the H100 SXM5 94 GB holds advantages in a narrow set of metrics. Its FP16 throughput of 267.6 TFLOPS (4:1) exceeds the Rubin GPU's 260.0 TFLOPS (2:1) by 7.6 TFLOPS, or about 2.9%. The H100 also has a higher base clock at 1350 MHz versus 700 MHz, which could indicate different sustained-load characteristics, though no power or thermal measurements are present in the database to confirm behavior. The H100's memory clock of 1313 MHz (5.3 Gbps effective) is lower than the Rubin's 2695 MHz (10.8 Gbps effective) in absolute terms, but the H100 uses HBM3 while Rubin uses HBM4.

The Rubin GPU wins on nearly every other recorded performance-related field. FP32 compute is 130.0 TFLOPS versus 66.91 TFLOPS, a 94.3% advantage. Texture rate is 2,031.2 GTexel/s versus 1,045.4 GTexel/s, also a 94.3% advantage. Pixel rate is 54.41 GPixel/s versus 47.52 GPixel/s, a 14.5% advantage. Memory bandwidth is 22.1 TB/s versus 3.36 TB/s, a 6.6x advantage. Memory capacity is 288 GB versus 94 GB, a 3.1x advantage. The bus interface advances from PCIe 5.0 x16 to PCIe 6.0 x16.

The shading unit count is 28,672 for Rubin versus 16,896 for H100, a 69.7% increase. TMUs number 896 versus 528, a 69.7% increase. Tensor cores are 896 versus 528, the same 69.7% increase. ROPs are identical at 24 for both, which explains why the pixel rate gap is much smaller than the texture or compute gaps.

For use-case analysis, the FP16 ratio difference matters. The H100's 4:1 FP16 ratio indicates its tensor and shader paths are organized for that throughput. The Rubin's 2:1 ratio with 260.0 TFLOPS suggests a different internal arrangement, but the database does not record further detail on how these ratios translate to actual workloads.

Architecture Differences

The process node differs substantially. The H100 uses a 5 nm process from TSMC, while the Rubin GPU uses a 3 nm process from the same foundry. Transistor count grows from 80,000 million on the H100 to 336,000 million on the Rubin GPU, a 4.2x increase. Die size expands from 814 mm² to 1456 mm², a 78.9% increase. Transistor density rises from 98.3M / mm² to 230.8M / mm², a 2.3x improvement, reflecting the tighter process geometry.

The chip identifiers differ: GH100 for H100 versus GR100 for Rubin. Architecture names are Hopper and Rubin respectively. Generation labels are "Server Hopper (Hxx)" and "Server Rubin (Rxx)". The H100 lists its predecessor as "Server Ada" and successor as "Server Blackwell". The Rubin GPU lists its predecessor as "Server Blackwell" and has no successor recorded. This places the two parts two full generations apart in the database's server lineup.

Memory architecture changes completely. The H100 uses HBM3 with a 5,120 bit bus and 3.36 TB/s bandwidth. The Rubin GPU uses HBM4 with a 16,384 bit bus and 22.1 TB/s bandwidth. The bus width increase of 3.2x combined with the memory clock increase from 1313 MHz to 2695 MHz produces the 6.6x bandwidth jump. Capacity moves from 94 GB to 288 GB.

Power specifications differ dramatically. The H100 is rated at 700 W TDP with a suggested PSU of 1100 W and an 8-pin EPS power connector. The Rubin GPU is rated at 2300 W TDP with a suggested PSU of 2700 W and no power connector listed in the database. Both are SXM modules with no display outputs. The bus interface changes from PCIe 5.0 x16 to PCIe 6.0 x16.

The API support fields show a difference. The H100 lists null values for DirectX, OpenGL, and Vulkan. The Rubin GPU lists "N/A" for all three, which the database records as distinct values. Neither part supports display outputs, consistent with their server accelerator positioning.

Texture mapping units scale from 528 to 896, and tensor cores scale identically from 528 to 896. Shading units scale from 16,896 to 28,672. The ROP count stays fixed at 24, which is an unusual constraint in the Rubin design and likely explains the relatively modest pixel rate improvement. Release dates are recorded as 2023-03-20 for the H100 and 2025-12-31 for the Rubin GPU.

FAQ

Q: How much faster is the Rubin GPU in FP32 compute compared to the H100?

A: The Rubin GPU delivers 130.0 TFLOPS FP32, while the H100 delivers 66.91 TFLOPS. That is a 94.3% higher FP32 throughput for the Rubin GPU.

Q: Which GPU has higher FP16 performance?

A: The H100 lists 267.6 TFLOPS FP16 (4:1), while the Rubin GPU lists 260.0 TFLOPS FP16 (2:1). The H100 is about 2.9% ahead in the recorded FP16 figures, though the ratios differ between the two parts.

Q: What is the memory bandwidth difference between the two?

A: The Rubin GPU has 22.1 TB/s of memory bandwidth from HBM4 on a 16,384 bit bus. The H100 has 3.36 TB/s from HBM3 on a 5,120 bit bus. The Rubin GPU provides approximately 6.6x the bandwidth.

Q: How do the transistor counts compare?

A: The H100 uses 80,000 million transistors on a 814 mm² die at 5 nm. The Rubin GPU uses 336,000 million transistors on a 1456 mm² die at 3 nm. That is a 4.2x transistor increase and a 78.9% die area increase.

Q: Do both GPUs have the same pixel output capability?

A: No. The H100 has a pixel rate of 47.52 GPixel/s, and the Rubin GPU has 54.41 GPixel/s. Both have 24 ROPs, but the Rubin GPU's higher clock contributes to the 14.5% pixel rate advantage.

Q: What are the power requirements recorded for each?

A: The H100 has a TDP of 700 W and a suggested PSU of 1100 W. The Rubin GPU has a TDP of 2300 W and a suggested PSU of 2700 W. The Rubin GPU consumes over 3x the TDP of the H100.

Specification Differences

| Field | H100 SXM5 94 GB | Rubin GPU |

|---|---|---|

| Chip | GH100 | GR100 |

| Architecture | Hopper | Rubin |

| Generation | Server Hopper (Hxx) | Server Rubin (Rxx) |

| Process Node | 5 nm | 3 nm |

| Transistors | 80,000 million | 336,000 million |

| Die Size | 814 mm² | 1456 mm² |

| Transistor Density | 98.3M / mm² | 230.8M / mm² |

| Base Clock | 1350 MHz | 700 MHz |

| Boost Clock | 1980 MHz | 2267 MHz |

| Memory Clock | 1313 MHz, 5.3 Gbps effective | 2695 MHz, 10.8 Gbps effective |

| Memory Size | 94 GB | 288 GB |

| Memory Type | HBM3 | HBM4 |

| Memory Bus Width | 5120 bit | 16384 bit |

| Memory Bandwidth | 3.36 TB/s | 22.1 TB/s |

| Shading Units | 16896 | 28672 |

| TMUs | 528 | 896 |

| Tensor Cores | 528 | 896 |

| Pixel Rate | 47.52 GPixel/s | 54.41 GPixel/s |

| Texture Rate | 1,045.4 GTexel/s | 2,031.2 GTexel/s |

| FP32 | 66.91 TFLOPS | 130.0 TFLOPS |

| FP16 | 267.6 TFLOPS (4:1) | 260.0 TFLOPS (2:1) |

| TDP | 700 W | 2300 W |

| Power Connectors | 8-pin EPS | None recorded |

| Suggested PSU | 1100 W | 2700 W |

| Bus Interface | PCIe 5.0 x16 | PCIe 6.0 x16 |

| Predecessor | Server Ada | Server Blackwell |

| Successor | Server Blackwell | None recorded |

| Release Date | 2023-03-20 | 2025-12-31 |

| DirectX | Null | N/A |

| OpenGL | Null | N/A |

| Vulkan | Null | N/A |

DETAILED SPECIFICATIONS

SPECIFICATION
H100 SXM5 94 GB
Rubin GPU
Core Specs
Shading Units
16,896
28,672 +69.7%
Shaders
16,896
28,672 +69.7%
TMUs
528
896 +69.7%
ROPs
24
24 0.0%
SM Count
132
224 +69.7%
Clocks
Base Clock
1350 MHz
700 MHz
Boost Clock
1980 MHz
2267 MHz
Memory Clock
1313 MHz 5.3 Gbps effective
2695 MHz 10.8 Gbps effective
Memory
Memory Size
94 GB
288 GB
VRAM (MB)
96,256
294,912 +206.4%
Memory Type
HBM3
HBM4
Memory Bus
5120 bit
16384 bit
Bandwidth
3.36 TB/s
22.1 TB/s
Cache
L1 Cache
256 KB (per SM)
256 KB (per SM)
L2 Cache
50 MB
128 MB
Performance
Pixel Rate
47.52 GPixel/s
54.41 GPixel/s
Texture Rate
1,045.4 GTexel/s
2,031.2 GTexel/s
FP32 (TFLOPS)
66.91 TFLOPS
130.0 TFLOPS
FP64 (TFLOPS)
33.45 TFLOPS (1:2)
32.50 TFLOPS (1:4)
FP16 (TFLOPS)
267.6 TFLOPS (4:1)
260.0 TFLOPS (2:1)
AI/RT
Tensor Cores
528
896 +69.7%
Power
TDP
700 W
2300 W
TDP (W)
700
2,300 +228.6%
Suggested PSU
1100 W
2700 W
Power Connectors
8-pin EPS
—
Architecture
Architecture
Hopper
Rubin
GPU Name
GH100
GR100
Generation
Server Hopper (Hxx)
Server Rubin (Rxx)
Process Size
5 nm
3 nm
Transistors
80,000 million
336,000 million
Die Size
814 mm²
1456 mm²
Foundry
TSMC
TSMC
Density
98.3M / mm²
230.8M / mm²
API Support
OpenCL
3.0
3.0
CUDA
9.0
10.7
Physical
Slot Width
SXM Module
SXM Module
Outputs
No outputs
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 6.0 x16
Other
Production
Active
Active
Predecessor
Server Ada
Server Blackwell
Successor
Server Blackwell
—
View H100 SXM5 94 GB Details View Rubin GPU Details