NVIDIA H100 SXM5 96 GB vs NVIDIA Rubin GPU Comparison

NVIDIA
GEFORCE

NVIDIA H100 SXM5 96 GB

CORE STATE GH100
VRAM 96 GB
CLOCK SPEED 1980 MHz
TDP 700 W
BUS WIDTH 5120 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

Rubin GPU

CORE STATE GR100
VRAM 288 GB
CLOCK SPEED 2267 MHz
TDP 2300 W
BUS WIDTH 16384 bit
ARCHITECTURE Rubin
nm
PROCESS 3 nm
LAUNCH DATE 2026

Analysis: NVIDIA H100 SXM5 96 GB vs NVIDIA Rubin GPU

Where Each One Wins

The recorded data shows no benchmark suite results for either GPU, so the win/loss split cannot be established from direct measurements. Instead, the division of strengths must be inferred from the architectural specifications and the memory, compute, and interface characteristics that each part was designed to deliver.

The NVIDIA H100 SXM5 96 GB belongs to the Hopper generation, a server platform aimed at the previous wave of accelerated computing. Its strengths lie in a mature 5 nm process, a 700 W power envelope, and a balanced FP32 to FP16 ratio that suits mixed-precision training and inference workloads where the 4:1 FP16 ratio is a deliberate design choice. The H100 carries 96 GB of HBM3 memory on a 5120 bit bus, producing 3.36 TB/s of bandwidth, which positions it for models that fit within that capacity and benefit from the established PCIe 5.0 x16 interface.

The NVIDIA Rubin GPU, by contrast, is a next-generation part on the 3 nm node with a 2300 W thermal design. It delivers 288 GB of HBM4 memory across a 16384 bit bus, yielding 22.1 TB/s of bandwidth, a figure that dwarfs the H100 by a factor of roughly 6.5. The Rubin also moves to PCIe 6.0 x16, doubling the interconnect generation. In terms of raw shading throughput, the Rubin’s FP32 output of 130.0 TFLOPS is nearly double the H100’s 66.91 TFLOPS, while its FP16 rating of 260.0 TFLOPS at a 2:1 ratio is slightly below the H100’s 267.6 TFLOPS at a 4:1 ratio. That distinction matters: the Rubin trades the higher FP16 ratio for more balanced performance across precision formats, whereas the H100 concentrates its tensor throughput into a narrower FP16 path.

For memory-bound workloads, the Rubin wins decisively. For compute-bound tasks that rely on FP16 tensor cores, the H100 retains a narrow edge in raw FP16 TFLOPS, though the Rubin’s architecture likely compensates through other means. The H100’s pixel rate of 47.52 GPixel/s trails the Rubin’s 54.41 GPixel/s, and the texture rate of 1,045.4 GTexel/s on the H100 is roughly half the Rubin’s 2,031.2 GTexel/s. The Rubin also has more shading units (28672 versus 16896), more TMUs (896 versus 528), and more tensor cores (896 versus 528), while both parts share the same 24 ROPs.

Architecture Differences

The two GPUs diverge at the most fundamental level: process node, die size, and transistor budget. The H100 uses TSMC’s 5 nm process, fitting 80,000 million transistors onto an 814 mm² die, for a density of 98.3 million transistors per square millimeter. The Rubin moves to TSMC’s 3 nm node, packing 336,000 million transistors onto a 1456 mm² die, achieving a density of 230.8 million transistors per square millimeter. That is a 4.2x increase in transistor count and a 2.35x increase in density, enabled by the smaller process and a much larger physical die.

The architectural generation differs as well. The H100 is built on Hopper, the server architecture that succeeded Server Ada and was itself succeeded by Server Blackwell. The Rubin is the first part on the Rubin architecture, succeeding Server Blackwell, and it has no successor listed in the database. This generational leap is reflected in the clock behavior. The H100 has a base clock of 1350 MHz and a boost clock of 1980 MHz. The Rubin has a lower base clock of 700 MHz but a higher boost clock of 2267 MHz, indicating a design that can ramp aggressively under load while idling at a lower frequency to manage the 2300 W power envelope.

Memory architecture is another major split. The H100 uses HBM3 with a 5120 bit bus and a memory clock of 1313 MHz, effective 5.3 Gbps, producing 3.36 TB/s. The Rubin uses HBM4 with a 16384 bit bus and a memory clock of 2695 MHz, effective 10.8 Gbps, producing 22.1 TB/s. The bus width triples, the memory generation advances, and the effective data rate doubles, resulting in the massive bandwidth advantage noted earlier.

The compute arrays also differ in composition. The H100 has 16896 shading units, 528 TMUs, 24 ROPs, and 528 tensor cores. The Rubin has 28672 shading units, 896 TMUs, 24 ROPs, and 896 tensor cores. Both parts have no dedicated RT cores listed. The interface moves from PCIe 5.0 x16 on the H100 to PCIe 6.0 x16 on the Rubin, and the power connectors shift from an 8-pin EPS on the H100 to none listed on the Rubin, with the suggested PSU rising from 1100 W to 2700 W. Both are SXM modules with no display outputs, and both have null or N/A API support for DirectX, OpenGL, and Vulkan, confirming their server-only roles.

Head-to-Head Benchmarks

No direct benchmark scores exist for either GPU in the database. The head-to-head benchmark array is empty, and both parts have an average benchmark score of zero, placing them at the 50th percentile against all GPUs by default. That percentile is not a meaningful comparison, as it applies uniformly to both parts without measured data. The nearest rivals list is also empty for both, so no external comparison points are available.

What the data does provide is a set of computed rate figures that serve as proxy benchmarks. In FP32 compute, the Rubin delivers 130.0 TFLOPS, which is 94.3% higher than the H100’s 66.91 TFLOPS. In FP16 compute, the H100 delivers 267.6 TFLOPS at a 4:1 ratio, while the Rubin delivers 260.0 TFLOPS at a 2:1 ratio, a difference of only 2.9% in favor of the H100. That narrow margin is notable because the H100 achieves it with fewer tensor cores (528 versus 896) and a lower boost clock (1980 MHz versus 2267 MHz), suggesting that the H100’s FP16 path is optimized for throughput per core, while the Rubin’s tensor cores are configured for a wider range of precisions.

Memory bandwidth is the clearest separator. The Rubin’s 22.1 TB/s is 6.58 times the H100’s 3.36 TB/s. For workloads that stream large tensors or require high-bandwidth attention mechanisms, this is a decisive advantage. Texture rate follows a similar pattern: the Rubin’s 2,031.2 GTexel/s is 94.3% higher than the H100’s 1,045.4 GTexel/s, matching the FP32 ratio because texture rate scales with the number of TMUs and clock speed. Pixel rate is closer: 54.41 GPixel/s on the Rubin versus 47.52 GPixel/s on the H100, a 14.5% difference, because both parts have only 24 ROPs.

The transistor and density figures also indicate architectural efficiency. The H100 achieves 66.91 FP32 TFLOPS from 80,000 million transistors, or roughly 0.84 TFLOPS per billion transistors. The Rubin achieves 130.0 FP32 TFLOPS from 336,000 million transistors, or roughly 0.39 TFLOPS per billion transistors. That lower efficiency per transistor is expected given the Rubin’s much larger memory subsystem and wider bus, but it shows that the H100 is a more compute-dense design per transistor, while the Rubin spends its transistor budget on capacity and bandwidth.

The Verdict

The data points to a clear division of roles. The NVIDIA H100 SXM5 96 GB is the part to select when FP16 tensor throughput is the priority, as it holds a 2.9% edge in FP16 TFLOPS over the Rubin, and when power and thermal limits are constrained, given its 700 W TDP versus the Rubin’s 2300 W. The H100 also suits workloads that fit within 96 GB of HBM3 memory and can operate on a PCIe 5.0 interface.

The NVIDIA Rubin GPU is the choice when memory capacity and bandwidth dominate. Its 288 GB of HBM4 and 22.1 TB/s of bandwidth are unmatched by the H100, and its FP32 throughput of 130.0 TFLOPS is nearly double. The Rubin also offers more shading units, more TMUs, more tensor cores, a higher boost clock, a larger die, and a newer process node. Its 1456 mm² die and 336,000 million transistors indicate a design that prioritizes scale over density.

For training large models that exceed 96 GB, the Rubin is the only viable option. For inference workloads that are bandwidth-bound, the Rubin’s 6.58x bandwidth advantage is transformative. For mixed-precision training where FP16 is the primary format and the model fits in 96 GB, the H100 remains competitive due to its FP16 ratio advantage and lower power draw. The Rubin’s 260.0 FP16 TFLOPS at a 2:1 ratio means it can sustain FP16 work while also supporting other precisions more evenly, whereas the H100’s 4:1 ratio is narrower.

Neither part has measured benchmark scores, so the verdict rests entirely on the specification differences. The H100 is the established, lower-power server part with a proven Hopper architecture. The Rubin is the successor architecture with a massive memory and bandwidth lead, a newer process, and a higher transistor count. The choice hinges on whether the workload is capacity-constrained or compute-constrained, and the data shows the Rubin wins on capacity, bandwidth, FP32, texture, and pixel rates, while the H100 wins on FP16 TFLOPS and power efficiency.

FAQ

Q: Which GPU has more memory bandwidth?

A: The NVIDIA Rubin GPU has 22.1 TB/s of bandwidth, which is 6.58 times the H100 SXM5’s 3.36 TB/s.

Q: How do the FP16 compute ratings compare?

A: The H100 delivers 267.6 TFLOPS at a 4:1 ratio, while the Rubin delivers 260.0 TFLOPS at a 2:1 ratio, giving the H100 a 2.9% advantage.

Q: What are the memory capacities of each part?

A: The H100 has 96 GB of HBM3, while the Rubin has 288 GB of HBM4.

Q: Which GPU uses a newer process node?

A: The Rubin uses TSMC’s 3 nm process, while the H100 uses TSMC’s 5 nm process.

Q: Do either of these GPUs have display outputs?

A: No, both are SXM modules with no display outputs listed.

Q: What is the transistor count difference?

A: The H100 has 80,000 million transistors on an 814 mm² die, while the Rubin has 336,000 million transistors on a 1456 mm² die.

Specification Differences

The two parts differ across nearly every measured field. The H100 uses the GH100 chip on the Hopper architecture, while the Rubin uses the GR100 chip on the Rubin architecture. The process node moves from 5 nm to 3 nm. Transistors increase from 80,000 million to 336,000 million, and die size increases from 814 mm² to 1456 mm². Transistor density rises from 98.3 million per mm² to 230.8 million per mm².

Clock speeds differ: the H100 has a base clock of 1350 MHz and a boost of 1980 MHz, while the Rubin has a base of 700 MHz and a boost of 2267 MHz. Memory clocks also differ: the H100 runs at 1313 MHz with 5.3 Gbps effective, while the Rubin runs at 2695 MHz with 10.8 Gbps effective. Memory type changes from HBM3 to HBM4, capacity from 96 GB to 288 GB, bus width from 5120 bit to 16384 bit, and bandwidth from 3.36 TB/s to 22.1 TB/s.

The compute arrays scale up: shading units go from 16896 to 28672, TMUs from 528 to 896, and tensor cores from 528 to 896. Both parts have 24 ROPs. Pixel rate rises from 47.52 GPixel/s to 54.41 GPixel/s, texture rate from 1,045.4 GTexel/s to 2,031.2 GTexel/s, and FP32 from 66.91 TFLOPS to 130.0 TFLOPS. FP16 changes from 267.6 TFLOPS at a 4:1 ratio to 260.0 TFLOPS at a 2:1 ratio.

Power figures diverge sharply: TDP goes from 700 W to 2300 W, and suggested PSU from 1100 W to 2700 W. The H100 uses an 8-pin EPS power connector, while the Rubin has no connector listed. The bus interface advances from PCIe 5.0 x16 to PCIe 6.0 x16. Both are SXM modules with no display outputs. The H100 has null API support for DirectX, OpenGL, and Vulkan, while the Rubin lists N/A for all three. Release dates are March 20, 2023 for the H100 and December 31, 2025 for the Rubin. The H100’s predecessor is Server Ada and its successor is Server Blackwell, while the Rubin’s predecessor is Server Blackwell and it has no successor listed.

DETAILED SPECIFICATIONS

SPECIFICATION
H100 SXM5 96 GB
Rubin GPU
Core Specs
Shading Units
16,896
28,672 +69.7%
Shaders
16,896
28,672 +69.7%
TMUs
528
896 +69.7%
ROPs
24
24 0.0%
SM Count
132
224 +69.7%
Clocks
Base Clock
1350 MHz
700 MHz
Boost Clock
1980 MHz
2267 MHz
Memory Clock
1313 MHz 5.3 Gbps effective
2695 MHz 10.8 Gbps effective
Memory
Memory Size
96 GB
288 GB
VRAM (MB)
98,304
294,912 +200.0%
Memory Type
HBM3
HBM4
Memory Bus
5120 bit
16384 bit
Bandwidth
3.36 TB/s
22.1 TB/s
Cache
L1 Cache
256 KB (per SM)
256 KB (per SM)
L2 Cache
50 MB
128 MB
Performance
Pixel Rate
47.52 GPixel/s
54.41 GPixel/s
Texture Rate
1,045.4 GTexel/s
2,031.2 GTexel/s
FP32 (TFLOPS)
66.91 TFLOPS
130.0 TFLOPS
FP64 (TFLOPS)
33.45 TFLOPS (1:2)
32.50 TFLOPS (1:4)
FP16 (TFLOPS)
267.6 TFLOPS (4:1)
260.0 TFLOPS (2:1)
AI/RT
Tensor Cores
528
896 +69.7%
Power
TDP
700 W
2300 W
TDP (W)
700
2,300 +228.6%
Suggested PSU
1100 W
2700 W
Power Connectors
8-pin EPS
—
Architecture
Architecture
Hopper
Rubin
GPU Name
GH100
GR100
Generation
Server Hopper (Hxx)
Server Rubin (Rxx)
Process Size
5 nm
3 nm
Transistors
80,000 million
336,000 million
Die Size
814 mm²
1456 mm²
Foundry
TSMC
TSMC
Density
98.3M / mm²
230.8M / mm²
API Support
OpenCL
3.0
3.0
CUDA
9.0
10.7
Physical
Slot Width
SXM Module
SXM Module
Outputs
No outputs
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 6.0 x16
Other
Production
Active
Active
Predecessor
Server Ada
Server Blackwell
Successor
Server Blackwell
—
View H100 SXM5 96 GB Details View Rubin GPU Details