NVIDIA H100 NVL 94 GB vs NVIDIA Rubin GPU Comparison

NVIDIA
GEFORCE

NVIDIA H100 NVL 94 GB

CORE STATE GH100
VRAM 94 GB
CLOCK SPEED 1785 MHz
TDP 400 W
BUS WIDTH 6016 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

Rubin GPU

CORE STATE GR100
VRAM 288 GB
CLOCK SPEED 2267 MHz
TDP 2300 W
BUS WIDTH 16384 bit
ARCHITECTURE Rubin
nm
PROCESS 3 nm
LAUNCH DATE 2026

Analysis: NVIDIA H100 NVL 94 GB vs NVIDIA Rubin GPU

Head-to-Head Benchmarks

The recorded database contains no benchmark scores for either the NVIDIA H100 NVL 94 GB or the NVIDIA Rubin GPU. Both entries show an average benchmark score of zero, with no head-to-head benchmark results listed. The percentile ranking for both parts is identical at 50, placing them at the median of all GPUs tracked. Without measured performance data, direct numerical comparisons of rendering or compute workloads cannot be derived from the database.

The available data instead allows a comparison of theoretical specifications, which point to a decisive advantage for the Rubin GPU in most raw throughput metrics. The Rubin GPU delivers 130.0 TFLOPS of FP32 compute versus 60.32 TFLOPS for the H100 NVL, a margin of approximately 2.16 times. In FP16 operations, the Rubin GPU reaches 260.0 TFLOPS (2:1 ratio), while the H100 NVL achieves 241.3 TFLOPS (4:1 ratio), a narrower lead of about 8 percent. Texture rate shows a larger gap: the Rubin GPU processes 2,031.2 GTexel/s against 942.5 GTexel/s for the H100, again roughly 2.16 times higher. Pixel rate favors the Rubin GPU at 54.41 GPixel/s versus 42.84 GPixel/s, a 27 percent advantage.

Memory bandwidth is where the Rubin GPU separates itself most dramatically. The Rubin GPU features 22.1 TB/s of bandwidth across a 16384-bit bus, while the H100 NVL provides 3.94 TB/s over a 6016-bit interface. That represents a 5.6 times difference in memory throughput. Memory capacity follows a similar pattern: 288 GB of HBM4 on the Rubin GPU versus 94 GB of HBM3 on the H100 NVL, a 3.06 times increase. Effective memory clock rates are 10.8 Gbps for the Rubin and 5.2 Gbps for the H100, roughly doubling the per-pin data rate.

Clock speeds tell a different story. The H100 NVL has a higher base clock at 1080 MHz versus 700 MHz for the Rubin GPU. However, the Rubin GPU boosts to 2267 MHz, surpassing the H100's 1785 MHz boost clock by 27 percent. The Rubin GPU's higher boost clock, combined with its larger shader count of 28672 versus 16896, explains its substantial FP32 lead. Shader count increases by 69.7 percent, while the boost clock rises by 27 percent, and the product of these two factors accounts for the observed compute advantage.

Where Each One Wins

The H100 NVL wins in several practical areas based on the data. Its power draw of 400 W is far lower than the Rubin GPU's 2300 W. The H100 NVL uses a dual-slot form factor with an 8-pin EPS power connector, while the Rubin GPU is an SXM Module with no listed power connector. The H100 NVL also specifies a suggested power supply of 800 W, compared to 2700 W for the Rubin GPU. For systems with existing PCIe 5.0 infrastructure, the H100 NVL's bus interface matches that standard, whereas the Rubin GPU requires PCIe 6.0 x16 support. The H100 NVL has defined physical dimensions of 267 mm length and 111 mm height, while the Rubin GPU's dimensions are not recorded.

The Rubin GPU wins decisively in every measured computational and memory metric. Its FP32 output of 130.0 TFLOPS is more than double the H100's 60.32 TFLOPS. FP16 compute is 260.0 TFLOPS versus 241.3 TFLOPS, a smaller but still positive margin. Texture fill rate doubles, pixel rate increases by over a quarter, and memory bandwidth multiplies by a factor of 5.6. The Rubin GPU also uses a smaller process node at 3 nm versus 5 nm, packs 336,000 million transistors into a 1456 mm² die, and achieves a transistor density of 230.8M per mm². The H100 NVL contains 80,000 million transistors on an 814 mm² die with a density of 98.3M per mm².

Release timing favors the Rubin GPU, which is dated 2025-12-31, while the H100 NVL was released on 2023-03-20. The Rubin GPU succeeds Server Blackwell, while the H100 NVL's successor is also Server Blackwell, placing them in adjacent generations but with the Rubin as the newer design.

The Verdict

The database shows a clear generational split. The NVIDIA Rubin GPU delivers substantially higher raw performance in every computed category: FP32, FP16, texture rate, pixel rate, memory bandwidth, and memory capacity. Its 3 nm process, 336,000 million transistors, and 22.1 TB/s bandwidth place it in a different performance class than the H100 NVL.

The H100 NVL remains the more practical option for systems with power and form-factor constraints. Its 400 W TDP, dual-slot design, and 800 W suggested power supply fit standard server configurations. The Rubin GPU's 2300 W TDP and 2700 W suggested power supply require specialized infrastructure that most deployments do not have.

For workloads that are memory-bandwidth bound, the Rubin GPU is the only choice based on the data: 22.1 TB/s versus 3.94 TB/s is a 5.6 times difference. For compute-bound FP32 tasks, the Rubin GPU provides 2.16 times the throughput. For FP16 workloads, the advantage narrows to roughly 8 percent, but the Rubin GPU still leads.

The H100 NVL wins on compatibility. It uses PCIe 5.0, which is already widespread, while the Rubin GPU requires PCIe 6.0. The H100 NVL also has defined physical measurements that allow straightforward integration, whereas the Rubin GPU's dimensions are unlisted.

Neither part has recorded benchmark scores, so the verdict rests on specification data. That data indicates the Rubin GPU is the superior performer in absolute terms, while the H100 NVL is the more deployable part. Users with existing PCIe 5.0 systems and power budgets near 800 W should select the H100 NVL. Users planning new infrastructure with PCIe 6.0 and high-power delivery should select the Rubin GPU.

FAQ

Q: Which GPU has higher FP32 performance?

A: The NVIDIA Rubin GPU achieves 130.0 TFLOPS, while the NVIDIA H100 NVL 94 GB delivers 60.32 TFLOPS. The Rubin GPU is approximately 2.16 times faster in FP32.

Q: How do the memory bandwidths compare?

A: The Rubin GPU provides 22.1 TB/s across a 16384-bit bus, while the H100 NVL offers 3.94 TB/s over a 6016-bit bus. The Rubin GPU has 5.6 times the memory bandwidth.

Q: What is the power consumption difference?

A: The H100 NVL has a TDP of 400 W with a suggested power supply of 800 W. The Rubin GPU has a TDP of 2300 W and a suggested power supply of 2700 W.

Q: Which GPU uses newer memory technology?

A: The Rubin GPU uses HBM4 memory with 288 GB capacity. The H100 NVL uses HBM3 memory with 94 GB capacity.

Q: What are the process nodes for each GPU?

A: The H100 NVL is built on a 5 nm process by TSMC. The Rubin GPU uses a 3 nm process, also by TSMC.

Q: Do either of these GPUs have display outputs?

A: No. Both the H100 NVL and the Rubin GPU list "No outputs" for display connections.

Architecture Differences

The two GPUs come from different architectural generations. The H100 NVL uses the GH100 chip based on the Hopper architecture, belonging to the Server Hopper (Hxx) generation. The Rubin GPU uses the GR100 chip based on the Rubin architecture, belonging to the Server Rubin (Rxx) generation.

The transistor counts differ by a factor of 4.2. The H100 NVL contains 80,000 million transistors on an 814 mm² die, yielding a density of 98.3M per mm². The Rubin GPU contains 336,000 million transistors on a 1456 mm² die, yielding a density of 230.8M per mm². The die size increases by 79 percent, while transistor density more than doubles, indicating a significant architecture density improvement beyond the node shrink alone.

Shader resources scale accordingly. The H100 NVL has 16896 shading units, 528 TMUs, and 24 ROPs. The Rubin GPU has 28672 shading units, 896 TMUs, and 24 ROPs. Shading units increase by 69.7 percent, TMUs by 69.7 percent, while ROPs remain constant at 24. Tensor core counts follow the TMU pattern: 528 on the H100 versus 896 on the Rubin, a 69.7 percent increase.

Memory architecture diverges sharply. The H100 NVL uses HBM3 with a 6016-bit bus and 5.2 Gbps effective memory clock. The Rubin GPU uses HBM4 with a 16384-bit bus and 10.8 Gbps effective memory clock. The bus width increases by 172 percent, and the effective clock doubles, combining to produce the 5.6 times bandwidth improvement.

Neither GPU has ray tracing cores listed in the database. The H100 NVL has no API support data recorded, while the Rubin GPU lists DirectX, OpenGL, and Vulkan as N/A. Both are compute-focused server parts without display outputs.

Specification Differences

The following fields differ between the two GPUs:

  • Chip: GH100 (H100 NVL) versus GR100 (Rubin GPU)
  • Architecture: Hopper versus Rubin
  • Generation: Server Hopper (Hxx) versus Server Rubin (Rxx)
  • Process node: 5 nm versus 3 nm
  • Transistors: 80,000 million versus 336,000 million
  • Die size: 814 mm² versus 1456 mm²
  • Transistor density: 98.3M / mm² versus 230.8M / mm²
  • Base clock: 1080 MHz versus 700 MHz
  • Boost clock: 1785 MHz versus 2267 MHz
  • Memory clock: 1310 MHz, 5.2 Gbps effective versus 2695 MHz, 10.8 Gbps effective
  • Memory size: 94 GB versus 288 GB
  • Memory type: HBM3 versus HBM4
  • Memory bus width: 6016 bit versus 16384 bit
  • Memory bandwidth: 3.94 TB/s versus 22.1 TB/s
  • Shading units: 16896 versus 28672
  • TMUs: 528 versus 896
  • Tensor cores: 528 versus 896
  • Pixel rate: 42.84 GPixel/s versus 54.41 GPixel/s
  • Texture rate: 942.5 GTexel/s versus 2,031.2 GTexel/s
  • FP32: 60.32 TFLOPS versus 130.0 TFLOPS
  • FP16: 241.3 TFLOPS (4:1) versus 260.0 TFLOPS (2:1)
  • TDP: 400 W versus 2300 W
  • Slot width: Dual-slot versus SXM Module
  • Power connectors: 8-pin EPS versus not listed
  • Suggested PSU: 800 W versus 2700 W
  • Bus interface: PCIe 5.0 x16 versus PCIe 6.0 x16
  • APIs: not listed versus N/A for DirectX, OpenGL, and Vulkan
  • Dimensions: 267 mm length, 111 mm height versus not listed
  • Release date: 2023-03-20 versus 2025-12-31
  • Predecessor: Server Ada versus Server Blackwell
  • Successor: Server Blackwell versus not listed

DETAILED SPECIFICATIONS

SPECIFICATION
H100 NVL 94 GB
Rubin GPU
Core Specs
Shading Units
16,896
28,672 +69.7%
Shaders
16,896
28,672 +69.7%
TMUs
528
896 +69.7%
ROPs
24
24 0.0%
SM Count
132
224 +69.7%
Clocks
Base Clock
1080 MHz
700 MHz
Boost Clock
1785 MHz
2267 MHz
Memory Clock
1310 MHz 5.2 Gbps effective
2695 MHz 10.8 Gbps effective
Memory
Memory Size
94 GB
288 GB
VRAM (MB)
96,256
294,912 +206.4%
Memory Type
HBM3
HBM4
Memory Bus
6016 bit
16384 bit
Bandwidth
3.94 TB/s
22.1 TB/s
Cache
L1 Cache
256 KB (per SM)
256 KB (per SM)
L2 Cache
50 MB
128 MB
Performance
Pixel Rate
42.84 GPixel/s
54.41 GPixel/s
Texture Rate
942.5 GTexel/s
2,031.2 GTexel/s
FP32 (TFLOPS)
60.32 TFLOPS
130.0 TFLOPS
FP64 (TFLOPS)
30.16 TFLOPS (1:2)
32.50 TFLOPS (1:4)
FP16 (TFLOPS)
241.3 TFLOPS (4:1)
260.0 TFLOPS (2:1)
AI/RT
Tensor Cores
528
896 +69.7%
Power
TDP
400 W
2300 W
TDP (W)
400
2,300 +475.0%
Suggested PSU
800 W
2700 W
Power Connectors
8-pin EPS
—
Architecture
Architecture
Hopper
Rubin
GPU Name
GH100
GR100
Generation
Server Hopper (Hxx)
Server Rubin (Rxx)
Process Size
5 nm
3 nm
Transistors
80,000 million
336,000 million
Die Size
814 mm²
1456 mm²
Foundry
TSMC
TSMC
Density
98.3M / mm²
230.8M / mm²
API Support
OpenCL
3.0
3.0
CUDA
9.0
10.7
Physical
Slot Width
Dual-slot
SXM Module
Length
267 mm 10.5 inches
—
Height
111 mm 4.4 inches
—
Outputs
No outputs
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 6.0 x16
Other
Production
Active
Active
Predecessor
Server Ada
Server Blackwell
Successor
Server Blackwell
—
View H100 NVL 94 GB Details View Rubin GPU Details