NVIDIA H200 SXM 141 GB vs NVIDIA Rubin GPU Comparison

NVIDIA
GEFORCE

NVIDIA H200 SXM 141 GB

CORE STATE GH100
VRAM 141 GB
CLOCK SPEED 1980 MHz
TDP 700 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2024
VS
NVIDIA
GEFORCE

Rubin GPU

CORE STATE GR100
VRAM 288 GB
CLOCK SPEED 2267 MHz
TDP 2300 W
BUS WIDTH 16384 bit
ARCHITECTURE Rubin
nm
PROCESS 3 nm
LAUNCH DATE 2026

Analysis: NVIDIA H200 SXM 141 GB vs NVIDIA Rubin GPU

Head-to-Head Benchmarks

The recorded database contains no head-to-head benchmark entries for the NVIDIA H200 SXM 141 GB and the NVIDIA Rubin GPU. Both parts carry an average benchmark score of zero and a percentile ranking of 50 against all GPUs in the database, which reflects the absence of measured performance data rather than parity in capability. The H200 SXM 141 GB belongs to the Server Hopper generation and has shipped as an active product, while the Rubin GPU is listed as active with a release date in the database, but neither has accumulated any recorded benchmark runs.

Without direct benchmark scores, the comparison must rely on the computed throughput figures stored in the specification fields. The Rubin GPU delivers 130.0 TFLOPS of FP32 compute and 260.0 TFLOPS of FP16 compute (2:1), against the H200 SXM 141 GB's 66.91 TFLOPS FP32 and 133.8 TFLOPS FP16 (2:1). Those numbers indicate the Rubin GPU holds a 1.94x advantage in both FP32 and FP16 throughput, essentially doubling the arithmetic rate of the older part. Texture rate follows a similar pattern: the Rubin GPU reaches 2,031.2 GTexel/s versus 1,045.4 GTexel/s for the H200 SXM 141 GB, again a 1.94x difference. Pixel rate is closer, with the Rubin GPU at 54.41 GPixel/s and the H200 SXM 141 GB at 47.52 GPixel/s, a modest 1.14x gap that reflects the identical 24 ROP count on both chips.

Memory bandwidth shows the largest single-specification divergence. The Rubin GPU pairs 288 GB of HBM4 with a 16384-bit bus to reach 22.1 TB/s, while the H200 SXM 141 GB uses 141 GB of HBM3e on a 6144-bit bus for 4.89 TB/s. That is a 4.52x bandwidth advantage for the Rubin GPU, a far larger relative gap than the compute or texture deltas. The H200 SXM 141 GB does hold one advantage in the clock domain: its base clock runs at 1500 MHz versus 700 MHz for the Rubin GPU, a 2.14x lead at idle frequency. The boost clocks reverse that relationship, with the Rubin GPU at 2267 MHz against 1980 MHz for the H200 SXM 141 GB, an 1.14x advantage under load. Memory clock also favors the newer part, with 2695 MHz and 10.8 Gbps effective on the Rubin GPU versus 1593 MHz and 6.4 Gbps effective on the H200 SXM 141 GB.

The transistor and density figures further separate the two designs. The Rubin GPU integrates 336,000 million transistors on a 1456 mm² die, while the H200 SXM 141 GB integrates 80,000 million transistors on an 814 mm² die. That places the Rubin GPU at 230.8M transistors per mm² versus 98.3M per mm² for the H200 SXM 141 GB, a 2.35x density improvement. The Rubin GPU uses TSMC's 3 nm process, while the H200 SXM 141 GB uses TSMC's 5 nm node.

FAQ

Q: Which GPU has more memory bandwidth?

A: The NVIDIA Rubin GPU delivers 22.1 TB/s over a 16384-bit HBM4 interface, versus 4.89 TB/s over a 6144-bit HBM3e interface on the NVIDIA H200 SXM 141 GB. That is a 4.52x bandwidth advantage for the Rubin GPU.

Q: What are the FP32 compute differences?

A: The Rubin GPU reaches 130.0 TFLOPS of FP32, while the H200 SXM 141 GB reaches 66.91 TFLOPS. The Rubin GPU provides 1.94x the FP32 throughput of the H200 SXM 141 GB.

Q: How do the memory capacities compare?

A: The Rubin GPU carries 288 GB of HBM4, while the H200 SXM 141 GB carries 141 GB of HBM3e. The Rubin GPU offers 2.04x the memory capacity.

Q: Which GPU has a higher boost clock?

A: The Rubin GPU boosts to 2267 MHz, while the H200 SXM 141 GB boosts to 1980 MHz. The Rubin GPU's boost clock is 1.14x higher, despite its base clock of 700 MHz being 2.14x lower than the H200 SXM 141 GB's 1500 MHz base.

Q: What process nodes do the two GPUs use?

A: The Rubin GPU is built on TSMC's 3 nm process, while the H200 SXM 141 GB is built on TSMC's 5 nm process. The 3 nm node contributes to a transistor density of 230.8M per mm² on the Rubin GPU versus 98.3M per mm² on the H200 SXM 141 GB.

Q: Do either of these GPUs have display outputs?

A: No. Both the NVIDIA H200 SXM 141 GB and the NVIDIA Rubin GPU list "No outputs" in their display output fields, and both have N/A values for DirectX, OpenGL, and Vulkan support.

Where Each One Wins

The NVIDIA Rubin GPU wins on every throughput metric recorded in the database. FP32 compute, FP16 compute, texture rate, pixel rate, memory bandwidth, and memory capacity all favor the newer part. The FP32 and FP16 figures show a 1.94x advantage, texture rate shows a 1.94x advantage, and memory bandwidth shows a 4.52x advantage. The Rubin GPU also carries 2.04x the memory capacity of the H200 SXM 141 GB, with 288 GB versus 141 GB. For any workload that stresses raw arithmetic throughput, texture filtering, or memory transfer rates, the recorded data points exclusively to the Rubin GPU.

The NVIDIA H200 SXM 141 GB wins in two narrower categories. Its base clock of 1500 MHz is 2.14x higher than the Rubin GPU's 700 MHz, which can translate to lower idle power draw and steadier operation at low utilization. The H200 SXM 141 GB also runs at a 700 W TDP with a suggested PSU of 1100 W, while the Rubin GPU lists a 2300 W TDP and a suggested PSU of 2700 W. System integration differences also favor the older part: the H200 SXM 141 GB uses PCIe 5.0 x16 and an 8-pin EPS connector, while the Rubin GPU uses PCIe 6.0 x16 and lists no power connector data. The H200 SXM 141 GB's pixel rate of 47.52 GPixel/s is within 1.14x of the Rubin GPU's 54.41 GPixel/s, so rasterization-bound tasks see a smaller absolute gap than compute or memory workloads.

Both GPUs share the SXM Module slot width, the same 24 ROP count, and no display outputs. Neither part has recorded benchmark scores, so the win split in the database stands at zero wins each. The specification comparison, however, clearly assigns compute, memory, and texture wins to the Rubin GPU, with the H200 SXM 141 GB retaining advantages only in base clock and power delivery requirements.

Specification Differences

The two GPUs differ across nearly every recorded field. The H200 SXM 141 GB uses the GH100 chip with the Hopper architecture and belongs to the Server Hopper (Hxx) generation. The Rubin GPU uses the GR100 chip with the Rubin architecture and belongs to the Server Rubin (Rxx) generation. The H200 SXM 141 GB is fabricated on TSMC's 5 nm process with 80,000 million transistors on an 814 mm² die, while the Rubin GPU is fabricated on TSMC's 3 nm process with 336,000 million transistors on a 1456 mm² die. Transistor density reads 98.3M per mm² for the H200 SXM 141 GB and 230.8M per mm² for the Rubin GPU.

Clock specifications diverge substantially. The H200 SXM 141 GB runs at 1500 MHz base and 1980 MHz boost, with memory at 1593 MHz and 6.4 Gbps effective. The Rubin GPU runs at 700 MHz base and 2267 MHz boost, with memory at 2695 MHz and 10.8 Gbps effective. Memory configuration differs completely: 141 GB of HBM3e on a 6144-bit bus with 4.89 TB/s bandwidth for the H200 SXM 141 GB, versus 288 GB of HBM4 on a 16384-bit bus with 22.1 TB/s bandwidth for the Rubin GPU.

Compute resources scale with the newer part. The H200 SXM 141 GB has 16896 shading units, 528 TMUs, 24 ROPs, and 528 tensor cores, while the Rubin GPU has 28672 shading units, 896 TMUs, 24 ROPs, and 896 tensor cores. Pixel rate is 47.52 GPixel/s for the H200 SXM 141 GB and 54.41 GPixel/s for the Rubin GPU. Texture rate is 1,045.4 GTexel/s for the H200 SXM 141 GB and 2,031.2 GTexel/s for the Rubin GPU. FP32 output is 66.91 TFLOPS versus 130.0 TFLOPS, and FP16 output is 133.8 TFLOPS versus 260.0 TFLOPS.

Power figures show the Rubin GPU as a far higher-consumption part: 2300 W TDP with a suggested PSU of 2700 W, versus 700 W TDP with a suggested PSU of 1100 W for the H200 SXM 141 GB. The H200 SXM 141 GB lists an 8-pin EPS power connector, while the Rubin GPU lists no connector data. The bus interface advances from PCIe 5.0 x16 on the H200 SXM 141 GB to PCIe 6.0 x16 on the Rubin GPU. Both use the SXM Module slot width and have no display outputs. The H200 SXM 141 GB's predecessor is Server Ada and its successor is Server Blackwell, while the Rubin GPU's predecessor is Server Blackwell and it has no successor listed.

Architecture Differences

The architecture split is a full generation apart. The H200 SXM 141 GB uses the Hopper architecture, built around the GH100 chip, while the Rubin GPU uses the newer Rubin architecture, built around the GR100 chip. The Rubin architecture doubles the shading unit count from 16896 to 28672, doubles the TMU count from 528 to 896, and doubles the tensor core count from 528 to 896. The ROP count remains unchanged at 24 on both parts.

The manufacturing process accounts for a significant portion of the performance gap. TSMC's 3 nm node on the Rubin GPU enables 336,000 million transistors, a 4.2x increase over the H200 SXM 141 GB's 80,000 million transistors on 5 nm. The die grows from 814 mm² to 1456 mm², a 1.79x area increase, while transistor density improves from 98.3M per mm² to 230.8M per mm², a 2.35x density gain. This density improvement allows the Rubin GPU to fit more compute units and a wider memory interface within a single package.

Memory architecture changes are the most pronounced. The H200 SXM 141 GB uses HBM3e with a 6144-bit bus, while the Rubin GPU uses HBM4 with a 16384-bit bus. The bus width grows 2.67x, and the memory clock rises from 1593 MHz to 2695 MHz, with effective data rate increasing from 6.4 Gbps to 10.8 Gbps. Combined with the larger capacity, these changes produce the 22.1 TB/s bandwidth figure, a 4.52x improvement over 4.89 TB/s. The Rubin GPU also doubles memory capacity from 141 GB to 288 GB.

Clock behavior differs by design. The H200 SXM 141 GB runs a 1500 MHz base clock, while the Rubin GPU runs a 700 MHz base clock, a 2.14x lower idle frequency. The boost clocks tell the opposite story: the Rubin GPU boosts to 2267 MHz, 1.14x higher than the H200 SXM 141 GB's 1980 MHz. The lower base clock on the Rubin GPU likely reflects the higher transistor count and the need to manage the 2300 W TDP budget, while the higher boost clock allows the architecture to reach peak throughput when thermal and power headroom allow.

The bus interface advances from PCIe 5.0 x16 on the H200 SXM 141 GB to PCIe 6.0 x16 on the Rubin GPU. Both parts are SXM modules with no display outputs and no graphics API support. The H200 SXM 141 GB carries an 8-pin EPS power connector, while the Rubin GPU lists no power connector data. The power envelope grows from 700 W to 2300 W, and the suggested PSU scales from 1100 W to 2700 W. These architecture differences position the Rubin GPU as a substantially larger, faster, and more power-hungry server accelerator, while the H200 SXM 141 GB remains a lower-power entry point within the same SXM form factor.

DETAILED SPECIFICATIONS

SPECIFICATION
H200 SXM 141 GB
Rubin GPU
Core Specs
Shading Units
16,896
28,672 +69.7%
Shaders
16,896
28,672 +69.7%
TMUs
528
896 +69.7%
ROPs
24
24 0.0%
SM Count
132
224 +69.7%
Clocks
Base Clock
1500 MHz
700 MHz
Boost Clock
1980 MHz
2267 MHz
Memory Clock
1593 MHz 6.4 Gbps effective
2695 MHz 10.8 Gbps effective
Memory
Memory Size
141 GB
288 GB
VRAM (MB)
144,384
294,912 +104.3%
Memory Type
HBM3e
HBM4
Memory Bus
6144 bit
16384 bit
Bandwidth
4.89 TB/s
22.1 TB/s
Cache
L1 Cache
256 KB (per SM)
256 KB (per SM)
L2 Cache
50 MB
128 MB
Performance
Pixel Rate
47.52 GPixel/s
54.41 GPixel/s
Texture Rate
1,045.4 GTexel/s
2,031.2 GTexel/s
FP32 (TFLOPS)
66.91 TFLOPS
130.0 TFLOPS
FP64 (TFLOPS)
33.45 TFLOPS (1:2)
32.50 TFLOPS (1:4)
FP16 (TFLOPS)
133.8 TFLOPS (2:1)
260.0 TFLOPS (2:1)
AI/RT
Tensor Cores
528
896 +69.7%
Power
TDP
700 W
2300 W
TDP (W)
700
2,300 +228.6%
Suggested PSU
1100 W
2700 W
Power Connectors
8-pin EPS
—
Architecture
Architecture
Hopper
Rubin
GPU Name
GH100
GR100
Generation
Server Hopper (Hxx)
Server Rubin (Rxx)
Process Size
5 nm
3 nm
Transistors
80,000 million
336,000 million
Die Size
814 mm²
1456 mm²
Foundry
TSMC
TSMC
Density
98.3M / mm²
230.8M / mm²
API Support
OpenCL
3.0
3.0
CUDA
9.0
10.7
Physical
Slot Width
SXM Module
SXM Module
Outputs
No outputs
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 6.0 x16
Other
Production
Active
Active
Predecessor
Server Ada
Server Blackwell
Successor
Server Blackwell
—
View H200 SXM 141 GB Details View Rubin GPU Details