NVIDIA GeForce RTX 5090 vs NVIDIA H200 NVL Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 5090

CORE STATE GB202
VRAM 32 GB
CLOCK SPEED 2407 MHz
TDP 575 W
BUS WIDTH 512 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

H200 NVL

CORE STATE GH100
VRAM 141 GB
CLOCK SPEED 1785 MHz
TDP 600 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2024

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
18,355
N/A
geekbench_opencl
334,370
334,891
geekbench_vulkan
376,728
N/A
passmark_directx_10
226
N/A
passmark_directx_11
341
N/A
passmark_directx_12
185
N/A
passmark_directx_9
395
N/A
passmark_g2d
1,413
N/A
passmark_g3d
39,650
N/A
passmark_gpu_compute
26,756
N/A

Analysis: NVIDIA GeForce RTX 5090 vs NVIDIA H200 NVL

Head-to-Head Benchmarks

The only directly comparable benchmark recorded in the database for these two accelerators is Geekbench OpenCL. In that test, the NVIDIA H200 NVL scores 334,891, while the GeForce RTX 5090 scores 334,370. The H200 NVL wins by a razor-thin margin of 0.2%. This is effectively a statistical tie, with both parts landing within a fraction of a percent of each other in raw compute throughput as measured by this workload.

Looking at the broader database context, the H200 NVL sits at the 100th percentile among all GPUs, meaning no recorded GPU scores higher. Its nearest rival, the NVIDIA B200, posts an average score of 345,482, which is 3.1% higher than the H200 NVL. The AMD Instinct MI300X trails by 5.3%, and the NVIDIA L40S is 13.2% behind. The H200 NVL's average benchmark score of 334,891 reflects only its single OpenCL entry, so the percentile ranking is based on that one measurement.

The GeForce RTX 5090, by contrast, holds the 92nd percentile across all GPUs. Its average benchmark score of 79,842 is pulled down by the inclusion of multiple Passmark tests, many of which are legacy DirectX workloads. Its nearest rivals in the database are older or lower-tier parts: the Tesla P100 PCIe 16 GB is 0.3% behind, the Tesla P100 PCIe 12 GB is 0.6% behind, and the AMD Radeon RX 6850M XT is 1.1% behind. The Pro Vega 64X is 1.4% ahead.

The RTX 5090's OpenCL score of 334,370 is essentially identical to the H200 NVL's 334,891, but its other benchmark results tell a more varied story. In 3DMark Steel Nomad DX12, it scores 18,355. In Geekbench Vulkan, it reaches 376,728, which is notably higher than its OpenCL result. Passmark results are mixed: 39,650 in G3D, 26,756 in GPU Compute, 1,413 in G2D, and much lower scores in the legacy DirectX tests (395 in DX9, 341 in DX11, 226 in DX10, 185 in DX12). These numbers indicate the RTX 5090 is heavily optimized for graphics-oriented workloads, while the H200 NVL has no graphics benchmark entries at all.

Where Each One Wins

The H200 NVL wins in the single shared benchmark, Geekbench OpenCL, by 0.2%. That is the only head-to-head victory recorded. However, the nature of the two products suggests different domains of strength based on the data available.

The RTX 5090 wins decisively in graphics and rendering workloads. Its 3DMark Steel Nomad DX12 score of 18,355 demonstrates strong DirectX 12 gaming performance. Vulkan performance is even more impressive at 376,728 in Geekbench Vulkan, which exceeds its own OpenCL score by roughly 12.7%. The Passmark G3D score of 39,650 and G2D score of 1,413 further confirm that this is a graphics-first product. The RTX 5090 also has display outputs (1x HDMI 2.1b, 3x DisplayPort 2.1b), meaning it can drive monitors directly, something the H200 NVL cannot do.

The H200 NVL has no display outputs and no graphics API support (DirectX, OpenGL, and Vulkan are all listed as N/A). Its strengths lie elsewhere. The 141 GB of HBM3e memory with 4.89 TB/s of bandwidth is a massive advantage for data-intensive workloads. The memory bus width of 6144 bit is 12 times wider than the RTX 5090's 512 bit. The H200 NVL also has a higher FP16 throughput of 120.6 TFLOPS compared to the RTX 5090's 104.8 TFLOPS, though the RTX 5090 matches that figure in FP32.

For compute workloads that favor OpenCL, the H200 NVL takes the narrow edge. For anything involving rasterization, ray tracing, or display output, the RTX 5090 is the clear winner based on its benchmark suite and feature set.

Architecture Differences

The two GPUs come from different architectural generations. The H200 NVL uses the GH100 chip built on the Hopper architecture, part of the Server Hopper (Hxx) generation. The RTX 5090 uses the GB202 chip on the Blackwell 2.0 architecture, part of the GeForce 50 series. Both are manufactured by TSMC on a 5 nm process, but the similarities end there.

The H200 NVL packs 80,000 million transistors on an 814 mm² die, resulting in a transistor density of 98.3 million per square millimeter. The RTX 5090 has 92,200 million transistors on a smaller 750 mm² die, giving it a higher density of 122.9 million per square millimeter. The RTX 5090 crams more transistors into a smaller area.

Clock speeds differ substantially. The H200 NVL runs at a base clock of 1365 MHz and a boost clock of 1785 MHz. The RTX 5090 runs significantly higher at 2017 MHz base and 2407 MHz boost. Memory clocks also diverge: the H200 NVL's HBM3e memory runs at 1593 MHz (6.4 Gbps effective), while the RTX 5090's GDDR7 runs at 1750 MHz (28 Gbps effective).

The compute unit counts tell a story of specialization. The H200 NVL has 16,896 shading units, 528 TMUs, and only 24 ROPs. The RTX 5090 has 21,760 shading units, 680 TMUs, and 176 ROPs. The massive difference in ROPs (176 vs 24) explains why the RTX 5090 achieves a pixel rate of 423.6 GPixel/s versus the H200 NVL's 42.84 GPixel/s. The RTX 5090 also has 170 RT cores and 680 tensor cores, while the H200 NVL lists 528 tensor cores and no RT core count.

The H200 NVL's FP32 throughput is 60.32 TFLOPS, while its FP16 throughput is 120.6 TFLOPS (2:1 ratio). The RTX 5090 delivers 104.8 TFLOPS in both FP32 and FP16 (1:1 ratio), meaning the RTX 5090 is 73.7% faster in FP32 but 13.1% slower in FP16. This reflects the H200 NVL's design priority toward mixed-precision compute.

Memory configurations are starkly different. The H200 NVL has 141 GB of HBM3e on a 6144 bit bus, delivering 4.89 TB/s of bandwidth. The RTX 5090 has 32 GB of GDDR7 on a 512 bit bus, delivering 1.79 TB/s. The H200 NVL has 4.4 times more memory capacity and 2.7 times more memory bandwidth.

Power and physical specifications also differ. The H200 NVL has a TDP of 600 W and uses an 8-pin EPS power connector, with a suggested PSU of 1000 W. The RTX 5090 has a TDP of 575 W, uses a single 16-pin connector, and suggests a 950 W PSU. Both are dual-slot cards. The H200 NVL measures 267 mm in length and 111 mm in height, while the RTX 5090 is longer at 304 mm and taller at 137 mm, with a width of 40 mm.

FAQ

Q: Which GPU is faster in OpenCL?

A: The H200 NVL scores 334,891 in Geekbench OpenCL, while the RTX 5090 scores 334,370. The H200 NVL wins by 0.2%.

Q: Can the H200 NVL output video to a display?

A: No. The H200 NVL has no display outputs. The RTX 5090 has 1x HDMI 2.1b and 3x DisplayPort 2.1b outputs.

Q: How much memory does each GPU have?

A: The H200 NVL has 141 GB of HBM3e memory with 4.89 TB/s bandwidth. The RTX 5090 has 32 GB of GDDR7 memory with 1.79 TB/s bandwidth.

Q: Which GPU has higher FP32 compute?

A: The RTX 5090 delivers 104.8 TFLOPS in FP32, compared to the H200 NVL's 60.32 TFLOPS. The RTX 5090 is 73.7% faster in this metric.

Q: Which GPU has higher FP16 compute?

A: The H200 NVL delivers 120.6 TFLOPS in FP16 (2:1 ratio), compared to the RTX 5090's 104.8 TFLOPS (1:1 ratio). The H200 NVL is 15.1% faster in FP16.

Q: What are the transistor counts?

A: The H200 NVL has 80,000 million transistors on an 814 mm² die. The RTX 5090 has 92,200 million transistors on a 750 mm² die.

Specification Differences

| Specification | NVIDIA H200 NVL | NVIDIA GeForce RTX 5090 |

|---|---|---|

| Architecture | Hopper | Blackwell 2.0 |

| Generation | Server Hopper (Hxx) | GeForce 50 |

| Chip | GH100 | GB202 |

| Transistors | 80,000 million | 92,200 million |

| Die Size | 814 mm² | 750 mm² |

| Transistor Density | 98.3M / mm² | 122.9M / mm² |

| Base Clock | 1365 MHz | 2017 MHz |

| Boost Clock | 1785 MHz | 2407 MHz |

| Memory Clock | 1593 MHz 6.4 Gbps effective | 1750 MHz 28 Gbps effective |

| Memory Size | 141 GB | 32 GB |

| Memory Type | HBM3e | GDDR7 |

| Memory Bus Width | 6144 bit | 512 bit |

| Memory Bandwidth | 4.89 TB/s | 1.79 TB/s |

| Shading Units | 16896 | 21760 |

| TMUs | 528 | 680 |

| ROPs | 24 | 176 |

| RT Cores | N/A | 170 |

| Tensor Cores | 528 | 680 |

| Pixel Rate | 42.84 GPixel/s | 423.6 GPixel/s |

| Texture Rate | 942.5 GTexel/s | 1,636.8 GTexel/s |

| FP32 | 60.32 TFLOPS | 104.8 TFLOPS |

| FP16 | 120.6 TFLOPS (2:1) | 104.8 TFLOPS (1:1) |

| TDP | 600 W | 575 W |

| Power Connectors | 8-pin EPS | 1x 16-pin |

| Suggested PSU | 1000 W | 950 W |

| Display Outputs | No outputs | 1x HDMI 2.1b, 3x DisplayPort 2.1b |

| DirectX | N/A | 12 Ultimate (12_2) |

| OpenGL | N/A | 4.6 |

| Vulkan | N/A | 1.4 |

| Length | 267 mm 10.5 inches | 304 mm 12 inches |

| Height | 111 mm 4.4 inches | 137 mm 5.4 inches |

| Width | N/A | 40 mm 1.6 inches |

| Release Date | 2024-11-17 | 2025-01-29 |

| Launch MSRP | N/A | 1,999 USD |

The Verdict

The data points to two fundamentally different products with a narrow overlap. In the only shared benchmark, Geekbench OpenCL, the H200 NVL edges out the RTX 5090 by 0.2%. That margin is negligible in practice, but it establishes that the H200 NVL holds its own in raw compute throughput despite its lower clock speeds and fewer shading units.

The RTX 5090 should be chosen by anyone needing graphics rendering, display output, or DirectX/Vulkan workloads. Its 3DMark Steel Nomad score of 18,355, Geekbench Vulkan score of 376,728, and Passmark G3D score of 39,650 demonstrate strong graphics performance. Its 176 ROPs and 423.6 GPixel/s pixel rate are orders of magnitude beyond the H200 NVL's 24 ROPs and 42.84 GPixel/s. The presence of RT cores, display outputs, and full API support makes it the only choice for interactive graphics.

The H200 NVL should be chosen for memory-bound compute tasks. Its 141 GB of HBM3e memory, 4.89 TB/s bandwidth, and higher FP16 throughput of 120.6 TFLOPS make it better suited for large-scale data processing and mixed-precision workloads. The 100th percentile ranking among all GPUs indicates that no other recorded GPU outperforms it on average.

The RTX 5090's 92nd percentile ranking is lower, but this is partly due to its benchmark mix including legacy Passmark tests. Its average score of 79,842 is dragged down by scores like 226 in Passmark DirectX 10 and 185 in Passmark DirectX 12, which are not representative of its modern compute capabilities.

For users who need both compute and graphics, the RTX 5090 is the more versatile part. For users who need maximum memory capacity and bandwidth in a server context, the H200 NVL is the clear choice. The 0.2% OpenCL difference is too small to be a deciding factor; the decision rests on memory capacity, FP16 throughput, and graphics capability.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 5090
H200 NVL
Core Specs
Shading Units
21,760
16,896 -22.4%
Shaders
21,760
16,896 -22.4%
TMUs
680
528 -22.4%
ROPs
176
24 -86.4%
SM Count
170
132 -22.4%
Clocks
Base Clock
2017 MHz
1365 MHz
Boost Clock
2407 MHz
1785 MHz
Memory Clock
1750 MHz 28 Gbps effective
1593 MHz 6.4 Gbps effective
Memory
Memory Size
32 GB
141 GB
VRAM (MB)
32,768
144,384 +340.6%
Memory Type
GDDR7
HBM3e
Memory Bus
512 bit
6144 bit
Bandwidth
1.79 TB/s
4.89 TB/s
Cache
L1 Cache
128 KB (per SM)
256 KB (per SM)
L2 Cache
96 MB
50 MB
Performance
Pixel Rate
423.6 GPixel/s
42.84 GPixel/s
Texture Rate
1,636.8 GTexel/s
942.5 GTexel/s
FP32 (TFLOPS)
104.8 TFLOPS
60.32 TFLOPS
FP64 (TFLOPS)
1.637 TFLOPS (1:64)
30.16 TFLOPS (1:2)
FP16 (TFLOPS)
104.8 TFLOPS (1:1)
120.6 TFLOPS (2:1)
AI/RT
RT Cores
170
Tensor Cores
680
528 -22.4%
Power
TDP
575 W
600 W
TDP (W)
575
600 +4.3%
Suggested PSU
950 W
1000 W
Power Connectors
1x 16-pin
8-pin EPS
Architecture
Architecture
Blackwell 2.0
Hopper
GPU Name
GB202
GH100
Generation
GeForce 50
Server Hopper (Hxx)
Process Size
5 nm
5 nm
Transistors
92,200 million
80,000 million
Die Size
750 mm²
814 mm²
Foundry
TSMC
TSMC
Density
122.9M / mm²
98.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
12.0
9.0
Shader Model
6.9
Physical
Slot Width
Dual-slot
Dual-slot
Length
304 mm 12 inches
267 mm 10.5 inches
Height
137 mm 5.4 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.1b3x DisplayPort 2.1b
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Launch Price
1,999 USD
Production
Active
Active
Predecessor
GeForce 40
Server Ada
Successor
GeForce 60
Server Blackwell
View GeForce RTX 5090 Details View H200 NVL Details