NVIDIA GeForce RTX 5090 D V2 vs NVIDIA H20 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 5090 D V2

CORE STATE GB202
VRAM 24 GB
CLOCK SPEED 2407 MHz
TDP 575 W
BUS WIDTH 384 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

H20

CORE STATE GH100
VRAM 96 GB
CLOCK SPEED 1980 MHz
TDP 500 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2024

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
16,504
N/A

Analysis: NVIDIA GeForce RTX 5090 D V2 vs NVIDIA H20

Head-to-Head Benchmarks

The data available for a direct comparison is limited. The GeForce RTX 5090 D V2 has a recorded average benchmark score of 16504 in the 3DMark Steel Nomad DX12 test, placing it at the 59th percentile among all GPUs in the database. The H20 has no recorded benchmark scores and sits at the 50th percentile, with an average score of 0. This means the RTX 5090 D V2 is the only one of the two with measurable gaming performance data.

The RTX 5090 D V2's score of 16504 places it in a tight competitive cluster. Its nearest rival, the NVIDIA T400, scores 16508, a delta of 0%. The AMD Radeon PRO W7500 trails by 0.5% with a score of 16415. The NVIDIA RTX PRO 6000 Blackwell is 0.6% behind at 16408, and the AMD Radeon RX 5700 XT scores 16361, a 0.9% deficit. These margins are all within a single percentage point, indicating that the RTX 5090 D V2 is positioned squarely alongside these cards in this specific DirectX 12 workload, rather than delivering a decisive victory over them.

In terms of raw compute capacity, the RTX 5090 D V2 delivers 104.8 TFLOPS of FP32 performance and 104.8 TFLOPS of FP16 performance on a 1:1 ratio. The H20 delivers 39.54 TFLOPS of FP32 and 79.07 TFLOPS of FP16 on a 2:1 ratio. The FP32 advantage for the RTX 5090 D V2 is substantial, roughly 2.65 times the H20's figure. However, the H20's FP16 output is closer, though still lower, at about 75% of the RTX 5090 D V2's FP16 capability. These numbers indicate different design priorities, with the RTX 5090 D V2 favoring general-purpose and gaming workloads, while the H20's architecture is geared toward specific server tasks.

Memory bandwidth is another major divider. The RTX 5090 D V2 accesses 24 GB of GDDR7 memory over a 384-bit bus, yielding 1.34 TB/s of bandwidth. The H20 accesses 96 GB of HBM3 memory over a 6144-bit bus, yielding 4.03 TB/s of bandwidth. The H20 offers three times the memory capacity and roughly three times the bandwidth. That bandwidth advantage is critical for data-intensive server workloads.

FAQ

Q: Which GPU has a higher FP32 compute throughput?

A: The NVIDIA GeForce RTX 5090 D V2, with 104.8 TFLOPS versus 39.54 TFLOPS for the NVIDIA H20.

Q: How do their memory capacities and bandwidths compare?

A: The H20 has 96 GB of HBM3 memory with 4.03 TB/s bandwidth, while the RTX 5090 D V2 has 24 GB of GDDR7 memory with 1.34 TB/s bandwidth.

Q: What is the transistor count and die size for each chip?

A: The RTX 5090 D V2 uses the GB202 chip with 92,200 million transistors on a 750 mm² die. The H20 uses the GH100 chip with 80,000 million transistors on an 814 mm² die.

Q: Do both GPUs support PCIe 5.0?

A: Yes, both use a PCIe 5.0 x16 bus interface.

Q: Are there any display outputs on the H20?

A: No, the H20 has no display outputs. The RTX 5090 D V2 has 1x HDMI 2.1b and 3x DisplayPort 2.1b outputs.

Q: Which GPU has a higher boost clock?

A: The RTX 5090 D V2 has a boost clock of 2407 MHz, compared to the H20's boost clock of 1980 MHz.

Where Each One Wins

The GeForce RTX 5090 D V2 is the clear choice for graphics-rendering and gaming tasks. It is the only one of the two with any benchmark data, and its 16504 score in 3DMark Steel Nomad DX12 confirms it can execute modern DirectX 12 workloads. It also leads heavily in pixel throughput at 423.6 GPixel/s versus 47.52 GPixel/s for the H20, and in texture rate at 1,636.8 GTexel/s versus 617.8 GTexel/s. The RTX 5090 D V2 also has 170 RT cores for ray tracing, a feature the H20 lacks entirely. Its 1:1 FP16 ratio means it does not sacrifice half-rate performance for half-precision math, which benefits general compute and graphics.

The H20 wins decisively in memory-centric server applications. Its 96 GB of HBM3 memory and 4.03 TB/s bandwidth dwarf the RTX 5090 D V2's 24 GB and 1.34 TB/s. For large language model inference or scientific simulations that require massive datasets resident in fast memory, the H20's capacity and bandwidth are the dominant factors. Its FP16 output of 79.07 TFLOPS, while lower than the RTX 5090 D V2's 104.8 TFLOPS, is still substantial and is delivered through a 2:1 ratio architecture optimized for tensor operations. The H20 also comes in an SXM Module form factor, which is designed for dense server integration, whereas the RTX 5090 D V2 is a dual-slot PCIe card.

Specification Differences

The two GPUs diverge on nearly every core specification. The RTX 5090 D V2 has 21,760 shading units, 680 TMUs, and 176 ROPs. The H20 has 9,984 shading units, 312 TMUs, and 24 ROPs. The RTX 5090 D V2 also has 680 tensor cores, while the H20 has 312. The RTX 5090 D V2 has 170 RT cores; the H20 has none. Clock speeds differ as well, with the RTX 5090 D V2 boosting to 2407 MHz versus the H20's 1980 MHz.

Memory specifications are fundamentally different. The RTX 5090 D V2 uses 24 GB of GDDR7 on a 384-bit bus with 1.34 TB/s bandwidth. The H20 uses 96 GB of HBM3 on a 6144-bit bus with 4.03 TB/s bandwidth. The power profiles also differ: the RTX 5090 D V2 has a TDP of 575 W and requires a 950 W suggested PSU, while the H20 has a TDP of 500 W and a 900 W suggested PSU. The RTX 5090 D V2 has a 1x 16-pin power connector; the H20 has no listed power connector, consistent with its SXM Module design.

Physical and interface differences are stark. The RTX 5090 D V2 is a dual-slot card measuring 304 mm in length, 137 mm in height, and 48 mm in width, with display outputs and support for DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4. The H20 is an SXM Module with no dimensions listed, no display outputs, and no graphics API support. Both use PCIe 5.0 x16.

Architecture Differences

The RTX 5090 D V2 is built on the Blackwell 2.0 architecture using the GB202 chip. The H20 is built on the Hopper architecture using the GH100 chip. Both are fabricated by TSMC on a 5 nm process, but the transistor counts and densities differ. The GB202 packs 92,200 million transistors into 750 mm² for a density of 122.9M per mm². The GH100 houses 80,000 million transistors in 814 mm² for a density of 98.3M per mm². The RTX 5090 D V2 achieves a higher transistor density on a smaller die.

The architectural focus diverges sharply. The Blackwell 2.0 design in the RTX 5090 D V2 includes dedicated RT cores for ray tracing, a feature absent from the H20. The RTX 5090 D V2 also has a 1:1 FP16 ratio, meaning its FP16 throughput matches its FP32 throughput. The H20's Hopper architecture uses a 2:1 FP16 ratio, which indicates a design that prioritizes tensor and mixed-precision workloads over general-purpose FP32 compute. The H20's tensor core count of 312 is lower than the RTX 5090 D V2's 680, but the H20's HBM3 memory subsystem is engineered for the massive bandwidth demands of AI training and inference.

The RTX 5090 D V2 is part of the GeForce 50-series, with a predecessor in the GeForce 40 series and a successor in the GeForce 60 series. The H20 belongs to the Server Hopper (Hxx) generation, with a predecessor in Server Ada and a successor in Server Blackwell. This places the two cards in different product families with different lifecycle targets: one for consumer and professional graphics, the other for datacenter acceleration.

The Verdict

The recorded data points to two entirely different tools. The GeForce RTX 5090 D V2 is a graphics-first card. It has a measurable 3DMark Steel Nomad DX12 score of 16504, a high FP32 throughput of 104.8 TFLOPS, a 423.6 GPixel/s pixel rate, and 170 RT cores. It is the only one of the two with display outputs, graphics API support, and a benchmark record. Any workload involving rendering, ray tracing, or DirectX 12 gaming belongs to this card.

The H20 is a memory-capacity specialist with no graphics benchmark data. Its 96 GB of HBM3 memory and 4.03 TB/s bandwidth are its defining characteristics. For server-side applications that need to hold very large models or datasets in memory, the H20 is the only option between the two. Its 79.07 TFLOPS of FP16 performance, delivered through a 2:1 ratio, indicates an architecture tuned for tensor operations rather than rasterization.

The choice depends entirely on the workload. For interactive graphics, gaming, or any task requiring a display output, the RTX 5090 D V2 is the functional choice. For datacenter deployments focused on memory-bound compute, the H20's capacity and bandwidth are unmatched by the RTX 5090 D V2. The RTX 5090 D V2 was released on 2025-08-14 with a launch MSRP of 2,299 USD. The H20 was released earlier on 2024-01-31. Both remain in active production. The data does not support a single winner; it supports two winners in separate domains.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 5090 D V2
H20
Core Specs
Shading Units
21,760
9,984 -54.1%
Shaders
21,760
9,984 -54.1%
TMUs
680
312 -54.1%
ROPs
176
24 -86.4%
SM Count
170
78 -54.1%
Clocks
Base Clock
2017 MHz
1830 MHz
Boost Clock
2407 MHz
1980 MHz
Memory Clock
1750 MHz 28 Gbps effective
1313 MHz 5.3 Gbps effective
Memory
Memory Size
24 GB
96 GB
VRAM (MB)
24,576
98,304 +300.0%
Memory Type
GDDR7
HBM3
Memory Bus
384 bit
6144 bit
Bandwidth
1.34 TB/s
4.03 TB/s
Cache
L1 Cache
128 KB (per SM)
256 KB (per SM)
L2 Cache
96 MB
60 MB
Performance
Pixel Rate
423.6 GPixel/s
47.52 GPixel/s
Texture Rate
1,636.8 GTexel/s
617.8 GTexel/s
FP32 (TFLOPS)
104.8 TFLOPS
39.54 TFLOPS
FP64 (TFLOPS)
1.637 TFLOPS (1:64)
19.77 TFLOPS (1:2)
FP16 (TFLOPS)
104.8 TFLOPS (1:1)
79.07 TFLOPS (2:1)
AI/RT
RT Cores
170
Tensor Cores
680
312 -54.1%
Power
TDP
575 W
500 W
TDP (W)
575
500 -13.0%
Suggested PSU
950 W
900 W
Power Connectors
1x 16-pin
Architecture
Architecture
Blackwell 2.0
Hopper
GPU Name
GB202
GH100
Generation
GeForce 50
Server Hopper (Hxx)
Process Size
5 nm
5 nm
Transistors
92,200 million
80,000 million
Die Size
750 mm²
814 mm²
Foundry
TSMC
TSMC
Density
122.9M / mm²
98.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
12.0
9.0
Shader Model
6.9
Physical
Slot Width
Dual-slot
SXM Module
Length
304 mm 12 inches
Height
137 mm 5.4 inches
Outputs
1x HDMI 2.1b3x DisplayPort 2.1b
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Launch Price
2,299 USD
Production
Active
Active
Predecessor
GeForce 40
Server Ada
Successor
GeForce 60
Server Blackwell
View GeForce RTX 5090 D V2 Details View H20 Details