NVIDIA GeForce RTX 5090 SE vs NVIDIA H20 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 5090 SE

CORE STATE GB202
VRAM 24 GB
CLOCK SPEED 2377 MHz
TDP 500 W
BUS WIDTH 384 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2026
VS
NVIDIA
GEFORCE

H20

CORE STATE GH100
VRAM 96 GB
CLOCK SPEED 1980 MHz
TDP 500 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2024

Analysis: NVIDIA GeForce RTX 5090 SE vs NVIDIA H20

Where Each One Wins

The recorded data separates these two NVIDIA accelerators into distinct roles, and the benchmark wins split accordingly. The GeForce RTX 5090 SE takes the consumer graphics workload category outright. Its shading unit count of 14,080 against 9,984 on the H20, combined with 440 TMUs versus 312, gives it a clear advantage in rasterization-heavy tasks. The pixel rate of 380.3 GPixel/s versus 47.52 GPixel/s shows a massive gap in fill-rate-bound scenarios. The RTX 5090 SE also carries 160 ROPs, while the H20 has only 24, which means traditional frame rendering favors the GeForce card without question.

The NVIDIA H20 wins in the compute and memory capacity domain. Its 96 GB of HBM3 memory dwarfs the 24 GB of GDDR7 on the RTX 5090 SE. The memory bus width of 6144 bit versus 384 bit, and the resulting bandwidth of 4.03 TB/s versus 1.34 TB/s, positions the H20 for large-model inference and training workloads that require rapid data movement across massive parameter sets. The H20 also shows a 2:1 FP16 ratio, delivering 79.07 TFLOPS, which exceeds its own FP32 figure of 39.54 TFLOPS. The RTX 5090 SE runs FP16 at a 1:1 ratio, matching its FP32 at 66.94 TFLOPS. This means the H20 has a higher peak FP16 throughput, which matters for mixed-precision AI workloads.

The use-case split is unambiguous. The RTX 5090 SE is built for rendering, display output, and consumer-facing graphics APIs. It supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, plus it has display outputs including 1x HDMI 2.1b and 3x DisplayPort 2.1b. The H20 has no display outputs and lists its DirectX, OpenGL, and Vulkan support as N/A. The H20 exists for server-side computation, specifically in the Server Hopper (Hxx) generation, where graphics output is irrelevant. The data confirms a clean division: one card wins on graphics throughput and feature completeness for end-user systems, the other wins on memory scale, FP16 compute density, and server integration.

Architecture Differences

The architecture split starts at the chip level. The RTX 5090 SE uses the GB202 chip under the Blackwell 2.0 architecture, while the H20 uses the GH100 chip under the Hopper architecture. Both are fabricated on a 5 nm process at TSMC, but the die layouts differ significantly. The GB202 die measures 750 mm² and holds 92,200 million transistors, giving a transistor density of 122.9M per mm². The GH100 die is larger at 814 mm² but contains fewer transistors at 80,000 million, resulting in a lower density of 98.3M per mm². This indicates the Blackwell design packs logic more tightly, which aligns with its higher shading unit count and FP32 throughput.

The memory architecture diverges completely. The RTX 5090 SE uses GDDR7 across a 384-bit bus, achieving 1.34 TB/s. The H20 uses HBM3 across a 6144-bit bus, achieving 4.03 TB/s. The HBM3 implementation provides three times the bandwidth and four times the capacity, but it comes in a different form factor. The RTX 5090 SE is a dual-slot card with a 1x 16-pin power connector, while the H20 is an SXM Module with no discrete power connector listed. Both share a 500 W TDP and a suggested PSU of 900 W, but the physical integration differs: the H20 slots into a server chassis, while the RTX 5090 SE installs into a standard PCIe slot.

The compute feature sets also differ. The RTX 5090 SE has 110 RT cores and 440 tensor cores. The H20 lists tensor cores at 312 but has no RT core count recorded. The shading units, TMUs, and ROPs all favor the RTX 5090 SE, as noted. Clock behavior shows another split: the RTX 5090 SE runs a base of 1740 MHz and boosts to 2377 MHz, while the H20 runs a higher base of 1830 MHz but a lower boost of 1980 MHz. The memory clocks differ as well, with the RTX 5090 SE at 1750 MHz (28 Gbps effective) versus the H20 at 1313 MHz (5.3 Gbps effective), though the HBM3 bus width compensates.

The API support marks a hard boundary. The RTX 5090 SE supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4. The H20 lists N/A for all three. This is not a minor omission; it reflects the intended usage. The RTX 5090 SE is a graphics card that can run games and rendering applications. The H20 is an accelerator without a display stack, focused on compute kernels. The release dates reinforce this: the RTX 5090 SE has a release date of 2025-12-31, while the H20 came out on 2024-01-31. The H20 also has a predecessor of Server Ada and a successor of Server Blackwell, while the RTX 5090 SE sits in the GeForce 50 generation with a predecessor of GeForce 40 and a successor of GeForce 60.

Head-to-Head Benchmarks

The database contains no direct head-to-head benchmark entries between these two products, and neither has recorded benchmark scores. The wins tally shows zero for both sides. However, the specification data allows for direct numerical comparisons that serve the same analytical purpose.

The most decisive gap appears in pixel throughput. The RTX 5090 SE delivers 380.3 GPixel/s, which is exactly 8 times the H20's 47.52 GPixel/s. This ratio comes from the ROP count difference: 160 versus 24. The H20's ROP count is unusually low for a chip of its size, which indicates it was never designed for traditional frame rasterization. Conversely, the RTX 5090 SE's texture rate of 1,045.9 GTexel/s versus 617.8 GTexel/s on the H20 shows a 1.7x advantage in texture-heavy scenes, driven by 440 TMUs versus 312.

The FP32 compute comparison favors the RTX 5090 SE. Its 66.94 TFLOPS is 1.7x higher than the H20's 39.54 TFLOPS. This means general-purpose single-precision workloads, including many scientific simulations and graphics shaders, will run faster on the GeForce card. The FP16 comparison flips the result. The H20 reaches 79.07 TFLOPS at FP16, while the RTX 5090 SE stays at 66.94 TFLOPS. The H20 advantage is 18%, and it comes from the 2:1 FP16 architecture, which doubles the throughput relative to FP32. The RTX 5090 SE's 1:1 ratio means FP16 offers no additional headroom.

The memory bandwidth gap is the largest single-number difference in the dataset. The H20's 4.03 TB/s is three times the RTX 5090 SE's 1.34 TB/s. For workloads that are bandwidth-bound, such as large matrix multiplications or embedding lookups, this threefold advantage dominates. The capacity gap is even starker: 96 GB versus 24 GB is a 4x difference. A model that requires 40 GB of resident weights runs entirely on the H20 but exceeds the RTX 5090 SE's capacity entirely. The bus width difference of 6144 bit versus 384 bit explains how the H20 achieves this despite its lower memory clock.

Clock speeds tell a secondary story. The RTX 5090 SE boosts to 2377 MHz, which is 20% higher than the H20's 1980 MHz boost. The base clocks are closer, with the H20 at 1830 MHz versus 1740 MHz on the RTX 5090 SE. This means the H20 starts closer to its peak but has less headroom, while the RTX 5090 SE has a wider clock range and reaches a higher absolute frequency under load. The pixel rate difference is so large that clocks become secondary, but they do explain part of the FP32 gap.

The Verdict

The data points to separate buyers for each product. The RTX 5090 SE is the choice for any workload that ends in pixels on a screen. Its 160 ROPs, 440 TMUs, 14,080 shading units, and 110 RT cores provide the hardware path for high-resolution rendering, ray tracing, and VR applications. The 66.94 TFLOPS FP32 throughput handles graphics shaders efficiently, and the display outputs (1x HDMI 2.1b, 3x DisplayPort 2.1b) connect directly to monitors. The 24 GB GDDR7 capacity is sufficient for contemporary game textures and professional rendering scenes. The launch MSRP is 1,499 USD.

The NVIDIA H20 is the choice for server-side compute, specifically AI training and inference. Its 96 GB HBM3 memory allows it to hold large language models and embedding tables that would not fit on the RTX 5090 SE. The 4.03 TB/s bandwidth feeds data to the 312 tensor cores faster, and the 79.07 TFLOPS FP16 throughput exceeds the GeForce card by 18%. The lack of display outputs and graphics API support is irrelevant in a server rack. The SXM Module form factor fits into existing Hopper server infrastructure, and the 500 W TDP matches the RTX 5090 SE, meaning power delivery requirements are identical at the system level.

Neither card wins outright because they serve different functions. The database shows zero head-to-head benchmarks, which is consistent with this conclusion: they do not compete in the same market segment. A buyer choosing between them is actually choosing between a graphics workstation and a compute accelerator. The RTX 5090 SE's pixel rate of 380.3 GPixel/s versus 47.52 GPixel/s is a 8x gap in favor of the GeForce card for display workloads. The H20's memory bandwidth of 4.03 TB/s versus 1.34 TB/s is a 3x gap in favor of the Hopper card for data movement. The correct choice depends entirely on whether the task requires rendering output or large-scale numerical computation.

FAQ

Q: Which card has higher FP32 compute?

A: The RTX 5090 SE has 66.94 TFLOPS FP32, which is 1.7x the H20's 39.54 TFLOPS.

Q: Which card has more memory bandwidth?

A: The H20 has 4.03 TB/s from HBM3 across a 6144-bit bus, compared to 1.34 TB/s from GDDR7 across a 384-bit bus on the RTX 5090 SE.

Q: Can the H20 output video to a display?

A: No. The H20 lists no display outputs and its DirectX, OpenGL, and Vulkan support are all N/A. The RTX 5090 SE has 1x HDMI 2.1b and 3x DisplayPort 2.1b.

Q: What is the FP16 performance difference?

A: The H20 reaches 79.07 TFLOPS FP16 with a 2:1 ratio, while the RTX 5090 SE has 66.94 TFLOPS FP16 at a 1:1 ratio. The H20 is 18% ahead.

Q: Which card has more transistors?

A: The RTX 5090 SE has 92,200 million transistors on a 750 mm² die. The H20 has 80,000 million transistors on a larger 814 mm² die.

Q: Do both cards use the same power requirement?

A: Yes, both have a 500 W TDP and a suggested PSU of 900 W. The RTX 5090 SE uses a 1x 16-pin connector, while the H20 is an SXM Module without a listed connector.

Specification Differences

| Specification | NVIDIA GeForce RTX 5090 SE | NVIDIA H20 |

|---|---|---|

| Chip | GB202 | GH100 |

| Architecture | Blackwell 2.0 | Hopper |

| Generation | GeForce 50 | Server Hopper (Hxx) |

| Transistors | 92,200 million | 80,000 million |

| Die Size | 750 mm² | 814 mm² |

| Transistor Density | 122.9M / mm² | 98.3M / mm² |

| Base Clock | 1740 MHz | 1830 MHz |

| Boost Clock | 2377 MHz | 1980 MHz |

| Memory Clock | 1750 MHz 28 Gbps effective | 1313 MHz 5.3 Gbps effective |

| Memory Size | 24 GB | 96 GB |

| Memory Type | GDDR7 | HBM3 |

| Memory Bus Width | 384 bit | 6144 bit |

| Memory Bandwidth | 1.34 TB/s | 4.03 TB/s |

| Shading Units | 14080 | 9984 |

| TMUs | 440 | 312 |

| ROPs | 160 | 24 |

| RT Cores | 110 | null |

| Tensor Cores | 440 | 312 |

| Pixel Rate | 380.3 GPixel/s | 47.52 GPixel/s |

| Texture Rate | 1,045.9 GTexel/s | 617.8 GTexel/s |

| FP32 | 66.94 TFLOPS | 39.54 TFLOPS |

| FP16 | 66.94 TFLOPS (1:1) | 79.07 TFLOPS (2:1) |

| Slot Width | Dual-slot | SXM Module |

| Power Connectors | 1x 16-pin | null |

| Display Outputs | 1x HDMI 2.1b, 3x DisplayPort 2.1b | No outputs |

| DirectX | 12 Ultimate (12_2) | N/A |

| OpenGL | 4.6 | N/A |

| Vulkan | 1.4 | N/A |

| Release Date | 2025-12-31 | 2024-01-31 |

| Predecessor | GeForce 40 | Server Ada |

| Successor | GeForce 60 | Server Blackwell |

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 5090 SE
H20
Core Specs
Shading Units
14,080
9,984 -29.1%
Shaders
14,080
9,984 -29.1%
TMUs
440
312 -29.1%
ROPs
160
24 -85.0%
SM Count
110
78 -29.1%
Clocks
Base Clock
1740 MHz
1830 MHz
Boost Clock
2377 MHz
1980 MHz
Memory Clock
1750 MHz 28 Gbps effective
1313 MHz 5.3 Gbps effective
Memory
Memory Size
24 GB
96 GB
VRAM (MB)
24,576
98,304 +300.0%
Memory Type
GDDR7
HBM3
Memory Bus
384 bit
6144 bit
Bandwidth
1.34 TB/s
4.03 TB/s
Cache
L1 Cache
128 KB (per SM)
256 KB (per SM)
L2 Cache
96 MB
60 MB
Performance
Pixel Rate
380.3 GPixel/s
47.52 GPixel/s
Texture Rate
1,045.9 GTexel/s
617.8 GTexel/s
FP32 (TFLOPS)
66.94 TFLOPS
39.54 TFLOPS
FP64 (TFLOPS)
1,045.9 GFLOPS (1:64)
19.77 TFLOPS (1:2)
FP16 (TFLOPS)
66.94 TFLOPS (1:1)
79.07 TFLOPS (2:1)
AI/RT
RT Cores
110
Tensor Cores
440
312 -29.1%
Power
TDP
500 W
500 W
TDP (W)
500
500 0.0%
Suggested PSU
900 W
900 W
Power Connectors
1x 16-pin
Architecture
Architecture
Blackwell 2.0
Hopper
GPU Name
GB202
GH100
Generation
GeForce 50
Server Hopper (Hxx)
Process Size
5 nm
5 nm
Transistors
92,200 million
80,000 million
Die Size
750 mm²
814 mm²
Foundry
TSMC
TSMC
Density
122.9M / mm²
98.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
12.0
9.0
Shader Model
6.9
Physical
Slot Width
Dual-slot
SXM Module
Length
267 mm 10.5 inches
Height
111 mm 4.4 inches
Outputs
1x HDMI 2.1b3x DisplayPort 2.1b
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Launch Price
1,499 USD
Production
Active
Active
Predecessor
GeForce 40
Server Ada
Successor
GeForce 60
Server Blackwell
View GeForce RTX 5090 SE Details View H20 Details