NVIDIA GeForce RTX 5090 SE vs NVIDIA H20 NVL16 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 5090 SE

CORE STATE GB202
VRAM 24 GB
CLOCK SPEED 2377 MHz
TDP 500 W
BUS WIDTH 384 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2026
VS
NVIDIA
GEFORCE

H20 NVL16

CORE STATE GH100
VRAM 96 GB
CLOCK SPEED 1980 MHz
TDP 400 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2025

Analysis: NVIDIA GeForce RTX 5090 SE vs NVIDIA H20 NVL16

FAQ

Q: What are the core architectural identities of the RTX 5090 SE and the H20 NVL16?

A: The RTX 5090 SE is built on the GB202 chip using the Blackwell 2.0 architecture, while the H20 NVL16 uses the GH100 chip with the Hopper architecture. Both are fabricated by TSMC on a 5 nm process node, but the RTX 5090 SE has a transistor count of 92,200 million on a 750 mm² die, whereas the H20 NVL16 has 80,000 million transistors on a larger 814 mm² die.

Q: How do the memory subsystems differ between these two NVIDIA GPUs?

A: The RTX 5090 SE carries 24 GB of GDDR7 memory on a 384-bit bus, delivering 1.34 TB/s of bandwidth. The H20 NVL16 instead uses 96 GB of HBM3 memory on a 6144-bit bus, achieving 4.03 TB/s. The H20 NVL16 has triple the capacity and roughly three times the bandwidth, making it a memory-focused design.

Q: Which GPU has higher raw FP32 compute performance?

A: The RTX 5090 SE delivers 66.94 TFLOPS of FP32 performance, while the H20 NVL16 provides 39.54 TFLOPS. That puts the RTX 5090 SE about 69% ahead of the H20 NVL16 in standard single-precision floating-point throughput, based on the recorded figures.

Q: What is the FP16 performance situation for each card?

A: The RTX 5090 SE offers 66.94 TFLOPS of FP16 with a 1:1 ratio, meaning FP16 throughput matches FP32. The H20 NVL16 provides 79.07 TFLOPS of FP16 with a 2:1 ratio, meaning it doubles its FP32 rate. In FP16, the H20 NVL16 leads by roughly 18% despite having lower FP32 performance.

Q: Are there differences in display outputs and API support?

A: Yes, the RTX 5090 SE includes 1x HDMI 2.1b and 3x DisplayPort 2.1b outputs, and supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The H20 NVL16 has no display outputs and lists N/A for DirectX, OpenGL, and Vulkan, reflecting its server-oriented design without graphics interface support.

Q: What are the power and physical form factor differences?

A: The RTX 5090 SE has a 500 W TDP, is dual-slot, measures 267 mm in length, 111 mm in height, and 40 mm in width, and uses a single 16-pin power connector with a suggested 900 W PSU. The H20 NVL16 has a 400 W TDP, is an SXM module, has no listed dimensions, and uses a suggested 800 W PSU with no power connector specified.

Architecture Differences

The RTX 5090 SE and the H20 NVL16 represent two distinct NVIDIA architectures with different design goals. The RTX 5090 SE uses the GB202 chip under Blackwell 2.0, a consumer-oriented architecture optimized for graphics workloads. The H20 NVL16 uses the GH100 chip under Hopper, a server-focused architecture designed for datacenter compute tasks.

The process node is identical at 5 nm from TSMC, but the physical implementations diverge. The RTX 5090 SE packs 92,200 million transistors into a 750 mm² die, yielding a transistor density of 122.9M per mm². The H20 NVL16 integrates 80,000 million transistors across an 814 mm² die, resulting in a lower density of 98.3M per mm². The larger die for the H20 NVL16 with fewer transistors suggests a design prioritizing memory bandwidth and server features over raw transistor packing.

The shading unit counts differ substantially. The RTX 5090 SE contains 14,080 shading units, 440 TMUs, and 160 ROPs. The H20 NVL16 has 9,984 shading units, 312 TMUs, and only 24 ROPs. The extremely low ROP count on the H20 NVL16 indicates it is not designed for rasterization-heavy graphics work, while the RTX 5090 SE's higher ROP count supports traditional rendering pipelines.

Ray tracing cores are present on the RTX 5090 SE at 110 units, while the H20 NVL16 lists null for RT cores, meaning it does not carry dedicated ray tracing hardware in the recorded data. Tensor cores also differ: the RTX 5090 SE has 440 tensor cores, and the H20 NVL16 has 312 tensor cores. However, the H20 NVL16's FP16 performance at 79.07 TFLOPS exceeds the RTX 5090 SE's 66.94 TFLOPS, indicating that the H20 NVL16's tensor core implementation is tuned for higher throughput in mixed-precision AI workloads.

The clock behavior separates the two as well. The RTX 5090 SE runs at a base of 1740 MHz and boosts to 2377 MHz. The H20 NVL16 has a base clock of 1830 MHz and a boost of 1980 MHz. The RTX 5090 SE has a higher boost clock, but the H20 NVL16 starts from a higher base, reflecting different thermal and power envelopes.

Where Each One Wins

The RTX 5090 SE wins in scenarios that depend on high FP32 throughput, graphics rendering, and display output. Its 66.94 TFLOPS FP32 performance is 69% above the H20 NVL16's 39.54 TFLOPS, giving it a clear advantage in standard compute tasks that use single-precision arithmetic. The presence of 160 ROPs versus just 24 on the H20 NVL16 means the RTX 5090 SE can handle pixel-heavy workloads far more effectively, with a pixel rate of 380.3 GPixel/s compared to 47.52 GPixel/s. The RTX 5090 SE also provides full API support for DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, plus HDMI and DisplayPort outputs, making it the only one of the two that can drive a display.

The H20 NVL16 wins in memory-bound and mixed-precision scenarios. Its 96 GB of HBM3 memory with 4.03 TB/s bandwidth dwarfs the RTX 5090 SE's 24 GB GDDR7 at 1.34 TB/s. For large datasets that exceed 24 GB, the H20 NVL16 is the only viable option between the two. Its FP16 throughput of 79.07 TFLOPS exceeds the RTX 5090 SE's 66.94 TFLOPS, giving it an edge in AI inference and training workloads that rely on FP16 arithmetic. The H20 NVL16's texture rate of 617.8 GTexel/s, while lower than the RTX 5090 SE's 1,045.9 GTexel/s, still represents substantial throughput for its class.

The power draw favors the H20 NVL16 in efficiency terms. At 400 W TDP versus 500 W for the RTX 5090 SE, the H20 NVL16 delivers its FP16 performance at lower power consumption. The suggested PSU of 800 W versus 900 W reinforces this difference. However, the RTX 5090 SE's higher FP32 output per watt appears favorable for single-precision compute relative to its power budget.

Specification Differences

The two GPUs differ across nearly every major specification field. The RTX 5090 SE uses the GB202 chip with Blackwell 2.0 architecture, while the H20 NVL16 uses the GH100 chip with Hopper architecture. The RTX 5090 SE is part of the GeForce 50 generation with a predecessor in GeForce 40 and successor in GeForce 60, whereas the H20 NVL16 belongs to the Server Hopper generation with predecessor Server Ada and successor Server Blackwell.

Transistor counts and die sizes are distinct: 92,200 million transistors on 750 mm² for the RTX 5090 SE versus 80,000 million on 814 mm² for the H20 NVL16. Transistor density follows at 122.9M per mm² versus 98.3M per mm².

Clock speeds show both differences and similarities. Base clocks are close at 1740 MHz for the RTX 5090 SE and 1830 MHz for the H20 NVL16, but boost clocks differ more significantly at 2377 MHz versus 1980 MHz. Memory clocks are completely different: the RTX 5090 SE runs at 1750 MHz with 28 Gbps effective, while the H20 NVL16 runs at 1313 MHz with 5.3 Gbps effective.

Memory capacity, type, and bus width all differ. The RTX 5090 SE has 24 GB GDDR7 on a 384-bit bus, and the H20 NVL16 has 96 GB HBM3 on a 6144-bit bus. Bandwidth follows accordingly at 1.34 TB/s versus 4.03 TB/s.

Compute resources diverge: shading units at 14,080 versus 9,984, TMUs at 440 versus 312, ROPs at 160 versus 24. RT cores are 110 on the RTX 5090 SE and null on the H20 NVL16. Tensor cores are 440 versus 312.

Pixel rate and texture rate differ substantially: 380.3 GPixel/s versus 47.52 GPixel/s, and 1,045.9 GTexel/s versus 617.8 GTexel/s. FP32 is 66.94 TFLOPS versus 39.54 TFLOPS, while FP16 is 66.94 TFLOPS (1:1) versus 79.07 TFLOPS (2:1).

Power specifications show 500 W TDP for the RTX 5090 SE versus 400 W for the H20 NVL16. The RTX 5090 SE is dual-slot with a 16-pin connector and 900 W suggested PSU, while the H20 NVL16 is an SXM module with no power connector listed and 800 W suggested PSU.

Display outputs and APIs are exclusive to the RTX 5090 SE. The H20 NVL16 has no outputs and N/A for all APIs. Dimensions are only listed for the RTX 5090 SE at 267 mm length, 111 mm height, and 40 mm width.

Release dates differ, with the RTX 5090 SE launching on 2025-12-31 and the H20 NVL16 on 2025-09-01. The RTX 5090 SE has a launch MSRP of 1,499 USD, while the H20 NVL16 has no launch MSRP recorded.

Head-to-Head Benchmarks

The recorded data contains no direct benchmark scores for either GPU, with both listing empty benchmark arrays and an average benchmark score of 0. The percentile versus all GPUs is 50 for both, indicating they sit at the median in the database's overall distribution, though this is based on the incomplete benchmark data.

Without direct head-to-head benchmark results, the comparison must rely on the specification-derived metrics. The largest win for the RTX 5090 SE comes in FP32 performance, where 66.94 TFLOPS compares to 39.54 TFLOPS for the H20 NVL16. This represents a 69% advantage, the most significant single-metric gap in the comparison.

The pixel rate provides another decisive RTX 5090 SE win: 380.3 GPixel/s versus 47.52 GPixel/s, a difference of roughly 8 times. The ROP count of 160 versus 24 explains this disparity and indicates the RTX 5090 SE is built for rasterization while the H20 NVL16 is not.

The H20 NVL16's biggest wins come in memory and FP16 compute. Its 4.03 TB/s bandwidth is 3 times the RTX 5090 SE's 1.34 TB/s, and its 96 GB capacity is 4 times the RTX 5090 SE's 24 GB. In FP16, the H20 NVL16 achieves 79.07 TFLOPS versus 66.94 TFLOPS, a 18% lead that flips the compute advantage toward the server card in mixed-precision workloads.

The texture rate favors the RTX 5090 SE at 1,045.9 GTexel/s versus 617.8 GTexel/s, a 69% advantage that aligns with its higher TMU count of 440 versus 312. The H20 NVL16's higher base clock of 1830 MHz versus 1740 MHz gives it a small initial speed advantage, but the RTX 5090 SE's boost clock of 2377 MHz versus 1980 MHz reverses that under load.

The Verdict

The data shows two GPUs built for different purposes. The RTX 5090 SE is a graphics-first card with 14,080 shading units, 160 ROPs, 110 RT cores, and full display and API support. Its FP32 performance of 66.94 TFLOPS and pixel rate of 380.3 GPixel/s make it the choice for rendering, gaming, and single-precision compute. The 24 GB GDDR7 memory is sufficient for standard graphics workloads but limits large-scale data processing.

The H20 NVL16 is a compute-first server module. Its 96 GB HBM3 memory and 4.03 TB/s bandwidth position it for large memory footprints, and its FP16 performance of 79.07 TFLOPS at 2:1 ratio makes it stronger for AI workloads that rely on reduced precision. The lack of display outputs and graphics APIs confirms it is not intended for interactive use.

For users needing graphics output, ray tracing, or maximum FP32 throughput, the RTX 5090 SE is the only valid option between the two. For datacenter deployments where memory capacity, bandwidth, and FP16 performance matter more than rasterization, the H20 NVL16 offers advantages in memory and mixed-precision compute, while drawing 100 W less power. The choice depends entirely on whether the workload is graphics-oriented or memory-intensive server compute, as the specification split makes each GPU dominant in its own domain.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 5090 SE
H20 NVL16
Core Specs
Shading Units
14,080
9,984 -29.1%
Shaders
14,080
9,984 -29.1%
TMUs
440
312 -29.1%
ROPs
160
24 -85.0%
SM Count
110
78 -29.1%
Clocks
Base Clock
1740 MHz
1830 MHz
Boost Clock
2377 MHz
1980 MHz
Memory Clock
1750 MHz 28 Gbps effective
1313 MHz 5.3 Gbps effective
Memory
Memory Size
24 GB
96 GB
VRAM (MB)
24,576
98,304 +300.0%
Memory Type
GDDR7
HBM3
Memory Bus
384 bit
6144 bit
Bandwidth
1.34 TB/s
4.03 TB/s
Cache
L1 Cache
128 KB (per SM)
256 KB (per SM)
L2 Cache
96 MB
60 MB
Performance
Pixel Rate
380.3 GPixel/s
47.52 GPixel/s
Texture Rate
1,045.9 GTexel/s
617.8 GTexel/s
FP32 (TFLOPS)
66.94 TFLOPS
39.54 TFLOPS
FP64 (TFLOPS)
1,045.9 GFLOPS (1:64)
19.77 TFLOPS (1:2)
FP16 (TFLOPS)
66.94 TFLOPS (1:1)
79.07 TFLOPS (2:1)
AI/RT
RT Cores
110
—
Tensor Cores
440
312 -29.1%
Power
TDP
500 W
400 W
TDP (W)
500
400 -20.0%
Suggested PSU
900 W
800 W
Power Connectors
1x 16-pin
—
Architecture
Architecture
Blackwell 2.0
Hopper
GPU Name
GB202
GH100
Generation
GeForce 50
Server Hopper (Hxx)
Process Size
5 nm
5 nm
Transistors
92,200 million
80,000 million
Die Size
750 mm²
814 mm²
Foundry
TSMC
TSMC
Density
122.9M / mm²
98.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
—
OpenGL
4.6
—
Vulkan
1.4
—
OpenCL
3.0
3.0
CUDA
12.0
9.0
Shader Model
6.9
—
Physical
Slot Width
Dual-slot
SXM Module
Length
267 mm 10.5 inches
—
Height
111 mm 4.4 inches
—
Outputs
1x HDMI 2.1b3x DisplayPort 2.1b
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Launch Price
1,499 USD
—
Production
Active
Active
Predecessor
GeForce 40
Server Ada
Successor
GeForce 60
Server Blackwell
View GeForce RTX 5090 SE Details View H20 NVL16 Details