NVIDIA GeForce RTX 5070 SUPER vs NVIDIA H20 NVL16 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 5070 SUPER

CORE STATE GB205
VRAM 18 GB
CLOCK SPEED 2512 MHz
TDP 275 W
BUS WIDTH 192 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2026
VS
NVIDIA
GEFORCE

H20 NVL16

CORE STATE GH100
VRAM 96 GB
CLOCK SPEED 1980 MHz
TDP 400 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2025

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
2,690
N/A

Analysis: NVIDIA GeForce RTX 5070 SUPER vs NVIDIA H20 NVL16

Where Each One Wins

The recorded data splits these two NVIDIA parts into entirely different roles. The GeForce RTX 5070 SUPER is a client-side graphics card with one benchmark entry: a 3DMark Steel Nomad DX12 score of 2690. That places it in the 18th percentile of all GPUs in the database, with nearest rivals including the Quadro K1100M, GT 1030, Arc Pro B50, and GT 440, all within a 1.7% delta. This is a real-time rendering part, built for DirectX 12 Ultimate workloads, with display outputs for HDMI 2.1b and DisplayPort 2.1b.

The H20 NVL16 has no benchmark scores recorded, no display outputs, and no DirectX, OpenGL, or Vulkan API support. Its measured strengths are entirely in compute and memory capacity. It carries 96 GB of HBM3 memory on a 6144-bit bus, delivering 4.03 TB/s of bandwidth, versus the RTX 5070 SUPER's 18 GB of GDDR7 on a 192-bit bus at 672.0 GB/s. The H20 NVL16 also has higher FP32 throughput at 39.54 TFLOPS and much higher FP16 throughput at 79.07 TFLOPS, thanks to a 2:1 FP16 ratio, whereas the RTX 5070 SUPER runs FP16 at 32.15 TFLOPS (1:1).

The wins are split by workload type. The RTX 5070 SUPER wins in any scenario that requires rasterization, ray tracing, or standard graphics APIs. It has 50 RT cores and 200 tensor cores, plus 80 ROPs and 200 TMUs. The H20 NVL16 has no RT core data at all and only 24 ROPs, which severely limits its pixel output. Its pixel rate is 47.52 GPixel/s versus 201.0 GPixel/s for the RTX 5070 SUPER. The H20 NVL16 wins in memory-bound compute and large-model inference, where 96 GB of HBM3 and 4.03 TB/s of bandwidth dwarf the RTX 5070 SUPER's 18 GB and 672.0 GB/s.

Architecture Differences

The RTX 5070 SUPER uses the GB205 chip built on Blackwell 2.0 architecture. It is manufactured by TSMC on a 5 nm process. The chip contains 31,100 million transistors on a 263 mm² die, giving a transistor density of 118.3M per mm². The H20 NVL16 uses the GH100 chip on Hopper architecture, also TSMC 5 nm, but with 80,000 million transistors on an 814 mm² die, for a density of 98.3M per mm². The H20 NVL16 die is more than three times larger physically and holds over twice as many transistors.

Memory architectures diverge sharply. The RTX 5070 SUPER uses GDDR7 with a 192-bit bus, while the H20 NVL16 uses HBM3 with a 6144-bit bus. The H20 NVL16's bandwidth advantage is roughly 6x, at 4.03 TB/s versus 672.0 GB/s. Clock speeds run in opposite directions: the RTX 5070 SUPER boosts to 2512 MHz, the H20 NVL16 to 1980 MHz. The H20 NVL16 compensates with more shading units, 9984 versus 6400, and more tensor cores, 312 versus 200.

The H20 NVL16 is a server module with SXM form factor, no display outputs, and a 400 W TDP. The RTX 5070 SUPER is a dual-slot card, 245 mm long, with a 16-pin power connector and a 275 W TDP. The H20 NVL16 lists an 800 W suggested PSU, while the RTX 5070 SUPER does not list one. The H20 NVL16 supports PCIe 5.0 x16, same as the RTX 5070 SUPER, but it has no graphics API support at all, marking it as a pure compute accelerator.

Head-to-Head Benchmarks

The database contains no direct head-to-head benchmark comparisons between these two products. The only benchmark score recorded for either part is the RTX 5070 SUPER's 3DMark Steel Nomad DX12 result of 2690. The H20 NVL16 has no benchmark entries and an average benchmark score of zero. The wins count is zero for both in the head-to-head table.

The absence of a shared benchmark suite is itself informative. The RTX 5070 SUPER can run DirectX 12 Ultimate and Vulkan 1.4 workloads; the H20 NVL16 cannot run either. Any graphics test that relies on rasterization or ray tracing would fail on the H20 NVL16 due to missing API support. Conversely, the H20 NVL16's FP16 throughput of 79.07 TFLOPS is more than double the RTX 5070 SUPER's 32.15 TFLOPS, indicating a clear advantage in mixed-precision compute tasks, though no recorded benchmark confirms this directly.

Comparing memory capacity shows a 5.3x gap: 96 GB versus 18 GB. The H20 NVL16's 4.03 TB/s bandwidth is 6x the RTX 5070 SUPER's 672.0 GB/s. The RTX 5070 SUPER's texture rate of 502.4 GTexel/s is lower than the H20 NVL16's 617.8 GTexel/s, but the RTX 5070 SUPER's pixel rate of 201.0 GPixel/s is over 4x the H20 NVL16's 47.52 GPixel/s. The RTX 5070 SUPER also has 80 ROPs versus 24, a 3.3x difference that directly impacts fill-rate-bound workloads.

The Verdict

The data defines the choice clearly. The RTX 5070 SUPER is the only option for any workload that ends in a display output. It has HDMI and DisplayPort connections, full DirectX 12 Ultimate and Vulkan 1.4 support, and 50 RT cores for ray tracing. Its 3DMark Steel Nomad score of 2690, sitting in the 18th percentile, indicates a capable mainstream graphics card for gaming and client rendering. The 18 GB GDDR7 frame buffer is substantial for a card of this class.

The H20 NVL16 is for server-side compute with no graphics component. It offers 96 GB of HBM3, 4.03 TB/s of bandwidth, and 39.54 TFLOPS of FP32 compute, with FP16 doubling to 79.07 TFLOPS. The 400 W TDP and SXM module form factor point to datacenter deployment, not desktop use. Its 9984 shading units and 312 tensor cores provide raw parallel throughput that the RTX 5070 SUPER cannot match, but without display outputs or graphics APIs, it cannot be used for conventional GPU workloads.

For a builder assembling a gaming or workstation PC, the RTX 5070 SUPER is the only viable choice. For a datacenter operator running inference or training workloads that fit within 96 GB of HBM3, the H20 NVL16 offers memory capacity and bandwidth no client card in this comparison can approach. The two parts do not compete; they occupy separate segments with no overlap in measured benchmark coverage.

FAQ

Q: Does the NVIDIA H20 NVL16 support DirectX or Vulkan?

A: No. The H20 NVL16 lists DirectX as N/A, OpenGL as N/A, and Vulkan as N/A.

Q: Which card has more memory bandwidth?

A: The H20 NVL16 has 4.03 TB/s from HBM3 memory on a 6144-bit bus. The RTX 5070 SUPER has 672.0 GB/s from GDDR7 on a 192-bit bus.

Q: What is the RTX 5070 SUPER's only recorded benchmark score?

A: The database lists a 3DMark Steel Nomad DX12 score of 2690, placing it in the 18th percentile of all GPUs.

Q: Can the H20 NVL16 output video to a display?

A: No. It has no display outputs, while the RTX 5070 SUPER has 1x HDMI 2.1b and 3x DisplayPort 2.1b.

Q: Which part has more FP16 compute throughput?

A: The H20 NVL16 delivers 79.07 TFLOPS at FP16 with a 2:1 ratio. The RTX 5070 SUPER delivers 32.15 TFLOPS at FP16 with a 1:1 ratio.

Q: What are the transistor counts for these two chips?

A: The RTX 5070 SUPER's GB205 chip has 31,100 million transistors. The H20 NVL16's GH100 chip has 80,000 million transistors.

Specification Differences

| Specification | NVIDIA GeForce RTX 5070 SUPER | NVIDIA H20 NVL16 |

|---|---|---|

| Architecture | Blackwell 2.0 | Hopper |

| Process Node | 5 nm | 5 nm |

| Transistors | 31,100 million | 80,000 million |

| Die Size | 263 mm² | 814 mm² |

| Transistor Density | 118.3M / mm² | 98.3M / mm² |

| Base Clock | 2325 MHz | 1830 MHz |

| Boost Clock | 2512 MHz | 1980 MHz |

| Memory Size | 18 GB | 96 GB |

| Memory Type | GDDR7 | HBM3 |

| Memory Bus Width | 192 bit | 6144 bit |

| Memory Bandwidth | 672.0 GB/s | 4.03 TB/s |

| Memory Clock | 1750 MHz 28 Gbps effective | 1313 MHz 5.3 Gbps effective |

| Shading Units | 6400 | 9984 |

| TMUs | 200 | 312 |

| ROPs | 80 | 24 |

| RT Cores | 50 | Not specified |

| Tensor Cores | 200 | 312 |

| Pixel Rate | 201.0 GPixel/s | 47.52 GPixel/s |

| Texture Rate | 502.4 GTexel/s | 617.8 GTexel/s |

| FP32 Performance | 32.15 TFLOPS | 39.54 TFLOPS |

| FP16 Performance | 32.15 TFLOPS (1:1) | 79.07 TFLOPS (2:1) |

| TDP | 275 W | 400 W |

| Slot Width | Dual-slot | SXM Module |

| Power Connectors | 1x 16-pin | Not specified |

| Suggested PSU | Not specified | 800 W |

| Bus Interface | PCIe 5.0 x16 | PCIe 5.0 x16 |

| Display Outputs | 1x HDMI 2.1b 3x DisplayPort 2.1b | No outputs |

| DirectX Support | 12 Ultimate (12_2) | N/A |

| OpenGL Support | 4.6 | N/A |

| Vulkan Support | 1.4 | N/A |

| Dimensions | 245 mm 9.6 inches | Not specified |

| Production Status | Active | Active |

| Release Date | 2025-12-31 | 2025-09-01 |

| Predecessor | Not specified | Server Ada |

| Successor | Not specified | Server Blackwell |

| Percentile vs All GPUs | 18 | 50 |

| Average Benchmark Score | 2690 | 0 |

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 5070 SUPER
H20 NVL16
Core Specs
Shading Units
6,400
9,984 +56.0%
Shaders
6,400
9,984 +56.0%
TMUs
200
312 +56.0%
ROPs
80
24 -70.0%
SM Count
—
78
Clocks
Base Clock
2325 MHz
1830 MHz
Boost Clock
2512 MHz
1980 MHz
Memory Clock
1750 MHz 28 Gbps effective
1313 MHz 5.3 Gbps effective
Memory
Memory Size
18 GB
96 GB
VRAM (MB)
18,432
98,304 +433.3%
Memory Type
GDDR7
HBM3
Memory Bus
192 bit
6144 bit
Bandwidth
672.0 GB/s
4.03 TB/s
Cache
L1 Cache
128 KB (per SM)
256 KB (per SM)
L2 Cache
48 MB
60 MB
Performance
Pixel Rate
201.0 GPixel/s
47.52 GPixel/s
Texture Rate
502.4 GTexel/s
617.8 GTexel/s
FP32 (TFLOPS)
32.15 TFLOPS
39.54 TFLOPS
FP64 (TFLOPS)
502.4 GFLOPS (1:64)
19.77 TFLOPS (1:2)
FP16 (TFLOPS)
32.15 TFLOPS (1:1)
79.07 TFLOPS (2:1)
AI/RT
RT Cores
50
—
Tensor Cores
200
312 +56.0%
Power
TDP
275 W
400 W
TDP (W)
275
400 +45.5%
Suggested PSU
—
800 W
Power Connectors
1x 16-pin
—
Architecture
Architecture
Blackwell 2.0
Hopper
GPU Name
GB205
GH100
Generation
GeForce 50
Server Hopper (Hxx)
Process Size
5 nm
5 nm
Transistors
31,100 million
80,000 million
Die Size
263 mm²
814 mm²
Foundry
TSMC
TSMC
Density
118.3M / mm²
98.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
—
OpenGL
4.6
—
Vulkan
1.4
—
OpenCL
3.0
3.0
CUDA
—
9.0
Shader Model
6.8
—
Physical
Slot Width
Dual-slot
SXM Module
Length
245 mm 9.6 inches
—
Height
115 mm 4.5 inches
—
Outputs
1x HDMI 2.1b 3x DisplayPort 2.1b
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 5.0 x16
Other
Production
Active
Active
Predecessor
—
Server Ada
Successor
—
Server Blackwell
View GeForce RTX 5070 SUPER Details View H20 NVL16 Details