NVIDIA GeForce RTX 4070 SUPER vs NVIDIA H20 NVL16 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 4070 SUPER

CORE STATE AD104
VRAM 12 GB
CLOCK SPEED 2475 MHz
TDP 220 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2024
VS
NVIDIA
GEFORCE

H20 NVL16

CORE STATE GH100
VRAM 96 GB
CLOCK SPEED 1980 MHz
TDP 400 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2025

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
4,627
N/A
geekbench_opencl
172,795
N/A
geekbench_vulkan
205,624
N/A
passmark_directx_10
167
N/A
passmark_directx_11
273
N/A
passmark_directx_12
110
N/A
passmark_directx_9
344
N/A
passmark_g2d
1,184
N/A
passmark_g3d
29,995
N/A
passmark_gpu_compute
17,108
N/A

Analysis: NVIDIA GeForce RTX 4070 SUPER vs NVIDIA H20 NVL16

Head-to-Head Benchmarks

The recorded database contains benchmark scores for the NVIDIA GeForce RTX 4070 SUPER, while the NVIDIA H20 NVL16 has no benchmark entries in the database. The RTX 4070 SUPER’s average benchmark score across all recorded tests is 43,223, placing it in the 83rd percentile of all GPUs in the database. The H20 NVL16 has an average benchmark score of 0, with a 50th percentile ranking, which reflects the absence of any recorded test results for that unit.

The RTX 4070 SUPER delivers a 3DMark Steel Nomad DX12 score of 4,627. In Geekbench OpenCL, it records 172,795 points, and in Geekbench Vulkan, it reaches 205,624 points. Passmark results show a DirectX 10 score of 167, DirectX 11 score of 273, DirectX 12 score of 110, and DirectX 9 score of 344. The Passmark G2D score is 1,184, while the G3D score is 29,995. For GPU compute, the Passmark score is 17,108.

Since the H20 NVL16 has no benchmark scores, no direct head-to-head delta can be calculated from the database. The RTX 4070 SUPER’s nearest rivals in the database, based on average score, include the NVIDIA Quadro M6000 24 GB at 43,262 (a delta of -0.1% relative to the 4070 SUPER), the NVIDIA GeForce RTX 5050 Mobile at 43,268 (-0.1%), the NVIDIA Quadro M6000 at 43,301 (-0.2%), and the NVIDIA GeForce RTX 4090 Mobile at 43,667 (-1%). These deltas indicate that the 4070 SUPER sits within 1% of these competing products in aggregate performance, with the 4090 Mobile being the fastest of the group by a single percentage point.

The data shows that the RTX 4070 SUPER’s average score of 43,223 is lower than all four nearest rivals, but the differences are small: 39 points behind the Quadro M6000 24 GB, 45 points behind the RTX 5050 Mobile, 78 points behind the Quadro M6000, and 444 points behind the RTX 4090 Mobile. In percentage terms, the 4070 SUPER trails by 0.1%, 0.1%, 0.2%, and 1% respectively. This clustering suggests that the 4070 SUPER performs within the same tier as these older or mobile-class parts in aggregate database metrics.

For the H20 NVL16, the absence of benchmark data means the database cannot substantiate any performance comparison. Its 50th percentile ranking is a default value assigned to unmeasured hardware, not an indication of measured capability. The RTX 4070 SUPER, by contrast, has a full suite of ten recorded benchmark results, giving the database a robust basis for its 83rd percentile placement.

Architecture Differences

The RTX 4070 SUPER uses the AD104 chip built on Ada Lovelace architecture, fabricated by TSMC on a 5 nm process. The die contains 35,800 million transistors across 294 mm², yielding a transistor density of 121.8M per mm². The H20 NVL16 uses the GH100 chip built on Hopper architecture, also fabricated by TSMC on a 5 nm process. The GH100 die holds 80,000 million transistors across 814 mm², with a transistor density of 98.3M per mm².

The H20 NVL16 has more shading units: 9,984 compared to 7,168 on the RTX 4070 SUPER. Texture mapping units also favor the H20 NVL16, with 312 versus 224. Raster operation units differ significantly: the RTX 4070 SUPER has 80 ROPs, while the H20 NVL16 has only 24 ROPs. The RTX 4070 SUPER includes 56 ray tracing cores and 224 tensor cores. The H20 NVL16 has 312 tensor cores but its ray tracing core count is not recorded in the database.

Clock speeds show the RTX 4070 SUPER running at a base of 1980 MHz and a boost of 2475 MHz. The H20 NVL16 runs at a base of 1830 MHz and a boost of 1980 MHz. Memory clocks are identical at 1313 MHz, but the effective data rates differ: the RTX 4070 SUPER reaches 21 Gbps effective, while the H20 NVL16 reaches 5.3 Gbps effective. The memory type and configuration are fundamentally different. The RTX 4070 SUPER uses 12 GB of GDDR6X on a 192-bit bus, delivering 504.2 GB/s of bandwidth. The H20 NVL16 uses 96 GB of HBM3 on a 6144-bit bus, delivering 4.03 TB/s of bandwidth.

Pixel rates show the RTX 4070 SUPER at 198.0 GPixel/s versus 47.52 GPixel/s for the H20 NVL16, a consequence of the H20’s lower ROP count. Texture rates are closer: 554.4 GTexel/s for the RTX 4070 SUPER and 617.8 GTexel/s for the H20 NVL16. Floating-point performance differs by workload. FP32 rates are 35.48 TFLOPS for the RTX 4070 SUPER and 39.54 TFLOPS for the H20 NVL16. FP16 rates are 35.48 TFLOPS (1:1 ratio) for the RTX 4070 SUPER and 79.07 TFLOPS (2:1 ratio) for the H20 NVL16, indicating that the H20 NVL16 is designed for double-rate half-precision compute.

The RTX 4070 SUPER supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The H20 NVL16 has no recorded graphics API support, listing DirectX, OpenGL, and Vulkan as N/A. The RTX 4070 SUPER includes display outputs (1x HDMI 2.1 and 3x DisplayPort 1.4a), while the H20 NVL16 has no display outputs. The bus interfaces also differ: PCIe 4.0 x16 for the RTX 4070 SUPER and PCIe 5.0 x16 for the H20 NVL16.

Power requirements are substantially different. The RTX 4070 SUPER has a TDP of 220 W with a suggested PSU of 550 W and a single 16-pin power connector. The H20 NVL16 has a TDP of 400 W with a suggested PSU of 800 W and no recorded power connectors, as it is an SXM module form factor. The RTX 4070 SUPER is a dual-slot card measuring 267 mm in length, 112 mm in height, and 42 mm in width. The H20 NVL16 is an SXM module with no recorded dimensions.

FAQ

Q: How does the memory bandwidth of the H20 NVL16 compare to the RTX 4070 SUPER?

A: The H20 NVL16 has a memory bandwidth of 4.03 TB/s, while the RTX 4070 SUPER has 504.2 GB/s. The H20 NVL16’s bandwidth is approximately eight times higher, due to its 6144-bit HBM3 bus versus the 192-bit GDDR6X bus on the RTX 4070 SUPER.

Q: Which GPU has higher FP32 compute performance?

A: The H20 NVL16 has a higher FP32 rate at 39.54 TFLOPS, compared to 35.48 TFLOPS for the RTX 4070 SUPER. The difference is about 11% in favor of the H20 NVL16.

Q: Does the RTX 4070 SUPER support ray tracing hardware?

A: Yes, the RTX 4070 SUPER is recorded with 56 ray tracing cores. The H20 NVL16 has no ray tracing core count listed in the database, and its graphics API support is marked as N/A, suggesting it is not designed for ray-traced workloads.

Q: What is the production status of each GPU?

A: The RTX 4070 SUPER is listed as end-of-life, while the H20 NVL16 is listed as active. The RTX 4070 SUPER was released on 2024-01-16, and the H20 NVL16 on 2025-09-01.

Q: How do the transistor counts and die sizes compare?

A: The H20 NVL16 has 80,000 million transistors on an 814 mm² die, while the RTX 4070 SUPER has 35,800 million transistors on a 294 mm² die. The H20 NVL16 has more than double the transistor count and nearly three times the die area, but the RTX 4070 SUPER has a higher transistor density at 121.8M per mm² versus 98.3M per mm².

Q: Can the H20 NVL16 be used in a standard desktop PC?

A: The database lists the H20 NVL16 as an SXM module with no display outputs, no power connectors, and no dimensions. The RTX 4070 SUPER is a dual-slot PCIe card with display outputs and a 16-pin connector. The H20 NVL16 is designed for server integration, not standard desktop use.

Specification Differences

The two GPUs differ in nearly every recorded specification. The chip differs: AD104 for the RTX 4070 SUPER versus GH100 for the H20 NVL16. Architecture differs: Ada Lovelace versus Hopper. The generation field shows GeForce 40 versus Server Hopper (Hxx). Transistor counts are 35,800 million versus 80,000 million. Die sizes are 294 mm² versus 814 mm². Transistor densities are 121.8M per mm² versus 98.3M per mm².

Clock speeds differ: base clocks are 1980 MHz versus 1830 MHz, boost clocks are 2475 MHz versus 1980 MHz. Memory effective rates are 21 Gbps versus 5.3 Gbps. Memory sizes are 12 GB versus 96 GB. Memory types are GDDR6X versus HBM3. Bus widths are 192 bit versus 6144 bit. Bandwidths are 504.2 GB/s versus 4.03 TB/s.

Shading units are 7,168 versus 9,984. TMUs are 224 versus 312. ROPs are 80 versus 24. RT cores are 56 versus null. Tensor cores are 224 versus 312. Pixel rates are 198.0 GPixel/s versus 47.52 GPixel/s. Texture rates are 554.4 GTexel/s versus 617.8 GTexel/s. FP32 is 35.48 TFLOPS versus 39.54 TFLOPS. FP16 is 35.48 TFLOPS versus 79.07 TFLOPS, with ratios of 1:1 versus 2:1.

TDP is 220 W versus 400 W. Slot width is dual-slot versus SXM module. Power connectors are 1x 16-pin versus null. Suggested PSU is 550 W versus 800 W. Bus interface is PCIe 4.0 x16 versus PCIe 5.0 x16. Display outputs are 1x HDMI 2.1, 3x DisplayPort 1.4a versus no outputs. DirectX support is 12 Ultimate (12_2) versus N/A. OpenGL support is 4.6 versus N/A. Vulkan support is 1.4 versus N/A.

Dimensions are recorded for the RTX 4070 SUPER only: 267 mm length, 112 mm height, 42 mm width. The H20 NVL16 has null dimensions. Production status is end-of-life versus active. Release dates are 2024-01-16 versus 2025-09-01. Predecessors are GeForce 30 versus Server Ada. Successors are GeForce 50 versus Server Blackwell. The RTX 4070 SUPER has a launch MSRP of 599 USD; the H20 NVL16 has no recorded launch MSRP.

The Verdict

The database presents two GPUs with fundamentally different design goals. The RTX 4070 SUPER is a client-side graphics card with full display output, DirectX 12 Ultimate support, and ray tracing hardware. Its benchmark suite shows it performing in the 83rd percentile of all GPUs, with an average score of 43,223. The H20 NVL16 is a server module with no display outputs, no graphics API support, and no recorded benchmarks. Its 50th percentile ranking is a placeholder, not a measured result.

The data indicates that the RTX 4070 SUPER is the appropriate choice for graphics workloads that require rasterization, ray tracing, and API compatibility. Its 56 RT cores and 224 tensor cores support real-time rendering and AI-accelerated graphics features. The H20 NVL16, with 312 tensor cores, 79.07 TFLOPS FP16 performance, and 4.03 TB/s memory bandwidth, is positioned for compute-heavy tasks, specifically those that benefit from half-precision throughput and massive memory capacity.

The RTX 4070 SUPER’s higher ROP count (80 versus 24) and pixel rate (198.0 GPixel/s versus 47.52 GPixel/s) confirm its rasterization focus. The H20 NVL16’s higher texture rate (617.8 GTexel/s versus 554.4 GTexel/s) and FP16 rate indicate a compute orientation. The H20 NVL16 also carries 96 GB of HBM3 memory, which is eight times the capacity of the RTX 4070 SUPER’s 12 GB, and offers bandwidth that is roughly eight times higher.

Power draw and form factor further separate the two. The RTX 4070 SUPER fits a standard dual-slot PCIe slot with a 220 W TDP and a 550 W suggested PSU. The H20 NVL16 is an SXM module with a 400 W TDP and an 800 W suggested PSU, requiring a server chassis. The RTX 4070 SUPER is end-of-life, while the H20 NVL16 is active, reflecting their respective product lifecycles.

For users with graphics-intensive applications, the RTX 4070 SUPER is the only option with recorded performance data. For server deployments focused on compute, the H20 NVL16’s specifications indicate a different role, but the database contains no benchmark results to quantify its performance. The verdict follows the data: the RTX 4070 SUPER is a measured, capable graphics card, while the H20 NVL16 is an unmeasured server accelerator whose strengths lie in memory and compute specifications rather than rendered benchmarks.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 4070 SUPER
H20 NVL16
Core Specs
Shading Units
7,168
9,984 +39.3%
Shaders
7,168
9,984 +39.3%
TMUs
224
312 +39.3%
ROPs
80
24 -70.0%
SM Count
56
78 +39.3%
Clocks
Base Clock
1980 MHz
1830 MHz
Boost Clock
2475 MHz
1980 MHz
Memory Clock
1313 MHz 21 Gbps effective
1313 MHz 5.3 Gbps effective
Memory
Memory Size
12 GB
96 GB
VRAM (MB)
12,288
98,304 +700.0%
Memory Type
GDDR6X
HBM3
Memory Bus
192 bit
6144 bit
Bandwidth
504.2 GB/s
4.03 TB/s
Cache
L1 Cache
128 KB (per SM)
256 KB (per SM)
L2 Cache
48 MB
60 MB
Performance
Pixel Rate
198.0 GPixel/s
47.52 GPixel/s
Texture Rate
554.4 GTexel/s
617.8 GTexel/s
FP32 (TFLOPS)
35.48 TFLOPS
39.54 TFLOPS
FP64 (TFLOPS)
554.4 GFLOPS (1:64)
19.77 TFLOPS (1:2)
FP16 (TFLOPS)
35.48 TFLOPS (1:1)
79.07 TFLOPS (2:1)
AI/RT
RT Cores
56
—
Tensor Cores
224
312 +39.3%
Power
TDP
220 W
400 W
TDP (W)
220
400 +81.8%
Suggested PSU
550 W
800 W
Power Connectors
1x 16-pin
—
Architecture
Architecture
Ada Lovelace
Hopper
GPU Name
AD104
GH100
Generation
GeForce 40
Server Hopper (Hxx)
Process Size
5 nm
5 nm
Transistors
35,800 million
80,000 million
Die Size
294 mm²
814 mm²
Foundry
TSMC
TSMC
Density
121.8M / mm²
98.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
—
OpenGL
4.6
—
Vulkan
1.4
—
OpenCL
3.0
3.0
CUDA
8.9
9.0
Shader Model
6.9
—
Physical
Slot Width
Dual-slot
SXM Module
Length
267 mm 10.5 inches
—
Height
112 mm 4.4 inches
—
Outputs
1x HDMI 2.13x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 5.0 x16
Other
Launch Price
599 USD
—
Production
End-of-life
Active
Predecessor
GeForce 30
Server Ada
Successor
GeForce 50
Server Blackwell
View GeForce RTX 4070 SUPER Details View H20 NVL16 Details