NVIDIA GeForce RTX 4070 SUPER vs NVIDIA H20 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 4070 SUPER

CORE STATE AD104
VRAM 12 GB
CLOCK SPEED 2475 MHz
TDP 220 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2024
VS
NVIDIA
GEFORCE

H20

CORE STATE GH100
VRAM 96 GB
CLOCK SPEED 1980 MHz
TDP 500 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2024

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
4,627
N/A
geekbench_opencl
172,795
N/A
geekbench_vulkan
205,624
N/A
passmark_directx_10
167
N/A
passmark_directx_11
273
N/A
passmark_directx_12
110
N/A
passmark_directx_9
344
N/A
passmark_g2d
1,184
N/A
passmark_g3d
29,995
N/A
passmark_gpu_compute
17,108
N/A

Analysis: NVIDIA GeForce RTX 4070 SUPER vs NVIDIA H20

Where Each One Wins

The recorded data presents a fundamentally asymmetric comparison. The NVIDIA GeForce RTX 4070 SUPER has ten benchmark entries across DirectX, OpenCL, Vulkan, and compute workloads, while the NVIDIA H20 has no recorded benchmark scores in the database. Consequently, the RTX 4070 SUPER holds every measured win in the dataset. Its average benchmark score stands at 43,223, placing it in the 83rd percentile of all GPUs. The H20, by contrast, has an average score of 0 and sits at the 50th percentile, a position that reflects its absence of measurement data rather than a performance verdict.

The RTX 4070 SUPER's benchmark profile shows particular strength in synthetic graphics tests. Its highest single score is 205,624 in Geekbench Vulkan, followed by 172,795 in Geekbench OpenCL. The Passmark G3D suite returns 29,995, while the GPU compute test yields 17,108. DirectX-specific results are lower in absolute terms: Passmark DirectX 9 scores 344, DirectX 11 scores 273, DirectX 10 scores 167, and DirectX 12 scores 110. The 2D graphics score is 1,184. The 3DMark Steel Nomad DX12 test produces 4,627 points.

For the H20, the absence of benchmark data means the database cannot assign a single performance win. Its 50th percentile ranking alongside a zero average score indicates that no workload measurements have been recorded. The use-case split therefore defaults entirely to the RTX 4070 SUPER for all measurable graphics and compute tasks. The H20's role in the comparison must be assessed through its architecture and memory subsystem rather than through direct performance numbers.

Architecture Differences

The two GPUs come from distinct architectural lineages. The RTX 4070 SUPER uses the AD104 chip built on Ada Lovelace architecture, fabricated on a 5 nm process by TSMC. The H20 uses the GH100 chip built on Hopper architecture, also on a 5 nm TSMC process. Both share the same process node, but their transistor counts diverge sharply. The AD104 packs 35,800 million transistors on a 294 mm² die, yielding a density of 121.8 million transistors per square millimeter. The GH100 houses 80,000 million transistors on an 814 mm² die, with a density of 98.3 million per square millimeter.

The H20's die is more than twice the physical size of the RTX 4070 SUPER's, and it carries over twice the transistor count. The transistor density is lower on the H20, indicating a less compact layout despite the larger scale. The RTX 4070 SUPER belongs to the GeForce 40-series with a release date of January 2024, while the H20 belongs to the Server Hopper generation with a release date later that same month. The H20 lists its predecessor as Server Ada and successor as Server Blackwell; the RTX 4070 SUPER lists GeForce 30 as predecessor and GeForce 50 as successor.

Clock behavior differs substantially. The RTX 4070 SUPER runs at a base clock of 1980 MHz and boosts to 2475 MHz. The H20 runs at a lower base clock of 1830 MHz and a lower boost clock of 1980 MHz. Memory clocks show a different pattern: both list a memory clock of 1313 MHz, but the effective data rate differs. The RTX 4070 SUPER operates at 21 Gbps effective, while the H20 operates at 5.3 Gbps effective. The H20 compensates with a dramatically wider memory bus and a different memory type.

The H20 uses 96 GB of HBM3 memory on a 6144-bit bus, producing 4.03 TB/s of bandwidth. The RTX 4070 SUPER uses 12 GB of GDDR6X on a 192-bit bus, producing 504.2 GB/s. The H20's bandwidth advantage is roughly eightfold in raw throughput. The H20's memory size is eight times larger as well.

The RTX 4070 SUPER has 7168 shading units, 224 texture mapping units, and 80 ROPs. The H20 has 9984 shading units, 312 TMUs, and only 24 ROPs. The H20's ROP count is a fraction of the RTX 4070 SUPER's, which limits its pixel throughput: 47.52 GPixel/s versus 198.0 GPixel/s. Texture rate favors the H20 at 617.8 GTexel/s versus 554.4 GTexel/s. The H20 has 312 tensor cores and no listed ray tracing cores; the RTX 4070 SUPER has 224 tensor cores and 56 ray tracing cores.

Floating-point performance splits by precision. In FP32, the H20 leads with 39.54 TFLOPS versus 35.48 TFLOPS on the RTX 4070 SUPER. In FP16, the H20 reaches 79.07 TFLOPS with a 2:1 ratio, while the RTX 4070 SUPER delivers 35.48 TFLOPS at a 1:1 ratio. The H20's FP16 output is more than double that of the RTX 4070 SUPER.

Power and interface requirements diverge as well. The RTX 4070 SUPER has a TDP of 220 W, uses a single 16-pin connector, and suggests a 550 W power supply. The H20 has a TDP of 500 W, lists no power connectors, and suggests a 900 W power supply. The RTX 4070 SUPER is a dual-slot card with a 267 mm length, 112 mm height, and 42 mm width. The H20 is an SXM module with no recorded physical dimensions.

The Verdict

Based strictly on recorded data, the RTX 4070 SUPER is the only GPU in this comparison with measurable performance. It delivers a full suite of benchmark scores across multiple APIs, ranking in the 83rd percentile of all GPUs. Its nearest rivals in the database include the NVIDIA Quadro M6000 24 GB at a 0.1% lower average score, the NVIDIA GeForce RTX 5050 Mobile at 0.1% lower, the NVIDIA Quadro M6000 at 0.2% lower, and the NVIDIA GeForce RTX 4090 Mobile at 1% lower. The RTX 4070 SUPER's average score of 43,223 sits slightly above each of these.

The H20 has no benchmark scores, no average score, and no nearest rivals. Its percentile ranking of 50th is the default position for an unmeasured GPU. The H20's architecture provides different capabilities: a much larger memory pool, higher bandwidth, more tensor cores, and higher FP16 throughput. But the database contains no test results to confirm how those capabilities translate into workload performance.

The RTX 4070 SUPER is marked end-of-life with a launch MSRP of 599 USD. The H20 is marked active with no launch MSRP. The RTX 4070 SUPER supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, and has display outputs including HDMI and DisplayPort. The H20 lists no API support and no display outputs, reflecting its server-oriented design.

For any workload covered by the recorded benchmarks, the RTX 4070 SUPER is the only option with data. For workloads that rely on massive memory capacity, high bandwidth, or dense FP16 compute, the H20's specifications suggest capability, but no measurement exists to quantify it.

FAQ

Q: Which GPU has a higher average benchmark score?

A: The NVIDIA GeForce RTX 4070 SUPER has an average benchmark score of 43,223. The NVIDIA H20 has an average benchmark score of 0.

Q: How does the RTX 4070 SUPER compare to its nearest rivals?

A: The RTX 4070 SUPER scores 0.1% higher than the Quadro M6000 24 GB and the RTX 5050 Mobile, 0.2% higher than the Quadro M6000, and 1% higher than the RTX 4090 Mobile in average benchmark score.

Q: What memory configurations do the two GPUs use?

A: The RTX 4070 SUPER uses 12 GB of GDDR6X on a 192-bit bus with 504.2 GB/s bandwidth. The H20 uses 96 GB of HBM3 on a 6144-bit bus with 4.03 TB/s bandwidth.

Q: Which GPU has higher FP16 compute throughput?

A: The H20 delivers 79.07 TFLOPS FP16 with a 2:1 ratio. The RTX 4070 SUPER delivers 35.48 TFLOPS FP16 at a 1:1 ratio.

Q: Does the H20 have any recorded benchmark results?

A: No, the H20 has an empty benchmark list and a 50th percentile ranking with a zero average score.

Q: What is the production status of each GPU?

A: The RTX 4070 SUPER is marked end-of-life. The H20 is marked active.

Head-to-Head Benchmarks

The head-to-head benchmark list is empty in the recorded data, so no direct comparison scores exist between the two GPUs. The RTX 4070 SUPER's individual benchmark scores provide the only numerical performance data in this comparison. Its Geekbench Vulkan result of 205,624 is the highest single score recorded for either GPU. The Geekbench OpenCL score of 172,795 follows. The Passmark G3D score is 29,995, and the GPU compute score is 17,108.

The RTX 4070 SUPER's DirectX results show a clear hierarchy by API generation. DirectX 9 scores 344, DirectX 11 scores 273, DirectX 12 scores 110, and DirectX 10 scores 167. The 3DMark Steel Nomad DX12 test returns 4,627. The Passmark 2D score is 1,184.

The H20 has no benchmark entries, so no wins can be attributed to it. The RTX 4070 SUPER wins all recorded workload categories by default. In raw specification terms, the H20 leads in shading units (9984 versus 7168), texture units (312 versus 224), tensor cores (312 versus 224), memory size, memory bandwidth, texture rate, FP32 throughput, and FP16 throughput. The RTX 4070 SUPER leads in ROPs (80 versus 24), pixel rate (198.0 GPixel/s versus 47.52 GPixel/s), base clock, boost clock, and transistor density.

Specification Differences

The two GPUs differ across nearly every recorded specification field. The RTX 4070 SUPER uses the AD104 chip on Ada Lovelace architecture; the H20 uses the GH100 chip on Hopper architecture. Both are fabricated on a 5 nm TSMC process. The RTX 4070 SUPER has 35,800 million transistors on a 294 mm² die; the H20 has 80,000 million transistors on an 814 mm² die. Transistor density is 121.8M per mm² for the RTX 4070 SUPER and 98.3M per mm² for the H20.

The RTX 4070 SUPER has a base clock of 1980 MHz and a boost clock of 2475 MHz. The H20 has a base clock of 1830 MHz and a boost clock of 1980 MHz. Memory clock is 1313 MHz for both, but effective rates are 21 Gbps for the RTX 4070 SUPER and 5.3 Gbps for the H20. Memory size is 12 GB GDDR6X versus 96 GB HBM3. Bus width is 192 bit versus 6144 bit. Bandwidth is 504.2 GB/s versus 4.03 TB/s.

Shading units number 7168 versus 9984. TMUs number 224 versus 312. ROPs number 80 versus 24. The RTX 4070 SUPER has 56 ray tracing cores and 224 tensor cores; the H20 has no listed ray tracing cores and 312 tensor cores. Pixel rate is 198.0 GPixel/s versus 47.52 GPixel/s. Texture rate is 554.4 GTexel/s versus 617.8 GTexel/s. FP32 is 35.48 TFLOPS versus 39.54 TFLOPS. FP16 is 35.48 TFLOPS (1:1) versus 79.07 TFLOPS (2:1).

TDP is 220 W versus 500 W. The RTX 4070 SUPER uses a 1x 16-pin power connector and suggests a 550 W power supply; the H20 lists no power connector and suggests a 900 W power supply. The RTX 4070 SUPER is dual-slot with dimensions of 267 mm length, 112 mm height, and 42 mm width; the H20 is an SXM module with no dimensions recorded. The RTX 4070 SUPER uses PCIe 4.0 x16, while the H20 uses PCIe 5.0 x16. Display outputs are 1x HDMI 2.1 and 3x DisplayPort 1.4a on the RTX 4070 SUPER, with no outputs on the H20. API support exists only on the RTX 4070 SUPER, which lists DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4; the H20 lists N/A for all three.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 4070 SUPER
H20
Core Specs
Shading Units
7,168
9,984 +39.3%
Shaders
7,168
9,984 +39.3%
TMUs
224
312 +39.3%
ROPs
80
24 -70.0%
SM Count
56
78 +39.3%
Clocks
Base Clock
1980 MHz
1830 MHz
Boost Clock
2475 MHz
1980 MHz
Memory Clock
1313 MHz 21 Gbps effective
1313 MHz 5.3 Gbps effective
Memory
Memory Size
12 GB
96 GB
VRAM (MB)
12,288
98,304 +700.0%
Memory Type
GDDR6X
HBM3
Memory Bus
192 bit
6144 bit
Bandwidth
504.2 GB/s
4.03 TB/s
Cache
L1 Cache
128 KB (per SM)
256 KB (per SM)
L2 Cache
48 MB
60 MB
Performance
Pixel Rate
198.0 GPixel/s
47.52 GPixel/s
Texture Rate
554.4 GTexel/s
617.8 GTexel/s
FP32 (TFLOPS)
35.48 TFLOPS
39.54 TFLOPS
FP64 (TFLOPS)
554.4 GFLOPS (1:64)
19.77 TFLOPS (1:2)
FP16 (TFLOPS)
35.48 TFLOPS (1:1)
79.07 TFLOPS (2:1)
AI/RT
RT Cores
56
Tensor Cores
224
312 +39.3%
Power
TDP
220 W
500 W
TDP (W)
220
500 +127.3%
Suggested PSU
550 W
900 W
Power Connectors
1x 16-pin
Architecture
Architecture
Ada Lovelace
Hopper
GPU Name
AD104
GH100
Generation
GeForce 40
Server Hopper (Hxx)
Process Size
5 nm
5 nm
Transistors
35,800 million
80,000 million
Die Size
294 mm²
814 mm²
Foundry
TSMC
TSMC
Density
121.8M / mm²
98.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
8.9
9.0
Shader Model
6.9
Physical
Slot Width
Dual-slot
SXM Module
Length
267 mm 10.5 inches
Height
112 mm 4.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 5.0 x16
Other
Launch Price
599 USD
Production
End-of-life
Active
Predecessor
GeForce 30
Server Ada
Successor
GeForce 50
Server Blackwell
View GeForce RTX 4070 SUPER Details View H20 Details