NVIDIA GeForce RTX 4070 SUPER vs NVIDIA GeForce RTX 4090 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 4070 SUPER

CORE STATE AD104
VRAM 12 GB
CLOCK SPEED 2475 MHz
TDP 220 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2024
VS
NVIDIA
GEFORCE

GeForce RTX 4090

CORE STATE AD102
VRAM 24 GB
CLOCK SPEED 2520 MHz
TDP 450 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
4,627
9,223
geekbench_opencl
172,795
255,416
geekbench_vulkan
205,624
271,631
passmark_directx_10
167
224
passmark_directx_11
273
326
passmark_directx_12
110
150
passmark_directx_9
344
397
passmark_g2d
1,184
1,299
passmark_g3d
29,995
38,194
passmark_gpu_compute
17,108
26,613

Analysis: NVIDIA GeForce RTX 4070 SUPER vs NVIDIA GeForce RTX 4090

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA GeForce RTX 4090 records an average benchmark score of 60,347, while the NVIDIA GeForce RTX 4070 SUPER records 43,223. That places the RTX 4090 roughly 39.6% higher in the database's aggregate measurements.

Q: How do the two cards compare in 3DMark Steel Nomad DX12?

A: The RTX 4090 scores 9,223 versus 4,627 for the RTX 4070 SUPER, a delta of 99.3%. This is the largest relative gap in the head-to-head set.

Q: Which card has more RT cores and tensor cores?

A: The RTX 4090 has 128 RT cores and 512 tensor cores. The RTX 4070 SUPER has 56 RT cores and 224 tensor cores. Both use the Ada Lovelace architecture.

Q: Do the two cards share the same memory type?

A: Yes, both use GDDR6X memory. However, the RTX 4090 has 24 GB on a 384-bit bus with 1.01 TB/s bandwidth, while the RTX 4070 SUPER has 12 GB on a 192-bit bus with 504.2 GB/s bandwidth.

Q: What is the power connector requirement for each card?

A: Both cards use a single 16-pin power connector. The RTX 4090 has a TDP of 450 W and a suggested PSU of 850 W, while the RTX 4070 SUPER has a TDP of 220 W and a suggested PSU of 550 W.

Q: Which card ranks higher in the database's percentile versus all GPUs?

A: The RTX 4090 sits in the 88th percentile, while the RTX 4070 SUPER sits in the 83rd percentile. Both belong to the GeForce 40-series and use the same Ada Lovelace architecture and TSMC 5 nm process.

Architecture Differences

Both cards are built on the Ada Lovelace architecture and fabricated by TSMC on a 5 nm process, but they are fundamentally different silicon. The RTX 4090 uses the AD102 chip with 76,300 million transistors on a 609 mm² die, giving a transistor density of 125.3M per mm². The RTX 4070 SUPER uses the AD104 chip with 35,800 million transistors on a 294 mm² die, a density of 121.8M per mm². The AD102 is more than twice the physical size and carries more than double the transistor count.

The compute resource disparity is large. The RTX 4090 has 16,384 shading units, 512 TMUs, and 176 ROPs. The RTX 4070 SUPER has 7,168 shading units, 224 TMUs, and 80 ROPs. That is a 2.29x advantage in shading units for the RTX 4090. The RTX 4090 also holds 128 RT cores versus 56, and 512 tensor cores versus 224.

Clock speeds are close. The RTX 4090 has a base clock of 2235 MHz and a boost clock of 2520 MHz. The RTX 4070 SUPER has a base of 1980 MHz and a boost of 2475 MHz. The RTX 4090 runs slightly higher in both states. Memory clocks are identical at 1313 MHz with 21 Gbps effective.

Memory architecture differs sharply. The RTX 4090's 384-bit bus delivers 1.01 TB/s of bandwidth, while the RTX 4070 SUPER's 192-bit bus delivers 504.2 GB/s. That is roughly a 2x bandwidth advantage for the RTX 4090, which matters for high-resolution texture streaming and compute workloads.

The FP32 throughput reflects the core count difference. The RTX 4090 reaches 82.58 TFLOPS, while the RTX 4070 SUPER reaches 35.48 TFLOPS. Both have a 1:1 FP16 ratio, so FP16 throughput matches FP32 on each card.

Physical design diverges as well. The RTX 4090 is a triple-slot card measuring 304 mm in length, 137 mm in height, and 61 mm in width. The RTX 4070 SUPER is dual-slot, measuring 267 mm long, 112 mm high, and 42 mm wide. The RTX 4090 is substantially larger in every dimension.

Both cards share the same API support set, including DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. Both use PCIe 4.0 x16 and provide the same display outputs, 1x HDMI 2.1 and 3x DisplayPort 1.4a. Both are end-of-life products in the database, both succeeded by the GeForce 50 generation.

The Verdict

The data points to the RTX 4090 as the stronger card in every recorded benchmark. It wins all 10 head-to-head tests in the database. The largest margins are in 3DMark Steel Nomad DX12 at 99.3% and PassMark GPU Compute at 55.6%. The smallest margin is PassMark G2D at 9.7%, a test that measures 2D performance rather than raw 3D throughput.

The RTX 4070 SUPER is the more constrained card in terms of memory, compute, and physical footprint. Its 12 GB frame buffer and 504.2 GB/s bandwidth are half of what the RTX 4090 offers. Its 220 W TDP and 550 W suggested PSU make it a lighter system integration burden, while the RTX 4090 requires 450 W and an 850 W PSU.

Who should pick the RTX 4090: anyone whose workload is dominated by the tests where its lead is largest, namely DX12 gaming, GPU compute, and OpenCL. The 99.3% advantage in 3DMark Steel Nomad DX12 and the 55.6% advantage in PassMark GPU Compute show a card that is not just faster but disproportionately faster in modern graphics and compute-heavy tasks. The 24 GB memory and 1.01 TB/s bandwidth also give it headroom that the 12 GB card cannot match.

Who should pick the RTX 4070 SUPER: the data does not show any benchmark where it wins. Its case rests on being a smaller, lower-power card with the same architecture and API coverage. For a system with a 550 W PSU recommendation and dual-slot clearance, the RTX 4070 SUPER fits where the RTX 4090 would not. But in pure performance terms, the database records no scenario where the RTX 4070 SUPER outpaces the RTX 4090.

The RTX 4090's nearest rivals in the database include the AMD Radeon Pro W6600M at a delta of -2.5%, the AMD Radeon PRO V710 at 2.9%, and the Intel Arc Pro A60 at 0%. The RTX 4070 SUPER's nearest rivals are the NVIDIA Quadro M6000 24 GB at -0.1%, the NVIDIA GeForce RTX 5050 Mobile at -0.1%, and the NVIDIA GeForce RTX 4090 Mobile at -1%. These proximity bands show that each card competes in a different performance tier, with the RTX 4090 sitting among workstation-class GPUs and the RTX 4070 SUPER among mobile and older desktop flagships.

Specification Differences

| Specification | NVIDIA GeForce RTX 4090 | NVIDIA GeForce RTX 4070 SUPER |

|---|---|---|

| Chip | AD102 | AD104 |

| Process Node | 5 nm (TSMC) | 5 nm (TSMC) |

| Transistors | 76,300 million | 35,800 million |

| Die Size | 609 mm² | 294 mm² |

| Transistor Density | 125.3M / mm² | 121.8M / mm² |

| Base Clock | 2235 MHz | 1980 MHz |

| Boost Clock | 2520 MHz | 2475 MHz |

| Memory Size | 24 GB GDDR6X | 12 GB GDDR6X |

| Memory Bus | 384 bit | 192 bit |

| Memory Bandwidth | 1.01 TB/s | 504.2 GB/s |

| Shading Units | 16,384 | 7,168 |

| TMUs | 512 | 224 |

| ROPs | 176 | 80 |

| RT Cores | 128 | 56 |

| Tensor Cores | 512 | 224 |

| Pixel Rate | 443.5 GPixel/s | 198.0 GPixel/s |

| Texture Rate | 1,290.2 GTexel/s | 554.4 GTexel/s |

| FP32 | 82.58 TFLOPS | 35.48 TFLOPS |

| FP16 | 82.58 TFLOPS (1:1) | 35.48 TFLOPS (1:1) |

| TDP | 450 W | 220 W |

| Slot Width | Triple-slot | Dual-slot |

| Suggested PSU | 850 W | 550 W |

| Dimensions (L×H×W) | 304 × 137 × 61 mm | 267 × 112 × 42 mm |

| Launch MSRP | 1,599 USD | 599 USD |

| Release Date | 2022-09-19 | 2024-01-16 |

Head-to-Head Benchmarks

The RTX 4090 wins every recorded head-to-head test. The biggest win is 3DMark Steel Nomad DX12, where the RTX 4090 scores 9,223 against 4,627, a 99.3% delta. This is the database's clearest signal that the RTX 4090 offers nearly double the DX12 performance of the RTX 4070 SUPER in this workload.

PassMark GPU Compute shows the second-largest gap. The RTX 4090 scores 26,613 versus 17,108, a 55.6% advantage. This test emphasizes general compute throughput, and the RTX 4090's 82.58 TFLOPS FP32 capacity against 35.48 TFLOPS explains the scale of the lead.

Geekbench OpenCL also favors the RTX 4090 heavily. Scores are 255,416 versus 172,795, a 47.8% delta. Geekbench Vulkan is closer but still decisive, with 271,631 versus 205,624, a 32.1% advantage. The RTX 4090's Vulkan score is its highest single benchmark result in the set.

PassMark DirectX 12 shows a 36.4% delta, with scores of 150 versus 110. PassMark DirectX 10 follows at 34.1%, scores 224 versus 167. PassMark G3D, a broad 3D graphics aggregate, has 38,194 versus 29,995, a 27.3% delta. PassMark DirectX 11 records 326 versus 273, a 19.4% delta. PassMark DirectX 9 shows 397 versus 344, a 15.4% delta.

The smallest margin is PassMark G2D, where the RTX 4090 leads 1,299 versus 1,184, a 9.7% delta. This 2D-oriented test is the only benchmark where the two cards are close, indicating that the RTX 4090's advantages are concentrated in 3D and compute workloads rather than basic graphics output.

The overall picture is one of consistent, across-the-board superiority for the RTX 4090. Its wins range from 9.7% to 99.3%, and every category from legacy DirectX 9 to modern DX12 and compute favors it. The RTX 4070 SUPER's closest recorded performance is in 2D, which is the least demanding test in the set. For any workload involving 3D rendering, ray tracing, or general GPU compute, the database shows the RTX 4090 with a commanding lead.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 4070 SUPER
RTX 4090
Core Specs
Shading Units
7,168
16,384 +128.6%
Shaders
7,168
16,384 +128.6%
TMUs
224
512 +128.6%
ROPs
80
176 +120.0%
SM Count
56
128 +128.6%
Clocks
Base Clock
1980 MHz
2235 MHz
Boost Clock
2475 MHz
2520 MHz
Memory Clock
1313 MHz 21 Gbps effective
1313 MHz 21 Gbps effective
Memory
Memory Size
12 GB
24 GB
VRAM (MB)
12,288
24,576 +100.0%
Memory Type
GDDR6X
GDDR6X
Memory Bus
192 bit
384 bit
Bandwidth
504.2 GB/s
1.01 TB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
48 MB
72 MB
Performance
Pixel Rate
198.0 GPixel/s
443.5 GPixel/s
Texture Rate
554.4 GTexel/s
1,290.2 GTexel/s
FP32 (TFLOPS)
35.48 TFLOPS
82.58 TFLOPS
FP64 (TFLOPS)
554.4 GFLOPS (1:64)
1,290.2 GFLOPS (1:64)
FP16 (TFLOPS)
35.48 TFLOPS (1:1)
82.58 TFLOPS (1:1)
AI/RT
RT Cores
56
128 +128.6%
Tensor Cores
224
512 +128.6%
Power
TDP
220 W
450 W
TDP (W)
220
450 +104.5%
Suggested PSU
550 W
850 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Ada Lovelace
Ada Lovelace
GPU Name
AD104
AD102
Generation
GeForce 40
GeForce 40
Process Size
5 nm
5 nm
Transistors
35,800 million
76,300 million
Die Size
294 mm²
609 mm²
Foundry
TSMC
TSMC
Density
121.8M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
8.9
Shader Model
6.9
6.8
Physical
Slot Width
Dual-slot
Triple-slot
Length
267 mm 10.5 inches
304 mm 12 inches
Height
112 mm 4.4 inches
137 mm 5.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Launch Price
599 USD
1,599 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 30
GeForce 30
Successor
GeForce 50
GeForce 50
View GeForce RTX 4070 SUPER Details View GeForce RTX 4090 Details