NVIDIA GeForce RTX 3090 vs NVIDIA GeForce RTX 4070 Ti SUPER Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 3090

CORE STATE GA102
VRAM 24 GB
CLOCK SPEED 1695 MHz
TDP 350 W
BUS WIDTH 384 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2020
VS
NVIDIA
GEFORCE

GeForce RTX 4070 Ti SUPER

CORE STATE AD103
VRAM 16 GB
CLOCK SPEED 2610 MHz
TDP 285 W
BUS WIDTH 256 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2024

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
5,118
5,569
geekbench_opencl
172,758
199,267
geekbench_vulkan
53,927
53,683
passmark_directx_10
182
181
passmark_directx_11
220
278
passmark_directx_12
110
119
passmark_directx_9
268
360
passmark_g2d
1,063
1,225
passmark_g3d
26,645
31,811
passmark_gpu_compute
15,356
18,372

Analysis: NVIDIA GeForce RTX 3090 vs NVIDIA GeForce RTX 4070 Ti SUPER

Head-to-Head Benchmarks

The benchmark data shows a decisive overall victory for the NVIDIA GeForce RTX 4070 Ti SUPER, which wins 8 of the 10 recorded head-to-head comparisons. The only two wins for the RTX 3090 come in specific API-focused tests, and even those are narrow margins. The largest single-test advantage belongs to the RTX 4070 Ti SUPER in the PassMark DirectX 9 test, where it scores 360 versus 268, a 34.3% lead. This pattern of large wins in legacy DirectX workloads carries over to the DirectX 11 test, where the 4070 Ti SUPER leads by 26.4% (278 versus 220).

The compute-oriented results reinforce the same conclusion. In PassMark GPU Compute, the 4070 Ti SUPER posts 18,372 against 15,356 for the 3090, a 19.6% advantage. The Geekbench OpenCL test shows a 15.3% lead (199,267 versus 172,758), and the PassMark G3D score favors the newer card by 19.4% (31,811 versus 26,645). Even the modern 3DMark Steel Nomad DX12 test, which typically stresses contemporary architectures, shows the 4070 Ti SUPER ahead by 8.8% (5,569 versus 5,118). The 2D-focused PassMark G2D test adds another win for the 4070 Ti SUPER, with a 15.2% margin (1,225 versus 1,063).

The RTX 3090’s two victories are statistically insignificant. In Geekbench Vulkan, the 3090 scores 53,927 versus 53,683 for the 4070 Ti SUPER, a delta of only 0.5%. The PassMark DirectX 10 result is even closer: 182 versus 181, again a 0.5% difference. These near-ties suggest that in Vulkan and DirectX 10 workloads, the two cards perform effectively at parity, but the magnitude of the 4070 Ti SUPER’s wins in other tests dwarfs these nominal reversals.

Looking at the aggregate metrics, the average benchmark score for the 4070 Ti SUPER is 31,087, which places it in the 76th percentile among all GPUs. The RTX 3090 averages 27,565, sitting in the 73rd percentile. The 4070 Ti SUPER’s nearest rivals in the database include the NVIDIA TITAN RTX at 31,676 (a 1.9% higher average) and the NVIDIA RTX PRO 4500 Blackwell at 31,532 (1.4% higher). The 3090’s nearest rivals include the AMD Radeon RX 6700 XT at 27,425 (0.5% lower) and the NVIDIA GeForce RTX 4070 Mobile at 27,435 (0.5% lower). These context points show that the 4070 Ti SUPER operates in a higher performance tier than the 3090, despite the latter’s older flagship status.

Where Each One Wins

The RTX 4070 Ti SUPER dominates in scenarios that leverage its newer architecture’s efficiency and per-clock execution. The DirectX 11 and DirectX 9 wins, with deltas of 26.4% and 34.3% respectively, indicate that older API titles, many of which are still widely played, run substantially faster on the 40-series card. The compute-heavy workloads, as captured by PassMark GPU Compute and Geekbench OpenCL, also favor the 4070 Ti SUPER by roughly 15% to 20%, making it the stronger choice for GPGPU tasks such as rendering, simulation, or machine learning inference that do not rely on specialized tensor paths.

The 3DMark Steel Nomad DX12 result, while a smaller win at 8.8%, still favors the 4070 Ti SUPER, suggesting that even current-generation DX12 games generally perform better on the newer card. The PassMark G3D score, which aggregates multiple DirectX tests, confirms a 19.4% overall advantage for the 4070 Ti SUPER. For users prioritizing raw frame rates in contemporary titles and general 3D acceleration, the data consistently points to the 4070 Ti SUPER.

The RTX 3090’s wins are confined to Geekbench Vulkan and PassMark DirectX 10, where it edges ahead by 0.5% in both cases. These are essentially statistical noise, but they do indicate that Vulkan-based applications, such as certain Linux-native games or emulators, may see no disadvantage on the 3090. The DirectX 10 test is more of a legacy curiosity, as few modern titles use that API, but the 3090’s slight edge there shows it retains competence in older software stacks.

For memory-intensive workloads, the 3090 offers a substantial capacity advantage with 24 GB versus 16 GB. The database records 936.2 GB/s bandwidth for the 3090 against 672.3 GB/s for the 4070 Ti SUPER. This makes the 3090 the better candidate for tasks that demand very large datasets residing in VRAM, such as high-resolution texture packs or multi-model AI inference batches that exceed 16 GB. However, the 4070 Ti SUPER’s higher raw compute throughput in most measured tests means that for memory demands under 16 GB, the newer card will generally finish faster.

The Verdict

The data supports a clear verdict: the NVIDIA GeForce RTX 4070 Ti SUPER is the superior performer in the vast majority of recorded benchmarks. Its 8 wins to 2, combined with an average score advantage of 3,522 points (31,087 versus 27,565), positions it as the more capable card for general gaming, 3D rendering, and compute workloads. The 4070 Ti SUPER also achieves this with a lower thermal design power of 285 W compared to 350 W for the 3090, and it lists a lower suggested PSU requirement of 600 W versus 750 W. These efficiency figures, paired with higher performance, make the 4070 Ti SUPER the rational choice for any system where power draw or PSU headroom matters.

The RTX 3090 remains relevant only for a narrow set of use cases. Its 24 GB memory capacity is the single biggest differentiator, and for workloads that genuinely require more than 16 GB of VRAM, the 3090 is the only option between these two. The 3090 also narrowly wins the Vulkan test, but the 0.5% margin is negligible. For anyone who does not need the extra memory, the 3090’s lower average score and higher power consumption make it the weaker pick.

The percentile ranking reinforces the gap: 76th percentile for the 4070 Ti SUPER versus 73rd for the 3090. The 4070 Ti SUPER’s nearest rivals in the database are all higher-scoring cards, whereas the 3090’s nearest rivals include mid-range options like the RX 6700 XT. This indicates that the 3090, despite its former flagship status, now sits closer to upper-mid-range performance territory, while the 4070 Ti SUPER maintains a position among stronger contemporary GPUs.

FAQ

Q: Which card wins more benchmarks head-to-head?

A: The NVIDIA GeForce RTX 4070 Ti SUPER wins 8 of the 10 recorded comparisons, with the RTX 3090 winning only 2 (Geekbench Vulkan and PassMark DirectX 10, both by 0.5%).

Q: What is the biggest performance gap between the two cards?

A: The largest delta is in the PassMark DirectX 9 test, where the 4070 Ti SUPER scores 360 versus 268 for the 3090, a 34.3% advantage.

Q: Does the RTX 3090 have any clear advantage?

A: Yes, in memory capacity and bandwidth. The 3090 has 24 GB of GDDR6X on a 384-bit bus providing 936.2 GB/s, compared to 16 GB on a 256-bit bus for the 4070 Ti SUPER at 672.3 GB/s.

Q: How do their average benchmark scores compare?

A: The 4070 Ti SUPER has an average benchmark score of 31,087, while the 3090 averages 27,565. The 4070 Ti SUPER ranks in the 76th percentile of all GPUs, the 3090 in the 73rd.

Q: Which card is more power-efficient according to the data?

A: The 4070 Ti SUPER has a TDP of 285 W and a suggested PSU of 600 W, while the 3090 has a TDP of 350 W and a suggested PSU of 750 W.

Q: Are there any tests where the two cards perform identically?

A: The closest results are in Geekbench Vulkan (53,683 versus 53,927) and PassMark DirectX 10 (181 versus 182), both showing a 0.5% difference, which is effectively a tie.

Architecture Differences

The two cards represent distinct architectural generations from NVIDIA. The RTX 4070 Ti SUPER is built on the Ada Lovelace architecture using the AD103 chip, fabricated on a 5 nm process at TSMC. The RTX 3090 uses the Ampere architecture with the GA102 chip, manufactured on an 8 nm process at Samsung. This process node difference is substantial: 5 nm versus 8 nm, which contributes to the 4070 Ti SUPER’s efficiency advantages.

Transistor counts differ dramatically. The AD103 chip packs 45,900 million transistors on a die size of 379 mm², yielding a transistor density of 121.1 million per mm². The GA102 has 28,300 million transistors on a larger 628 mm² die, giving a density of just 45.1 million per mm². The smaller, denser chip in the 4070 Ti SUPER explains how it achieves higher clock speeds and better performance per watt.

The stream processor configurations also differ. The 4070 Ti SUPER has 8,448 shading units, 264 texture mapping units, and 96 raster output units. The 3090 has more of each: 10,496 shading units, 328 TMUs, and 112 ROPs. Despite having fewer cores, the 4070 Ti SUPER achieves higher FP32 throughput at 44.10 TFLOPS versus 35.58 TFLOPS for the 3090, a direct result of the much higher clock speeds (2,340 MHz base and 2,610 MHz boost versus 1,395 MHz base and 1,695 MHz boost).

Ray tracing and tensor core counts follow the same pattern. The 4070 Ti SUPER has 66 RT cores and 264 tensor cores, while the 3090 has 82 RT cores and 328 tensor cores. Again, the newer architecture’s higher clocks compensate for the lower core counts in raw throughput. Both cards support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The 4070 Ti SUPER uses a 1x 16-pin power connector, while the 3090 uses a 1x 12-pin connector. Both are triple-slot cards, and both offer identical display outputs: 1x HDMI 2.1 and 3x DisplayPort 1.4a.

Specification Differences

The most obvious specification gap is memory. The 4070 Ti SUPER has 16 GB of GDDR6X on a 256-bit bus, while the 3090 has 24 GB of GDDR6X on a 384-bit bus. This gives the 3090 a bandwidth advantage of 936.2 GB/s versus 672.3 GB/s for the 4070 Ti SUPER. Memory clock speeds also differ: the 4070 Ti SUPER runs at 1,313 MHz (21 Gbps effective), and the 3090 at 1,219 MHz (19.5 Gbps effective).

Clocks are a significant differentiator. The 4070 Ti SUPER has a base clock of 2,340 MHz and a boost clock of 2,610 MHz. The 3090’s base clock is 1,395 MHz with a boost of 1,695 MHz. This nearly 1,000 MHz advantage in boost clock explains the 4070 Ti SUPER’s higher FP32, FP16, pixel, and texture rates despite fewer cores. Pixel rate for the 4070 Ti SUPER is 250.6 GPixel/s versus 189.8 GPixel/s for the 3090. Texture rate is 689.0 GTexel/s versus 556.0 GTexel/s.

Power specifications favor the 4070 Ti SUPER: 285 W TDP versus 350 W for the 3090, and a suggested PSU of 600 W versus 750 W. Physical dimensions are close, with the 4070 Ti SUPER measuring 310 mm in length, 140 mm in height, and 61 mm in width. The 3090 is longer at 336 mm, with the same 140 mm height and 61 mm width. Both cards are end-of-life products. The 4070 Ti SUPER launched on 2024-01-23 with a launch MSRP of 799 USD, and the 3090 launched on 2020-08-31 with a launch MSRP of 1,499 USD. The 4070 Ti SUPER belongs to the GeForce 40-series with a predecessor in the GeForce 30 series and a successor in the GeForce 50 series; the 3090 belongs to the GeForce 30-series with a predecessor in the GeForce 20 series and a successor in the GeForce 40 series.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 3090
RTX 4070 Ti SUPER
Core Specs
Shading Units
10,496
8,448 -19.5%
Shaders
10,496
8,448 -19.5%
TMUs
328
264 -19.5%
ROPs
112
96 -14.3%
SM Count
82
66 -19.5%
Clocks
Base Clock
1395 MHz
2340 MHz
Boost Clock
1695 MHz
2610 MHz
Memory Clock
1219 MHz 19.5 Gbps effective
1313 MHz 21 Gbps effective
Memory
Memory Size
24 GB
16 GB
VRAM (MB)
24,576
16,384 -33.3%
Memory Type
GDDR6X
GDDR6X
Memory Bus
384 bit
256 bit
Bandwidth
936.2 GB/s
672.3 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
6 MB
48 MB
Performance
Pixel Rate
189.8 GPixel/s
250.6 GPixel/s
Texture Rate
556.0 GTexel/s
689.0 GTexel/s
FP32 (TFLOPS)
35.58 TFLOPS
44.10 TFLOPS
FP64 (TFLOPS)
556.0 GFLOPS (1:64)
689.0 GFLOPS (1:64)
FP16 (TFLOPS)
35.58 TFLOPS (1:1)
44.10 TFLOPS (1:1)
AI/RT
RT Cores
82
66 -19.5%
Tensor Cores
328
264 -19.5%
Power
TDP
350 W
285 W
TDP (W)
350
285 -18.6%
Suggested PSU
750 W
600 W
Power Connectors
1x 12-pin
1x 16-pin
Architecture
Architecture
Ampere
Ada Lovelace
GPU Name
GA102
AD103
Generation
GeForce 30
GeForce 40
Process Size
8 nm
5 nm
Transistors
28,300 million
45,900 million
Die Size
628 mm²
379 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
121.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
8.9
Shader Model
6.8
6.9
Physical
Slot Width
Triple-slot
Triple-slot
Length
336 mm 13.2 inches
310 mm 12.2 inches
Height
140 mm 5.5 inches
140 mm 5.5 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Launch Price
1,499 USD
799 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 20
GeForce 30
Successor
GeForce 40
GeForce 50
View GeForce RTX 3090 Details View GeForce RTX 4070 Ti SUPER Details