NVIDIA GeForce RTX 3090 vs NVIDIA GeForce RTX 4070 Mobile Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 3090

CORE STATE GA102
VRAM 24 GB
CLOCK SPEED 1695 MHz
TDP 350 W
BUS WIDTH 384 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2020
VS
NVIDIA
GEFORCE

GeForce RTX 4070 Mobile

CORE STATE AD106
VRAM 8 GB
CLOCK SPEED 1695 MHz
TDP 115 W
BUS WIDTH 128 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
5,118
N/A
geekbench_opencl
172,758
109,197
geekbench_vulkan
53,927
108,367
passmark_directx_10
182
116
passmark_directx_11
220
179
passmark_directx_12
110
85
passmark_directx_9
268
223
passmark_g2d
1,063
763
passmark_g3d
26,645
19,587
passmark_gpu_compute
15,356
8,399

Analysis: NVIDIA GeForce RTX 3090 vs NVIDIA GeForce RTX 4070 Mobile

The NVIDIA GeForce RTX 3090 and the NVIDIA GeForce RTX 4070 Mobile represent two distinct philosophies in GPU design: one a desktop behemoth built for raw throughput, the other a mobile chip engineered to deliver competitive performance within a strict power envelope. While their average benchmark scores are remarkably close—the RTX 3090 averages 27,565, and the RTX 4070 Mobile averages 27,435, a mere 0.5% difference—the data reveals a fascinating split in their strengths. The desktop card dominates traditional rasterization and compute workloads, while the mobile chip shows a surprising edge in a specific modern API test, indicating that the choice between them hinges entirely on the target use case and platform constraints.

Where Each One Wins

The benchmark results paint a clear picture of two different performance profiles. The RTX 3090 is the decisive winner in 8 of the 9 head-to-head tests, with its largest margins coming in compute-heavy and legacy API workloads. Its lead in Passmark’s GPU compute test is a staggering 82.8%, and it also holds a 58.2% advantage in Geekbench OpenCL, underscoring its massive FP32 throughput of 35.58 TFLOPS compared to the mobile chip’s 15.62 TFLOPS. This makes the RTX 3090 the clear choice for any application that leverages raw compute power, such as rendering, scientific simulation, or machine learning inference.

Conversely, the RTX 4070 Mobile’s only victory is in the Geekbench Vulkan test, where it scores 108,367 against the RTX 3090’s 53,927, representing a 50.2% advantage for the mobile part. This is a significant outlier in the data, suggesting that the Ada Lovelace architecture’s implementation of Vulkan is markedly more efficient than that of the older Ampere design. For users prioritizing modern cross-platform graphics APIs in games or professional applications, this result implies the RTX 4070 Mobile could provide a smoother experience despite its lower raw compute specs. The RTX 3090 wins the Passmark G3D test by 36% (26,645 vs 19,587), indicating superior DirectX gaming performance, but the Vulkan result cannot be ignored.

Architecture Differences

The architectural chasm between these two GPUs is vast, explaining their divergent performance characteristics. The RTX 3090 is built on Samsung’s 8 nm process, housing 28,300 million transistors on a massive 628 mm² die, resulting in a transistor density of 45.1M per mm². In contrast, the RTX 4070 Mobile utilizes TSMC’s 5 nm node, packing 22,900 million transistors into a much smaller 188 mm² die, achieving a far higher density of 121.8M per mm². This process advantage is foundational to the mobile chip’s efficiency, allowing it to deliver competitive performance at a 115 W TDP compared to the desktop card’s 350 W.

The core configurations diverge sharply. The RTX 3090 features 10,496 shading units, 328 TMUs, 112 ROPs, 82 RT cores, and 328 tensor cores, while the RTX 4070 Mobile ships with just 4,608 shading units, 144 TMUs, 48 ROPs, 36 RT cores, and 144 tensor cores. Despite having less than half the shaders and TMUs, the RTX 4070 Mobile manages to stay within 0.5% of the RTX 3090’s average benchmark score, a signal of the architectural efficiency of Ada Lovelace. Memory configurations also differ fundamentally: the RTX 3090 uses 24 GB of GDDR6X on a 384-bit bus for 936.2 GB/s of bandwidth, while the RTX 4070 Mobile uses 8 GB of GDDR6 on a 128-bit bus, limiting bandwidth to 256.0 GB/s. Both support DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, but the RTX 3090 connects via PCIe 4.0 x16, whereas the mobile part uses a narrower PCIe 4.0 x8 interface.

Head-to-Head Benchmarks

The head-to-head results offer a detailed look at where each GPU excels. The RTX 3090’s most dominant win is in Passmark GPU Compute, where its score of 15,356 dwarfs the RTX 4070 Mobile’s 8,399, a 82.8% difference that highlights the desktop card’s sheer parallel processing capability. In Geekbench OpenCL, the RTX 3090 again dominates, scoring 172,758 versus 109,197, a 58.2% lead that reinforces its compute superiority. Even in legacy DirectX tests, the older card prevails: it leads by 56.9% in Passmark DirectX 10 (182 vs 116), by 22.9% in DirectX 11 (220 vs 179), by 29.4% in DirectX 12 (110 vs 85), and by 20.2% in DirectX 9 (268 vs 223).

The RTX 3090 also wins in synthetic gaming and 2D tests. In Passmark G3D, it scores 26,645 against 19,587, a 36% advantage, and in the G2D test it leads 1,063 to 763, a 39.3% margin. These results suggest that for most DirectX-based gaming workloads, the RTX 3090 provides substantially higher frame rates. However, the RTX 4070 Mobile’s striking 50.2% victory in Geekbench Vulkan (108,367 vs 53,927) is the single most important counterpoint. This result implies that in Vulkan-native titles or applications, the mobile GPU’s architecture is far more efficient, potentially offering a better experience despite its lower memory bandwidth and fewer cores. This single win prevents the RTX 3090 from being a universal recommendation, as the Vulkan API is increasingly prevalent in modern games and professional tools.

FAQ

Q: Which GPU has a higher average benchmark score?

A: The NVIDIA GeForce RTX 3090 has a slightly higher average score of 27,565, while the RTX 4070 Mobile averages 27,435. The delta is only 0.5%, making them statistically near-identical in overall performance.

Q: Why is the RTX 3090’s score so much higher in compute tests?

A: The RTX 3090’s FP32 throughput of 35.58 TFLOPS is more than double the RTX 4070 Mobile’s 15.62 TFLOPS. This is reflected in the 82.8% lead in Passmark GPU Compute and the 58.2% lead in Geekbench OpenCL.

Q: Where does the RTX 4070 Mobile outperform the RTX 3090?

A: The RTX 4070 Mobile wins only in the Geekbench Vulkan test, scoring 108,367 to the RTX 3090’s 53,927, a 50.2% advantage. This suggests a more efficient Vulkan implementation in the Ada Lovelace architecture.

Q: How do their memory configurations compare?

A: The RTX 3090 has 24 GB of GDDR6X on a 384-bit bus, yielding 936.2 GB/s of bandwidth. The RTX 4070 Mobile has 8 GB of GDDR6 on a 128-bit bus, providing 256.0 GB/s, making the desktop card’s bandwidth advantage substantial.

Q: What are the power consumption differences?

A: The RTX 3090 has a TDP of 350 W and requires a 750 W power supply, while the RTX 4070 Mobile has a TDP of 115 W. The desktop card is triple-slot and uses a 12-pin connector, whereas the mobile chip is an IGP with no power connectors.

Q: Which GPU is better for DirectX gaming based on the data?

A: The RTX 3090 wins all Passmark DirectX tests, with leads of 56.9% in DirectX 10, 22.9% in DirectX 11, 29.4% in DirectX 12, and 20.2% in DirectX 9, indicating superior performance in traditional gaming APIs.

The Verdict

The data presents a nuanced picture that defies a simple “better” or “worse” verdict. The RTX 3090 is the overwhelming winner in raw performance across nearly every benchmark category, particularly compute and DirectX workloads. Its 82.8% lead in GPU compute and 36% lead in G3D make it the superior choice for users who need maximum throughput for rendering, simulation, or high-end DirectX gaming. The desktop card’s 24 GB of memory and 936.2 GB/s bandwidth also provide a significant advantage for large datasets and high-resolution textures. For a stationary desktop workstation or gaming rig where power consumption and physical space are not constraints, the RTX 3090 is clearly the more capable part.

However, the RTX 4070 Mobile’s 50.2% victory in Geekbench Vulkan is a crucial data point that cannot be dismissed. This result indicates that the Ada Lovelace architecture is significantly more efficient in this modern API, which is becoming the standard for many new games and cross-platform applications. Combined with its 115 W TDP and IGP form factor, the RTX 4070 Mobile is the logical choice for a laptop or compact system where portability and efficiency are paramount. Its average score is within 0.5% of the RTX 3090, meaning that for users primarily engaged in Vulkan-based workloads, the mobile chip could deliver comparable or better performance without the need for a massive power supply and cooling solution. The verdict depends on the platform: for a desktop, the RTX 3090’s brute force is unmatched; for a portable system, the RTX 4070 Mobile offers a compelling efficiency-driven alternative.

Specification Differences

| Specification | NVIDIA GeForce RTX 3090 | NVIDIA GeForce RTX 4070 Mobile |

|:--- |:--- |:--- |

| Architecture | Ampere | Ada Lovelace |

| Process Node | 8 nm | 5 nm |

| Foundry | Samsung | TSMC |

| Transistors | 28,300 million | 22,900 million |

| Die Size | 628 mm² | 188 mm² |

| Transistor Density | 45.1M / mm² | 121.8M / mm² |

| Memory Size | 24 GB | 8 GB |

| Memory Type | GDDR6X | GDDR6 |

| Memory Bus Width | 384 bit | 128 bit |

| Memory Bandwidth | 936.2 GB/s | 256.0 GB/s |

| Memory Clock | 1219 MHz (19.5 Gbps effective) | 2000 MHz (16 Gbps effective) |

| Shading Units | 10496 | 4608 |

| TMUs | 328 | 144 |

| ROPs | 112 | 48 |

| RT Cores | 82 | 36 |

| Tensor Cores | 328 | 144 |

| Pixel Rate | 189.8 GPixel/s | 81.36 GPixel/s |

| Texture Rate | 556.0 GTexel/s | 244.1 GTexel/s |

| FP32 Performance | 35.58 TFLOPS | 15.62 TFLOPS |

| TDP | 350 W | 115 W |

| Slot Width | Triple-slot | IGP |

| Power Connectors | 1x 12-pin | None |

| Suggested PSU | 750 W | None |

| Bus Interface | PCIe 4.0 x16 | PCIe 4.0 x8 |

| Base Clock | 1395 MHz | 1395 MHz |

| Boost Clock | 1695 MHz | 1695 MHz |

| Release Date | 2020-08-31 | 2023-01-02 |

| Production Status | End-of-life | Active |

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 3090
RTX 4070 Mobile
Core Specs
Shading Units
10,496
4,608 -56.1%
Shaders
10,496
4,608 -56.1%
TMUs
328
144 -56.1%
ROPs
112
48 -57.1%
SM Count
82
36 -56.1%
Clocks
Base Clock
1395 MHz
1395 MHz
Boost Clock
1695 MHz
1695 MHz
Memory Clock
1219 MHz 19.5 Gbps effective
2000 MHz 16 Gbps effective
Memory
Memory Size
24 GB
8 GB
VRAM (MB)
24,576
8,192 -66.7%
Memory Type
GDDR6X
GDDR6
Memory Bus
384 bit
128 bit
Bandwidth
936.2 GB/s
256.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
6 MB
32 MB
Performance
Pixel Rate
189.8 GPixel/s
81.36 GPixel/s
Texture Rate
556.0 GTexel/s
244.1 GTexel/s
FP32 (TFLOPS)
35.58 TFLOPS
15.62 TFLOPS
FP64 (TFLOPS)
556.0 GFLOPS (1:64)
244.1 GFLOPS (1:64)
FP16 (TFLOPS)
35.58 TFLOPS (1:1)
15.62 TFLOPS (1:1)
AI/RT
RT Cores
82
36 -56.1%
Tensor Cores
328
144 -56.1%
Power
TDP
350 W
115 W
TDP (W)
350
115 -67.1%
Suggested PSU
750 W
Power Connectors
1x 12-pin
None
Architecture
Architecture
Ampere
Ada Lovelace
GPU Name
GA102
AD106
Generation
GeForce 30
GeForce 40 Mobile
Process Size
8 nm
5 nm
Transistors
28,300 million
22,900 million
Die Size
628 mm²
188 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
121.8M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Triple-slot
IGP
Length
336 mm 13.2 inches
Height
140 mm 5.5 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
Portable Device Dependent
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x8
Other
Launch Price
1,499 USD
Production
End-of-life
Active
Predecessor
GeForce 20
GeForce 30 Mobile
Successor
GeForce 40
GeForce 50 Mobile
View GeForce RTX 3090 Details View GeForce RTX 4070 Mobile Details