NVIDIA GeForce RTX 4080 SUPER vs NVIDIA GeForce RTX 4090 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 4080 SUPER

CORE STATE AD103
VRAM 16 GB
CLOCK SPEED 2550 MHz
TDP 320 W
BUS WIDTH 256 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2024
VS
NVIDIA
GEFORCE

GeForce RTX 4090

CORE STATE AD102
VRAM 24 GB
CLOCK SPEED 2520 MHz
TDP 450 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
6,600
9,223
geekbench_opencl
219,065
255,416
geekbench_vulkan
260,075
271,631
passmark_directx_10
193
224
passmark_directx_11
301
326
passmark_directx_12
134
150
passmark_directx_9
381
397
passmark_g2d
1,270
1,299
passmark_g3d
34,245
38,194
passmark_gpu_compute
19,822
26,613

Analysis: NVIDIA GeForce RTX 4080 SUPER vs NVIDIA GeForce RTX 4090

The NVIDIA GeForce RTX 4090 is the definitive performance king in this comparison, winning all ten head-to-head benchmarks against the RTX 4080 SUPER. The data shows a clear hierarchy within the Ada Lovelace lineup, with the 4090 delivering leads that range from a marginal 2.3% to a commanding 39.7%, depending on the workload. While the RTX 4080 SUPER is no slouch, its results consistently place it a tier below, making the choice between them a matter of how much absolute performance is required versus what the rest of the system can support.

Head-to-Head Benchmarks

The most significant gap between these two cards appears in the modern DirectX 12 workload of 3DMark Steel Nomad. The RTX 4090 scores 9223, which is a substantial 39.7% higher than the 6600 achieved by the RTX 4080 SUPER. This is the largest delta in the entire benchmark suite and indicates that the 4090’s additional hardware resources scale exceptionally well in newer, more demanding rendering paths.

Compute performance shows a similar story of dominance. In the Passmark GPU Compute test, the RTX 4090 scores 26613, which is 34.3% ahead of the 4080 SUPER’s 19822. This suggests that the 4090 is not just a gaming card but also a significantly more capable compute accelerator, making it a stronger choice for tasks like rendering or scientific workloads that leverage raw FP32 throughput.

The lead narrows considerably in other API tests, but the 4090 still wins. In Geekbench OpenCL, the 4090 posts 255416 versus 219065 for the 4080 SUPER, a 16.6% advantage. The Vulkan result is much closer, with the 4090 winning 271631 to 260075, a slim 4.4% margin. This indicates that while the 4090 has more raw power, the performance gap can shrink depending on the API and driver optimization.

DirectX legacy tests show the same pattern of 4090 superiority. In Passmark DirectX 10, the 4090 wins by 16.1% (224 vs 193), and in DirectX 11, it leads by 8.3% (326 vs 301). The DirectX 12 Passmark result shows an 11.9% gap (150 vs 134), and even the older DirectX 9 test gives the 4090 a 4.2% edge (397 vs 381).

The overall graphics score in Passmark G3D has the 4090 at 38194, which is 11.5% higher than the 34245 of the 4080 SUPER. Even in the less compute-intensive Passmark G2D test, the 4090 wins, though only by 2.3% (1299 vs 1270). Across every single metric in the suite, the RTX 4090 holds the lead, with a perfect 10-0 win record.

The Verdict

The data is unequivocal: the NVIDIA GeForce RTX 4090 is the superior graphics card for anyone who demands the absolute highest frame rates and compute throughput. Its victories are not narrow; they are often substantial, particularly in DirectX 12 and compute benchmarks where it can be nearly 40% faster. The 4090’s average benchmark score of 60347 places it in the 88th percentile of all GPUs, while the 4080 SUPER’s 54209 average puts it at the 86th percentile. This difference in percentile underscores that the 4090 is not just a little better, but a meaningful step up in the overall hierarchy.

The RTX 4080 SUPER, however, is the more sensible pick for a larger group of users. Its performance is still very high, and its benchmark scores are close to the RTX 4080, as shown in its nearest rivals list where it is only -0.1% behind the non-SUPER 4080. For a system where power consumption and PSU requirements are a concern, the 4080 SUPER has a clear advantage, with a 320 W TDP and 700 W suggested PSU versus the 4090’s 450 W TDP and 850 W suggested PSU. If the nearly 40% performance gain in the most demanding scenarios is not essential, the 4080 SUPER is the more balanced option, especially for those with power supply constraints.

Architecture Differences

Both cards are built on the same fundamental architecture, but they use different physical chips that scale the design to different levels. The RTX 4090 is based on the AD102 chip, while the RTX 4080 SUPER uses the AD103. This difference in the silicon is the primary driver of the performance gap. The AD102 chip is significantly larger, with a die size of 609 mm², compared to the 379 mm² of the AD103. This allows the 4090 to pack in far more transistors, 76,300 million versus 45,900 million on the 4080 SUPER. Interestingly, the process node is identical at 5 nm from TSMC, but the 4090 has a slightly higher transistor density of 125.3M / mm² versus the 4080 SUPER’s 121.1M / mm².

The core configuration differences are stark. The RTX 4090 features 16384 shading units, while the 4080 SUPER has 10240. This 60% advantage in shader count is the core reason for the 4090’s higher FP32 throughput of 82.58 TFLOPS versus 52.22 TFLOPS on the 4080 SUPER. This pattern repeats across the board: the 4090 has 512 texture mapping units (TMUs) versus 320, and 176 raster output units (ROPs) versus 112. Ray tracing and AI hardware also scale up, with the 4090 having 128 RT cores and 512 tensor cores, while the 4080 SUPER has 80 and 320 respectively. Both cards support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

Specification Differences

The most visible specification difference is memory. The RTX 4090 comes with 24 GB of GDDR6X memory on a 384-bit bus, yielding a bandwidth of 1.01 TB/s. The RTX 4080 SUPER has 16 GB of the same memory type but on a 256-bit bus, resulting in a lower bandwidth of 736.3 GB/s. While the 4080 SUPER has a slightly higher memory clock at 1438 MHz (23 Gbps effective) compared to the 4090’s 1313 MHz (21 Gbps effective), the 4090’s wider bus gives it a massive bandwidth advantage. This makes the 4090 better suited for 4K gaming and large datasets.

Clock speeds are another point of differentiation, though the 4080 SUPER is slightly faster here. The 4080 SUPER has a base clock of 2295 MHz and a boost clock of 2550 MHz, while the 4090 has a base of 2235 MHz and a boost of 2520 MHz. This shows that the 4080 SUPER is running at a higher frequency, but it cannot compensate for the 4090’s massive core count advantage. The power profiles are also different, with the 4090 having a 450 W TDP and the 4080 SUPER at 320 W. Both require a single 16-pin power connector, but the 4090 mandates a larger 850 W PSU versus the 700 W recommended for the 4080 SUPER. The physical dimensions are similar, with both cards being triple-slot, but the 4080 SUPER is slightly longer at 310 mm versus the 4090’s 304 mm.

FAQ

Q: Which card is faster in 3DMark Steel Nomad?

A: The NVIDIA GeForce RTX 4090 is significantly faster, scoring 9223 compared to the RTX 4080 SUPER’s 6600, a 39.7% advantage.

Q: Is the RTX 4080 SUPER a good alternative for a smaller power supply?

A: Yes, the data indicates it is. The RTX 4080 SUPER has a 320 W TDP and a suggested PSU of 700 W, while the RTX 4090 has a 450 W TDP and requires an 850 W PSU.

Q: How does the memory configuration differ between the two cards?

A: The RTX 4090 has 24 GB of GDDR6X on a 384-bit bus, providing 1.01 TB/s of bandwidth. The RTX 4080 SUPER has 16 GB on a 256-bit bus, providing 736.3 GB/s.

Q: What is the performance gap in compute workloads?

A: In the Passmark GPU Compute test, the RTX 4090 scores 26613, which is 34.3% higher than the RTX 4080 SUPER’s 19822, showing a major lead in compute tasks.

Q: Are there any benchmarks where the RTX 4080 SUPER wins?

A: No. Based on the data, the RTX 4090 wins all ten head-to-head benchmarks, ranging from a 2.3% lead in Passmark G2D to a 39.7% lead in 3DMark Steel Nomad.

Q: Which card has a higher boost clock speed?

A: The RTX 4080 SUPER has a higher boost clock at 2550 MHz, compared to the RTX 4090’s 2520 MHz. However, this does not translate into a performance win for the 4080 SUPER.

Where Each One Wins

The RTX 4090 is the clear winner for users who prioritize maximum performance in the most demanding scenarios. Its 39.7% lead in 3DMark Steel Nomad makes it the top choice for high-refresh-rate 4K gaming and future DirectX 12 titles that can fully utilize its hardware. The 34.3% advantage in GPU compute makes it the superior option for 3D rendering, video editing, and other professional workloads that rely heavily on FP32 and tensor core performance. With 24 GB of memory and 1.01 TB/s of bandwidth, it is also better equipped for large datasets and high-resolution texture packs.

The RTX 4080 SUPER wins in the context of system requirements and efficiency. Its 320 W TDP and 700 W PSU recommendation make it far easier to integrate into existing systems without requiring a major power supply upgrade. While it trails the 4090 in every benchmark, its performance is still exceptionally high, and its average score of 54209 places it just 0.1% behind the original RTX 4080, indicating it is a very capable card in its own right. For a user who does not need the absolute peak performance and prefers a more manageable power footprint, the 4080 SUPER is the logical choice. It offers a more balanced package, sacrificing raw speed for lower system demands while still delivering a premium experience.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 4080 SUPER
RTX 4090
Core Specs
Shading Units
10,240
16,384 +60.0%
Shaders
10,240
16,384 +60.0%
TMUs
320
512 +60.0%
ROPs
112
176 +57.1%
SM Count
80
128 +60.0%
Clocks
Base Clock
2295 MHz
2235 MHz
Boost Clock
2550 MHz
2520 MHz
Memory Clock
1438 MHz 23 Gbps effective
1313 MHz 21 Gbps effective
Memory
Memory Size
16 GB
24 GB
VRAM (MB)
16,384
24,576 +50.0%
Memory Type
GDDR6X
GDDR6X
Memory Bus
256 bit
384 bit
Bandwidth
736.3 GB/s
1.01 TB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
64 MB
72 MB
Performance
Pixel Rate
285.6 GPixel/s
443.5 GPixel/s
Texture Rate
816.0 GTexel/s
1,290.2 GTexel/s
FP32 (TFLOPS)
52.22 TFLOPS
82.58 TFLOPS
FP64 (TFLOPS)
816.0 GFLOPS (1:64)
1,290.2 GFLOPS (1:64)
FP16 (TFLOPS)
52.22 TFLOPS (1:1)
82.58 TFLOPS (1:1)
AI/RT
RT Cores
80
128 +60.0%
Tensor Cores
320
512 +60.0%
Power
TDP
320 W
450 W
TDP (W)
320
450 +40.6%
Suggested PSU
700 W
850 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Ada Lovelace
Ada Lovelace
GPU Name
AD103
AD102
Generation
GeForce 40
GeForce 40
Process Size
5 nm
5 nm
Transistors
45,900 million
76,300 million
Die Size
379 mm²
609 mm²
Foundry
TSMC
TSMC
Density
121.1M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
8.9
Shader Model
6.9
6.8
Physical
Slot Width
Triple-slot
Triple-slot
Length
310 mm 12.2 inches
304 mm 12 inches
Height
140 mm 5.5 inches
137 mm 5.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Launch Price
999 USD
1,599 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 30
GeForce 30
Successor
GeForce 50
GeForce 50
View GeForce RTX 4080 SUPER Details View GeForce RTX 4090 Details