GPU Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 3090 Ti

CORE STATE GA102
VRAM 24 GB
CLOCK SPEED 1860 MHz
TDP 450 W
BUS WIDTH 384 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2022
VS
NVIDIA
GEFORCE

L4

CORE STATE AD104
VRAM 24 GB
CLOCK SPEED 2040 MHz
TDP 72 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
5,741
N/A
geekbench_opencl
168,491
140,838
geekbench_vulkan
215,633
121,306

Analysis: NVIDIA GeForce RTX 3090 Ti vs NVIDIA L4

The NVIDIA GeForce RTX 3090 Ti and NVIDIA L4 are both 24 GB GPUs from NVIDIA, but they occupy opposite ends of the design spectrum: the RTX 3090 Ti is a high-power desktop compute card built on Ampere, while the L4 is a low-power server accelerator on Ada Lovelace. Benchmark data from Geekbench shows the RTX 3090 Ti leading decisively in raw compute, yet the L4 counters with drastically lower power consumption, a compact single-slot form factor, and an active production status. The following analysis draws exclusively from the provided benchmark and specification data to break down where each card stands.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The RTX 3090 Ti scores 131,911 on average, while the L4 scores 128,665. That puts the RTX 3090 Ti 2.5% ahead of the L4, according to the nearest-rival delta.

Q: What are the memory configurations?

A: Both have 24 GB, but the RTX 3090 Ti uses GDDR6X on a 384-bit bus with 1.01 TB/s bandwidth, whereas the L4 uses GDDR6 on a 192-bit bus with 300.1 GB/s.

Q: How do their power requirements differ?

A: The RTX 3090 Ti has a 450 W TDP and requires an 850 W suggested PSU, while the L4 has a 72 W TDP and only needs a 250 W suggested PSU.

Q: Which card can output video?

A: The RTX 3090 Ti has 1x HDMI 2.1 and 3x DisplayPort 1.4a outputs; the L4 has no display outputs at all.

Q: What are the process nodes and foundries?

A: The RTX 3090 Ti is built on Samsung's 8 nm process, while the L4 uses TSMC's 5 nm process.

Q: When were they released?

A: The RTX 3090 Ti launched on January 26, 2022, and the L4 followed on March 20, 2023.

Architecture Differences

The two GPUs are built on different architectures and manufacturing processes. The RTX 3090 Ti uses the GA102 chip on Ampere architecture, fabricated on Samsung's 8 nm node. It packs 28,300 million transistors on a 628 mm² die, giving a transistor density of 45.1 million per mm². The L4 uses the AD104 chip on Ada Lovelace, made by TSMC on a 5 nm process. It contains 35,800 million transistors on a much smaller 294 mm² die, achieving a far higher density of 121.8 million per mm².

Core counts differ significantly. The RTX 3090 Ti has 10,752 shading units, 336 TMUs, 112 ROPs, 84 RT cores, and 336 tensor cores. The L4, despite the newer architecture, has fewer: 7,424 shading units, 240 TMUs, 80 ROPs, 60 RT cores, and 240 tensor cores. Clock behavior also diverges. The RTX 3090 Ti runs at a base of 1560 MHz and boosts to 1860 MHz. The L4 has a much lower base clock of 795 MHz but a higher boost of 2040 MHz, reflecting its power-constrained design.

Memory architecture is another major split. The RTX 3090 Ti uses 24 GB of GDDR6X across a 384-bit bus, delivering 1.01 TB/s of bandwidth. The L4 also has 24 GB, but it is GDDR6 on a 192-bit bus, yielding only 300.1 GB/s. The memory clock is 1313 MHz (21 Gbps effective) on the RTX 3090 Ti versus 1563 MHz (12.5 Gbps effective) on the L4.

Power and physical design reinforce their different roles. The RTX 3090 Ti draws 450 W, requires a triple-slot cooler, a 1x 16-pin power connector, and measures 336 mm in length. The L4 is a single-slot card with no power connector, a 72 W TDP, and a 169 mm length. The RTX 3090 Ti offers display outputs; the L4 has none. Both support PCIe 4.0 x16 and the same API set (DirectX 12 Ultimate, OpenGL 4.6, Vulkan 1.4). The RTX 3090 Ti is end-of-life, while the L4 is active production. The RTX 3090 Ti's predecessor is GeForce 20 and successor GeForce 40; the L4's predecessor is Server Ampere and successor Server Hopper.

Head-to-Head Benchmarks

In the only two direct benchmark comparisons available, the RTX 3090 Ti wins both by wide margins. In Geekbench OpenCL, the RTX 3090 Ti scores 205,978 against the L4's 140,838 — a 46.3% advantage. In Geekbench Vulkan, the RTX 3090 Ti scores 184,015 versus 116,491, a 58% lead. These are not marginal differences; they show the RTX 3090 Ti delivering roughly 1.5 to 1.6 times the compute throughput in these tests.

The average benchmark scores tell a similar but more muted story. The RTX 3090 Ti averages 131,911, while the L4 averages 128,665. The nearest-rival data confirms this: the RTX 3090 Ti is 2.5% ahead of the L4, and the L4 is 2.5% behind the RTX 3090 Ti. Both cards sit in the 97th percentile of all GPUs, meaning each is among the top 3% of performers overall, but the RTX 3090 Ti holds a consistent edge.

It is worth noting that the RTX 3090 Ti has three benchmark entries (including 3DMark Steel Nomad DX12 with a score of 5,741), while the L4 only has two (both Geekbench tests). The inclusion of that extra 3DMark test, which the L4 does not have, could influence the average, but the head-to-head results are unambiguous.

The Verdict

The data points to a clear split: the RTX 3090 Ti is the faster compute card, but the L4 is the more efficient and deployment-friendly option. For anyone prioritizing raw performance in OpenCL or Vulkan workloads, the RTX 3090 Ti is the obvious choice — it leads by 46–58% in the direct comparisons and holds a 2.5% higher average score. Its 40.00 TFLOPS FP32 and 40.00 TFLOPS FP16 (1:1) exceed the L4's 30.29 TFLOPS in both. The RTX 3090 Ti also offers higher pixel and texture rates: 208.3 GPixel/s versus 163.2, and 625.0 GTexel/s versus 489.6.

However, the L4 wins on power and form factor. At 72 W TDP, it uses just 16% of the RTX 3090 Ti's 450 W envelope. It is a single-slot, 169 mm card with no power connector, making it far easier to install in dense server environments. The L4 is also actively produced, while the RTX 3090 Ti is end-of-life. The L4's newer 5 nm process and higher transistor density (121.8M/mm² vs 45.1M/mm²) indicate a more modern design, but that efficiency does not translate into compute superiority in these benchmarks.

For a workstation or desktop user needing maximum compute throughput and display outputs, the RTX 3090 Ti is the data-supported pick. For a power-constrained server deployment where compute density per watt matters more than raw speed, the L4 is the sensible choice. The RTX 3090 Ti had a launch MSRP of 1,999 USD; the L4 has no listed launch MSRP.

Specification Differences

The following table lists the fields where the two GPUs differ, based on the provided data.

| Field | RTX 3090 Ti | L4 |

|-------|-------------|----|

| Chip | GA102 | AD104 |

| Architecture | Ampere | Ada Lovelace |

| Process Node | 8 nm | 5 nm |

| Foundry | Samsung | TSMC |

| Transistors | 28,300 million | 35,800 million |

| Die Size | 628 mm² | 294 mm² |

| Transistor Density | 45.1M / mm² | 121.8M / mm² |

| Base Clock | 1560 MHz | 795 MHz |

| Boost Clock | 1860 MHz | 2040 MHz |

| Memory Clock | 1313 MHz (21 Gbps effective) | 1563 MHz (12.5 Gbps effective) |

| Memory Type | GDDR6X | GDDR6 |

| Bus Width | 384 bit | 192 bit |

| Bandwidth | 1.01 TB/s | 300.1 GB/s |

| Shading Units | 10752 | 7424 |

| TMUs | 336 | 240 |

| ROPs | 112 | 80 |

| RT Cores | 84 | 60 |

| Tensor Cores | 336 | 240 |

| Pixel Rate | 208.3 GPixel/s | 163.2 GPixel/s |

| Texture Rate | 625.0 GTexel/s | 489.6 GTexel/s |

| FP32 | 40.00 TFLOPS | 30.29 TFLOPS |

| FP16 | 40.00 TFLOPS (1:1) | 30.29 TFLOPS (1:1) |

| TDP | 450 W | 72 W |

| Slot Width | Triple-slot | Single-slot |

| Power Connectors | 1x 16-pin | None |

| Suggested PSU | 850 W | 250 W |

| Display Outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a | No outputs |

| Dimensions | 336 mm x 140 mm x 61 mm | 169 mm x 56 mm (width not listed) |

| Production Status | End-of-life | Active |

| Release Date | 2022-01-26 | 2023-03-20 |

| Predecessor | GeForce 20 | Server Ampere |

| Successor | GeForce 40 | Server Hopper |

Where Each One Wins

The RTX 3090 Ti wins on raw compute. It is ahead in both Geekbench OpenCL (46.3%) and Vulkan (58%) tests. Its higher shading unit count (10,752 vs 7,424), larger memory bus (384-bit vs 192-bit), and greater bandwidth (1.01 TB/s vs 300.1 GB/s) underpin these wins. It also has more RT cores (84 vs 60) and tensor cores (336 vs 240), making it the stronger choice for any workload that leverages those units. Its 40.00 TFLOPS FP32 and FP16 outputs double the L4's 30.29 TFLOPS in the same precision.

The L4 wins on efficiency and deployment. Its 72 W TDP is a fraction of the RTX 3090 Ti's 450 W, and it requires no external power connector. The single-slot, 169 mm length makes it suitable for space-constrained servers. It is actively produced, unlike the end-of-life RTX 3090 Ti. The L4's higher boost clock (2040 MHz vs 1860 MHz) and denser transistor packing (121.8M/mm² vs 45.1M/mm²) show a more modern design, but these advantages do not translate into benchmark victories in the available data. For applications where power draw, physical size, or production availability are the primary constraints, the L4 is the clear winner. For applications where compute speed is the sole criterion, the RTX 3090 Ti leads without contest.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 3090 Ti
L4
Core Specs
Shading Units
10,752
7,424 -31.0%
Shaders
10,752
7,424 -31.0%
TMUs
336
240 -28.6%
ROPs
112
80 -28.6%
SM Count
84
60 -28.6%
Clocks
Base Clock
1560 MHz
795 MHz
Boost Clock
1860 MHz
2040 MHz
Memory Clock
1313 MHz 21 Gbps effective
1563 MHz 12.5 Gbps effective
Memory
Memory Size
24 GB
24 GB
VRAM (MB)
24,576
24,576 0.0%
Memory Type
GDDR6X
GDDR6
Memory Bus
384 bit
192 bit
Bandwidth
1.01 TB/s
300.1 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
6 MB
48 MB
Performance
Pixel Rate
208.3 GPixel/s
163.2 GPixel/s
Texture Rate
625.0 GTexel/s
489.6 GTexel/s
FP32 (TFLOPS)
40.00 TFLOPS
30.29 TFLOPS
FP64 (TFLOPS)
625.0 GFLOPS (1:64)
473.3 GFLOPS (1:64)
FP16 (TFLOPS)
40.00 TFLOPS (1:1)
30.29 TFLOPS (1:1)
AI/RT
RT Cores
84
60 -28.6%
Tensor Cores
336
240 -28.6%
Power
TDP
450 W
72 W
TDP (W)
450
72 -84.0%
Suggested PSU
850 W
250 W
Power Connectors
1x 16-pin
None
Architecture
Architecture
Ampere
Ada Lovelace
GPU Name
GA102
AD104
Generation
GeForce 30
Server Ada (Lxx)
Process Size
8 nm
5 nm
Transistors
28,300 million
35,800 million
Die Size
628 mm²
294 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
121.8M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Triple-slot
Single-slot
Length
336 mm 13.2 inches
169 mm 6.7 inches
Height
140 mm 5.5 inches
56 mm 2.2 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Launch Price
1,999 USD
Production
End-of-life
Active
Predecessor
GeForce 20
Server Ampere
Successor
GeForce 40
Server Hopper
View GeForce RTX 3090 Ti Details View L4 Details