NVIDIA GeForce RTX 3090 Ti vs NVIDIA L40 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 3090 Ti

CORE STATE GA102
VRAM 24 GB
CLOCK SPEED 1860 MHz
TDP 450 W
BUS WIDTH 384 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2022
VS
NVIDIA
GEFORCE

L40

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2490 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
5,741
N/A
geekbench_opencl
174,441
330,926
geekbench_vulkan
215,633
237,295

Analysis: NVIDIA GeForce RTX 3090 Ti vs NVIDIA L40

The Verdict

The recorded benchmark data gives the NVIDIA L40 a clear overall advantage over the GeForce RTX 3090 Ti. Across the two head-to-head tests in the database, the L40 wins both, with an average benchmark score of 284,111 versus 131,938 for the RTX 3090 Ti. The L40 also sits in the 99th percentile of all GPUs, while the RTX 3090 Ti ranks in the 95th. That is a substantial gap, and it aligns with the architectural differences between the two cards.

For compute-heavy workloads, the L40 is the obvious pick. Its 48 GB of GDDR6 memory and 864.0 GB/s of bandwidth give it a strong capacity advantage over the RTX 3090 Ti's 24 GB of GDDR6X, which delivers 1.01 TB/s. The L40 also has more than double the shading units, texture units, and ray tracing cores, which translates directly into higher raw throughput in the database's OpenCL test. That test shows the L40 at 330,926 points, which is 89.7% ahead of the RTX 3090 Ti's 174,441.

However, the RTX 3090 Ti does have one notable edge: memory bandwidth. At 1.01 TB/s, it is about 17% faster than the L40's 864.0 GB/s. That can matter in bandwidth-limited scenarios, but the L40's far larger memory pool and higher compute throughput make it the better choice for most professional workloads, especially those that need to fit large datasets in VRAM.

For gaming or consumer use, the RTX 3090 Ti is the more conventional option. It is a GeForce product with a triple-slot design and HDMI output, and its launch MSRP was 1,999 USD. But the data does not favor it in raw compute. The L40 is built for servers, with a dual-slot form factor and four DisplayPort outputs. If the workload is compute-heavy and memory capacity is critical, the L40 wins. If the workload is more about raw bandwidth and the user is already in a GeForce ecosystem, the RTX 3090 Ti remains a capable fallback, but the benchmark results show it is clearly behind in overall compute performance.

Architecture Differences

The two cards come from different NVIDIA generations and foundries. The L40 uses the AD102 chip on the Ada Lovelace architecture, fabricated on a 5 nm process at TSMC. The RTX 3090 Ti uses the GA102 chip on the Ampere architecture, built on an 8 nm process at Samsung. This is a major node difference, and it shows in transistor density: the L40 packs 76,300 million transistors into a 609 mm² die, for a density of 125.3M per mm². The RTX 3090 Ti has 28,300 million transistors on a 628 mm² die, at 45.1M per mm². The L40's die is slightly smaller but holds nearly three times the transistors.

Core counts follow the same pattern. The L40 has 18,176 shading units, 568 texture mapping units, and 192 render output units. The RTX 3090 Ti has 10,752 shading units, 336 TMUs, and 112 ROPs. Ray tracing cores also differ: 142 on the L40 versus 84 on the RTX 3090 Ti, and tensor cores are 568 versus 336. This is not a minor refinement; it is a generational leap in compute resources.

Clock speeds tell a more nuanced story. The RTX 3090 Ti has a higher base clock at 1560 MHz versus 735 MHz for the L40, and the boost clocks are closer: 1860 MHz for the RTX 3090 Ti and 2490 MHz for the L40. The L40's boost clock is actually higher, which helps it overcome its lower base clock. Memory clocks also differ, with the L40 running at 2250 MHz (18 Gbps effective) and the RTX 3090 Ti at 1313 MHz (21 Gbps effective). The RTX 3090 Ti's higher effective memory speed, combined with GDDR6X, gives it the bandwidth advantage.

Power and physical design are also distinct. The L40 is rated at 300 W and fits a dual-slot width, while the RTX 3090 Ti draws 450 W and is triple-slot. The L40's suggested PSU is 700 W, the RTX 3090 Ti's is 850 W. Both use a single 16-pin power connector. The L40 is 267 mm long and 111 mm tall, while the RTX 3090 Ti is 336 mm long, 140 mm tall, and 61 mm wide. The L40 has four DisplayPort 1.4a outputs; the RTX 3090 Ti has one HDMI 2.1 and three DisplayPort 1.4a.

Both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The L40 is classified under the "Server Ada (Lxx)" generation, while the RTX 3090 Ti belongs to the GeForce 30 series. The L40's predecessor is Server Ampere and its successor is Server Hopper. The RTX 3090 Ti's predecessor is GeForce 20 and its successor is GeForce 40. Both are end-of-life, with the L40 released on October 12, 2022 and the RTX 3090 Ti on January 26, 2022.

FAQ

Q: Which card has more memory?

A: The NVIDIA L40 has 48 GB of GDDR6 memory, while the GeForce RTX 3090 Ti has 24 GB of GDDR6X. The L40 offers double the capacity.

Q: Which card is faster in the database's OpenCL benchmark?

A: The L40 scores 330,926 in Geekbench OpenCL, which is 89.7% higher than the RTX 3090 Ti's 174,441. The L40 wins that test by a wide margin.

Q: Does the RTX 3090 Ti have any advantage in memory bandwidth?

A: Yes, the RTX 3090 Ti delivers 1.01 TB/s of bandwidth, which is higher than the L40's 864.0 GB/s. The RTX 3090 Ti uses GDDR6X at 21 Gbps effective, while the L40 uses GDDR6 at 18 Gbps effective.

Q: What is the difference in power consumption?

A: The L40 has a 300 W TDP and a suggested PSU of 700 W. The RTX 3090 Ti has a 450 W TDP and a suggested PSU of 850 W. The RTX 3090 Ti requires more power.

Q: Which card is better for large compute workloads?

A: Based on the data, the L40 is better for compute. It has more shading units, tensor cores, and ray tracing cores, plus double the memory capacity. Its average benchmark score of 284,111 is far above the RTX 3090 Ti's 131,938.

Q: Do both cards support the same graphics APIs?

A: Yes, both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. They also both use a PCIe 4.0 x16 interface.

Specification Differences

| Specification | NVIDIA L40 | NVIDIA GeForce RTX 3090 Ti |

|---|---|---|

| Chip | AD102 | GA102 |

| Architecture | Ada Lovelace | Ampere |

| Generation | Server Ada (Lxx) | GeForce 30 |

| Process Node | 5 nm | 8 nm |

| Foundry | TSMC | Samsung |

| Transistors | 76,300 million | 28,300 million |

| Die Size | 609 mm² | 628 mm² |

| Transistor Density | 125.3M / mm² | 45.1M / mm² |

| Base Clock | 735 MHz | 1560 MHz |

| Boost Clock | 2490 MHz | 1860 MHz |

| Memory Clock | 2250 MHz (18 Gbps effective) | 1313 MHz (21 Gbps effective) |

| Memory Size | 48 GB | 24 GB |

| Memory Type | GDDR6 | GDDR6X |

| Memory Bus Width | 384 bit | 384 bit |

| Memory Bandwidth | 864.0 GB/s | 1.01 TB/s |

| Shading Units | 18,176 | 10,752 |

| TMUs | 568 | 336 |

| ROPs | 192 | 112 |

| RT Cores | 142 | 84 |

| Tensor Cores | 568 | 336 |

| Pixel Rate | 478.1 GPixel/s | 208.3 GPixel/s |

| Texture Rate | 1,414.3 GTexel/s | 625.0 GTexel/s |

| FP32 | 90.52 TFLOPS | 40.00 TFLOPS |

| FP16 | 90.52 TFLOPS (1:1) | 40.00 TFLOPS (1:1) |

| TDP | 300 W | 450 W |

| Slot Width | Dual-slot | Triple-slot |

| Power Connectors | 1x 16-pin | 1x 16-pin |

| Suggested PSU | 700 W | 850 W |

| Display Outputs | 4x DisplayPort 1.4a | 1x HDMI 2.1, 3x DisplayPort 1.4a |

| Length | 267 mm (10.5 inches) | 336 mm (13.2 inches) |

| Height | 111 mm (4.4 inches) | 140 mm (5.5 inches) |

| Width | Not specified | 61 mm (2.4 inches) |

| Release Date | 2022-10-12 | 2022-01-26 |

| Predecessor | Server Ampere | GeForce 20 |

| Successor | Server Hopper | GeForce 40 |

| Launch MSRP | Not specified | 1,999 USD |

Head-to-Head Benchmarks

The database records two head-to-head benchmark comparisons between the L40 and the RTX 3090 Ti. The L40 wins both.

The first test is Geekbench OpenCL. The L40 scores 330,926, and the RTX 3090 Ti scores 174,441. That is a delta of 89.7% in favor of the L40. This is the largest gap between the two cards and it reflects the L40's massive compute advantage. The L40 has 18,176 shading units versus 10,752, and its FP32 throughput is 90.52 TFLOPS versus 40.00 TFLOPS. The OpenCL benchmark appears to scale with raw compute resources, and the L40 has more than double the shading units and tensor cores.

The second test is Geekbench Vulkan. The L40 scores 237,295, and the RTX 3090 Ti scores 215,633. The delta here is 10% in favor of the L40. This is a much narrower margin. The Vulkan test may be more sensitive to driver overhead or memory bandwidth, and the RTX 3090 Ti's higher bandwidth of 1.01 TB/s helps it stay competitive. Still, the L40 pulls ahead, likely due to its higher boost clock of 2490 MHz and its larger core count.

In total, the L40 records 2 wins and the RTX 3090 Ti records 0 wins in the head-to-head section. The L40's average benchmark score of 284,111 is more than double the RTX 3090 Ti's 131,938. The nearest rivals to the L40 are the NVIDIA RTX 6000 Ada Generation at 287,237 (1.1% ahead), the NVIDIA L40S at 295,763 (3.9% ahead), the NVIDIA L20 at 251,147 (13.1% behind), and the AMD Instinct MI300X at 317,994 (10.7% ahead). The RTX 3090 Ti's nearest rivals are the NVIDIA L4 at 131,072 (0.7% behind), the NVIDIA RTX 4000 Ada Generation at 135,218 (2.4% ahead), the NVIDIA A10M at 135,230 (2.4% ahead), and the AMD Radeon PRO W6800 at 135,396 (2.6% ahead).

The data shows the L40 is not just faster than the RTX 3090 Ti; it is in a different performance class. The RTX 3090 Ti's average score places it near the L4 and RTX 4000 Ada Generation, while the L40 sits near the RTX 6000 Ada and L40S. The 89.7% OpenCL gap is the clearest demonstration of the architectural difference between Ada Lovelace and Ampere. Even in Vulkan, where the RTX 3090 Ti's bandwidth helps, the L40 still holds a 10% lead. For buyers choosing between these two, the benchmark results point firmly to the L40 for compute performance, with the RTX 3090 Ti's only consolation being its higher memory bandwidth and its status as a GeForce product with HDMI output.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 3090 Ti
L40
Core Specs
Shading Units
10,752
18,176 +69.0%
Shaders
10,752
18,176 +69.0%
TMUs
336
568 +69.0%
ROPs
112
192 +71.4%
SM Count
84
142 +69.0%
Clocks
Base Clock
1560 MHz
735 MHz
Boost Clock
1860 MHz
2490 MHz
Memory Clock
1313 MHz 21 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
24 GB
48 GB
VRAM (MB)
24,576
49,152 +100.0%
Memory Type
GDDR6X
GDDR6
Memory Bus
384 bit
384 bit
Bandwidth
1.01 TB/s
864.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
6 MB
96 MB
Performance
Pixel Rate
208.3 GPixel/s
478.1 GPixel/s
Texture Rate
625.0 GTexel/s
1,414.3 GTexel/s
FP32 (TFLOPS)
40.00 TFLOPS
90.52 TFLOPS
FP64 (TFLOPS)
625.0 GFLOPS (1:64)
1,414.3 GFLOPS (1:64)
FP16 (TFLOPS)
40.00 TFLOPS (1:1)
90.52 TFLOPS (1:1)
AI/RT
RT Cores
84
142 +69.0%
Tensor Cores
336
568 +69.0%
Power
TDP
450 W
300 W
TDP (W)
450
300 -33.3%
Suggested PSU
850 W
700 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Ampere
Ada Lovelace
GPU Name
GA102
AD102
Generation
GeForce 30
Server Ada (Lxx)
Process Size
8 nm
5 nm
Transistors
28,300 million
76,300 million
Die Size
628 mm²
609 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Triple-slot
Dual-slot
Length
336 mm 13.2 inches
267 mm 10.5 inches
Height
140 mm 5.5 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
4x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Launch Price
1,999 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 20
Server Ampere
Successor
GeForce 40
Server Hopper
View GeForce RTX 3090 Ti Details View L40 Details