NVIDIA GeForce RTX 3090 Ti vs NVIDIA L40S Comparison
NVIDIA GeForce RTX 3090 Ti
L40S
PERFORMANCE BENCHMARKS
Analysis: NVIDIA GeForce RTX 3090 Ti vs NVIDIA L40S
The Verdict
The benchmark database clearly separates these two NVIDIA cards. The NVIDIA L40S wins both recorded head-to-head tests, with an 89.6% lead in Geekbench OpenCL and a 20.9% lead in Geekbench Vulkan. Its average benchmark score of 295,763 places it at the 99th percentile among all GPUs, while the RTX 3090 Ti sits at the 95th percentile with an average of 131,938.
The L40S is the choice for compute-heavy workloads where raw FP32 throughput and large memory capacity matter. Its nearest rivals are the RTX 6000 Ada Generation (3% behind), the L40 (4.1% behind), the Instinct MI300X (7% ahead), and the H200 NVL (11.7% ahead), so it sits in a competitive server-class bracket. The RTX 3090 Ti, by contrast, competes with workstation cards like the L4 (0.7% behind), RTX 4000 Ada Generation (2.4% ahead), A10M (2.4% ahead), and Radeon PRO W6800 (2.6% ahead), marking it as a lower-tier performer in the same database.
For users seeking a dual-purpose card that can handle both rendering and general compute, the RTX 3090 Ti still holds relevance, but the data shows it trails in every shared benchmark. The L40S is the stronger pick for AI inference, scientific simulation, and any workload that benefits from 48 GB of GDDR6 memory. The RTX 3090 Ti is the pick only when the smaller 24 GB footprint and lower transistor count are acceptable, and even then, its performance ceiling is measurably lower.
Architecture Differences
The L40S uses the AD102 chip on a 5 nm TSMC process, while the RTX 3090 Ti uses the GA102 chip on an 8 nm Samsung process. This node difference is substantial: the L40S packs 76,300 million transistors into a 609 mm² die, yielding a transistor density of 125.3 million per mm². The RTX 3090 Ti has 28,300 million transistors on a 628 mm² die, with a density of 45.1 million per mm². The L40S achieves more than double the transistor density despite a slightly smaller die.
The L40S belongs to the Ada Lovelace generation (Server Ada class), while the RTX 3090 Ti is from the Ampere generation (GeForce 30 series). Both share the same PCIe 4.0 x16 bus interface, the same display outputs (1x HDMI 2.1 and 3x DisplayPort 1.4a), and identical API support: DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.
Core counts differ sharply. The L40S has 18,176 shading units, 568 texture mapping units, 192 ROPs, 142 ray tracing cores, and 568 tensor cores. The RTX 3090 Ti has 10,752 shading units, 336 TMUs, 112 ROPs, 84 RT cores, and 336 tensor cores. The L40S nearly doubles the shading units and RT cores, which directly explains its large FP32 advantage.
Memory architecture also diverges. The L40S uses 48 GB of GDDR6 on a 384-bit bus with 864.0 GB/s bandwidth. The RTX 3090 Ti uses 24 GB of GDDR6X on the same 384-bit bus but achieves 1.01 TB/s bandwidth. The L40S trades bandwidth for capacity, a sensible trade for large models that do not fit in 24 GB. Clock speeds differ too: the L40S has a base of 1110 MHz and boost of 2520 MHz, while the RTX 3090 Ti starts at 1560 MHz and boosts to 1860 MHz. The L40S compensates for its lower base clock with a much higher boost ceiling.
Power characteristics are not part of the benchmark scores, but the recorded TDP figures show the L40S at 300 W versus 450 W for the RTX 3090 Ti. The L40S also requires a 700 W suggested PSU versus 850 W for the RTX 3090 Ti. Both use a single 16-pin power connector.
Head-to-Head Benchmarks
The Geekbench OpenCL test shows the largest gap. The L40S scores 330,727 against 174,441 for the RTX 3090 Ti, a delta of 89.6%. This is nearly double the performance, consistent with the L40S having 18,176 shading units versus 10,752, and 91.61 TFLOPS FP32 against 40.00 TFLOPS. OpenCL workloads that scale with raw shader count and FP32 throughput will see the L40S dominate.
Geekbench Vulkan narrows the gap but still favors the L40S. The L40S scores 260,799 versus 215,633, a 20.9% delta. Vulkan performance is less dependent on raw FP32 and more sensitive to driver overhead and memory latency. The RTX 3090 Ti's higher memory bandwidth (1.01 TB/s versus 864.0 GB/s) may help close the gap, but the L40S still wins decisively.
The database records 2 wins for the L40S and 0 for the RTX 3090 Ti. No benchmark in the shared set favors the older card. The closest margin is Vulkan at 20.9%, which is still a comfortable lead. The OpenCL result is a rout. Users should interpret these numbers as representative of general compute performance, not gaming-specific workloads, since neither card carries gaming-focused benchmark entries in this comparison.
Specification Differences
| Specification | NVIDIA L40S | NVIDIA GeForce RTX 3090 Ti |
|---|---|---|
| Architecture | Ada Lovelace | Ampere |
| Generation | Server Ada (Lxx) | GeForce 30 |
| Process node | 5 nm | 8 nm |
| Foundry | TSMC | Samsung |
| Transistors | 76,300 million | 28,300 million |
| Die size | 609 mm² | 628 mm² |
| Transistor density | 125.3M / mm² | 45.1M / mm² |
| Base clock | 1110 MHz | 1560 MHz |
| Boost clock | 2520 MHz | 1860 MHz |
| Memory size | 48 GB | 24 GB |
| Memory type | GDDR6 | GDDR6X |
| Memory bandwidth | 864.0 GB/s | 1.01 TB/s |
| Shading units | 18,176 | 10,752 |
| TMUs | 568 | 336 |
| ROPs | 192 | 112 |
| RT cores | 142 | 84 |
| Tensor cores | 568 | 336 |
| Pixel rate | 483.8 GPixel/s | 208.3 GPixel/s |
| Texture rate | 1,431.4 GTexel/s | 625.0 GTexel/s |
| FP32 | 91.61 TFLOPS | 40.00 TFLOPS |
| FP16 | 91.61 TFLOPS (1:1) | 40.00 TFLOPS (1:1) |
| TDP | 300 W | 450 W |
| Slot width | Dual-slot | Triple-slot |
| Suggested PSU | 700 W | 850 W |
| Length | 267 mm | 336 mm |
| Height | 111 mm | 140 mm |
| Width | Not recorded | 61 mm |
| Release date | 2022-10-12 | 2022-01-26 |
| Predecessor | Server Ampere | GeForce 20 |
| Successor | Server Hopper | GeForce 40 |
| Production status | End-of-life | End-of-life |
| Launch MSRP | Not recorded | 1,999 USD |
The key differences beyond raw compute are memory capacity, memory type, and physical dimensions. The L40S is shorter (267 mm versus 336 mm) and shorter in height (111 mm versus 140 mm), and it fits in a dual-slot design versus the RTX 3090 Ti's triple-slot. The L40S also has a lower TDP and PSU requirement, which matters for dense server deployments.
FAQ
Q: Which card has more memory?
A: The NVIDIA L40S has 48 GB of GDDR6, while the RTX 3090 Ti has 24 GB of GDDR6X. Both use a 384-bit memory bus.
Q: How much faster is the L40S in OpenCL workloads?
A: The L40S scores 330,727 in Geekbench OpenCL versus 174,441 for the RTX 3090 Ti, a delta of 89.6%.
Q: Does the RTX 3090 Ti win any benchmark in this comparison?
A: No. The database records 2 wins for the L40S and 0 for the RTX 3090 Ti across the shared tests.
Q: What is the FP32 compute difference?
A: The L40S delivers 91.61 TFLOPS FP32, while the RTX 3090 Ti delivers 40.00 TFLOPS. The L40S is more than double.
Q: Are both cards still in production?
A: No. Both are marked as end-of-life in the database. The L40S released on 2022-10-12 and the RTX 3090 Ti on 2022-01-26.
Q: Which card has higher memory bandwidth?
A: The RTX 3090 Ti has 1.01 TB/s bandwidth due to GDDR6X memory, while the L40S has 864.0 GB/s with GDDR6. The L40S compensates with twice the capacity.
Where Each One Wins
The L40S wins in every measured category. Its 89.6% OpenCL advantage makes it the clear choice for compute-heavy applications like machine learning training, scientific simulation, and large-scale data processing. The 48 GB memory capacity allows models and datasets that would exceed the RTX 3090 Ti's 24 GB limit. The 91.61 TFLOPS FP32 throughput is more than double the RTX 3090 Ti's 40.00 TFLOPS, and the pixel rate of 483.8 GPixel/s versus 208.3 GPixel/s shows superiority in rasterization-heavy tasks as well. Texture rate follows the same pattern: 1,431.4 GTexel/s versus 625.0 GTexel/s.
The RTX 3090 Ti wins in no recorded benchmark, but it retains two technical advantages from the specification sheet. Its memory bandwidth of 1.01 TB/s exceeds the L40S's 864.0 GB/s, which can benefit certain bandwidth-bound workloads that fit within 24 GB. Its base clock of 1560 MHz is higher than the L40S's 1110 MHz, though the L40S's boost clock of 2520 MHz far exceeds the 3090 Ti's 1860 MHz. For users who already own the RTX 3090 Ti and have workloads under 24 GB, the card remains functional, but the data does not show any scenario where it outperforms the L40S.
The use-case split is therefore clear. Choose the L40S for AI model training, large-batch inference, rendering with heavy scene complexity, or any workload that needs more than 24 GB of memory. Choose the RTX 3090 Ti only if the lower transistor budget, smaller memory footprint, and higher power draw are acceptable constraints, and even then, expect lower performance across the board. The RTX 3090 Ti's 95th percentile ranking is respectable, but the L40S at the 99th percentile occupies a different performance tier entirely.