NVIDIA GeForce RTX 5090 vs NVIDIA L40S Comparison
NVIDIA GeForce RTX 5090
L40S
PERFORMANCE BENCHMARKS
Analysis: NVIDIA GeForce RTX 5090 vs NVIDIA L40S
Where Each One Wins
The benchmark data splits these two cards cleanly along compute and graphics workloads. The NVIDIA L40S and NVIDIA GeForce RTX 5090 share only two direct head-to-head measurements, and the RTX 5090 wins both. However, the L40S occupies a different performance tier when viewed against its own nearest rivals, while the RTX 5090's average benchmark score sits in a much lower percentile due to the inclusion of lightweight legacy tests.
In the Geekbench OpenCL test, the RTX 5090 scores 334,370 against the L40S's 330,727, a narrow 1.1% margin. That result is essentially a tie, and it suggests that raw compute throughput in an OpenCL workload is comparable between the two. The L40S trails by only 3,643 points, which is within run-to-run variance for most benchmark suites. In contrast, the Geekbench Vulkan test shows a decisive RTX 5090 victory: 376,728 versus 260,799, a 30.8% gap. This is the largest differential in the entire comparison, and it points to the RTX 5090's substantial advantage in graphics API workloads.
The L40S, however, sits at the 99th percentile among all GPUs in the database, with an average benchmark score of 295,763. Its nearest rivals include the NVIDIA RTX 6000 Ada Generation at 287,237 (3% behind) and the NVIDIA L40 at 284,111 (4.1% behind). The L40S also outperforms the AMD Instinct MI300X, which scores 317,994, meaning the L40S is 7% ahead of that accelerator. Only the NVIDIA H200 NVL scores higher at 334,891, putting the L40S 11.7% behind that part.
The RTX 5090's average benchmark score of 79,842 is dragged down by several Passmark legacy tests: DirectX 9 at 395, DirectX 10 at 226, DirectX 11 at 341, and DirectX 12 at 185. These low scores pull the average far below the Geekbench results. The RTX 5090's nearest rivals are older data-center and mobile parts: the Tesla P100 PCIe 16 GB at 79,605 (0.3% behind), the Tesla P100 PCIe 12 GB at 79,396 (0.6% behind), and the AMD Radeon RX 6850M XT at 78,940 (1.1% behind). The AMD Radeon Pro Vega 64X sits 1.4% ahead at 80,959. This clustering shows the RTX 5090's average is not representative of its peak capability, and the Passmark G3D score of 39,650 and GPU Compute score of 26,756 are far more indicative of its real performance.
Architecture Differences
The L40S uses the AD102 chip on the Ada Lovelace architecture, fabricated by TSMC on a 5 nm process. It packs 76,300 million transistors into a 609 mm² die, resulting in a transistor density of 125.3 million per square millimeter. The RTX 5090 uses the GB202 chip on the Blackwell 2.0 architecture, also fabricated by TSMC on 5 nm, with 92,200 million transistors on a 750 mm² die, giving a density of 122.9 million per square millimeter. The L40S has a slightly higher transistor density despite being a smaller die, but the RTX 5090 has roughly 21% more transistors overall.
The L40S belongs to the Server Ada generation, while the RTX 5090 is part of the GeForce 50-series. The L40S is marked as end-of-life in the database, with a release date of October 12, 2022, and its predecessor is Server Ampere, successor Server Hopper. The RTX 5090 is active, released January 29, 2025, with the GeForce 40 series as predecessor and GeForce 60 as successor.
Core counts differ significantly. The L40S has 18,176 shading units, 568 texture mapping units, 192 raster operation units, 142 ray tracing cores, and 568 tensor cores. The RTX 5090 has 21,760 shading units, 680 TMUs, 176 ROPs, 170 RT cores, and 680 tensor cores. The RTX 5090 leads in shading units by 19.7%, in TMUs by 19.7%, and in RT cores by 19.7%. The L40S counters with 192 ROPs versus 176, a 9.1% advantage in pixel throughput.
Clock speeds also diverge. The L40S has a base clock of 1110 MHz and a boost clock of 2520 MHz. The RTX 5090's base clock is 2017 MHz, boost 2407 MHz. The L40S boosts 4.7% higher, but the RTX 5090's base clock is 81.7% higher, which matters for sustained workloads that do not reach boost frequencies.
The memory subsystems are fundamentally different. The L40S uses 48 GB of GDDR6 on a 384-bit bus with a bandwidth of 864.0 GB/s. The RTX 5090 uses 32 GB of GDDR7 on a 512-bit bus with 1.79 TB/s of bandwidth. The RTX 5090 has 107.2% more memory bandwidth, despite having 33.3% less capacity. The L40S's memory clock is 2250 MHz (18 Gbps effective), while the RTX 5090 runs at 1750 MHz (28 Gbps effective).
The bus interface differs: the L40S uses PCIe 4.0 x16, the RTX 5090 uses PCIe 5.0 x16. Power requirements also separate them. The L40S has a TDP of 300 W with a suggested PSU of 700 W. The RTX 5090 has a TDP of 575 W and a suggested PSU of 950 W. Both are dual-slot cards with a single 16-pin power connector. The L40S measures 267 mm in length and 111 mm in height; the RTX 5090 is 304 mm long, 137 mm high, and 40 mm wide.
FAQ
Q: Which card has more memory bandwidth?
A: The RTX 5090. Its GDDR7 memory on a 512-bit bus delivers 1.79 TB/s, while the L40S's GDDR6 on a 384-bit bus provides 864.0 GB/s. The RTX 5090 has 107.2% higher bandwidth.
Q: Does the L40S have more VRAM?
A: Yes. The L40S carries 48 GB of GDDR6, compared to 32 GB of GDDR7 on the RTX 5090. That is 50% more capacity, which matters for large model inference and rendering scenes that exceed 32 GB.
Q: Which card wins in Vulkan compute?
A: The RTX 5090 wins decisively. In the Geekbench Vulkan test, it scores 376,728 versus the L40S's 260,799, a 30.8% advantage. This is the largest single benchmark gap between the two.
Q: How close are they in OpenCL?
A: Nearly identical. The RTX 5090 scores 334,370 and the L40S scores 330,727 in Geekbench OpenCL, a 1.1% difference. This suggests the L40S remains competitive in compute-oriented OpenCL workloads.
Q: What is the transistor count difference?
A: The RTX 5090 has 92,200 million transistors on the GB202 chip, while the L40S has 76,300 million on the AD102 chip. The RTX 5090 has roughly 20.8% more transistors, though the L40S has a higher transistor density at 125.3 million per mm² versus 122.9 million per mm².
Q: Which card has a higher boost clock?
A: The L40S. Its boost clock is 2520 MHz, while the RTX 5090's is 2407 MHz. However, the RTX 5090's base clock is 2017 MHz versus 1110 MHz for the L40S.
Specification Differences
| Specification | NVIDIA L40S | NVIDIA GeForce RTX 5090 |
|---|---|---|
| Chip | AD102 | GB202 |
| Architecture | Ada Lovelace | Blackwell 2.0 |
| Generation | Server Ada (Lxx) | GeForce 50 |
| Process Node | 5 nm | 5 nm |
| Transistors | 76,300 million | 92,200 million |
| Die Size | 609 mm² | 750 mm² |
| Transistor Density | 125.3M / mm² | 122.9M / mm² |
| Base Clock | 1110 MHz | 2017 MHz |
| Boost Clock | 2520 MHz | 2407 MHz |
| Memory Clock | 2250 MHz (18 Gbps effective) | 1750 MHz (28 Gbps effective) |
| Memory Size | 48 GB | 32 GB |
| Memory Type | GDDR6 | GDDR7 |
| Memory Bus Width | 384 bit | 512 bit |
| Memory Bandwidth | 864.0 GB/s | 1.79 TB/s |
| Shading Units | 18,176 | 21,760 |
| TMUs | 568 | 680 |
| ROPs | 192 | 176 |
| RT Cores | 142 | 170 |
| Tensor Cores | 568 | 680 |
| Pixel Rate | 483.8 GPixel/s | 423.6 GPixel/s |
| Texture Rate | 1,431.4 GTexel/s | 1,636.8 GTexel/s |
| FP32 | 91.61 TFLOPS | 104.8 TFLOPS |
| FP16 | 91.61 TFLOPS (1:1) | 104.8 TFLOPS (1:1) |
| TDP | 300 W | 575 W |
| Suggested PSU | 700 W | 950 W |
| Bus Interface | PCIe 4.0 x16 | PCIe 5.0 x16 |
| Length | 267 mm (10.5 inches) | 304 mm (12 inches) |
| Height | 111 mm (4.4 inches) | 137 mm (5.4 inches) |
| Width | Not specified | 40 mm (1.6 inches) |
| Production Status | End-of-life | Active |
| Release Date | 2022-10-12 | 2025-01-29 |
Head-to-Head Benchmarks
The database contains only two direct comparisons between the L40S and RTX 5090, both from Geekbench. The first, OpenCL, is a near dead heat. The RTX 5090 posts 334,370, and the L40S follows with 330,727. The delta is 1.1%, which puts the L40S within a rounding error of the newer card. For compute workloads that use OpenCL, an end-of-life server card from 2022 essentially matches a 2025 flagship in this test.
The second test, Vulkan, is not close. The RTX 5090 scores 376,728, while the L40S manages 260,799. That is a 30.8% deficit for the L40S. Vulkan exercises the graphics pipeline more directly, and the RTX 5090's 21,760 shading units, 680 TMUs, and 170 RT cores combine with its 1.79 TB/s memory bandwidth to deliver a substantial edge. The L40S's higher pixel rate of 483.8 GPixel/s versus 423.6 GPixel/s does not compensate for the raw shader throughput difference.
The RTX 5090 also holds a clear lead in FP32 compute. Its 104.8 TFLOPS versus 91.61 TFLOPS for the L40S is a 14.4% advantage, and the same ratio applies to FP16 because both cards run 1:1. The texture rate favors the RTX 5090 as well: 1,636.8 GTexel/s versus 1,431.4 GTexel/s, a 14.4% margin. The L40S counters with a 14.2% higher pixel rate, but that metric matters more for rasterization-bound scenes than for compute.
The RTX 5090's additional benchmark results add context. Its Passmark G3D score of 39,650 and GPU Compute score of 26,756 are far above its legacy DirectX scores, which range from 185 to 395. Those legacy tests are not relevant to modern workloads but they pull the average down to 79,842, placing the RTX 5090 at the 92nd percentile. The L40S, with only two benchmark entries, sits at the 99th percentile with an average of 295,763. This percentile gap is misleading: it reflects the different test suites each card was subjected to, not a raw performance disparity.
In the two tests where the cards directly compete, the RTX 5090 wins both, but the OpenCL result is effectively a tie. The Vulkan result is a decisive victory. When the full specification sheet is considered, the RTX 5090 has more shaders, more RT cores, more tensor cores, higher memory bandwidth, and a higher FP32 throughput. The L40S offers more VRAM, higher pixel rate, higher boost clock, and a lower power draw. For workloads that fit within 32 GB and favor Vulkan or modern graphics APIs, the RTX 5090 is the clear choice. For large-memory compute tasks that rely on OpenCL and need the capacity, the L40S remains competitive, but it cannot match the RTX 5090's peak graphics performance.