NVIDIA GeForce RTX 5090 vs NVIDIA L40 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 5090

CORE STATE GB202
VRAM 32 GB
CLOCK SPEED 2407 MHz
TDP 575 W
BUS WIDTH 512 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

L40

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2490 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
18,355
N/A
geekbench_opencl
334,370
330,926
geekbench_vulkan
376,728
237,295
passmark_directx_10
226
N/A
passmark_directx_11
341
N/A
passmark_directx_12
185
N/A
passmark_directx_9
395
N/A
passmark_g2d
1,413
N/A
passmark_g3d
39,650
N/A
passmark_gpu_compute
26,756
N/A

Analysis: NVIDIA GeForce RTX 5090 vs NVIDIA L40

Head-to-Head Benchmarks

The only two directly comparable benchmark results in the database are Geekbench OpenCL and Vulkan tests, and the RTX 5090 wins both. In Geekbench OpenCL, the RTX 5090 scores 334,370 against the L40’s 330,926, a margin of roughly 1% in favor of the newer card. That is a narrow gap, close enough that it could be considered a statistical tie in real-world workloads. The bigger divergence shows up in Vulkan: the RTX 5090 posts 376,728 versus the L40’s 237,295, a 37% advantage. That is a decisive swing, indicating that the RTX 5090 has a substantially stronger graphics API performance profile, particularly in Vulkan-based applications.

The L40’s average benchmark score across all recorded tests is 284,111, while the RTX 5090’s average is 79,842. That comparison is misleading, however, because the two cards are tested under different benchmark suites. The L40 only has Geekbench entries in the database, while the RTX 5090 has a broader set including Passmark and 3DMark tests. When looking at the shared tests, the RTX 5090 leads by 1% in OpenCL and 37% in Vulkan. The L40’s percentile ranking is 99, meaning it outperforms 99% of all GPUs in the database, while the RTX 5090 sits at the 92nd percentile. This discrepancy reflects the different benchmark pools each card is measured against, not necessarily a raw performance hierarchy.

Where Each One Wins

The RTX 5090 wins outright in both shared benchmarks. Its Vulkan score is the standout, with a 37% lead over the L40. That makes it the clear choice for workloads that rely heavily on Vulkan, such as certain game engines, compute frameworks, and emulation environments. The OpenCL result is closer, but the RTX 5090 still edges ahead. If a workload is OpenCL-bound, the difference is nearly negligible, so other factors like memory capacity or power draw would matter more.

The L40 does not win any of the head-to-head tests, but its strengths lie elsewhere. It has 48 GB of GDDR6 memory on a 384-bit bus, delivering 864.0 GB/s of bandwidth. The RTX 5090 has 32 GB of GDDR7 on a 512-bit bus, with 1.79 TB/s bandwidth. The L40’s larger memory pool is a clear advantage for models or datasets that exceed 32 GB. The RTX 5090’s bandwidth advantage is substantial, but the L40’s capacity allows it to handle larger working sets without spilling to system memory. For applications that are memory-capacity-bound rather than bandwidth-bound, the L40 wins.

Architecture Differences

The two GPUs come from different architectural generations. The L40 uses the AD102 chip based on Ada Lovelace, built on a 5 nm process at TSMC. The RTX 5090 uses the GB202 chip based on Blackwell 2.0, also on a 5 nm process at TSMC. The transistor counts differ significantly: the L40 has 76,300 million transistors on a 609 mm² die, while the RTX 5090 packs 92,200 million on a 750 mm² die. Transistor density is slightly higher on the L40 at 125.3 million per mm² versus 122.9 million per mm² on the RTX 5090. That means the L40 is a denser design, while the RTX 5090 uses its larger die to add more functional units.

The shading unit counts tell the story. The L40 has 18,176 shading units, 568 texture mapping units, and 192 ROPs. The RTX 5090 has 21,760 shading units, 680 TMUs, and 176 ROPs. The RTX 5090 has more shading and texture hardware, but fewer ROPs. The L40 compensates with a higher pixel rate of 478.1 GPixel/s versus 423.6 GPixel/s on the RTX 5090. The texture rate favors the RTX 5090 at 1,636.8 GTexel/s versus 1,414.3 GTexel/s. The RTX 5090 also has more RT cores (170 versus 142) and more tensor cores (680 versus 568). In raw FP32 throughput, the RTX 5090 delivers 104.8 TFLOPS against the L40’s 90.52 TFLOPS, a 16% advantage.

Clock behavior differs as well. The L40 has a base clock of 735 MHz and a boost of 2490 MHz, while the RTX 5090 runs at a much higher base of 2017 MHz and a boost of 2407 MHz. The L40’s low base clock suggests it is designed for sustained server workloads with power limits in mind. The RTX 5090’s high base clock indicates it can maintain high frequency under load without the same headroom. Memory technology also differs: the L40 uses GDDR6 at 18 Gbps effective, while the RTX 5090 uses GDDR7 at 28 Gbps effective. The bus width increases from 384-bit to 512-bit, and bandwidth jumps from 864.0 GB/s to 1.79 TB/s.

The Verdict

The data points to a clear split: the RTX 5090 is the faster card in shared benchmarks, with a 37% Vulkan lead and a 1% OpenCL edge. Its architecture is newer, with more shading units, tensor cores, and RT cores, plus higher clock speeds and memory bandwidth. For applications that prioritize raw compute throughput, particularly in Vulkan or FP32-heavy workloads, the RTX 5090 is the better choice.

The L40 is not without merit, but its advantages are narrower. It has more memory (48 GB versus 32 GB), a higher pixel rate, and a higher transistor density. It also draws significantly less power at 300 W versus 575 W, with a lower suggested PSU of 700 W versus 950 W. That makes it a more power-efficient option for server environments where density and cooling matter. The L40’s production status is end-of-life, while the RTX 5090 is active, which matters for procurement decisions. If a workload fits within 32 GB, the RTX 5090 is the stronger performer. If it needs more than 32 GB, the L40 is the only option that can handle it.

The RTX 5090’s launch MSRP is 1,999 USD. The L40 has no recorded launch MSRP in the database.

FAQ

Q: Which GPU performs better in Vulkan workloads?

A: The RTX 5090 leads by a significant margin. In Geekbench Vulkan, it scores 376,728 against the L40’s 237,295, a 37% advantage.

Q: How do they compare in OpenCL performance?

A: The RTX 5090 edges out the L40 with a score of 334,370 versus 330,926, a difference of about 1%.

Q: Which card has more memory and how does that affect use cases?

A: The L40 has 48 GB of GDDR6, while the RTX 5090 has 32 GB of GDDR7. The L40’s larger capacity allows it to handle datasets and models that exceed 32 GB, whereas the RTX 5090’s higher bandwidth (1.79 TB/s versus 864.0 GB/s) benefits workloads that are bandwidth-sensitive but fit within its memory limit.

Q: What are the power requirements for each card?

A: The L40 has a 300 W TDP and a suggested PSU of 700 W. The RTX 5090 has a 575 W TDP and a suggested PSU of 950 W. The L40 is substantially more power-efficient.

Q: Which card has more compute units?

A: The RTX 5090 has 21,760 shading units, 680 tensor cores, and 170 RT cores. The L40 has 18,176 shading units, 568 tensor cores, and 142 RT cores. The RTX 5090 also achieves higher FP32 throughput at 104.8 TFLOPS versus 90.52 TFLOPS.

Q: What is the production status of each card?

A: The L40 is marked as end-of-life, while the RTX 5090 is active. The L40 was released on 2022-10-12, and the RTX 5090 on 2025-01-29.

Specification Differences

| Field | NVIDIA L40 | NVIDIA GeForce RTX 5090 |

|---|---|---|

| Chip | AD102 | GB202 |

| Architecture | Ada Lovelace | Blackwell 2.0 |

| Process Node | 5 nm | 5 nm |

| Transistors | 76,300 million | 92,200 million |

| Die Size | 609 mm² | 750 mm² |

| Transistor Density | 125.3M / mm² | 122.9M / mm² |

| Base Clock | 735 MHz | 2017 MHz |

| Boost Clock | 2490 MHz | 2407 MHz |

| Memory Size | 48 GB | 32 GB |

| Memory Type | GDDR6 | GDDR7 |

| Memory Bus Width | 384 bit | 512 bit |

| Memory Bandwidth | 864.0 GB/s | 1.79 TB/s |

| Shading Units | 18176 | 21760 |

| TMUs | 568 | 680 |

| ROPs | 192 | 176 |

| RT Cores | 142 | 170 |

| Tensor Cores | 568 | 680 |

| Pixel Rate | 478.1 GPixel/s | 423.6 GPixel/s |

| Texture Rate | 1,414.3 GTexel/s | 1,636.8 GTexel/s |

| FP32 Performance | 90.52 TFLOPS | 104.8 TFLOPS |

| TDP | 300 W | 575 W |

| Suggested PSU | 700 W | 950 W |

| Bus Interface | PCIe 4.0 x16 | PCIe 5.0 x16 |

| Display Outputs | 4x DisplayPort 1.4a | 1x HDMI 2.1b, 3x DisplayPort 2.1b |

| Length | 267 mm (10.5 inches) | 304 mm (12 inches) |

| Height | 111 mm (4.4 inches) | 137 mm (5.4 inches) |

| Width | Not specified | 40 mm (1.6 inches) |

| Production Status | End-of-life | Active |

| Release Date | 2022-10-12 | 2025-01-29 |

| Predecessor | Server Ampere | GeForce 40 |

| Successor | Server Hopper | GeForce 60 |

| Launch MSRP | Not specified | 1,999 USD |

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 5090
L40
Core Specs
Shading Units
21,760
18,176 -16.5%
Shaders
21,760
18,176 -16.5%
TMUs
680
568 -16.5%
ROPs
176
192 +9.1%
SM Count
170
142 -16.5%
Clocks
Base Clock
2017 MHz
735 MHz
Boost Clock
2407 MHz
2490 MHz
Memory Clock
1750 MHz 28 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
32 GB
48 GB
VRAM (MB)
32,768
49,152 +50.0%
Memory Type
GDDR7
GDDR6
Memory Bus
512 bit
384 bit
Bandwidth
1.79 TB/s
864.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
96 MB
96 MB
Performance
Pixel Rate
423.6 GPixel/s
478.1 GPixel/s
Texture Rate
1,636.8 GTexel/s
1,414.3 GTexel/s
FP32 (TFLOPS)
104.8 TFLOPS
90.52 TFLOPS
FP64 (TFLOPS)
1.637 TFLOPS (1:64)
1,414.3 GFLOPS (1:64)
FP16 (TFLOPS)
104.8 TFLOPS (1:1)
90.52 TFLOPS (1:1)
AI/RT
RT Cores
170
142 -16.5%
Tensor Cores
680
568 -16.5%
Power
TDP
575 W
300 W
TDP (W)
575
300 -47.8%
Suggested PSU
950 W
700 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Blackwell 2.0
Ada Lovelace
GPU Name
GB202
AD102
Generation
GeForce 50
Server Ada (Lxx)
Process Size
5 nm
5 nm
Transistors
92,200 million
76,300 million
Die Size
750 mm²
609 mm²
Foundry
TSMC
TSMC
Density
122.9M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
12.0
8.9
Shader Model
6.9
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
304 mm 12 inches
267 mm 10.5 inches
Height
137 mm 5.4 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.1b3x DisplayPort 2.1b
4x DisplayPort 1.4a
Bus Interface
PCIe 5.0 x16
PCIe 4.0 x16
Other
Launch Price
1,999 USD
Production
Active
End-of-life
Predecessor
GeForce 40
Server Ampere
Successor
GeForce 60
Server Hopper
View GeForce RTX 5090 Details View L40 Details