NVIDIA GeForce RTX 5090 D vs NVIDIA L40 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 5090 D

CORE STATE GB202
VRAM 32 GB
CLOCK SPEED 2407 MHz
TDP 575 W
BUS WIDTH 512 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

L40

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2490 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
14,326
N/A
geekbench_opencl
310,674
330,926
geekbench_vulkan
376,915
237,295
passmark_directx_10
231
N/A
passmark_directx_11
371
N/A
passmark_directx_12
219
N/A
passmark_directx_9
434
N/A
passmark_g2d
1,487
N/A
passmark_g3d
44,065
N/A
passmark_gpu_compute
28,396
N/A

Analysis: NVIDIA GeForce RTX 5090 D vs NVIDIA L40

Head-to-Head Benchmarks

The recorded data shows a split decision between these two NVIDIA accelerators, with each claiming one victory in the shared benchmark suite. In Geekbench OpenCL, the NVIDIA L40 scores 330,926 against the GeForce RTX 5090 D's 310,674, giving the L40 a 6.5% advantage. This result aligns with the L40's positioning in the database, where its average benchmark score of 284,111 places it in the 99th percentile of all GPUs. The RTX 5090 D, by contrast, holds a 92nd percentile ranking with an average score of 77,712, though that figure is heavily influenced by its broader test set including older DirectX and compute workloads.

The Geekbench Vulkan test flips the outcome decisively. The RTX 5090 D posts 376,915 points, while the L40 manages 237,295, a 37% margin in favor of the newer card. This is a substantial gap, one that suggests the Blackwell architecture's Vulkan driver stack and hardware scheduling deliver far better results in this particular API. The L40's Vulkan score is notably lower than its OpenCL result, indicating the Ada Lovelace design may prioritize compute APIs differently. For the RTX 5090 D, the Vulkan result is its strongest recorded benchmark, exceeding its OpenCL score by over 21%.

What the head-to-head table does not capture is the RTX 5090 D's additional benchmark coverage. The database includes ten tests for the GeForce card, spanning 3DMark Steel Nomad DX12, multiple Passmark DirectX versions, Passmark G2D, G3D, and GPU compute, plus the two Geekbench tests. In 3DMark Steel Nomad DX12, the RTX 5090 D scores 14,326, a result that has no direct L40 counterpart but indicates strong modern DirectX 12 performance. Passmark G3D shows 44,065 points, while GPU compute reaches 28,396. The DirectX 9 score of 434 and DirectX 11 score of 371 from Passmark suggest legacy API efficiency, though these figures are not directly comparable to the L40's limited test set.

Where Each One Wins

The L40 demonstrates its strength in OpenCL compute workloads, a domain where its 48 GB memory configuration and server-oriented design shine. The 6.5% OpenCL lead over the RTX 5090 D, combined with its 99th percentile ranking against all GPUs, positions it as the preferred choice for general-purpose compute tasks that rely on this API. The L40's average benchmark score of 284,111 also sits close to its nearest rivals: the RTX 6000 Ada Generation at 287,237 (1.1% lower), the L40S at 295,763 (3.9% higher), and the L20 at 251,147 (13.1% below). This cluster of scores indicates the L40 operates in a high-performance server tier where small percentage differences separate top contenders.

The RTX 5090 D wins decisively in Vulkan, making it the better option for graphics-oriented workloads that leverage this modern cross-platform API. The 37% margin over the L40 is the largest performance gap recorded in any shared test. Beyond Vulkan, the GeForce card's additional benchmark results paint a picture of a versatile graphics processor. Its Passmark DirectX 12 score of 219, while lower than DirectX 11's 371, still reflects capable modern API support. The G2D score of 1,487 indicates strong 2D rendering throughput, useful for desktop and interface workloads. The RTX 5090 D's 92nd percentile ranking, though lower than the L40's 99th, must be weighed against its broader test coverage, which includes legacy APIs that drag down the average.

For users prioritizing raw compute in OpenCL, the L40 is the data-backed choice. For those needing Vulkan performance or a wider range of DirectX capabilities, the RTX 5090 D holds the advantage. The L40's nearest rival comparisons reinforce its compute focus: it trails the AMD Instinct MI300X by 10.7% in average score (317,994 versus 284,111) but leads the L20 by 13.1%, showing a clear hierarchy within NVIDIA's server lineup. The RTX 5090 D's nearest rivals are quite different, with the AMD Radeon RX 6650M XT just 1.1% behind (76,904), the RX 6850M XT 1.6% ahead (78,940), and Tesla P100 variants trailing by 2.1% and 2.4%. This places the GeForce card in a mid-range performance band relative to its peers, despite its high-end feature set.

Architecture Differences

The two GPUs come from different architectural generations. The L40 is built on Ada Lovelace, using the AD102 chip, while the RTX 5090 D employs Blackwell 2.0 with the GB202 chip. Both are fabricated on a 5 nm process at TSMC, but the transistor counts diverge significantly. The L40 packs 76,300 million transistors on a 609 mm² die, yielding a density of 125.3 million transistors per square millimeter. The RTX 5090 D scales up to 92,200 million transistors across a larger 750 mm² die, though its density drops slightly to 122.9 million per square millimeter. This means the Blackwell chip adds roughly 16 billion transistors while increasing die area by about 23%, a trade-off that favors raw scale over density.

Memory architecture presents a clear generational shift. The L40 uses 48 GB of GDDR6 on a 384-bit bus, delivering 864.0 GB/s of bandwidth. The RTX 5090 D steps up to 32 GB of GDDR7 on a 512-bit bus, more than doubling bandwidth to 1.79 TB/s. The memory type change from GDDR6 to GDDR7 is significant: the newer standard operates at 28 Gbps effective, while the L40's GDDR6 runs at 18 Gbps. Despite having 16 GB less capacity, the RTX 5090 D's wider bus and faster memory give it a substantial bandwidth advantage, which likely contributes to its Vulkan and DirectX performance gains.

Compute resources differ in both count and configuration. The L40 features 18,176 shading units, 568 texture mapping units, 192 ROPs, 142 ray tracing cores, and 568 tensor cores. The RTX 5090 D increases shading units to 21,760, TMUs to 680, RT cores to 170, and tensor cores to 680, but reduces ROPs to 176. The higher shader and tensor core counts on the Blackwell card explain its FP32 and FP16 throughput of 104.8 TFLOPS, compared to the L40's 90.52 TFLOPS. The pixel rate tells a different story: the L40 achieves 478.1 GPixel/s versus the RTX 5090 D's 423.6 GPixel/s, a function of the L40's higher ROP count. Texture rate favors the newer card, 1,636.8 GTexel/s versus 1,414.3 GTexel/s, driven by its additional TMUs.

Clock speeds reveal another divergence. The L40 has a low base clock of 735 MHz but boosts to 2,490 MHz, a 3.4x multiplier. The RTX 5090 D starts at 2,017 MHz base and boosts to 2,407 MHz, a much smaller 1.19x ratio. The L40's boosting behavior suggests a power-adaptive design that idles low but ramps aggressively under load. The RTX 5090 D's higher base clock indicates sustained performance without relying on boost headroom. Power consumption scales accordingly: the L40 is rated at 300 W TDP with a 700 W suggested PSU, while the RTX 5090 D draws 575 W TDP and recommends a 950 W PSU. Both use a single 16-pin power connector.

Connectivity and output options also differ. The L40 uses PCIe 4.0 x16, while the RTX 5090 D upgrades to PCIe 5.0 x16, doubling the interface bandwidth potential. Display outputs show the L40's server focus: four DisplayPort 1.4a ports. The RTX 5090 D offers one HDMI 2.1b and three DisplayPort 2.1b ports, supporting the latest consumer display standards. Physical dimensions vary as well: the L40 measures 267 mm long and 111 mm tall, while the RTX 5090 D is longer at 304 mm, taller at 137 mm, and adds a width dimension of 48 mm, both being dual-slot cards.

FAQ

Q: Which GPU has higher raw FP32 compute performance?

A: The RTX 5090 D leads with 104.8 TFLOPS FP32, compared to the L40's 90.52 TFLOPS, a 15.8% advantage for the Blackwell card.

Q: How do their memory bandwidth figures compare?

A: The RTX 5090 D delivers 1.79 TB/s from its 512-bit GDDR7 interface, while the L40 provides 864.0 GB/s over a 384-bit GDDR6 bus. The newer card offers roughly 107% more bandwidth.

Q: What is the L40's advantage in the OpenCL benchmark?

A: The L40 scores 330,926 in Geekbench OpenCL versus 310,674 for the RTX 5090 D, a 6.5% lead that reflects its compute-oriented design and larger memory capacity.

Q: Why does the RTX 5090 D win so decisively in Vulkan?

A: The RTX 5090 D scores 376,915 in Geekbench Vulkan, 37% higher than the L40's 237,295. This likely stems from the Blackwell architecture's newer Vulkan driver optimizations and higher shader core count.

Q: Which card has more memory?

A: The L40 has 48 GB of GDDR6, while the RTX 5090 D has 32 GB of GDDR7. The L40's larger capacity suits memory-intensive workloads, though the RTX 5090 D's memory is faster.

Q: What are their transistor counts and die sizes?

A: The L40 contains 76,300 million transistors on a 609 mm² die, while the RTX 5090 D has 92,200 million transistors on a 750 mm² die. The RTX 5090 D adds about 20.8% more transistors despite a 23% larger die area.

Specification Differences

| Specification | NVIDIA L40 | NVIDIA GeForce RTX 5090 D |

|---|---|---|

| Chip | AD102 | GB202 |

| Architecture | Ada Lovelace | Blackwell 2.0 |

| Process Node | 5 nm | 5 nm |

| Transistors | 76,300 million | 92,200 million |

| Die Size | 609 mm² | 750 mm² |

| Transistor Density | 125.3M / mm² | 122.9M / mm² |

| Base Clock | 735 MHz | 2017 MHz |

| Boost Clock | 2490 MHz | 2407 MHz |

| Memory Size | 48 GB | 32 GB |

| Memory Type | GDDR6 | GDDR7 |

| Memory Bus Width | 384 bit | 512 bit |

| Memory Bandwidth | 864.0 GB/s | 1.79 TB/s |

| Memory Clock | 2250 MHz, 18 Gbps effective | 1750 MHz, 28 Gbps effective |

| Shading Units | 18,176 | 21,760 |

| TMUs | 568 | 680 |

| ROPs | 192 | 176 |

| RT Cores | 142 | 170 |

| Tensor Cores | 568 | 680 |

| Pixel Rate | 478.1 GPixel/s | 423.6 GPixel/s |

| Texture Rate | 1,414.3 GTexel/s | 1,636.8 GTexel/s |

| FP32 | 90.52 TFLOPS | 104.8 TFLOPS |

| FP16 | 90.52 TFLOPS (1:1) | 104.8 TFLOPS (1:1) |

| TDP | 300 W | 575 W |

| Suggested PSU | 700 W | 950 W |

| Bus Interface | PCIe 4.0 x16 | PCIe 5.0 x16 |

| Display Outputs | 4x DisplayPort 1.4a | 1x HDMI 2.1b, 3x DisplayPort 2.1b |

| Dimensions | 267 mm x 111 mm | 304 mm x 137 mm x 48 mm |

| Production Status | End-of-life | Active |

| Release Date | 2022-10-12 | 2025-01-29 |

| Predecessor | Server Ampere | GeForce 40 |

| Successor | Server Hopper | GeForce 60 |

| Launch MSRP | Not available | 2,299 USD |

The specification table highlights the generational divide. The RTX 5090 D is newer, active, and carries a launch MSRP of 2,299 USD, while the L40 is end-of-life with no listed launch MSRP. The L40's advantages lie in memory capacity, ROP count, pixel rate, and lower power draw. The RTX 5090 D counters with higher shader and tensor core counts, faster memory, greater bandwidth, higher FP32 throughput, and a newer PCIe interface. Both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, indicating parity in API feature levels despite architectural differences. The L40's release date of October 2022 places it firmly in the Ada generation, while the RTX 5090 D's January 2025 launch marks the Blackwell era's consumer debut.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 5090 D
L40
Core Specs
Shading Units
21,760
18,176 -16.5%
Shaders
21,760
18,176 -16.5%
TMUs
680
568 -16.5%
ROPs
176
192 +9.1%
SM Count
170
142 -16.5%
Clocks
Base Clock
2017 MHz
735 MHz
Boost Clock
2407 MHz
2490 MHz
Memory Clock
1750 MHz 28 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
32 GB
48 GB
VRAM (MB)
32,768
49,152 +50.0%
Memory Type
GDDR7
GDDR6
Memory Bus
512 bit
384 bit
Bandwidth
1.79 TB/s
864.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
96 MB
96 MB
Performance
Pixel Rate
423.6 GPixel/s
478.1 GPixel/s
Texture Rate
1,636.8 GTexel/s
1,414.3 GTexel/s
FP32 (TFLOPS)
104.8 TFLOPS
90.52 TFLOPS
FP64 (TFLOPS)
1.637 TFLOPS (1:64)
1,414.3 GFLOPS (1:64)
FP16 (TFLOPS)
104.8 TFLOPS (1:1)
90.52 TFLOPS (1:1)
AI/RT
RT Cores
170
142 -16.5%
Tensor Cores
680
568 -16.5%
Power
TDP
575 W
300 W
TDP (W)
575
300 -47.8%
Suggested PSU
950 W
700 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Blackwell 2.0
Ada Lovelace
GPU Name
GB202
AD102
Generation
GeForce 50
Server Ada (Lxx)
Process Size
5 nm
5 nm
Transistors
92,200 million
76,300 million
Die Size
750 mm²
609 mm²
Foundry
TSMC
TSMC
Density
122.9M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
12.0
8.9
Shader Model
6.9
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
304 mm 12 inches
267 mm 10.5 inches
Height
137 mm 5.4 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.1b3x DisplayPort 2.1b
4x DisplayPort 1.4a
Bus Interface
PCIe 5.0 x16
PCIe 4.0 x16
Other
Launch Price
2,299 USD
Production
Active
End-of-life
Predecessor
GeForce 40
Server Ampere
Successor
GeForce 60
Server Hopper
View GeForce RTX 5090 D Details View L40 Details