NVIDIA GeForce RTX 5090 D vs NVIDIA L40S Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 5090 D

CORE STATE GB202
VRAM 32 GB
CLOCK SPEED 2407 MHz
TDP 575 W
BUS WIDTH 512 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
14,326
N/A
geekbench_opencl
310,674
330,727
geekbench_vulkan
376,915
260,799
passmark_directx_10
231
N/A
passmark_directx_11
371
N/A
passmark_directx_12
219
N/A
passmark_directx_9
434
N/A
passmark_g2d
1,487
N/A
passmark_g3d
44,065
N/A
passmark_gpu_compute
28,396
N/A

Analysis: NVIDIA GeForce RTX 5090 D vs NVIDIA L40S

Head-to-Head Benchmarks

The recorded data shows a split decision between the NVIDIA L40S and the NVIDIA GeForce RTX 5090 D, with each card claiming one decisive victory in the two common benchmark workloads. The L40S wins the OpenCL compute test, while the RTX 5090 D dominates the Vulkan graphics test by a significantly larger margin.

In Geekbench OpenCL, the L40S scores 330,727 points against 310,674 for the RTX 5090 D. That is a 6.5% advantage for the server-oriented card. This result aligns with the L40S's positioning near the top of the database's overall rankings, where it sits in the 99th percentile of all GPUs. Its average benchmark score of 295,763 places it ahead of the NVIDIA RTX 6000 Ada Generation (287,237, a 3% gap) and the NVIDIA L40 (284,111, a 4.1% gap). The RTX 5090 D, by contrast, sits in the 92nd percentile with an average score of 77,712, a figure that pulls it into a different competitive tier entirely.

The Vulkan result flips the script completely. Here the RTX 5090 D scores 376,915, which is 30.8% higher than the L40S's 260,799. That is a massive swing, more than four times larger in percentage terms than the L40S's OpenCL lead. The RTX 5090 D's Vulkan score is its strongest recorded benchmark, and it demonstrates that the Blackwell architecture's graphics pipeline is substantially more efficient in this API.

Looking at the broader rival context, the L40S's nearest competitors in the database include the AMD Instinct MI300X (317,994, trailing by 7%) and the NVIDIA H200 NVL (334,891, leading by 11.7%). These figures frame the L40S as a mid-pack compute leader, not an outlier. The RTX 5090 D's nearest rivals are entirely different, including the AMD Radeon RX 6650M XT (76,904, just 1.1% behind) and the AMD Radeon RX 6850M XT (78,940, leading by 1.6%). The massive disparity in average scores between the two cards, 295,763 versus 77,712, reflects the different benchmark suites each card was tested with, not necessarily a direct performance ranking.

The win count stands at one each. The L40S takes OpenCL, the RTX 5090 D takes Vulkan. But the margin asymmetry matters. A 6.5% win in one test and a 30.8% loss in the other means the RTX 5090 D's Vulkan advantage is the single most decisive data point in this comparison.

The Verdict

The data supports a clear verdict: the NVIDIA L40S is the compute-focused workhorse, while the NVIDIA GeForce RTX 5090 D is the graphics-oriented specialist. Users whose workloads rely on OpenCL compute should choose the L40S, as it delivers a 6.5% higher score in that test and carries 48 GB of GDDR6 memory versus 32 GB of GDDR7. The L40S also holds the 99th percentile ranking among all GPUs, meaning it sits at the very top of the database's performance distribution.

For Vulkan-based rendering or graphics workloads, the RTX 5090 D is the unambiguous pick. Its 30.8% lead in that benchmark is overwhelming, and its 1.79 TB/s memory bandwidth, enabled by a 512-bit bus, is nearly double the L40S's 864.0 GB/s. The RTX 5090 D also has more shading units (21,760 versus 18,176), more tensor cores (680 versus 568), and more RT cores (170 versus 142). These architectural advantages translate directly into the Vulkan result.

Neither card is a general-purpose winner. The database records one win apiece, and the choice depends entirely on the target API and workload type. The L40S is end-of-life production status, while the RTX 5090 D is active, which may influence availability but does not change the recorded performance figures. The L40S's launch MSRP is not listed, while the RTX 5090 D carries a launch MSRP of 2,299 USD.

Where Each One Wins

The L40S wins in OpenCL compute tasks. Its 330,727 score versus 310,674 demonstrates a 6.5% edge that is consistent with its server-oriented Ada Lovelace architecture. The card's 48 GB of GDDR6 memory on a 384-bit bus provides 864.0 GB/s of bandwidth, which is ample for large dataset processing. Its 91.61 TFLOPS of FP32 and FP16 performance, delivered at a 300 W TDP, makes it the efficiency pick for compute-heavy environments. The L40S also has a higher pixel rate at 483.8 GPixel/s, though its texture rate of 1,431.4 GTexel/s is lower than the RTX 5090 D's 1,636.8 GTexel/s.

The RTX 5090 D wins in Vulkan graphics and API-specific workloads. Its 376,915 score represents a 30.8% advantage, which is the largest delta in any shared benchmark. The Blackwell 2.0 architecture with 21,760 shading units and 680 TMUs powers this result. The card's 104.8 TFLOPS of FP32 and FP16 performance is 14.4% higher than the L40S's figures. Memory bandwidth is the other major differentiator: 1.79 TB/s from GDDR7 on a 512-bit bus eclipses the L40S's 864.0 GB/s. The RTX 5090 D also supports PCIe 5.0 x16, while the L40S is limited to PCIe 4.0 x16, which matters for data transfer in modern systems.

The RTX 5090 D also wins on raw transistor count and die size. It packs 92,200 million transistors on a 750 mm² die, compared to 76,300 million on 609 mm² for the L40S. The transistor density is slightly lower on the newer chip (122.9M per mm² versus 125.3M per mm²), but the absolute resource pool is larger.

FAQ

Q: Which card has higher FP32 compute performance?

A: The RTX 5090 D leads with 104.8 TFLOPS versus 91.61 TFLOPS for the L40S, a difference of 14.4%.

Q: What is the memory configuration difference?

A: The L40S has 48 GB of GDDR6 on a 384-bit bus with 864.0 GB/s bandwidth. The RTX 5090 D has 32 GB of GDDR7 on a 512-bit bus with 1.79 TB/s bandwidth.

Q: Which card ranks higher in the database's overall GPU percentile?

A: The L40S sits in the 99th percentile of all GPUs, while the RTX 5090 D sits in the 92nd percentile.

Q: What is the TDP difference between the two cards?

A: The L40S has a 300 W TDP, while the RTX 5090 D has a 575 W TDP. The suggested PSU ratings are 700 W and 950 W respectively.

Q: Do both cards support the same API versions?

A: Yes, both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

Q: What is the release timeline for these cards?

A: The L40S was released on October 12, 2022, and is end-of-life. The RTX 5090 D was released on January 29, 2025, and is active.

Architecture Differences

The two cards represent different NVIDIA architectures. The L40S uses the AD102 chip built on the Ada Lovelace architecture, while the RTX 5090 D uses the GB202 chip on the Blackwell 2.0 architecture. Both are fabricated by TSMC on a 5 nm process, but the transistor counts diverge substantially: 76,300 million for the L40S versus 92,200 million for the RTX 5090 D. Die sizes follow the same trend, with the L40S at 609 mm² and the RTX 5090 D at 750 mm². Transistor density is marginally higher on the older chip, 125.3M per mm² versus 122.9M per mm².

The processing resource distribution differs. The RTX 5090 D has 21,760 shading units, 680 TMUs, and 176 ROPs. The L40S has 18,176 shading units, 568 TMUs, and 192 ROPs. The RTX 5090 D leads in shading units and TMUs, but the L40S has more ROPs. RT core counts are 170 versus 142 in favor of the RTX 5090 D, and tensor cores are 680 versus 568 in the same direction.

Memory architecture is a major generational shift. The L40S uses GDDR6 at 2250 MHz (18 Gbps effective), while the RTX 5090 D uses GDDR7 at 1750 MHz (28 Gbps effective). The RTX 5090 D's 512-bit bus delivers 1.79 TB/s versus 864.0 GB/s for the L40S's 384-bit bus. The RTX 5090 D also supports PCIe 5.0 x16, a full generation ahead of the L40S's PCIe 4.0 x16.

Display outputs differ. The L40S offers 1x HDMI 2.1 and 3x DisplayPort 1.4a. The RTX 5090 D offers 1x HDMI 2.1b and 3x DisplayPort 2.1b. The newer card's DisplayPort support is a notable upgrade for high-refresh-rate displays.

Specification Differences

The core clock profiles differ significantly. The L40S has a base clock of 1110 MHz and a boost clock of 2520 MHz. The RTX 5090 D has a base clock of 2017 MHz and a boost clock of 2407 MHz. The RTX 5090 D runs much faster at base, but the L40S has a higher boost ceiling.

Memory size and type are direct contrasts: 48 GB GDDR6 for the L40S, 32 GB GDDR7 for the RTX 5090 D. Bus width is 384-bit versus 512-bit, and bandwidth is 864.0 GB/s versus 1.79 TB/s.

The pixel rate favors the L40S at 483.8 GPixel/s, while the texture rate favors the RTX 5090 D at 1,636.8 GTexel/s. FP32 and FP16 performance both favor the RTX 5090 D at 104.8 TFLOPS versus 91.61 TFLOPS for the L40S.

Power requirements are substantially higher for the RTX 5090 D: 575 W TDP versus 300 W, and a suggested PSU of 950 W versus 700 W. Both cards are dual-slot with a single 16-pin power connector.

Physical dimensions differ. The L40S measures 267 mm (10.5 inches) in length and 111 mm (4.4 inches) in height. The RTX 5090 D is larger at 304 mm (12 inches) in length, 137 mm (5.4 inches) in height, and 48 mm (1.9 inches) in width.

The bus interface is PCIe 4.0 x16 for the L40S and PCIe 5.0 x16 for the RTX 5090 D. Production status is end-of-life for the L40S and active for the RTX 5090 D. Release dates are October 12, 2022, for the L40S and January 29, 2025, for the RTX 5090 D. The L40S's predecessor is Server Ampere and its successor is Server Hopper. The RTX 5090 D's predecessor is GeForce 40 and its successor is GeForce 60. The RTX 5090 D has a listed launch MSRP of 2,299 USD, while the L40S has no listed MSRP.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 5090 D
L40S
Core Specs
Shading Units
21,760
18,176 -16.5%
Shaders
21,760
18,176 -16.5%
TMUs
680
568 -16.5%
ROPs
176
192 +9.1%
SM Count
170
142 -16.5%
Clocks
Base Clock
2017 MHz
1110 MHz
Boost Clock
2407 MHz
2520 MHz
Memory Clock
1750 MHz 28 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
32 GB
48 GB
VRAM (MB)
32,768
49,152 +50.0%
Memory Type
GDDR7
GDDR6
Memory Bus
512 bit
384 bit
Bandwidth
1.79 TB/s
864.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
96 MB
48 MB
Performance
Pixel Rate
423.6 GPixel/s
483.8 GPixel/s
Texture Rate
1,636.8 GTexel/s
1,431.4 GTexel/s
FP32 (TFLOPS)
104.8 TFLOPS
91.61 TFLOPS
FP64 (TFLOPS)
1.637 TFLOPS (1:64)
1,431.4 GFLOPS (1:64)
FP16 (TFLOPS)
104.8 TFLOPS (1:1)
91.61 TFLOPS (1:1)
AI/RT
RT Cores
170
142 -16.5%
Tensor Cores
680
568 -16.5%
Power
TDP
575 W
300 W
TDP (W)
575
300 -47.8%
Suggested PSU
950 W
700 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Blackwell 2.0
Ada Lovelace
GPU Name
GB202
AD102
Generation
GeForce 50
Server Ada (Lxx)
Process Size
5 nm
5 nm
Transistors
92,200 million
76,300 million
Die Size
750 mm²
609 mm²
Foundry
TSMC
TSMC
Density
122.9M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
12.0
8.9
Shader Model
6.9
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
304 mm 12 inches
267 mm 10.5 inches
Height
137 mm 5.4 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.1b3x DisplayPort 2.1b
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 5.0 x16
PCIe 4.0 x16
Other
Launch Price
2,299 USD
Production
Active
End-of-life
Predecessor
GeForce 40
Server Ampere
Successor
GeForce 60
Server Hopper
View GeForce RTX 5090 D Details View L40S Details