NVIDIA GeForce RTX 4090 D vs NVIDIA L40 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 4090 D

CORE STATE AD102
VRAM 24 GB
CLOCK SPEED 2520 MHz
TDP 425 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

L40

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2490 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
8,587
N/A
geekbench_opencl
278,621
330,926
geekbench_vulkan
246,941
237,295

Analysis: NVIDIA GeForce RTX 4090 D vs NVIDIA L40

The NVIDIA L40 and NVIDIA GeForce RTX 4090 D are both built on the same AD102 chip and Ada Lovelace architecture, yet they target entirely different workloads. The L40 is a server-class compute card with 48 GB of memory, while the RTX 4090 D is a desktop GeForce part with 24 GB. Benchmark data from the FACT PACK shows a split decision: the L40 wins one test decisively, and the RTX 4090 D wins the other by a narrow margin. This page breaks down those results, the architectural differences, and what each card is best suited for according to the numbers.

Head-to-Head Benchmarks

The only two shared benchmarks in the data are Geekbench OpenCL and Geekbench Vulkan, and they tell opposite stories. In Geekbench OpenCL, the NVIDIA L40 scores 330,926 against the RTX 4090 D's 278,621. That is an 18.8% advantage for the L40, a substantial margin that reflects its compute-oriented design. OpenCL is a general-purpose compute API, and the L40's higher shading unit count (18,176 vs. 14,592) and higher FP32 throughput (90.52 TFLOPS vs. 73.54 TFLOPS) align with this result.

The Geekbench Vulkan test flips the outcome. The RTX 4090 D scores 246,941, while the L40 trails at 237,295. The delta is only 3.9% in favor of the RTX 4090 D, a much tighter gap than the OpenCL difference. Vulkan is a graphics-centric API, and the RTX 4090 D's higher boost clock (2520 MHz vs. 2490 MHz) and its GeForce-tuned driver stack likely contribute to this win. Still, the margin is small enough that it does not suggest a major gaming advantage—it merely shows the RTX 4090 D is competitive in this specific test.

Looking at the broader benchmark context, the L40's average benchmark score is 284,111, while the RTX 4090 D's is 178,050. That gap is large, but it is important to note that the RTX 4090 D's average includes a third benchmark (3DMark Steel Nomad DX12 with a score of 8,587) that the L40 does not have in its results. The L40 sits at the 99th percentile among all GPUs, while the RTX 4090 D is at the 98th percentile—both are elite performers, but the L40 edges ahead in overall standing.

In terms of nearest rivals, the L40's average score of 284,111 is 1.1% below the NVIDIA RTX 6000 Ada Generation (287,237) and 3.9% below the NVIDIA L40S (295,763). Conversely, it is 13.1% above the NVIDIA L20 (251,147) and 10.7% below the AMD Instinct MI300X (317,994). The RTX 4090 D's average of 178,050 is 2.2% below the NVIDIA RTX PRO 5000 Blackwell (182,109), 3.1% below the NVIDIA A100 SXM4 80 GB (183,725), 3.6% below the NVIDIA RTX 5000 Ada Generation (184,664), and 4.9% below the NVIDIA A100 SXM4 40 GB (187,147). These figures place the RTX 4090 D in a lower tier than the L40 when averaged across all tests, despite its Vulkan win.

Architecture Differences

Both cards use the AD102 chip, fabricated by TSMC on a 5 nm process, with 76,300 million transistors on a 609 mm² die. The transistor density is identical at 125.3M per mm². The differences begin with the memory subsystem. The L40 ships with 48 GB of GDDR6 memory, while the RTX 4090 D has 24 GB of GDDR6X. Both use a 384-bit bus, but the memory types and effective speeds differ. The L40's memory runs at 2250 MHz with 18 Gbps effective, yielding 864.0 GB/s of bandwidth. The RTX 4090 D's memory runs at 1313 MHz with 21 Gbps effective, yielding 1.01 TB/s of bandwidth. The RTX 4090 D has a clear bandwidth advantage despite having half the capacity.

The compute resources are heavily skewed toward the L40. It has 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The RTX 4090 D has 14,592 shading units, 456 TMUs, 176 ROPs, 114 RT cores, and 456 tensor cores. The L40 leads in every computational category, which explains its higher FP32 and FP16 performance: both are 90.52 TFLOPS for the L40, versus 73.54 TFLOPS for the RTX 4090 D. Pixel rate and texture rate follow the same pattern—the L40 hits 478.1 GPixel/s and 1,414.3 GTexel/s, while the RTX 4090 D reaches 443.5 GPixel/s and 1,149.1 GTexel/s.

Clock speeds are where the RTX 4090 D fights back. Its base clock is 2280 MHz and boost clock is 2520 MHz, compared to the L40's 735 MHz base and 2490 MHz boost. The RTX 4090 D starts at a much higher frequency, though the boost clocks are close. Power consumption also differs significantly: the L40 is rated at 300 W TDP, while the RTX 4090 D is 425 W. The L40 is a dual-slot card at 267 mm long and 111 mm tall, while the RTX 4090 D is a triple-slot card at 304 mm long, 137 mm tall, and 61 mm wide. Both use a single 16-pin power connector, but the suggested PSU is 700 W for the L40 and 800 W for the RTX 4090 D.

The display outputs differ as well. The L40 has four DisplayPort 1.4a outputs. The RTX 4090 D has one HDMI 2.1 and three DisplayPort 1.4a outputs. Both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The L40 is part of the Server Ada (Lxx) generation and was released on 2022-10-12, while the RTX 4090 D is from the GeForce 40 series and released on 2023-12-27. The L40's predecessor is Server Ampere and its successor is Server Hopper; the RTX 4090 D's predecessor is GeForce 30 and its successor is GeForce 50. Both are marked as end-of-life.

Where Each One Wins

The L40 wins decisively in general-purpose compute. Its 18.8% OpenCL advantage over the RTX 4090 D is backed by superior raw specs: 18,176 shading units, 568 tensor cores, and 90.52 TFLOPS of FP32 performance. The 48 GB of GDDR6 memory also gives it a capacity edge, which is critical for large datasets, AI model training, or rendering scenes that exceed 24 GB. The L40's average benchmark score of 284,111 places it at the 99th percentile, and it outperforms the RTX 4090 D's average score by roughly 60% when comparing the two averages directly. In a server context, the L40 is also more power-efficient per unit of compute, with a 300 W TDP versus the RTX 4090 D's 425 W, and it fits in a dual-slot form factor.

The RTX 4090 D wins in the Vulkan graphics test, albeit by a narrow 3.9% margin. Its higher boost clock (2520 MHz) and faster memory bandwidth (1.01 TB/s vs. 864.0 GB/s) give it an edge in latency-sensitive or graphics-heavy workloads. The RTX 4090 D also has a more consumer-friendly output configuration with an HDMI 2.1 port, which is absent on the L40. Its smaller memory capacity (24 GB) is a limitation for large compute tasks, but for gaming or real-time graphics, that capacity is rarely a bottleneck. The launch MSRP of the RTX 4090 D is 1,599 USD, a fact that can be noted once, though this analysis will not discuss value.

In the nearest-rival context, the L40 sits close to the RTX 6000 Ada Generation (1.1% below) and L40S (3.9% below), while the RTX 4090 D sits below several datacenter parts like the A100 SXM4 variants. This suggests the L40 is positioned as a professional compute card that competes with other workstation GPUs, while the RTX 4090 D is a GeForce product that falls short of datacenter-class averages but still ranks in the 98th percentile.

The Verdict

The data indicates a clear split: the NVIDIA L40 is the stronger choice for compute-intensive workloads, while the NVIDIA GeForce RTX 4090 D is better suited for graphics-oriented tasks. The L40's 18.8% OpenCL win is significant, and its 48 GB memory capacity is double that of the RTX 4090 D. Its lower TDP (300 W vs. 425 W) and dual-slot design make it more suitable for dense server environments. The L40's average benchmark score of 284,111 at the 99th percentile places it among the top GPUs, and its performance relative to rivals like the RTX 6000 Ada Generation (1.1% lower) and L40S (3.9% lower) confirms it is a high-tier compute product.

The RTX 4090 D, by contrast, wins the Vulkan test by 3.9%, but that margin is small. Its average score of 178,050 at the 98th percentile is still excellent, but it trails the L40 by a wide margin in the overall average. The RTX 4090 D's higher memory bandwidth (1.01 TB/s) and boost clock (2520 MHz) give it a niche in graphics-heavy applications, and its HDMI 2.1 output is a practical advantage for consumer displays. However, its 24 GB memory capacity and 425 W TDP are drawbacks for anyone needing large memory pools or efficient power usage.

Who should pick which? The L40 is for users who prioritize compute throughput, large memory capacity, and power efficiency—think AI inference, scientific computing, or rendering large scenes. The RTX 4090 D is for users who need a high-end GeForce card with competitive graphics performance and a more consumer-friendly feature set, but they should accept lower compute scores and higher power draw. The data does not support the RTX 4090 D as a compute replacement for the L40, nor the L40 as a graphics replacement for the RTX 4090 D. They are different tools, and the benchmarks reflect that.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA L40 has an average benchmark score of 284,111, while the NVIDIA GeForce RTX 4090 D has an average of 178,050.

Q: How much faster is the L40 in Geekbench OpenCL?

A: The L40 scores 330,926 compared to the RTX 4090 D's 278,621, a delta of 18.8% in favor of the L40.

Q: Does the RTX 4090 D win any benchmark?

A: Yes, it wins Geekbench Vulkan with a score of 246,941 against the L40's 237,295, a 3.9% advantage.

Q: What is the memory capacity difference?

A: The L40 has 48 GB of GDDR6 memory, while the RTX 4090 D has 24 GB of GDDR6X memory.

Q: Which card has more shading units?

A: The L40 has 18,176 shading units, while the RTX 4090 D has 14,592.

Q: What is the TDP of each card?

A: The L40 has a TDP of 300 W, and the RTX 4090 D has a TDP of 425 W.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 4090 D
L40
Core Specs
Shading Units
14,592
18,176 +24.6%
Shaders
14,592
18,176 +24.6%
TMUs
456
568 +24.6%
ROPs
176
192 +9.1%
SM Count
114
142 +24.6%
Clocks
Base Clock
2280 MHz
735 MHz
Boost Clock
2520 MHz
2490 MHz
Memory Clock
1313 MHz 21 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
24 GB
48 GB
VRAM (MB)
24,576
49,152 +100.0%
Memory Type
GDDR6X
GDDR6
Memory Bus
384 bit
384 bit
Bandwidth
1.01 TB/s
864.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
72 MB
96 MB
Performance
Pixel Rate
443.5 GPixel/s
478.1 GPixel/s
Texture Rate
1,149.1 GTexel/s
1,414.3 GTexel/s
FP32 (TFLOPS)
73.54 TFLOPS
90.52 TFLOPS
FP64 (TFLOPS)
1,149.1 GFLOPS (1:64)
1,414.3 GFLOPS (1:64)
FP16 (TFLOPS)
73.54 TFLOPS (1:1)
90.52 TFLOPS (1:1)
AI/RT
RT Cores
114
142 +24.6%
Tensor Cores
456
568 +24.6%
Power
TDP
425 W
300 W
TDP (W)
425
300 -29.4%
Suggested PSU
800 W
700 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Ada Lovelace
Ada Lovelace
GPU Name
AD102
AD102
Generation
GeForce 40
Server Ada (Lxx)
Process Size
5 nm
5 nm
Transistors
76,300 million
76,300 million
Die Size
609 mm²
609 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Triple-slot
Dual-slot
Length
304 mm 12 inches
267 mm 10.5 inches
Height
137 mm 5.4 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
4x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Launch Price
1,599 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 30
Server Ampere
Successor
GeForce 50
Server Hopper
View GeForce RTX 4090 D Details View L40 Details