NVIDIA GeForce RTX 4090 D vs NVIDIA L40S Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 4090 D

CORE STATE AD102
VRAM 24 GB
CLOCK SPEED 2520 MHz
TDP 425 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
8,587
N/A
geekbench_opencl
278,621
330,727
geekbench_vulkan
246,941
260,799

Analysis: NVIDIA GeForce RTX 4090 D vs NVIDIA L40S

The NVIDIA L40S and the NVIDIA GeForce RTX 4090 D are both built on the same AD102 chip and Ada Lovelace architecture, yet they are engineered for entirely different roles. The L40S is a server-focused accelerator with 48 GB of GDDR6 memory, while the RTX 4090 D is a consumer-oriented card with 24 GB of GDDR6X. Both are now end-of-life products, but their benchmark data reveals a clear performance hierarchy. The L40S wins both head-to-head tests, with an 18.7% lead in Geekbench OpenCL and a 5.6% lead in Geekbench Vulkan over the RTX 4090 D. However, the RTX 4090 D has a significantly higher boost clock and a lower average benchmark score, making the choice between them dependent on workload type rather than raw speed alone.

The Verdict

The data points to a split decision based on workload requirements, not a universal winner. The NVIDIA L40S is the superior choice for compute-intensive, memory-hungry server tasks, as evidenced by its 18.7% advantage in Geekbench OpenCL (330,727 vs 278,621) and its 5.6% lead in Geekbench Vulkan (260,799 vs 246,941). Its 48 GB of GDDR6 memory is double the RTX 4090 D's 24 GB, making it the pick for large datasets and AI inference workloads where memory capacity is the bottleneck.

Conversely, the RTX 4090 D is the better fit for high-frequency, latency-sensitive rendering or gaming scenarios, despite losing both head-to-head benchmarks. Its base clock of 2280 MHz is more than double the L40S's 1110 MHz, and it achieves a 1.01 TB/s memory bandwidth versus the L40S's 864.0 GB/s. These specs suggest the RTX 4090 D can sustain higher instantaneous throughput in short bursts, even if its aggregate compute scores are lower.

The percentile data reinforces this split. The L40S sits at the 99th percentile of all GPUs with an average benchmark score of 295,763, while the RTX 4090 D is at the 98th percentile with an average score of 178,050. The L40S's nearest rival, the AMD Instinct MI300X, beats it by 7%, but the L40S is 4.1% faster than the NVIDIA L40. The RTX 4090 D, meanwhile, trails the NVIDIA RTX PRO 5000 Blackwell by 2.2% and the NVIDIA A100 SXM4 80 GB by 3.1%. For buyers, the L40S is the data-center workhorse, while the RTX 4090 D is a high-end desktop part with a launch MSRP of 1,599 USD — stated once here for reference.

Architecture Differences

Both GPUs share the same fundamental architecture: a 5 nm TSMC process, 76,300 million transistors, and a 609 mm² die size, yielding a transistor density of 125.3M per mm². They also share the same AD102 chip and support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The differences appear in their execution resources and memory subsystems.

The L40S deploys 18,176 shading units, 568 TMUs, and 192 ROPs, alongside 142 RT cores and 568 tensor cores. The RTX 4090 D, by contrast, has 14,592 shading units, 456 TMUs, and 176 ROPs, with 114 RT cores and 456 tensor cores. This gives the L40S a 24.5% advantage in shading units and a 24.5% advantage in tensor cores — a direct contributor to its superior FP32 and FP16 performance of 91.61 TFLOPS each, versus the RTX 4090 D's 73.54 TFLOPS for both.

Memory architecture diverges sharply. The L40S uses 48 GB of GDDR6 on a 384-bit bus, running at 18 Gbps effective, for a bandwidth of 864.0 GB/s. The RTX 4090 D uses 24 GB of GDDR6X on the same 384-bit bus, but at 21 Gbps effective, yielding a bandwidth of 1.01 TB/s. The GDDR6X memory is faster per pin, but the L40S's larger capacity allows it to hold more data on-chip, trading raw speed for capacity.

The power and physical profiles also differ. The L40S has a TDP of 300 W, is dual-slot, and requires a 700 W suggested PSU. The RTX 4090 D has a TDP of 425 W, is triple-slot, and needs an 800 W suggested PSU. Both use a single 16-pin power connector and a PCIe 4.0 x16 interface, but the RTX 4090 D is longer at 304 mm versus the L40S's 267 mm, and taller at 137 mm versus 111 mm.

Head-to-Head Benchmarks

The two available head-to-head tests show the L40S winning decisively in one and marginally in the other. In Geekbench OpenCL, the L40S scores 330,727 against the RTX 4090 D's 278,621, a delta of 18.7%. This is a large gap, likely driven by the L40S's higher shading unit count (18,176 vs 14,592) and its 91.61 TFLOPS FP32 throughput, which is 24.5% higher than the RTX 4090 D's 73.54 TFLOPS. OpenCL workloads that scale with shader count and raw FLOPs will consistently favor the L40S.

In Geekbench Vulkan, the L40S wins again, but by a narrower 5.6% margin: 260,799 versus 246,941. The smaller delta suggests that Vulkan's driver overhead or memory access patterns partially mitigate the L40S's compute advantage. The RTX 4090 D's faster 1.01 TB/s bandwidth and higher base clock (2280 MHz vs 1110 MHz) may help it close the gap in memory-bound sub-tests, even though the L40S still prevails overall.

The RTX 4090 D has no wins in these head-to-head results, but its single benchmark — 3DMark Steel Nomad DX12 — scores 8,587, which is not compared directly against the L40S. This absence of a 3DMark result for the L40S means the RTX 4090 D's gaming-oriented performance cannot be directly quantified against the server card in this dataset. The L40S's 18.7% OpenCL lead is the largest margin, while its 5.6% Vulkan lead shows it is not invincible in every API.

Specification Differences

The specification sheet reveals where the two diverge, beyond the core counts already discussed. The L40S has a base clock of 1110 MHz, while the RTX 4090 D starts at 2280 MHz — a 105% higher base frequency. Both boost to 2520 MHz, so the L40S's boost clock is a 127% increase over its base, whereas the RTX 4090 D's boost is only a 10.5% increase. This suggests the L40S relies on sustained boost under load, while the RTX 4090 D runs closer to its maximum clock at idle.

Memory clocks differ: the L40S runs at 2250 MHz with 18 Gbps effective, while the RTX 4090 D runs at 1313 MHz with 21 Gbps effective. The L40S's higher memory clock is offset by the RTX 4090 D's faster effective data rate, resulting in the bandwidth gap noted earlier. Pixel and texture rates follow the core counts: the L40S achieves 483.8 GPixel/s and 1,431.4 GTexel/s, versus the RTX 4090 D's 443.5 GPixel/s and 1,149.1 GTexel/s.

The L40S is rated at 300 W TDP, 125 W lower than the RTX 4090 D's 425 W, yet it delivers higher FP32 and FP16 throughput. This efficiency gap is notable, but the RTX 4090 D compensates with a higher pixel rate per watt in some scenarios. The L40S is dual-slot and 267 mm long; the RTX 4090 D is triple-slot, 304 mm long, and 61 mm wide. Both offer the same display outputs: 1x HDMI 2.1 and 3x DisplayPort 1.4a.

Release dates differ by over a year: the L40S launched on 2022-10-12, while the RTX 4090 D came on 2023-12-27. Their predecessor and successor lines also differ — the L40S follows Server Ampere and precedes Server Hopper, while the RTX 4090 D follows GeForce 30 and precedes GeForce 50.

FAQ

Q: Which GPU has higher raw compute performance?

A: The NVIDIA L40S, with 91.61 TFLOPS FP32 and FP16, versus the RTX 4090 D's 73.54 TFLOPS for both. This is a 24.5% advantage for the L40S, reflected in its 18.7% OpenCL benchmark lead.

Q: Does the RTX 4090 D have any memory advantage?

A: Yes, in bandwidth. Its GDDR6X memory delivers 1.01 TB/s, exceeding the L40S's 864.0 GB/s from GDDR6. However, the L40S has double the capacity at 48 GB versus 24 GB, which is critical for large models.

Q: Why is the RTX 4090 D's average benchmark score lower despite a higher base clock?

A: The RTX 4090 D averages 178,050 across all benchmarks, while the L40S averages 295,763. The L40S's higher core counts (18,176 vs 14,592 shading units) and 91.61 TFLOPS compute outweigh the RTX 4090 D's 2280 MHz base clock, which only helps in short bursts.

Q: Which card is more power-efficient?

A: The L40S, rated at 300 W TDP versus the RTX 4090 D's 425 W. The L40S delivers higher FP32 performance at a lower power draw, making it the better choice for dense server deployments where power density is a concern.

Q: Are there any benchmarks where the RTX 4090 D wins?

A: In the provided head-to-head data, the RTX 4090 D wins zero tests. Its only standalone benchmark is 3DMark Steel Nomad DX12 (8,587), but there is no L40S result to compare against, so no direct win can be confirmed.

Q: What is the significance of the launch MSRP for the RTX 4090 D?

A: The RTX 4090 D has a launch MSRP of 1,599 USD. The L40S has no listed launch MSRP, so a direct price comparison is impossible from the data, but the RTX 4090 D's consumer positioning is evident from its GeForce 40-series generation.

Where Each One Wins

The L40S wins in any scenario that stresses raw compute throughput or memory capacity. Its 48 GB GDDR6 pool is ideal for deep learning training batches, large language model inference, or scientific simulation datasets that would overflow the RTX 4090 D's 24 GB. The 18.7% OpenCL lead and 24.5% higher FP32 rate make it the clear choice for GPU-accelerated analytics, rendering farms, or any workload that scales linearly with shader and tensor core count. Its 300 W TDP and dual-slot form factor also allow higher density in server chassis compared to the RTX 4090 D's 425 W triple-slot design.

The RTX 4090 D wins in scenarios where memory bandwidth and clock speed matter more than capacity. Its 1.01 TB/s bandwidth is 16.9% higher than the L40S's, which benefits real-time ray tracing, high-resolution texture streaming, or low-latency inference where data must move quickly rather than sit in large pools. Its 2280 MHz base clock ensures it reaches peak performance faster, making it suitable for interactive workloads where response time is critical. The 5.6% Vulkan delta shows it is competitive in modern graphics APIs, and its 3DMark Steel Nomad score of 8,587 suggests gaming or DX12-based rendering is its natural habitat.

The data does not support a single winner. The L40S dominates compute and capacity; the RTX 4090 D excels at bandwidth and clock-driven tasks. For a server rack running batch jobs, the L40S is the statistical pick. For a desktop workstation with latency-sensitive interactive rendering, the RTX 4090 D's higher clocks and bandwidth make it a reasonable alternative, despite losing both head-to-head tests. The 99th percentile ranking of the L40S versus the 98th percentile of the RTX 4090 D underscores that both are elite parts, but their strengths are orthogonal. The L40S's 48 GB memory is its trump card; the RTX 4090 D's 1.01 TB/s bandwidth is its counter. Choose based on which bottleneck you hit first.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 4090 D
L40S
Core Specs
Shading Units
14,592
18,176 +24.6%
Shaders
14,592
18,176 +24.6%
TMUs
456
568 +24.6%
ROPs
176
192 +9.1%
SM Count
114
142 +24.6%
Clocks
Base Clock
2280 MHz
1110 MHz
Boost Clock
2520 MHz
2520 MHz
Memory Clock
1313 MHz 21 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
24 GB
48 GB
VRAM (MB)
24,576
49,152 +100.0%
Memory Type
GDDR6X
GDDR6
Memory Bus
384 bit
384 bit
Bandwidth
1.01 TB/s
864.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
72 MB
48 MB
Performance
Pixel Rate
443.5 GPixel/s
483.8 GPixel/s
Texture Rate
1,149.1 GTexel/s
1,431.4 GTexel/s
FP32 (TFLOPS)
73.54 TFLOPS
91.61 TFLOPS
FP64 (TFLOPS)
1,149.1 GFLOPS (1:64)
1,431.4 GFLOPS (1:64)
FP16 (TFLOPS)
73.54 TFLOPS (1:1)
91.61 TFLOPS (1:1)
AI/RT
RT Cores
114
142 +24.6%
Tensor Cores
456
568 +24.6%
Power
TDP
425 W
300 W
TDP (W)
425
300 -29.4%
Suggested PSU
800 W
700 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Ada Lovelace
Ada Lovelace
GPU Name
AD102
AD102
Generation
GeForce 40
Server Ada (Lxx)
Process Size
5 nm
5 nm
Transistors
76,300 million
76,300 million
Die Size
609 mm²
609 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Triple-slot
Dual-slot
Length
304 mm 12 inches
267 mm 10.5 inches
Height
137 mm 5.4 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Launch Price
1,599 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 30
Server Ampere
Successor
GeForce 50
Server Hopper
View GeForce RTX 4090 D Details View L40S Details