NVIDIA GeForce RTX 5090 D vs NVIDIA L4 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 5090 D

CORE STATE GB202
VRAM 32 GB
CLOCK SPEED 2407 MHz
TDP 575 W
BUS WIDTH 512 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

L4

CORE STATE AD104
VRAM 24 GB
CLOCK SPEED 2040 MHz
TDP 72 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
14,326
N/A
geekbench_opencl
310,674
140,838
geekbench_vulkan
376,915
121,306
passmark_directx_10
231
N/A
passmark_directx_11
371
N/A
passmark_directx_12
219
N/A
passmark_directx_9
434
N/A
passmark_g2d
1,487
N/A
passmark_g3d
44,065
N/A
passmark_gpu_compute
28,396
N/A

Analysis: NVIDIA GeForce RTX 5090 D vs NVIDIA L4

Head-to-Head Benchmarks

The recorded database contains two common benchmark results for this pair, and both decisively favor the NVIDIA GeForce RTX 5090 D. In the Geekbench OpenCL test, the RTX 5090 D scores 310,674 against the L4's 140,838, a delta of 54.7%. The Vulkan result is even more lopsided: the RTX 5090 D reaches 376,915 while the L4 manages 121,306, putting the L4 67.8% behind. In both tests, the RTX 5090 D more than doubles the L4's output, indicating a fundamental performance gap rather than a close contest.

The L4's own benchmark average of 131,072 places it in the 95th percentile of all GPUs in the database, which is a strong showing for a low-profile server card. Its nearest rivals in the database include the GeForce RTX 3090 Ti at 131,938 (0.7% higher), the RTX 4000 Ada Generation at 135,218 (3.1% higher), and the A10M at 135,230 (3.1% higher). The L4 sits within a tight cluster of professional cards, meaning its compute throughput is competitive with much larger, higher-power solutions from the previous generation.

The RTX 5090 D, despite its higher raw scores in these two tests, has a much lower average benchmark score of 77,712 across all recorded tests, and it sits in the 92nd percentile. This counterintuitive result stems from the fact that the RTX 5090 D's benchmark suite includes many older DirectX tests (PassMark DX9, DX10, DX11, DX12) where its scores are relatively modest, dragging down the arithmetic average. Its nearest rivals in the database are the Radeon RX 6650M XT at 76,904 (1.1% higher), the Radeon RX 6850M XT at 78,940 (1.6% lower), and the Tesla P100 variants at 79,396 and 79,605 (2.1% and 2.4% lower respectively). These are all mobile or datacenter parts, not desktop flagships, which suggests the average is heavily influenced by the heterogeneous test mix.

Where Each One Wins

The RTX 5090 D wins both head-to-head tests in the database, specifically the OpenCL and Vulkan compute workloads. These are modern, general-purpose GPU compute APIs, and the advantage is substantial: 54.7% in OpenCL and 67.8% in Vulkan. This indicates that for any workload using these APIs, the RTX 5090 D is the clear choice, likely due to its much larger shader array and higher clock speeds.

The L4, while losing both direct comparisons, still demonstrates its own strengths within the database. Its average benchmark score of 131,072 is higher than the RTX 5090 D's 77,712, and its percentile rank (95th) is above the RTX 5090 D's (92nd). This is because the L4's recorded benchmarks are only the two Geekbench tests, both of which are compute-heavy but not representative of the full spectrum of graphics workloads. The L4's performance is tightly clustered with the RTX 3090 Ti, RTX 4000 Ada, and A10M, making it a balanced performer within its own segment. For tasks that prioritize raw compute per watt, the L4's 72 W TDP against the RTX 5090 D's 575 W suggests a vastly different efficiency profile, though the database does not provide direct efficiency metrics.

Architecture Differences

The two GPUs come from different architectural generations. The L4 is built on Ada Lovelace, using the AD104 chip, manufactured on a 5 nm process at TSMC. It packs 35,800 million transistors into a 294 mm² die, yielding a transistor density of 121.8 million per mm². The RTX 5090 D uses the Blackwell 2.0 architecture with the GB202 chip, also on a 5 nm TSMC process, but with 92,200 million transistors across a 750 mm² die, for a density of 122.9 million per mm². The density figures are nearly identical, but the RTX 5090 D has roughly 2.6 times the transistor count and 2.55 times the die area.

The memory subsystems are completely different. The L4 uses 24 GB of GDDR6 on a 192-bit bus, delivering 300.1 GB/s of bandwidth. The RTX 5090 D uses 32 GB of GDDR7 on a 512-bit bus, with 1.79 TB/s of bandwidth, nearly six times the L4's memory throughput. Clock speeds also differ: the L4 has a 795 MHz base and 2040 MHz boost, while the RTX 5090 D runs at 2017 MHz base and 2407 MHz boost. Memory clocks are 1563 MHz (12.5 Gbps effective) for the L4 and 1750 MHz (28 Gbps effective) for the RTX 5090 D.

The compute resources scale accordingly. The L4 has 7,424 shading units, 240 TMUs, 80 ROPs, 60 RT cores, and 240 Tensor cores. The RTX 5090 D has 21,760 shading units, 680 TMUs, 176 ROPs, 170 RT cores, and 680 Tensor cores. This translates to pixel rates of 163.2 GPixel/s for the L4 versus 423.6 GPixel/s for the RTX 5090 D, and texture rates of 489.6 GTexel/s versus 1,636.8 GTexel/s. FP32 throughput is 30.29 TFLOPS for the L4 and 104.8 TFLOPS for the RTX 5090 D, with both offering 1:1 FP16 to FP32 ratios.

Physical and power characteristics diverge sharply. The L4 is a single-slot, 169 mm long, 56 mm tall card with no power connectors and a 72 W TDP, requiring only a 250 W suggested PSU. It has no display outputs. The RTX 5090 D is a dual-slot card measuring 304 mm by 137 mm by 48 mm, uses a single 16-pin power connector, draws 575 W, and requires a 950 W PSU. It offers 1x HDMI 2.1b and 3x DisplayPort 2.1b outputs. The bus interface is PCIe 4.0 x16 for the L4 and PCIe 5.0 x16 for the RTX 5090 D.

FAQ

Q: Which GPU has the higher average benchmark score in the database?

A: The L4 has a higher average score of 131,072, compared to the RTX 5090 D's 77,712. This is because the L4 is only recorded with two Geekbench compute tests, while the RTX 5090 D includes several legacy PassMark DirectX tests that score lower.

Q: What is the memory bandwidth difference between the two?

A: The RTX 5090 D has 1.79 TB/s of bandwidth across a 512-bit bus with GDDR7 memory, while the L4 has 300.1 GB/s across a 192-bit bus with GDDR6. The RTX 5090 D's bandwidth is roughly six times higher.

Q: How do their transistor counts compare?

A: The RTX 5090 D contains 92,200 million transistors, while the L4 has 35,800 million. Both are manufactured on a 5 nm process at TSMC.

Q: Which card has a higher boost clock?

A: The RTX 5090 D boosts to 2407 MHz, while the L4 boosts to 2040 MHz. The base clocks are 2017 MHz and 795 MHz respectively.

Q: Do both cards support the same graphics APIs?

A: Yes, both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

Q: What is the power connector requirement for each card?

A: The L4 has no power connectors and a 72 W TDP, while the RTX 5090 D uses a single 16-pin connector and has a 575 W TDP. The suggested PSU ratings are 250 W and 950 W respectively.

Specification Differences

| Specification | NVIDIA L4 | NVIDIA GeForce RTX 5090 D |

| --- | --- | --- |

| Architecture | Ada Lovelace | Blackwell 2.0 |

| Chip | AD104 | GB202 |

| Generation | Server Ada (Lxx) | GeForce 50 |

| Transistors | 35,800 million | 92,200 million |

| Die Size | 294 mm² | 750 mm² |

| Transistor Density | 121.8M / mm² | 122.9M / mm² |

| Base Clock | 795 MHz | 2017 MHz |

| Boost Clock | 2040 MHz | 2407 MHz |

| Memory Clock | 1563 MHz (12.5 Gbps effective) | 1750 MHz (28 Gbps effective) |

| Memory Size | 24 GB | 32 GB |

| Memory Type | GDDR6 | GDDR7 |

| Memory Bus Width | 192 bit | 512 bit |

| Memory Bandwidth | 300.1 GB/s | 1.79 TB/s |

| Shading Units | 7,424 | 21,760 |

| TMUs | 240 | 680 |

| ROPs | 80 | 176 |

| RT Cores | 60 | 170 |

| Tensor Cores | 240 | 680 |

| Pixel Rate | 163.2 GPixel/s | 423.6 GPixel/s |

| Texture Rate | 489.6 GTexel/s | 1,636.8 GTexel/s |

| FP32 | 30.29 TFLOPS | 104.8 TFLOPS |

| FP16 | 30.29 TFLOPS (1:1) | 104.8 TFLOPS (1:1) |

| TDP | 72 W | 575 W |

| Slot Width | Single-slot | Dual-slot |

| Power Connectors | None | 1x 16-pin |

| Suggested PSU | 250 W | 950 W |

| Bus Interface | PCIe 4.0 x16 | PCIe 5.0 x16 |

| Display Outputs | No outputs | 1x HDMI 2.1b, 3x DisplayPort 2.1b |

| Length | 169 mm (6.7 inches) | 304 mm (12 inches) |

| Height | 56 mm (2.2 inches) | 137 mm (5.4 inches) |

| Width | Not specified | 48 mm (1.9 inches) |

| Release Date | 2023-03-20 | 2025-01-29 |

| Predecessor | Server Ampere | GeForce 40 |

| Successor | Server Hopper | GeForce 60 |

| Launch MSRP | Not specified | 2,299 USD |

The Verdict

The data presents a clear split. For raw compute performance in OpenCL and Vulkan, the RTX 5090 D is the overwhelming winner, with deltas of 54.7% and 67.8% respectively. It also offers more memory, wider bandwidth, and higher clock speeds across the board. The RTX 5090 D's 32 GB of GDDR7 with 1.79 TB/s bandwidth versus the L4's 24 GB of GDDR6 with 300.1 GB/s makes it the choice for memory-bound workloads. Its 104.8 TFLOPS FP32 throughput is more than three times the L4's 30.29 TFLOPS.

However, the L4's record in the database is not without merit. Its average benchmark score of 131,072 is higher than the RTX 5090 D's 77,712, and its 95th percentile ranking exceeds the 92nd percentile of the RTX 5090 D. The L4 is a single-slot, 72 W card that requires no external power connectors, making it suitable for dense server deployments. The RTX 5090 D, with a 575 W TDP and a 950 W suggested PSU, is a power-hungry dual-slot desktop part with display outputs.

The decision hinges on the use case. For a datacenter environment prioritizing compute density and low power draw, the L4 is the logical pick, especially given its competitive standing among the RTX 3090 Ti and RTX 4000 Ada in the database. For a workstation or desktop requiring maximum compute throughput, the RTX 5090 D is the clear choice, provided the power and cooling infrastructure can support it. The RTX 5090 D also has the advantage of a 2025 release date and a successor (GeForce 60) already listed, while the L4's 2023 release and Server Hopper successor indicate a different product lifecycle. Ultimately, the database shows two specialized tools: one for efficiency and density, the other for peak performance.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 5090 D
L4
Core Specs
Shading Units
21,760
7,424 -65.9%
Shaders
21,760
7,424 -65.9%
TMUs
680
240 -64.7%
ROPs
176
80 -54.5%
SM Count
170
60 -64.7%
Clocks
Base Clock
2017 MHz
795 MHz
Boost Clock
2407 MHz
2040 MHz
Memory Clock
1750 MHz 28 Gbps effective
1563 MHz 12.5 Gbps effective
Memory
Memory Size
32 GB
24 GB
VRAM (MB)
32,768
24,576 -25.0%
Memory Type
GDDR7
GDDR6
Memory Bus
512 bit
192 bit
Bandwidth
1.79 TB/s
300.1 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
96 MB
48 MB
Performance
Pixel Rate
423.6 GPixel/s
163.2 GPixel/s
Texture Rate
1,636.8 GTexel/s
489.6 GTexel/s
FP32 (TFLOPS)
104.8 TFLOPS
30.29 TFLOPS
FP64 (TFLOPS)
1.637 TFLOPS (1:64)
473.3 GFLOPS (1:64)
FP16 (TFLOPS)
104.8 TFLOPS (1:1)
30.29 TFLOPS (1:1)
AI/RT
RT Cores
170
60 -64.7%
Tensor Cores
680
240 -64.7%
Power
TDP
575 W
72 W
TDP (W)
575
72 -87.5%
Suggested PSU
950 W
250 W
Power Connectors
1x 16-pin
None
Architecture
Architecture
Blackwell 2.0
Ada Lovelace
GPU Name
GB202
AD104
Generation
GeForce 50
Server Ada (Lxx)
Process Size
5 nm
5 nm
Transistors
92,200 million
35,800 million
Die Size
750 mm²
294 mm²
Foundry
TSMC
TSMC
Density
122.9M / mm²
121.8M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
12.0
8.9
Shader Model
6.9
6.8
Physical
Slot Width
Dual-slot
Single-slot
Length
304 mm 12 inches
169 mm 6.7 inches
Height
137 mm 5.4 inches
56 mm 2.2 inches
Outputs
1x HDMI 2.1b3x DisplayPort 2.1b
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 4.0 x16
Other
Launch Price
2,299 USD
Production
Active
Active
Predecessor
GeForce 40
Server Ampere
Successor
GeForce 60
Server Hopper
View GeForce RTX 5090 D Details View L4 Details