NVIDIA GeForce RTX 5090 D vs NVIDIA L20 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 5090 D

CORE STATE GB202
VRAM 32 GB
CLOCK SPEED 2407 MHz
TDP 575 W
BUS WIDTH 512 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

L20

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 275 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
14,326
N/A
geekbench_opencl
310,674
274,276
geekbench_vulkan
376,915
228,018
passmark_directx_10
231
N/A
passmark_directx_11
371
N/A
passmark_directx_12
219
N/A
passmark_directx_9
434
N/A
passmark_g2d
1,487
N/A
passmark_g3d
44,065
N/A
passmark_gpu_compute
28,396
N/A

Analysis: NVIDIA GeForce RTX 5090 D vs NVIDIA L20

Head-to-Head Benchmarks

The recorded data contains two direct comparisons between the NVIDIA L20 and the NVIDIA GeForce RTX 5090 D, both in Geekbench compute workloads. The RTX 5090 D wins both, but the margins differ sharply depending on the API. In Geekbench OpenCL, the RTX 5090 D scores 310,674 against the L20’s 274,276, a lead of 11.7%. That is a solid but not overwhelming advantage in a compute-heavy, cross-platform workload. In Geekbench Vulkan, the gap widens dramatically: the RTX 5090 D scores 376,915 versus 228,018 for the L20, a 39.5% difference. Vulkan’s lower overhead and the RTX 5090 D’s much larger shading core count likely explain the larger separation, though the database does not break down the workload composition further.

Looking at the broader benchmark context, the L20’s average score across all recorded tests is 251,147, placing it in the 99th percentile of all GPUs in the database. The RTX 5090 D’s average score is only 77,712, but that figure is skewed because it includes several Passmark tests (DirectX 9, 10, 11, 12, G2D, G3D, and GPU compute) where the RTX 5090 D records scores in the hundreds or low thousands, dragging the mean down. The L20 has only two recorded benchmarks, both Geekbench, so its average is purely compute-oriented. Direct comparison of the two averages is therefore misleading; the head-to-head Geekbench numbers are the only apples-to-apples data available.

The nearest rivals in the database further contextualize each card. The L20 sits 11.6% above the NVIDIA PG506-232, 14.2% above the AMD Radeon PRO W7900D, but 11.6% below the NVIDIA L40 and 12.6% below the NVIDIA RTX 6000 Ada Generation. The RTX 5090 D, by contrast, is nearly tied with the AMD Radeon RX 6650M XT (1.1% ahead) and trails the AMD Radeon RX 6850M XT by 1.6%, the NVIDIA Tesla P100 PCIe 12 GB by 2.1%, and the Tesla P100 PCIe 16 GB by 2.4%. Those rival deltas for the RTX 5090 D are all based on its distorted average, so they should be read with caution; the Geekbench head-to-head results are far more representative of the RTX 5090 D’s actual compute strength.

Architecture Differences

The two GPUs come from different NVIDIA architectures and different market segments. The L20 uses the AD102 chip on the Ada Lovelace architecture, fabricated by TSMC on a 5 nm process. The RTX 5090 D uses the GB202 chip on the Blackwell 2.0 architecture, also TSMC 5 nm. Transistor counts differ substantially: the L20 packs 76,300 million transistors on a 609 mm² die, for a density of 125.3 million per mm². The RTX 5090 D has 92,200 million transistors on a 750 mm² die, a slightly lower density of 122.9 million per mm². The RTX 5090 D is physically larger and denser in absolute terms, but the L20 is marginally more efficient in packing transistors per square millimeter.

Core counts follow the architecture generational leap. The L20 has 11,776 shading units, 368 TMUs, 128 ROPs, 92 RT cores, and 368 tensor cores. The RTX 5090 D nearly doubles the shading units to 21,760, with 680 TMUs, 176 ROPs, 170 RT cores, and 680 tensor cores. Clock behavior is noteworthy: the L20 has a lower base clock (1440 MHz) but a higher boost clock (2520 MHz), while the RTX 5090 D starts at 2017 MHz base and boosts to 2407 MHz. So the L20’s boost exceeds the RTX 5090 D’s boost by 113 MHz, but the RTX 5090 D’s base clock is 577 MHz higher, giving it a higher floor under sustained load.

Memory configurations diverge completely. The L20 uses 48 GB of GDDR6 on a 384-bit bus, with 2250 MHz memory clock (18 Gbps effective) and 864.0 GB/s bandwidth. The RTX 5090 D uses 32 GB of GDDR7 on a 512-bit bus, with 1750 MHz memory clock (28 Gbps effective) and 1.79 TB/s bandwidth. The RTX 5090 D has 48% more bus width and nearly double the bandwidth, but the L20 has 50% more capacity. For workloads that need to hold large datasets, the L20 wins; for streaming throughput, the RTX 5090 D wins decisively. The memory type difference (GDDR6 vs GDDR7) drives the bandwidth gap, as the effective data rate is 28 Gbps versus 18 Gbps.

Pixel and texture fill rates reflect the core count difference. The L20 achieves 322.6 GPixel/s and 927.4 GTexel/s; the RTX 5090 D reaches 423.6 GPixel/s and 1,636.8 GTexel/s. Compute throughput in FP32 is 59.35 TFLOPS for the L20 and 104.8 TFLOPS for the RTX 5090 D, a 76.5% advantage for the newer card. Both have FP16 at 1:1 ratio with FP32, so the same ratio holds in half precision. Power envelopes differ accordingly: the L20 is rated at 275 W with a 600 W suggested PSU, while the RTX 5090 D is rated at 575 W with a 950 W suggested PSU. Both use a single 16-pin power connector and dual-slot cooling. The L20 is shorter at 267 mm (10.5 inches) versus 304 mm (12 inches), and lower at 111 mm (4.4 inches) versus 137 mm (5.4 inches); the RTX 5090 D adds a width spec of 48 mm (1.9 inches) while the L20 has no width listed.

Bus interface and display outputs differ as well. The L20 uses PCIe 4.0 x16, while the RTX 5090 D uses PCIe 5.0 x16. Display outputs: the L20 has four DisplayPort 1.4a, the RTX 5090 D has one HDMI 2.1b and three DisplayPort 2.1b. Both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. Release dates are separated by roughly 14 months: the L20 launched on 2023-11-15, the RTX 5090 D on 2025-01-29. The L20’s predecessor is Server Ampere and successor is Server Hopper; the RTX 5090 D’s predecessor is GeForce 40 and successor is GeForce 60.

Where Each One Wins

The RTX 5090 D wins in raw compute throughput. Its FP32 of 104.8 TFLOPS is 76.5% higher than the L20’s 59.35 TFLOPS. Texture rate is 76.5% higher (1,636.8 vs 927.4 GTexel/s), pixel rate is 31.3% higher (423.6 vs 322.6 GPixel/s). Memory bandwidth is 107% higher (1.79 TB/s vs 864.0 GB/s). These are all throughput metrics where the RTX 5090 D’s larger core count and faster GDDR7 memory give it a clear edge. For rendering, simulation, and any workload that scales with raw shading or texture throughput, the RTX 5090 D is the stronger choice.

The L20 wins in memory capacity and power efficiency per watt. It holds 48 GB versus 32 GB, a 50% advantage, which matters for large language model inference, massive datasets, or multi-tenant virtualized workloads that need to fit entire models in VRAM. Its TDP is 275 W versus 575 W, so it draws less than half the power for roughly 57% of the FP32 throughput (59.35 vs 104.8 TFLOPS). The L20 also boosts higher (2520 MHz vs 2407 MHz), which helps in lightly threaded or latency-sensitive tasks where clock speed matters more than core count. Its smaller physical footprint (267 mm vs 304 mm length) and lower height (111 mm vs 137 mm) make it easier to fit in dense server chassis.

The Geekbench results align with these architectural splits. In OpenCL, the RTX 5090 D leads by 11.7%, which is modest given its 84.7% more shading units. That suggests the L20’s higher boost clock and possibly better driver optimization for compute in OpenCL narrow the gap. In Vulkan, the RTX 5090 D leads by 39.5%, which is closer to the raw core count difference. Vulkan tends to expose more hardware parallelism, so the RTX 5090 D’s larger pool of shaders, RT cores, and tensor cores shows through more clearly. For users who primarily run OpenCL workloads, the L20 is surprisingly competitive; for Vulkan-centric applications, the RTX 5090 D runs away.

FAQ

Q: Which GPU has higher FP32 compute throughput?

A: The RTX 5090 D, at 104.8 TFLOPS, which is 76.5% higher than the L20’s 59.35 TFLOPS.

Q: Which GPU has more memory capacity?

A: The L20, with 48 GB, versus 32 GB on the RTX 5090 D.

Q: What is the memory bandwidth difference?

A: The RTX 5090 D offers 1.79 TB/s, which is 107% higher than the L20’s 864.0 GB/s.

Q: How do they compare in Geekbench OpenCL?

A: The RTX 5090 D scores 310,674 versus 274,276 for the L20, a lead of 11.7%.

Q: How do they compare in Geekbench Vulkan?

A: The RTX 5090 D scores 376,915 versus 228,018 for the L20, a lead of 39.5%.

Q: Which GPU has a higher boost clock?

A: The L20, at 2520 MHz, versus 2407 MHz for the RTX 5090 D.

The Verdict

The data points to a clear split by use case. For compute-heavy rendering, Vulkan-based applications, and workloads that demand maximum throughput in FP32, texture fill, or memory bandwidth, the RTX 5090 D is the superior card. Its 76.5% higher FP32, 76.5% higher texture rate, and 107% higher bandwidth are decisive. The Geekbench Vulkan result (39.5% ahead) confirms this. The RTX 5090 D also has a much higher base clock (2017 MHz vs 1440 MHz), which helps sustained performance under load. Its launch MSRP is 2,299 USD.

For memory-bound workloads, the L20 is the better choice. Its 48 GB capacity is 50% larger, which is critical for models or datasets that exceed 32 GB. It also draws 275 W versus 575 W, making it far more efficient per watt for FP32 compute (roughly 216 GFLOPS per watt versus 182 GFLOPS per watt, calculated from the recorded figures). The L20’s higher boost clock (2520 MHz) gives it an edge in latency-sensitive tasks, and its OpenCL score is only 11.7% behind the RTX 5090 D despite having 46% fewer shading units. For server environments where power density, cooling, and physical space are constrained, the L20’s smaller footprint and lower power draw are significant advantages.

The percentile rankings reinforce this: the L20 sits in the 99th percentile of all GPUs, while the RTX 5090 D is in the 92nd. That difference is partly due to the RTX 5090 D’s Passmark scores pulling its average down, but it also reflects the L20’s specialization as a compute-optimized server part. Users who need raw speed in a workstation context should choose the RTX 5090 D. Users who need large VRAM capacity, lower power consumption, or server-specific features (PCIe 4.0, DisplayPort 1.4a outputs, compact dimensions) should choose the L20. The two are not direct competitors; they serve different segments of the same silicon family lineage.

Specification Differences

| Field | NVIDIA L20 | NVIDIA GeForce RTX 5090 D |

|---|---|---|

| Architecture | Ada Lovelace | Blackwell 2.0 |

| Chip | AD102 | GB202 |

| Generation | Server Ada (Lxx) | GeForce 50 |

| Transistors | 76,300 million | 92,200 million |

| Die Size | 609 mm² | 750 mm² |

| Transistor Density | 125.3M / mm² | 122.9M / mm² |

| Base Clock | 1440 MHz | 2017 MHz |

| Boost Clock | 2520 MHz | 2407 MHz |

| Memory Clock | 2250 MHz (18 Gbps effective) | 1750 MHz (28 Gbps effective) |

| Memory Size | 48 GB | 32 GB |

| Memory Type | GDDR6 | GDDR7 |

| Memory Bus Width | 384 bit | 512 bit |

| Memory Bandwidth | 864.0 GB/s | 1.79 TB/s |

| Shading Units | 11776 | 21760 |

| TMUs | 368 | 680 |

| ROPs | 128 | 176 |

| RT Cores | 92 | 170 |

| Tensor Cores | 368 | 680 |

| Pixel Rate | 322.6 GPixel/s | 423.6 GPixel/s |

| Texture Rate | 927.4 GTexel/s | 1,636.8 GTexel/s |

| FP32 | 59.35 TFLOPS | 104.8 TFLOPS |

| FP16 | 59.35 TFLOPS (1:1) | 104.8 TFLOPS (1:1) |

| TDP | 275 W | 575 W |

| Suggested PSU | 600 W | 950 W |

| Bus Interface | PCIe 4.0 x16 | PCIe 5.0 x16 |

| Display Outputs | 4x DisplayPort 1.4a | 1x HDMI 2.1b, 3x DisplayPort 2.1b |

| Length | 267 mm (10.5 inches) | 304 mm (12 inches) |

| Height | 111 mm (4.4 inches) | 137 mm (5.4 inches) |

| Width | Not listed | 48 mm (1.9 inches) |

| Release Date | 2023-11-15 | 2025-01-29 |

| Predecessor | Server Ampere | GeForce 40 |

| Successor | Server Hopper | GeForce 60 |

| Launch MSRP | Not listed | 2,299 USD |

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 5090 D
L20
Core Specs
Shading Units
21,760
11,776 -45.9%
Shaders
21,760
11,776 -45.9%
TMUs
680
368 -45.9%
ROPs
176
128 -27.3%
SM Count
170
92 -45.9%
Clocks
Base Clock
2017 MHz
1440 MHz
Boost Clock
2407 MHz
2520 MHz
Memory Clock
1750 MHz 28 Gbps effective
2250 MHz 18 Gbps effective
Memory
Memory Size
32 GB
48 GB
VRAM (MB)
32,768
49,152 +50.0%
Memory Type
GDDR7
GDDR6
Memory Bus
512 bit
384 bit
Bandwidth
1.79 TB/s
864.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
96 MB
96 MB
Performance
Pixel Rate
423.6 GPixel/s
322.6 GPixel/s
Texture Rate
1,636.8 GTexel/s
927.4 GTexel/s
FP32 (TFLOPS)
104.8 TFLOPS
59.35 TFLOPS
FP64 (TFLOPS)
1.637 TFLOPS (1:64)
927.4 GFLOPS (1:64)
FP16 (TFLOPS)
104.8 TFLOPS (1:1)
59.35 TFLOPS (1:1)
AI/RT
RT Cores
170
92 -45.9%
Tensor Cores
680
368 -45.9%
Power
TDP
575 W
275 W
TDP (W)
575
275 -52.2%
Suggested PSU
950 W
600 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Blackwell 2.0
Ada Lovelace
GPU Name
GB202
AD102
Generation
GeForce 50
Server Ada (Lxx)
Process Size
5 nm
5 nm
Transistors
92,200 million
76,300 million
Die Size
750 mm²
609 mm²
Foundry
TSMC
TSMC
Density
122.9M / mm²
125.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
12.0
8.9
Shader Model
6.9
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
304 mm 12 inches
267 mm 10.5 inches
Height
137 mm 5.4 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.1b3x DisplayPort 2.1b
4x DisplayPort 1.4a
Bus Interface
PCIe 5.0 x16
PCIe 4.0 x16
Other
Launch Price
2,299 USD
Production
Active
Active
Predecessor
GeForce 40
Server Ampere
Successor
GeForce 60
Server Hopper
View GeForce RTX 5090 D Details View L20 Details