NVIDIA L40S vs NVIDIA RTX PRO 5000 Blackwell Comparison

NVIDIA
GEFORCE

NVIDIA L40S

CORE STATE AD102
VRAM 48 GB
CLOCK SPEED 2520 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022
VS
NVIDIA
GEFORCE

RTX PRO 5000 Blackwell

CORE STATE GB202
VRAM 48 GB
CLOCK SPEED 2377 MHz
TDP 300 W
BUS WIDTH 384 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2025

PERFORMANCE BENCHMARKS

geekbench_opencl
330,727
254,116
geekbench_vulkan
260,799
282,631
3dmark_3dmark_steel_nomad_dx12
N/A
9,579.5

Analysis: NVIDIA L40S vs NVIDIA RTX PRO 5000 Blackwell

FAQ

Q: How do the two GPUs compare in raw compute benchmarks?

A: The NVIDIA L40S wins the Geekbench OpenCL test with a score of 330,727, which is 30.1% higher than the RTX PRO 5000 Blackwell’s 254,116. In the Geekbench Vulkan test, the RTX PRO 5000 Blackwell takes the lead with 282,631 versus 260,799, a 7.7% advantage for the newer card.

Q: What are the memory specifications of each card?

A: Both cards have 48 GB of memory on a 384-bit bus. The L40S uses GDDR6 with 864.0 GB/s bandwidth, while the RTX PRO 5000 Blackwell uses GDDR7 with 1.34 TB/s bandwidth — a substantial memory speed advantage for the Blackwell card.

Q: Which card has a higher transistor count and larger die?

A: The RTX PRO 5000 Blackwell packs 92,200 million transistors on a 750 mm² die, while the L40S has 76,300 million transistors on a 609 mm² die. The L40S has a slightly higher transistor density at 125.3M per mm² versus 122.9M per mm² for the Blackwell card.

Q: What are the core counts for each GPU?

A: The L40S has 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The RTX PRO 5000 Blackwell has 14,080 shading units, 440 TMUs, 160 ROPs, 110 RT cores, and 440 tensor cores — notably lower in every category.

Q: What is the production status and release timeline for these cards?

A: The L40S is end-of-life and was released on 2022-10-12. The RTX PRO 5000 Blackwell is active and was released on 2025-03-17, with a launch MSRP of 5,099 USD.

Q: What is the average benchmark score and performance percentile for each?

A: The L40S has an average benchmark score of 295,763 and sits at the 99th percentile of all GPUs. The RTX PRO 5000 Blackwell has an average benchmark score of 182,109 and sits at the 98th percentile — despite a much lower average, it remains near the top of the overall GPU distribution.

Architecture Differences

The two cards represent distinct architectural generations. The L40S is built on Ada Lovelace with the AD102 chip, while the RTX PRO 5000 Blackwell uses the newer Blackwell 2.0 architecture with the GB202 chip. Both are fabricated on a 5 nm process at TSMC, but the transistor counts differ significantly: the L40S contains 76,300 million transistors, while the Blackwell card contains 92,200 million — a 15,900 million transistor increase. The die sizes also diverge, with the L40S at 609 mm² and the RTX PRO 5000 Blackwell at 750 mm².

The memory architecture marks a major generational shift. The L40S uses GDDR6 memory running at 18 Gbps effective, while the RTX PRO 5000 Blackwell uses GDDR7 at 28 Gbps effective. This yields a bandwidth jump from 864.0 GB/s to 1.34 TB/s — a 55% improvement in memory throughput. Both cards maintain a 384-bit bus and 48 GB capacity, so the bandwidth difference comes down to the faster memory type and clock speeds.

Core configurations also differ substantially. The L40S has 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The RTX PRO 5000 Blackwell has 14,080 shading units, 440 TMUs, 160 ROPs, 110 RT cores, and 440 tensor cores. The L40S leads in every core count category, which directly impacts its higher FP32 and FP16 throughput of 91.61 TFLOPS compared to 66.94 TFLOPS on the Blackwell card.

Clock speeds tell a different story. The RTX PRO 5000 Blackwell has a higher base clock at 1740 MHz versus 1110 MHz on the L40S, but the L40S boosts higher at 2520 MHz versus 2377 MHz. The Blackwell card’s higher base clock suggests better sustained performance at lower loads, but the L40S has more headroom under boost conditions.

Interface and display outputs also differ. The L40S uses PCIe 4.0 x16, while the RTX PRO 5000 Blackwell uses PCIe 5.0 x16 — doubling the bus bandwidth for data transfer. For display outputs, the L40S offers 1x HDMI 2.1 and 3x DisplayPort 1.4a, while the Blackwell card offers 4x DisplayPort 2.1b, supporting newer display standards.

Head-to-Head Benchmarks

The two available head-to-head benchmark results show a split decision. In Geekbench OpenCL, the L40S scores 330,727 against 254,116 for the RTX PRO 5000 Blackwell — a 30.1% margin in favor of the older card. This is a decisive victory, reflecting the L40S’s higher shading unit count (18,176 versus 14,080) and higher FP32 throughput (91.61 TFLOPS versus 66.94 TFLOPS). The OpenCL workload appears to scale strongly with raw compute resources, and the L40S has a clear advantage there.

In Geekbench Vulkan, the tables turn. The RTX PRO 5000 Blackwell scores 282,631 against 260,799 for the L40S — a 7.7% lead for the newer card. Despite having fewer cores and lower FP32 performance, the Blackwell card wins this workload. The likely explanation lies in its faster GDDR7 memory (1.34 TB/s versus 864.0 GB/s) and PCIe 5.0 interface, which can reduce data transfer bottlenecks. Vulkan workloads often benefit from memory bandwidth and efficient command processing, areas where the Blackwell architecture appears stronger.

The deltaPct values in the nearest rivals data add context. The L40S’s closest rival is the AMD Instinct MI300X, which is 7% ahead, and the NVIDIA H200 NVL, which is 11.7% ahead. The RTX PRO 5000 Blackwell’s nearest rivals include the NVIDIA A100 SXM4 80 GB (0.9% ahead), the RTX 5000 Ada Generation (1.4% ahead), and the GeForce RTX 4090 D (2.3% behind). This shows the Blackwell card competes closely with prior-generation flagship accelerators, while the L40S sits in a slightly higher performance tier.

When comparing average benchmark scores, the L40S looks stronger on paper: 295,763 versus 182,109 for the Blackwell card. However, this average includes different benchmark suites — the L40S has two Geekbench results, while the Blackwell card includes a 3DMark Steel Nomad DX12 score of 9,579.5, which is not comparable to Geekbench scores. The raw average should be interpreted cautiously.

Specification Differences

The specification sheets reveal several key differences between the two cards. The L40S uses the AD102 chip with Ada Lovelace architecture, while the RTX PRO 5000 Blackwell uses the GB202 chip with Blackwell 2.0 architecture. Process nodes are identical at 5 nm TSMC, but transistor density differs slightly: 125.3M per mm² for the L40S versus 122.9M per mm² for the Blackwell card.

Memory specifications diverge significantly. The L40S has GDDR6 at 18 Gbps effective with 864.0 GB/s bandwidth, while the RTX PRO 5000 Blackwell has GDDR7 at 28 Gbps effective with 1.34 TB/s bandwidth. Both have 48 GB capacity and a 384-bit bus, so the bandwidth gap is entirely due to memory type and speed.

Core counts favor the L40S across the board: 18,176 versus 14,080 shading units, 568 versus 440 TMUs, 192 versus 160 ROPs, 142 versus 110 RT cores, and 568 versus 440 tensor cores. This translates to higher pixel rate (483.8 GPixel/s versus 380.3 GPixel/s), higher texture rate (1,431.4 GTexel/s versus 1,045.9 GTexel/s), and higher FP32/FP16 performance (91.61 TFLOPS versus 66.94 TFLOPS).

Clock speeds differ in interesting ways. The L40S has a base clock of 1110 MHz and boost clock of 2520 MHz, while the RTX PRO 5000 Blackwell has a base clock of 1740 MHz and boost clock of 2377 MHz. The higher base clock on the Blackwell card suggests better idle-to-load transition performance, but the L40S has a higher ceiling.

Power and physical specs are largely identical: both are 300 W TDP, dual-slot, use a single 16-pin power connector, suggest a 700 W PSU, and measure 267 mm in length and 111 mm in height. The Blackwell card is slightly thicker at 40 mm (1.6 inches) versus no width listed for the L40S. The bus interface differs: PCIe 4.0 x16 for the L40S versus PCIe 5.0 x16 for the Blackwell card.

Where Each One Wins

The L40S wins decisively in raw compute throughput. Its 91.61 TFLOPS FP32 performance is 36.8% higher than the Blackwell card’s 66.94 TFLOPS, and this shows in the Geekbench OpenCL result where it leads by 30.1%. For workloads that hammer shader cores and tensor units — such as AI training, scientific computing, or dense matrix operations — the L40S is the stronger choice. Its higher pixel rate (483.8 GPixel/s) and texture rate (1,431.4 GTexel/s) also make it better suited for graphics-heavy rendering tasks at high resolutions.

The RTX PRO 5000 Blackwell wins in memory-bound scenarios. Its GDDR7 memory delivers 1.34 TB/s bandwidth, a 55% improvement over the L40S’s 864.0 GB/s. This explains its 7.7% victory in Geekbench Vulkan, where memory bandwidth often becomes the limiting factor. The newer Blackwell architecture also brings PCIe 5.0 support, which is beneficial for data-intensive workloads that transfer large datasets between CPU and GPU. The higher base clock of 1740 MHz suggests better responsiveness in bursty or latency-sensitive applications.

For display-centric use cases, the RTX PRO 5000 Blackwell has an edge with 4x DisplayPort 2.1b outputs, supporting higher bandwidth display connections than the L40S’s 1x HDMI 2.1 and 3x DisplayPort 1.4a. The Blackwell card’s active production status also means ongoing availability and driver support, while the L40S being end-of-life could limit long-term deployment plans.

The Verdict

The choice between these two cards depends entirely on workload priorities. If raw compute performance is the primary requirement, the data clearly favors the NVIDIA L40S. Its 30.1% lead in Geekbench OpenCL, higher FP32 throughput (91.61 TFLOPS versus 66.94 TFLOPS), and more numerous cores (18,176 shading units versus 14,080) make it the stronger option for compute-heavy tasks like AI training, scientific simulation, and high-resolution rendering. The L40S also holds a higher average benchmark score (295,763 versus 182,109) and a higher performance percentile (99th versus 98th), indicating it sits at the very top of the GPU hierarchy.

However, the RTX PRO 5000 Blackwell is not without its merits. Its 7.7% win in Geekbench Vulkan demonstrates that the newer architecture and faster GDDR7 memory (1.34 TB/s versus 864.0 GB/s) can overcome a core-count deficit in certain workloads. The Blackwell card’s PCIe 5.0 interface, 4x DisplayPort 2.1b outputs, and higher base clock (1740 MHz versus 1110 MHz) make it a better fit for workstation environments where display connectivity, data transfer speed, and modern feature support matter more than raw TFLOPS.

The production status is a practical consideration. The L40S is end-of-life, which could complicate procurement and long-term support. The RTX PRO 5000 Blackwell is active with a launch MSRP of 5,099 USD, indicating ongoing availability. For builders who need a card that will remain supported and available, the Blackwell card is the safer choice. For those who prioritize maximum compute throughput above all else, the L40S remains the stronger performer despite its age.

In practical terms: pick the L40S for compute density and raw performance in AI or rendering workloads. Pick the RTX PRO 5000 Blackwell for memory bandwidth, modern display support, and a forward-looking platform with PCIe 5.0. The benchmark split — one win each — reflects a genuine trade-off, not a clear winner.

DETAILED SPECIFICATIONS

SPECIFICATION
L40S
RTX PRO 5000 Blackwell
Core Specs
Shading Units
18,176
14,080 -22.5%
Shaders
18,176
14,080 -22.5%
TMUs
568
440 -22.5%
ROPs
192
160 -16.7%
SM Count
142
110 -22.5%
Clocks
Base Clock
1110 MHz
1740 MHz
Boost Clock
2520 MHz
2377 MHz
Memory Clock
2250 MHz 18 Gbps effective
1750 MHz 28 Gbps effective
Memory
Memory Size
48 GB
48 GB
VRAM (MB)
49,152
49,152 0.0%
Memory Type
GDDR6
GDDR7
Memory Bus
384 bit
384 bit
Bandwidth
864.0 GB/s
1.34 TB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
48 MB
96 MB
Performance
Pixel Rate
483.8 GPixel/s
380.3 GPixel/s
Texture Rate
1,431.4 GTexel/s
1,045.9 GTexel/s
FP32 (TFLOPS)
91.61 TFLOPS
66.94 TFLOPS
FP64 (TFLOPS)
1,431.4 GFLOPS (1:64)
1,045.9 GFLOPS (1:64)
FP16 (TFLOPS)
91.61 TFLOPS (1:1)
66.94 TFLOPS (1:1)
AI/RT
RT Cores
142
110 -22.5%
Tensor Cores
568
440 -22.5%
Power
TDP
300 W
300 W
TDP (W)
300
300 0.0%
Suggested PSU
700 W
700 W
Power Connectors
1x 16-pin
1x 16-pin
Architecture
Architecture
Ada Lovelace
Blackwell 2.0
GPU Name
AD102
GB202
Generation
Server Ada (Lxx)
Blackwell PRO W (x000)
Process Size
5 nm
5 nm
Transistors
76,300 million
92,200 million
Die Size
609 mm²
750 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
122.9M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
12.0
Shader Model
6.8
6.9
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
4x DisplayPort 2.1b
Bus Interface
PCIe 4.0 x16
PCIe 5.0 x16
Other
Launch Price
—
5,099 USD
Production
End-of-life
Active
Predecessor
Server Ampere
Workstation Ada
Successor
Server Hopper
—
View L40S Details View RTX PRO 5000 Blackwell Details