AMD Radeon Pro VII vs NVIDIA L40S Comparison
AMD Radeon Pro VII
L40S
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon Pro VII vs NVIDIA L40S
The recorded data shows a lopsided contest: the NVIDIA L40S wins both shared benchmark tests against the AMD Radeon Pro VII, and it wins them by margins rarely seen in this database. The Radeon Pro VII counterpunches in exactly one metric, memory bandwidth, where its HBM2 subsystem delivers more raw throughput than the L40S despite being the older and far smaller chip. This page walks through the architecture, the benchmark numbers, and the narrow conditions under which the Radeon Pro VII still makes sense.
The Verdict
The database places the L40S in the 99th percentile of all GPUs, with an average benchmark score of 295763. The Radeon Pro VII sits in the 93rd percentile at 97131. That is a substantial gap in class, and the head-to-head results confirm it: the L40S leads by 266.9 percent in Geekbench OpenCL and 180.8 percent in Geekbench Vulkan. For any workload captured by these compute-oriented tests, the L40S is the clear pick.
The Radeon Pro VII's case rests on its memory subsystem and launch MSRP of 1,899 USD, stated here once for the record. Its HBM2 memory moves 1.02 TB/s across a 4096-bit bus, comfortably ahead of the L40S at 864.0 GB/s. For bandwidth-bound workloads insensitive to raw compute throughput, that advantage is the one concrete reason to prefer it. Both cards are end-of-life, so availability rather than capability may decide real-world purchases.
Architecture Differences
These two products come from different eras and different design philosophies. The L40S is built on the AD102 die, Ada Lovelace architecture, a 5 nm TSMC process, 76,300 million transistors, and a 609 mm² die with a density of 125.3M per mm². The Radeon Pro VII uses the Vega 20 die, GCN 5.1 architecture, a 7 nm TSMC process, 13,230 million transistors, and a 331 mm² die at 40.0M per mm². The L40S carries roughly 5.8 times the transistor count on a die nearly twice the area, and the generational node jump from 7 nm to 5 nm shows up directly in that density figure.
The compute resource gap is stark. The L40S fields 18176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The Radeon Pro VII has 3840 shading units, 240 TMUs, 64 ROPs, and no RT cores or tensor cores listed at all. Ray tracing and dedicated matrix hardware simply are not part of the Vega 20 feature set as recorded here. The resulting throughput numbers diverge accordingly: 91.61 TFLOPS FP32 for the L40S versus 13.06 TFLOPS for the Radeon Pro VII, and 91.61 TFLOPS FP16 at a 1:1 rate versus 26.11 TFLOPS at a 2:1 rate.
Clocking follows opposite strategies. The L40S runs a 1110 MHz base and a 2520 MHz boost, a wide range typical of modern boost behavior. The Radeon Pro VII runs 1400 MHz base and 1700 MHz boost, a tighter band. Memory clocks diverge even more: 18 Gbps effective GDDR6 on the L40S versus 2 Gbps effective HBM2 on the Radeon Pro VII, with the Radeon Pro VII compensating through sheer bus width, 4096 bits against 384, to take the bandwidth crown.
Memory capacity also favors the L40S: 48 GB of GDDR6 against 16 GB of HBM2. Both use PCIe 4.0 x16 interfaces. The L40S carries a 300 W TDP with a single 16-pin connector and a suggested 700 W PSU; the Radeon Pro VII is rated at 250 W with one 6-pin plus one 8-pin connector and a suggested 600 W PSU. Physically, both are dual-slot cards of the same height, 111 mm, but the Radeon Pro VII is longer at 305 mm versus 267 mm. Display output philosophies differ sharply: the L40S offers one HDMI 2.1 plus three DisplayPort 1.4a outputs, while the Radeon Pro VII dedicates six mini-DisplayPort 1.4a connectors to it. API support is more modern on the L40S: DirectX 12 Ultimate (12_2) and Vulkan 1.4 against DirectX 12 (12_1) and Vulkan 1.3, with OpenGL 4.6 common to both.
Head-to-Head Benchmarks
Only two tests overlap across both cards in the database, and the L40S sweeps them.
Geekbench OpenCL: the L40S scores 330727, the Radeon Pro VII scores 90148, a 266.9 percent margin for the L40S. To put that in context, the L40S's OpenCL result sits above its own nearest rivals cluster, which includes the RTX 6000 Ada Generation at 287237 average and the L40 at 284111, and even against the AMD Instinct MI300X at 317994 average the L40S holds a 7 percent edge per the recorded delta. The Radeon Pro VII's OpenCL score of 90148, by contrast, lands it just below the AMD Radeon Instinct MI60 at 92466 average and slightly under the NVIDIA Quadro RTX 6000 at 101872 average.
Geekbench Vulkan: the L40S scores 260799 against 92862, a 180.8 percent win. Notably, the Radeon Pro VII's Vulkan result exceeds its own OpenCL result, indicating its strongest recorded compute path is Vulkan, yet it still trails the L40S by nearly a factor of three. The Radeon Pro VII also posts a Geekbench Metal score of 108383 on a platform the L40S does not record, which is its best single benchmark number but not comparable head-to-head.
The aggregate picture: 2 wins for the L40S, 0 for the Radeon Pro VII, average scores of 295763 versus 97131. In the L40S's rival table, only the NVIDIA H200 NVL at 334891 average beats it, by 11.7 percent. The Radeon Pro VII's rivals are far closer: the AMD Radeon RX 7900M sits at 97487, just 0.4 percent ahead, while the Radeon Pro VII leads the NVIDIA RTX A4500 by 6 percent and the MI60 by 5 percent. In other words, the L40S competes with datacenter flagships while the Radeon Pro VII trades blows with midrange professional cards.
FAQ
Q: Which card is faster overall?
A: The NVIDIA L40S, decisively. It wins both shared tests, leads by 266.9 percent in Geekbench OpenCL and 180.8 percent in Geekbench Vulkan, and holds a 295763 average score versus 97131.
Q: Does the Radeon Pro VII win anything?
A: Yes, memory bandwidth. Its HBM2 memory delivers 1.02 TB/s over a 4096-bit bus, ahead of the L40S at 864.0 GB/s, though the L40S counters with 48 GB of capacity against 16 GB.
Q: How much more compute throughput does the L40S have?
A: FP32 is 91.61 TFLOPS versus 13.06 TFLOPS, and FP16 is 91.61 TFLOPS at 1:1 versus 26.11 TFLOPS at 2:1. The L40S also has 568 tensor cores and 142 RT cores where the database lists none for the Radeon Pro VII.
Q: Are both cards still in production?
A: No. Both are recorded as end-of-life. The Radeon Pro VII launched on May 12, 2020, and the L40S on October 12, 2022.
Q: How do their power requirements compare?
A: The L40S has a 300 W TDP, a single 16-pin connector, and a suggested 700 W PSU. The Radeon Pro VII has a 250 W TDP, one 6-pin plus one 8-pin connector, and a suggested 600 W PSU.
Q: How do they rank against the broader GPU field?
A: The L40S sits in the 99th percentile of all GPUs in the database; the Radeon Pro VII sits in the 93rd.
Where Each One Wins
The L40S wins every general compute scenario the database measures. OpenCL and Vulkan workloads, any task leaning on FP32 or FP16 throughput, anything requiring ray tracing or tensor hardware, and any job that needs large memory capacity, since 48 GB triples the Radeon Pro VII's 16 GB. Its rival positioning, within a few percent of the RTX 6000 Ada Generation and L40, ahead of the Instinct MI300X by 7 percent, and behind only the H200 NVL by 11.7 percent, confirms it as a compute-class part. It is also the shorter card at 267 mm, which eases integration in space-constrained chassis, and its more modern API support (DirectX 12 Ultimate, Vulkan 1.4) future-proofs software compatibility.
The Radeon Pro VII wins in a narrower set of circumstances. Its 1.02 TB/s of HBM2 bandwidth leads the pairing, which matters for bandwidth-saturated workloads where arithmetic intensity is low and raw compute is not the bottleneck. Its six mini-DisplayPort 1.4a outputs make it a natural fit for multi-monitor walls and dense visualization setups that exceed the L40S's four outputs. Its 250 W TDP and 600 W suggested PSU are lighter demands than the L40S's 300 W and 700 W. And in its own competitive bracket, the data shows it trading within single-digit percentages of the RX 7900M, MI60, Quadro RTX 6000, and RTX A4500, so buyers comparing against that tier rather than against the L40S's tier will find it far more evenly matched.
The summary judgment from the recorded data is simple: for compute, the L40S dominates; for bandwidth density and display count, the Radeon Pro VII retains two defensible niches.