AMD Radeon PRO V620 vs NVIDIA L40S Comparison
AMD Radeon PRO V620
L40S
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon PRO V620 vs NVIDIA L40S
# NVIDIA L40S vs AMD Radeon PRO V620
The NVIDIA L40S and AMD Radeon PRO V620 are both dual-slot, 300 W server accelerators aimed at professional workloads, but they represent vastly different generations and design philosophies. The L40S, built on NVIDIA's Ada Lovelace architecture, delivers an average benchmark score of 295,763, placing it in the 99th percentile of all GPUs. The Radeon PRO V620, based on AMD's RDNA 2.0 architecture, achieves an average score of 136,472, sitting in the 96th percentile. In direct head-to-head testing, the L40S wins both available benchmarks, with a 157.2% advantage in Geekbench OpenCL and an 80.7% advantage in Geekbench Vulkan. These are not close competitors; the data positions them in different performance tiers entirely.
Where Each One Wins
The benchmark results are unambiguous: the NVIDIA L40S wins every test in the comparison suite. In Geekbench OpenCL, the L40S scores 330,727 against the Radeon PRO V620's 128,580, a 157.2% delta. In Geekbench Vulkan, the L40S posts 260,799 versus 144,364, an 80.7% advantage. With two wins for the L40S and zero for the AMD card, the use-case split is one-sided.
However, the Radeon PRO V620 is not without a position in the market. Its average score of 136,472 places it within 0.5% of the AMD Radeon Pro W6800X Duo (135,774) and 0.8% of the AMD Radeon PRO W6800 (135,396). It also sits just 0.9% ahead of both the NVIDIA A10M (135,230) and the NVIDIA RTX 4000 Ada Generation (135,218). This clustering suggests the V620 is competitive with a specific tier of workstation and server GPUs, but that tier is far below the L40S. The L40S, by contrast, sits 3% ahead of the NVIDIA RTX 6000 Ada Generation (287,237) and 4.1% ahead of the NVIDIA L40 (284,111), while trailing the AMD Instinct MI300X (317,994) by 7% and the NVIDIA H200 NVL (334,891) by 11.7%.
For workloads that stress compute throughput, the L40S is the clear choice. For applications where the V620's particular feature set or availability matters, it remains a viable option, but the data shows no benchmark where it outperforms the L40S.
Architecture Differences
The two GPUs diverge sharply at the architectural level. The NVIDIA L40S uses the AD102 chip built on TSMC's 5 nm process, packing 76,300 million transistors into a 609 mm² die, yielding a transistor density of 125.3 million per mm². The AMD Radeon PRO V620 uses the Navi 21 chip on TSMC's 7 nm process, with 26,800 million transistors on a 520 mm² die, for a density of 51.5 million per mm². The L40S has nearly three times the transistor count on a slightly larger die, reflecting the more advanced process node.
The memory subsystems differ substantially. The L40S offers 48 GB of GDDR6 on a 384-bit bus, delivering 864.0 GB/s of bandwidth. The V620 provides 32 GB of GDDR6 on a 256-bit bus, with 512.0 GB/s of bandwidth. The L40S also runs its memory at 2250 MHz (18 Gbps effective), compared to the V620's 2000 MHz (16 Gbps effective). The L40S's 69% bandwidth advantage matters for large datasets and memory-bound workloads.
Compute resources heavily favor the L40S. It has 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The V620 has 4,608 shading units, 288 TMUs, 128 ROPs, and 72 RT cores, with no tensor cores listed. The L40S's FP32 throughput is 91.61 TFLOPS, while the V620 manages 20.28 TFLOPS. In FP16, the L40S maintains 91.61 TFLOPS at a 1:1 ratio, while the V620 reaches 40.55 TFLOPS at a 2:1 ratio. The L40S also leads in pixel rate (483.8 GPixel/s vs 281.6 GPixel/s) and texture rate (1,431.4 GTexel/s vs 633.6 GTexel/s).
Clock speeds tell a different story. The V620's base clock of 1825 MHz and boost clock of 2200 MHz are higher than the L40S's 1110 MHz base and 2520 MHz boost. The V620's higher base clock reflects its simpler architecture, but the L40S's boost clock ultimately exceeds it.
The Verdict
The data supports a clear verdict: the NVIDIA L40S is the superior performer in every measured benchmark. Its 157.2% OpenCL lead and 80.7% Vulkan lead are decisive margins that no workload characteristic of the V620 can bridge. The L40S's 99th percentile ranking, 4.1% ahead of the NVIDIA L40 and 3% ahead of the RTX 6000 Ada Generation, places it among the top-tier server accelerators. The Radeon PRO V620's 96th percentile ranking and its proximity to the W6800X Duo, W6800, A10M, and RTX 4000 Ada (all within 0.9%) define it as a mid-tier option.
Who should pick which? Based strictly on the data, any workload requiring maximum FP32, FP16, memory bandwidth, or ray tracing performance should choose the L40S. It offers 48 GB of memory versus 32 GB, 864.0 GB/s versus 512.0 GB/s, and 91.61 TFLOPS FP32 versus 20.28 TFLOPS. The V620's only advantages are its higher base clock and its dual 8-pin power connectors instead of a single 16-pin, but neither translates into a benchmark win. The V620 is a reasonable choice for systems already aligned with AMD's RDNA ecosystem, but the data shows no scenario where it matches the L40S.
FAQ
Q: Which GPU has higher raw compute performance?
A: The NVIDIA L40S delivers 91.61 TFLOPS FP32 and 91.61 TFLOPS FP16, while the AMD Radeon PRO V620 delivers 20.28 TFLOPS FP32 and 40.55 TFLOPS FP16.
Q: How much memory and bandwidth does each card offer?
A: The L40S has 48 GB of GDDR6 on a 384-bit bus with 864.0 GB/s bandwidth. The V620 has 32 GB of GDDR6 on a 256-bit bus with 512.0 GB/s bandwidth.
Q: What is the benchmark score difference?
A: The L40S averages 295,763 across benchmarks, compared to the V620's 136,472. In Geekbench OpenCL, the L40S scores 330,727 versus 128,580 (157.2% higher), and in Geekbench Vulkan it scores 260,799 versus 144,364 (80.7% higher).
Q: How do these cards compare to their nearest rivals?
A: The L40S is 3% ahead of the NVIDIA RTX 6000 Ada Generation and 4.1% ahead of the NVIDIA L40, while trailing the AMD Instinct MI300X by 7% and the NVIDIA H200 NVL by 11.7%. The V620 is 0.5% ahead of the AMD Radeon Pro W6800X Duo, 0.8% ahead of the AMD Radeon PRO W6800, 0.9% ahead of the NVIDIA A10M, and 0.9% ahead of the NVIDIA RTX 4000 Ada Generation.
Q: Do both cards support the same APIs?
A: Yes, both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.
Q: What are the physical differences in power and size?
A: Both are dual-slot, 300 W cards with a 700 W suggested PSU and PCIe 4.0 x16 interface. The L40S is 267 mm long and 111 mm high, while the V620 is 267 mm long, 120 mm high, and 50 mm wide. The L40S uses a 1x 16-pin connector; the V620 uses 2x 8-pin connectors.
Head-to-Head Benchmarks
The Geekbench OpenCL test shows the most extreme gap. The L40S scores 330,727, while the V620 scores 128,580, producing a 157.2% delta. This test typically stresses general-purpose compute across a wide range of operations, and the L40S's 91.61 TFLOPS FP32 throughput, 568 tensor cores, and 864.0 GB/s memory bandwidth overwhelm the V620's 20.28 TFLOPS and 512.0 GB/s. The V620's nearest rivals in this range — the W6800X Duo at 135,774, the W6800 at 135,396, the A10M at 135,230, and the RTX 4000 Ada at 135,218 — all cluster near its score, confirming that the V620 is not an outlier but rather part of a performance tier that the L40S surpasses by more than double.
The Geekbench Vulkan test narrows the gap but still favors the L40S decisively. The L40S scores 260,799 against the V620's 144,364, an 80.7% advantage. Vulkan workloads often leverage graphics and compute simultaneously, and here the L40S's 142 RT cores, 1,431.4 GTexel/s texture rate, and 483.8 GPixel/s pixel rate provide substantial headroom. The V620's 72 RT cores and 633.6 GTexel/s texture rate are respectable for its class, but they cannot compensate for the architectural gap. Notably, the L40S's Vulkan score of 260,799 is lower than its OpenCL score of 330,727, while the V620's Vulkan score of 144,364 is higher than its OpenCL score of 128,580. This suggests the V620 is relatively better optimized for Vulkan, but still far behind.
The aggregate data reinforces the same conclusion. The L40S's average benchmark score of 295,763 places it in the 99th percentile of all GPUs, while the V620's 136,472 sits in the 96th percentile. The percentile difference of 3 points understates the performance gap, because the L40S is competing with far more powerful accelerators at the top of the distribution.
Specification Differences
| Specification | NVIDIA L40S | AMD Radeon PRO V620 |
|---|---|---|
| Architecture | Ada Lovelace | RDNA 2.0 |
| Chip | AD102 | Navi 21 |
| Process Node | 5 nm | 7 nm |
| Transistors | 76,300 million | 26,800 million |
| Die Size | 609 mm² | 520 mm² |
| Transistor Density | 125.3M / mm² | 51.5M / mm² |
| Base Clock | 1110 MHz | 1825 MHz |
| Boost Clock | 2520 MHz | 2200 MHz |
| Memory Clock | 2250 MHz (18 Gbps effective) | 2000 MHz (16 Gbps effective) |
| Memory Size | 48 GB GDDR6 | 32 GB GDDR6 |
| Memory Bus Width | 384 bit | 256 bit |
| Memory Bandwidth | 864.0 GB/s | 512.0 GB/s |
| Shading Units | 18,176 | 4,608 |
| TMUs | 568 | 288 |
| ROPs | 192 | 128 |
| RT Cores | 142 | 72 |
| Tensor Cores | 568 | — |
| Pixel Rate | 483.8 GPixel/s | 281.6 GPixel/s |
| Texture Rate | 1,431.4 GTexel/s | 633.6 GTexel/s |
| FP32 Performance | 91.61 TFLOPS | 20.28 TFLOPS |
| FP16 Performance | 91.61 TFLOPS (1:1) | 40.55 TFLOPS (2:1) |
| Power Connectors | 1x 16-pin | 2x 8-pin |
| Dimensions | 267 mm × 111 mm | 267 mm × 120 mm × 50 mm |
| Display Outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a | No outputs |
| Release Date | 2022-10-12 | 2021-11-03 |
| Production Status | End-of-life | End-of-life |
| Predecessor | Server Ampere | Radeon Pro Vega |
| Successor | Server Hopper | — |
The L40S leads in every compute-relevant specification except base clock and physical width. Its 48 GB memory capacity and 864.0 GB/s bandwidth are critical for large model inference and rendering workloads. The V620's lack of display outputs and absence of tensor cores further differentiate it as a pure compute accelerator, whereas the L40S includes display connectivity. Both cards are end-of-life, with the L40S released on 2022-10-12 and the V620 on 2021-11-03. The L40S's successor is listed as Server Hopper, while the V620 has no listed successor.