AMD Radeon PRO V620 vs NVIDIA L40 Comparison
AMD Radeon PRO V620
L40
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon PRO V620 vs NVIDIA L40
The NVIDIA L40 is the clear performance leader in this comparison, dominating the AMD Radeon PRO V620 across every available benchmark metric. The data shows a decisive advantage for the NVIDIA part, with the L40 posting an average benchmark score of 284,111 compared to the V620’s 136,472, placing the L40 in the 99th percentile of all GPUs versus the V620’s 96th percentile.
Head-to-Head Benchmarks
The most striking result comes from the Geekbench OpenCL test, where the NVIDIA L40 scores 330,926 points against the AMD Radeon PRO V620’s 128,580 points. This represents a 157.4% advantage for the L40, meaning it delivers more than two and a half times the raw compute performance in this workload. The margin is so large that it dwarfs the differences seen between the two cards’ nearest rivals.
In the Geekbench Vulkan test, the gap narrows somewhat but remains substantial. The L40 scores 237,295 points while the V620 manages 144,364 points, giving the NVIDIA card a 64.4% lead. This suggests that while the V620 is relatively stronger in Vulkan than in OpenCL, it still cannot match the L40’s absolute performance in either API.
Looking at the rival landscape provides additional context. The L40’s average score of 284,111 places it just 1.1% behind the NVIDIA RTX 6000 Ada Generation (287,237) and 3.9% behind the NVIDIA L40S (295,763). It sits 13.1% ahead of the NVIDIA L20 (251,147) and 10.7% behind the AMD Instinct MI300X (317,994). The V620, by contrast, sits in a much lower performance tier, with its 136,472 average score coming in just 0.5% ahead of the AMD Radeon Pro W6800X Duo (135,774) and 0.8% ahead of the AMD Radeon PRO W6800 (135,396). The V620 also edges out the NVIDIA A10M (135,230) and NVIDIA RTX 4000 Ada Generation (135,218) by 0.9% each.
The win tally is unambiguous: the L40 takes both head-to-head benchmarks, giving it 2 wins against 0 for the V620. In the OpenCL test, the L40’s score is more than 2.5 times higher than the V620’s, while in Vulkan it is roughly 1.6 times higher. No benchmark in the data shows the AMD card winning or even coming close.
Architecture Differences
The two GPUs are built on fundamentally different architectures and manufacturing processes. The NVIDIA L40 uses the AD102 chip based on the Ada Lovelace architecture, fabricated on a 5 nm process at TSMC. The AMD Radeon PRO V620 uses the Navi 21 chip based on RDNA 2.0, also from TSMC but on a larger 7 nm node. This process advantage helps explain the L40’s enormous transistor count of 76,300 million, nearly three times the V620’s 26,800 million transistors, despite the L40’s die size of 609 mm² being only modestly larger than the V620’s 520 mm². The resulting transistor density tells the story: the L40 packs 125.3 million transistors per mm², versus just 51.5 million per mm² for the V620.
Memory configurations differ substantially. The L40 carries 48 GB of GDDR6 memory on a 384-bit bus, delivering 864.0 GB/s of bandwidth at 18 Gbps effective speed. The V620 has 32 GB of GDDR6 on a 256-bit bus, yielding 512.0 GB/s at 16 Gbps effective. The L40’s memory bandwidth advantage of 68.8% is significant for large datasets and compute workloads.
Compute resources show a stark disparity. The L40 features 18,176 shading units, 568 TMUs, and 192 ROPs, along with 142 RT cores and 568 tensor cores. The V620 has only 4,608 shading units, 288 TMUs, and 128 ROPs, with 72 RT cores and no tensor cores at all. These differences translate directly into peak rates: the L40 achieves 90.52 TFLOPS FP32 and 90.52 TFLOPS FP16 (1:1 ratio), while the V620 manages 20.28 TFLOPS FP32 and 40.55 TFLOPS FP16 (2:1 ratio). The L40 also leads in pixel rate (478.1 GPixel/s vs 281.6 GPixel/s) and texture rate (1,414.3 GTexel/s vs 633.6 GTexel/s).
Clock speeds tell a different story, however. The V620 runs at a base clock of 1825 MHz and a boost of 2200 MHz, compared to the L40’s 735 MHz base and 2490 MHz boost. The AMD card’s higher base clock reflects its simpler architecture, but the L40’s superior IPC and massive core count more than compensate. Power draw is identical at 300 W TDP for both cards, with both suggesting a 700 W power supply, though the L40 uses a single 16-pin connector while the V620 requires two 8-pin connectors.
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA L40 has an average benchmark score of 284,111, compared to 136,472 for the AMD Radeon PRO V620. This gives the L40 the 99th percentile ranking among all GPUs, while the V620 sits in the 96th percentile.
Q: How much faster is the L40 in OpenCL?
A: The L40 scores 330,926 in Geekbench OpenCL, which is 157.4% higher than the V620’s 128,580. This is the largest performance gap between the two cards in any tested workload.
Q: Does the Radeon PRO V620 win in any benchmark?
A: No. The head-to-head data shows the L40 winning both the Geekbench OpenCL and Geekbench Vulkan tests. The V620 has zero wins in the available benchmark comparisons.
Q: What is the memory bandwidth difference?
A: The L40 provides 864.0 GB/s of memory bandwidth from 48 GB of GDDR6 on a 384-bit bus. The V620 offers 512.0 GB/s from 32 GB of GDDR6 on a 256-bit bus, giving the L40 a 68.8% bandwidth advantage.
Q: Do both cards support the same APIs?
A: Yes, both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. However, the L40 adds tensor cores, which are absent from the V620, enabling different compute capabilities.
Q: What are the physical size differences?
A: Both cards are dual-slot and 267 mm long. The L40 is 111 mm tall while the V620 is 120 mm tall and 50 mm wide. The L40 has four DisplayPort 1.4a outputs, while the V620 has no display outputs at all.
The Verdict
The data is unequivocal: the NVIDIA L40 is the superior GPU in every measured category. Its average benchmark score of 284,111 is more than double the V620’s 136,472, and it wins both head-to-head tests by margins of 157.4% and 64.4%. The L40 offers more than four times the FP32 compute (90.52 TFLOPS vs 20.28 TFLOPS), 50% more memory capacity (48 GB vs 32 GB), and 68.8% more memory bandwidth. It also benefits from a more advanced 5 nm process, higher transistor density, and the presence of tensor cores that the AMD card lacks entirely.
The V620 does have advantages in base clock speed (1825 MHz vs 735 MHz) and a lower profile in terms of physical height, but these do not translate into competitive performance. The V620’s nearest rivals are all in the 135,000-136,000 score range, while the L40 competes with cards scoring between 251,000 and 318,000. The performance tiers are simply different.
For any workload that stresses raw compute, memory bandwidth, or ray tracing, the L40 is the clear choice. The V620’s only conceivable role in a comparison is as a lower-cost alternative, but even that consideration is outside the scope of the performance data presented. The verdict is straightforward: the NVIDIA L40 is the definitive winner, with the AMD Radeon PRO V620 trailing by a wide margin across all benchmark results.
Specification Differences
The key differing specifications between the two cards are numerous. The L40 uses the AD102 chip on 5 nm with 76,300 million transistors and a 609 mm² die, while the V620 uses Navi 21 on 7 nm with 26,800 million transistors and a 520 mm² die. The L40 has a base clock of 735 MHz and boost of 2490 MHz, versus the V620’s 1825 MHz base and 2200 MHz boost. Memory differs with 48 GB GDDR6 at 18 Gbps effective on a 384-bit bus for the L40, versus 32 GB GDDR6 at 16 Gbps on a 256-bit bus for the V620. Compute units show 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores for the L40, against 4,608 shading units, 288 TMUs, 128 ROPs, 72 RT cores, and no tensor cores for the V620. Peak rates are 90.52 TFLOPS FP32 and FP16 for the L40, while the V620 reaches 20.28 TFLOPS FP32 and 40.55 TFLOPS FP16. Power connectors are 1x 16-pin for the L40 and 2x 8-pin for the V620. The L40 measures 111 mm tall with 4x DisplayPort 1.4a outputs, while the V620 is 120 mm tall and 50 mm wide with no display outputs. The V620 was released on 2021-11-03, with the L40 following on 2022-10-12.
Where Each One Wins
The NVIDIA L40 wins in every performance category where data exists. Its 90.52 TFLOPS FP32 and FP16 throughput makes it suitable for compute-heavy tasks like AI inference, scientific simulation, and rendering, especially with tensor cores available. The 48 GB memory capacity and 864.0 GB/s bandwidth support large datasets and high-resolution textures. The 4x DisplayPort 1.4a outputs enable direct display connectivity, which the V620 lacks entirely. The L40’s 99th percentile ranking places it among the top GPUs globally.
The AMD Radeon PRO V620 has no benchmark wins to claim. Its higher base clock of 1825 MHz does not compensate for its lower core count and older architecture. The card’s 32 GB memory and 512.0 GB/s bandwidth are adequate for moderate workloads, but they trail the L40 by significant margins. The V620’s 96th percentile ranking is respectable, but it sits in a completely different performance class. The only areas where the V620 could be considered preferable are its smaller physical height (120 mm vs 111 mm is actually taller, so this is not an advantage) and its 2x 8-pin power connectors being more common in existing infrastructure. Neither factor appears in benchmark results, and the data shows no scenario where the V620 outperforms the L40. For any user prioritizing compute performance, memory bandwidth, or ray tracing capabilities, the L40 is the only logical choice from these two options.