AMD Radeon PRO V620 vs NVIDIA L4 Comparison

AMD
RADEON

AMD Radeon PRO V620

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 2200 MHz
TDP 300 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

L4

CORE STATE AD104
VRAM 24 GB
CLOCK SPEED 2040 MHz
TDP 72 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

geekbench_opencl
128,580
140,838
geekbench_vulkan
144,364
121,306

Analysis: AMD Radeon PRO V620 vs NVIDIA L4

The AMD Radeon PRO V620 and NVIDIA L4 are closely matched accelerators, but the data shows a clear split: the AMD card wins in Vulkan compute by a substantial 19% margin, while the NVIDIA card wins in OpenCL by 8.7%. The V620 is the pick for Vulkan-heavy workloads, while the L4 is the pick for OpenCL-centric tasks and scenarios where its dramatically lower power draw is paramount.

The Verdict

The benchmark results indicate a tie in wins, with each card taking one of the two tested workloads. The AMD Radeon PRO V620 delivers a decisive victory in the Geekbench Vulkan test, scoring 144,364 against the NVIDIA L4's 121,306, a 19% advantage. Conversely, the NVIDIA L4 wins the Geekbench OpenCL test, scoring 140,838 versus the V620's 128,580, an 8.7% lead. For users prioritizing Vulkan-based compute or rendering, the V620 is the clear choice. For those standardized on OpenCL or requiring minimal power consumption, the L4 is superior. The L4's 72 W TDP versus the V620's 300 W TDP is a massive differentiator, making the L4 far more suitable for dense, power-constrained server deployments. The V620, with its 32 GB of memory, is better suited for large dataset workloads.

Architecture Differences

The two cards are built on fundamentally different architectures from different eras. The AMD Radeon PRO V620 uses the Navi 21 chip based on the RDNA 2.0 architecture, fabricated on a 7 nm process at TSMC. In contrast, the NVIDIA L4 uses the AD104 chip based on the Ada Lovelace architecture, fabricated on a more advanced 5 nm process, also at TSMC. The L4's newer process node allows for a significantly higher transistor density of 121.8M / mm², compared to the V620's 51.5M / mm², though the V620's larger 520 mm² die size accommodates more total transistors at 26,800 million versus the L4's 35,800 million on a smaller 294 mm² die.

Architecturally, the V620 is a pure RDNA 2.0 design with 4608 shading units, 288 TMUs, and 128 ROPs. It has 72 ray tracing cores but no dedicated tensor cores. The NVIDIA L4, by contrast, has a higher count of 7424 shading units but fewer TMUs (240) and ROPs (80). It also has 60 ray tracing cores and, crucially, 240 tensor cores, which are absent from the AMD card. This difference is critical for AI and machine learning workloads that leverage tensor core acceleration. The L4 also supports FP16 at a 1:1 ratio with FP32, whereas the V620's FP16 performance is 2:1, indicating a different compute philosophy.

Head-to-Head Benchmarks

The two benchmark tests reveal opposite strengths. In the Geekbench OpenCL test, the NVIDIA L4 scores 140,838, outperforming the AMD Radeon PRO V620's 128,580 by 8.7%. This suggests the L4 has better raw compute throughput in this API, likely benefiting from its higher FP32 performance of 30.29 TFLOPS versus the V620's 20.28 TFLOPS. The L4's lead here is significant and aligns with its higher shading unit count.

The tables turn completely in the Geekbench Vulkan test. The AMD Radeon PRO V620 scores 144,364, a massive 19% improvement over the NVIDIA L4's 121,306. This is a substantial margin that indicates the V620's RDNA 2.0 architecture is more efficient in Vulkan workloads. Despite having lower theoretical FP32 throughput, the V620's Vulkan score is higher, suggesting better driver optimization or architectural efficiency for this API. The average benchmark scores reflect the split: the V620 averages 136,472 across both tests, while the L4 averages 131,072, putting the V620 at the 96th percentile versus the L4's 95th percentile of all GPUs.

FAQ

Q: Which card has higher benchmark scores overall?

A: The AMD Radeon PRO V620 has a higher average benchmark score of 136,472, compared to the NVIDIA L4's 131,072. However, the L4 wins the OpenCL test, and the V620 wins the Vulkan test.

Q: What is the difference in power consumption?

A: The NVIDIA L4 has a TDP of 72 W, which is dramatically lower than the AMD Radeon PRO V620's 300 W TDP. The L4 also requires no power connectors and suggests a 250 W PSU, while the V620 requires 2x 8-pin connectors and a 700 W PSU.

Q: Are these cards suitable for AI workloads?

A: The NVIDIA L4 is equipped with 240 tensor cores, which are designed for AI acceleration. The AMD Radeon PRO V620 does not have tensor cores.

Q: How does the memory configuration compare?

A: The AMD Radeon PRO V620 has 32 GB of GDDR6 memory with a 256-bit bus and 512.0 GB/s bandwidth. The NVIDIA L4 has 24 GB of GDDR6 memory with a 192-bit bus and 300.1 GB/s bandwidth.

Q: What is the physical size difference?

A: The AMD Radeon PRO V620 is a dual-slot card measuring 267 mm in length. The NVIDIA L4 is a single-slot card measuring 169 mm in length, making it significantly more compact.

Q: What is the production status of each card?

A: The AMD Radeon PRO V620 is marked as end-of-life, while the NVIDIA L4 is marked as active.

Where Each One Wins

The AMD Radeon PRO V620 wins decisively in Vulkan-based applications. Its Geekbench Vulkan score of 144,364 is 19% higher than the L4's, making it the better choice for software stacks that leverage this API. Its larger 32 GB memory pool and higher 512.0 GB/s bandwidth also give it an edge in scenarios with very large datasets that exceed the L4's 24 GB capacity. The V620 is the winner for users who prioritize Vulkan compute and need maximum memory capacity.

The NVIDIA L4 wins in OpenCL workloads, with an 8.7% higher score in that specific test. Its 30.29 TFLOPS FP32 performance is substantially higher than the V620's 20.28 TFLOPS, which likely contributes to its OpenCL advantage. The L4's inclusion of 240 tensor cores makes it the only choice for AI inference and training tasks that can utilize them. Its extremely low 72 W TDP, single-slot design, and shorter 169 mm length make it ideal for high-density server installations where space and power are at a premium.

Specification Differences

The two cards differ significantly across nearly every specification. The AMD Radeon PRO V620 uses the RDNA 2.0 architecture on a 7 nm process, while the NVIDIA L4 uses Ada Lovelace on a 5 nm process. The V620 has a larger die (520 mm²) but fewer transistors (26,800 million) compared to the L4's smaller die (294 mm²) with more transistors (35,800 million). The L4 has a much higher transistor density at 121.8M / mm² versus 51.5M / mm².

In terms of compute, the L4 has more shading units (7424 vs 4608) and higher FP32 performance (30.29 TFLOPS vs 20.28 TFLOPS). However, the V620 has more TMUs (288 vs 240) and ROPs (128 vs 80). The V620 has 72 ray tracing cores, while the L4 has 60. The most significant architectural difference is that the L4 has 240 tensor cores, while the V620 has none. The L4 also achieves FP16 performance at a 1:1 ratio with FP32, whereas the V620's FP16 performance is 2:1.

Memory is another major differentiator: the V620 has 32 GB of GDDR6 on a 256-bit bus offering 512.0 GB/s bandwidth, while the L4 has 24 GB on a 192-bit bus with 300.1 GB/s bandwidth. The V620's base clock is 1825 MHz with a boost of 2200 MHz, while the L4 has a much lower base clock of 795 MHz but a boost of 2040 MHz. The V620 is a dual-slot card at 267 mm long, while the L4 is a single-slot card at 169 mm long. The V620 uses 2x 8-pin power connectors and has a 300 W TDP, whereas the L4 has no power connectors and a 72 W TDP. The V620 is end-of-life and was released in 2021, while the L4 is active and was released in 2023.

DETAILED SPECIFICATIONS

SPECIFICATION
PRO V620
L4
Core Specs
Shading Units
4,608
7,424 +61.1%
Shaders
4,608
7,424 +61.1%
TMUs
288
240 -16.7%
ROPs
128
80 -37.5%
Compute Units
72
SM Count
60
Clocks
Base Clock
1825 MHz
795 MHz
Boost Clock
2200 MHz
2040 MHz
Memory Clock
2000 MHz 16 Gbps effective
1563 MHz 12.5 Gbps effective
Memory
Memory Size
32 GB
24 GB
VRAM (MB)
32,768
24,576 -25.0%
Memory Type
GDDR6
GDDR6
Memory Bus
256 bit
192 bit
Bandwidth
512.0 GB/s
300.1 GB/s
Cache
L1 Cache
128 KB per Array
128 KB (per SM)
L2 Cache
4 MB
48 MB
L3 Cache
128 MB
L0 Cache
32 KB per WGP
Performance
Pixel Rate
281.6 GPixel/s
163.2 GPixel/s
Texture Rate
633.6 GTexel/s
489.6 GTexel/s
FP32 (TFLOPS)
20.28 TFLOPS
30.29 TFLOPS
FP64 (TFLOPS)
1,267.2 GFLOPS (1:16)
473.3 GFLOPS (1:64)
FP16 (TFLOPS)
40.55 TFLOPS (2:1)
30.29 TFLOPS (1:1)
AI/RT
RT Cores
72
60 -16.7%
Tensor Cores
240
Power
TDP
300 W
72 W
TDP (W)
300
72 -76.0%
Suggested PSU
700 W
250 W
Power Connectors
2x 8-pin
None
Architecture
Architecture
RDNA 2.0
Ada Lovelace
GPU Name
Navi 21
AD104
Generation
Radeon Pro Navi (Navi II Series)
Server Ada (Lxx)
Process Size
7 nm
5 nm
Transistors
26,800 million
35,800 million
Die Size
520 mm²
294 mm²
Foundry
TSMC
TSMC
Density
51.5M / mm²
121.8M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.1
3.0
CUDA
8.9
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Single-slot
Length
267 mm 10.5 inches
169 mm 6.7 inches
Height
120 mm 4.7 inches
56 mm 2.2 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
Active
Predecessor
Radeon Pro Vega
Server Ampere
Successor
Server Hopper
View Radeon PRO V620 Details View L4 Details