AMD Radeon PRO V620 vs NVIDIA A10G Comparison
AMD Radeon PRO V620
A10G
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon PRO V620 vs NVIDIA A10G
The NVIDIA A10G and AMD Radeon PRO V620 are both end-of-life server accelerators targeting compute and rendering workloads, yet they deliver distinctly different performance profiles. The A10G, based on the GA102 chip in NVIDIA's Ampere architecture, leads in both benchmark categories recorded, winning 2 out of 2 head-to-head tests. Its most decisive victory comes in Geekbench OpenCL, where it scores 158,063 against the Radeon PRO V620's 128,580, a 22.9% advantage. The Vulkan result is far closer: the A10G scores 145,863 versus 144,364, a marginal 1% lead that puts the two within striking distance of each other. The overall average benchmark score reinforces this split — the A10G averages 151,963, placing it in the 97th percentile of all GPUs, while the Radeon PRO V620 averages 136,472, sitting in the 96th percentile. These numbers indicate a clear but not overwhelming overall performance edge for the NVIDIA part, with the caveat that the gap nearly vanishes in Vulkan workloads.
Head-to-Head Benchmarks
The Geekbench OpenCL test is where the A10G establishes its dominance. Scoring 158,063 against the Radeon PRO V620's 128,580, the NVIDIA card is 22.9% faster — a substantial margin that reflects its raw compute throughput. The A10G's FP32 performance of 31.52 TFLOPS nearly doubles the Radeon's 20.28 TFLOPS, and while OpenCL workloads vary, this test appears to favor the higher single-precision throughput. The A10G also benefits from a lower boost clock (1710 MHz vs. 2200 MHz) but compensates with 9,216 shading units, twice the Radeon's 4,608. The result is a decisive win that places the A10G well ahead of its closest rivals: it is 1.1% faster than the NVIDIA Tesla V100 PCIe 32 GB (avg score 150,305) and 9.3% faster than the AMD Instinct MI100 (avg score 139,035), though it trails the AMD Radeon Pro W6800X (avg score 160,671) by 5.4% and the NVIDIA A100 PCIe 40 GB (avg score 162,504) by 6.5%.
The Vulkan benchmark tells a different story. The A10G wins again, scoring 145,863 against 144,364, but the delta is just 1% — a statistical tie in practical terms. This near-parity suggests that the Radeon PRO V620's architecture handles Vulkan's API overhead and render paths more efficiently relative to its raw compute specs. The Radeon's higher texture rate (633.6 GTexel/s vs. 492.5 GTexel/s) and pixel rate (281.6 GPixel/s vs. 164.2 GPixel/s) likely contribute, as these metrics directly impact graphics-oriented tasks. Despite losing, the Radeon PRO V620's Vulkan score places it close to its own rivals: it is 0.5% ahead of the AMD Radeon Pro W6800X Duo (avg score 135,774), 0.8% ahead of the AMD Radeon PRO W6800 (avg score 135,396), and 0.9% ahead of both the NVIDIA A10M (avg score 135,230) and NVIDIA RTX 4000 Ada Generation (avg score 135,218). The A10G's Vulkan score of 145,863, by contrast, is above all of these, reinforcing its overall 97th percentile ranking.
Where Each One Wins
The A10G wins outright in OpenCL, making it the stronger choice for compute-heavy workloads that rely on OpenCL kernels — typical of scientific simulation, machine learning inference, and data processing tasks. Its 31.52 TFLOPS FP32 throughput and 24 GB of GDDR6 memory with 600.2 GB/s bandwidth provide a wide pipeline for such workloads. The Radeon PRO V620, despite losing overall, has clear strengths in graphics-oriented or Vulkan-accelerated tasks. Its Vulkan score is only 1% behind, and its higher pixel rate (281.6 GPixel/s) and texture rate (633.6 GTexel/s) indicate superior rasterization and texture-fetch capabilities. This makes it competitive for rendering workloads, where the A10G's lower pixel rate (164.2 GPixel/s) could become a bottleneck. The Radeon also holds an advantage in memory capacity: 32 GB versus 24 GB, which is critical for large datasets that exceed the A10G's frame buffer. However, the Radeon's memory bandwidth is lower at 512.0 GB/s versus 600.2 GB/s, so the A10G moves data faster once it fits in memory. In FP16 workloads, the Radeon PRO V620 offers 40.55 TFLOPS at a 2:1 ratio versus the A10G's 31.52 TFLOPS at 1:1 — a meaningful edge for mixed-precision AI inference, though the A10G's dedicated 288 tensor cores (which the Radeon lacks) can accelerate certain neural network ops.
Architecture Differences
The A10G is built on NVIDIA's Ampere architecture, fabricated on Samsung's 8 nm process with 28,300 million transistors on a 628 mm² die, yielding a transistor density of 45.1M per mm². The Radeon PRO V620 uses AMD's RDNA 2.0 architecture, manufactured by TSMC on a 7 nm process with 26,800 million transistors on a 520 mm² die, achieving a higher density of 51.5M per mm². The node advantage gives AMD a more compact design, but NVIDIA compensates with a larger die and more total transistors. The A10G features 9,216 shading units, 288 TMUs, and 96 ROPs, alongside 72 RT cores and 288 tensor cores — the latter being absent from the Radeon PRO V620. The Radeon counters with 4,608 shading units (half the A10G), 288 TMUs (identical), and 128 ROPs (33% more). Both have 72 RT cores, but the Radeon's ray-tracing implementation is part of RDNA 2.0's unified compute units, whereas the A10G uses dedicated RT cores. The Radeon's FP16 throughput of 40.55 TFLOPS at a 2:1 ratio indicates a packed math path, while the A10G's FP16 is 1:1 with FP32 at 31.52 TFLOPS, reflecting NVIDIA's design choice to prioritize consistent throughput across precisions. The A10G's memory subsystem uses a 384-bit bus with 24 GB of GDDR6 at 12.5 Gbps effective, achieving 600.2 GB/s; the Radeon uses a narrower 256-bit bus with 32 GB of GDDR6 at 16 Gbps effective, but only reaches 512.0 GB/s due to the reduced bus width. Both support PCIe 4.0 x16 and share identical API support for DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.
Specification Differences
The two cards diverge on nearly every core specification. The process node differs: 8 nm (Samsung) for the A10G versus 7 nm (TSMC) for the Radeon PRO V620. Transistor counts are close but not equal — 28,300 million for NVIDIA versus 26,800 million for AMD — and die sizes are 628 mm² versus 520 mm², respectively. The A10G's base clock is 1320 MHz with a boost of 1710 MHz, while the Radeon runs much higher at 1825 MHz base and 2200 MHz boost. Memory configurations are a major split: the A10G offers 24 GB on a 384-bit bus with 600.2 GB/s bandwidth, while the Radeon offers 32 GB on a 256-bit bus with 512.0 GB/s. Shading units favor NVIDIA at 9,216 versus 4,608, but ROPs favor AMD at 128 versus 96; TMUs are tied at 288. The A10G has 288 tensor cores, the Radeon has none. FP32 throughput is 31.52 TFLOPS versus 20.28 TFLOPS; FP16 is 31.52 TFLOPS (1:1) versus 40.55 TFLOPS (2:1). Pixel rate is 164.2 GPixel/s versus 281.6 GPixel/s, and texture rate is 492.5 GTexel/s versus 633.6 GTexel/s. Power draw differs substantially: the A10G is rated at 150 W TDP with a single 8-pin EPS connector and a suggested 450 W PSU, whereas the Radeon draws 300 W with dual 8-pin connectors and a suggested 700 W PSU. Physical dimensions are close in length (both 267 mm / 10.5 inches), but the A10G is single-slot at 112 mm height, while the Radeon is dual-slot at 120 mm height and 50 mm width. Neither card has display outputs. Release dates differ by about seven months: the A10G launched on 2021-04-11, and the Radeon on 2021-11-03. The A10G's predecessor is Tesla Turing and its successor is Server Ada; the Radeon's predecessor is Radeon Pro Vega with no listed successor.
FAQ
Q: Which card has higher raw FP32 compute performance?
A: The NVIDIA A10G, with 31.52 TFLOPS, is over 50% ahead of the AMD Radeon PRO V620's 20.28 TFLOPS. This is reflected in the OpenCL benchmark, where the A10G leads by 22.9%.
Q: Does the Radeon PRO V620 have any performance advantage?
A: Yes, in FP16 throughput it offers 40.55 TFLOPS at a 2:1 ratio, versus the A10G's 31.52 TFLOPS at 1:1. It also has higher pixel rate (281.6 GPixel/s vs. 164.2 GPixel/s) and texture rate (633.6 GTexel/s vs. 492.5 GTexel/s), which explains its near-parity Vulkan score.
Q: How do the two compare in memory capacity and bandwidth?
A: The Radeon PRO V620 has more capacity at 32 GB, but the A10G has higher bandwidth at 600.2 GB/s versus 512.0 GB/s. The A10G's 384-bit bus versus the Radeon's 256-bit bus is the key difference.
Q: Are these cards still in production?
A: No, both are end-of-life. The A10G was released on 2021-04-11, and the Radeon PRO V620 on 2021-11-03.
Q: Which card has tensor cores?
A: Only the NVIDIA A10G, which includes 288 tensor cores. The AMD Radeon PRO V620 has no tensor core equivalent in the FACT PACK.
Q: What is the power requirement difference?
A: The A10G has a 150 W TDP with a suggested 450 W PSU, while the Radeon PRO V620 has a 300 W TDP with a suggested 700 W PSU. The Radeon also requires dual 8-pin connectors versus the A10G's single 8-pin EPS.