AMD Radeon PRO V620 vs NVIDIA A100 PCIe 40 GB Comparison

AMD
RADEON

AMD Radeon PRO V620

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 2200 MHz
TDP 300 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

A100 PCIe 40 GB

CORE STATE GA100
VRAM 40 GB
CLOCK SPEED 1410 MHz
TDP 250 W
BUS WIDTH 5120 bit
ARCHITECTURE Ampere
nm
PROCESS 7 nm
LAUNCH DATE 2020

PERFORMANCE BENCHMARKS

geekbench_opencl
128,580
178,627
geekbench_vulkan
144,364
146,380

Analysis: AMD Radeon PRO V620 vs NVIDIA A100 PCIe 40 GB

The Verdict

The benchmark data places these two accelerators in different tiers despite their shared server-oriented positioning. The NVIDIA A100 PCIe 40 GB holds a decisive lead in the aggregate, with an average benchmark score of 162,504 against the AMD Radeon PRO V620's 136,472 — a gap of roughly 19.1%. The A100 also sits at the 97th percentile of all GPUs, one point above the V620's 96th percentile. In the two available head-to-head tests, the A100 wins both: a commanding 38.9% margin in Geekbench OpenCL and a narrower 1.4% edge in Geekbench Vulkan.

The intended user splits cleanly. The A100 is the choice for workloads that stress general-purpose compute through OpenCL, where its 178,627 score versus 128,580 represents a 50,047-point advantage. The V620, however, is not without merit; its Vulkan result of 144,364 trails the A100's 146,380 by only 2,016 points, suggesting that graphics-adjacent or Vulkan-optimized tasks narrow the gap considerably. The data indicates that buyers prioritizing raw compute throughput — particularly in OpenCL-heavy environments — should select the A100, while those whose workloads lean on Vulkan and can tolerate lower OpenCL performance may find the V620 sufficient, especially given its higher base and boost clocks.

The A100's nearest rivals bracket its performance tightly: it is 1.1% ahead of the AMD Radeon Pro W6800X, 1.4% behind the AMD Radeon PRO W7800, and 1.6% behind the NVIDIA RTX A5500. The V620, by contrast, sits in a cluster where its closest competitors are within 0.9% — the AMD Radeon Pro W6800X Duo (0.5% ahead), the AMD Radeon PRO W6800 (0.8% ahead), the NVIDIA A10M (0.9% ahead), and the NVIDIA RTX 4000 Ada Generation (0.9% ahead). This indicates the V620 is a mid-pack performer among its peers, whereas the A100 is near the top of its class.

FAQ

Q: Which card wins in OpenCL performance?

A: The NVIDIA A100 PCIe 40 GB wins decisively, scoring 178,627 in Geekbench OpenCL versus 128,580 for the AMD Radeon PRO V620 — a 38.9% advantage.

Q: How close are the two cards in Vulkan performance?

A: The gap narrows dramatically. The A100 scores 146,380 in Geekbench Vulkan, while the V620 scores 144,364, a difference of only 1.4%.

Q: What is the memory configuration difference?

A: The A100 uses 40 GB of HBM2e on a 5120-bit bus with 1.56 TB/s bandwidth. The V620 uses 32 GB of GDDR6 on a 256-bit bus with 512.0 GB/s bandwidth. The A100's bandwidth is roughly three times higher.

Q: Which card has higher clock speeds?

A: The AMD Radeon PRO V620 runs at a base clock of 1825 MHz and a boost clock of 2200 MHz. The NVIDIA A100 PCIe 40 GB operates at a base clock of 765 MHz and a boost clock of 1410 MHz. The V620's clocks are substantially higher.

Q: Do both cards have the same physical dimensions?

A: Both are dual-slot cards with a length of 267 mm (10.5 inches). The V620 is taller at 120 mm (4.7 inches) versus 111 mm (4.4 inches) for the A100, and the V620 adds a 50 mm (2 inches) width dimension not listed for the A100.

Q: What is the transistor count difference?

A: The A100's GA100 chip contains 54,200 million transistors on an 826 mm² die. The V620's Navi 21 chip contains 26,800 million transistors on a 520 mm² die. The A100 has roughly double the transistors and a larger die area.

Architecture Differences

The NVIDIA A100 PCIe 40 GB is built on the Ampere architecture using the GA100 chip, fabricated on a 7 nm process by TSMC. It packs 54,200 million transistors into an 826 mm² die, yielding a transistor density of 65.6 million per mm². The AMD Radeon PRO V620 uses the RDNA 2.0 architecture with the Navi 21 chip, also on TSMC's 7 nm node, but with 26,800 million transistors on a 520 mm² die — a density of 51.5 million per mm². The A100's design emphasizes massive compute throughput with its 6912 shading units and 432 tensor cores, while the V620 counters with 4608 shading units and 72 ray tracing cores, a feature the A100 lacks entirely.

The A100's memory subsystem is fundamentally different: HBM2e across a 5120-bit bus delivers 1.56 TB/s of bandwidth, compared to the V620's GDDR6 on a 256-bit bus at 512.0 GB/s. This threefold bandwidth advantage is a hallmark of the A100's server-compute heritage. The V620, meanwhile, relies on higher clock speeds — a 1825 MHz base and 2200 MHz boost versus 765 MHz and 1410 MHz for the A100 — to compensate for its narrower memory interface. The V620 also supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, while the A100 lists no API support in the data, reflecting its compute-first, display-less design.

The A100's tensor cores (432 of them) position it for AI and machine learning workloads, a capability absent from the V620's spec sheet. Conversely, the V620's ray tracing cores (72) give it hardware acceleration for graphics tasks the A100 cannot perform. Both cards are end-of-life products with no display outputs, but their architectural priorities diverge sharply: the A100 maximizes parallel compute and memory bandwidth, while the V620 balances compute with graphics features and clock-speed-driven performance.

Specification Differences

| Specification | NVIDIA A100 PCIe 40 GB | AMD Radeon PRO V620 |

|---|---|---|

| Architecture | Ampere | RDNA 2.0 |

| Transistors | 54,200 million | 26,800 million |

| Die Size | 826 mm² | 520 mm² |

| Transistor Density | 65.6M / mm² | 51.5M / mm² |

| Base Clock | 765 MHz | 1825 MHz |

| Boost Clock | 1410 MHz | 2200 MHz |

| Memory Size | 40 GB | 32 GB |

| Memory Type | HBM2e | GDDR6 |

| Memory Bus Width | 5120 bit | 256 bit |

| Memory Bandwidth | 1.56 TB/s | 512.0 GB/s |

| Memory Clock | 1215 MHz (2.4 Gbps effective) | 2000 MHz (16 Gbps effective) |

| Shading Units | 6912 | 4608 |

| TMUs | 432 | 288 |

| ROPs | 160 | 128 |

| Ray Tracing Cores | None | 72 |

| Tensor Cores | 432 | None |

| FP32 Performance | 19.49 TFLOPS | 20.28 TFLOPS |

| FP16 Performance | 77.97 TFLOPS (4:1) | 40.55 TFLOPS (2:1) |

| Pixel Rate | 225.6 GPixel/s | 281.6 GPixel/s |

| Texture Rate | 609.1 GTexel/s | 633.6 GTexel/s |

| TDP | 250 W | 300 W |

| Power Connectors | 8-pin EPS | 2x 8-pin |

| Suggested PSU | 600 W | 700 W |

| Height | 111 mm (4.4 inches) | 120 mm (4.7 inches) |

| Width | Not listed | 50 mm (2 inches) |

| Release Date | 2020-06-21 | 2021-11-03 |

| Predecessor | Tesla Turing | Radeon Pro Vega |

| Successor | Server Ada | None listed |

The FP32 figures are nearly equivalent — 19.49 TFLOPS for the A100 versus 20.28 TFLOPS for the V620 — yet the A100's FP16 output of 77.97 TFLOPS is nearly double the V620's 40.55 TFLOPS, indicating the A100's enhanced precision-flexible tensor throughput. The V620 has higher pixel and texture rates, at 281.6 GPixel/s and 633.6 GTexel/s respectively, compared to the A100's 225.6 GPixel/s and 609.1 GTexel/s.

Head-to-Head Benchmarks

The most significant performance differential appears in Geekbench OpenCL, where the NVIDIA A100 PCIe 40 GB scores 178,627 against the AMD Radeon PRO V620's 128,580. This 38.9% delta is the largest of any metric in the comparison and drives the A100's overall advantage. The OpenCL result aligns with the architectural data: the A100's 1.56 TB/s memory bandwidth, 432 tensor cores, and 6912 shading units provide a formidable foundation for compute-heavy workloads, while the V620's higher clocks and 20.28 TFLOPS FP32 cannot overcome its narrower memory pipeline and fewer shading units.

In Geekbench Vulkan, the gap nearly vanishes. The A100 scores 146,380, and the V620 scores 144,364, a marginal 1.4% difference. This near-parity suggests that Vulkan workloads — which often favor graphics-oriented architectures with ray tracing support and higher clock speeds — allow the V620 to leverage its 2200 MHz boost clock and 72 ray tracing cores effectively. The V620's 281.6 GPixel/s pixel rate and 633.6 GTexel/s texture rate, both higher than the A100's, may also contribute to closing the gap in this API.

The aggregate picture reinforces the A100's superiority: its average benchmark score of 162,504 is 19.1% higher than the V620's 136,472. The A100 wins both head-to-head tests, with winsA equal to 2 and winsB equal to 0. However, the data also reveals that the V620 is not a distant also-ran; its Vulkan result sits within 2,016 points of the A100, and its nearest rivals — the AMD Radeon Pro W6800X Duo, AMD Radeon PRO W6800, NVIDIA A10M, and NVIDIA RTX 4000 Ada Generation — all sit within 0.9% of its average score. The A100, by contrast, edges out its closest competitor by 1.1% and trails only three cards in its vicinity, none by more than 2.2%. This places the A100 as a top-tier compute accelerator and the V620 as a solid, clock-driven alternative whose strengths emerge primarily in Vulkan-centric tasks.

DETAILED SPECIFICATIONS

SPECIFICATION
PRO V620
A100 PCIe 40 GB
Core Specs
Shading Units
4,608
6,912 +50.0%
Shaders
4,608
6,912 +50.0%
TMUs
288
432 +50.0%
ROPs
128
160 +25.0%
Compute Units
72
—
SM Count
—
108
Clocks
Base Clock
1825 MHz
765 MHz
Boost Clock
2200 MHz
1410 MHz
Memory Clock
2000 MHz 16 Gbps effective
1215 MHz 2.4 Gbps effective
Memory
Memory Size
32 GB
40 GB
VRAM (MB)
32,768
40,960 +25.0%
Memory Type
GDDR6
HBM2e
Memory Bus
256 bit
5120 bit
Bandwidth
512.0 GB/s
1.56 TB/s
Cache
L1 Cache
128 KB per Array
192 KB (per SM)
L2 Cache
4 MB
40 MB
L3 Cache
128 MB
—
L0 Cache
32 KB per WGP
—
Performance
Pixel Rate
281.6 GPixel/s
225.6 GPixel/s
Texture Rate
633.6 GTexel/s
609.1 GTexel/s
FP32 (TFLOPS)
20.28 TFLOPS
19.49 TFLOPS
FP64 (TFLOPS)
1,267.2 GFLOPS (1:16)
9.746 TFLOPS (1:2)
FP16 (TFLOPS)
40.55 TFLOPS (2:1)
77.97 TFLOPS (4:1)
AI/RT
RT Cores
72
—
Tensor Cores
—
432
BF16
—
311.84 TFLOPS (16:1)
TF32
—
155.92 TFLOPs (8:1)
Power
TDP
300 W
250 W
TDP (W)
300
250 -16.7%
Suggested PSU
700 W
600 W
Power Connectors
2x 8-pin
8-pin EPS
Architecture
Architecture
RDNA 2.0
Ampere
GPU Name
Navi 21
GA100
Generation
Radeon Pro Navi (Navi II Series)
Server Ampere (Axx)
Process Size
7 nm
7 nm
Transistors
26,800 million
54,200 million
Die Size
520 mm²
826 mm²
Foundry
TSMC
TSMC
Density
51.5M / mm²
65.6M / mm²
API Support
DirectX
12 Ultimate (12_2)
—
OpenGL
4.6
—
Vulkan
1.4
—
OpenCL
2.1
3.0
CUDA
—
8.0
Shader Model
6.8
—
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
120 mm 4.7 inches
111 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Radeon Pro Vega
Tesla Turing
Successor
—
Server Ada
View Radeon PRO V620 Details View A100 PCIe 40 GB Details