AMD Radeon PRO V620 vs NVIDIA GB10 Comparison
AMD Radeon PRO V620
GB10
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon PRO V620 vs NVIDIA GB10
The AMD Radeon PRO V620 and NVIDIA GB10 are two very different solutions that happen to share a performance tier, but the data shows a clear victor in raw compute benchmarks. The AMD Radeon PRO V620 wins both head-to-head tests, but the NVIDIA GB10 counters with a vastly larger memory pool, a much lower power draw, and a completely different form factor that makes it a unique product for specific workloads. The benchmark results indicate that the PRO V620 is the stronger pure compute card, while the GB10 is a specialized, high-capacity compute module with a different design philosophy.
Head-to-Head Benchmarks
The direct comparison between these two GPUs is straightforward, as the FACT PACK includes only two benchmark tests. In the Geekbench OpenCL test, the AMD Radeon PRO V620 scores 128580, defeating the NVIDIA GB10’s 120137 by a margin of 7%. While this is not a dominant victory, it establishes the AMD card as the faster option in this general-purpose compute API. The GB10, despite having a newer architecture and more shading units, cannot match the PRO V620’s raw throughput in this test.
The gap widens significantly in the Geekbench Vulkan test. Here, the AMD Radeon PRO V620 scores 144364, while the NVIDIA GB10 scores 114648. This result gives the AMD card a decisive 25.9% advantage. This substantial delta suggests that the PRO V620’s architecture is particularly well-suited to the Vulkan API, or that the GB10’s drivers or design are less optimized for this workload. The data is clear: in both available head-to-head metrics, the AMD Radeon PRO V620 is the winner.
The average benchmark scores reinforce this narrative. The AMD Radeon PRO V620 has an average benchmark score of 136472, placing it in the 96th percentile of all GPUs. The NVIDIA GB10, with an average score of 117393, sits just behind in the 95th percentile. The delta between their averages is significant. Looking at the nearest rivals, the PRO V620’s score is 0.5% higher than the AMD Radeon Pro W6800X Duo and 0.8% higher than the AMD Radeon PRO W6800. The GB10, on the other hand, is 1.3% behind the AMD Radeon PRO W7700 and only 0.3% ahead of the NVIDIA RTX 4000 SFF Ada Generation. This places both cards in a similar competitive bracket, but the AMD card sits at the top of that bracket.
FAQ
Q: Which GPU has the higher average benchmark score?
A: The AMD Radeon PRO V620 has a higher average benchmark score of 136472, compared to the NVIDIA GB10’s 117393.
Q: How much faster is the AMD Radeon PRO V620 in the Vulkan test?
A: The AMD Radeon PRO V620 scores 144364 in Geekbench Vulkan, which is 25.9% higher than the NVIDIA GB10’s score of 114648.
Q: Does the NVIDIA GB10 offer any advantage in memory capacity?
A: Yes, the NVIDIA GB10 features 128 GB of LPDDR5X memory, which is four times the 32 GB of GDDR6 memory found on the AMD Radeon PRO V620.
Q: What is the power consumption difference between the two cards?
A: The NVIDIA GB10 has a TDP of 140 W, while the AMD Radeon PRO V620 has a TDP of 300 W. The GB10 also requires no power connectors and a 300 W power supply, compared to the PRO V620’s two 8-pin connectors and 700 W suggested PSU.
Q: Which GPU has a higher boost clock speed?
A: The NVIDIA GB10 has a higher boost clock of 2418 MHz, compared to the AMD Radeon PRO V620’s boost clock of 2200 MHz.
Q: Are both GPUs in the same performance percentile?
A: Yes, they are close. The AMD Radeon PRO V620 is in the 96th percentile, while the NVIDIA GB10 is in the 95th percentile of all GPUs.
Architecture Differences
The two GPUs are built on fundamentally different architectures and for different purposes. The AMD Radeon PRO V620 is based on the RDNA 2.0 architecture and uses the Navi 21 chip, manufactured on a 7 nm process at TSMC. It is a large, traditional graphics card with a die size of 520 mm² and 26,800 million transistors. In contrast, the NVIDIA GB10 uses the newer Blackwell 2.0 architecture with the GB20B chip, built on a smaller 5 nm process. Its die size is 382 mm², and while its transistor count is listed as unknown, the smaller die and newer node indicate a different design strategy.
The compute configurations differ significantly. The AMD card has 4608 shading units, 288 texture mapping units, and 128 ROPs. It also includes 72 ray tracing cores. The NVIDIA GB10 has more shading units (6144) and more TMUs (384), but far fewer ROPs (48). It also has fewer ray tracing cores (48) but includes 384 tensor cores, which the AMD card lacks entirely. This suggests the GB10 is designed with AI and machine learning workloads in mind, while the PRO V620 is a more general-purpose rasterization and compute card. The FP32 performance reflects this: the GB10 achieves 29.71 TFLOPS compared to the PRO V620’s 20.28 TFLOPS. However, the GB10’s FP16 performance is the same as its FP32 (29.71 TFLOPS at 1:1), while the AMD card achieves 40.55 TFLOPS FP16 with a 2:1 ratio.
The memory subsystems are also worlds apart. The AMD Radeon PRO V620 uses 32 GB of GDDR6 memory on a 256-bit bus, delivering a massive 512.0 GB/s of bandwidth. The NVIDIA GB10, however, uses 128 GB of LPDDR5X memory on the same 256-bit bus, but its bandwidth is only 273.2 GB/s. This means the AMD card has nearly double the memory bandwidth, which is crucial for high-resolution textures and compute tasks, while the GB10 prioritizes capacity for large datasets that exceed the PRO V620’s limits. The form factor also differs: the PRO V620 is a dual-slot, 267 mm PCIe 4.0 x16 card with no display outputs, while the GB10 is an IGP (integrated graphics processor) module using PCIe 5.0 x16, measuring just 150 mm, and includes a single HDMI output.
The Verdict
The data points to a clear split in use cases. For raw compute performance in OpenCL and Vulkan, the AMD Radeon PRO V620 is the unambiguous winner. It leads in both head-to-head benchmarks, with a 7% advantage in OpenCL and a substantial 25.9% lead in Vulkan. Its higher memory bandwidth of 512.0 GB/s, compared to the GB10’s 273.2 GB/s, also makes it better suited for tasks that are sensitive to memory throughput. The PRO V620’s higher pixel rate (281.6 GPixel/s vs 116.1 GPixel/s) and higher ROP count (128 vs 48) further cement its position as a stronger traditional graphics and compute processor.
However, the NVIDIA GB10 is not a failure; it is a different tool. Its 128 GB of memory is a massive advantage for workloads that require holding large models or datasets entirely in VRAM. Its lower TDP of 140 W, versus the PRO V620’s 300 W, makes it far more power-efficient and easier to integrate into dense server environments. The GB10 also has a higher FP32 throughput (29.71 TFLOPS) and a much higher texture rate (928.5 GTexel/s vs 633.6 GTexel/s), indicating that in certain non-benchmarked compute tasks, it could outperform the AMD card. The GB10 is also an active product with a successor (Server Rubin) and a launch MSRP of 3,999 USD, while the PRO V620 is end-of-life.
Specification Differences
The two cards differ in nearly every major specification category. The process node is a key difference: the AMD Radeon PRO V620 uses a 7 nm process, while the NVIDIA GB10 uses a 5 nm process. The die size also differs, with the AMD card at 520 mm² and the NVIDIA card at 382 mm². The PRO V620 has a base clock of 1825 MHz and a boost clock of 2200 MHz, while the GB10 has a lower base clock of 1665 MHz but a higher boost clock of 2418 MHz.
The memory is a major point of divergence. The AMD card has 32 GB of GDDR6 memory with a bandwidth of 512.0 GB/s and a memory clock of 2000 MHz (16 Gbps effective). The NVIDIA card has 128 GB of LPDDR5X memory with a bandwidth of 273.2 GB/s and a memory clock of 1067 MHz (8.5 Gbps effective). The compute units also differ: the AMD card has 4608 shading units, 288 TMUs, 128 ROPs, and 72 RT cores, while the NVIDIA card has 6144 shading units, 384 TMUs, 48 ROPs, 48 RT cores, and 384 tensor cores. The PRO V620’s FP32 performance is 20.28 TFLOPS, while the GB10’s is 29.71 TFLOPS.
The power and physical requirements are starkly different. The AMD Radeon PRO V620 has a TDP of 300 W, requires two 8-pin power connectors, and a 700 W power supply. It is a dual-slot card measuring 267 mm in length, 120 mm in height, and 50 mm in width. The NVIDIA GB10, in contrast, has a TDP of 140 W, needs no power connectors, and suggests a 300 W PSU. It is an IGP module with dimensions of 150 mm by 51 mm by 150 mm. The AMD card uses a PCIe 4.0 x16 interface, while the GB10 uses PCIe 5.0 x16. Finally, the AMD card has no display outputs, while the NVIDIA card has one HDMI output. The AMD card supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the NVIDIA card lists N/A for all three APIs.
Where Each One Wins
The AMD Radeon PRO V620 wins in scenarios that demand high memory bandwidth and strong API-level performance. Its 512.0 GB/s bandwidth is a major asset for compute tasks that stream large amounts of data, such as real-time ray tracing or high-resolution rendering. The PRO V620’s 25.9% lead in Vulkan and 7% lead in OpenCL make it the clear choice for applications built on these APIs. Its higher pixel rate (281.6 GPixel/s) and larger ROP count (128) also give it an edge in traditional graphics and rasterization workloads. For users needing a proven, high-performance compute card for general GPU tasks, the PRO V620 is the winner.
The NVIDIA GB10 wins in scenarios that prioritize memory capacity and power efficiency. The 128 GB of memory allows it to handle datasets that would exhaust the PRO V620’s 32 GB frame buffer, making it ideal for large-scale AI inference or scientific computing where the entire model must fit in VRAM. Its lower TDP of 140 W and no power connector requirement make it far easier to deploy in dense, power-constrained server racks. The GB10’s higher FP32 throughput (29.71 TFLOPS) and higher texture rate (928.5 GTexel/s) also suggest it could be faster in specific compute tasks that are not captured by the Geekbench tests. The inclusion of 384 tensor cores and a newer Blackwell 2.0 architecture points to a card designed for AI acceleration, making it the superior option for those specific workloads.