AMD Radeon PRO V620 vs NVIDIA B300 SXM6 AC Comparison

AMD
RADEON

AMD Radeon PRO V620

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 2200 MHz
TDP 300 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

B300 SXM6 AC

CORE STATE GB110
VRAM 288 GB
CLOCK SPEED 2032 MHz
TDP 1100 W
BUS WIDTH 8192 bit
ARCHITECTURE Blackwell Ultra
nm
PROCESS 5 nm
LAUNCH DATE 2025

PERFORMANCE BENCHMARKS

geekbench_opencl
128,580
369,831
geekbench_vulkan
144,364
N/A

Analysis: AMD Radeon PRO V620 vs NVIDIA B300 SXM6 AC

The NVIDIA B300 SXM6 AC and AMD Radeon PRO V620 are separated by four years of silicon evolution, yet they occupy vastly different positions in the server accelerator market. The B300 SXM6 AC delivers a Geekbench OpenCL score of 369,831, placing it in the 100th percentile of all GPUs, while the Radeon PRO V620 scores 128,580 in the same test, sitting at the 96th percentile. A single head-to-head benchmark shows the NVIDIA part winning by 187.6%, a decisive margin that reflects fundamental architectural and market differences. The B300 is an active, current-generation Blackwell Ultra compute monster with 288 GB of HBM3e, while the V620 is an end-of-life RDNA 2.0 workstation card with 32 GB of GDDR6. The data tells a story of specialization: one product targets massive AI and HPC workloads, the other targeted professional rendering and virtualization at a more modest scale.

The Verdict

The benchmark data is unambiguous: the NVIDIA B300 SXM6 AC is in a completely different performance class. Its OpenCL score of 369,831 is 187.6% higher than the Radeon PRO V620's 128,580. The B300 also sits 7% above the NVIDIA B200 (345,482), 10.4% above the H200 NVL (334,891), and 16.3% ahead of the AMD Instinct MI300X (317,994). These deltas show the B300 is not merely iterative; it is the top performer in the entire GPU database, holding the 100th percentile ranking.

The Radeon PRO V620, by contrast, is clustered tightly with its nearest rivals: it is just 0.5% behind the Radeon Pro W6800X Duo (135,774), 0.8% behind the PRO W6800 (135,396), and 0.9% behind both the NVIDIA A10M (135,230) and RTX 4000 Ada Generation (135,218). This indicates the V620 is a solid mid-range performer for its generation, but it is now end-of-life and cannot compete with current flagship silicon.

Who should pick which? Strictly from the data, organizations requiring maximum compute density for AI training, large-scale inference, or HPC simulation should choose the B300 SXM6 AC. Its 288 GB memory capacity, 8.19 TB/s bandwidth, and 76.99 TFLOPS FP32 throughput are unmatched in this comparison. The V620 makes sense only for legacy deployments or tasks that specifically require its feature set, such as DirectX 12 Ultimate support or its 72 ray-tracing cores, which the B300 lacks entirely. The V620's 96th percentile ranking is respectable, but the 3.7x performance gap in OpenCL leaves no practical scenario where the older card wins on raw compute.

FAQ

Q: How much faster is the NVIDIA B300 SXM6 AC than the AMD Radeon PRO V620 in OpenCL?

A: The B300 scores 369,831 compared to the V620's 128,580, which represents a 187.6% advantage for the NVIDIA part. This is the only benchmark shared between the two cards.

Q: What is the memory configuration difference between these two GPUs?

A: The B300 uses 288 GB of HBM3e with an 8192-bit bus and 8.19 TB/s bandwidth. The V620 uses 32 GB of GDDR6 with a 256-bit bus and 512.0 GB/s bandwidth. The B300 has 9x the capacity and 16x the bandwidth.

Q: Which card has a higher transistor density?

A: The B300 is built on a 5 nm TSMC process with 208,000 million transistors on a 1628 mm² die, yielding 127.8M transistors per mm². The V620 uses a 7 nm process with 26,800 million transistors on a 520 mm² die, yielding 51.5M per mm².

Q: Does the Radeon PRO V620 support modern graphics APIs?

A: Yes, the V620 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The B300 lists "N/A" for DirectX, OpenGL, and Vulkan, indicating it is not designed for graphics workloads.

Q: What is the power consumption difference?

A: The B300 has a TDP of 1100 W with a suggested PSU of 1500 W, while the V620 has a TDP of 300 W with a suggested PSU of 700 W. The B300 consumes 3.7x more power.

Q: How does the B300 compare to other NVIDIA accelerators in its class?

A: The B300 is 7% faster than the B200 (345,482), 10.4% faster than the H200 NVL (334,891), and 25% faster than the L40S (295,763) in OpenCL. It holds the 100th percentile ranking globally.

Architecture Differences

The NVIDIA B300 SXM6 AC is built on the GB110 chip using the Blackwell Ultra architecture, fabricated on a 5 nm TSMC process. It packs 208,000 million transistors into a 1628 mm² die, achieving a density of 127.8M transistors per mm². This is a compute-first design with no display outputs and no graphics API support. The architecture emphasizes tensor operations, featuring 592 tensor cores alongside 18,944 shading units and 592 TMUs. The B300's memory subsystem uses HBM3e across an 8192-bit bus, providing 8.19 TB/s of bandwidth. Its FP32 throughput is 76.99 TFLOPS, with FP16 also rated at 76.99 TFLOPS at a 1:1 ratio.

The AMD Radeon PRO V620 uses the Navi 21 chip with RDNA 2.0 architecture, fabricated on a 7 nm TSMC process. It contains 26,800 million transistors on a 520 mm² die, with a density of 51.5M per mm². This is a graphics-capable design supporting DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, though it has no display outputs. The V620 includes 72 ray-tracing cores, a feature entirely absent from the B300. Its compute layout features 4,608 shading units, 288 TMUs, and 128 ROPs. The memory system uses GDDR6 on a 256-bit bus, delivering 512.0 GB/s. FP32 performance is 20.28 TFLOPS, while FP16 reaches 40.55 TFLOPS at a 2:1 ratio.

The process node difference is significant: 5 nm versus 7 nm gives the B300 a 2.5x higher transistor density. The B300 also has 4.1x more shading units and 2.1x more TMUs, but the V620 has 5.3x more ROPs (128 vs 24). The B300's pixel rate is lower at 48.77 GPixel/s versus 281.6 GPixel/s for the V620, reflecting the NVIDIA part's compute orientation. Texture rates are closer: 1,202.9 GTexel/s for the B300 versus 633.6 GTexel/s for the V620.

Specification Differences

The most striking difference is memory capacity: the B300 offers 288 GB of HBM3e versus 32 GB of GDDR6 for the V620. Bus width differs dramatically at 8192-bit versus 256-bit, and bandwidth is 8.19 TB/s versus 512.0 GB/s. Clock speeds show the V620 has a higher boost clock at 2200 MHz versus 2032 MHz for the B300, though base clocks are closer at 1825 MHz versus 1665 MHz.

Compute resources differ substantially: the B300 has 18,944 shading units, 592 TMUs, and 24 ROPs, while the V620 has 4,608 shading units, 288 TMUs, and 128 ROPs. The B300 features 592 tensor cores with no RT cores, while the V620 has 72 RT cores and no tensor cores. FP32 throughput is 76.99 TFLOPS for the B300 versus 20.28 TFLOPS for the V620. FP16 performance is 76.99 TFLOPS (1:1) for the B300 versus 40.55 TFLOPS (2:1) for the V620.

Power and physical specifications diverge sharply. The B300 is an SXM module with a 1100 W TDP and 1500 W suggested PSU, while the V620 is a dual-slot card with a 300 W TDP and 700 W suggested PSU. The V620 uses two 8-pin power connectors; the B300 lists none. The V620 has measureable dimensions of 267 mm length, 120 mm height, and 50 mm width. The B300 has no listed dimensions. Bus interfaces differ: PCIe 6.0 x16 for the B300 versus PCIe 4.0 x16 for the V620.

Production status and release dates confirm their lifecycles: the B300 is active, released on 2025-09-10, while the V620 is end-of-life, released on 2021-11-03. The B300's predecessor is "Server Hopper" and successor is "Server Rubin," while the V620's predecessor is "Radeon Pro Vega" with no successor listed.

Head-to-Head Benchmarks

The sole shared benchmark is Geekbench OpenCL, where the NVIDIA B300 SXM6 AC scores 369,831 against the AMD Radeon PRO V620's 128,580. This yields a delta of 187.6%, meaning the B300 is nearly three times faster. The margin is so large that it dwarfs the differences between the B300 and its own nearest rivals, which range from 7% to 25%. For context, the B300's advantage over the V620 is 11.5x larger than its advantage over the B200.

The B300 also has a second benchmark result in its profile, but only the OpenCL score of 369,831 is listed. The V620 has two results: OpenCL at 128,580 and Vulkan at 144,364. The Vulkan score is 12.3% higher than its OpenCL score, suggesting the V620's RDNA 2.0 architecture is better optimized for Vulkan workloads. However, the B300 does not provide a Vulkan score, so no cross-comparison is possible.

Average benchmark scores reinforce the gap: the B300 averages 369,831, while the V620 averages 136,472. The B300's percentile ranking of 100 means it outperforms every other GPU in the database, while the V620's 96th percentile places it in the top 4% but far from the top. The V620's nearest rival deltas are all under 1%, indicating a tightly competitive field around the 135,000-136,000 score range, whereas the B300's closest competitor is 7% behind.

Where Each One Wins

The NVIDIA B300 SXM6 AC wins decisively in raw compute performance. Its OpenCL score of 369,831 is 187.6% higher than the V620's, making it the clear choice for any workload that prioritizes raw FP32 throughput, memory bandwidth, or tensor operations. The 288 GB memory capacity and 8.19 TB/s bandwidth enable datasets that would be impossible to fit in the V620's 32 GB frame buffer. The B300 also wins on transistor density (127.8M per mm² versus 51.5M), suggesting better power efficiency per transistor despite its 1100 W TDP. Its 76.99 TFLOPS FP32 and FP16 performance is unmatched in this comparison.

The AMD Radeon PRO V620 wins in specific architectural features. It has 5.3x more ROPs (128 versus 24), giving it a much higher pixel rate of 281.6 GPixel/s compared to 48.77 GPixel/s for the B300. It also has 72 ray-tracing cores, which the B300 lacks entirely. The V620 supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, making it suitable for graphics and visualization tasks that the B300 cannot handle. Its lower TDP of 300 W and compact dual-slot design with 267 mm length make it easier to integrate into existing systems, while the B300 requires an SXM module and 1500 W PSU.

The data suggests the B300 is the winner for AI training, large-scale inference, and HPC simulation, where tensor cores and massive memory bandwidth dominate. The V620 retains relevance for legacy virtualization, professional rendering, or any workload requiring ray tracing or modern graphics APIs. The 96th percentile ranking for the V620 shows it remains competitive within its generation, but the 100th percentile B300 represents a generational leap that the older card cannot bridge. The V620's Vulkan score of 144,364, which is 12.3% above its OpenCL score, hints at untapped potential in Vulkan-based workloads, but no comparable data exists for the B300 to validate this advantage.

DETAILED SPECIFICATIONS

SPECIFICATION
PRO V620
B300 SXM6 AC
Core Specs
Shading Units
4,608
18,944 +311.1%
Shaders
4,608
18,944 +311.1%
TMUs
288
592 +105.6%
ROPs
128
24 -81.3%
Compute Units
72
—
SM Count
—
148
Clocks
Base Clock
1825 MHz
1665 MHz
Boost Clock
2200 MHz
2032 MHz
Memory Clock
2000 MHz 16 Gbps effective
2000 MHz 8 Gbps effective
Memory
Memory Size
32 GB
288 GB
VRAM (MB)
32,768
294,912 +800.0%
Memory Type
GDDR6
HBM3e
Memory Bus
256 bit
8192 bit
Bandwidth
512.0 GB/s
8.19 TB/s
Cache
L1 Cache
128 KB per Array
256 KB (per SM)
L2 Cache
4 MB
126 MB
L3 Cache
128 MB
—
L0 Cache
32 KB per WGP
—
Performance
Pixel Rate
281.6 GPixel/s
48.77 GPixel/s
Texture Rate
633.6 GTexel/s
1,202.9 GTexel/s
FP32 (TFLOPS)
20.28 TFLOPS
76.99 TFLOPS
FP64 (TFLOPS)
1,267.2 GFLOPS (1:16)
1,202.9 GFLOPS (1:64)
FP16 (TFLOPS)
40.55 TFLOPS (2:1)
76.99 TFLOPS (1:1)
AI/RT
RT Cores
72
—
Tensor Cores
—
592
Power
TDP
300 W
1100 W
TDP (W)
300
1,100 +266.7%
Suggested PSU
700 W
1500 W
Power Connectors
2x 8-pin
—
Architecture
Architecture
RDNA 2.0
Blackwell Ultra
GPU Name
Navi 21
GB110
Generation
Radeon Pro Navi (Navi II Series)
Server Blackwell (Bxx)
Process Size
7 nm
5 nm
Transistors
26,800 million
208,000 million
Die Size
520 mm²
1628 mm²
Foundry
TSMC
TSMC
Density
51.5M / mm²
127.8M / mm²
API Support
DirectX
12 Ultimate (12_2)
—
OpenGL
4.6
—
Vulkan
1.4
—
OpenCL
2.1
3.0
CUDA
—
10.3
Shader Model
6.8
—
Physical
Slot Width
Dual-slot
SXM Module
Length
267 mm 10.5 inches
—
Height
120 mm 4.7 inches
—
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 6.0 x16
Other
Production
End-of-life
Active
Predecessor
Radeon Pro Vega
Server Hopper
Successor
—
Server Rubin
View Radeon PRO V620 Details View B300 SXM6 AC Details