AMD Radeon PRO V620 vs NVIDIA B200 Comparison

AMD
RADEON

AMD Radeon PRO V620

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 2200 MHz
TDP 300 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

B200

CORE STATE GB100
VRAM 90 GB
CLOCK SPEED 1965 MHz
TDP 1000 W
BUS WIDTH 4096 bit
ARCHITECTURE Blackwell
nm
PROCESS 5 nm
LAUNCH DATE

PERFORMANCE BENCHMARKS

geekbench_opencl
128,580
345,482
geekbench_vulkan
144,364
N/A

Analysis: AMD Radeon PRO V620 vs NVIDIA B200

# NVIDIA B200 vs AMD Radeon PRO V620

The NVIDIA B200 is the definitive choice for compute-dense workloads, delivering a Geekbench OpenCL score of 345,482 — a 168.7% advantage over the AMD Radeon PRO V620’s 128,580 in the same test. The B200 sits at the 100th percentile among all GPUs, while the V620 ranks at the 96th, making the gap between them a matter of class rather than mere performance tier. The Radeon PRO V620 remains a viable option only for legacy deployments or scenarios where its end-of-life status and lower power envelope are acceptable, but the data offers no benchmark category where it surpasses the B200.

The Verdict

Choose the NVIDIA B200 if your workload prioritizes raw compute throughput, AI acceleration, or memory bandwidth above all else. Its 74.45 TFLOPS FP32 and 1,191.2 TFLOPS FP16 (16:1) figures dwarf the V620’s 20.28 TFLOPS FP32 and 40.55 TFLOPS FP16 (2:1). The B200 also commands a 4.10 TB/s memory bandwidth from 90 GB of HBM3e on a 4096-bit bus, versus the V620’s 512.0 GB/s from 32 GB of GDDR6 on a 256-bit bus — a 700% bandwidth advantage that no amount of architectural efficiency can offset.

Choose the AMD Radeon PRO V620 if you require a dual-slot card with standard 2x 8-pin power connectors, a 300 W TDP, and a 267 mm length that fits conventional server chassis. It supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the B200 lists no API support. However, benchmark results show the V620’s best score (144,364 in Geekbench Vulkan) still trails the B200’s single OpenCL result by 139.4%. The V620 is end-of-life production, while the B200 is active — a practical consideration for long-term procurement.

The verdict is unambiguous: the B200 wins the only head-to-head benchmark available (Geekbench OpenCL) by 168.7%, and its nearest rivals list shows it beating the NVIDIA H200 NVL by 3.2% and AMD Instinct MI300X by 8.6%, while only the NVIDIA B300 SXM6 AC (6.6% faster) outranks it. The V620, by contrast, trades blows with cards like the AMD Radeon Pro W6800X Duo (0.5% delta) and NVIDIA RTX 4000 Ada Generation (0.9% delta), placing it in a completely different performance stratum.

Architecture Differences

The NVIDIA B200 is built on the Blackwell architecture, fabricated on a 5 nm process at TSMC with 104,000 million transistors on the GB100 chip. It features 18,944 shading units, 592 TMUs, and 592 tensor cores, with only 24 ROPs — a configuration clearly optimized for compute rather than rasterization. The B200’s 1,163.3 GTexel/s texture rate and 47.16 GPixel/s pixel rate reflect this design focus, as does its SXM Module form factor with no display outputs.

The AMD Radeon PRO V620 uses the RDNA 2.0 architecture on a 7 nm process with the Navi 21 chip, containing 26,800 million transistors on a 520 mm² die with a transistor density of 51.5M per mm². It has 4,608 shading units, 288 TMUs, 128 ROPs, and 72 ray tracing cores — a more balanced configuration that still targets compute but retains traditional graphics features. The V620’s 281.6 GPixel/s pixel rate is nearly 6x higher than the B200’s, and its 633.6 GTexel/s texture rate is approximately 45% of the B200’s, indicating the architectural trade-off between compute density and graphics throughput.

The B200’s base clock of 700 MHz with a 1965 MHz boost is notably lower than the V620’s 1825 MHz base and 2200 MHz boost, yet the B200 still achieves 3.67x higher FP32 throughput due to its massive shader count. The B200 uses HBM3e memory at 8 Gbps effective speed, while the V620 uses GDDR6 at 16 Gbps effective — the B200’s wider 4096-bit bus compensates for the lower per-pin speed. The B200 has no listed tensor core equivalent on the V620, which lacks dedicated AI acceleration hardware.

FAQ

Q: Which GPU has better raw compute performance?

A: The NVIDIA B200 delivers 74.45 TFLOPS FP32 and 1,191.2 TFLOPS FP16 (16:1), compared to the V620’s 20.28 TFLOPS FP32 and 40.55 TFLOPS FP16 (2:1). In Geekbench OpenCL, the B200 scores 345,482 versus 128,580 for the V620, a 168.7% difference.

Q: Are there any benchmark tests where the AMD Radeon PRO V620 wins?

A: No. In the only head-to-head benchmark available (Geekbench OpenCL), the B200 wins outright. The V620 has a separate Geekbench Vulkan score of 144,364, but no Vulkan result exists for the B200, so no direct comparison is possible in that test.

Q: What memory configurations do these cards use?

A: The B200 features 90 GB of HBM3e on a 4096-bit bus with 4.10 TB/s bandwidth. The V620 has 32 GB of GDDR6 on a 256-bit bus with 512.0 GB/s bandwidth. The B200’s bandwidth advantage is roughly 8x.

Q: How do these cards compare in terms of power requirements?

A: The B200 has a TDP of 1000 W with a suggested PSU of 1400 W, while the V620 has a TDP of 300 W with a 700 W suggested PSU. The B200 uses an SXM Module slot width with no power connectors listed, while the V620 is dual-slot with 2x 8-pin connectors.

Q: Which card has better production status and availability?

A: The NVIDIA B200 is marked as Active in production, while the AMD Radeon PRO V620 is End-of-life. The V620 was released on 2021-11-03, and its predecessor is Radeon Pro Vega. The B200’s predecessor is Server Hopper, and its successor is Server Rubin.

Q: Can either card output video to displays?

A: No. Both the NVIDIA B200 and AMD Radeon PRO V620 list "No outputs" for display connectivity, indicating they are designed exclusively for compute/server use rather than workstation graphics.

Specification Differences

| Specification | NVIDIA B200 | AMD Radeon PRO V620 |

|---|---|---|

| Architecture | Blackwell | RDNA 2.0 |

| Process Node | 5 nm | 7 nm |

| Transistors | 104,000 million | 26,800 million |

| Die Size | Not listed | 520 mm² |

| Transistor Density | Not listed | 51.5M / mm² |

| Base Clock | 700 MHz | 1825 MHz |

| Boost Clock | 1965 MHz | 2200 MHz |

| Memory Speed | 8 Gbps effective | 16 Gbps effective |

| Memory Size | 90 GB | 32 GB |

| Memory Type | HBM3e | GDDR6 |

| Memory Bus Width | 4096 bit | 256 bit |

| Memory Bandwidth | 4.10 TB/s | 512.0 GB/s |

| Shading Units | 18,944 | 4,608 |

| TMUs | 592 | 288 |

| ROPs | 24 | 128 |

| Ray Tracing Cores | Not listed | 72 |

| Tensor Cores | 592 | Not listed |

| Pixel Rate | 47.16 GPixel/s | 281.6 GPixel/s |

| Texture Rate | 1,163.3 GTexel/s | 633.6 GTexel/s |

| FP32 Performance | 74.45 TFLOPS | 20.28 TFLOPS |

| FP16 Performance | 1,191.2 TFLOPS (16:1) | 40.55 TFLOPS (2:1) |

| TDP | 1000 W | 300 W |

| Slot Width | SXM Module | Dual-slot |

| Power Connectors | Not listed | 2x 8-pin |

| Suggested PSU | 1400 W | 700 W |

| Bus Interface | PCIe 5.0 x16 | PCIe 4.0 x16 |

| Dimensions (L×H×W) | Not listed | 267 mm × 120 mm × 50 mm |

| Production Status | Active | End-of-life |

| Release Date | Not listed | 2021-11-03 |

| Predecessor | Server Hopper | Radeon Pro Vega |

| Successor | Server Rubin | Not listed |

| DirectX Support | Not listed | 12 Ultimate (12_2) |

| OpenGL Support | Not listed | 4.6 |

| Vulkan Support | Not listed | 1.4 |

| Geekbench OpenCL | 345,482 | 128,580 |

| Geekbench Vulkan | Not listed | 144,364 |

| Percentile vs All GPUs | 100 | 96 |

Head-to-Head Benchmarks

The sole head-to-head benchmark available is Geekbench OpenCL, where the NVIDIA B200 scores 345,482 against the AMD Radeon PRO V620’s 128,580. The 168.7% delta represents a massive performance gulf that eclipses any architectural nuance. To contextualize this, the B200’s nearest rivals include the NVIDIA H200 NVL at 334,891 (3.2% slower) and the AMD Instinct MI300X at 317,994 (8.6% slower); the B200 even comes within 6.6% of the newer NVIDIA B300 SXM6 AC, which scores 369,831. Meanwhile, the V620’s closest competitors — the AMD Radeon Pro W6800X Duo at 135,774 (0.5% delta), AMD Radeon PRO W6800 at 135,396 (0.8%), NVIDIA A10M at 135,230 (0.9%), and NVIDIA RTX 4000 Ada Generation at 135,218 (0.9%) — all cluster within 1% of its score, showing that the V620 is essentially tied with mid-range workstation cards.

The B200’s FP32 throughput of 74.45 TFLOPS is 3.67x the V620’s 20.28 TFLOPS, while its FP16 performance of 1,191.2 TFLOPS (16:1) is 29.4x the V620’s 40.55 TFLOPS (2:1) — a disparity that widens dramatically in mixed-precision AI workloads. Memory bandwidth tells a similar story: 4.10 TB/s versus 512.0 GB/s means the B200 can feed its compute units nearly 8x faster. The V620 does counter in pixel rate (281.6 GPixel/s versus 47.16 GPixel/s) and ROP count (128 versus 24), but these metrics are irrelevant for compute-first applications where neither card has display outputs.

Where Each One Wins

NVIDIA B200 wins in: raw compute (FP32 and FP16), memory capacity and bandwidth, tensor core acceleration, and overall benchmark dominance. Its 100th percentile ranking versus the V620’s 96th places it in the top tier of all GPUs. The B200’s PCIe 5.0 x16 interface (versus the V620’s PCIe 4.0 x16) provides double the host interconnect bandwidth for data transfer. Its Active production status and successor path from Server Hopper to Server Rubin indicate an ongoing product lifecycle. The B200 also holds a 3.2% edge over the H200 NVL and 8.6% over the MI300X, showing competitive strength against other flagship accelerators.

AMD Radeon PRO V620 wins in: physical compatibility and graphics API support. Its dual-slot 267 mm length, 120 mm height, and 50 mm width fit standard PCIe slots, and its 300 W TDP with 2x 8-pin connectors makes it deployable in systems with modest power delivery. It supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, which the B200 lacks entirely. The V620’s 72 ray tracing cores and 128 ROPs give it a graphics-rendering capability that the B200 does not offer, though neither card has display outputs. Its 16 Gbps effective memory speed is double the B200’s 8 Gbps, though the B200’s wider bus makes this irrelevant in practice. For legacy software stacks that require Vulkan or DirectX, the V620 is the only option of the two — but its end-of-life status means any deployment must account for eventual obsolescence.

DETAILED SPECIFICATIONS

SPECIFICATION
PRO V620
B200
Core Specs
Shading Units
4,608
18,944 +311.1%
Shaders
4,608
18,944 +311.1%
TMUs
288
592 +105.6%
ROPs
128
24 -81.3%
Compute Units
72
SM Count
148
Clocks
Base Clock
1825 MHz
700 MHz
Boost Clock
2200 MHz
1965 MHz
Memory Clock
2000 MHz 16 Gbps effective
2000 MHz 8 Gbps effective
Memory
Memory Size
32 GB
90 GB
VRAM (MB)
32,768
92,160 +181.3%
Memory Type
GDDR6
HBM3e
Memory Bus
256 bit
4096 bit
Bandwidth
512.0 GB/s
4.10 TB/s
Cache
L1 Cache
128 KB per Array
256 KB (per SM)
L2 Cache
4 MB
50 MB
L3 Cache
128 MB
L0 Cache
32 KB per WGP
Performance
Pixel Rate
281.6 GPixel/s
47.16 GPixel/s
Texture Rate
633.6 GTexel/s
1,163.3 GTexel/s
FP32 (TFLOPS)
20.28 TFLOPS
74.45 TFLOPS
FP64 (TFLOPS)
1,267.2 GFLOPS (1:16)
37.22 TFLOPS (1:2)
FP16 (TFLOPS)
40.55 TFLOPS (2:1)
1,191.2 TFLOPS (16:1)
AI/RT
RT Cores
72
Tensor Cores
592
Power
TDP
300 W
1000 W
TDP (W)
300
1,000 +233.3%
Suggested PSU
700 W
1400 W
Power Connectors
2x 8-pin
Architecture
Architecture
RDNA 2.0
Blackwell
GPU Name
Navi 21
GB100
Generation
Radeon Pro Navi (Navi II Series)
Server Blackwell (Bxx)
Process Size
7 nm
5 nm
Transistors
26,800 million
104,000 million
Die Size
520 mm²
Foundry
TSMC
TSMC
Density
51.5M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
2.1
3.0
CUDA
10.0
Shader Model
6.8
Physical
Slot Width
Dual-slot
SXM Module
Length
267 mm 10.5 inches
Height
120 mm 4.7 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 5.0 x16
Other
Production
End-of-life
Active
Predecessor
Radeon Pro Vega
Server Hopper
Successor
Server Rubin
View Radeon PRO V620 Details View B200 Details