AMD Radeon PRO V620 vs NVIDIA H200 NVL Comparison

AMD
RADEON

AMD Radeon PRO V620

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 2200 MHz
TDP 300 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

H200 NVL

CORE STATE GH100
VRAM 141 GB
CLOCK SPEED 1785 MHz
TDP 600 W
BUS WIDTH 6144 bit
ARCHITECTURE Hopper
nm
PROCESS 5 nm
LAUNCH DATE 2024

PERFORMANCE BENCHMARKS

geekbench_opencl
128,580
334,891
geekbench_vulkan
144,364
N/A

Analysis: AMD Radeon PRO V620 vs NVIDIA H200 NVL

Where Each One Wins

The benchmark data splits these two accelerators into entirely different performance strata. The NVIDIA H200 NVL wins the only shared benchmark, Geekbench OpenCL, with a score of 334,891 against the AMD Radeon PRO V620’s 128,580. That is a 160.5% delta, meaning the H200 NVL delivers roughly two and a half times the OpenCL throughput of the V620. There is no shared Vulkan result; the AMD card has a Geekbench Vulkan score of 144,364, which is higher than its own OpenCL score by about 12.3%, but the NVIDIA part has no Vulkan data in the pack, so no cross-comparison is possible on that API.

The H200 NVL sits at the 100th percentile of all GPUs in the database, while the Radeon PRO V620 sits at the 96th percentile. That percentile gap is small in rank terms, but the raw score gap is enormous. The H200 NVL’s nearest rival, the NVIDIA B200, scores 345,482, which is 3.1% higher; the B300 SXM6 AC scores 369,831, which is 9.4% higher. The H200 NVL beats the AMD Instinct MI300X by 5.3% (317,994) and the NVIDIA L40S by 13.2% (295,763). The Radeon PRO V620, by contrast, is clustered tightly with its nearest rivals: the AMD Radeon Pro W6800X Duo scores 135,774 (0.5% higher), the AMD Radeon PRO W6800 scores 135,396 (0.8% higher), the NVIDIA A10M scores 135,230 (0.9% higher), and the NVIDIA RTX 4000 Ada Generation scores 135,218 (0.9% higher). The V620 is essentially at parity with those four cards, whereas the H200 NVL is in a different league entirely.

In terms of use cases, the data suggests the H200 NVL is for compute workloads where raw OpenCL throughput is paramount—large-scale server inference, scientific computing, or any task that can saturate 141 GB of HBM3e memory. The Radeon PRO V620, with its 32 GB of GDDR6 and 512.0 GB/s bandwidth, is a more modest server accelerator, appropriate for workloads that fit within that memory pool and do not require the H200 NVL’s massive bandwidth. The V620 does have a Vulkan path, which the H200 NVL lacks entirely, so for any Vulkan-based rendering or compute task, the AMD part is the only option in this pairing.

Architecture Differences

The two chips come from different architectural lineages. The NVIDIA H200 NVL uses the GH100 chip on the Hopper architecture, fabricated on a 5 nm process at TSMC. It packs 80,000 million transistors on an 814 mm² die, yielding a transistor density of 98.3M per mm². The AMD Radeon PRO V620 uses the Navi 21 chip on the RDNA 2.0 architecture, also TSMC-fabricated but on a 7 nm process. It has 26,800 million transistors on a 520 mm² die, giving a density of 51.5M per mm². The H200 NVL has roughly three times the transistor count and a 56.5% larger die, but the 5 nm node gives it a density advantage of nearly 2x.

Memory architecture is a major differentiator. The H200 NVL has 141 GB of HBM3e on a 6144-bit bus, delivering 4.89 TB/s of bandwidth. The V620 has 32 GB of GDDR6 on a 256-bit bus, delivering 512.0 GB/s. That is a 9.55x bandwidth advantage for the NVIDIA part and a 4.4x capacity advantage. The memory clock differs accordingly: the H200 NVL runs at 1593 MHz (6.4 Gbps effective), while the V620 runs at 2000 MHz (16 Gbps effective). The V620’s memory is faster per-pin, but the H200 NVL’s massive bus width overwhelms that advantage.

Compute resources also diverge sharply. The H200 NVL has 16,896 shading units, 528 TMUs, and 24 ROPs, plus 528 tensor cores. The V620 has 4,608 shading units, 288 TMUs, and 128 ROPs, plus 72 ray-tracing cores. The NVIDIA part has 3.67x the shader count and 1.83x the TMUs, but the AMD part has 5.33x the ROPs. The V620 also has dedicated ray-tracing hardware, which the H200 NVL lacks entirely (its RT core field is null). The H200 NVL’s tensor cores are its analogue to the V620’s RT cores, but they serve different purposes—tensor math versus ray traversal.

Clock speeds favor AMD. The V620 has a base clock of 1825 MHz and a boost of 2200 MHz, versus the H200 NVL’s 1365 MHz base and 1785 MHz boost. That is a 33.7% higher base clock and 23.2% higher boost clock for the AMD part. Despite the clock disadvantage, the H200 NVL’s raw throughput is far higher: 60.32 TFLOPS FP32 versus 20.28 TFLOPS, and 120.6 TFLOPS FP16 versus 40.55 TFLOPS (both at 2:1 ratio). The pixel rate tells a different story: the V620 outputs 281.6 GPixel/s versus the H200 NVL’s 42.84 GPixel/s, because the AMD part has far more ROPs. Texture rate favors NVIDIA at 942.5 GTexel/s versus 633.6 GTexel/s.

Power and interface also differ. The H200 NVL has a 600 W TDP with an 8-pin EPS connector and a suggested 1000 W PSU; the V620 has a 300 W TDP with 2x 8-pin connectors and a suggested 700 W PSU. The H200 NVL uses PCIe 5.0 x16, while the V620 uses PCIe 4.0 x16. Both are dual-slot cards with no display outputs. The H200 NVL is 267 mm long and 111 mm high; the V620 is also 267 mm long but 120 mm high and 50 mm wide. The H200 NVL lists no API support (DirectX, OpenGL, Vulkan all N/A), while the V620 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

The Verdict

The data points to a clear split. The NVIDIA H200 NVL is the compute monster: 160.5% ahead in OpenCL, 100th percentile, with 141 GB of HBM3e and 4.89 TB/s bandwidth. Any workload that can use OpenCL and needs massive memory capacity or bandwidth should pick the H200 NVL. The 600 W TDP and 1000 W PSU requirement are the costs of that performance, but the benchmark delta justifies them for compute-heavy tasks.

The AMD Radeon PRO V620 is for a different niche. It is end-of-life, sits at the 96th percentile, and its 128,580 OpenCL score is 2.6x lower than the H200 NVL’s. However, it is the only card here with Vulkan support and ray-tracing cores. For any Vulkan-based workload, the V620 is the only choice in this pair, since the H200 NVL has no Vulkan path. Its 32 GB of GDDR6 is smaller but still substantial, and its 300 W TDP is half the NVIDIA part’s, with a 700 W PSU suggestion. The V620 also has a much higher pixel rate (281.6 GPixel/s vs 42.84 GPixel/s), which could matter for rasterization-style tasks despite the absence of display outputs.

There is no scenario where the V620 wins on raw compute. But for Vulkan workloads, lower power envelopes, or applications that need ray-tracing hardware, the V620 is the functional option. The H200 NVL is the performance leader without qualification.

FAQ

Q: Which GPU has the higher OpenCL benchmark score?

A: The NVIDIA H200 NVL scores 334,891 in Geekbench OpenCL, which is 160.5% higher than the AMD Radeon PRO V620’s 128,580.

Q: Does the AMD Radeon PRO V620 support Vulkan while the NVIDIA H200 NVL does not?

A: Yes. The V620 has a Geekbench Vulkan score of 144,364 and supports Vulkan 1.4, while the H200 NVL lists Vulkan as N/A and has no Vulkan benchmark data.

Q: What is the memory capacity and bandwidth difference?

A: The H200 NVL has 141 GB of HBM3e with 4.89 TB/s bandwidth on a 6144-bit bus. The V620 has 32 GB of GDDR6 with 512.0 GB/s bandwidth on a 256-bit bus.

Q: Which card has more shading units and tensor cores?

A: The H200 NVL has 16,896 shading units and 528 tensor cores. The V620 has 4,608 shading units and no tensor cores, but it does have 72 ray-tracing cores.

Q: What are the power requirements for each card?

A: The H200 NVL has a 600 W TDP and suggests a 1000 W PSU. The V620 has a 300 W TDP and suggests a 700 W PSU.

Q: How does the NVIDIA H200 NVL compare to its nearest rivals?

A: The H200 NVL is 3.1% behind the NVIDIA B200 (345,482), 9.4% behind the NVIDIA B300 SXM6 AC (369,831), 5.3% ahead of the AMD Instinct MI300X (317,994), and 13.2% ahead of the NVIDIA L40S (295,763).

Head-to-Head Benchmarks

The only shared benchmark is Geekbench OpenCL. The NVIDIA H200 NVL scores 334,891, and the AMD Radeon PRO V620 scores 128,580. The delta is 160.5% in favor of NVIDIA. To put that in context, the H200 NVL’s score is more than double the V620’s, and the difference is larger than the gap between the H200 NVL and any of its nearest rivals. The closest competitor to the H200 NVL is the NVIDIA B200 at 345,482, which is only 3.1% higher. The V620’s nearest rival, the AMD Radeon Pro W6800X Duo, scores 135,774, just 0.5% higher. The H200 NVL’s OpenCL score is 2.6x the V620’s, but the V620 is only 0.9% behind the NVIDIA RTX 4000 Ada Generation (135,218). This suggests the V620 is competitive within its own tier, but that tier is far below the H200 NVL.

The V620 does have a Geekbench Vulkan score of 144,364. That is 12.3% higher than its own OpenCL score, indicating the card performs better under Vulkan in this test. However, the H200 NVL has no Vulkan benchmark, so the only head-to-head comparison available is OpenCL. Within that single data point, the H200 NVL’s victory is decisive. The 160.5% delta means the H200 NVL delivers 2.6x the performance, which is a magnitude of difference that dwarfs the 0.5-0.9% gaps seen among the V620’s rivals.

Specification Differences

The two cards differ in nearly every specification category. The NVIDIA H200 NVL uses a GH100 chip on Hopper architecture (5 nm, TSMC), while the AMD Radeon PRO V620 uses Navi 21 on RDNA 2.0 (7 nm, TSMC). Transistor count is 80,000 million versus 26,800 million; die size is 814 mm² versus 520 mm²; density is 98.3M/mm² versus 51.5M/mm². Base clocks are 1365 MHz versus 1825 MHz, boost clocks 1785 MHz versus 2200 MHz. Memory is 141 GB HBM3e versus 32 GB GDDR6; bus width is 6144-bit versus 256-bit; bandwidth is 4.89 TB/s versus 512.0 GB/s; memory clock is 1593 MHz (6.4 Gbps) versus 2000 MHz (16 Gbps). Shading units are 16,896 versus 4,608; TMUs are 528 versus 288; ROPs are 24 versus 128; tensor cores are 528 versus none; RT cores are none versus 72. Pixel rate is 42.84 GPixel/s versus 281.6 GPixel/s; texture rate is 942.5 GTexel/s versus 633.6 GTexel/s; FP32 is 60.32 TFLOPS versus 20.28 TFLOPS; FP16 is 120.6 TFLOPS versus 40.55 TFLOPS. TDP is 600 W versus 300 W; power connectors are 8-pin EPS versus 2x 8-pin; suggested PSU is 1000 W versus 700 W; bus interface is PCIe 5.0 x16 versus PCIe 4.0 x16. Dimensions: the H200 NVL is 267 mm x 111 mm, the V620 is 267 mm x 120 mm x 50 mm. The H200 NVL has no API support listed (DirectX, OpenGL, Vulkan all N/A), while the V620 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. Production status is Active versus End-of-life; release dates are November 2024 versus November 2021.

DETAILED SPECIFICATIONS

SPECIFICATION
PRO V620
H200 NVL
Core Specs
Shading Units
4,608
16,896 +266.7%
Shaders
4,608
16,896 +266.7%
TMUs
288
528 +83.3%
ROPs
128
24 -81.3%
Compute Units
72
SM Count
132
Clocks
Base Clock
1825 MHz
1365 MHz
Boost Clock
2200 MHz
1785 MHz
Memory Clock
2000 MHz 16 Gbps effective
1593 MHz 6.4 Gbps effective
Memory
Memory Size
32 GB
141 GB
VRAM (MB)
32,768
144,384 +340.6%
Memory Type
GDDR6
HBM3e
Memory Bus
256 bit
6144 bit
Bandwidth
512.0 GB/s
4.89 TB/s
Cache
L1 Cache
128 KB per Array
256 KB (per SM)
L2 Cache
4 MB
50 MB
L3 Cache
128 MB
L0 Cache
32 KB per WGP
Performance
Pixel Rate
281.6 GPixel/s
42.84 GPixel/s
Texture Rate
633.6 GTexel/s
942.5 GTexel/s
FP32 (TFLOPS)
20.28 TFLOPS
60.32 TFLOPS
FP64 (TFLOPS)
1,267.2 GFLOPS (1:16)
30.16 TFLOPS (1:2)
FP16 (TFLOPS)
40.55 TFLOPS (2:1)
120.6 TFLOPS (2:1)
AI/RT
RT Cores
72
Tensor Cores
528
Power
TDP
300 W
600 W
TDP (W)
300
600 +100.0%
Suggested PSU
700 W
1000 W
Power Connectors
2x 8-pin
8-pin EPS
Architecture
Architecture
RDNA 2.0
Hopper
GPU Name
Navi 21
GH100
Generation
Radeon Pro Navi (Navi II Series)
Server Hopper (Hxx)
Process Size
7 nm
5 nm
Transistors
26,800 million
80,000 million
Die Size
520 mm²
814 mm²
Foundry
TSMC
TSMC
Density
51.5M / mm²
98.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
2.1
3.0
CUDA
9.0
Shader Model
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
120 mm 4.7 inches
111 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 5.0 x16
Other
Production
End-of-life
Active
Predecessor
Radeon Pro Vega
Server Ada
Successor
Server Blackwell
View Radeon PRO V620 Details View H200 NVL Details