AMD Radeon PRO V620 vs NVIDIA CMP 40HX Comparison

AMD
RADEON

AMD Radeon PRO V620

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 2200 MHz
TDP 300 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

CMP 40HX

CORE STATE TU106
VRAM 8 GB
CLOCK SPEED 1650 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
128,580
93,395
geekbench_vulkan
144,364
77,879

Analysis: AMD Radeon PRO V620 vs NVIDIA CMP 40HX

Where Each One Wins

The data separates these two cards into sharply different roles, and the split is not subtle. The AMD Radeon PRO V620 wins both recorded benchmark tests outright, with a 2-0 record in the head-to-head comparison. That makes it the clear choice for compute-oriented workloads that stress raw throughput and memory capacity. The NVIDIA CMP 40HX, despite its lower scores, still occupies a distinct niche, one defined by efficiency and physical compactness rather than absolute performance.

Look at the benchmark averages. The Radeon PRO V620 posts a 136472 average score across the two tests, while the CMP 40HX sits at 85637. That is a 59.4% gap in favor of the AMD part, and it places the two cards in different performance tiers entirely. The Radeon PRO V620 ranks in the 96th percentile among all GPUs in the database, whereas the CMP 40HX lands in the 93rd percentile. The percentile difference looks modest, but the score gap tells the real story: the AMD card is operating in a higher performance class.

Where the CMP 40HX wins is in the practical, almost mundane metrics. It draws 185 W of power against 300 W for the AMD card. It requires a single 8-pin power connector instead of two. Its suggested power supply is 450 W versus 700 W. It measures 229 mm in length, 111 mm in height, and 35 mm in width, while the Radeon PRO V620 stretches to 267 mm, 120 mm, and 50 mm. For a system builder constrained by chassis space or power delivery, the CMP 40HX is the easier fit. The trade-off is performance, and the benchmark data shows exactly how much performance is sacrificed.

The use-case split becomes clear: the Radeon PRO V620 is for workloads that need maximum compute throughput, large memory buffers, and the headroom to handle demanding parallel tasks. The CMP 40HX is for environments where power density, heat output, and physical footprint take priority over raw numbers. Neither card has display outputs, so both are purely compute or mining accelerators, but they serve different scales of operation.

Architecture Differences

The architectural gap between these two is generational and fundamental. The AMD Radeon PRO V620 uses the Navi 21 chip built on RDNA 2.0 architecture, fabricated on a 7 nm process at TSMC. The NVIDIA CMP 40HX uses the TU106 chip on Turing architecture, also from TSMC but on a 12 nm process. That process difference alone explains a great deal: the AMD chip packs 26,800 million transistors into a 520 mm² die, achieving a transistor density of 51.5 million per square millimeter. The NVIDIA chip contains 10,800 million transistors on a 445 mm² die, for a density of 24.3 million per square millimeter. The AMD part is more than twice as dense, and that density translates directly into compute resources.

The shading unit count tells the same story. The Radeon PRO V620 has 4608 shading units, 288 texture mapping units, and 128 render output units. The CMP 40HX has 2304 shading units, 144 TMUs, and 64 ROPs. Every resource on the AMD card is exactly double the NVIDIA card's count. The ray tracing cores follow the same pattern: 72 on the AMD side versus 36 on the NVIDIA side. The NVIDIA card does have 288 tensor cores, a feature the AMD card lacks entirely, but those tensor cores do not appear in the recorded benchmark scores, and their absence from the AMD architecture means the comparison rests on the traditional compute paths.

Clock speeds differ substantially. The Radeon PRO V620 runs at a base of 1825 MHz and boosts to 2200 MHz, while the CMP 40HX operates at 1470 MHz base and 1650 MHz boost. The AMD card's higher clocks compound its resource advantage. Memory clocks also favor AMD: 2000 MHz with 16 Gbps effective transfer versus 1750 MHz with 14 Gbps effective. Both use GDDR6, both use a 256-bit bus, but the AMD card achieves 512.0 GB/s bandwidth against 448.0 GB/s for the NVIDIA part.

The API support is identical: both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. That means software compatibility is not a differentiator. The differences are internal, and they are stark.

Head-to-Head Benchmarks

The two recorded benchmark tests show a consistent and decisive pattern. In Geekbench OpenCL, the AMD Radeon PRO V620 scores 128580, while the NVIDIA CMP 40HX scores 93395. The AMD card leads by 37.7%. This is a large margin, but it is not the largest in the comparison.

The Vulkan test is where the gap widens dramatically. The Radeon PRO V620 scores 144364, and the CMP 40HX scores 77879. The delta is 85.4% in favor of the AMD card. That is nearly double the OpenCL margin, and it suggests that the AMD architecture handles Vulkan's compute and rendering paths with substantially greater efficiency, or that the NVIDIA card's Turing design is less well-suited to this particular workload. The numbers do not reveal the cause, only the effect: in Vulkan, the AMD card is in a different league.

Breaking down the averages, the Radeon PRO V620's average of 136472 is 59.4% higher than the CMP 40HX's 85637. The nearest rivals in the database contextualize these scores. The Radeon PRO V620 sits between the AMD Radeon Pro W6800X Duo at 135774 (0.5% behind) and the AMD Radeon PRO W6800 at 135396 (0.8% behind), with the NVIDIA A10M at 135230 (0.9% behind) and the NVIDIA RTX 4000 Ada Generation at 135218 (0.9% behind) close behind. The CMP 40HX, by contrast, finds itself surrounded by lower-tier cards: the AMD Radeon PRO W7600 at 87108 is 1.7% ahead, the NVIDIA Quadro GP100 at 87445 is 2.1% ahead, while the AMD Radeon PRO W6600 at 81995 trails by 4.4% and the AMD Radeon Pro Vega 64X at 80959 trails by 5.8%.

The implication is that the Radeon PRO V620 competes with professional workstation GPUs at the top of the stack, while the CMP 40HX sits among mid-range professional cards. The delta percentages against their respective rivals show that the AMD card is tightly clustered with its peers, within 1% either way, whereas the CMP 40HX has more breathing room, with rivals spread across a wider band.

Specification Differences

The specification sheet reveals a comprehensive divergence. The most obvious difference is memory: the Radeon PRO V620 has 32 GB of GDDR6, while the CMP 40HX has 8 GB. Both use a 256-bit bus, but the AMD card's higher memory clock yields 512.0 GB/s versus 448.0 GB/s. For workloads that need large datasets resident in VRAM, the 32 GB capacity is a decisive advantage.

Compute resources scale accordingly. The AMD card has 4608 shading units, 288 TMUs, and 128 ROPs. The NVIDIA card has 2304 shading units, 144 TMUs, and 64 ROPs. The pixel rate is 281.6 GPixel/s for AMD versus 105.6 GPixel/s for NVIDIA. The texture rate is 633.6 GTexel/s versus 237.6 GTexel/s. Floating-point performance follows: FP32 is 20.28 TFLOPS for AMD against 7.603 TFLOPS for NVIDIA, and FP16 is 40.55 TFLOPS against 15.21 TFLOPS, both using the 2:1 ratio.

Power and physical specifications differ just as much. The AMD card consumes 300 W, needs two 8-pin connectors, and suggests a 700 W power supply. The NVIDIA card draws 185 W, uses a single 8-pin connector, and suggests a 450 W PSU. The bus interface is another key differentiator: the Radeon PRO V620 uses PCIe 4.0 x16, while the CMP 40HX uses PCIe 1.0 x4. That older, narrower interface could bottleneck data transfer in some systems, though the recorded benchmarks do not isolate this effect.

Dimensions reinforce the size difference. The AMD card is 267 mm long, 120 mm tall, and 50 mm wide. The NVIDIA card is 229 mm long, 111 mm tall, and 35 mm wide. Both are dual-slot. The transistor counts are 26,800 million versus 10,800 million, and the die sizes are 520 mm² versus 445 mm². The release dates are close: the CMP 40HX launched on 2021-02-24, and the Radeon PRO V620 followed on 2021-11-03. Both are end-of-life products. The CMP 40HX has a launch MSRP of 699 USD; the Radeon PRO V620 has no recorded launch MSRP.

FAQ

Q: Which card has more memory, and does it matter for benchmarks?

A: The AMD Radeon PRO V620 has 32 GB of GDDR6, while the NVIDIA CMP 40HX has 8 GB. The capacity difference does not directly appear in the Geekbench scores, but the AMD card's higher memory bandwidth (512.0 GB/s versus 448.0 GB/s) contributes to its 37.7% lead in OpenCL and 85.4% lead in Vulkan.

Q: Is the NVIDIA CMP 40HX ever the better choice based on the data?

A: The benchmark data shows no performance wins for the CMP 40HX. However, it draws 185 W versus 300 W, requires only one 8-pin connector instead of two, and is physically smaller (229 mm versus 267 mm long). For constrained power or space environments, it is the more practical card despite lower scores.

Q: How do these cards compare to their nearest rivals in the database?

A: The Radeon PRO V620 averages 136472, sitting within 0.9% of the AMD Radeon Pro W6800X Duo, AMD Radeon PRO W6800, NVIDIA A10M, and NVIDIA RTX 4000 Ada Generation. The CMP 40HX averages 85637, with the AMD Radeon PRO W7600 1.7% ahead and the AMD Radeon PRO W6600 4.4% behind.

Q: What is the performance gap in the recorded benchmarks?

A: In Geekbench OpenCL, the AMD card scores 128580 against 93395, a 37.7% lead. In Geekbench Vulkan, the AMD card scores 144364 against 77879, an 85.4% lead. The average scores are 136472 and 85637, a 59.4% difference.

Q: Why does the AMD card have such a large advantage in Vulkan?

A: The data shows an 85.4% delta in Vulkan, the largest margin in the comparison. The AMD card has 4608 shading units and 72 ray tracing cores, while the NVIDIA card has 2304 shading units and 36 ray tracing cores, along with 288 tensor cores. The Vulkan workload appears to favor the AMD architecture's resource allocation.

Q: Are both cards still viable for modern systems?

A: Both are end-of-life products. The Radeon PRO V620 supports PCIe 4.0 x16, while the CMP 40HX uses PCIe 1.0 x4, which may limit data throughput. Both support DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, so software compatibility is current, but the NVIDIA card's older interface and lower compute resources make it less suited to demanding workloads.

The Verdict

The recorded data points to a straightforward conclusion. The AMD Radeon PRO V620 is the superior performer by every measured metric. It wins both head-to-head benchmarks, achieves a 59.4% higher average score, and ranks in the 96th percentile against all GPUs, compared to the 93rd percentile for the CMP 40HX. Its nearest rivals are top-tier professional cards like the AMD Radeon Pro W6800X Duo and NVIDIA RTX 4000 Ada Generation, all within 0.9% of its average. For anyone prioritizing compute throughput, memory capacity, or Vulkan performance, the Radeon PRO V620 is the clear selection.

The NVIDIA CMP 40HX, however, is not without justification. Its 185 W power draw, single 8-pin connector, 450 W suggested PSU, and smaller physical footprint make it a lower-impact installation. Its nearest rivals in the database, such as the AMD Radeon PRO W7600 and NVIDIA Quadro GP100, are within 2.1% of its score, so it performs consistently within its tier. The card's 8 GB memory and 7.603 TFLOPS FP32 performance are sufficient for lighter compute tasks, and its 288 tensor cores offer capabilities the AMD card lacks, though those cores do not appear in the recorded benchmarks.

The choice depends on what the workload demands. If the priority is maximum compute power, the Radeon PRO V620 wins without qualification. If the priority is minimal power draw, compact size, and adequate performance for less demanding tasks, the CMP 40HX fills that role. The data does not equivocate: one card is built for throughput, the other for efficiency, and the benchmarks quantify the trade-off precisely.

DETAILED SPECIFICATIONS

SPECIFICATION
PRO V620
CMP 40HX
Core Specs
Shading Units
4,608
2,304 -50.0%
Shaders
4,608
2,304 -50.0%
TMUs
288
144 -50.0%
ROPs
128
64 -50.0%
Compute Units
72
—
SM Count
—
36
Clocks
Base Clock
1825 MHz
1470 MHz
Boost Clock
2200 MHz
1650 MHz
Memory Clock
2000 MHz 16 Gbps effective
1750 MHz 14 Gbps effective
Memory
Memory Size
32 GB
8 GB
VRAM (MB)
32,768
8,192 -75.0%
Memory Type
GDDR6
GDDR6
Memory Bus
256 bit
256 bit
Bandwidth
512.0 GB/s
448.0 GB/s
Cache
L1 Cache
128 KB per Array
64 KB (per SM)
L2 Cache
4 MB
4 MB
L3 Cache
128 MB
—
L0 Cache
32 KB per WGP
—
Performance
Pixel Rate
281.6 GPixel/s
105.6 GPixel/s
Texture Rate
633.6 GTexel/s
237.6 GTexel/s
FP32 (TFLOPS)
20.28 TFLOPS
7.603 TFLOPS
FP64 (TFLOPS)
1,267.2 GFLOPS (1:16)
237.6 GFLOPS (1:32)
FP16 (TFLOPS)
40.55 TFLOPS (2:1)
15.21 TFLOPS (2:1)
AI/RT
RT Cores
72
36 -50.0%
Tensor Cores
—
288
Power
TDP
300 W
185 W
TDP (W)
300
185 -38.3%
Suggested PSU
700 W
450 W
Power Connectors
2x 8-pin
1x 8-pin
Architecture
Architecture
RDNA 2.0
Turing
GPU Name
Navi 21
TU106
Generation
Radeon Pro Navi (Navi II Series)
Mining GPUs
Process Size
7 nm
12 nm
Transistors
26,800 million
10,800 million
Die Size
520 mm²
445 mm²
Foundry
TSMC
TSMC
Density
51.5M / mm²
24.3M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.1
3.0
CUDA
—
7.5
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
229 mm 9 inches
Height
120 mm 4.7 inches
111 mm 4.4 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 1.0 x4
Other
Launch Price
—
699 USD
Production
End-of-life
End-of-life
Predecessor
Radeon Pro Vega
—
View Radeon PRO V620 Details View CMP 40HX Details