GPU Comparison

AMD
RADEON

AMD Instinct MI100

CORE STATE Arcturus
VRAM 32 GB
CLOCK SPEED 1502 MHz
TDP 300 W
BUS WIDTH 4096 bit
ARCHITECTURE CDNA 1.0
nm
PROCESS 7 nm
LAUNCH DATE 2020
VS
AMD
RADEON

Radeon PRO V620

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 2200 MHz
TDP 300 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
139,035
128,580
geekbench_vulkan
N/A
144,364

Analysis: AMD Instinct MI100 vs AMD Radeon PRO V620

AMD Instinct MI100 and AMD Radeon PRO V620 are both end-of-life workstation accelerators from AMD, built on the 7 nm TSMC process, yet they represent fundamentally different architectural philosophies. The data shows a single direct benchmark comparison and a wider field of average scores, placing both GPUs in the 96th percentile of all GPUs tested. The MI100, based on the CDNA 1.0 architecture with the Arcturus chip, is engineered for compute density, while the Radeon PRO V620, using the RDNA 2.0 architecture with the Navi 21 chip, is a more versatile graphics and compute hybrid. This analysis breaks down their performance, architectural differences, and practical use cases based solely on the provided facts.

Head-to-Head Benchmarks

The only direct head-to-head benchmark available is the Geekbench OpenCL test, which measures raw compute throughput. In this test, the AMD Instinct MI100 scores 139035 points, while the AMD Radeon PRO V620 scores 128580 points. This results in an 8.1% advantage for the MI100, a clear and decisive win for the compute-focused Instinct card. The data shows the MI100’s lead is not marginal; it is a substantial gap that reflects its higher FP32 throughput.

Looking at the broader average benchmark score, the MI100 again leads, averaging 139035 points across all its tested benchmarks. The Radeon PRO V620 averages 136472 points, putting the MI100 1.9% ahead in this aggregate metric. This average score is significant because it places the MI100 in the same performance tier as the NVIDIA Tesla V100 PCIe 16 GB (138063, 0.7% behind) and the Tesla V100 SXM2 32 GB (137731, 0.9% behind). The Radeon PRO V620, meanwhile, sits 0.5% above the AMD Radeon Pro W6800X Duo (135774) and 0.8% above the AMD Radeon PRO W6800 (135396).

However, the Radeon PRO V620 has a notable strength that the MI100 lacks entirely: a Vulkan benchmark score. The V620 achieves 144364 points in Geekbench Vulkan, which is not only higher than its own OpenCL score but also higher than the MI100’s OpenCL result. This indicates that while the MI100 dominates in OpenCL compute, the V620 is capable of higher performance in graphics-API-driven workloads. The MI100 has no Vulkan or graphics API support listed, making this a non-competitive area where the V620 wins by default.

In summary, the MI100 is the outright winner in the single shared compute benchmark, and it also holds a lead in average score. The V620’s only win is in Vulkan, which is a category where the MI100 does not participate. The data suggests a clear split: compute-heavy OpenCL tasks favor the MI100, while graphics-oriented or Vulkan-based tasks are the exclusive domain of the V620.

Where Each One Wins

The benchmark data paints a clear picture of two distinct use-case profiles. The AMD Instinct MI100 wins in raw compute throughput. Its 23.07 TFLOPS FP32 performance is 13.8% higher than the V620’s 20.28 TFLOPS, and its OpenCL score of 139035 is 8.1% higher. This makes the MI100 the stronger choice for workloads that rely heavily on parallel floating-point math, such as scientific simulation, machine learning inference, or any task that can leverage OpenCL’s compute model. Its memory subsystem, with 1.23 TB/s of bandwidth from HBM2 on a 4096-bit bus, is 2.4 times faster than the V620’s 512.0 GB/s GDDR6 implementation, giving it a massive advantage in memory-bound kernels where data transfer speed is the bottleneck.

The AMD Radeon PRO V620 wins in graphics-adjacent compute and API compatibility. Its Vulkan score of 144364 demonstrates superior performance in that API, which is increasingly used for cross-platform compute and rendering. The V620 also supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, while the MI100 lists no supported APIs. This makes the V620 the only viable option for any workload that requires a graphics API, even if it is used for non-rendering compute. Furthermore, the V620’s 72 ray accelerators (RT cores) provide hardware support for ray tracing, which the MI100 lacks entirely. Its higher pixel rate of 281.6 GPixel/s, compared to the MI100’s 96.13 GPixel/s, also points to stronger rasterization capabilities, despite both cards having no display outputs. The V620 is the winner in any scenario that needs a modern graphics feature set or hybrid compute-render workloads.

The wins are not evenly distributed. The MI100 wins on pure compute density and memory bandwidth. The V620 wins on API support, ray tracing, and pixel throughput. For a data center server dedicated to compute, the MI100 is superior. For a workstation handling a mix of compute and graphics, the V620 is the only option with the necessary feature set.

Architecture Differences

The fundamental divergence lies in the chip architecture. The MI100 uses CDNA 1.0 (Compute DNA), a design optimized exclusively for compute, while the V620 uses RDNA 2.0 (Radeon DNA), a graphics-first architecture with compute capabilities. This is reflected in their physical characteristics. The MI100’s Arcturus chip has 25,600 million transistors on a 750 mm² die, yielding a transistor density of 34.1M / mm². The V620’s Navi 21 chip has 26,800 million transistors on a smaller 520 mm² die, resulting in a much higher density of 51.5M / mm². The V620’s denser design is a result of its more complex graphics pipeline.

The compute resource allocation differs significantly. The MI100 has 7680 shading units, 480 TMUs, and 64 ROPs. The V620 has 4608 shading units, 288 TMUs, and 128 ROPs. The MI100 has 66.7% more shaders and 66.7% more TMUs than the V620, but the V620 has exactly double the ROPs. This explains the MI100’s higher texture rate (721.0 GTexel/s vs 633.6 GTexel/s) and the V620’s higher pixel rate (281.6 GPixel/s vs 96.13 GPixel/s). The V620 also includes 72 ray accelerators, a feature absent from the MI100. Neither card has tensor cores.

The memory architecture is another major differentiator. The MI100 uses 32 GB of HBM2 on a 4096-bit bus, achieving 1.23 TB/s of bandwidth. The V620 uses 32 GB of GDDR6 on a 256-bit bus, achieving 512.0 GB/s. The MI100’s bandwidth is 2.4 times higher, a critical advantage for compute workloads. The clocks also differ: the MI100 has a base clock of 1000 MHz and a boost of 1502 MHz, while the V620 starts at 1825 MHz base and boosts to 2200 MHz. The V620’s higher clocks partially compensate for its lower core count, but not enough to overcome the MI100’s raw compute advantage.

Both cards share the same 300 W TDP, dual-slot design, 2x 8-pin power connectors, 700 W suggested PSU, and PCIe 4.0 x16 interface. However, their physical dimensions differ slightly: the MI100 is 267 mm long and 111 mm high, while the V620 is the same length (267 mm) but taller at 120 mm and 50 mm wide. The MI100 was released on 2020-11-15, and the V620 on 2021-11-03, roughly a year later. The MI100’s predecessor is the Radeon Instinct, while the V620’s is the Radeon Pro Vega.

FAQ

Q: Which GPU has higher raw compute performance in OpenCL?

A: The AMD Instinct MI100 scores 139035 in Geekbench OpenCL, which is 8.1% higher than the Radeon PRO V620’s score of 128580. The MI100 also has higher FP32 throughput at 23.07 TFLOPS versus 20.28 TFLOPS.

Q: Does the Radeon PRO V620 support any APIs that the Instinct MI100 does not?

A: Yes. The V620 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The MI100 lists no supported graphics APIs, making the V620 the only option for those workloads.

Q: How does memory bandwidth compare between the two cards?

A: The MI100 has 1.23 TB/s of bandwidth from 32 GB of HBM2 on a 4096-bit bus. The V620 has 512.0 GB/s from 32 GB of GDDR6 on a 256-bit bus. The MI100’s bandwidth is 2.4 times higher.

Q: Which GPU has more shading units and TMUs?

A: The MI100 has 7680 shading units and 480 TMUs. The V620 has 4608 shading units and 288 TMUs. The MI100 has significantly more of both.

Q: Does either GPU support hardware ray tracing?

A: Only the Radeon PRO V620 has ray accelerators, with 72 RT cores. The Instinct MI100 has no RT cores listed.

Q: What is the average benchmark score for each GPU?

A: The MI100 has an average benchmark score of 139035, while the V620 averages 136472. The MI100 is 1.9% ahead in this aggregate metric.

Specification Differences

| Specification | AMD Instinct MI100 | AMD Radeon PRO V620 |

|:---------------|:-------------------|:--------------------|

| Chip | Arcturus | Navi 21 |

| Architecture | CDNA 1.0 | RDNA 2.0 |

| Transistors | 25,600 million | 26,800 million |

| Die Size | 750 mm² | 520 mm² |

| Transistor Density | 34.1M / mm² | 51.5M / mm² |

| Base Clock | 1000 MHz | 1825 MHz |

| Boost Clock | 1502 MHz | 2200 MHz |

| Memory Type | HBM2 | GDDR6 |

| Memory Bus Width | 4096 bit | 256 bit |

| Memory Bandwidth | 1.23 TB/s | 512.0 GB/s |

| Shading Units | 7680 | 4608 |

| TMUs | 480 | 288 |

| ROPs | 64 | 128 |

| RT Cores | N/A | 72 |

| Pixel Rate | 96.13 GPixel/s | 281.6 GPixel/s |

| Texture Rate | 721.0 GTexel/s | 633.6 GTexel/s |

| FP32 | 23.07 TFLOPS | 20.28 TFLOPS |

| FP16 | 46.14 TFLOPS (2:1) | 40.55 TFLOPS (2:1) |

| DirectX Support | N/A | 12 Ultimate (12_2) |

| OpenGL Support | N/A | 4.6 |

| Vulkan Support | N/A | 1.4 |

| Height | 111 mm (4.4 inches) | 120 mm (4.7 inches) |

| Width | N/A | 50 mm (2 inches) |

| Release Date | 2020-11-15 | 2021-11-03 |

| Predecessor | Radeon Instinct | Radeon Pro Vega |

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI100
PRO V620
Core Specs
Shading Units
7,680
4,608 -40.0%
Shaders
7,680
4,608 -40.0%
TMUs
480
288 -40.0%
ROPs
64
128 +100.0%
Compute Units
120
72 -40.0%
Clocks
Base Clock
1000 MHz
1825 MHz
Boost Clock
1502 MHz
2200 MHz
Memory Clock
1200 MHz 2.4 Gbps effective
2000 MHz 16 Gbps effective
Memory
Memory Size
32 GB
32 GB
VRAM (MB)
32,768
32,768 0.0%
Memory Type
HBM2
GDDR6
Memory Bus
4096 bit
256 bit
Bandwidth
1.23 TB/s
512.0 GB/s
Cache
L1 Cache
16 KB (per CU)
128 KB per Array
L2 Cache
8 MB
4 MB
L3 Cache
128 MB
L0 Cache
32 KB per WGP
Performance
Pixel Rate
96.13 GPixel/s
281.6 GPixel/s
Texture Rate
721.0 GTexel/s
633.6 GTexel/s
FP32 (TFLOPS)
23.07 TFLOPS
20.28 TFLOPS
FP64 (TFLOPS)
11.54 TFLOPS (1:2)
1,267.2 GFLOPS (1:16)
FP16 (TFLOPS)
46.14 TFLOPS (2:1)
40.55 TFLOPS (2:1)
AI/RT
RT Cores
72
Power
TDP
300 W
300 W
TDP (W)
300
300 0.0%
Suggested PSU
700 W
700 W
Power Connectors
2x 8-pin
2x 8-pin
Architecture
Architecture
CDNA 1.0
RDNA 2.0
GPU Name
Arcturus
Navi 21
Generation
Instinct (MIx)
Radeon Pro Navi (Navi II Series)
Process Size
7 nm
7 nm
Transistors
25,600 million
26,800 million
Die Size
750 mm²
520 mm²
Foundry
TSMC
TSMC
Density
34.1M / mm²
51.5M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
2.1
2.1
Shader Model
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
120 mm 4.7 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Radeon Instinct
Radeon Pro Vega
View Instinct MI100 Details View Radeon PRO V620 Details