AMD Radeon Instinct MI60 vs NVIDIA A10G Comparison

AMD
RADEON

AMD Radeon Instinct MI60

CORE STATE Vega 20
VRAM 32 GB
CLOCK SPEED 1800 MHz
TDP 300 W
BUS WIDTH 4096 bit
ARCHITECTURE GCN 5.1
nm
PROCESS 7 nm
LAUNCH DATE 2018
VS
NVIDIA
GEFORCE

A10G

CORE STATE GA102
VRAM 24 GB
CLOCK SPEED 1710 MHz
TDP 150 W
BUS WIDTH 384 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

geekbench_opencl
92,488
158,063
geekbench_vulkan
92,444
145,863

Analysis: AMD Radeon Instinct MI60 vs NVIDIA A10G

Head-to-Head Benchmarks

The recorded data shows a decisive performance advantage for the NVIDIA A10G in both compute APIs tested. In Geekbench OpenCL, the A10G scores 158,063 against 92,488 for the AMD Radeon Instinct MI60, a 70.9% delta. That is not a marginal gap; it is a generational-class difference in raw throughput for general compute workloads. The Vulkan result follows the same pattern: the A10G posts 145,863 versus 92,444, a 57.8% lead. Neither test favors the MI60, so the head-to-head win count stands at 2 wins for NVIDIA and 0 for AMD.

Context from the database's nearest-rival comparisons strengthens this conclusion. The A10G's average benchmark score is 151,963, placing it in the 97th percentile of all GPUs. Its closest competitors include the NVIDIA Tesla V100 PCIe 32 GB at 150,305 (1.1% behind), the AMD Radeon Pro W6800X at 160,671 (5.4% ahead of the A10G), and the NVIDIA A100 PCIe 40 GB at 162,504 (6.5% ahead). The A10G sits comfortably in that tier. For the MI60, the average score is 92,466, in the 93rd percentile. Its nearest rivals are the NVIDIA RTX A4500 at 91,671 (0.9% behind), the RTX A4500 Mobile at 91,134 (1.5% behind), the AMD Radeon Pro VII at 97,131 (4.8% ahead), and the AMD Radeon RX 7900M at 97,487 (5.2% ahead). The MI60 is essentially at parity with mid-range workstation GPUs, while the A10G is competing with top-tier accelerators. The 70.9% OpenCL delta between the two cards is far larger than the deltas within their respective rival clusters, meaning the A10G's advantage is not a statistical fluke but a real performance tier separation.

Breaking down the individual tests, the OpenCL gap is larger than the Vulkan gap. This suggests the A10G's architecture handles the OpenCL compute path particularly well, possibly due to its driver maturity or compute-specific resource allocation. The Vulkan delta, while smaller, still exceeds 57%, so there is no workload type in these benchmarks where the MI60 closes the gap to a competitive level. The MI60's best result in either test is 92,488, which is barely half of the A10G's Vulkan score of 145,863. In practical terms, any application relying on these APIs would see the A10G finish rendering or computation in roughly 60% to 70% less time, assuming linear scaling.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA A10G has an average benchmark score of 151,963, which is 64.3% higher than the AMD Radeon Instinct MI60's 92,466. The A10G also ranks higher in the database, at the 97th percentile of all GPUs compared to the MI60's 93rd percentile.

Q: How do the two GPUs compare in Vulkan performance?

A: The A10G scores 145,863 in Geekbench Vulkan, while the MI60 scores 92,444. The A10G leads by 57.8%. This is a consistent advantage, though slightly smaller than the 70.9% lead seen in OpenCL.

Q: What memory configuration does each card use?

A: The NVIDIA A10G uses 24 GB of GDDR6 on a 384-bit bus, delivering 600.2 GB/s of bandwidth. The AMD Radeon Instinct MI60 uses 32 GB of HBM2 on a 4096-bit bus, delivering 1.02 TB/s of bandwidth, nearly double the A10G's bandwidth.

Q: Which card has a higher peak FP32 performance?

A: The A10G has 31.52 TFLOPS of FP32 performance, while the MI60 has 14.75 TFLOPS. The A10G is more than twice as fast in single-precision compute, which explains its large lead in the benchmark scores.

Q: Are both cards end-of-life products?

A: Yes, both the NVIDIA A10G and the AMD Radeon Instinct MI60 are marked as end-of-life in the database. The A10G was released in April 2021, while the MI60 was released in November 2018.

Q: What is the power draw difference?

A: The A10G has a TDP of 150 W, while the MI60 has a TDP of 300 W. The A10G delivers substantially higher benchmark performance while drawing half the power, making it a more efficient compute solution per watt.

Architecture Differences

The two cards come from fundamentally different design philosophies. The NVIDIA A10G uses the GA102 chip built on the Ampere architecture, fabricated on Samsung's 8 nm process. It packs 28,300 million transistors into a 628 mm² die, giving a transistor density of 45.1 million per square millimeter. The MI60 uses the Vega 20 chip on AMD's GCN 5.1 architecture, built on TSMC's 7 nm process. It contains 13,230 million transistors on a much smaller 331 mm² die, with a density of 40.0 million per square millimeter. The A10G uses a larger, more transistor-rich chip, while the MI60 relies on a denser but smaller design.

Compute resource allocation differs dramatically. The A10G has 9,216 shading units, 288 texture mapping units, and 96 render output units. It also includes 72 ray tracing cores and 288 tensor cores, making it a full-featured modern accelerator. The MI60 has 4,096 shading units, 256 TMUs, and 64 ROPs, with no ray tracing cores and no tensor cores at all. The A10G's shading unit count is more than double, which directly explains its FP32 throughput advantage (31.52 TFLOPS versus 14.75 TFLOPS). Interestingly, the MI60 has higher FP16 performance per FLOP ratio at 29.49 TFLOPS (2:1 ratio) versus the A10G's 31.52 TFLOPS (1:1 ratio), meaning the MI60's FP16 performance is only slightly lower despite its much smaller FP32 core count.

Both cards support PCIe 4.0 x16, but their memory subsystems diverge sharply. The A10G uses GDDR6 with a 384-bit bus and 600.2 GB/s bandwidth. The MI60 uses HBM2 with a 4096-bit bus and 1.02 TB/s bandwidth. The MI60's memory bandwidth is superior, which could help in memory-bound workloads, but the benchmark data shows the A10G still wins decisively in compute tests. The A10G also has a higher boost clock at 1710 MHz versus the MI60's 1800 MHz base and boost, though the MI60's boost clock is actually higher (1800 MHz) than the A10G's (1710 MHz). Neither clock advantage translates into a benchmark win for the MI60.

Feature support also differs. The A10G supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The MI60 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.3. The A10G has no display outputs, while the MI60 has a single mini-DisplayPort 1.4a. The MI60's display output is notable for a compute card, as it allows direct video output for debugging or visualization, whereas the A10G is strictly a headless accelerator.

Specification Differences

The specification table shows several clear divergences. The process node differs: the A10G uses 8 nm Samsung, the MI60 uses 7 nm TSMC. Transistor counts are 28,300 million versus 13,230 million, and die sizes are 628 mm² versus 331 mm². The A10G has a lower base clock (1320 MHz) but also a lower boost clock (1710 MHz) compared to the MI60's 1200 MHz base and 1800 MHz boost. Memory size is 24 GB GDDR6 for the A10G versus 32 GB HBM2 for the MI60, with bus widths of 384-bit and 4096-bit respectively. Bandwidth is 600.2 GB/s versus 1.02 TB/s.

The compute units differ as detailed above: shading units 9,216 versus 4,096, TMUs 288 versus 256, ROPs 96 versus 64. The A10G includes 72 RT cores and 288 tensor cores, while the MI60 has none. Pixel and texture rates are higher on the A10G (164.2 GPixel/s and 492.5 GTexel/s) versus the MI60 (115.2 GPixel/s and 460.8 GTexel/s). FP32 is 31.52 TFLOPS versus 14.75 TFLOPS, and FP16 is 31.52 TFLOPS versus 29.49 TFLOPS.

Power and physical specs also vary. The A10G has a TDP of 150 W, is single-slot, and uses an 8-pin EPS connector, with a suggested PSU of 450 W. The MI60 has a TDP of 300 W, is dual-slot, uses 1x 6-pin plus 1x 8-pin connectors, and requires a 700 W PSU. Both cards are 267 mm long and roughly 111-112 mm high, so they fit similar chassis. Release dates are April 2021 for the A10G and November 2018 for the MI60, a gap of roughly two and a half years. The A10G's predecessor is Tesla Turing and its successor is Server Ada; the MI60's predecessor is FirePro Data Center and it has no listed successor.

Where Each One Wins

The NVIDIA A10G wins in every benchmark category recorded. It is faster in OpenCL by 70.9% and faster in Vulkan by 57.8%. Its average score is 151,963, placing it in the 97th percentile, while the MI60 sits at 92,466 in the 93rd percentile. The A10G's FP32 performance is more than double, and its pixel fill rate is 42.5% higher. It also has ray tracing and tensor cores, which the MI60 completely lacks, making the A10G suitable for workloads that use these features, such as RT-accelerated rendering or tensor-based machine learning inference.

The AMD MI60 does have specific advantages in the specification sheet. Its 32 GB of HBM2 memory doubles the A10G's 24 GB capacity, and its 1.02 TB/s bandwidth is 70% higher. This could matter for very large models or datasets that exceed the A10G's memory, or for memory-bandwidth-bound operations. Its FP16 performance (29.49 TFLOPS) is close to the A10G's (31.52 TFLOPS), so mixed-precision workloads would not see a huge gap. The MI60 also has a display output, which the A10G lacks, allowing direct video output without a separate GPU. Its lower transistor count and smaller die might imply easier manufacturing yields, though both are end-of-life.

The MI60's power draw is double that of the A10G, so the A10G wins on efficiency per watt. The MI60's higher memory bandwidth does not translate to benchmark victories, suggesting the A10G's compute throughput dominates in the tested workloads. For any user prioritizing raw compute speed, ray tracing, tensor operations, or power efficiency, the A10G is the clear choice. For users needing more memory capacity or direct display output, the MI60 has niche appeal, but its compute scores lag significantly.

The Verdict

The data is unambiguous. The NVIDIA A10G outperforms the AMD Radeon Instinct MI60 in every recorded benchmark, with deltas of 70.9% in OpenCL and 57.8% in Vulkan. The A10G's average score is 151,963 against 92,466, a 64.3% overall advantage. This places the A10G at the 97th percentile versus the MI60's 93rd, and the A10G's nearest rivals are high-end accelerators like the A100 and W6800X, while the MI60 competes with mid-range cards like the RTX A4500.

Who should pick the A10G? Anyone running compute-heavy workloads, especially FP32 or FP16 with tensor cores, or any application that uses ray tracing. Its 150 W TDP makes it far easier to cool and power, and its single-slot design is more flexible in dense servers. The 24 GB GDDR6 memory is smaller than the MI60's 32 GB, but the A10G's compute advantage likely outweighs this for most tasks. The A10G is also newer by roughly two and a half years, which shows in its architecture.

Who should pick the MI60? Users who absolutely need 32 GB of HBM2 memory and 1.02 TB/s bandwidth, or who require a display output on the compute card. Its FP16 performance is close to the A10G's, so mixed-precision workloads are not a total loss. However, the MI60's 14.75 TFLOPS FP32 is less than half the A10G's, and its lack of RT and tensor cores limits its feature set. The MI60's dual-slot design and 300 W TDP also make it less practical in power-constrained environments.

The verdict is straightforward: the NVIDIA A10G is the superior compute accelerator in this comparison. The MI60's memory advantages do not overcome the A10G's massive compute lead in the recorded benchmarks. Choose the A10G unless the specific requirement for 32 GB HBM2 or a display output dictates otherwise.

DETAILED SPECIFICATIONS

SPECIFICATION
Instinct MI60
A10G
Core Specs
Shading Units
4,096
9,216 +125.0%
Shaders
4,096
9,216 +125.0%
TMUs
256
288 +12.5%
ROPs
64
96 +50.0%
Compute Units
64
—
SM Count
—
72
Clocks
Base Clock
1200 MHz
1320 MHz
Boost Clock
1800 MHz
1710 MHz
Memory Clock
1000 MHz 2 Gbps effective
1563 MHz 12.5 Gbps effective
Memory
Memory Size
32 GB
24 GB
VRAM (MB)
32,768
24,576 -25.0%
Memory Type
HBM2
GDDR6
Memory Bus
4096 bit
384 bit
Bandwidth
1.02 TB/s
600.2 GB/s
Cache
L1 Cache
16 KB (per CU)
128 KB (per SM)
L2 Cache
4 MB
6 MB
Performance
Pixel Rate
115.2 GPixel/s
164.2 GPixel/s
Texture Rate
460.8 GTexel/s
492.5 GTexel/s
FP32 (TFLOPS)
14.75 TFLOPS
31.52 TFLOPS
FP64 (TFLOPS)
7.373 TFLOPS (1:2)
985.0 GFLOPS (1:32)
FP16 (TFLOPS)
29.49 TFLOPS (2:1)
31.52 TFLOPS (1:1)
AI/RT
RT Cores
—
72
Tensor Cores
—
288
Power
TDP
300 W
150 W
TDP (W)
300
150 -50.0%
Suggested PSU
700 W
450 W
Power Connectors
1x 6-pin + 1x 8-pin
8-pin EPS
Architecture
Architecture
GCN 5.1
Ampere
GPU Name
Vega 20
GA102
Generation
Radeon Instinct (MIx)
Server Ampere (Axx)
Process Size
7 nm
8 nm
Transistors
13,230 million
28,300 million
Die Size
331 mm²
628 mm²
Foundry
TSMC
Samsung
Density
40.0M / mm²
45.1M / mm²
API Support
DirectX
12 (12_1)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.3
1.4
OpenCL
2.1
3.0
CUDA
—
8.6
Shader Model
6.7
6.8
Physical
Slot Width
Dual-slot
Single-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
112 mm 4.4 inches
Outputs
1x mini-DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
FirePro Data Center
Tesla Turing
Successor
—
Server Ada
View Radeon Instinct MI60 Details View A10G Details