AMD Radeon Instinct MI60 vs NVIDIA Tesla P40 Comparison
AMD Radeon Instinct MI60
Tesla P40
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon Instinct MI60 vs NVIDIA Tesla P40
# FAQ
Q: Which GPU has the higher average benchmark score?
A: The AMD Radeon Instinct MI60 records an average benchmark score of 92,466, while the NVIDIA Tesla P40 averages 65,095. This places the MI60 in the 93rd percentile of all GPUs, compared to the P40's 89th percentile.
Q: How large is the performance gap in OpenCL workloads?
A: In the Geekbench OpenCL test, the MI60 scores 92,488 versus the P40's 62,017, a difference of 49.1%. This is the larger of the two head-to-head deltas recorded.
Q: Does the MI60 also lead in Vulkan performance?
A: Yes, the MI60 scores 92,444 in Geekbench Vulkan, while the P40 reaches 68,172. The MI60 leads by 35.6%, making it the winner in both recorded benchmark tests.
Q: What is the memory configuration difference?
A: The MI60 has 32 GB of HBM2 memory on a 4096-bit bus, delivering 1.02 TB/s of bandwidth. The P40 uses 24 GB of GDDR5 on a 384-bit bus, providing 347.1 GB/s.
Q: Which card has a higher transistor density?
A: The MI60, built on TSMC's 7 nm process, packs 13,230 million transistors into a 331 mm² die, yielding 40.0 million transistors per mm². The P40, on 16 nm, has 11,800 million transistors across a 471 mm² die, for 25.1 million per mm².
Q: What are the power connector requirements?
A: The MI60 uses one 6-pin plus one 8-pin connector, with a suggested PSU of 700 W. The P40 uses a single 8-pin EPS connector and recommends a 600 W PSU.
---
Where Each One Wins
Based on the recorded data, the AMD Radeon Instinct MI60 wins every benchmark category where both cards were tested. It takes both Geekbench OpenCL and Vulkan, giving it 2 wins out of 2 head-to-head tests. The NVIDIA Tesla P40 records 0 wins in this comparison.
The MI60's strongest margin comes in OpenCL, where it outperforms the P40 by 49.1%. This suggests particularly large advantages in compute-heavy, general-purpose workloads that stress raw floating-point throughput and memory bandwidth. The Vulkan delta of 35.6% is also substantial, indicating that the MI60 holds its lead across graphics-oriented APIs as well.
However, the P40 is not without areas of relative strength. Its pixel rate of 147.0 GPixel/s exceeds the MI60's 115.2 GPixel/s, which points to faster rasterization throughput. Its 96 ROPs versus the MI60's 64 ROPs supports this advantage. For tasks heavily dependent on pixel fill, such as certain rendering passes, the P40's architecture may be more efficient despite its lower overall compute scores.
The P40 also operates at a higher base clock (1303 MHz versus 1200 MHz) and a smaller die area per transistor, which may translate to different thermal behavior in sustained workloads. Yet in aggregate benchmark performance, the MI60 is the clear winner in every recorded test.
---
Architecture Differences
The MI60 is built on AMD's GCN 5.1 architecture, using the Vega 20 chip, while the P40 relies on NVIDIA's Pascal architecture with the GP102 chip. These are fundamentally different designs separated by two process generations.
The MI60 uses a 7 nm TSMC process, allowing for 13,230 million transistors in a 331 mm² die. The P40 is on 16 nm TSMC, with 11,800 million transistors in a larger 471 mm² die. This gives the MI60 a transistor density of 40.0M per mm², compared to 25.1M per mm² for the P40. The smaller process node is a key reason the MI60 achieves higher performance with a similar transistor budget.
Memory architecture differs dramatically. The MI60 uses HBM2 on a 4096-bit bus, achieving 1.02 TB/s of bandwidth. The P40 uses GDDR5 on a 384-bit bus, delivering 347.1 GB/s. The MI60's bandwidth advantage is roughly 3x, which heavily influences compute-heavy benchmarks. The P40's 24 GB capacity is lower than the MI60's 32 GB, but both are generous for their eras.
The MI60 has 4096 shading units, 256 TMUs, and 64 ROPs. The P40 has 3840 shading units, 240 TMUs, and 96 ROPs. The MI60 leads in shader and texture throughput, while the P40 has more ROPs. This explains why the P40 wins on pixel rate (147.0 GPixel/s) but loses on texture rate (367.4 GTexel/s versus 460.8 GTexel/s).
FP16 support is another major divider. The MI60 delivers 29.49 TFLOPS FP16 at a 2:1 ratio, while the P40 only reaches 183.7 GFLOPS FP16 at a 1:64 ratio. This makes the MI60 vastly more capable in workloads that leverage half-precision arithmetic, a common pattern in AI inference and certain scientific simulations.
The MI60 supports PCIe 4.0 x16, while the P40 is limited to PCIe 3.0 x16. The MI60 also has a mini-DisplayPort 1.4a output, whereas the P40 has no display outputs at all. In terms of API support, both support DirectX 12 (12_1) and OpenGL 4.6, but the P40 lists Vulkan 1.4 while the MI60 lists Vulkan 1.3.
---
Specification Differences
The two cards differ across nearly every specification category.
- Process node: MI60 is 7 nm; P40 is 16 nm.
- Transistor count: MI60 has 13,230 million; P40 has 11,800 million.
- Die size: MI60 is 331 mm²; P40 is 471 mm².
- Transistor density: MI60 is 40.0M / mm²; P40 is 25.1M / mm².
- Base clock: MI60 is 1200 MHz; P40 is 1303 MHz.
- Boost clock: MI60 is 1800 MHz; P40 is 1531 MHz.
- Memory clock: MI60 is 1000 MHz (2 Gbps effective); P40 is 1808 MHz (7.2 Gbps effective).
- Memory size: MI60 is 32 GB; P40 is 24 GB.
- Memory type: MI60 is HBM2; P40 is GDDR5.
- Memory bus width: MI60 is 4096 bit; P40 is 384 bit.
- Memory bandwidth: MI60 is 1.02 TB/s; P40 is 347.1 GB/s.
- Shading units: MI60 has 4096; P40 has 3840.
- TMUs: MI60 has 256; P40 has 240.
- ROPs: MI60 has 64; P40 has 96.
- Pixel rate: MI60 is 115.2 GPixel/s; P40 is 147.0 GPixel/s.
- Texture rate: MI60 is 460.8 GTexel/s; P40 is 367.4 GTexel/s.
- FP32 performance: MI60 is 14.75 TFLOPS; P40 is 11.76 TFLOPS.
- FP16 performance: MI60 is 29.49 TFLOPS (2:1); P40 is 183.7 GFLOPS (1:64).
- TDP: MI60 is 300 W; P40 is 250 W.
- Power connectors: MI60 uses 1x 6-pin + 1x 8-pin; P40 uses 8-pin EPS.
- Suggested PSU: MI60 is 700 W; P40 is 600 W.
- Bus interface: MI60 is PCIe 4.0 x16; P40 is PCIe 3.0 x16.
- Display outputs: MI60 has 1x mini-DisplayPort 1.4a; P40 has no outputs.
- Vulkan support: MI60 is 1.3; P40 is 1.4.
- Release date: MI60 is 2018-11-17; P40 is 2016-09-12.
- Predecessor: MI60 follows FirePro Data Center; P40 follows Tesla Maxwell.
- Successor: MI60 has none listed; P40 is succeeded by Tesla Volta.
- Launch MSRP: P40 has a launch MSRP of 5,699 USD; MI60 has none listed.
---
Head-to-Head Benchmarks
The recorded head-to-head data covers two tests, both won by the AMD Radeon Instinct MI60.
In Geekbench OpenCL, the MI60 scores 92,488 against the P40's 62,017. This is a 49.1% delta, the largest margin in this comparison. The MI60's advantage here aligns with its 1.02 TB/s memory bandwidth, which is nearly three times the P40's 347.1 GB/s. OpenCL workloads often scale with memory throughput and raw FP32 compute, and the MI60 leads in both (14.75 TFLOPS versus 11.76 TFLOPS).
In Geekbench Vulkan, the MI60 scores 92,444 and the P40 scores 68,172. The delta is 35.6%, still a decisive win for the MI60. While the P40's higher pixel rate (147.0 GPixel/s) might suggest an edge in certain render paths, the Vulkan result shows the MI60's overall compute and texture throughput (460.8 GTexel/s) dominate the P40's 367.4 GTexel/s.
Context from the nearest rivals reinforces these results. The MI60's average score of 92,466 sits just 0.9% above the NVIDIA RTX A4500 (91,671) and 1.5% above the RTX A4500 Mobile (91,134). It trails the AMD Radeon Pro VII (97,131) by 4.8% and the AMD Radeon RX 7900M (97,487) by 5.2%. This places the MI60 in a competitive band near modern workstation GPUs.
The P40's average of 65,095 is 1.4% above the AMD Radeon Pro WX 9100 (64,212) and 2% above both the NVIDIA CMP 30HX (63,842) and AMD Radeon RX 9060 XT LP (63,830). It sits 1.4% below the AMD Radeon VII (66,004). The P40 is therefore competitive with mid-range cards of its generation but far behind the MI60.
---
The Verdict
The data clearly favors the AMD Radeon Instinct MI60 for anyone prioritizing raw compute performance. It wins both recorded benchmarks by substantial margins, offers over three times the memory bandwidth, and delivers nearly double the FP16 throughput. Its 32 GB HBM2 memory also provides more capacity for large datasets than the P40's 24 GB GDDR5.
The MI60 is the pick for compute-heavy workloads such as machine learning inference, scientific simulation, or any task that can exploit FP16 acceleration. Its PCIe 4.0 interface and modern 7 nm process also suggest better efficiency per transistor.
The NVIDIA Tesla P40 remains relevant for specific rasterization-focused tasks. Its higher pixel rate (147.0 GPixel/s) and larger ROP count (96 versus 64) give it an edge in fill-rate-limited scenarios, though this is not reflected in the aggregated benchmark scores. Its lower TDP (250 W versus 300 W) and simpler power connector requirement (8-pin EPS) may make it easier to deploy in certain server chassis.
For general-purpose compute, however, the MI60 is the superior choice. Its average benchmark score is 42% higher than the P40's, and it leads in both OpenCL and Vulkan. The P40's only advantages are pixel throughput, ROP count, and a lower power draw. Unless those specific traits are critical, the MI60 is the stronger accelerator according to all recorded metrics.