AMD Radeon Pro 560 vs NVIDIA Tesla K20m Comparison
AMD Radeon Pro 560
Tesla K20m
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon Pro 560 vs NVIDIA Tesla K20m
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA Tesla K20m has a higher average benchmark score of 19089, while the AMD Radeon Pro 560 scores 17551. The K20m sits in the 64th percentile of all GPUs, slightly ahead of the Radeon Pro 560's 61st percentile.
Q: How large is the performance gap in OpenCL workloads?
A: In the Geekbench OpenCL test, the Tesla K20m scores 16241 versus 15504 for the Radeon Pro 560, a 4.8% advantage for the NVIDIA card. This is a relatively narrow margin, indicating the two are close in raw compute tasks.
Q: Which GPU wins in Vulkan performance?
A: The Tesla K20m dominates Vulkan, scoring 21936 against 16232 for the Radeon Pro 560, a 35.1% lead. This is the largest single-benchmark gap between the two cards in the recorded data.
Q: What is the transistor count difference between the two chips?
A: The Tesla K20m's GK110 chip packs 7,080 million transistors on a 561 mm² die, while the Radeon Pro 560's Polaris 21 has 3,000 million transistors on a 123 mm² die. The NVIDIA chip has more than double the transistors and nearly five times the die area.
Q: Which card supports more modern graphics APIs?
A: The Radeon Pro 560 supports DirectX 12 (12_0) and Vulkan 1.3, while the Tesla K20m supports DirectX 12 (11_0) and Vulkan 1.2.175. Both offer OpenGL 4.6, but the AMD card has a more recent API feature set.
Q: What is the power consumption difference?
A: The Tesla K20m has a TDP of 225 W and requires a 550 W suggested PSU with dual power connectors, whereas the Radeon Pro 560 has a 75 W TDP with no power connectors, making it an integrated-class part (IGP slot width).
Architecture Differences
The two GPUs come from entirely different architectural generations and design philosophies. The NVIDIA Tesla K20m uses the Kepler architecture, built on a 28 nm process at TSMC, with the GK110 chip. This is a compute-oriented design from the Tesla Kepler line, targeting high-throughput workloads. Its sheer scale is notable: 7,080 million transistors spread across a 561 mm² die, yielding a transistor density of 12.6 million transistors per mm². The chip includes 2,496 shading units, 208 texture mapping units, and 40 ROPs, reflecting a design optimized for massive parallel processing.
The AMD Radeon Pro 560, in contrast, uses the GCN 4.0 architecture with the Polaris 21 chip, fabricated on a 14 nm process at GlobalFoundries. This is a much smaller and denser design: 3,000 million transistors on a 123 mm² die, giving a transistor density of 24.4 million per mm², nearly double that of the Kepler chip. The Radeon Pro 560 has 1,024 shading units, 64 TMUs, and only 16 ROPs, a configuration suited for efficiency and mobile or embedded integration rather than raw compute throughput.
The memory subsystems diverge sharply. The K20m uses 5 GB of GDDR5 on a 320-bit bus, delivering 208.0 GB/s of bandwidth. The Radeon Pro 560 has 4 GB of GDDR5 on a 128-bit bus, with just 81.28 GB/s. This 2.5x bandwidth difference is a major architectural gap that affects any memory-bound workload. Clock speeds are also distinct: the K20m runs memory at 1300 MHz (5.2 Gbps effective), while the AMD card runs at 1270 MHz (5.1 Gbps effective), though the bus width makes the effective bandwidth far more important.
Feature support reveals their different eras. The K20m supports DirectX 12 (11_0) and Vulkan 1.2.175, while the Radeon Pro 560 supports DirectX 12 (12_0) and Vulkan 1.3. The AMD card also offers FP16 at 1.858 TFLOPS (1:1 ratio), a capability the K20m lacks entirely, as its FP16 field is null. The K20m has no display outputs, positioning it as a pure compute accelerator, whereas the Radeon Pro 560 has portable-device-dependent outputs, designed for integrated use in laptops or all-in-one systems.
Head-to-Head Benchmarks
The recorded head-to-head benchmarks cover two tests: OpenCL and Vulkan. In OpenCL, the Tesla K20m scores 16241 against the Radeon Pro 560's 15504, a 4.8% lead. This is a modest win, suggesting that in pure compute workloads, the older Kepler architecture still holds an edge, likely due to its far larger shader count (2,496 versus 1,024) and higher memory bandwidth (208.0 GB/s versus 81.28 GB/s). The narrow margin, however, indicates that the GCN 4.0 architecture's efficiency partially compensates for its smaller scale.
The Vulkan test tells a different story. The K20m scores 21936, while the Radeon Pro 560 manages only 16232, giving the NVIDIA card a 35.1% advantage. This is a decisive win, and it is striking because Vulkan is a modern API that the Radeon Pro 560 ostensibly supports with a newer version (1.3 versus 1.2.175). Despite the AMD card's more recent API implementation, the K20m's raw compute resources and driver optimization for this workload deliver a massive performance gap.
Across these two tests, the Tesla K20m wins both, giving it a 2-0 record in head-to-head matchups. The average benchmark scores reinforce this: 19089 for the K20m versus 17551 for the Radeon Pro 560, an 8.8% overall difference. Looking at the nearest rivals, the K20m sits between the GeForce RTX 4050 Mobile (19049, 0.2% lower) and the GeForce GTX 780 (19164, 0.4% higher), while the Radeon Pro 560 is close to the AMD Radeon Pro 460 (17509, 0.2% lower) and the Tesla K40c (17468, 0.5% lower). This places the two cards in adjacent performance tiers, but with the K20m consistently on top.
Specification Differences
The following fields differ between the two GPUs:
- Process Node: 28 nm (TSMC) versus 14 nm (GlobalFoundries)
- Transistors: 7,080 million versus 3,000 million
- Die Size: 561 mm² versus 123 mm²
- Transistor Density: 12.6M / mm² versus 24.4M / mm²
- Memory Clock: 1300 MHz (5.2 Gbps effective) versus 1270 MHz (5.1 Gbps effective)
- Memory Size: 5 GB versus 4 GB
- Memory Bus Width: 320 bit versus 128 bit
- Memory Bandwidth: 208.0 GB/s versus 81.28 GB/s
- Shading Units: 2496 versus 1024
- TMUs: 208 versus 64
- ROPs: 40 versus 16
- Pixel Rate: 36.71 GPixel/s versus 14.51 GPixel/s
- Texture Rate: 146.8 GTexel/s versus 58.05 GTexel/s
- FP32: 3.524 TFLOPS versus 1.858 TFLOPS
- FP16: null versus 1.858 TFLOPS (1:1)
- TDP: 225 W versus 75 W
- Slot Width: Dual-slot versus IGP
- Power Connectors: 1x 6-pin + 1x 8-pin versus None
- Suggested PSU: 550 W versus null
- Bus Interface: PCIe 2.0 x16 versus PCIe 3.0 x8
- Display Outputs: No outputs versus Portable Device Dependent
- DirectX Support: 12 (11_0) versus 12 (12_0)
- Vulkan Support: 1.2.175 versus 1.3
- Release Date: 2013-01-04 versus 2017-04-17
- Launch MSRP: 3,199 USD versus null
The Verdict
The data points to a clear overall winner in raw performance: the NVIDIA Tesla K20m. It wins both head-to-head benchmarks, holds a higher average score (19089 versus 17551), and delivers more than double the FP32 throughput (3.524 TFLOPS versus 1.858 TFLOPS). Its memory bandwidth advantage is enormous, at 208.0 GB/s versus 81.28 GB/s, which explains its dominance in Vulkan and its edge in OpenCL. For any workload that stresses compute or memory bandwidth, the K20m is the stronger choice.
However, the Radeon Pro 560 has its own domain. Its 75 W TDP versus 225 W means it consumes a third of the power, and its IGP slot width with no power connectors makes it suitable for integrated systems where the K20m's dual-slot size and 550 W PSU requirement would be impractical. The Radeon also supports more modern APIs: DirectX 12 (12_0) and Vulkan 1.3, plus FP16 compute, which the K20m lacks. For users needing a low-power, API-current GPU in a portable or embedded environment, the Radeon Pro 560 is the only viable option.
Where Each One Wins
NVIDIA Tesla K20m wins in:
- OpenCL compute (4.8% ahead of the Radeon Pro 560)
- Vulkan performance (35.1% ahead)
- Average benchmark score (19089 versus 17551)
- Memory bandwidth (208.0 GB/s versus 81.28 GB/s)
- Raw shader throughput (2,496 units versus 1,024)
- Pixel and texture rates (36.71 GPixel/s and 146.8 GTexel/s versus 14.51 and 58.05)
AMD Radeon Pro 560 wins in:
- Power efficiency (75 W TDP versus 225 W, no external power needed)
- Physical integration (IGP slot, no power connectors, portable-device-dependent outputs)
- API modernity (DirectX 12_0, Vulkan 1.3, FP16 support)
- Transistor density (24.4M / mm² versus 12.6M / mm²)
- PCIe generation (PCIe 3.0 x8 versus PCIe 2.0 x16)
The K20m is a high-performance accelerator for compute-heavy tasks where power and space are not constraints. The Radeon Pro 560 is a low-power, modern-API solution for integrated systems, sacrificing raw performance for efficiency and compatibility. The choice hinges on whether the workload prioritizes speed or system integration.