AMD Radeon Pro 570 vs NVIDIA Tesla M40 Comparison

AMD
RADEON

AMD Radeon Pro 570

CORE STATE Ellesmere
VRAM 4 GB
CLOCK SPEED 1105 MHz
TDP 150 W
BUS WIDTH 256 bit
ARCHITECTURE GCN 4.0
nm
PROCESS 14 nm
LAUNCH DATE 2017
VS
NVIDIA
GEFORCE

Tesla M40

CORE STATE GM200
VRAM 12 GB
CLOCK SPEED 1112 MHz
TDP 250 W
BUS WIDTH 384 bit
ARCHITECTURE Maxwell 2.0
nm
PROCESS 28 nm
LAUNCH DATE 2015

PERFORMANCE BENCHMARKS

geekbench_metal
39,945
N/A
geekbench_opencl
27,702
39,192
geekbench_vulkan
31,974
44,602

Analysis: AMD Radeon Pro 570 vs NVIDIA Tesla M40

The Verdict

The recorded data positions the NVIDIA Tesla M40 as the clear performance leader in this pairing. Across the two shared benchmark tests, the M40 wins both, with a 41.5% advantage in Geekbench OpenCL and a 39.5% advantage in Geekbench Vulkan. Its average benchmark score of 41,897 places it in the 83rd percentile of all GPUs, while the AMD Radeon Pro 570 averages 33,207, landing in the 78th percentile. The gap is substantial, but the intended use cases are entirely different.

The Tesla M40 is a dual-slot, 250 W add-in board with no display outputs, designed for compute workloads in a server or workstation chassis. It is a datacenter-oriented accelerator, not a card for a desktop workstation with a monitor attached. The Radeon Pro 570 is an integrated graphics processor, listed with an IGP slot width, no power connectors, and display outputs described as portable device dependent. It belongs to a Mac-focused product family, the Radeon Pro Mac 500 Series, and appears engineered for mobile or compact systems where power draw and physical footprint are constrained.

Who should pick which? Strictly from the data, anyone prioritizing raw compute throughput in OpenCL or Vulkan should choose the Tesla M40. Its 12 GB of GDDR5 memory, 384-bit bus, and 288.4 GB/s of bandwidth give it a decisive edge in memory-heavy workloads. The Radeon Pro 570 is the choice only when the system requires an integrated solution with no external power connectors, a 150 W TDP, and portable-device-dependent outputs, such as a laptop or an all-in-one Mac. The M40 cannot serve that role at all, as it has no display outputs. The data does not show a scenario where the Radeon Pro 570 wins on performance, but it wins on integration and system compatibility.

Architecture Differences

The two GPUs come from different architectural generations and foundries. The NVIDIA Tesla M40 uses the GM200 chip on Maxwell 2.0 architecture, built on TSMC's 28 nm process. The AMD Radeon Pro 570 uses Ellesmere on GCN 4.0, fabricated by GlobalFoundries on a 14 nm process. The process node difference is significant: 28 nm versus 14 nm, and the transistor density reflects this. The M40 packs 8,000 million transistors across a 601 mm² die, yielding a density of 13.3 million transistors per square millimeter. The Radeon Pro 570 fits 5,700 million transistors on a 232 mm² die, achieving 24.6 million per square millimeter, a much denser design.

The compute resources diverge sharply. The M40 has 3,072 shading units, 192 texture mapping units, and 96 render output units. The Radeon Pro 570 has 1,792 shading units, 112 TMUs, and only 32 ROPs. Neither GPU includes ray tracing cores or tensor cores. The M40's pixel rate is 106.8 GPixel/s and its texture rate is 213.5 GTexel/s. The Radeon Pro 570's pixel rate is 35.36 GPixel/s and its texture rate is 123.8 GTexel/s.

The M40 does not list a half-precision FP16 figure in the database, while the Radeon Pro 570 lists FP16 at 3.960 TFLOPS with a 1:1 ratio to its FP32 output. The M40's FP32 throughput is 6.832 TFLOPS, versus 3.960 TFLOPS for the AMD part. Memory configurations also differ: 12 GB of GDDR5 on a 384-bit bus for the M40, 4 GB of GDDR5 on a 256-bit bus for the Radeon Pro 570. Bandwidth is 288.4 GB/s versus 217.0 GB/s.

API support is close but not identical. Both support DirectX 12 and OpenGL 4.6. The M40 supports Vulkan 1.4; the Radeon Pro 570 supports Vulkan 1.3. The M40's DirectX feature level is 12_1, while the Radeon Pro 570 is 12_0. The M40 also uses an 8-pin EPS power connector and lists a suggested 600 W power supply, while the Radeon Pro 570 has no power connectors and no suggested PSU rating.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA Tesla M40, with an average benchmark score of 41,897 compared to the AMD Radeon Pro 570's 33,207. The M40 also sits in the 83rd percentile of all GPUs, while the Radeon Pro 570 sits in the 78th.

Q: How large is the performance gap in the shared tests?

A: In Geekbench OpenCL, the M40 scores 39,192 against 27,702 for the Radeon Pro 570, a 41.5% delta. In Geekbench Vulkan, the M40 scores 44,602 against 31,974, a 39.5% delta. The M40 wins both tests.

Q: Can the Tesla M40 be used in a system with a display attached?

A: No. The database lists its display outputs as "No outputs," meaning it cannot drive a monitor. The Radeon Pro 570 lists "Portable Device Dependent" outputs, indicating it can drive displays in portable systems.

Q: What is the power requirement difference?

A: The Tesla M40 has a 250 W TDP and requires an 8-pin EPS power connector, with a suggested 600 W power supply. The Radeon Pro 570 has a 150 W TDP, no power connectors, and no suggested PSU listed.

Q: How do their memory subsystems compare?

A: The M40 has 12 GB of GDDR5 on a 384-bit bus with 288.4 GB/s bandwidth. The Radeon Pro 570 has 4 GB of GDDR5 on a 256-bit bus with 217.0 GB/s bandwidth. The M40 offers triple the capacity and roughly a third more bandwidth.

Q: Which GPU supports a newer Vulkan version?

A: The NVIDIA Tesla M40 supports Vulkan 1.4, while the AMD Radeon Pro 570 supports Vulkan 1.3. Both support DirectX 12 and OpenGL 4.6, but the M40's DirectX feature level is 12_1 versus 12_0 for the AMD part.

Specification Differences

The database records the following fields where the two GPUs differ. The M40 is built on a 28 nm node by TSMC; the Radeon Pro 570 uses 14 nm by GlobalFoundries. The M40 has 8,000 million transistors on a 601 mm² die; the Radeon Pro 570 has 5,700 million on 232 mm². Transistor density is 13.3M per mm² versus 24.6M per mm².

Clocks differ modestly: the M40 runs at 948 MHz base and 1112 MHz boost, with memory at 1502 MHz (6 Gbps effective). The Radeon Pro 570 runs at 1000 MHz base and 1105 MHz boost, with memory at 1695 MHz (6.8 Gbps effective). The M40's shading unit count is 3,072 versus 1,792; TMUs are 192 versus 112; ROPs are 96 versus 32.

Pixel rate is 106.8 GPixel/s versus 35.36 GPixel/s. Texture rate is 213.5 GTexel/s versus 123.8 GTexel/s. FP32 is 6.832 TFLOPS versus 3.960 TFLOPS. The Radeon Pro 570 lists FP16 at 3.960 TFLOPS (1:1); the M40 lists no FP16 figure.

Memory capacity is 12 GB versus 4 GB, both GDDR5. Bus width is 384-bit versus 256-bit. Bandwidth is 288.4 GB/s versus 217.0 GB/s. The M40 has a 250 W TDP, dual-slot width, an 8-pin EPS connector, and a suggested 600 W PSU. The Radeon Pro 570 has a 150 W TDP, IGP slot width, no power connectors, and no suggested PSU.

The M40's display outputs are "No outputs"; the Radeon Pro 570's are "Portable Device Dependent." The M40 supports Vulkan 1.4 and DirectX 12_1; the Radeon Pro 570 supports Vulkan 1.3 and DirectX 12_0. Both support OpenGL 4.6 and PCIe 3.0 x16. The M40 is 267 mm (10.5 inches) long; the Radeon Pro 570 has no listed dimensions. Release dates differ: the M40 launched on 2015-11-09, the Radeon Pro 570 on 2017-06-04. Both are end-of-life.

Head-to-Head Benchmarks

The head-to-head data contains only two shared tests, and the NVIDIA Tesla M40 dominates both. In Geekbench OpenCL, the M40 scores 39,192 against the Radeon Pro 570's 27,702. That is a delta of 41.5%, meaning the M40 performs roughly 40% better in this compute workload. The score gap of 11,490 points is larger than the entire average score of many lower-tier GPUs in the database.

In Geekbench Vulkan, the M40 scores 44,602 against 31,974, a delta of 39.5%. The gap here is 12,628 points. Interestingly, the M40's Vulkan score is higher than its OpenCL score, while the Radeon Pro 570's Vulkan score is also higher than its OpenCL score, but the relative gap between the two remains nearly identical.

The M40's nearest rivals in the database include the NVIDIA Tesla M40 24 GB (avg score 41,707, delta 0.5%), the NVIDIA GeForce RTX 3080 Ti (41,187, delta 1.7%), and the AMD Radeon Pro 5300 (40,870, delta 2.5%). The AMD Radeon RX 7650 GRE scores 42,723, which is 1.9% higher than the M40. These figures show the M40 sits in a tight cluster of strong performers, all within roughly 2.5% of each other.

The Radeon Pro 570's nearest rivals are far lower in absolute terms. The NVIDIA GeForce RTX 3050 Mobile averages 33,170 (delta 0.1%), the NVIDIA T550 Mobile averages 33,161 (delta 0.1%), the NVIDIA P104-100 averages 32,982 (delta 0.7%), and the NVIDIA T600 Mobile averages 32,849 (delta 1.1%). Every one of these rivals is within 1.1% of the Radeon Pro 570, indicating that the AMD part is competitive within its own tier, but that tier is roughly 25% below the M40's tier.

The database also records a Geekbench Metal score for the Radeon Pro 570 of 39,945, a test that the M40 does not have. This suggests the AMD part may be better suited to Metal-based workloads on macOS, but no direct comparison is possible from the data.

Where Each One Wins

The NVIDIA Tesla M40 wins on every directly comparable metric. In raw compute, it leads by 41.5% in OpenCL and 39.5% in Vulkan. Its FP32 output of 6.832 TFLOPS is 72.5% higher than the Radeon Pro 570's 3.960 TFLOPS. Its pixel rate is 3 times higher, its texture rate is 72.5% higher, and its ROP count is 3 times higher. Memory capacity is 3 times larger, bandwidth is 32.9% higher, and the bus is 50% wider. For compute-heavy tasks like GPU rendering, scientific simulation, or machine learning inference where OpenCL and Vulkan are the relevant APIs, the M40 is the only rational choice from this data.

The AMD Radeon Pro 570 wins in integration and efficiency. It has a 150 W TDP versus 250 W, it requires no power connectors, and it is listed as an IGP with portable-device-dependent outputs. It can fit in systems where the M40 physically cannot: the M40 is a dual-slot, 267 mm card with no display outputs, while the Radeon Pro 570 is an integrated solution. The Radeon Pro 570 also has a denser transistor layout, 24.6M per mm² versus 13.3M, and a smaller die, 232 mm² versus 601 mm², which reflects its design for lower-power, space-constrained platforms. Its FP16 capability at 1:1 ratio may benefit certain half-precision workloads, though the M40 has no listed FP16 figure to compare.

The data also suggests the Radeon Pro 570's nearest competitive set is mobile and low-power discrete GPUs, all within 1.1% of its average score. The M40's nearest set is desktop and high-end workstation parts, all within 2.5%. The M40 belongs to a higher performance class entirely. The Radeon Pro 570 belongs to a class where the NVIDIA RTX 3050 Mobile, T550 Mobile, and T600 Mobile compete, and it matches them closely. Choose the Radeon Pro 570 for a portable Mac or compact system where power, space, and display output matter more than raw compute. Choose the Tesla M40 for a server or workstation where performance is the only priority.

DETAILED SPECIFICATIONS

SPECIFICATION
Pro 570
Tesla M40
Core Specs
Shading Units
1,792
3,072 +71.4%
Shaders
1,792
3,072 +71.4%
TMUs
112
192 +71.4%
ROPs
32
96 +200.0%
Compute Units
28
Clocks
Base Clock
1000 MHz
948 MHz
Boost Clock
1105 MHz
1112 MHz
Memory Clock
1695 MHz 6.8 Gbps effective
1502 MHz 6 Gbps effective
Memory
Memory Size
4 GB
12 GB
VRAM (MB)
4,096
12,288 +200.0%
Memory Type
GDDR5
GDDR5
Memory Bus
256 bit
384 bit
Bandwidth
217.0 GB/s
288.4 GB/s
Cache
L1 Cache
16 KB (per CU)
48 KB (per SMM)
L2 Cache
2 MB
3 MB
Performance
Pixel Rate
35.36 GPixel/s
106.8 GPixel/s
Texture Rate
123.8 GTexel/s
213.5 GTexel/s
FP32 (TFLOPS)
3.960 TFLOPS
6.832 TFLOPS
FP64 (TFLOPS)
247.5 GFLOPS (1:16)
213.5 GFLOPS (1:32)
FP16 (TFLOPS)
3.960 TFLOPS (1:1)
Power
TDP
150 W
250 W
TDP (W)
150
250 +66.7%
Suggested PSU
600 W
Power Connectors
None
8-pin EPS
Architecture
Architecture
GCN 4.0
Maxwell 2.0
GPU Name
Ellesmere
GM200
Generation
Radeon Pro Mac (500 Series)
Tesla Maxwell (Mxx)
Process Size
14 nm
28 nm
Transistors
5,700 million
8,000 million
Die Size
232 mm²
601 mm²
Foundry
GlobalFoundries
TSMC
Density
24.6M / mm²
13.3M / mm²
API Support
DirectX
12 (12_0)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.3
1.4
OpenCL
2.1
3.0
CUDA
5.2
Shader Model
6.7
6.8
Physical
Slot Width
IGP
Dual-slot
Length
267 mm 10.5 inches
Outputs
Portable Device Dependent
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 3.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Tesla Kepler
Successor
Tesla Pascal
View Radeon Pro 570 Details View Tesla M40 Details