AMD Radeon Pro Vega 56 vs NVIDIA Tesla P40 Comparison

AMD
RADEON

AMD Radeon Pro Vega 56

CORE STATE Vega 10
VRAM 8 GB
CLOCK SPEED 1250 MHz
TDP 210 W
BUS WIDTH 2048 bit
ARCHITECTURE GCN 5.0
nm
PROCESS 14 nm
LAUNCH DATE 2017
VS
NVIDIA
GEFORCE

Tesla P40

CORE STATE GP102
VRAM 24 GB
CLOCK SPEED 1531 MHz
TDP 250 W
BUS WIDTH 384 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

geekbench_metal
63,145
N/A
geekbench_opencl
61,930
62,017
geekbench_vulkan
66,004
68,172

Analysis: AMD Radeon Pro Vega 56 vs NVIDIA Tesla P40

# Head-to-Head Benchmarks

The head-to-head data between the NVIDIA Tesla P40 and the AMD Radeon Pro Vega 56 reveals a surprisingly narrow contest, with the Tesla P40 taking both available benchmark comparisons but by margins that tell a more complex story than the raw win count suggests.

In the Geekbench OpenCL test, the Tesla P40 scores 62,017 against the Radeon Pro Vega 56's 61,930. That delta of just 0.1% is essentially a statistical tie — the kind of margin that could flip with driver revisions or thermal conditions. The data shows both cards landing within 87 points of each other, which given the architectural differences between Pascal and GCN 5.0, makes for a fascinating equivalence in raw compute throughput.

The Vulkan benchmark paints a clearer picture of separation. Here the Tesla P40 posts 70,237 while the Radeon Pro Vega 56 manages 66,004. That 6.4% advantage for NVIDIA is the largest gap in any metric between these two, suggesting the Pascal architecture's handling of Vulkan's explicit graphics and compute workloads gives it a real edge. Interestingly, the Radeon Pro Vega 56 counters with a Geekbench Metal score of 73,356 — a test the Tesla P40 cannot run at all, as it has no display outputs and is not designed for Apple's Metal API ecosystem.

Looking at average benchmark scores across all tests each card supports, the Radeon Pro Vega 56 edges ahead with 67,097 versus the Tesla P40's 66,127. That 1.5% aggregate advantage, as reflected in the nearestRivals data, shows AMD's card performing slightly better when all available workloads are averaged, even though it loses the two direct head-to-head comparisons. Both cards sit at the 91st percentile versus all GPUs, placing them in the same performance tier despite their different design philosophies.

The nearestRivals data contextualizes these cards further. The Tesla P40's average score sits 0.5% below the NVIDIA Tesla T4, 0.9% below the NVIDIA GeForce RTX 4090, and 1.8% below the NVIDIA Quadro P6000. The Radeon Pro Vega 56, meanwhile, sits 0.3% below the Quadro P6000, 0.5% above the Tesla T4, and 0.9% above the RTX 4090. This means the AMD card actually outperforms NVIDIA's flagship consumer GPU in average benchmark score, while the Tesla P40 trails it slightly — a notable inversion given the RTX 4090's reputation.

# The Verdict

The data points to a nuanced verdict that depends entirely on workload priority. For users who need maximum OpenCL and Vulkan performance, the NVIDIA Tesla P40 is the choice — it wins both head-to-head tests, with the Vulkan margin being particularly decisive at 6.4%. The Tesla P40 also brings 24 GB of GDDR5 memory versus the Radeon Pro Vega 56's 8 GB of HBM2, a capacity difference that could matter significantly for large dataset workloads, even if the AMD card's bandwidth is higher at 402.4 GB/s versus 347.1 GB/s.

However, the Radeon Pro Vega 56 counters with a higher average benchmark score of 67,097 versus 66,127, driven largely by its Metal performance of 73,356 — a test the Tesla P40 cannot participate in. For macOS environments or Metal-based compute, the AMD card is the only option between these two. The Radeon Pro Vega 56 also draws less power at 210 W versus 250 W, and its HBM2 memory delivers superior bandwidth despite the smaller capacity.

The production status for both is end-of-life, so this comparison is about legacy systems rather than new purchases. The Tesla P40 had a launch MSRP of 5,699 USD. The Radeon Pro Vega 56 has no listed launch MSRP.

# Where Each One Wins

The NVIDIA Tesla P40 wins in scenarios that leverage its strengths in traditional compute APIs. The 6.4% Vulkan advantage suggests better performance in Vulkan-based rendering and compute workloads. The card's 11.76 TFLOPS FP32 performance versus 8.960 TFLOPS for the AMD card indicates a 31% theoretical peak compute advantage. The Tesla P40 also offers substantially more memory — 24 GB versus 8 GB — which is decisive for models or datasets that exceed the AMD card's capacity. Its 96 ROPs versus 64, and 240 TMUs versus 224, give it higher pixel rate at 147.0 GPixel/s versus 80.00 GPixel/s, and texture rate at 367.4 GTexel/s versus 280.0 GTexel/s.

The Radeon Pro Vega 56 wins in memory bandwidth, with 402.4 GB/s from its 2048-bit HBM2 interface versus the Tesla P40's 347.1 GB/s from a 384-bit GDDR5 bus. That 16% bandwidth advantage could benefit memory-bound workloads. The AMD card's FP16 performance is dramatically better — 17.92 TFLOPS at 2:1 ratio versus the Tesla P40's 183.7 GFLOPS at 1:64 ratio. For applications that use FP16 tensor-style operations, the AMD card is nearly 100x faster. The Radeon Pro Vega 56 also has actual display outputs — 1x HDMI 2.0b and 3x DisplayPort 1.4a — while the Tesla P40 has none, making the AMD card viable for workstation display tasks. The AMD card's Metal support, evidenced by its 73,356 Metal score, gives it exclusive access to Apple ecosystem workloads.

# FAQ

Q: Which card wins in OpenCL performance?

A: The NVIDIA Tesla P40 wins by a razor-thin margin, scoring 62,017 versus the Radeon Pro Vega 56's 61,930 — a delta of just 0.1%, effectively a tie.

Q: How significant is the Vulkan performance gap?

A: The Tesla P40's 70,237 Vulkan score beats the Radeon Pro Vega 56's 66,004 by 6.4%, which is the largest performance difference in any directly comparable test.

Q: Can the Radeon Pro Vega 56 be used in a Mac?

A: The data indicates yes — it has a Geekbench Metal score of 73,356 and is listed under the "Radeon Pro Mac" generation, with display outputs including 1x HDMI 2.0b and 3x DisplayPort 1.4a. The Tesla P40 has no display outputs and no Metal benchmark.

Q: Which card has more memory?

A: The NVIDIA Tesla P40 has 24 GB of GDDR5, three times the Radeon Pro Vega 56's 8 GB of HBM2.

Q: Which card is more power-efficient?

A: The Radeon Pro Vega 56 has a lower TDP at 210 W versus the Tesla P40's 250 W, though the Tesla P40 does deliver higher FP32 performance at 11.76 TFLOPS versus 8.960 TFLOPS.

Q: How do these cards compare to the NVIDIA GeForce RTX 4090?

A: In average benchmark score, the Radeon Pro Vega 56 is 0.9% ahead of the RTX 4090, while the Tesla P40 is 0.9% behind it. Both cards trail the NVIDIA Quadro P6000, with the Tesla P40 1.8% behind and the Radeon Pro Vega 56 0.3% behind.

# Architecture Differences

The two cards represent fundamentally different architectural approaches. The NVIDIA Tesla P40 uses the GP102 chip built on TSMC's 16 nm process, part of the Pascal architecture. It packs 11,800 million transistors into a 471 mm² die, yielding a transistor density of 25.1M per mm². The Radeon Pro Vega 56 uses AMD's Vega 10 chip on GlobalFoundries' 14 nm process, implementing the GCN 5.0 architecture. That chip contains 12,500 million transistors across a 495 mm² die, with a nearly identical transistor density of 25.3M per mm².

The compute configurations differ notably. The Tesla P40 has 3,840 shading units, 240 TMUs, and 96 ROPs, while the Radeon Pro Vega 56 has 3,584 shading units, 224 TMUs, and only 64 ROPs. This gives NVIDIA a clear advantage in pixel throughput, but AMD's HBM2 memory interface is significantly wider at 2048 bits versus 384 bits.

FP16 performance is perhaps the starkest architectural difference. The Tesla P40's FP16 throughput is 183.7 GFLOPS at a 1:64 ratio, meaning it's severely deprioritized. The Radeon Pro Vega 56 delivers 17.92 TFLOPS FP16 at a 2:1 ratio, making it a much more capable card for half-precision workloads. This reflects AMD's design choice to support fast FP16 in GCN 5.0, while NVIDIA reserved that capability for its Volta architecture and tensor cores.

The API support also diverges. Both support DirectX 12 (12_1) and OpenGL 4.6, but NVIDIA offers Vulkan 1.4 while AMD is limited to Vulkan 1.3. The Tesla P40 has no display outputs, reflecting its server/compute orientation, while the Radeon Pro Vega 56 includes 1x HDMI 2.0b and 3x DisplayPort 1.4a. The Tesla P40's predecessor was Tesla Maxwell and its successor Tesla Volta, while the Radeon Pro Vega 56 has no listed predecessor or successor in this data.

# Specification Differences

The core clock speeds differ substantially: the Tesla P40 runs at 1303 MHz base and 1531 MHz boost, while the Radeon Pro Vega 56 operates at 1138 MHz base and 1250 MHz boost. Memory clocks also diverge — the Tesla P40's GDDR5 runs at 1808 MHz (7.2 Gbps effective), while the Radeon Pro Vega 56's HBM2 operates at 786 MHz (1572 Mbps effective).

Memory capacity is a major differentiator: 24 GB GDDR5 on a 384-bit bus for NVIDIA versus 8 GB HBM2 on a 2048-bit bus for AMD. Bandwidth favors AMD at 402.4 GB/s versus 347.1 GB/s. The Tesla P40 has a dual-slot form factor with an 8-pin EPS power connector and a 600 W suggested PSU, while the Radeon Pro Vega 56 is listed as IGP (integrated graphics processor) with no power connectors and no suggested PSU.

Physical dimensions are only available for the Tesla P40: 267 mm (10.5 inches) in length and 111 mm (4.4 inches) in height. The Tesla P40's launch MSRP was 5,699 USD; the Radeon Pro Vega 56 has no listed launch MSRP. Release dates show the Tesla P40 arriving on September 12, 2016, while the Radeon Pro Vega 56 followed on August 13, 2017. Both are end-of-life products, and both use PCIe 3.0 x16 interfaces.

DETAILED SPECIFICATIONS

SPECIFICATION
Pro Vega 56
Tesla P40
Core Specs
Shading Units
3,584
3,840 +7.1%
Shaders
3,584
3,840 +7.1%
TMUs
224
240 +7.1%
ROPs
64
96 +50.0%
Compute Units
56
—
SM Count
—
30
Clocks
Base Clock
1138 MHz
1303 MHz
Boost Clock
1250 MHz
1531 MHz
Memory Clock
786 MHz 1572 Mbps effective
1808 MHz 7.2 Gbps effective
Memory
Memory Size
8 GB
24 GB
VRAM (MB)
8,192
24,576 +200.0%
Memory Type
HBM2
GDDR5
Memory Bus
2048 bit
384 bit
Bandwidth
402.4 GB/s
347.1 GB/s
Cache
L1 Cache
16 KB (per CU)
48 KB (per SM)
L2 Cache
4 MB
3 MB
Performance
Pixel Rate
80.00 GPixel/s
147.0 GPixel/s
Texture Rate
280.0 GTexel/s
367.4 GTexel/s
FP32 (TFLOPS)
8.960 TFLOPS
11.76 TFLOPS
FP64 (TFLOPS)
560.0 GFLOPS (1:16)
367.4 GFLOPS (1:32)
FP16 (TFLOPS)
17.92 TFLOPS (2:1)
183.7 GFLOPS (1:64)
Power
TDP
210 W
250 W
TDP (W)
210
250 +19.0%
Suggested PSU
—
600 W
Power Connectors
None
8-pin EPS
Architecture
Architecture
GCN 5.0
Pascal
GPU Name
Vega 10
GP102
Generation
Radeon Pro Mac (Vega Series)
Tesla Pascal (Pxx)
Process Size
14 nm
16 nm
Transistors
12,500 million
11,800 million
Die Size
495 mm²
471 mm²
Foundry
GlobalFoundries
TSMC
Density
25.3M / mm²
25.1M / mm²
API Support
DirectX
12 (12_1)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.3
1.4
OpenCL
2.1
3.0
CUDA
—
6.1
Shader Model
6.7
6.8
Physical
Slot Width
IGP
Dual-slot
Length
—
267 mm 10.5 inches
Height
—
111 mm 4.4 inches
Outputs
1x HDMI 2.0b3x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 3.0 x16
Other
Launch Price
—
5,699 USD
Production
End-of-life
End-of-life
Predecessor
—
Tesla Maxwell
Successor
—
Tesla Volta
View Radeon Pro Vega 56 Details View Tesla P40 Details