NVIDIA Tesla M60 vs NVIDIA Tesla P4 Comparison

NVIDIA
GEFORCE

NVIDIA Tesla M60

CORE STATE GM204
VRAM 8 GB
CLOCK SPEED 1178 MHz
TDP 300 W
BUS WIDTH 256 bit
ARCHITECTURE Maxwell 2.0
nm
PROCESS 28 nm
LAUNCH DATE 2015
VS
NVIDIA
GEFORCE

Tesla P4

CORE STATE GP104
VRAM 8 GB
CLOCK SPEED 1114 MHz
TDP 75 W
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

geekbench_opencl
29,506
34,947
geekbench_vulkan
31,473
40,309

Analysis: NVIDIA Tesla M60 vs NVIDIA Tesla P4

The Verdict

The data places the NVIDIA Tesla P4 clearly ahead of the NVIDIA Tesla M60 in both recorded benchmark disciplines. The P4 wins the Geekbench OpenCL test with a score of 34947 against 29506, a delta of 18.4 percent, and the Vulkan test with 40309 against 31473, a delta of 28.1 percent. The database records two benchmark wins for the P4 and zero for the M60.

The P4 also holds a higher percentile ranking across all GPUs, sitting at the 81st percentile versus the M60's 75th. Its average benchmark score of 37628 places it within 0.1 percent of the NVIDIA GeForce RTX 4070 (37648) and 0.3 percent ahead of the AMD Radeon RX Vega 56 (37507). The M60's average of 30490 places it level with the NVIDIA CMP 70HX (30476, 0 percent delta) and 0.2 percent ahead of the AMD Radeon RX 6700 (30433). In short, the P4 competes in a higher performance tier entirely.

Who should pick which? The P4 is the choice for any workload that stresses compute throughput per watt and modern API performance, given its 18.4 to 28.1 percent lead in the recorded tests, its 16 nm process, and its 75 W power draw. The M60, by contrast, is a larger, older dual-slot card drawing 300 W, and its only advantage in the specification sheet is a slightly higher pixel rate (75.39 GPixel/s vs 71.30 GPixel/s) and a higher boost clock (1178 MHz vs 1114 MHz). Those two numbers do not translate into a benchmark victory in the recorded data.

For buyers or system integrators choosing between these two end-of-life Tesla accelerators for inference or virtualized workloads, the P4 is the stronger candidate on raw scores, efficiency, and API support. The M60 may still be considered only if the specific workload depends on pixel throughput characteristics that the P4 cannot match, but the benchmark record does not support that choice as a general rule.

FAQ

Q: Which GPU is faster in the recorded benchmarks?

A: The NVIDIA Tesla P4 wins both recorded tests. In Geekbench OpenCL it scores 34947 against the M60's 29506, a lead of 18.4 percent. In Geekbench Vulkan it scores 40309 against 31473, a lead of 28.1 percent.

Q: How do these cards compare to other GPUs in the database?

A: The P4 sits at the 81st percentile of all GPUs, with an average score of 37628, which puts it 0.1 percent behind the GeForce RTX 4070 and 0.3 percent ahead of the Radeon RX Vega 56. The M60 sits at the 75th percentile with an average of 30490, level with the CMP 70HX and 0.2 percent ahead of the Radeon RX 6700.

Q: What are the power requirements of each card?

A: The Tesla P4 has a 75 W TDP, requires no power connectors, and lists a suggested PSU of 250 W. The Tesla M60 has a 300 W TDP, requires one 8-pin connector, and lists a suggested PSU of 700 W.

Q: Do both cards support the same APIs?

A: Yes, both list DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. Neither card has display outputs.

Q: Which card has higher memory bandwidth?

A: The Tesla P4 has 192.3 GB/s of memory bandwidth, while the Tesla M60 has 160.4 GB/s. Both have 8 GB of GDDR5 memory on a 256 bit bus.

Q: Are both cards still in production?

A: No, both are marked as end-of-life in the database. The P4 was released in September 2016 and the M60 in August 2015.

Architecture Differences

The two cards represent different architecture generations from NVIDIA. The Tesla P4 is built on the Pascal architecture using the GP104 chip, fabricated on TSMC's 16 nm process. The Tesla M60 uses the older Maxwell 2.0 architecture with the GM204 chip, fabricated on TSMC's 28 nm process. This process gap is substantial: the P4 packs 7,200 million transistors into a 314 mm² die, yielding a transistor density of 22.9M per mm². The M60 contains 5,200 million transistors on a larger 398 mm² die, for a density of 13.1M per mm². The P4 achieves higher density on a smaller die with more transistors, a direct consequence of the newer node.

The P4 also carries more execution resources. It has 2560 shading units, 160 texture mapping units, and 64 ROPs. The M60 has 2048 shading units, 128 TMUs, and 64 ROPs. Neither card has ray tracing cores or tensor cores, as both predate those features in the database. The P4's compute rates reflect its larger shader count: 5.704 TFLOPS FP32 and 89.12 GFLOPS FP16 at a 1:64 ratio. The M60 delivers 4.825 TFLOPS FP32 and lists no FP16 figure at all.

Memory architecture is similar in capacity and bus width, but not in speed. Both use 8 GB of GDDR5 on a 256 bit bus. The P4 runs its memory at 1502 MHz with 6 Gbps effective and reaches 192.3 GB/s. The M60 runs at 1253 MHz with 5 Gbps effective and reaches 160.4 GB/s. The P4's 19.9 percent bandwidth advantage aligns with its overall benchmark lead.

Physical and power characteristics differ sharply. The P4 is a single-slot card measuring 168 mm (6.6 inches) with no power connectors and a 75 W TDP. The M60 is a dual-slot card measuring 267 mm (10.5 inches), requiring one 8-pin connector and drawing 300 W. The M60's suggested PSU is 700 W versus 250 W for the P4. Both use a PCIe 3.0 x16 bus interface and have no display outputs.

The generational relationship is explicit in the database: the M60's successor is Tesla Pascal, which is the P4's generation. The P4's predecessor is Tesla Maxwell, the M60's generation. The two cards are thus direct neighbors across a generational boundary, with the P4 representing the newer design in almost every measurable way.

Specification Differences

The following fields differ between the two cards in the database:

  • Architecture: Pascal (P4) vs Maxwell 2.0 (M60)
  • Chip: GP104 (P4) vs GM204 (M60)
  • Generation: Tesla Pascal (Pxx) vs Tesla Maxwell (Mxx)
  • Process node: 16 nm (P4) vs 28 nm (M60)
  • Transistors: 7,200 million (P4) vs 5,200 million (M60)
  • Die size: 314 mm² (P4) vs 398 mm² (M60)
  • Transistor density: 22.9M per mm² (P4) vs 13.1M per mm² (M60)
  • Base clock: 886 MHz (P4) vs 557 MHz (M60)
  • Boost clock: 1114 MHz (P4) vs 1178 MHz (M60)
  • Memory clock: 1502 MHz / 6 Gbps effective (P4) vs 1253 MHz / 5 Gbps effective (M60)
  • Memory bandwidth: 192.3 GB/s (P4) vs 160.4 GB/s (M60)
  • Shading units: 2560 (P4) vs 2048 (M60)
  • TMUs: 160 (P4) vs 128 (M60)
  • Pixel rate: 71.30 GPixel/s (P4) vs 75.39 GPixel/s (M60)
  • Texture rate: 178.2 GTexel/s (P4) vs 150.8 GTexel/s (M60)
  • FP32: 5.704 TFLOPS (P4) vs 4.825 TFLOPS (M60)
  • FP16: 89.12 GFLOPS at 1:64 (P4) vs not listed (M60)
  • TDP: 75 W (P4) vs 300 W (M60)
  • Slot width: Single-slot (P4) vs Dual-slot (M60)
  • Power connectors: None (P4) vs 1x 8-pin (M60)
  • Suggested PSU: 250 W (P4) vs 700 W (M60)
  • Length: 168 mm / 6.6 inches (P4) vs 267 mm / 10.5 inches (M60)
  • Release date: 2016-09-12 (P4) vs 2015-08-29 (M60)
  • Predecessor/Successor: Tesla Maxwell/Tesla Volta (P4) vs Tesla Kepler/Tesla Pascal (M60)

Fields that are identical: manufacturer (NVIDIA), foundry (TSMC), memory size (8 GB), memory type (GDDR5), memory bus width (256 bit), ROPs (64), DirectX (12_1), OpenGL (4.6), Vulkan (1.4), bus interface (PCIe 3.0 x16), display outputs (none), production status (end-of-life).

Head-to-Head Benchmarks

The database records two direct comparisons between these cards, and the Tesla P4 wins both by substantial margins.

In Geekbench OpenCL, the P4 scores 34947 against the M60's 29506. That is an 18.4 percent advantage. In absolute terms, the gap is 5441 points. The P4's FP32 throughput of 5.704 TFLOPS versus 4.825 TFLOPS for the M60 helps explain the margin: the P4 has roughly 18 percent more FP32 compute, and its OpenCL lead is nearly identical at 18.4 percent. The M60's higher boost clock of 1178 MHz versus 1114 MHz does not compensate for its lower shader count and older architecture.

In Geekbench Vulkan, the P4 extends its lead further. It scores 40309 against 31473, a delta of 28.1 percent, or 8836 points. This is a larger margin than the OpenCL result, suggesting the P4's Pascal architecture handles the Vulkan 1.4 API workload more efficiently than the M60's Maxwell 2.0 design, despite both cards listing the same API version support. The P4's 192.3 GB/s memory bandwidth versus 160.4 GB/s may also factor into the wider Vulkan gap, as Vulkan workloads often stress memory access patterns.

The M60's only recorded specification wins are pixel rate (75.39 GPixel/s vs 71.30 GPixel/s, a 5.7 percent advantage) and boost clock (1178 MHz vs 1114 MHz, a 5.7 percent advantage). Neither translates into a benchmark victory. The P4 counters with a higher texture rate (178.2 GTexel/s vs 150.8 GTexel/s, an 18.2 percent advantage), higher memory bandwidth, and more shading units.

The percentile data reinforces the P4's position. At the 81st percentile, the P4 ranks above the M60's 75th percentile. The P4's nearest rivals include the GeForce RTX 4070 (0.1 percent behind), the Radeon RX Vega 56 (0.3 percent ahead), and the Radeon PRO W6400 (1.3 percent ahead). The M60's nearest rivals are the CMP 70HX (0 percent delta), the Radeon RX 6700 (0.2 percent behind), and the Radeon RX 6800 (1.3 percent behind). The P4 competes with modern consumer and workstation GPUs; the M60 competes with cards that are one to two tiers lower in the database's performance ranking.

The overall verdict is unambiguous. The Tesla P4 wins both recorded benchmarks, holds a higher percentile rank, delivers more compute and memory bandwidth, and draws 225 W less power. The M60's dual-slot size and 8-pin connector requirement make it a heavier integration burden as well. The data supports the P4 as the superior accelerator in this pairing.

DETAILED SPECIFICATIONS

SPECIFICATION
Tesla M60
Tesla P4
Core Specs
Shading Units
2,048
2,560 +25.0%
Shaders
2,048
2,560 +25.0%
TMUs
128
160 +25.0%
ROPs
64
64 0.0%
SM Count
20
Clocks
Base Clock
557 MHz
886 MHz
Boost Clock
1178 MHz
1114 MHz
Memory Clock
1253 MHz 5 Gbps effective
1502 MHz 6 Gbps effective
Memory
Memory Size
8 GB
8 GB
VRAM (MB)
8,192
8,192 0.0%
Memory Type
GDDR5
GDDR5
Memory Bus
256 bit
256 bit
Bandwidth
160.4 GB/s
192.3 GB/s
Cache
L1 Cache
48 KB (per SMM)
48 KB (per SM)
L2 Cache
2 MB
2 MB
Performance
Pixel Rate
75.39 GPixel/s
71.30 GPixel/s
Texture Rate
150.8 GTexel/s
178.2 GTexel/s
FP32 (TFLOPS)
4.825 TFLOPS
5.704 TFLOPS
FP64 (TFLOPS)
150.8 GFLOPS (1:32)
178.2 GFLOPS (1:32)
FP16 (TFLOPS)
89.12 GFLOPS (1:64)
Power
TDP
300 W
75 W
TDP (W)
300
75 -75.0%
Suggested PSU
700 W
250 W
Power Connectors
1x 8-pin
None
Architecture
Architecture
Maxwell 2.0
Pascal
GPU Name
GM204
GP104
Generation
Tesla Maxwell (Mxx)
Tesla Pascal (Pxx)
Process Size
28 nm
16 nm
Transistors
5,200 million
7,200 million
Die Size
398 mm²
314 mm²
Foundry
TSMC
TSMC
Density
13.1M / mm²
22.9M / mm²
API Support
DirectX
12 (12_1)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
5.2
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Single-slot
Length
267 mm 10.5 inches
168 mm 6.6 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 3.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Tesla Kepler
Tesla Maxwell
Successor
Tesla Pascal
Tesla Volta
View Tesla M60 Details View Tesla P4 Details