NVIDIA P104-100 vs NVIDIA Tesla M60 Comparison
NVIDIA P104-100
Tesla M60
PERFORMANCE BENCHMARKS
Analysis: NVIDIA P104-100 vs NVIDIA Tesla M60
The NVIDIA P104-100 and NVIDIA Tesla M60 are both end-of-life, dual-slot, no-display-output accelerators from NVIDIA, but they target very different workloads. The data shows a clear performance hierarchy, with the P104-100 taking every benchmark in the head-to-head comparison. However, the Tesla M60 counters with a larger memory pool and a more modern bus interface, making the choice highly dependent on the specific application.
Head-to-Head Benchmarks
The head-to-head data contains two benchmark entries, and the NVIDIA P104-100 wins both decisively. In the Geekbench OpenCL test, the P104-100 scores 52,368 against the Tesla M60’s 29,506. This translates to a 77.5% advantage for the P104-100, a massive margin that indicates a fundamental difference in compute throughput. The P104-100’s average benchmark score of 32,982 further reinforces this, as it sits considerably above the Tesla M60’s 32,982 average.
The second test, Geekbench Vulkan, shows a narrower but still significant gap. The P104-100 scores 45,165, while the Tesla M60 manages 31,473. This is a 43.5% lead for the P104-100. While the Vulkan delta is smaller than the OpenCL gap, it still places the P104-100 in a different performance tier. The P104-100 also claims two total wins in the head-to-head, while the Tesla M60 has zero.
Looking at the broader context, the P104-100’s percentile ranking of 77 among all GPUs is slightly higher than the Tesla M60’s 75. The P104-100’s nearest rivals are the NVIDIA T600 Mobile (0.4% slower), the NVIDIA T550 Mobile (0.5% faster), and the NVIDIA GeForce RTX 3050 Mobile (0.6% faster). The Tesla M60’s nearest rivals are the NVIDIA CMP 70HX (exact tie at 0% delta), the AMD Radeon RX 6700 (0.2% slower), and the AMD Radeon RX 6800 (1.3% slower). The data indicates that the P104-100 competes in the same sphere as modern mobile gaming GPUs, while the Tesla M60 sits alongside desktop gaming cards from a previous generation.
The FP32 compute figures align with the benchmark results. The P104-100 delivers 6.655 TFLOPS, while the Tesla M60 produces 4.825 TFLOPS. This 27.5% raw compute advantage for the P104-100 is the underlying driver behind its benchmark wins. The texture rate also favors the P104-100, with 208.0 GTexel/s versus 150.8 GTexel/s for the Tesla M60. The pixel rate is another win for the P104-100 at 110.9 GPixel/s, compared to 75.39 GPixel/s for the Tesla M60.
FAQ
Q: Which GPU is faster in Geekbench OpenCL?
A: The NVIDIA P104-100 is significantly faster, scoring 52,368 compared to the Tesla M60’s 29,506. This is a 77.5% performance advantage.
Q: How do the two compare in Vulkan performance?
A: The P104-100 also wins in Vulkan, with a score of 45,165 versus 31,473 for the Tesla M60. The P104-100 holds a 43.5% lead in this test.
Q: What is the difference in average benchmark scores?
A: The P104-100 has an average benchmark score of 32,982, while the Tesla M60 has an average score of 30,490. The P104-100 is approximately 8.2% higher based on these averages.
Q: Which GPU has a higher memory bandwidth?
A: The P104-100 has a much higher memory bandwidth of 320.3 GB/s, while the Tesla M60 offers 160.4 GB/s. The P104-100’s bandwidth is exactly double that of the Tesla M60.
Q: Do both GPUs support the same APIs?
A: Yes, both support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. They are identical in their API feature set.
Q: Which GPU has a higher pixel fill rate?
A: The P104-100 has a higher pixel rate at 110.9 GPixel/s, while the Tesla M60 achieves 75.39 GPixel/s. The P104-100 is ahead by roughly 47% in this metric.
The Verdict
The benchmark data is unambiguous: the NVIDIA P104-100 is the superior performer in both tested workloads. Its 77.5% lead in OpenCL and 43.5% lead in Vulkan are substantial, and its higher FP32 compute (6.655 TFLOPS vs 4.825 TFLOPS) supports the idea that it is a more powerful compute device. The P104-100 also has a higher percentile ranking (77 vs 75), indicating it sits higher in the overall GPU performance distribution.
However, the Tesla M60 is not without merit. It offers 8 GB of memory, which is double the P104-100’s 4 GB. For workloads that require large datasets to reside in GPU memory, this capacity advantage could be critical, even if the bandwidth is lower. The Tesla M60 also uses a PCIe 3.0 x16 interface, while the P104-100 is limited to PCIe 1.0 x4. This is a stark difference that could impact data transfer speeds between the CPU and GPU, potentially bottlenecking the P104-100 in scenarios with frequent host-device communication.
The choice comes down to raw compute versus memory capacity and interface bandwidth. For pure compute throughput, the P104-100 is the clear winner based on the benchmark scores. For applications that are memory-capacity bound or rely heavily on PCIe transfers, the Tesla M60’s larger memory and faster bus interface might be more practical, despite its lower compute performance. The data suggests the P104-100 is aimed at mining or compute tasks where shader throughput is paramount, while the Tesla M60 is a more balanced, albeit older, accelerator for virtualized workloads where memory and I/O matter more.
Specification Differences
The two cards differ on several key specification points. The most obvious is memory: the P104-100 has 4 GB of GDDR5X on a 256-bit bus, while the Tesla M60 has 8 GB of GDDR5 on the same 256-bit bus. The P104-100’s memory runs at 10 Gbps effective, yielding a bandwidth of 320.3 GB/s. The Tesla M60’s memory is slower at 5 Gbps effective, resulting in a bandwidth of 160.4 GB/s.
Clock speeds also diverge significantly. The P104-100 has a base clock of 1607 MHz and a boost clock of 1733 MHz. The Tesla M60 has a much lower base clock of 557 MHz, but a boost clock of 1178 MHz. This large gap between base and boost on the Tesla M60 suggests a power-constrained design that relies on boosting to reach performance targets.
The power requirements are distinct. The P104-100 does not have a TDP listed, but its suggested PSU is 200 W. The Tesla M60 has a TDP of 300 W and a suggested PSU of 700 W. The bus interface is another differentiator: the P104-100 uses PCIe 1.0 x4, while the Tesla M60 uses PCIe 3.0 x16.
Shading unit counts are close, with the P104-100 at 1920 and the Tesla M60 at 2048. The texture mapping units are 120 for the P104-100 and 128 for the Tesla M60. Both have 64 ROPs. The P104-100 has a higher process node density at 22.9M transistors per mm², while the Tesla M60 is at 13.1M per mm².
Architecture Differences
The architectural gap is generational. The P104-100 is built on the Pascal architecture using the GP104 chip, fabricated on a 16 nm process at TSMC. The Tesla M60 uses the older Maxwell 2.0 architecture with the GM204 chip, also from TSMC but on a 28 nm process. This process difference is a major factor in the P104-100’s efficiency and clock speed advantage.
Transistor counts differ, with the P104-100 housing 7,200 million transistors on a 314 mm² die. The Tesla M60 has 5,200 million transistors on a larger 398 mm² die. Despite having fewer transistors, the Tesla M60’s die is larger, which is a direct consequence of the older, less dense 28 nm manufacturing process.
Memory technology also separates them. The P104-100 uses GDDR5X, while the Tesla M60 uses GDDR5. The P104-100 supports FP16 at 104.0 GFLOPS with a 1:64 ratio, while the Tesla M60 does not list an FP16 capability. Both share the same API support for DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4, and both lack dedicated ray tracing or tensor cores.
The release dates reflect their positions in NVIDIA’s product cycle. The P104-100 was released later, and its generation is listed as "Mining GPUs," indicating a specific purpose-built design. The Tesla M60 is from the "Tesla Maxwell (Mxx)" generation and has a predecessor of "Tesla Kepler" and a successor of "Tesla Pascal."
Where Each One Wins
The P104-100 wins in every benchmark category tested, making it the choice for raw compute performance. Its 77.5% advantage in OpenCL and 43.5% in Vulkan show it is a far more capable compute device. Its higher FP32 throughput (6.655 TFLOPS), texture rate (208.0 GTexel/s), and pixel rate (110.9 GPixel/s) all point to superior shader and fill-rate performance. This makes the P104-100 the better option for tasks like general-purpose GPU compute, mining, or any workload that is heavily dependent on floating-point math and texture operations.
The Tesla M60’s strengths lie outside the direct compute benchmarks. Its 8 GB of memory is double the P104-100’s 4 GB, making it the better choice for workloads that need to hold large models or datasets in VRAM. Its PCIe 3.0 x16 interface is also a significant advantage over the P104-100’s PCIe 1.0 x4, which could be a severe bottleneck for data transfer. In virtualized environments or cloud streaming scenarios, where multiple users share the GPU and memory capacity is more important than peak FLOPs, the Tesla M60’s memory and bus interface make it a more practical, if slower, solution. The data does not show the Tesla M60 winning any compute tests, so its value is entirely in its capacity and connectivity rather than its speed.