NVIDIA P104-100 vs NVIDIA Tesla M40 Comparison
NVIDIA P104-100
Tesla M40
PERFORMANCE BENCHMARKS
Analysis: NVIDIA P104-100 vs NVIDIA Tesla M40
Head-to-Head Benchmarks
The recorded data includes two direct benchmark comparisons between the NVIDIA Tesla M40 and the NVIDIA P104-100. The results show a clear split between compute workloads and graphics API performance, with the newer Pascal-based card taking both recorded wins.
In Geekbench OpenCL, the NVIDIA P104-100 scores 52,368 against the Tesla M40's 39,192. This represents a 25.2% advantage for the P104-100, a substantial margin that reflects the architectural efficiency of the Pascal generation. The OpenCL test stresses raw compute throughput across the GPU, and the P104-100's higher clock speeds appear to overcome the Tesla M40's larger shading unit count.
The Geekbench Vulkan result is much closer. The P104-100 scores 45,165 while the Tesla M40 posts 44,602, a delta of only 1.2%. Vulkan's low-level API design tends to expose different hardware characteristics than OpenCL, and here the two cards are nearly identical in practical terms. This suggests that for graphics-oriented workloads using modern APIs, users would be hard pressed to notice a difference between the two in controlled testing.
Looking at the broader database context, the Tesla M40 holds an average benchmark score of 41,897 across all recorded tests, placing it at the 83rd percentile among all GPUs. Its nearest rivals in the database include the NVIDIA Tesla M40 24 GB at 41,707 (0.5% slower), the NVIDIA GeForce RTX 3080 Ti at 41,187 (1.7% slower), and the AMD Radeon Pro 5300 at 40,870 (2.5% slower). Interestingly, the AMD Radeon RX 7650 GRE beats it by 1.9% with a score of 42,723.
The P104-100's average benchmark score sits at 32,982, placing it at the 77th percentile. This is notably lower than the Tesla M40's average, despite the P104-100 winning both direct head-to-head tests. The reason becomes clear when examining the benchmark lists: the P104-100 has a third recorded test, 3DMark Steel Nomad DX12, where it scores 1,413. This additional workload, presumably a demanding modern gaming test, drags down the average and highlights that the P104-100's wins are concentrated in specific compute and API workloads rather than across the board. Its nearest rivals include the NVIDIA T600 Mobile at 32,849 (0.4% lower), the NVIDIA T550 Mobile at 33,161 (0.5% higher), and the NVIDIA GeForce RTX 3050 Mobile at 33,170 (0.6% higher).
The direct comparison data shows the P104-100 winning both recorded head-to-head tests, but the average benchmark scores tell a more nuanced story about overall capability.
Architecture Differences
The two cards come from different NVIDIA generations and are built on different process nodes. The Tesla M40 uses the GM200 chip based on Maxwell 2.0 architecture, fabricated on a 28 nm process at TSMC. The P104-100 uses the GP104 chip based on Pascal architecture, also from TSMC but on a 16 nm process. This node shrink is significant: the P104-100 packs 7,200 million transistors into a 314 mm² die, achieving a transistor density of 22.9 million per square millimeter. The Tesla M40, by contrast, contains 8,000 million transistors on a much larger 601 mm² die, yielding only 13.3 million per square millimeter. The Pascal design is clearly more efficient in terms of transistor packing, which helps explain how the P104-100 achieves competitive performance with far fewer resources.
Core configuration differs substantially. The Tesla M40 fields 3,072 shading units, 192 texture mapping units, and 96 raster operation pipelines. The P104-100 has 1,920 shading units, 120 TMUs, and 64 ROPs. Despite having 37.5% fewer shading units, the P104-100 nearly matches the Tesla M40 in raw FP32 throughput: 6.655 TFLOPS versus 6.832 TFLOPS. The Pascal architecture's higher clock speeds (1,607 MHz base and 1,733 MHz boost versus 948 MHz base and 1,112 MHz boost) close the gap that the Tesla M40's wider configuration would otherwise create.
Memory architecture also diverges. The Tesla M40 uses 12 GB of GDDR5 on a 384-bit bus, delivering 288.4 GB/s of bandwidth. The P104-100 uses 4 GB of GDDR5X on a 256-bit bus, achieving 320.3 GB/s. The P104-100's memory runs at 1,251 MHz with 10 Gbps effective speed, while the Tesla M40's memory runs at 1,502 MHz with 6 Gbps effective. The GDDR5X technology and higher effective data rate allow the narrower 256-bit bus to outpace the wider 384-bit bus in raw bandwidth.
The P104-100 also lists FP16 performance at 104.0 GFLOPS with a 1:64 ratio, indicating heavily reduced half-precision throughput. The Tesla M40 does not record an FP16 figure in the database. Both cards support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4, and neither has display outputs, making both unsuitable for direct monitor connection. Both are dual-slot cards with 267 mm (10.5 inches) length.
The bus interface presents another notable difference. The Tesla M40 uses PCIe 3.0 x16, while the P104-100 uses PCIe 1.0 x4. This is a dramatic reduction in host interface bandwidth for the P104-100, which may limit performance in scenarios that require frequent CPU-GPU communication, such as certain data transfer patterns or workloads that spill to system memory.
Where Each One Wins
The benchmark data indicates that the P104-100 wins in OpenCL compute tasks by a decisive margin. Its 25.2% advantage in the Geekbench OpenCL test suggests that the combination of Pascal's architecture, higher clocks, and faster memory bandwidth translates into superior raw compute execution. This would make the P104-100 the stronger choice for compute-oriented workloads that rely on OpenCL, including many scientific and data processing applications.
The Vulkan results are nearly identical, with the P104-100 leading by just 1.2%. For Vulkan-based rendering workloads, the two cards are effectively interchangeable in performance. Users would not observe meaningful differences in frame rates or render times.
The Tesla M40's strengths lie elsewhere. Its higher average benchmark score of 41,897 versus 32,982 indicates that across the full set of recorded tests, the Tesla M40 is the more consistent performer. The 12 GB memory capacity is four times larger than the P104-100's 4 GB, which matters for workloads with large working sets, such as machine learning model training, large dataset processing, or rendering scenes that exceed 4 GB of video memory. The 384-bit memory bus, while yielding slightly lower bandwidth than the P104-100, provides a wider path for memory-intensive operations that benefit from bus width.
The Tesla M40 also has a higher pixel rate at 106.8 GPixel/s compared to 110.9 GPixel/s for the P104-100, and a texture rate of 213.5 GTexel/s versus 208.0 GTexel/s. These are close figures, but the Tesla M40 edges ahead in texture throughput while the P104-100 leads in pixel throughput.
For the P104-100, the key wins are compute efficiency and bandwidth. Its 320.3 GB/s of memory bandwidth exceeds the Tesla M40's 288.4 GB/s, and its higher boost clock of 1,733 MHz allows rapid execution of clock-bound workloads. The smaller die and lower transistor count also suggest lower power draw, though the database does not record a TDP for the P104-100. The Tesla M40 is rated at 250 W with a suggested PSU of 600 W, while the P104-100 lists a suggested PSU of 200 W, indicating a substantially lower power envelope.
Specification Differences
The following specifications differ between the two cards:
| Specification | NVIDIA Tesla M40 | NVIDIA P104-100 |
|---|---|---|
| Architecture | Maxwell 2.0 | Pascal |
| Process node | 28 nm | 16 nm |
| Transistors | 8,000 million | 7,200 million |
| Die size | 601 mm² | 314 mm² |
| Transistor density | 13.3M / mm² | 22.9M / mm² |
| Base clock | 948 MHz | 1,607 MHz |
| Boost clock | 1,112 MHz | 1,733 MHz |
| Memory clock | 1,502 MHz (6 Gbps effective) | 1,251 MHz (10 Gbps effective) |
| Memory size | 12 GB | 4 GB |
| Memory type | GDDR5 | GDDR5X |
| Memory bus | 384 bit | 256 bit |
| Memory bandwidth | 288.4 GB/s | 320.3 GB/s |
| Shading units | 3,072 | 1,920 |
| TMUs | 192 | 120 |
| ROPs | 96 | 64 |
| Pixel rate | 106.8 GPixel/s | 110.9 GPixel/s |
| Texture rate | 213.5 GTexel/s | 208.0 GTexel/s |
| FP32 | 6.832 TFLOPS | 6.655 TFLOPS |
| FP16 | Not recorded | 104.0 GFLOPS (1:64) |
| TDP | 250 W | Not recorded |
| Power connectors | 8-pin EPS | 1x 8-pin |
| Suggested PSU | 600 W | 200 W |
| Bus interface | PCIe 3.0 x16 | PCIe 1.0 x4 |
| Release date | 2015-11-09 | 2017-12-11 |
| Generation | Tesla Maxwell (Mxx) | Mining GPUs |
Both cards share the same manufacturer, DirectX 12 (12_1) support, OpenGL 4.6, Vulkan 1.4, dual-slot width, no display outputs, 267 mm length, and end-of-life production status.
FAQ
Q: Which card is faster in OpenCL compute benchmarks?
A: The NVIDIA P104-100 scores 52,368 in Geekbench OpenCL compared to the Tesla M40's 39,192, a 25.2% advantage.
Q: How do the two cards compare in Vulkan performance?
A: The P104-100 leads by 1.2% in Geekbench Vulkan with a score of 45,165 versus 44,602 for the Tesla M40, making them nearly equivalent in this workload.
Q: Which card has more memory capacity?
A: The Tesla M40 has 12 GB of GDDR5 memory, while the P104-100 has 4 GB of GDDR5X memory. The P104-100 has higher bandwidth at 320.3 GB/s despite the smaller capacity, compared to 288.4 GB/s for the Tesla M40.
Q: Why does the Tesla M40 have a higher average benchmark score if it loses both head-to-head tests?
A: The Tesla M40's average score is 41,897 across its recorded benchmarks, while the P104-100 averages 32,982. The P104-100 has an additional recorded 3DMark Steel Nomad DX12 test scoring 1,413, which lowers its average despite winning the two Geekbench comparisons.
Q: What are the clock speed differences between these cards?
A: The P104-100 runs at 1,607 MHz base and 1,733 MHz boost, while the Tesla M40 runs at 948 MHz base and 1,112 MHz boost. The P104-100's memory runs at 1,251 MHz with 10 Gbps effective speed, while the Tesla M40's memory runs at 1,502 MHz with 6 Gbps effective.
Q: Do these cards support display output?
A: Neither card has display outputs, so both require a separate GPU for video output if used in a system.
The Verdict
The data supports a clear division of purpose. The NVIDIA P104-100 is the stronger choice for OpenCL compute workloads, delivering a 25.2% advantage in the direct comparison. Its higher clocks, faster memory bandwidth, and more efficient 16 nm Pascal architecture allow it to outperform the Tesla M40 despite having fewer shading units. The 1.2% Vulkan margin is effectively a tie, so users focused on Vulkan rendering can pick either card without meaningful performance penalty.
The NVIDIA Tesla M40, however, offers advantages that the head-to-head tests do not capture. Its 12 GB memory capacity is four times the P104-100's 4 GB, which is critical for workloads that exceed 4 GB of working set. Its higher average benchmark score of 41,897 versus 32,982 indicates better overall consistency across the full test suite. The PCIe 3.0 x16 interface provides far greater host bandwidth than the P104-100's PCIe 1.0 x4, which matters for data transfer-heavy tasks.
For compute users prioritizing raw OpenCL throughput and memory bandwidth, the P104-100 is the data-backed pick. For users needing large memory capacity, broader workload consistency, or a full-bandwidth PCIe interface, the Tesla M40 is the better fit. The P104-100's classification as a mining GPU and its limited 4 GB memory make it a specialized tool, while the Tesla M40's 12 GB capacity and higher percentile ranking suggest broader applicability despite its older architecture.