NVIDIA P104-100 vs NVIDIA Tesla M40 24 GB Comparison
NVIDIA P104-100
Tesla M40 24 GB
PERFORMANCE BENCHMARKS
Analysis: NVIDIA P104-100 vs NVIDIA Tesla M40 24 GB
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA Tesla M40 24 GB has an average benchmark score of 41,707, while the NVIDIA P104-100 averages 32,982. The Tesla M40 also sits in the 83rd percentile of all GPUs, versus the 77th percentile for the P104-100.
Q: How do the two cards compare in Geekbench OpenCL performance?
A: The P104-100 is substantially ahead in this workload, scoring 52,368 versus the Tesla M40's 37,439. That represents a 28.5% advantage for the P104-100 in OpenCL compute.
Q: Which card wins in Vulkan performance?
A: The Tesla M40 24 GB takes the Vulkan test with a score of 45,975, narrowly edging out the P104-100's 45,165. The margin is just 1.8%.
Q: What memory configurations do these cards offer?
A: The Tesla M40 24 GB comes with 24 GB of GDDR5 on a 384-bit bus, delivering 288.4 GB/s of bandwidth. The P104-100 has 4 GB of GDDR5X on a 256-bit bus, which provides 320.3 GB/s of bandwidth.
Q: Are these cards suitable for gaming or display output?
A: Neither card has display outputs. Both are compute-focused accelerators with no display connectivity, and both support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4.
Q: What are the power requirements for each card?
A: The Tesla M40 24 GB has a TDP of 250 W and requires a 600 W power supply with an 8-pin EPS connector. The P104-100 has no listed TDP but suggests a 200 W power supply and uses a single 8-pin connector.
The Verdict
The data presents two very different tools for different jobs. The Tesla M40 24 GB is the higher-tier card overall, with an average benchmark score of 41,707 and an 83rd percentile ranking. Its nearest rivals include the Tesla M40 (0.5% lower), the GeForce RTX 3080 Ti (1.3% lower), and the Radeon RX 7650 GRE (2.4% higher). This places it in solidly upper-midrange territory for compute workloads.
The P104-100, by contrast, averages 32,982 and sits in the 77th percentile. Its nearest rivals are mobile-class parts: the T600 Mobile (0.4% higher), the T550 Mobile (0.5% lower), the RTX 3050 Mobile (0.6% lower), and the Radeon Pro 570 (0.7% lower). This card competes with laptop-class GPUs rather than desktop flagships.
For anyone needing large memory capacity, the choice is clear: the Tesla M40's 24 GB frame buffer is unmatched by the P104-100's 4 GB. For raw OpenCL throughput, the P104-100 wins decisively. For Vulkan workloads, the Tesla M40 has a slight edge. Pick the Tesla M40 if you need capacity, stability in the 83rd percentile, or Vulkan performance. Pick the P104-100 if OpenCL compute throughput is your priority and 4 GB of memory is sufficient.
Head-to-Head Benchmarks
The benchmark record contains two direct comparisons, and each card claims one victory. The split is exactly even, but the margins tell different stories.
Geekbench OpenCL: P104-100 wins by a landslide. The P104-100 scores 52,368 against the Tesla M40's 37,439. That is a 28.5% gap, which is a massive difference in compute throughput. The P104-100's higher clock speeds and GDDR5X memory bandwidth of 320.3 GB/s likely contribute to this result, though the database does not break down the score by component. What matters is the outcome: for OpenCL-based compute tasks, the P104-100 is the clear pick.
Geekbench Vulkan: Tesla M40 wins by a hair. The Tesla M40 scores 45,975 versus 45,165 for the P104-100. The 1.8% margin is close enough to be within run-to-run variance, but the recorded data gives the win to the Tesla M40. Vulkan results show both cards performing at a similar level in this API.
The average benchmark scores reinforce the split. The Tesla M40's average of 41,707 is dragged down by its weaker OpenCL result, while the P104-100's average of 32,982 reflects the absence of a Vulkan-heavy workload in its score set. The P104-100 also has a 3DMark Steel Nomad DX12 score of 1,413, which is not available for the Tesla M40, so a direct comparison in that test is impossible.
Specification Differences
The two cards diverge sharply on memory and interface specifications.
Memory capacity and type: The Tesla M40 24 GB packs 24 GB of GDDR5 across a 384-bit bus, while the P104-100 has 4 GB of GDDR5X on a 256-bit bus. The Tesla's larger frame buffer is its defining feature; the P104's smaller capacity but faster GDDR5X memory gives it a bandwidth advantage of 320.3 GB/s versus 288.4 GB/s.
Clock speeds: The Tesla M40 runs at a base of 948 MHz and boosts to 1112 MHz, with memory at 1502 MHz (6 Gbps effective). The P104-100 runs much higher, with a base of 1607 MHz and a boost of 1733 MHz, plus memory at 1251 MHz (10 Gbps effective). The P104's clocks are roughly 70% higher on the core.
Compute resources: The Tesla M40 has 3,072 shading units, 192 texture mapping units, and 96 ROPs. The P104-100 has 1,920 shading units, 120 TMUs, and 64 ROPs. Despite having fewer units, the P104's higher clocks nearly equalize throughput.
Power and interface: The Tesla M40 draws 250 W TDP and needs a 600 W power supply with an 8-pin EPS connector. The P104-100 has no TDP listed, suggests a 200 W power supply, and uses a single 8-pin connector. The bus interface also differs: the Tesla uses PCIe 3.0 x16, while the P104 is limited to PCIe 1.0 x4. That interface difference is significant for data transfer-intensive workloads.
Physical dimensions: Both cards are dual-slot and 267 mm (10.5 inches) long, so they will fit in the same chassis spaces.
Architecture Differences
The architecture gap between these cards reflects two generations of NVIDIA design.
Process node and die size: The Tesla M40 uses the GM200 chip on a 28 nm TSMC process, measuring 601 mm² with 8,000 million transistors. The P104-100 uses the GP104 chip on a 16 nm TSMC process, measuring 314 mm² with 7,200 million transistors. The P104's die is nearly half the size yet packs almost as many transistors, giving it a transistor density of 22.9M per mm² versus 13.3M per mm² for the Tesla.
Shading architecture: The Tesla M40 is built on Maxwell 2.0, while the P104-100 uses Pascal. Pascal's key advantage is higher clock scaling: the P104's boost of 1733 MHz is far above the Tesla's 1112 MHz, which compensates for its fewer shading units. The result in raw FP32 is nearly identical: 6.832 TFLOPS for the Tesla versus 6.655 TFLOPS for the P104.
Memory architecture: The Tesla's GDDR5 on a 384-bit bus provides broad bandwidth with lower clock speeds. The P104's GDDR5X on a 256-bit bus achieves higher bandwidth with fewer memory chips. The P104 also supports FP16 at 104.0 GFLOPS (1:64 ratio), while the Tesla lists no FP16 capability.
Feature set: Both support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. Neither has ray tracing cores or tensor cores. The Tesla's generation is listed as "Tesla Maxwell (Mxx)" with a predecessor of Tesla Kepler and successor of Tesla Pascal. The P104-100 is classified under "Mining GPUs" with no predecessor or successor listed. The Tesla M40 was released on 2015-11-09, while the P104-100 came later on 2017-12-11.
Where Each One Wins
Choose the Tesla M40 24 GB for memory-bound workloads. The 24 GB frame buffer is six times larger than the P104's 4 GB. Any task that needs to hold large datasets, big textures, or multi-model inference batches will hit the P104's capacity limit long before the Tesla. The Tesla also wins the Vulkan benchmark with its 1.8% edge, and its 83rd percentile ranking places it above the P104's 77th percentile in the overall GPU hierarchy. The PCIe 3.0 x16 interface is also a substantial advantage over the P104's PCIe 1.0 x4 link for host-device data transfers.
Choose the P104-100 for OpenCL compute throughput. The 28.5% lead in Geekbench OpenCL is the largest margin in any head-to-head test. If your software stack uses OpenCL, this card delivers significantly more compute per second. The P104's higher bandwidth (320.3 GB/s) and much higher boost clock (1733 MHz) make it the better raw compute engine, provided 4 GB of memory is sufficient. Its lower power requirement (200 W suggested PSU versus 600 W) also makes it easier to slot into existing systems, though the PCIe 1.0 x4 interface will bottleneck any workload that frequently moves data between the card and system memory.
Mixed workloads: For Vulkan-based rendering or compute, the two are nearly tied. The Tesla wins by 1.8%, but that margin is small enough that other factors, such as memory capacity or power constraints, should decide the purchase. For DX12 workloads, the P104 has a recorded 3DMark Steel Nomad score of 1,413, but no comparable Tesla M40 result exists in the database, so no direct comparison is possible.
Final call: The Tesla M40 24 GB is the better all-round card by percentile, memory capacity, and interface. The P104-100 is a specialist that wins big in OpenCL but loses on capacity and interface speed. Match the card to your actual workload, not to the higher average score.