NVIDIA Tesla M40 vs NVIDIA Tesla P4 Comparison

NVIDIA
GEFORCE

NVIDIA Tesla M40

CORE STATE GM200
VRAM 12 GB
CLOCK SPEED 1112 MHz
TDP 250 W
BUS WIDTH 384 bit
ARCHITECTURE Maxwell 2.0
nm
PROCESS 28 nm
LAUNCH DATE 2015
VS
NVIDIA
GEFORCE

Tesla P4

CORE STATE GP104
VRAM 8 GB
CLOCK SPEED 1114 MHz
TDP 75 W
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

geekbench_opencl
39,192
34,947
geekbench_vulkan
44,602
40,309

Analysis: NVIDIA Tesla M40 vs NVIDIA Tesla P4

The NVIDIA Tesla M40 and NVIDIA Tesla P4 are both end-of-life datacenter accelerators from NVIDIA, but they represent different architectural generations and performance tiers. Benchmark data shows the Tesla M40 holds a decisive lead in compute workloads, while the Tesla P4 offers a dramatically more efficient physical footprint. The M40 wins both available head-to-head benchmarks, yet the P4’s design philosophy targets a different deployment scenario entirely.

Head-to-Head Benchmarks

The Tesla M40 dominates the Tesla P4 in every recorded benchmark test. In Geekbench OpenCL, the M40 scores 39,192 points against the P4’s 34,947 points, producing a 12.1% performance advantage. The Vulkan results follow the same pattern, with the M40 reaching 44,602 points versus the P4’s 40,309 points, a 10.7% lead. These are not marginal differences; they represent a consistent, measurable gap across both APIs.

This performance delta is substantial enough to place the M40 in a higher competitive tier. The M40’s average benchmark score of 41,897 places it at the 83rd percentile of all GPUs, while the P4’s 37,628 average score sits at the 81st percentile. The M40’s nearest rival, the Tesla M40 24 GB, averages 41,707 points (0.5% lower), meaning the standard 12 GB M40 is essentially on par with its higher-memory sibling. The P4’s closest competitor, the GeForce RTX 4070, averages 37,648 points — a mere 0.1% higher — indicating the P4 performs right at the edge of modern consumer graphics cards.

The raw specs explain this gap. The M40 packs 3,072 shading units, 192 texture mapping units, and 96 ROPs, versus the P4’s 2,560 shaders, 160 TMUs, and 64 ROPs. The M40 also delivers 6.832 TFLOPS of FP32 compute compared to the P4’s 5.704 TFLOPS. Pixel fill rates tell a similar story: the M40 outputs 106.8 GPixel/s, while the P4 manages 71.30 GPixel/s. Texture fill rates are 213.5 GTexel/s for the M40 versus 178.2 GTexel/s for the P4. Every compute-oriented metric favors the older Maxwell card.

Memory bandwidth further widens the gap. The M40 uses a 384-bit bus with 12 GB of GDDR5, achieving 288.4 GB/s of bandwidth. The P4 is limited to a 256-bit bus, 8 GB of GDDR5, and 192.3 GB/s. That is a 96.1 GB/s difference, which heavily impacts memory-bound workloads. The M40’s larger frame buffer and wider interface give it a structural advantage in large dataset processing.

The Verdict

The data points to a clear conclusion: the Tesla M40 is the superior compute performer. It wins both benchmarks, posts higher raw scores, and achieves a better percentile ranking. If raw throughput is the sole criterion, the M40 is the only choice between these two.

However, the P4 wins on efficiency and physical integration. Its 75 W TDP is one-third of the M40’s 250 W, and it requires no external power connectors. The P4 is a single-slot, 168 mm card, while the M40 is a dual-slot, 267 mm unit. The P4’s suggested power supply is 250 W versus the M40’s 600 W. For dense server deployments where power and space are constrained, the P4 presents a viable low-profile option.

The verdict depends on workload priority. For batch inference, rendering, or any task that can tolerate a larger card, the M40’s 12.1% OpenCL lead and 10.7% Vulkan lead justify its size. For edge deployments, multi-GPU arrays, or systems with strict power budgets, the P4’s 75 W draw makes it the only rational fit. The M40 is the performance winner; the P4 is the efficiency winner. Neither card is a general-purpose replacement for the other.

FAQ

Q: Which card has a higher average benchmark score?

A: The NVIDIA Tesla M40 has an average benchmark score of 41,897, while the NVIDIA Tesla P4 scores 37,628. The M40 leads by 4,269 points.

Q: How does the M40 compare to its nearest rival, the Tesla M40 24 GB?

A: The standard 12 GB M40 scores 41,897 on average, which is 0.5% higher than the Tesla M40 24 GB’s average of 41,707.

Q: What is the P4’s closest competitor in the benchmark database?

A: The NVIDIA GeForce RTX 4070 averages 37,648 points, which is just 0.1% higher than the P4’s 37,628 average. The AMD Radeon RX Vega 56 is also close at 37,507 points, 0.3% lower.

Q: Does the P4 support FP16 compute?

A: Yes, the Tesla P4 offers FP16 performance of 89.12 GFLOPS, but at a ratio of 1:64 relative to its FP32 throughput. The M40 has no recorded FP16 capability.

Q: What are the memory specifications for each card?

A: The M40 has 12 GB of GDDR5 on a 384-bit bus with 288.4 GB/s bandwidth. The P4 has 8 GB of GDDR5 on a 256-bit bus with 192.3 GB/s bandwidth.

Q: Which card has a higher pixel fill rate?

A: The Tesla M40 achieves 106.8 GPixel/s, compared to the P4’s 71.30 GPixel/s. The M40 leads by 35.5 GPixel/s.

Specification Differences

The two cards diverge on nearly every physical and compute specification. The M40 uses the GM200 chip on a 28 nm process from TSMC, with 8,000 million transistors on a 601 mm² die. The P4 uses the GP104 chip on a 16 nm process, also from TSMC, with 7,200 million transistors on a 314 mm² die. The P4’s smaller process allows a transistor density of 22.9M per mm², versus the M40’s 13.3M per mm².

Clock speeds are similar at the boost level: the M40 runs at 948 MHz base and 1112 MHz boost, while the P4 runs at 886 MHz base and 1114 MHz boost. The M40 has a higher base clock by 62 MHz, but the P4 boosts 2 MHz higher. Memory clocks are identical at 1502 MHz with 6 Gbps effective speed.

The M40’s shading units (3072), TMUs (192), and ROPs (96) all exceed the P4’s 2560, 160, and 64 respectively. FP32 compute is 6.832 TFLOPS for the M40 versus 5.704 TFLOPS for the P4. The M40’s TDP of 250 W requires a dual-slot cooler and an 8-pin EPS connector, with a suggested 600 W power supply. The P4’s 75 W TDP needs no power connector, fits in a single slot, and suggests a 250 W power supply.

Physical dimensions differ significantly: the M40 is 267 mm long, while the P4 is 168 mm long. The M40 has no display outputs, as does the P4. Both use a PCIe 3.0 x16 bus interface. The M40 was released on November 9, 2015; the P4 followed on September 12, 2016.

Architecture Differences

The M40 is built on Maxwell 2.0 architecture, while the P4 uses Pascal. This generational shift brings several key changes. The M40’s GM200 chip is a large, power-hungry design optimized for maximum throughput. The P4’s GP104 chip is a smaller, more efficient Pascal design that emphasizes performance per watt.

The process node difference is the most fundamental architectural split: 28 nm for Maxwell versus 16 nm for Pascal. This allows the P4 to pack 7,200 million transistors into less than half the die area of the M40 (314 mm² versus 601 mm²). The P4’s transistor density of 22.9M per mm² is nearly double the M40’s 13.3M per mm².

Memory architecture also reflects the generational change. Both use GDDR5, but the M40’s 384-bit bus is wider, enabling 288.4 GB/s bandwidth. The P4’s 256-bit bus is narrower, capping bandwidth at 192.3 GB/s. The M40’s 12 GB capacity is 50% larger than the P4’s 8 GB.

The P4 introduces FP16 compute support at 89.12 GFLOPS, though at a severely reduced 1:64 ratio. The M40 has no FP16 capability listed. Neither card includes ray tracing cores or tensor cores. Both support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4.

The M40’s predecessor is Tesla Kepler, and its successor is Tesla Pascal. The P4’s predecessor is Tesla Maxwell, and its successor is Tesla Volta. This places the two cards on opposite sides of the Pascal generation, with the M40 being the older architecture and the P4 being the newer one. The M40 is designated as Tesla Maxwell (Mxx), while the P4 is Tesla Pascal (Pxx).

Where Each One Wins

The Tesla M40 wins in every compute-heavy scenario. Its 12.1% OpenCL lead and 10.7% Vulkan lead make it the clear choice for tasks that demand maximum raw throughput. The higher shading unit count (3072 versus 2560) and greater FP32 output (6.832 TFLOPS versus 5.704 TFLOPS) benefit workloads like scientific simulation, large-scale rendering, and batch inference. The 288.4 GB/s memory bandwidth supports large datasets that would bottleneck the P4’s 192.3 GB/s interface. The 12 GB frame buffer allows larger models or higher resolution textures to reside in memory without swapping.

The Tesla P4 wins in power-constrained and space-constrained environments. Its 75 W TDP is 175 W lower than the M40’s, and it requires no external power connector. The single-slot design at 168 mm length allows for dense server configurations where the M40’s 267 mm dual-slot footprint would not fit. The suggested 250 W power supply is less than half the M40’s 600 W requirement, making the P4 viable in systems with limited power delivery.

For multi-GPU arrays, the P4’s lower power draw enables more cards per server. For edge inference or real-time processing where latency matters more than raw throughput, the P4’s 5.704 TFLOPS is sufficient for many tasks. The P4’s FP16 support, though limited, provides some capability for mixed-precision workloads that the M40 entirely lacks.

The performance percentile gap — 83rd for the M40 versus 81st for the P4 — is small in absolute terms, but the benchmark deltas are consistent. The M40 is the compute winner. The P4 is the efficiency and density winner. Choose the M40 for maximum performance per card; choose the P4 for maximum performance per watt and per rack unit.

DETAILED SPECIFICATIONS

SPECIFICATION
Tesla M40
Tesla P4
Core Specs
Shading Units
3,072
2,560 -16.7%
Shaders
3,072
2,560 -16.7%
TMUs
192
160 -16.7%
ROPs
96
64 -33.3%
SM Count
20
Clocks
Base Clock
948 MHz
886 MHz
Boost Clock
1112 MHz
1114 MHz
Memory Clock
1502 MHz 6 Gbps effective
1502 MHz 6 Gbps effective
Memory
Memory Size
12 GB
8 GB
VRAM (MB)
12,288
8,192 -33.3%
Memory Type
GDDR5
GDDR5
Memory Bus
384 bit
256 bit
Bandwidth
288.4 GB/s
192.3 GB/s
Cache
L1 Cache
48 KB (per SMM)
48 KB (per SM)
L2 Cache
3 MB
2 MB
Performance
Pixel Rate
106.8 GPixel/s
71.30 GPixel/s
Texture Rate
213.5 GTexel/s
178.2 GTexel/s
FP32 (TFLOPS)
6.832 TFLOPS
5.704 TFLOPS
FP64 (TFLOPS)
213.5 GFLOPS (1:32)
178.2 GFLOPS (1:32)
FP16 (TFLOPS)
89.12 GFLOPS (1:64)
Power
TDP
250 W
75 W
TDP (W)
250
75 -70.0%
Suggested PSU
600 W
250 W
Power Connectors
8-pin EPS
None
Architecture
Architecture
Maxwell 2.0
Pascal
GPU Name
GM200
GP104
Generation
Tesla Maxwell (Mxx)
Tesla Pascal (Pxx)
Process Size
28 nm
16 nm
Transistors
8,000 million
7,200 million
Die Size
601 mm²
314 mm²
Foundry
TSMC
TSMC
Density
13.3M / mm²
22.9M / mm²
API Support
DirectX
12 (12_1)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
5.2
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Single-slot
Length
267 mm 10.5 inches
168 mm 6.6 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 3.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Tesla Kepler
Tesla Maxwell
Successor
Tesla Pascal
Tesla Volta
View Tesla M40 Details View Tesla P4 Details