NVIDIA Tesla M40 24 GB vs NVIDIA Tesla P4 Comparison

NVIDIA
GEFORCE

NVIDIA Tesla M40 24 GB

CORE STATE GM200
VRAM 24 GB
CLOCK SPEED 1112 MHz
TDP 250 W
BUS WIDTH 384 bit
ARCHITECTURE Maxwell 2.0
nm
PROCESS 28 nm
LAUNCH DATE 2015
VS
NVIDIA
GEFORCE

Tesla P4

CORE STATE GP104
VRAM 8 GB
CLOCK SPEED 1114 MHz
TDP 75 W
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

geekbench_opencl
37,439
34,947
geekbench_vulkan
45,975
40,309

Analysis: NVIDIA Tesla M40 24 GB vs NVIDIA Tesla P4

NVIDIA Tesla M40 24 GB and NVIDIA Tesla P4 represent two distinct generations of datacenter compute cards, with the M40 built on Maxwell 2.0 and the P4 on Pascal. In direct head-to-head benchmarks, the M40 24 GB wins both recorded tests: Geekbench OpenCL (37439 vs 34947, a 7.1% delta) and Geekbench Vulkan (45975 vs 40309, a 14.1% delta). The data shows a clear overall performance advantage for the M40, which also holds a better percentile rank across all GPUs (83rd vs 81st) and a higher average benchmark score (41707 vs 37628). However, the P4 counters with a dramatically lower 75 W TDP versus 250 W, a smaller physical footprint, and a more modern architecture, making the comparison more nuanced than raw score alone.

Head-to-Head Benchmarks

The two recorded benchmark tests both favor the NVIDIA Tesla M40 24 GB. In Geekbench OpenCL, the M40 scores 37439 against the P4’s 34947, yielding a 7.1% advantage. This margin is consistent with the M40’s higher shading unit count (3072 vs 2560) and greater memory bandwidth (288.4 GB/s vs 192.3 GB/s). The Vulkan test shows a larger gap: the M40 reaches 45975 while the P4 manages 40309, a 14.1% delta. The wider Vulkan spread suggests the M40’s raw compute resources—particularly its 96 ROPs versus the P4’s 64—scale better under API-level workloads that emphasize pixel throughput.

When placed against their nearest rivals, both cards occupy similar performance tiers. The M40’s average score of 41707 is 1.3% above the GeForce RTX 3080 Ti (41187) and 2% above the AMD Radeon Pro 5300 (40870), while trailing the AMD Radeon RX 7650 GRE by 2.4% (42723). The P4’s average of 37628 sits nearly level with the RTX 4070 (37648, -0.1%) and RX Vega 56 (37507, +0.3%), and it edges out the Radeon PRO W6400 by 1.3% (37157). Notably, the M40’s nearest rival is the non-24GB Tesla M40, which scores 41897—just 0.5% higher—indicating the extra memory does not materially change compute performance in these tests.

The 14.1% Vulkan win for the M40 is its largest margin over the P4, while the 7.1% OpenCL win is more modest. Both deltas exceed the typical differences seen among the cards’ nearest rivals, suggesting the architectural gap between Maxwell and Pascal is smaller than the resource gap (3072 vs 2560 shading units, 192 vs 160 TMUs). The P4’s higher boost clock (1114 MHz vs 1112 MHz) does not compensate for its fewer execution units, and its lower pixel rate (71.30 GPixel/s vs 106.8 GPixel/s) directly explains the Vulkan shortfall.

Where Each One Wins

The NVIDIA Tesla M40 24 GB wins on raw compute throughput. Its fp32 performance of 6.832 TFLOPS exceeds the P4’s 5.704 TFLOPS by roughly 20%, a gap mirrored in texture rate (213.5 GTexel/s vs 178.2 GTexel/s) and pixel rate (106.8 GPixel/s vs 71.30 GPixel/s). For workloads that saturate shading units—such as dense matrix operations, raycasting without dedicated RT cores, or high-resolution image processing—the M40’s extra 512 shading units and 32 TMUs provide a direct advantage. The 24 GB memory capacity also gives it a clear edge for datasets that exceed 8 GB, allowing larger batch sizes in inference or rendering scenes that would spill to system memory on the P4.

The NVIDIA Tesla P4 wins on efficiency and physical integration. Its 75 W TDP is one-third of the M40’s 250 W, and it requires no power connectors, drawing all power from the PCIe slot. The single-slot, 168 mm (6.6 inches) design contrasts sharply with the M40’s dual-slot, 267 mm (10.5 inches) footprint, making the P4 suitable for dense servers where board space and airflow are constrained. The P4’s suggested PSU of 250 W versus the M40’s 600 W further underscores its lower system-level demands. For edge inference or always-on analytics where power density matters more than peak throughput, the P4 is the logical pick despite its lower scores.

The benchmark data does not show any test where the P4 outperforms the M40, so the use-case split is purely architectural. The M40 is the choice for maximum compute per card, while the P4 is the choice for minimum power and space per card. Neither has display outputs, so both are strictly for compute or virtualized workloads. The P4’s Pascal generation does bring a smaller die (314 mm² vs 601 mm²) and higher transistor density (22.9M / mm² vs 13.3M / mm²), which explains its efficiency but does not translate into a performance win in these benchmarks.

Architecture Differences

The M40 uses the GM200 chip on Maxwell 2.0 architecture, fabricated on TSMC’s 28 nm process. It packs 8,000 million transistors across a 601 mm² die, yielding a transistor density of 13.3M / mm². The P4 uses the GP104 chip on Pascal architecture, built on TSMC’s 16 nm process, with 7,200 million transistors on a 314 mm² die—a density of 22.9M / mm². The node shrink from 28 nm to 16 nm allows the P4 to achieve higher density despite fewer total transistors, which is the primary driver of its 75 W TDP versus 250 W.

Both chips lack RT cores and tensor cores, so neither supports hardware-accelerated ray tracing or tensor operations. The M40’s Maxwell 2.0 generation is listed as Tesla Maxwell (Mxx), with a predecessor of Tesla Kepler and successor of Tesla Pascal. The P4’s Pascal generation is Tesla Pascal (Pxx), with a predecessor of Tesla Maxwell and successor of Tesla Volta. This lineage places the P4 one generation ahead, but the M40 compensates with a larger silicon budget.

The M40 has more of every execution resource: 3072 shading units, 192 TMUs, and 96 ROPs, versus the P4’s 2560 shading units, 160 TMUs, and 64 ROPs. Clock speeds are nearly identical—base 948 MHz vs 886 MHz, boost 1112 MHz vs 1114 MHz—so the M40’s performance lead comes almost entirely from parallel width. Memory clocks are the same at 1502 MHz (6 Gbps effective), but the M40’s 384-bit bus versus the P4’s 256-bit bus gives it 288.4 GB/s bandwidth versus 192.3 GB/s. The M40’s fp32 output of 6.832 TFLOPS exceeds the P4’s 5.704 TFLOPS, while the P4 lists a token fp16 rate of 89.12 GFLOPS (1:64), indicating negligible half-precision support; the M40 has no fp16 listing.

Specification Differences

The two cards differ in nearly every physical and memory specification. The M40 has 24 GB GDDR5 on a 384-bit bus with 288.4 GB/s bandwidth; the P4 has 8 GB GDDR5 on a 256-bit bus with 192.3 GB/s bandwidth. Shading units are 3072 vs 2560, TMUs are 192 vs 160, and ROPs are 96 vs 64. Pixel rate is 106.8 GPixel/s vs 71.30 GPixel/s, and texture rate is 213.5 GTexel/s vs 178.2 GTexel/s. FP32 compute is 6.832 TFLOPS vs 5.704 TFLOPS, with the P4 adding an fp16 figure of 89.12 GFLOPS (1:64) that the M40 lacks.

Power and physical specs diverge sharply. The M40 consumes 250 W with an 8-pin EPS connector and a suggested 600 W PSU, in a dual-slot form factor measuring 267 mm (10.5 inches). The P4 consumes 75 W with no power connectors and a suggested 250 W PSU, in a single-slot form factor measuring 168 mm (6.6 inches). Both use PCIe 3.0 x16 and have no display outputs. Transistor count is 8,000 million vs 7,200 million, die size is 601 mm² vs 314 mm², and density is 13.3M / mm² vs 22.9M / mm². Process node is 28 nm vs 16 nm. Both support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. Production status for both is end-of-life, with release dates of November 2015 for the M40 and September 2016 for the P4.

FAQ

Q: Which card has a higher average benchmark score?

A: The NVIDIA Tesla M40 24 GB has an average benchmark score of 41707, while the NVIDIA Tesla P4 scores 37628, a difference of about 10.8% in favor of the M40.

Q: What is the largest performance delta between the two in a single test?

A: The largest delta is in Geekbench Vulkan, where the M40 scores 45975 versus the P4’s 40309, a 14.1% advantage.

Q: Does the P4 win any of the recorded head-to-head benchmarks?

A: No. The P4 wins zero tests, while the M40 wins both Geekbench OpenCL and Geekbench Vulkan.

Q: How do the TDPs compare, and what does that imply for system requirements?

A: The M40 has a 250 W TDP with a suggested 600 W PSU and an 8-pin EPS connector, while the P4 has a 75 W TDP with a suggested 250 W PSU and no power connectors. The P4 requires far less power infrastructure.

Q: Are there any differences in API support?

A: No. Both cards support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4.

Q: Which card has a higher transistor density?

A: The Pascal-based P4 has a transistor density of 22.9M / mm², versus 13.3M / mm² for the Maxwell-based M40, due to the 16 nm process versus 28 nm.

DETAILED SPECIFICATIONS

SPECIFICATION
Tesla M40 24 GB
Tesla P4
Core Specs
Shading Units
3,072
2,560 -16.7%
Shaders
3,072
2,560 -16.7%
TMUs
192
160 -16.7%
ROPs
96
64 -33.3%
SM Count
20
Clocks
Base Clock
948 MHz
886 MHz
Boost Clock
1112 MHz
1114 MHz
Memory Clock
1502 MHz 6 Gbps effective
1502 MHz 6 Gbps effective
Memory
Memory Size
24 GB
8 GB
VRAM (MB)
24,576
8,192 -66.7%
Memory Type
GDDR5
GDDR5
Memory Bus
384 bit
256 bit
Bandwidth
288.4 GB/s
192.3 GB/s
Cache
L1 Cache
48 KB (per SMM)
48 KB (per SM)
L2 Cache
3 MB
2 MB
Performance
Pixel Rate
106.8 GPixel/s
71.30 GPixel/s
Texture Rate
213.5 GTexel/s
178.2 GTexel/s
FP32 (TFLOPS)
6.832 TFLOPS
5.704 TFLOPS
FP64 (TFLOPS)
213.5 GFLOPS (1:32)
178.2 GFLOPS (1:32)
FP16 (TFLOPS)
89.12 GFLOPS (1:64)
Power
TDP
250 W
75 W
TDP (W)
250
75 -70.0%
Suggested PSU
600 W
250 W
Power Connectors
8-pin EPS
None
Architecture
Architecture
Maxwell 2.0
Pascal
GPU Name
GM200
GP104
Generation
Tesla Maxwell (Mxx)
Tesla Pascal (Pxx)
Process Size
28 nm
16 nm
Transistors
8,000 million
7,200 million
Die Size
601 mm²
314 mm²
Foundry
TSMC
TSMC
Density
13.3M / mm²
22.9M / mm²
API Support
DirectX
12 (12_1)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
5.2
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Single-slot
Length
267 mm 10.5 inches
168 mm 6.6 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 3.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Tesla Kepler
Tesla Maxwell
Successor
Tesla Pascal
Tesla Volta
View Tesla M40 24 GB Details View Tesla P4 Details