NVIDIA P104-100 vs NVIDIA Quadro M5000 Comparison

NVIDIA
GEFORCE

NVIDIA P104-100

CORE STATE GP104
VRAM 4 GB
CLOCK SPEED 1733 MHz
TDP
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2017
VS
NVIDIA
GEFORCE

Quadro M5000

CORE STATE GM204
VRAM 8 GB
CLOCK SPEED 1038 MHz
TDP 150 W
BUS WIDTH 256 bit
ARCHITECTURE Maxwell 2.0
nm
PROCESS 28 nm
LAUNCH DATE 2015

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
1,413
N/A
geekbench_opencl
52,368
29,481
geekbench_vulkan
45,165
32,931

Analysis: NVIDIA P104-100 vs NVIDIA Quadro M5000

The NVIDIA P104-100 and NVIDIA Quadro M5000 represent two distinct approaches to GPU design from the same manufacturer, separated by two generations of architecture and targeting fundamentally different workloads. The P104-100 is a Pascal-based mining card with no display outputs, stripped of video connectivity to maximize compute throughput, while the Quadro M5000 is a Maxwell 2.0 professional workstation card with full display support and a larger memory pool. Benchmark data shows a clear performance hierarchy between the two, but the choice is not simply about raw speed — it depends on what the user needs the GPU to do.

Head-to-Head Benchmarks

The P104-100 dominates the Quadro M5000 in every benchmark test where both were evaluated. In Geekbench OpenCL, the P104-100 scores 52,368 against the Quadro M5000’s 29,481, a decisive 77.6% advantage. This is not a marginal gap; it is a crushing lead that reflects the fundamental architectural and clock speed differences between the two cards. The P104-100’s base clock of 1607 MHz and boost clock of 1733 MHz are nearly double the Quadro M5000’s 861 MHz base and 1038 MHz boost, and that clock advantage translates directly into compute throughput.

The Vulkan benchmark tells a similar story, though with a narrower margin. The P104-100 scores 45,165 versus the Quadro M5000’s 32,931, a 37.2% lead. The smaller delta in Vulkan compared to OpenCL suggests that the Maxwell architecture on the Quadro M5000 holds up relatively better in API-bound workloads, but it still falls significantly short. Across the two shared benchmarks, the P104-100 wins both, giving it a 2-0 head-to-head record.

Looking at absolute performance percentiles, the P104-100 sits at the 77th percentile of all GPUs, while the Quadro M5000 sits at the 76th. The P104-100’s average benchmark score across all tests is 32,982, and its nearest rivals include the NVIDIA T600 Mobile at 32,849 (0.4% slower), the NVIDIA T550 Mobile at 33,161 (0.5% faster), and the NVIDIA GeForce RTX 3050 Mobile at 33,170 (0.6% faster). The Quadro M5000’s average score is 31,206, placing it near the NVIDIA GRID M60-1Q at 31,220 (0% delta), the NVIDIA GeForce RTX 4070 Ti SUPER at 31,087 (0.4% faster), and the NVIDIA RTX PRO 4500 Blackwell at 31,532 (1% faster).

The P104-100’s 77.6% advantage in OpenCL is its single largest win, driven by its 6.655 TFLOPS of FP32 performance against the Quadro M5000’s 4.252 TFLOPS — a 56.5% raw compute advantage. The P104-100 also delivers a texture rate of 208.0 GTexel/s versus 132.9 GTexel/s on the Quadro M5000, a 56.5% lead, and a pixel rate of 110.9 GPixel/s versus 66.43 GPixel/s, a 67% lead. These figures align closely with the observed benchmark deltas, confirming that the P104-100’s advantage is consistent across different types of workloads.

The Verdict

The data is unambiguous: the NVIDIA P104-100 outperforms the NVIDIA Quadro M5000 in every measured benchmark. For anyone whose primary criterion is raw compute performance, the P104-100 is the superior choice. It delivers a 77.6% higher OpenCL score and a 37.2% higher Vulkan score, with a higher average benchmark score of 32,982 versus 31,206. The P104-100 also has a higher percentile ranking at 77 versus 76, though the gap there is minimal.

However, the Quadro M5000 is not without its own advantages. It offers 8 GB of GDDR5 memory versus the P104-100’s 4 GB of GDDR5X, which is a critical differentiator for workloads that require large datasets to reside in VRAM. The Quadro M5000 also has display outputs (1x DVI and 4x DisplayPort 1.2), while the P104-100 has no outputs at all — it cannot drive a monitor. The Quadro M5000 also supports PCIe 3.0 x16, whereas the P104-100 is limited to PCIe 1.0 x4, a severe bandwidth constraint for host-device communication.

The verdict depends on the use case. If the workload is compute-bound and does not require display output or large memory capacity, the P104-100 is the clear winner. If the workload requires 8 GB of memory, display connectivity, or benefits from the Quadro M5000’s professional feature set, the Quadro M5000 becomes the more practical choice despite its lower raw performance. The Quadro M5000 also has a lower TDP at 150 W versus the P104-100’s unspecified TDP, and uses a single 6-pin power connector instead of an 8-pin, making it easier to integrate into existing systems with lower power supply requirements — the suggested PSU is 450 W versus 200 W for the P104-100.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA P104-100 has an average benchmark score of 32,982, while the NVIDIA Quadro M5000 has an average score of 31,206, giving the P104-100 a 5.7% lead in aggregate performance.

Q: How much faster is the P104-100 in Geekbench OpenCL?

A: The P104-100 scores 52,368 versus the Quadro M5000’s 29,481, a 77.6% advantage — the largest delta between the two cards in any shared benchmark.

Q: Does the Quadro M5000 have any advantages over the P104-100?

A: Yes, the Quadro M5000 offers 8 GB of GDDR5 memory versus 4 GB of GDDR5X, has display outputs (1x DVI and 4x DisplayPort 1.2) while the P104-100 has none, and supports PCIe 3.0 x16 versus the P104-100’s PCIe 1.0 x4.

Q: What is the memory bandwidth difference between the two cards?

A: The P104-100 has a memory bandwidth of 320.3 GB/s from its GDDR5X memory, while the Quadro M5000 has 211.6 GB/s from its GDDR5 memory — a 51.4% advantage for the P104-100.

Q: Which card has a higher pixel fill rate?

A: The P104-100 achieves 110.9 GPixel/s, compared to the Quadro M5000’s 66.43 GPixel/s, a 67% difference in favor of the P104-100.

Q: Are both cards still in production?

A: No, both are marked as end-of-life. The P104-100 was released in December 2017, while the Quadro M5000 was released in June 2015.

Specification Differences

The two cards differ in nearly every major specification category except for a few shared traits. Both have 64 ROPs, a 256-bit memory bus, dual-slot cooling, a length of 267 mm (10.5 inches), and support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4.

Memory configuration is a key differentiator: the P104-100 has 4 GB of GDDR5X running at 1251 MHz (10 Gbps effective) with 320.3 GB/s bandwidth, while the Quadro M5000 has 8 GB of GDDR5 at 1653 MHz (6.6 Gbps effective) with 211.6 GB/s bandwidth. The P104-100 has fewer shading units (1920 versus 2048) and fewer TMUs (120 versus 128), but compensates with far higher clocks: 1607 MHz base and 1733 MHz boost versus 861 MHz base and 1038 MHz boost.

Power and connectivity also diverge significantly. The P104-100 uses a single 8-pin power connector and has a suggested PSU of 200 W, while the Quadro M5000 uses a single 6-pin connector with a suggested PSU of 450 W and a rated TDP of 150 W. The P104-100 has no display outputs, whereas the Quadro M5000 provides 1x DVI and 4x DisplayPort 1.2. The bus interface differs as well: the P104-100 uses PCIe 1.0 x4, while the Quadro M5000 uses PCIe 3.0 x16. The Quadro M5000 also has a defined height of 111 mm (4.4 inches), while the P104-100’s height is unspecified.

Architecture Differences

The architectural gap between the two cards is substantial. The P104-100 is built on the Pascal architecture using the GP104 chip, fabricated on a 16 nm process at TSMC. The Quadro M5000 uses the Maxwell 2.0 architecture with the GM204 chip, fabricated on a 28 nm process, also at TSMC. This process shrink is a primary driver of the performance difference — the P104-100 packs 7,200 million transistors into a 314 mm² die, achieving a transistor density of 22.9 million per mm², while the Quadro M5000 has 5,200 million transistors on a 398 mm² die, with a density of 13.1 million per mm².

The clock speed disparity is directly attributable to the architectural and process improvements. The P104-100’s base clock of 1607 MHz is 86.6% higher than the Quadro M5000’s 861 MHz, and its boost clock of 1733 MHz is 67% higher than 1038 MHz. This translates into significant compute advantages: the P104-100 delivers 6.655 TFLOPS of FP32 performance versus 4.252 TFLOPS, and its texture rate of 208.0 GTexel/s is 56.5% higher than the Quadro M5000’s 132.9 GTexel/s.

The FP16 situation also differs. The P104-100 has an FP16 rate of 104.0 GFLOPS with a 1:64 ratio, while the Quadro M5000 does not list any FP16 capability. Both cards lack RT cores and tensor cores, but the P104-100’s memory speed advantage — GDDR5X at 10 Gbps effective versus GDDR5 at 6.6 Gbps — gives it a 51.4% bandwidth lead. The P104-100’s generation is listed as "Mining GPUs," which explains its lack of display outputs and its PCIe 1.0 x4 interface; it was designed for compute-only environments. The Quadro M5000, part of the Quadro Maxwell (Mx000) generation, is a professional workstation card with full display support and a predecessor of Quadro Kepler and successor of Quadro Pascal.

DETAILED SPECIFICATIONS

SPECIFICATION
P104-100
Quadro M5000
Core Specs
Shading Units
1,920
2,048 +6.7%
Shaders
1,920
2,048 +6.7%
TMUs
120
128 +6.7%
ROPs
64
64 0.0%
SM Count
15
Clocks
Base Clock
1607 MHz
861 MHz
Boost Clock
1733 MHz
1038 MHz
Memory Clock
1251 MHz 10 Gbps effective
1653 MHz 6.6 Gbps effective
Memory
Memory Size
4 GB
8 GB
VRAM (MB)
4,096
8,192 +100.0%
Memory Type
GDDR5X
GDDR5
Memory Bus
256 bit
256 bit
Bandwidth
320.3 GB/s
211.6 GB/s
Cache
L1 Cache
48 KB (per SM)
48 KB (per SMM)
L2 Cache
2 MB
2 MB
Performance
Pixel Rate
110.9 GPixel/s
66.43 GPixel/s
Texture Rate
208.0 GTexel/s
132.9 GTexel/s
FP32 (TFLOPS)
6.655 TFLOPS
4.252 TFLOPS
FP64 (TFLOPS)
208.0 GFLOPS (1:32)
132.9 GFLOPS (1:32)
FP16 (TFLOPS)
104.0 GFLOPS (1:64)
Power
TDP
150 W
TDP (W)
150
Suggested PSU
200 W
450 W
Power Connectors
1x 8-pin
1x 6-pin
Architecture
Architecture
Pascal
Maxwell 2.0
GPU Name
GP104
GM204
Generation
Mining GPUs
Quadro Maxwell (Mx000)
Process Size
16 nm
28 nm
Transistors
7,200 million
5,200 million
Die Size
314 mm²
398 mm²
Foundry
TSMC
TSMC
Density
22.9M / mm²
13.1M / mm²
API Support
DirectX
12 (12_1)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
6.1
5.2
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
Outputs
No outputs
1x DVI4x DisplayPort 1.2
Bus Interface
PCIe 1.0 x4
PCIe 3.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Quadro Kepler
Successor
Quadro Pascal
View P104-100 Details View Quadro M5000 Details