NVIDIA P106-100 vs NVIDIA Quadro M4000M Comparison
NVIDIA P106-100
Quadro M4000M
PERFORMANCE BENCHMARKS
Analysis: NVIDIA P106-100 vs NVIDIA Quadro M4000M
FAQ
Q: How does the NVIDIA P106-100 compare to the NVIDIA Quadro M4000M in overall benchmark scores?
A: The P106-100 has an average benchmark score of 23,249, placing it in the 68th percentile of all GPUs. The Quadro M4000M has an average score of 20,480, placing it in the 65th percentile. The P106-100 leads in both shared head-to-head tests.
Q: What are the head-to-head benchmark results between the two cards?
A: In Geekbench OpenCL, the P106-100 scores 35,951 versus 19,989 for the Quadro M4000M, a 79.9% advantage. In Geekbench Vulkan, the P106-100 scores 32,897 versus 20,971, a 56.9% advantage. The P106-100 wins both comparisons.
Q: Which card has a newer manufacturing process?
A: The P106-100 is built on a 16 nm process at TSMC, while the Quadro M4000M uses a 28 nm process at TSMC. The P106-100 also has a higher transistor density at 22.0M per mm² compared to 13.1M per mm².
Q: How does the memory configuration differ?
A: The P106-100 has 6 GB of GDDR5 memory on a 192-bit bus with 192.2 GB/s bandwidth. The Quadro M4000M has 4 GB of GDDR5 on a wider 256-bit bus but with 160.4 GB/s bandwidth. The P106-100 has more capacity and higher bandwidth despite the narrower bus.
Q: What are the display output capabilities of each card?
A: The P106-100 has no display outputs, as it is designed for mining workloads. The Quadro M4000M's display outputs are portable device dependent, meaning they vary based on the laptop or mobile workstation implementation.
Q: Which card has a higher peak FP32 performance?
A: The P106-100 delivers 4.375 TFLOPS of FP32 compute, while the Quadro M4000M delivers 2.593 TFLOPS. The P106-100 has a roughly 69% higher FP32 throughput based on these figures.
Where Each One Wins
The recorded data shows a clear split between the two cards, though the P106-100 dominates in every measured benchmark. The P106-100 wins both head-to-head tests, giving it a 2-0 record in direct comparisons. Its wins are substantial: a 79.9% advantage in Geekbench OpenCL and a 56.9% advantage in Geekbench Vulkan. This makes the P106-100 the stronger choice for raw compute throughput in OpenCL workloads and for Vulkan graphics workloads, even though it lacks any display outputs.
The Quadro M4000M does not win any benchmark in the database. However, its strengths lie elsewhere. As a mobile workstation GPU in the Quadro Maxwell-M series, it is designed for portable devices. It uses an MXM module form factor, has no power connectors of its own, and consumes 100 W TDP, which is 20 W less than the P106-100. The Quadro M4000M also uses a PCIe 3.0 x16 bus interface, whereas the P106-100 is limited to PCIe 1.0 x16. For scenarios where system integration, power efficiency, and a standard PCIe interface matter more than raw compute, the Quadro M4000M has structural advantages. In terms of pure benchmark scores, the P106-100 is ahead in all recorded tests.
Architecture Differences
The two GPUs come from different NVIDIA architectures. The P106-100 is based on the GP106 chip using the Pascal architecture, released in the Mining GPUs generation. The Quadro M4000M is based on the GM204 chip using the Maxwell 2.0 architecture, released in the Quadro Maxwell-M generation. This generational gap is significant: Pascal is the newer architecture, and the data reflects that in both compute and graphics performance.
The manufacturing process differs substantially. The P106-100 uses a 16 nm TSMC process, while the Quadro M4000M uses a 28 nm TSMC process. The P106-100 packs 4,400 million transistors on a 200 mm² die, yielding a transistor density of 22.0M per mm². The Quadro M4000M has 5,200 million transistors on a much larger 398 mm² die, resulting in a lower density of 13.1M per mm². The smaller, denser Pascal chip achieves higher performance despite having fewer total transistors.
Both chips use the same fundamental compute resources: 1280 shading units and 80 texture mapping units. They differ in render output units, with the P106-100 having 48 ROPs and the Quadro M4000M having 64 ROPs. The Quadro M4000M's advantage in ROP count does not translate into benchmark wins, as the P106-100's higher clock speeds and newer architecture compensate.
Neither card has ray tracing cores or tensor cores. Both support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The P106-100 has no display outputs, reflecting its mining-oriented design, while the Quadro M4000M is portable device dependent for display output, reflecting its mobile workstation role.
Specification Differences
The two cards differ across nearly every specification category. The P106-100 uses the GP106 chip with Pascal architecture, while the Quadro M4000M uses the GM204 chip with Maxwell 2.0 architecture. The process node is 16 nm for the P106-100 versus 28 nm for the Quadro M4000M. Transistor count is 4,400 million versus 5,200 million, and die size is 200 mm² versus 398 mm². Transistor density is 22.0M per mm² versus 13.1M per mm².
Clock speeds differ significantly. The P106-100 has a base clock of 1506 MHz and a boost clock of 1709 MHz. The Quadro M4000M has a base clock of 975 MHz and a boost clock of 1013 MHz. Memory clocks also differ: the P106-100 runs at 2002 MHz (8 Gbps effective), while the Quadro M4000M runs at 1253 MHz (5 Gbps effective).
Memory capacity is 6 GB for the P106-100 versus 4 GB for the Quadro M4000M. The memory bus is 192-bit versus 256-bit. Memory bandwidth is 192.2 GB/s versus 160.4 GB/s. The P106-100 has more bandwidth despite the narrower bus due to higher memory clock speeds.
Compute resources are identical in shading units (1280) and TMUs (80), but ROPs differ: 48 for the P106-100 versus 64 for the Quadro M4000M. Pixel rate is 82.03 GPixel/s versus 64.83 GPixel/s. Texture rate is 136.7 GTexel/s versus 81.04 GTexel/s. FP32 performance is 4.375 TFLOPS versus 2.593 TFLOPS. The P106-100 also lists FP16 performance at 68.36 GFLOPS (1:64), while the Quadro M4000M has no recorded FP16 figure.
Power and form factor differ. The P106-100 has a TDP of 120 W, uses a dual-slot design, requires a 1x 6-pin power connector, and suggests a 300 W PSU. It is 250 mm long (9.8 inches). The Quadro M4000M has a TDP of 100 W, uses an MXM module form factor, has no power connectors, and its dimensions are not recorded. The bus interface is PCIe 1.0 x16 for the P106-100 versus PCIe 3.0 x16 for the Quadro M4000M. The P106-100 has no display outputs; the Quadro M4000M's outputs are portable device dependent.
Release dates place the Quadro M4000M first, with a release date of 2015-08-17, while the P106-100 followed on 2017-06-18. The Quadro M4000M has a recorded predecessor (Quadro Kepler-M) and successor (Quadro Pascal-M), while the P106-100 has neither. Both cards are end-of-life in production status.
Head-to-Head Benchmarks
The database contains two direct head-to-head benchmark comparisons between the P106-100 and the Quadro M4000M, and the P106-100 wins both decisively.
The first and largest gap is in Geekbench OpenCL. The P106-100 scores 35,951, while the Quadro M4000M scores 19,989. This is a 79.9% advantage for the P106-100. The margin is explained by the combination of higher clock speeds (1506 MHz base and 1709 MHz boost versus 975 MHz base and 1013 MHz boost), higher FP32 throughput (4.375 TFLOPS versus 2.593 TFLOPS), and higher memory bandwidth (192.2 GB/s versus 160.4 GB/s). The OpenCL workload appears to scale strongly with raw compute and memory throughput, both of which favor the P106-100.
The second head-to-head test is Geekbench Vulkan. The P106-100 scores 32,897, and the Quadro M4000M scores 20,971. The delta is 56.9% in favor of the P106-100. This is a smaller margin than the OpenCL result, but still a substantial win. The Vulkan workload may be less sensitive to raw FP32 compute and more dependent on other factors such as driver efficiency and memory subsystem behavior. The Quadro M4000M's wider 256-bit memory bus and higher ROP count (64 versus 48) may narrow the gap in graphics-oriented tasks, but they are not enough to overcome the P106-100's architectural and clock speed advantages.
Context from the nearest rivals helps interpret these scores. The P106-100's average benchmark score of 23,249 places it just 0.1% behind the AMD Radeon RX 6600M (23,273) and 0.1% behind the AMD Radeon R9 M290X (23,276). It is effectively tied with the AMD Radeon Pro Vega 16 at 23,250. The Quadro M4000M's average score of 20,480 is 0.3% behind the NVIDIA GeForce RTX 3070 Mobile (20,534), 0.4% behind the Intel Arc B570 (20,556), and 0.5% behind the Intel Arc A750 (20,582). The P106-100's nearest rivals are all within 0.3% of its score, while the Quadro M4000M's nearest rivals are all within 0.9%, indicating that the P106-100 sits in a tighter competitive cluster at a higher absolute performance level.
Both cards support the same API feature set: DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The performance gap in the recorded tests is therefore not due to missing API support on either side. The P106-100's 16 nm process, higher clocks, and larger memory bandwidth give it a clear performance edge. The Quadro M4000M's only structural advantages are its lower TDP (100 W versus 120 W), its MXM form factor for mobile integration, and its PCIe 3.0 x16 interface versus the P106-100's PCIe 1.0 x16. For raw performance as measured in these benchmarks, the P106-100 is the stronger GPU by a wide margin.