NVIDIA P104-100 vs NVIDIA Tesla P4 Comparison
NVIDIA P104-100
Tesla P4
PERFORMANCE BENCHMARKS
Analysis: NVIDIA P104-100 vs NVIDIA Tesla P4
Head-to-Head Benchmarks
The recorded data shows a clear, consistent winner across the two shared benchmark tests. The NVIDIA P104-100 outperforms the Tesla P4 in every head-to-head comparison available, though the margin varies significantly depending on the workload.
In the Geekbench OpenCL test, the P104-100 scores 52,368 against the Tesla P4's 34,947. That is a 33.3% advantage for the P104-100, a substantial gap that indicates a major difference in raw compute throughput for general-purpose GPU tasks. This is the largest single victory for either card in the comparison.
The Vulkan test tells a similar story but with a narrower margin. The P104-100 scores 45,165, while the Tesla P4 manages 40,309. The delta here is 10.8% in favor of the P104-100. While still a decisive win, the smaller gap suggests that the Tesla P4 is relatively more competitive in graphics-oriented API workloads than in pure compute tasks.
The overall benchmark average tells a more nuanced story. The Tesla P4 records an average benchmark score of 37,628 across its tested workloads, which places it at the 81st percentile of all GPUs in the database. The P104-100, despite winning both direct head-to-head tests, has a lower average score of 32,982 and sits at the 77th percentile. This discrepancy is due to the fact that the P104-100's average includes a third benchmark, the 3DMark Steel Nomad DX12 test, where it scores 1,413. That result drags down its average considerably, indicating that the P104-100 is not well-suited to modern DX12 rasterization workloads.
Looking at the nearest rivals in the database, the Tesla P4's average score of 37,628 is nearly identical to the NVIDIA GeForce RTX 4070 at 37,648, a delta of only -0.1%. It also sits just 0.3% above the AMD Radeon RX Vega 56 and 1.3% above the AMD Radeon PRO W6400. The P104-100, by contrast, sits in a lower performance tier. Its average of 32,982 is 0.4% above the NVIDIA T600 Mobile, 0.5% below the NVIDIA T550 Mobile, and 0.6% below the NVIDIA GeForce RTX 3050 Mobile. These figures frame the two cards as belonging to different performance classes, despite the P104-100's wins in the direct comparisons.
Architecture Differences
Both cards are built on the same fundamental architecture. They share the GP104 chip, the Pascal architecture, a 16 nm process node from TSMC, 7,200 million transistors, and a die size of 314 mm². The transistor density is identical at 22.9M per mm². This means the underlying silicon is the same, but NVIDIA configured it very differently for two distinct market segments.
The Tesla P4 is a compute-focused professional card from the Tesla Pascal generation. It carries the full complement of 2,560 shading units, 160 texture mapping units, and 64 raster operations units. The P104-100, from the Mining GPUs generation, has fewer active compute units: 1,920 shading units and 120 TMUs, though it retains the same 64 ROPs. This reduction in shader and texture hardware is a key reason why the P104-100 cannot match the Tesla P4 in every scenario, despite its higher clock speeds.
Memory configuration is another major divergence. The Tesla P4 uses 8 GB of GDDR5 memory on a 256-bit bus, yielding 192.3 GB/s of bandwidth. The P104-100 uses 4 GB of GDDR5X on the same 256-bit bus, but with a much higher memory clock of 1,251 MHz (10 Gbps effective), producing 320.3 GB/s of bandwidth. The P104-100 has 67% more memory bandwidth, which explains its dominance in OpenCL compute tasks that are bandwidth-sensitive.
Clock speeds are also substantially different. The Tesla P4 runs at a base clock of 886 MHz and a boost clock of 1,114 MHz. The P104-100 runs at a base of 1,607 MHz and a boost of 1,733 MHz. That is a significant clock advantage for the P104-100, roughly 55% higher at base and 55% higher at boost. This clock difference, combined with the faster memory, drives the P104-100's superior raw throughput numbers: 6.655 TFLOPS FP32 versus 5.704 TFLOPS for the Tesla P4, and 208.0 GTexel/s versus 178.2 GTexel/s in texture rate. The pixel rate also favors the P104-100 at 110.9 GPixel/s versus 71.30 GPixel/s.
The cards also differ in their physical and electrical design. The Tesla P4 is a single-slot card, 168 mm long, with no power connectors, drawing just 75 W and requiring a 250 W power supply. The P104-100 is a dual-slot card, 267 mm long, with a single 8-pin power connector, and its TDP is not recorded. The suggested power supply is 200 W. The P104-100 also uses a PCIe 1.0 x4 interface, while the Tesla P4 uses PCIe 3.0 x16. This is a critical limitation for the P104-100 in systems that rely on PCIe bandwidth for data transfer, though it matters less for compute workloads that keep data local to the GPU.
Where Each One Wins
The P104-100 wins decisively in raw compute throughput. Its higher clocks and faster GDDR5X memory give it a 33.3% lead in OpenCL performance, which reflects its strength in general-purpose GPU compute, number crunching, and bandwidth-hungry tasks. The 320.3 GB/s memory bandwidth is the standout specification here, and the data shows it translates directly into a large benchmark advantage. If your workload is FP32 compute, the P104-100 is the faster card by a clear margin.
The Tesla P4 wins in overall consistency and in the areas that matter for a professional compute card. Its average benchmark score of 37,628 is higher than the P104-100's 32,982, and it sits at the 81st percentile versus the 77th. The Tesla P4 also has double the memory capacity at 8 GB versus 4 GB, which is critical for workloads that need to hold larger datasets on the GPU. Its 75 W TDP with no power connectors makes it far easier to integrate into dense servers or workstations where power and space are constrained. The PCIe 3.0 x16 interface also provides much better host connectivity than the P104-100's PCIe 1.0 x4.
The Vulkan gap is narrower at 10.8%, indicating that the Tesla P4's extra shading units (2,560 versus 1,920) partially compensate for the P104-100's clock advantage in graphics-oriented workloads. If you are running Vulkan-based applications, the Tesla P4 is closer in performance than the raw clock numbers would suggest, though it still loses.
The P104-100's 3DMark Steel Nomad DX12 score of 1,413 is the only data point for that test, and it is far below what its compute scores would imply. This suggests the card is not a good choice for modern DX12 gaming or rasterization, and its mining-oriented design shows in that result.
Specification Differences
The two cards differ on nearly every specification field except the ones tied to the underlying silicon. Here is the complete list of differences:
- Generation: Tesla P4 is "Tesla Pascal (Pxx)"; P104-100 is "Mining GPUs"
- Base clock: 886 MHz vs 1,607 MHz
- Boost clock: 1,114 MHz vs 1,733 MHz
- Memory clock: 1,502 MHz (6 Gbps effective) vs 1,251 MHz (10 Gbps effective)
- Memory size: 8 GB vs 4 GB
- Memory type: GDDR5 vs GDDR5X
- Memory bandwidth: 192.3 GB/s vs 320.3 GB/s
- Shading units: 2,560 vs 1,920
- TMUs: 160 vs 120
- Pixel rate: 71.30 GPixel/s vs 110.9 GPixel/s
- Texture rate: 178.2 GTexel/s vs 208.0 GTexel/s
- FP32: 5.704 TFLOPS vs 6.655 TFLOPS
- FP16: 89.12 GFLOPS (1:64) vs 104.0 GFLOPS (1:64)
- TDP: 75 W vs not recorded
- Slot width: Single-slot vs Dual-slot
- Power connectors: None vs 1x 8-pin
- Suggested PSU: 250 W vs 200 W
- Bus interface: PCIe 3.0 x16 vs PCIe 1.0 x4
- Dimensions: 168 mm vs 267 mm
- Release date: 2016-09-12 vs 2017-12-11
- Predecessor: Tesla Maxwell vs not recorded
- Successor: Tesla Volta vs not recorded
Identical between the two: GP104 chip, Pascal architecture, 16 nm process, TSMC foundry, 7,200 million transistors, 314 mm² die size, 22.9M / mm² transistor density, 64 ROPs, no RT cores, no tensor cores, DirectX 12 (12_1), OpenGL 4.6, Vulkan 1.4, no display outputs, and end-of-life production status.
FAQ
Q: Which card is faster in compute workloads?
A: The P104-100. It wins the Geekbench OpenCL test by 33.3% (52,368 versus 34,947) and the Vulkan test by 10.8% (45,165 versus 40,309). Its FP32 throughput is also higher at 6.655 TFLOPS versus 5.704 TFLOPS.
Q: Why does the Tesla P4 have a higher average benchmark score if it loses both head-to-head tests?
A: The P104-100's average of 32,982 includes a 3DMark Steel Nomad DX12 score of 1,413, which is far below its other results. The Tesla P4's average of 37,628 comes from only the OpenCL and Vulkan tests, where it scores 34,947 and 40,309 respectively.
Q: Which card has more memory and why does that matter?
A: The Tesla P4 has 8 GB of GDDR5, while the P104-100 has 4 GB of GDDR5X. The larger capacity on the Tesla P4 allows it to hold larger datasets, but the P104-100's GDDR5X provides much higher bandwidth at 320.3 GB/s versus 192.3 GB/s.
Q: Are these cards suitable for gaming?
A: The data is not favorable for either, but the P104-100 is clearly worse. Its 3DMark Steel Nomad DX12 score of 1,413 is very low, and neither card has display outputs, so they cannot drive a monitor directly. The Tesla P4 has no DX12 benchmark recorded.
Q: What are the power and physical requirements of each card?
A: The Tesla P4 is a 75 W single-slot card, 168 mm long, with no power connectors and a 250 W suggested PSU. The P104-100 is a dual-slot card, 267 mm long, requires one 8-pin power connector, and has a 200 W suggested PSU. Its TDP is not recorded.
Q: How do these cards compare to other GPUs in the database?
A: The Tesla P4 sits at the 81st percentile, nearly matching the GeForce RTX 4070 (-0.1%) and beating the RX Vega 56 by 0.3%. The P104-100 sits at the 77th percentile, roughly matching the T600 Mobile (0.4% above) and trailing the RTX 3050 Mobile by 0.6%.
The Verdict
The data points to two very different tools. If your priority is raw compute throughput, memory bandwidth, and you can tolerate a dual-slot card with an 8-pin connector, the P104-100 is the faster option. It wins both shared benchmarks outright, with a massive 33.3% lead in OpenCL and a solid 10.8% lead in Vulkan. Its 320.3 GB/s memory bandwidth and 6.655 TFLOPS FP32 make it a strong choice for compute-heavy tasks that fit within its 4 GB memory limit.
However, the Tesla P4 is the more balanced and practical professional card. It has a higher average benchmark score (37,628 versus 32,982), a higher percentile ranking (81st versus 77th), double the memory (8 GB versus 4 GB), a much friendlier physical profile (single-slot, 75 W, no power connectors), and a proper PCIe 3.0 x16 interface. Its lower clock speeds and slower memory mean it loses the compute race, but its 2,560 shading units and 64 ROPs keep it competitive in Vulkan, and its larger memory makes it viable for bigger workloads.
The P104-100 is a mining-oriented card, and the data reflects that. It excels at compute, struggles with DX12, and has a crippled PCIe interface. The Tesla P4 is a proper data center card designed for sustained professional use. Choose the P104-100 if you need maximum FP32 performance and bandwidth for a specific compute task and do not care about memory capacity, PCIe bandwidth, or power efficiency. Choose the Tesla P4 if you need a well-rounded, efficient, and easier-to-integrate card that still delivers respectable performance and much more memory. For most use cases that require a professional GPU, the Tesla P4 is the safer recommendation based on the recorded data.