NVIDIA P104-100 vs NVIDIA Quadro GV100 Comparison
NVIDIA P104-100
Quadro GV100
PERFORMANCE BENCHMARKS
Analysis: NVIDIA P104-100 vs NVIDIA Quadro GV100
The NVIDIA Quadro GV100 and the NVIDIA P104-100 are two very different products that share a manufacturer and little else. The GV100 is a professional workstation card built on the Volta architecture, designed for compute-heavy tasks, while the P104-100 is a mining-specific GPU stripped of display outputs and built on the older Pascal architecture. The data shows a massive performance gulf between them, but the story is more nuanced than one simply being "better" than the other. This analysis breaks down where each card wins, what the architectural differences mean, and who should consider which based on the benchmark results.
Where Each One Wins
The Quadro GV100 is the clear winner in every head-to-head benchmark recorded, but the nature of those wins reveals its intended use case. The GV100 dominates in compute-oriented workloads, specifically in OpenCL and Vulkan APIs. These are general-purpose compute interfaces, not gaming-specific tests, which aligns with the card's professional positioning. The data shows the GV100 winning 2 out of 2 head-to-head comparisons, with the P104-100 failing to secure a single victory.
The P104-100's situation is different. It wins nowhere in the direct comparison, but its benchmark profile suggests it was built for a very specific, narrow task: cryptocurrency mining. With no display outputs and a PCIe 1.0 x4 bus interface, it was never intended for interactive use, gaming, or even professional visualization. Its inclusion in the "Mining GPUs" generation and its lack of any DirectX or OpenGL benchmark scores in the head-to-head data point to a card that was optimized for raw, repetitive compute tasks rather than diverse workloads. The GV100, by contrast, has a full suite of benchmark scores across DirectX 9 through 12, OpenGL, and compute, indicating a general-purpose compute and graphics capability.
The use-case split is stark: the GV100 is a versatile, high-end compute and visualization tool, while the P104-100 is a single-purpose device whose only advantage would be in a scenario where its lower power draw (suggested PSU of 200 W vs. 600 W) and potentially lower cost (though no launch MSRP is available) might matter, provided the workload is simple enough to not require the GV100's massive memory and compute resources. The data does not support any scenario where the P104-100 wins on performance.
Architecture Differences
The architectural gap between these two GPUs is generational and fundamental. The Quadro GV100 is built on the Volta architecture using a 12 nm process node at TSMC, while the P104-100 uses the Pascal architecture on a 16 nm node, also at TSMC. This process advantage is part of why the GV100 packs significantly more hardware: 21,100 million transistors on an 815 mm² die, compared to 7,200 million transistors on a 314 mm² die for the P104-100. The transistor density is also higher on the GV100, at 25.9M per mm² versus 22.9M per mm².
The compute capabilities diverge sharply. The GV100 features 5,120 shading units, 320 texture mapping units, and 128 ROPs, alongside 640 dedicated tensor cores. The P104-100 has 1,920 shading units, 120 TMUs, and 64 ROPs, with no tensor cores at all. This absence of tensor cores is critical: the GV100's 33.32 TFLOPS of FP16 performance (listed as 2:1 ratio) is enabled by these cores, while the P104-100's FP16 performance is a paltry 104.0 GFLOPS (1:64 ratio), a 320x difference in raw half-precision throughput. The FP32 performance tells a similar story: the GV100 delivers 16.66 TFLOPS compared to the P104-100's 6.655 TFLOPS.
Memory architecture is another major divider. The GV100 uses 32 GB of HBM2 memory on a 4096-bit bus, yielding 868.4 GB/s of bandwidth. The P104-100 uses 4 GB of GDDR5X on a 256-bit bus, providing 320.3 GB/s. This is a 2.7x advantage in bandwidth for the GV100, and a 8x advantage in capacity, making it far more suitable for large datasets. The P104-100's PCIe 1.0 x4 interface, versus the GV100's PCIe 3.0 x16, further cements the mining card's role as a secondary compute device, not a primary system component.
Head-to-Head Benchmarks
The two head-to-head benchmarks show a decisive, near-comical margin in favor of the Quadro GV100. In Geekbench OpenCL, the GV100 scores 150,004 against the P104-100's 52,368, a delta of 186.4%. In Geekbench Vulkan, the GV100 scores 139,526 against 45,165, a delta of 208.9%. These are not incremental gains; they represent a more than doubling of performance in the Vulkan test and nearly tripling in OpenCL.
What do these numbers mean in context? The GV100's average benchmark score is 35,520, which places it in the 80th percentile of all GPUs. Its nearest rivals include the NVIDIA GeForce RTX 5070 Ti Mobile (avg score 35,435, just 0.2% behind), the AMD Radeon Pro Duo (35,860, 0.9% ahead), and the NVIDIA T1000 (36,289, 2.1% ahead). This places the GV100 in a competitive tier with modern mobile cards and older dual-GPU workstations. The P104-100, meanwhile, has an average benchmark score of 32,982, sitting in the 77th percentile. Its nearest rivals are the NVIDIA T600 Mobile (32,849, 0.4% behind), the NVIDIA T550 Mobile (33,161, 0.5% ahead), and the NVIDIA GeForce RTX 3050 Mobile (33,170, 0.6% ahead). The P104-100 is effectively on par with entry-level mobile graphics solutions, despite being a desktop card.
The deltaPct values in the head-to-head tests are the most important takeaway. A 186.4% lead in OpenCL means the GV100 is not just faster; it is in a completely different performance class. The data suggests that the P104-100's architecture, with its limited FP16 and small memory bus, is bottlenecked in ways that the GV100 simply is not. For any compute workload that can utilize the GV100's tensor cores or its massive HBM2 bandwidth, the P104-100 would be a non-starter.
Specification Differences
The specifications where the two cards differ are numerous and define their distinct purposes:
- Chip & Architecture: GV100 (Volta) vs. GP104 (Pascal)
- Generation: Quadro Volta (Vx000) vs. Mining GPUs
- Process Node: 12 nm vs. 16 nm
- Transistors: 21,100 million vs. 7,200 million
- Die Size: 815 mm² vs. 314 mm²
- Transistor Density: 25.9M / mm² vs. 22.9M / mm²
- Base Clock: 1132 MHz vs. 1607 MHz
- Boost Clock: 1627 MHz vs. 1733 MHz
- Memory Speed: 848 MHz / 1696 Mbps effective vs. 1251 MHz / 10 Gbps effective
- Memory Size: 32 GB vs. 4 GB
- Memory Type: HBM2 vs. GDDR5X
- Memory Bus Width: 4096 bit vs. 256 bit
- Memory Bandwidth: 868.4 GB/s vs. 320.3 GB/s
- Shading Units: 5120 vs. 1920
- TMUs: 320 vs. 120
- ROPs: 128 vs. 64
- Tensor Cores: 640 vs. None
- Pixel Rate: 208.3 GPixel/s vs. 110.9 GPixel/s
- Texture Rate: 520.6 GTexel/s vs. 208.0 GTexel/s
- FP32 Performance: 16.66 TFLOPS vs. 6.655 TFLOPS
- FP16 Performance: 33.32 TFLOPS (2:1) vs. 104.0 GFLOPS (1:64)
- TDP: 250 W vs. (not specified)
- Suggested PSU: 600 W vs. 200 W
- Bus Interface: PCIe 3.0 x16 vs. PCIe 1.0 x4
- Display Outputs: 4x DisplayPort 1.4a vs. No outputs
- Release Date: 2018-03-26 vs. 2017-12-11
- Launch MSRP: 8,999 USD vs. (not available)
The P104-100 has a higher base and boost clock, but this is meaningless given the massive core count and memory advantages of the GV100. The lack of a TDP for the P104-100 is notable, but the suggested PSU of 200 W vs. 600 W indicates a much lower power envelope.
FAQ
Q: Which card has more memory and bandwidth?
A: The Quadro GV100 has 32 GB of HBM2 memory on a 4096-bit bus, providing 868.4 GB/s of bandwidth. The P104-100 has 4 GB of GDDR5X on a 256-bit bus, providing 320.3 GB/s.
Q: Does the P104-100 support display outputs?
A: No. The P104-100 has "No outputs" listed for its display outputs, making it unsuitable for any use case requiring a monitor connection. The GV100 has 4x DisplayPort 1.4a outputs.
Q: What is the performance difference in Vulkan?
A: In Geekbench Vulkan, the GV100 scores 139,526 compared to the P104-100's 45,165, a delta of 208.9% in favor of the GV100.
Q: What is the FP16 compute capability of each card?
A: The GV100 delivers 33.32 TFLOPS of FP16 performance (2:1 ratio), thanks to its 640 tensor cores. The P104-100 delivers only 104.0 GFLOPS (1:64 ratio) and has no tensor cores.
Q: Which card has a higher average benchmark score?
A: The GV100 has an average benchmark score of 35,520, while the P104-100 has an average score of 32,982. The GV100 sits in the 80th percentile of all GPUs, while the P104-100 sits in the 77th percentile.
Q: What are the bus interfaces of these cards?
A: The GV100 uses PCIe 3.0 x16, a standard full-bandwidth interface. The P104-100 uses PCIe 1.0 x4, a severely limited interface that would bottleneck even modest data transfers.
The Verdict
The data is unambiguous: the NVIDIA Quadro GV100 is a vastly superior product in every measurable way. For professionals working with large datasets, machine learning, or high-end visualization, the GV100's 32 GB of HBM2 memory, 640 tensor cores, and 16.66 TFLOPS of FP32 performance make it a capable if expensive tool. Its launch MSRP of 8,999 USD reflects its position at the top of the workstation stack. The benchmark results show it performing on par with modern mobile RTX 50-series cards and older dual-GPU workstations, which is impressive for a card released in 2018.
The NVIDIA P104-100, on the other hand, is a relic of the cryptocurrency mining boom. Its lack of display outputs, limited PCIe interface, and small 4 GB memory pool make it useless for any modern gaming, professional, or even general-purpose computing task. Its only potential advantage is power consumption, with a suggested PSU of 200 W versus 600 W for the GV100. However, its performance is on par with entry-level mobile GPUs like the T600 Mobile or RTX 3050 Mobile, making it a poor choice even for compute tasks where power efficiency is paramount. The data suggests the P104-100 should be avoided, while the GV100 remains a relevant, high-performance option for those who need its specific capabilities.