NVIDIA P104-100 vs NVIDIA Quadro GV100 Comparison

NVIDIA
GEFORCE

NVIDIA P104-100

CORE STATE GP104
VRAM 4 GB
CLOCK SPEED 1733 MHz
TDP
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2017
VS
NVIDIA
GEFORCE

Quadro GV100

CORE STATE GV100
VRAM 32 GB
CLOCK SPEED 1627 MHz
TDP 250 W
BUS WIDTH 4096 bit
ARCHITECTURE Volta
nm
PROCESS 12 nm
LAUNCH DATE 2018

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
1,413
N/A
geekbench_opencl
52,368
150,004
geekbench_vulkan
45,165
139,526
passmark_directx_10
N/A
140
passmark_directx_11
N/A
168
passmark_directx_12
N/A
84
passmark_directx_9
N/A
207
passmark_g2d
N/A
836
passmark_g3d
N/A
19,650
passmark_gpu_compute
N/A
9,069

Analysis: NVIDIA P104-100 vs NVIDIA Quadro GV100

The NVIDIA Quadro GV100 and the NVIDIA P104-100 are two very different products that share a manufacturer and little else. The GV100 is a professional workstation card built on the Volta architecture, designed for compute-heavy tasks, while the P104-100 is a mining-specific GPU stripped of display outputs and built on the older Pascal architecture. The data shows a massive performance gulf between them, but the story is more nuanced than one simply being "better" than the other. This analysis breaks down where each card wins, what the architectural differences mean, and who should consider which based on the benchmark results.

Where Each One Wins

The Quadro GV100 is the clear winner in every head-to-head benchmark recorded, but the nature of those wins reveals its intended use case. The GV100 dominates in compute-oriented workloads, specifically in OpenCL and Vulkan APIs. These are general-purpose compute interfaces, not gaming-specific tests, which aligns with the card's professional positioning. The data shows the GV100 winning 2 out of 2 head-to-head comparisons, with the P104-100 failing to secure a single victory.

The P104-100's situation is different. It wins nowhere in the direct comparison, but its benchmark profile suggests it was built for a very specific, narrow task: cryptocurrency mining. With no display outputs and a PCIe 1.0 x4 bus interface, it was never intended for interactive use, gaming, or even professional visualization. Its inclusion in the "Mining GPUs" generation and its lack of any DirectX or OpenGL benchmark scores in the head-to-head data point to a card that was optimized for raw, repetitive compute tasks rather than diverse workloads. The GV100, by contrast, has a full suite of benchmark scores across DirectX 9 through 12, OpenGL, and compute, indicating a general-purpose compute and graphics capability.

The use-case split is stark: the GV100 is a versatile, high-end compute and visualization tool, while the P104-100 is a single-purpose device whose only advantage would be in a scenario where its lower power draw (suggested PSU of 200 W vs. 600 W) and potentially lower cost (though no launch MSRP is available) might matter, provided the workload is simple enough to not require the GV100's massive memory and compute resources. The data does not support any scenario where the P104-100 wins on performance.

Architecture Differences

The architectural gap between these two GPUs is generational and fundamental. The Quadro GV100 is built on the Volta architecture using a 12 nm process node at TSMC, while the P104-100 uses the Pascal architecture on a 16 nm node, also at TSMC. This process advantage is part of why the GV100 packs significantly more hardware: 21,100 million transistors on an 815 mm² die, compared to 7,200 million transistors on a 314 mm² die for the P104-100. The transistor density is also higher on the GV100, at 25.9M per mm² versus 22.9M per mm².

The compute capabilities diverge sharply. The GV100 features 5,120 shading units, 320 texture mapping units, and 128 ROPs, alongside 640 dedicated tensor cores. The P104-100 has 1,920 shading units, 120 TMUs, and 64 ROPs, with no tensor cores at all. This absence of tensor cores is critical: the GV100's 33.32 TFLOPS of FP16 performance (listed as 2:1 ratio) is enabled by these cores, while the P104-100's FP16 performance is a paltry 104.0 GFLOPS (1:64 ratio), a 320x difference in raw half-precision throughput. The FP32 performance tells a similar story: the GV100 delivers 16.66 TFLOPS compared to the P104-100's 6.655 TFLOPS.

Memory architecture is another major divider. The GV100 uses 32 GB of HBM2 memory on a 4096-bit bus, yielding 868.4 GB/s of bandwidth. The P104-100 uses 4 GB of GDDR5X on a 256-bit bus, providing 320.3 GB/s. This is a 2.7x advantage in bandwidth for the GV100, and a 8x advantage in capacity, making it far more suitable for large datasets. The P104-100's PCIe 1.0 x4 interface, versus the GV100's PCIe 3.0 x16, further cements the mining card's role as a secondary compute device, not a primary system component.

Head-to-Head Benchmarks

The two head-to-head benchmarks show a decisive, near-comical margin in favor of the Quadro GV100. In Geekbench OpenCL, the GV100 scores 150,004 against the P104-100's 52,368, a delta of 186.4%. In Geekbench Vulkan, the GV100 scores 139,526 against 45,165, a delta of 208.9%. These are not incremental gains; they represent a more than doubling of performance in the Vulkan test and nearly tripling in OpenCL.

What do these numbers mean in context? The GV100's average benchmark score is 35,520, which places it in the 80th percentile of all GPUs. Its nearest rivals include the NVIDIA GeForce RTX 5070 Ti Mobile (avg score 35,435, just 0.2% behind), the AMD Radeon Pro Duo (35,860, 0.9% ahead), and the NVIDIA T1000 (36,289, 2.1% ahead). This places the GV100 in a competitive tier with modern mobile cards and older dual-GPU workstations. The P104-100, meanwhile, has an average benchmark score of 32,982, sitting in the 77th percentile. Its nearest rivals are the NVIDIA T600 Mobile (32,849, 0.4% behind), the NVIDIA T550 Mobile (33,161, 0.5% ahead), and the NVIDIA GeForce RTX 3050 Mobile (33,170, 0.6% ahead). The P104-100 is effectively on par with entry-level mobile graphics solutions, despite being a desktop card.

The deltaPct values in the head-to-head tests are the most important takeaway. A 186.4% lead in OpenCL means the GV100 is not just faster; it is in a completely different performance class. The data suggests that the P104-100's architecture, with its limited FP16 and small memory bus, is bottlenecked in ways that the GV100 simply is not. For any compute workload that can utilize the GV100's tensor cores or its massive HBM2 bandwidth, the P104-100 would be a non-starter.

Specification Differences

The specifications where the two cards differ are numerous and define their distinct purposes:

  • Chip & Architecture: GV100 (Volta) vs. GP104 (Pascal)
  • Generation: Quadro Volta (Vx000) vs. Mining GPUs
  • Process Node: 12 nm vs. 16 nm
  • Transistors: 21,100 million vs. 7,200 million
  • Die Size: 815 mm² vs. 314 mm²
  • Transistor Density: 25.9M / mm² vs. 22.9M / mm²
  • Base Clock: 1132 MHz vs. 1607 MHz
  • Boost Clock: 1627 MHz vs. 1733 MHz
  • Memory Speed: 848 MHz / 1696 Mbps effective vs. 1251 MHz / 10 Gbps effective
  • Memory Size: 32 GB vs. 4 GB
  • Memory Type: HBM2 vs. GDDR5X
  • Memory Bus Width: 4096 bit vs. 256 bit
  • Memory Bandwidth: 868.4 GB/s vs. 320.3 GB/s
  • Shading Units: 5120 vs. 1920
  • TMUs: 320 vs. 120
  • ROPs: 128 vs. 64
  • Tensor Cores: 640 vs. None
  • Pixel Rate: 208.3 GPixel/s vs. 110.9 GPixel/s
  • Texture Rate: 520.6 GTexel/s vs. 208.0 GTexel/s
  • FP32 Performance: 16.66 TFLOPS vs. 6.655 TFLOPS
  • FP16 Performance: 33.32 TFLOPS (2:1) vs. 104.0 GFLOPS (1:64)
  • TDP: 250 W vs. (not specified)
  • Suggested PSU: 600 W vs. 200 W
  • Bus Interface: PCIe 3.0 x16 vs. PCIe 1.0 x4
  • Display Outputs: 4x DisplayPort 1.4a vs. No outputs
  • Release Date: 2018-03-26 vs. 2017-12-11
  • Launch MSRP: 8,999 USD vs. (not available)

The P104-100 has a higher base and boost clock, but this is meaningless given the massive core count and memory advantages of the GV100. The lack of a TDP for the P104-100 is notable, but the suggested PSU of 200 W vs. 600 W indicates a much lower power envelope.

FAQ

Q: Which card has more memory and bandwidth?

A: The Quadro GV100 has 32 GB of HBM2 memory on a 4096-bit bus, providing 868.4 GB/s of bandwidth. The P104-100 has 4 GB of GDDR5X on a 256-bit bus, providing 320.3 GB/s.

Q: Does the P104-100 support display outputs?

A: No. The P104-100 has "No outputs" listed for its display outputs, making it unsuitable for any use case requiring a monitor connection. The GV100 has 4x DisplayPort 1.4a outputs.

Q: What is the performance difference in Vulkan?

A: In Geekbench Vulkan, the GV100 scores 139,526 compared to the P104-100's 45,165, a delta of 208.9% in favor of the GV100.

Q: What is the FP16 compute capability of each card?

A: The GV100 delivers 33.32 TFLOPS of FP16 performance (2:1 ratio), thanks to its 640 tensor cores. The P104-100 delivers only 104.0 GFLOPS (1:64 ratio) and has no tensor cores.

Q: Which card has a higher average benchmark score?

A: The GV100 has an average benchmark score of 35,520, while the P104-100 has an average score of 32,982. The GV100 sits in the 80th percentile of all GPUs, while the P104-100 sits in the 77th percentile.

Q: What are the bus interfaces of these cards?

A: The GV100 uses PCIe 3.0 x16, a standard full-bandwidth interface. The P104-100 uses PCIe 1.0 x4, a severely limited interface that would bottleneck even modest data transfers.

The Verdict

The data is unambiguous: the NVIDIA Quadro GV100 is a vastly superior product in every measurable way. For professionals working with large datasets, machine learning, or high-end visualization, the GV100's 32 GB of HBM2 memory, 640 tensor cores, and 16.66 TFLOPS of FP32 performance make it a capable if expensive tool. Its launch MSRP of 8,999 USD reflects its position at the top of the workstation stack. The benchmark results show it performing on par with modern mobile RTX 50-series cards and older dual-GPU workstations, which is impressive for a card released in 2018.

The NVIDIA P104-100, on the other hand, is a relic of the cryptocurrency mining boom. Its lack of display outputs, limited PCIe interface, and small 4 GB memory pool make it useless for any modern gaming, professional, or even general-purpose computing task. Its only potential advantage is power consumption, with a suggested PSU of 200 W versus 600 W for the GV100. However, its performance is on par with entry-level mobile GPUs like the T600 Mobile or RTX 3050 Mobile, making it a poor choice even for compute tasks where power efficiency is paramount. The data suggests the P104-100 should be avoided, while the GV100 remains a relevant, high-performance option for those who need its specific capabilities.

DETAILED SPECIFICATIONS

SPECIFICATION
P104-100
Quadro GV100
Core Specs
Shading Units
1,920
5,120 +166.7%
Shaders
1,920
5,120 +166.7%
TMUs
120
320 +166.7%
ROPs
64
128 +100.0%
SM Count
15
80 +433.3%
Clocks
Base Clock
1607 MHz
1132 MHz
Boost Clock
1733 MHz
1627 MHz
Memory Clock
1251 MHz 10 Gbps effective
848 MHz 1696 Mbps effective
Memory
Memory Size
4 GB
32 GB
VRAM (MB)
4,096
32,768 +700.0%
Memory Type
GDDR5X
HBM2
Memory Bus
256 bit
4096 bit
Bandwidth
320.3 GB/s
868.4 GB/s
Cache
L1 Cache
48 KB (per SM)
128 KB (per SM)
L2 Cache
2 MB
6 MB
Performance
Pixel Rate
110.9 GPixel/s
208.3 GPixel/s
Texture Rate
208.0 GTexel/s
520.6 GTexel/s
FP32 (TFLOPS)
6.655 TFLOPS
16.66 TFLOPS
FP64 (TFLOPS)
208.0 GFLOPS (1:32)
8.330 TFLOPS (1:2)
FP16 (TFLOPS)
104.0 GFLOPS (1:64)
33.32 TFLOPS (2:1)
AI/RT
Tensor Cores
640
Power
TDP
250 W
TDP (W)
250
Suggested PSU
200 W
600 W
Power Connectors
1x 8-pin
1x 8-pin
Architecture
Architecture
Pascal
Volta
GPU Name
GP104
GV100
Generation
Mining GPUs
Quadro Volta (Vx000)
Process Size
16 nm
12 nm
Transistors
7,200 million
21,100 million
Die Size
314 mm²
815 mm²
Foundry
TSMC
TSMC
Density
22.9M / mm²
25.9M / mm²
API Support
DirectX
12 (12_1)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
6.1
7.0
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
Outputs
No outputs
4x DisplayPort 1.4a
Bus Interface
PCIe 1.0 x4
PCIe 3.0 x16
Other
Launch Price
8,999 USD
Production
End-of-life
End-of-life
Predecessor
Quadro Pascal
Successor
Quadro Turing
View P104-100 Details View Quadro GV100 Details