NVIDIA P104-100 vs NVIDIA Quadro RTX 8000 Comparison
NVIDIA P104-100
Quadro RTX 8000
PERFORMANCE BENCHMARKS
Analysis: NVIDIA P104-100 vs NVIDIA Quadro RTX 8000
Where Each One Wins
The benchmark data splits these two NVIDIA cards into completely different performance tiers, with the Quadro RTX 8000 dominating every recorded test. Out of the two head-to-head benchmarks, the Quadro RTX 8000 wins both, leaving the P104-100 with zero wins in direct comparison. This is not a close contest; it is a generational and architectural gap expressed through raw compute results.
In Geekbench OpenCL, the Quadro RTX 8000 scores 101,883 against the P104-100’s 52,368. That is a 48.6% margin in favor of the Turing card, meaning the RTX 8000 delivers roughly double the compute throughput in this workload. The P104-100 is not merely slower; it is decisively outclassed in general-purpose GPU compute, which is the primary use case for both cards.
The Vulkan results are even more lopsided. The Quadro RTX 8000 posts 122,637, while the P104-100 manages 45,165. The delta percentage here is 63.2% in favor of the RTX 8000, indicating that the Turing architecture’s advantage grows in graphics API workloads. Vulkan is a low-level API that rewards raw shader throughput, memory bandwidth, and driver efficiency, all areas where the RTX 8000’s TU102 chip excels.
Where does the P104-100 actually win? The data shows no benchmark victories. Its strongest recorded result is the Geekbench OpenCL score of 52,368, which places it in the 77th percentile of all GPUs in the database. That is a respectable position, but the Quadro RTX 8000, despite a slightly lower overall percentile of 74, achieves far higher absolute scores in every shared test. The percentile difference is misleading: the RTX 8000’s average benchmark score of 28,421 is dragged down by its inclusion of older DirectX 9, 10, 11, and 12 tests, where it scores between 79 and 211. Those legacy tests are not part of the P104-100’s benchmark suite, so the average comparison is not apples-to-apples.
In practical terms, the P104-100’s wins are limited to what its Pascal architecture can do efficiently: 6.655 TFLOPS of FP32, 208.0 GTexel/s of texture fill, and 110.9 GPixel/s of pixel fill. Those numbers are competitive for its era and class, but they are dwarfed by the RTX 8000’s 16.31 TFLOPS FP32, 509.8 GTexel/s, and 169.9 GPixel/s. The P104-100 also has a 1:64 FP16 ratio, meaning its half-precision throughput is negligible at 104.0 GFLOPS, while the RTX 8000 delivers 32.62 TFLOPS FP16 at a 2:1 ratio. For any workload that touches half-precision, the RTX 8000 is in another universe.
The Verdict
The data is unambiguous: the NVIDIA Quadro RTX 8000 is the superior card in every measured benchmark. If the choice is between these two, the RTX 8000 is the only rational pick for anyone needing maximum compute performance, graphics capability, or memory capacity. The P104-100, despite being a capable mining-oriented Pascal card, cannot compete on any metric that appears in the database.
Who should choose the P104-100? Strictly speaking, only someone constrained by the fact that the P104-100 has no display outputs and was designed for mining workloads. But even in that niche, the benchmark data shows it loses to the RTX 8000 in both OpenCL and Vulkan. The P104-100’s nearest rivals in the database are mobile and low-power cards like the NVIDIA T600 Mobile (0.4% ahead), T550 Mobile (0.5% behind), and RTX 3050 Mobile (0.6% behind). That places the P104-100 in the performance range of thin-and-light laptop GPUs, not workstation-class hardware.
The Quadro RTX 8000’s nearest rivals, by contrast, include the AMD Radeon RX 570 (1.2% faster), AMD Radeon R9 M295X (0.6% faster), and AMD FirePro S7150 (1.1% slower). These are older desktop and mobile parts, and the RTX 8000 still trades blows with them in average score despite having 48 GB of GDDR6 memory and 72 RT cores. The RTX 8000’s average benchmark score of 28,421 is lower than the P104-100’s 32,982, but that is an artifact of the different benchmark suites each card was tested with.
For professional visualization, AI inference, or any task requiring large memory footprints, the RTX 8000 is the clear winner. Its 48 GB memory capacity is 12 times larger than the P104-100’s 4 GB. Its 672.0 GB/s bandwidth is more than double the P104-100’s 320.3 GB/s. Its 4608 shading units are 2.4 times the P104-100’s 1920. Every architectural advantage points the same direction.
The verdict is simple: the Quadro RTX 8000 wins on every recorded benchmark and every relevant specification. The P104-100 is only viable if the workload is so narrow that the RTX 8000’s features are irrelevant, but the data does not support that scenario.
Head-to-Head Benchmarks
The database records two direct comparisons between these cards, and both are decisive wins for the Quadro RTX 8000.
Geekbench OpenCL: The RTX 8000 scores 101,883 against the P104-100’s 52,368. The delta is 48.6% in favor of the RTX 8000. This test measures general-purpose compute on the GPU, and the Turing architecture’s 4608 shading units, 576 tensor cores, and 16.31 TFLOPS FP32 throughput produce a score nearly double that of the Pascal card. The P104-100’s 6.655 TFLOPS FP32 and 1920 shading units simply cannot keep pace. OpenCL workloads that are memory-bound also favor the RTX 8000, which has 672.0 GB/s of bandwidth versus 320.3 GB/s on the P104-100.
Geekbench Vulkan: The RTX 8000 scores 122,637, while the P104-100 scores 45,165. The delta is 63.2%, the largest margin in the head-to-head set. Vulkan is a low-overhead graphics API that exposes raw hardware capabilities, and the RTX 8000’s 72 RT cores and 576 tensor cores, while not directly used in standard Vulkan rasterization, indicate a much more capable compute and graphics pipeline. The P104-100’s Vulkan score of 45,165 is less than half of its own OpenCL score, suggesting that Pascal’s Vulkan driver path is less optimized or that the card’s hardware lacks the features needed to scale in this API. The RTX 8000’s Vulkan score is actually higher than its OpenCL score, showing that Turing’s architecture is particularly well-suited to modern low-level APIs.
The wins are 2 for the RTX 8000, 0 for the P104-100. There are no other shared benchmarks in the database. The P104-100 has one additional benchmark, 3DMark Steel Nomad DX12, where it scores 1413, but the RTX 8000 does not have a recorded result for that test, so no comparison is possible. The RTX 8000 has additional Passmark tests (DirectX 9, 10, 11, 12, G2D, G3D, and GPU Compute), but the P104-100 has no corresponding scores, so those cannot be compared directly either.
The largest win in absolute terms is the Vulkan test, where the RTX 8000 leads by 77,472 points. The largest percentage win is also Vulkan at 63.2%. Both metrics point to the same conclusion: the RTX 8000 is not just faster, it is categorically more capable in modern graphics and compute workloads.
FAQ
Q: Which card wins in Geekbench OpenCL?
A: The NVIDIA Quadro RTX 8000 wins with a score of 101,883 against the P104-100’s 52,368, a 48.6% advantage.
Q: How much faster is the RTX 8000 in Vulkan?
A: The RTX 8000 scores 122,637 in Geekbench Vulkan, which is 63.2% higher than the P104-100’s 45,165.
Q: Does the P104-100 have any benchmark wins over the RTX 8000?
A: No. The database records zero wins for the P104-100 and two wins for the RTX 8000 in head-to-head tests.
Q: What is the memory capacity difference?
A: The RTX 8000 has 48 GB of GDDR6 memory on a 384-bit bus, while the P104-100 has 4 GB of GDDR5X on a 256-bit bus. The RTX 8000’s bandwidth is 672.0 GB/s versus 320.3 GB/s.
Q: Are these cards still in production?
A: Both are end-of-life. The P104-100 was released in December 2017, and the RTX 8000 was released in August 2018.
Q: Which card has ray tracing and tensor cores?
A: Only the RTX 8000. It has 72 RT cores and 576 tensor cores. The P104-100 has neither.
Architecture Differences
The two cards come from different NVIDIA architectures and generations. The P104-100 is based on the GP104 chip, built on the Pascal architecture, and belongs to the Mining GPUs generation. It uses a 16 nm process at TSMC, packs 7,200 million transistors on a 314 mm² die, and has a transistor density of 22.9 million per mm². The RTX 8000 uses the TU102 chip, built on the Turing architecture, from the Quadro Turing (Tx000) generation. It is fabricated on a 12 nm process, also at TSMC, but contains 18,600 million transistors on a much larger 754 mm² die, with a density of 24.7 million per mm².
The compute core counts differ substantially. The P104-100 has 1920 shading units, 120 texture mapping units, and 64 ROPs. The RTX 8000 has 4608 shading units, 288 TMUs, and 96 ROPs. The RTX 8000 also adds dedicated hardware that the P104-100 lacks entirely: 72 RT cores for ray tracing and 576 tensor cores for AI acceleration. This is a fundamental architectural difference; Pascal has no equivalent hardware, while Turing was designed with these features as core capabilities.
FP16 performance illustrates the architectural gap. The P104-100 delivers 104.0 GFLOPS of FP16 at a 1:64 ratio, meaning its half-precision throughput is a tiny fraction of its FP32 rate. The RTX 8000 delivers 32.62 TFLOPS of FP16 at a 2:1 ratio, meaning it is twice as fast in FP16 as in FP32. For any workload that uses half-precision, such as certain AI inference or graphics effects, the RTX 8000 is over 300 times faster than the P104-100.
The memory subsystems are also architecturally different. The P104-100 uses 4 GB of GDDR5X with a 256-bit bus and 320.3 GB/s bandwidth. The RTX 8000 uses 48 GB of GDDR6 with a 384-bit bus and 672.0 GB/s bandwidth. The RTX 8000 also has a higher memory clock at 1750 MHz (14 Gbps effective) versus 1251 MHz (10 Gbps effective) on the P104-100.
The API support differs as well. The P104-100 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The RTX 8000 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The 12_2 feature level on the RTX 8000 includes ray tracing and mesh shaders, which the P104-100 cannot expose.
Specification Differences
The following specifications differ between the two cards, with the RTX 8000 leading in every compute and memory category.
Process node: 16 nm on the P104-100 versus 12 nm on the RTX 8000, both fabricated by TSMC.
Transistors: 7,200 million on the P104-100 versus 18,600 million on the RTX 8000. Die size is 314 mm² versus 754 mm².
Base clock: 1607 MHz on the P104-100 versus 1395 MHz on the RTX 8000. Boost clock: 1733 MHz on the P104-100 versus 1770 MHz on the RTX 8000. The P104-100 has a higher base clock, but the RTX 8000 boosts higher.
Memory: 4 GB GDDR5X on a 256-bit bus with 320.3 GB/s bandwidth on the P104-100. The RTX 8000 has 48 GB GDDR6 on a 384-bit bus with 672.0 GB/s bandwidth. Memory clock is 1251 MHz (10 Gbps effective) versus 1750 MHz (14 Gbps effective).
Shading units: 1920 versus 4608. TMUs: 120 versus 288. ROPs: 64 versus 96. RT cores: none versus 72. Tensor cores: none versus 576.
Pixel rate: 110.9 GPixel/s versus 169.9 GPixel/s. Texture rate: 208.0 GTexel/s versus 509.8 GTexel/s. FP32: 6.655 TFLOPS versus 16.31 TFLOPS. FP16: 104.0 GFLOPS (1:64) versus 32.62 TFLOPS (2:1).
TDP: not recorded for the P104-100, versus 260 W for the RTX 8000. Suggested PSU: 200 W versus 600 W. Power connectors: 1x 8-pin on the P104-100, versus 1x 6-pin + 1x 8-pin on the RTX 8000.
Bus interface: PCIe 1.0 x4 on the P104-100, versus PCIe 3.0 x16 on the RTX 8000. This is a massive difference; the P104-100 is severely bandwidth-limited on the bus, which may explain some of its lower benchmark scores in API-bound tests.
Display outputs: The P104-100 has no outputs. The RTX 8000 has 4x DisplayPort 1.4a and 1x USB Type-C.
Dimensions: Both are 267 mm (10.5 inches) long. The RTX 8000 adds a height specification of 111 mm (4.4 inches), which is not recorded for the P104-100. Both are dual-slot cards.
DirectX support: 12 (12_1) on the P104-100 versus 12 Ultimate (12_2) on the RTX 8000. OpenGL is 4.6 on both. Vulkan is 1.4 on both.
Release dates: The P104-100 launched on December 11, 2017. The RTX 8000 launched on August 12, 2018. Both are end-of-life. The RTX 8000 has a recorded launch MSRP of 9,999 USD; the P104-100 has no recorded MSRP.