NVIDIA GeForce RTX 3090 vs NVIDIA P104-100 Comparison
NVIDIA GeForce RTX 3090
P104-100
PERFORMANCE BENCHMARKS
Analysis: NVIDIA GeForce RTX 3090 vs NVIDIA P104-100
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA P104-100 has a higher average benchmark score at 32,982, compared to the NVIDIA GeForce RTX 3090 at 27,565. However, this aggregate figure is heavily influenced by the specific test sets available for each card.
Q: How do the two cards compare in the 3DMark Steel Nomad DX12 test?
A: The RTX 3090 is the clear winner, scoring 5,118 against the P104-100's 1,413. This represents a 72.4% advantage for the RTX 3090 in this test.
Q: What is the difference in compute performance between the two cards?
A: In Geekbench OpenCL, the RTX 3090 scores 172,758, while the P104-100 scores 52,368. The RTX 3090 is 69.7% ahead in this test.
Q: Are the two cards close in any benchmark?
A: The closest margin is in Geekbench Vulkan, where the RTX 3090 scores 53,927 versus the P104-100's 45,165. Here, the RTX 3090 leads by only 16.2%.
Q: Which card has a higher percentile ranking among all GPUs?
A: The P104-100 sits at the 77th percentile, while the RTX 3090 is at the 73rd percentile. This suggests the P104-100 outperforms a larger share of the GPU population in the database's aggregate ranking.
Q: What is the fabrication process difference between the two?
A: The P104-100 uses a 16 nm process at TSMC, while the RTX 3090 uses an 8 nm process at Samsung. The RTX 3090 also packs significantly more transistors: 28,300 million versus 7,200 million.
Architecture Differences
The two GPUs come from different NVIDIA generations and are built for entirely different purposes. The P104-100 is part of the Mining GPUs generation, based on the Pascal architecture with the GP104 chip. It was released in December 2017. The RTX 3090 belongs to the GeForce 30 series, uses the Ampere architecture with the GA102 chip, and launched in August 2020. The architectural gap spans roughly three years of development.
The process nodes differ substantially. The P104-100 is fabricated on a 16 nm process at TSMC, which yields a transistor density of 22.9 million transistors per square millimeter. The RTX 3090 moves to an 8 nm process at Samsung, achieving a much higher density of 45.1 million transistors per square millimeter. This allows the RTX 3090 to house 28,300 million transistors on a 628 mm² die, compared to the P104-100's 7,200 million on a 314 mm² die.
The RTX 3090 introduces dedicated hardware that the Pascal-based P104-100 lacks entirely. The RTX 3090 includes 82 RT cores for ray tracing and 328 tensor cores for AI acceleration. The P104-100 has no such units listed in its specifications. This is a fundamental architectural difference: one is a compute-focused mining card, the other is a full-featured consumer flagship with specialized processing blocks.
Memory architecture also diverges. The P104-100 uses 4 GB of GDDR5X on a 256-bit bus, delivering 320.3 GB/s of bandwidth. The RTX 3090 uses 24 GB of GDDR6X on a 384-bit bus, providing 936.2 GB/s. The RTX 3090's memory clock is 1219 MHz with 19.5 Gbps effective speed, while the P104-100's memory runs at 1251 MHz with 10 Gbps effective. The RTX 3090's advantage in bandwidth is roughly triple that of the P104-100.
Compute capabilities differ sharply. The RTX 3090 has 10,496 shading units, 328 TMUs, and 112 ROPs. The P104-100 has 1,920 shading units, 120 TMUs, and 64 ROPs. The RTX 3090 also has a major FP16 advantage: it delivers 35.58 TFLOPS with a 1:1 ratio, while the P104-100 manages only 104.0 GFLOPS with a 1:64 ratio. FP32 performance is 35.58 TFLOPS versus 6.655 TFLOPS in favor of the RTX 3090.
Head-to-Head Benchmarks
The head-to-head results show a consistent sweep for the RTX 3090, winning all three recorded tests. The margin varies considerably by workload, which is informative for understanding where each card's strengths lie.
In 3DMark Steel Nomad DX12, the RTX 3090 scores 5,118 against the P104-100's 1,413. The delta is 72.4% in favor of the RTX 3090. This is a modern DX12 workload that stresses the full feature set of the newer architecture, including the RT and tensor cores. The P104-100, despite being a capable card in its own right, falls far behind here.
In Geekbench OpenCL, the gap is similar: the RTX 3090 posts 172,758 versus 52,368, a 69.7% advantage. OpenCL compute performance is heavily dependent on raw shader throughput and memory bandwidth, both of which favor the RTX 3090 by wide margins. The RTX 3090 has over five times the shading units and nearly three times the memory bandwidth.
The closest contest is Geekbench Vulkan. Here the RTX 3090 scores 53,927, while the P104-100 scores 45,165. The RTX 3090's lead narrows to 16.2%. This suggests that the P104-100's Pascal architecture is relatively efficient in Vulkan workloads, closing the gap substantially compared to DX12 or OpenCL. It also indicates that Vulkan may not fully utilize the RTX 3090's extra compute resources, or that driver overhead plays a role in this specific API.
The overall win count is 3 to 0 in favor of the RTX 3090. The P104-100 does not win any benchmark in the direct comparison. However, the magnitude of the losses varies dramatically: from a 72.4% deficit in Steel Nomad to a much more competitive 16.2% deficit in Vulkan. This is a meaningful distinction for users who prioritize specific APIs.
Specification Differences
The specification table shows differences across nearly every major category. Process node: 16 nm for the P104-100 versus 8 nm for the RTX 3090. Foundry: TSMC versus Samsung. Transistors: 7,200 million versus 28,300 million. Die size: 314 mm² versus 628 mm². Transistor density: 22.9M per mm² versus 45.1M per mm².
Clock speeds: the P104-100 has a base clock of 1607 MHz and a boost of 1733 MHz. The RTX 3090 has a lower base of 1395 MHz but a boost of 1695 MHz. Memory clocks differ as well: 1251 MHz with 10 Gbps effective for the P104-100, versus 1219 MHz with 19.5 Gbps effective for the RTX 3090.
Memory configuration: 4 GB versus 24 GB, GDDR5X versus GDDR6X, 256-bit versus 384-bit bus, and 320.3 GB/s versus 936.2 GB/s bandwidth. The RTX 3090 has substantially more capacity and bandwidth.
Compute units: 1,920 shading units versus 10,496, 120 TMUs versus 328, and 64 ROPs versus 112. The RTX 3090 adds 82 RT cores and 328 tensor cores, which the P104-100 does not have. Pixel rate is 110.9 GPixel/s versus 189.8 GPixel/s. Texture rate is 208.0 GTexel/s versus 556.0 GTexel/s. FP32 is 6.655 TFLOPS versus 35.58 TFLOPS. FP16 is 104.0 GFLOPS versus 35.58 TFLOPS.
Power and physical specs: the P104-100 has no listed TDP, while the RTX 3090 has a 350 W TDP. The P104-100 is dual-slot with a 1x 8-pin connector and a 200 W suggested PSU. The RTX 3090 is triple-slot with a 1x 12-pin connector and a 750 W suggested PSU. The P104-100 is 267 mm long, while the RTX 3090 is 336 mm long, 140 mm high, and 61 mm wide. The P104-100 has no display outputs, while the RTX 3090 has 1x HDMI 2.1 and 3x DisplayPort 1.4a.
Bus interface: the P104-100 uses PCIe 1.0 x4, while the RTX 3090 uses PCIe 4.0 x16. DirectX support: 12 (12_1) versus 12 Ultimate (12_2). OpenGL is 4.6 for both, and Vulkan is 1.4 for both.
The Verdict
The data points to a straightforward conclusion for raw performance: the RTX 3090 is the dominant card in every head-to-head benchmark. It wins all three tests, with margins ranging from 16.2% to 72.4%. The RTX 3090 also brings modern features like RT cores, tensor cores, and a much larger memory pool. Its 24 GB of GDDR6X memory and 936.2 GB/s bandwidth are decisive for memory-heavy workloads.
However, the P104-100 has its own story. Its average benchmark score of 32,982 is higher than the RTX 3090's 27,565, and its percentile ranking is 77th versus 73rd. This suggests that in the database's aggregate scoring, the P104-100 performs better relative to other GPUs. This is likely because the P104-100's benchmark set is limited to three tests, two of which are compute-oriented, while the RTX 3090's set includes older DirectX 9, 10, and 11 tests where it scores low. The P104-100's Vulkan score of 45,165 is within 16.2% of the RTX 3090, making it competitive in that specific API.
For users who need maximum performance in modern DX12 titles, the RTX 3090 is the clear choice. For users who work primarily with Vulkan compute or who value relative standing in aggregate scores, the P104-100 holds up better than expected. The P104-100's lack of display outputs means it cannot serve as a traditional graphics card, while the RTX 3090 is a full-featured consumer GPU.
Where Each One Wins
The RTX 3090 wins in all three direct benchmark comparisons. Its biggest victory is in 3DMark Steel Nomad DX12, where it leads by 72.4%. This test likely benefits from the RTX 3090's RT cores, tensor cores, and higher memory bandwidth. The RTX 3090 also excels in Geekbench OpenCL, leading by 69.7%, which reflects its massive shader count and memory throughput. In Geekbench Vulkan, the RTX 3090 wins by a smaller 16.2% margin, but it still takes the top spot.
The P104-100's strengths are best understood through its aggregate numbers. It holds a higher average benchmark score (32,982 versus 27,565) and a higher percentile ranking (77 versus 73). Its nearest rivals are mobile and workstation GPUs like the NVIDIA T600 Mobile, T550 Mobile, and RTX 3050 Mobile, with deltas under 1%. This places it in a different performance tier than the RTX 3090, whose nearest rivals include the RTX 4070 Mobile and RX 6700 XT.
In terms of use cases, the RTX 3090 is suited for demanding DX12 gaming, ray tracing workloads, and compute tasks that leverage its tensor cores. The P104-100, given its mining origins and lack of display outputs, is suited for compute-only environments where Vulkan is the primary API. Its lower power requirements (200 W suggested PSU versus 750 W) and dual-slot design make it easier to deploy in dense systems, though its PCIe 1.0 x4 interface may bottleneck data transfer in some scenarios. The RTX 3090's PCIe 4.0 x16 interface is far more capable for modern systems.