NVIDIA GeForce RTX 3090 vs NVIDIA P106-100 Comparison
NVIDIA GeForce RTX 3090
P106-100
PERFORMANCE BENCHMARKS
Analysis: NVIDIA GeForce RTX 3090 vs NVIDIA P106-100
Head-to-Head Benchmarks
The recorded data shows a decisive sweep in favor of the NVIDIA GeForce RTX 3090 across all three shared benchmark tests. The largest margin appears in the 3DMark Steel Nomad DX12 test, where the RTX 3090 scores 5,118 against the P106-100's 899, a delta of 469.3%. This is not a marginal gap; it represents a fundamentally different performance class in modern DirectX 12 workloads.
The Geekbench OpenCL test tells a similar story, with the RTX 3090 reaching 172,758 points versus 35,951 for the P106-100, a 380.5% advantage. This metric reflects general-purpose compute throughput, where the RTX 3090's massive shader array and memory subsystem dominate. The smallest relative win for the RTX 3090 comes in Geekbench Vulkan, where it scores 53,927 against 32,897, a still substantial 63.9% lead. Notably, this is the only test where the P106-100 manages to exceed half of the RTX 3090's score, suggesting that Vulkan's lower-level API overhead narrows the architectural gap somewhat, though it does not come close to closing it.
The aggregate benchmark picture reinforces this hierarchy. The RTX 3090 holds an average benchmark score of 27,565 across all recorded tests, while the P106-100 averages 23,249. The RTX 3090 sits at the 73rd percentile among all GPUs in the database, while the P106-100 sits at the 68th percentile. The nearest rivals for each card illustrate their respective competitive neighborhoods: the RTX 3090 trades blows with the GeForce RTX 4070 Mobile (0.5% ahead) and the AMD Radeon RX 6700 XT (0.5% ahead), while the P106-100 is essentially tied with the AMD Radeon Pro Vega 16 (0% delta) and trails the AMD Radeon RX 6600M by just 0.1%.
Architecture Differences
The two cards come from entirely different architectural generations and design philosophies. The RTX 3090 is built on the Ampere architecture using the GA102 chip, fabricated on an 8 nm process at Samsung. It packs 28,300 million transistors into a 628 mm² die, yielding a transistor density of 45.1 million per square millimeter. The P106-100, by contrast, uses the older Pascal architecture with the GP106 chip, built on a 16 nm process at TSMC. Its 4,400 million transistors occupy a 200 mm² die, with a density of 22.0 million per square millimeter. The RTX 3090 therefore offers more than twice the transistor density and over six times the raw transistor count.
Memory architecture diverges sharply. The RTX 3090 carries 24 GB of GDDR6X on a 384-bit bus, delivering 936.2 GB/s of bandwidth. The P106-100 has 6 GB of GDDR5 on a 192-bit bus, with 192.2 GB/s of bandwidth. That is a nearly fivefold difference in memory bandwidth, which directly impacts texture-heavy and large-dataset workloads. The RTX 3090 also features 82 RT cores and 328 Tensor cores, while the P106-100 has neither, reflecting its origin as a mining-focused product with no ray tracing or AI acceleration hardware.
Compute resources differ by an order of magnitude. The RTX 3090 has 10,496 shading units, 328 TMUs, and 112 ROPs, while the P106-100 has 1,280 shading units, 80 TMUs, and 48 ROPs. The RTX 3090's FP32 throughput is rated at 35.58 TFLOPS, while the P106-100 manages 4.375 TFLOPS. FP16 performance is even more lopsided: the RTX 3090 achieves 35.58 TFLOPS at a 1:1 ratio with FP32, while the P106-100 offers only 68.36 GFLOPS at a 1:64 ratio, a difference of roughly 520 times.
The process node and foundry differences also imply distinct power and thermal profiles. The RTX 3090 is rated at 350 W TDP with a 750 W suggested PSU, while the P106-100 is rated at 120 W with a 300 W suggested PSU. The RTX 3090 uses a 1x 12-pin power connector, while the P106-100 uses a 1x 6-pin. The RTX 3090 is triple-slot and 336 mm long, while the P106-100 is dual-slot and 250 mm long. Critically, the P106-100 has no display outputs at all, confirming its purpose as a compute-only or mining card.
Where Each One Wins
The RTX 3090 wins every recorded workload category, but the magnitude of its wins varies by workload type. In DirectX 12 rasterization (3DMark Steel Nomad), the margin is extreme at 469.3%. In OpenCL compute, the margin is also very large at 380.5%. The Vulkan test shows the closest relative result at 63.9%, which suggests that the P106-100's Pascal architecture retains some efficiency in lower-level API scenarios, even if absolute scores remain far behind.
For use-case analysis, the RTX 3090 is clearly suited for demanding gaming, ray tracing, AI inference, and content creation workloads. Its 24 GB memory capacity and 936.2 GB/s bandwidth can handle large textures, high-resolution rendering, and large model datasets. The presence of RT and Tensor cores enables hardware-accelerated ray tracing and DLSS-style features, which the P106-100 cannot perform at all. The RTX 3090's passmark scores reinforce this: it records 26,645 in G3D and 15,356 in GPU compute, with individual DirectX tests ranging from 110 (DX12) to 268 (DX9). The P106-100 has no such passmark entries in the database, so its gaming-specific behavior cannot be directly measured here, but its lack of display outputs and absence of RT/Tensor cores make it unsuitable for interactive graphics work.
The P106-100, despite its age and limitations, still manages a 68th percentile ranking among all GPUs, which is respectable given its 2017 release and mining-oriented design. Its 1280 shading units and 4.375 TFLOPS of FP32 performance can handle compute tasks that do not require modern API features or large memory footprints. Its 192.2 GB/s bandwidth and 6 GB capacity are sufficient for many embedded or server-side compute workloads, particularly in environments where power draw and physical space are constrained. The dual-slot form factor and 120 W TDP make it significantly easier to deploy in dense systems compared to the RTX 3090's triple-slot, 350 W footprint.
FAQ
Q: How much faster is the RTX 3090 than the P106-100 in DirectX 12?
A: In the 3DMark Steel Nomad DX12 test, the RTX 3090 scores 5,118 compared to the P106-100's 899, a 469.3% advantage.
Q: Does the P106-100 support ray tracing or Tensor cores?
A: No. The P106-100 has no RT cores and no Tensor cores. The RTX 3090 has 82 RT cores and 328 Tensor cores.
Q: What is the memory bandwidth difference between the two cards?
A: The RTX 3090 has 936.2 GB/s of bandwidth from 24 GB of GDDR6X on a 384-bit bus. The P106-100 has 192.2 GB/s from 6 GB of GDDR5 on a 192-bit bus.
Q: Can the P106-100 be used for display output?
A: No. The P106-100 has no display outputs. The RTX 3090 offers 1x HDMI 2.1 and 3x DisplayPort 1.4a.
Q: What is the average benchmark score for each card?
A: The RTX 3090 averages 27,565 across all recorded tests, while the P106-100 averages 23,249. The RTX 3090 sits at the 73rd percentile of all GPUs, the P106-100 at the 68th.
Q: How do the power requirements compare?
A: The RTX 3090 has a 350 W TDP and suggests a 750 W PSU, while the P106-100 has a 120 W TDP and suggests a 300 W PSU.
Specification Differences
| Specification | NVIDIA GeForce RTX 3090 | NVIDIA P106-100 |
|---|---|---|
| Architecture | Ampere | Pascal |
| Chip | GA102 | GP106 |
| Process Node | 8 nm | 16 nm |
| Foundry | Samsung | TSMC |
| Transistors | 28,300 million | 4,400 million |
| Die Size | 628 mm² | 200 mm² |
| Transistor Density | 45.1M / mm² | 22.0M / mm² |
| Base Clock | 1395 MHz | 1506 MHz |
| Boost Clock | 1695 MHz | 1709 MHz |
| Memory Size | 24 GB | 6 GB |
| Memory Type | GDDR6X | GDDR5 |
| Memory Bus | 384 bit | 192 bit |
| Memory Bandwidth | 936.2 GB/s | 192.2 GB/s |
| Shading Units | 10496 | 1280 |
| TMUs | 328 | 80 |
| ROPs | 112 | 48 |
| RT Cores | 82 | None |
| Tensor Cores | 328 | None |
| FP32 Performance | 35.58 TFLOPS | 4.375 TFLOPS |
| FP16 Performance | 35.58 TFLOPS (1:1) | 68.36 GFLOPS (1:64) |
| Pixel Rate | 189.8 GPixel/s | 82.03 GPixel/s |
| Texture Rate | 556.0 GTexel/s | 136.7 GTexel/s |
| TDP | 350 W | 120 W |
| Slot Width | Triple-slot | Dual-slot |
| Power Connectors | 1x 12-pin | 1x 6-pin |
| Suggested PSU | 750 W | 300 W |
| Bus Interface | PCIe 4.0 x16 | PCIe 1.0 x16 |
| Display Outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a | No outputs |
| DirectX Support | 12 Ultimate (12_2) | 12 (12_1) |
| Length | 336 mm | 250 mm |
| Release Date | 2020-08-31 | 2017-06-18 |
| Launch MSRP | 1,499 USD | None |
The Verdict
The data is unambiguous: the RTX 3090 outperforms the P106-100 in every measured category, with margins ranging from 63.9% in Vulkan to 469.3% in DirectX 12. The RTX 3090 is the only choice for any workload requiring modern graphics features, ray tracing, AI acceleration, or high-bandwidth memory. Its 24 GB capacity and 936.2 GB/s bandwidth handle datasets that the P106-100's 6 GB and 192.2 GB/s cannot approach. The RTX 3090's 82 RT cores and 328 Tensor cores open capabilities that the P106-100 simply lacks at the hardware level.
The P106-100, however, retains a niche for compute-only deployments. Its 68th percentile ranking is not trivial, and its 120 W TDP with a 300 W PSU requirement makes it far easier to integrate into power-constrained or space-constrained systems. Its Pascal architecture, while old, still delivers 4.375 TFLOPS of FP32 throughput and supports DirectX 12 (12_1) and Vulkan 1.4, which covers a range of server-side or embedded compute tasks. The absence of display outputs is a non-issue for such roles.
For a buyer choosing between these two, the decision hinges entirely on workload. If the task involves rendering, gaming, ray tracing, AI inference, or any interactive graphics, the RTX 3090 is the only viable option, despite its higher power draw and physical footprint. If the task is purely computational, with no display requirements and strict power or space limits, the P106-100 offers a lower-cost entry point into that space. The benchmark data offers no scenario where the P106-100 wins on performance; its advantages are purely operational, not performance-based.