NVIDIA GeForce RTX 3090 vs NVIDIA P104-100 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 3090

CORE STATE GA102
VRAM 24 GB
CLOCK SPEED 1695 MHz
TDP 350 W
BUS WIDTH 384 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2020
VS
NVIDIA
GEFORCE

P104-100

CORE STATE GP104
VRAM 4 GB
CLOCK SPEED 1733 MHz
TDP
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2017

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
5,118
1,413
geekbench_opencl
172,758
52,368
geekbench_vulkan
53,927
45,165
passmark_directx_10
182
N/A
passmark_directx_11
220
N/A
passmark_directx_12
110
N/A
passmark_directx_9
268
N/A
passmark_g2d
1,063
N/A
passmark_g3d
26,645
N/A
passmark_gpu_compute
15,356
N/A

Analysis: NVIDIA GeForce RTX 3090 vs NVIDIA P104-100

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA P104-100 has a higher average benchmark score at 32,982, compared to the NVIDIA GeForce RTX 3090 at 27,565. However, this aggregate figure is heavily influenced by the specific test sets available for each card.

Q: How do the two cards compare in the 3DMark Steel Nomad DX12 test?

A: The RTX 3090 is the clear winner, scoring 5,118 against the P104-100's 1,413. This represents a 72.4% advantage for the RTX 3090 in this test.

Q: What is the difference in compute performance between the two cards?

A: In Geekbench OpenCL, the RTX 3090 scores 172,758, while the P104-100 scores 52,368. The RTX 3090 is 69.7% ahead in this test.

Q: Are the two cards close in any benchmark?

A: The closest margin is in Geekbench Vulkan, where the RTX 3090 scores 53,927 versus the P104-100's 45,165. Here, the RTX 3090 leads by only 16.2%.

Q: Which card has a higher percentile ranking among all GPUs?

A: The P104-100 sits at the 77th percentile, while the RTX 3090 is at the 73rd percentile. This suggests the P104-100 outperforms a larger share of the GPU population in the database's aggregate ranking.

Q: What is the fabrication process difference between the two?

A: The P104-100 uses a 16 nm process at TSMC, while the RTX 3090 uses an 8 nm process at Samsung. The RTX 3090 also packs significantly more transistors: 28,300 million versus 7,200 million.

Architecture Differences

The two GPUs come from different NVIDIA generations and are built for entirely different purposes. The P104-100 is part of the Mining GPUs generation, based on the Pascal architecture with the GP104 chip. It was released in December 2017. The RTX 3090 belongs to the GeForce 30 series, uses the Ampere architecture with the GA102 chip, and launched in August 2020. The architectural gap spans roughly three years of development.

The process nodes differ substantially. The P104-100 is fabricated on a 16 nm process at TSMC, which yields a transistor density of 22.9 million transistors per square millimeter. The RTX 3090 moves to an 8 nm process at Samsung, achieving a much higher density of 45.1 million transistors per square millimeter. This allows the RTX 3090 to house 28,300 million transistors on a 628 mm² die, compared to the P104-100's 7,200 million on a 314 mm² die.

The RTX 3090 introduces dedicated hardware that the Pascal-based P104-100 lacks entirely. The RTX 3090 includes 82 RT cores for ray tracing and 328 tensor cores for AI acceleration. The P104-100 has no such units listed in its specifications. This is a fundamental architectural difference: one is a compute-focused mining card, the other is a full-featured consumer flagship with specialized processing blocks.

Memory architecture also diverges. The P104-100 uses 4 GB of GDDR5X on a 256-bit bus, delivering 320.3 GB/s of bandwidth. The RTX 3090 uses 24 GB of GDDR6X on a 384-bit bus, providing 936.2 GB/s. The RTX 3090's memory clock is 1219 MHz with 19.5 Gbps effective speed, while the P104-100's memory runs at 1251 MHz with 10 Gbps effective. The RTX 3090's advantage in bandwidth is roughly triple that of the P104-100.

Compute capabilities differ sharply. The RTX 3090 has 10,496 shading units, 328 TMUs, and 112 ROPs. The P104-100 has 1,920 shading units, 120 TMUs, and 64 ROPs. The RTX 3090 also has a major FP16 advantage: it delivers 35.58 TFLOPS with a 1:1 ratio, while the P104-100 manages only 104.0 GFLOPS with a 1:64 ratio. FP32 performance is 35.58 TFLOPS versus 6.655 TFLOPS in favor of the RTX 3090.

Head-to-Head Benchmarks

The head-to-head results show a consistent sweep for the RTX 3090, winning all three recorded tests. The margin varies considerably by workload, which is informative for understanding where each card's strengths lie.

In 3DMark Steel Nomad DX12, the RTX 3090 scores 5,118 against the P104-100's 1,413. The delta is 72.4% in favor of the RTX 3090. This is a modern DX12 workload that stresses the full feature set of the newer architecture, including the RT and tensor cores. The P104-100, despite being a capable card in its own right, falls far behind here.

In Geekbench OpenCL, the gap is similar: the RTX 3090 posts 172,758 versus 52,368, a 69.7% advantage. OpenCL compute performance is heavily dependent on raw shader throughput and memory bandwidth, both of which favor the RTX 3090 by wide margins. The RTX 3090 has over five times the shading units and nearly three times the memory bandwidth.

The closest contest is Geekbench Vulkan. Here the RTX 3090 scores 53,927, while the P104-100 scores 45,165. The RTX 3090's lead narrows to 16.2%. This suggests that the P104-100's Pascal architecture is relatively efficient in Vulkan workloads, closing the gap substantially compared to DX12 or OpenCL. It also indicates that Vulkan may not fully utilize the RTX 3090's extra compute resources, or that driver overhead plays a role in this specific API.

The overall win count is 3 to 0 in favor of the RTX 3090. The P104-100 does not win any benchmark in the direct comparison. However, the magnitude of the losses varies dramatically: from a 72.4% deficit in Steel Nomad to a much more competitive 16.2% deficit in Vulkan. This is a meaningful distinction for users who prioritize specific APIs.

Specification Differences

The specification table shows differences across nearly every major category. Process node: 16 nm for the P104-100 versus 8 nm for the RTX 3090. Foundry: TSMC versus Samsung. Transistors: 7,200 million versus 28,300 million. Die size: 314 mm² versus 628 mm². Transistor density: 22.9M per mm² versus 45.1M per mm².

Clock speeds: the P104-100 has a base clock of 1607 MHz and a boost of 1733 MHz. The RTX 3090 has a lower base of 1395 MHz but a boost of 1695 MHz. Memory clocks differ as well: 1251 MHz with 10 Gbps effective for the P104-100, versus 1219 MHz with 19.5 Gbps effective for the RTX 3090.

Memory configuration: 4 GB versus 24 GB, GDDR5X versus GDDR6X, 256-bit versus 384-bit bus, and 320.3 GB/s versus 936.2 GB/s bandwidth. The RTX 3090 has substantially more capacity and bandwidth.

Compute units: 1,920 shading units versus 10,496, 120 TMUs versus 328, and 64 ROPs versus 112. The RTX 3090 adds 82 RT cores and 328 tensor cores, which the P104-100 does not have. Pixel rate is 110.9 GPixel/s versus 189.8 GPixel/s. Texture rate is 208.0 GTexel/s versus 556.0 GTexel/s. FP32 is 6.655 TFLOPS versus 35.58 TFLOPS. FP16 is 104.0 GFLOPS versus 35.58 TFLOPS.

Power and physical specs: the P104-100 has no listed TDP, while the RTX 3090 has a 350 W TDP. The P104-100 is dual-slot with a 1x 8-pin connector and a 200 W suggested PSU. The RTX 3090 is triple-slot with a 1x 12-pin connector and a 750 W suggested PSU. The P104-100 is 267 mm long, while the RTX 3090 is 336 mm long, 140 mm high, and 61 mm wide. The P104-100 has no display outputs, while the RTX 3090 has 1x HDMI 2.1 and 3x DisplayPort 1.4a.

Bus interface: the P104-100 uses PCIe 1.0 x4, while the RTX 3090 uses PCIe 4.0 x16. DirectX support: 12 (12_1) versus 12 Ultimate (12_2). OpenGL is 4.6 for both, and Vulkan is 1.4 for both.

The Verdict

The data points to a straightforward conclusion for raw performance: the RTX 3090 is the dominant card in every head-to-head benchmark. It wins all three tests, with margins ranging from 16.2% to 72.4%. The RTX 3090 also brings modern features like RT cores, tensor cores, and a much larger memory pool. Its 24 GB of GDDR6X memory and 936.2 GB/s bandwidth are decisive for memory-heavy workloads.

However, the P104-100 has its own story. Its average benchmark score of 32,982 is higher than the RTX 3090's 27,565, and its percentile ranking is 77th versus 73rd. This suggests that in the database's aggregate scoring, the P104-100 performs better relative to other GPUs. This is likely because the P104-100's benchmark set is limited to three tests, two of which are compute-oriented, while the RTX 3090's set includes older DirectX 9, 10, and 11 tests where it scores low. The P104-100's Vulkan score of 45,165 is within 16.2% of the RTX 3090, making it competitive in that specific API.

For users who need maximum performance in modern DX12 titles, the RTX 3090 is the clear choice. For users who work primarily with Vulkan compute or who value relative standing in aggregate scores, the P104-100 holds up better than expected. The P104-100's lack of display outputs means it cannot serve as a traditional graphics card, while the RTX 3090 is a full-featured consumer GPU.

Where Each One Wins

The RTX 3090 wins in all three direct benchmark comparisons. Its biggest victory is in 3DMark Steel Nomad DX12, where it leads by 72.4%. This test likely benefits from the RTX 3090's RT cores, tensor cores, and higher memory bandwidth. The RTX 3090 also excels in Geekbench OpenCL, leading by 69.7%, which reflects its massive shader count and memory throughput. In Geekbench Vulkan, the RTX 3090 wins by a smaller 16.2% margin, but it still takes the top spot.

The P104-100's strengths are best understood through its aggregate numbers. It holds a higher average benchmark score (32,982 versus 27,565) and a higher percentile ranking (77 versus 73). Its nearest rivals are mobile and workstation GPUs like the NVIDIA T600 Mobile, T550 Mobile, and RTX 3050 Mobile, with deltas under 1%. This places it in a different performance tier than the RTX 3090, whose nearest rivals include the RTX 4070 Mobile and RX 6700 XT.

In terms of use cases, the RTX 3090 is suited for demanding DX12 gaming, ray tracing workloads, and compute tasks that leverage its tensor cores. The P104-100, given its mining origins and lack of display outputs, is suited for compute-only environments where Vulkan is the primary API. Its lower power requirements (200 W suggested PSU versus 750 W) and dual-slot design make it easier to deploy in dense systems, though its PCIe 1.0 x4 interface may bottleneck data transfer in some scenarios. The RTX 3090's PCIe 4.0 x16 interface is far more capable for modern systems.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 3090
P104-100
Core Specs
Shading Units
10,496
1,920 -81.7%
Shaders
10,496
1,920 -81.7%
TMUs
328
120 -63.4%
ROPs
112
64 -42.9%
SM Count
82
15 -81.7%
Clocks
Base Clock
1395 MHz
1607 MHz
Boost Clock
1695 MHz
1733 MHz
Memory Clock
1219 MHz 19.5 Gbps effective
1251 MHz 10 Gbps effective
Memory
Memory Size
24 GB
4 GB
VRAM (MB)
24,576
4,096 -83.3%
Memory Type
GDDR6X
GDDR5X
Memory Bus
384 bit
256 bit
Bandwidth
936.2 GB/s
320.3 GB/s
Cache
L1 Cache
128 KB (per SM)
48 KB (per SM)
L2 Cache
6 MB
2 MB
Performance
Pixel Rate
189.8 GPixel/s
110.9 GPixel/s
Texture Rate
556.0 GTexel/s
208.0 GTexel/s
FP32 (TFLOPS)
35.58 TFLOPS
6.655 TFLOPS
FP64 (TFLOPS)
556.0 GFLOPS (1:64)
208.0 GFLOPS (1:32)
FP16 (TFLOPS)
35.58 TFLOPS (1:1)
104.0 GFLOPS (1:64)
AI/RT
RT Cores
82
Tensor Cores
328
Power
TDP
350 W
TDP (W)
350
Suggested PSU
750 W
200 W
Power Connectors
1x 12-pin
1x 8-pin
Architecture
Architecture
Ampere
Pascal
GPU Name
GA102
GP104
Generation
GeForce 30
Mining GPUs
Process Size
8 nm
16 nm
Transistors
28,300 million
7,200 million
Die Size
628 mm²
314 mm²
Foundry
Samsung
TSMC
Density
45.1M / mm²
22.9M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Triple-slot
Dual-slot
Length
336 mm 13.2 inches
267 mm 10.5 inches
Height
140 mm 5.5 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 1.0 x4
Other
Launch Price
1,499 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 20
Successor
GeForce 40
View GeForce RTX 3090 Details View P104-100 Details