NVIDIA GeForce RTX 3070 Ti vs NVIDIA P104-100 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 3070 Ti

CORE STATE GA104
VRAM 8 GB
CLOCK SPEED 1770 MHz
TDP 290 W
BUS WIDTH 256 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

P104-100

CORE STATE GP104
VRAM 4 GB
CLOCK SPEED 1733 MHz
TDP
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2017

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
3,478
1,413
geekbench_opencl
119,718
52,368
geekbench_vulkan
139,541
45,165
passmark_directx_10
155
N/A
passmark_directx_11
192
N/A
passmark_directx_12
91
N/A
passmark_directx_9
261
N/A
passmark_g2d
1,055
N/A
passmark_g3d
23,356
N/A
passmark_gpu_compute
11,601
N/A

Analysis: NVIDIA GeForce RTX 3070 Ti vs NVIDIA P104-100

The NVIDIA P104-100 and the NVIDIA GeForce RTX 3070 Ti represent two very different moments in NVIDIA’s GPU lineup. The P104-100 is a mining-specific board built on the Pascal architecture, while the RTX 3070 Ti is a mainstream gaming card from the Ampere generation. The recorded benchmark data shows a decisive performance gap, but the two cards also differ fundamentally in their intended roles, feature sets, and system requirements. This analysis separates the raw numbers from the architectural context to explain what each card actually offers.

Where Each One Wins

The data is unambiguous in one respect: the RTX 3070 Ti wins every single head-to-head benchmark recorded in the database. Out of three comparative tests, the RTX 3070 Ti takes all three, with the P104-100 recording zero wins. However, that does not mean the P104-100 has no purpose. Its win condition is not performance, but rather the specific niche it was built for: a mining GPU with no display outputs, a PCIe 1.0 x4 interface, and a design that prioritizes compute throughput over graphics features.

In the benchmark results, the P104-100 shows a 77th percentile ranking against all GPUs, while the RTX 3070 Ti sits at the 75th percentile. That is an unusual inversion: the slower card in direct comparison actually has a higher percentile rank in the overall database. The reason is that the P104-100’s average benchmark score of 32,982 is higher than the RTX 3070 Ti’s average of 29,945. This is because the P104-100 is only measured on two Geekbench tests (OpenCL and Vulkan), where it scores 52,368 and 45,165 respectively, plus a single 3DMark Steel Nomad DX12 score of 1,413. The RTX 3070 Ti, in contrast, has a much larger benchmark suite including multiple Passmark tests, and its lower-scoring DirectX tests (like Passmark DirectX 9 at 261 and Passmark DirectX 10 at 155) drag its average down.

The RTX 3070 Ti wins in every workload category that matters for modern gaming and compute. In 3DMark Steel Nomad DX12, it achieves 3,478 versus the P104-100’s 1,413, a delta of -59.4%. In Geekbench OpenCL, the RTX 3070 Ti scores 119,718 against 52,368, a -56.3% delta. In Geekbench Vulkan, it hits 139,541 versus 45,165, a -67.6% delta. Those are massive margins, all exceeding 55%.

But the P104-100 does have one clear advantage in the data: its nearest rivals are all low-power mobile or workstation GPUs. The database shows the P104-100 sitting within 0.7% of the AMD Radeon Pro 570, the NVIDIA GeForce RTX 3050 Mobile, the NVIDIA T550 Mobile, and the NVIDIA T600 Mobile. That places it in a completely different competitive tier than the RTX 3070 Ti, whose nearest rivals include the AMD Radeon RX 6800, the NVIDIA GeForce RTX 2080 Ti, the NVIDIA GeForce RTX 5070 Mobile, and the AMD Radeon RX 6700. The RTX 3070 Ti is competing with high-end desktop parts from both AMD and NVIDIA, while the P104-100 is in the range of entry-level laptop GPUs.

The Verdict

The verdict from the recorded data is clear: the NVIDIA GeForce RTX 3070 Ti is the superior card for any application that relies on DirectX 12, Vulkan, or OpenCL performance. Its three benchmark wins are not narrow margins; they are dominant leads of 56% to 68%. The RTX 3070 Ti also brings a full feature set that the P104-100 lacks entirely: 48 ray tracing cores, 192 tensor cores, DirectX 12 Ultimate support, and display outputs (1x HDMI 2.1 and 3x DisplayPort 1.4a). The P104-100 has no display outputs, no ray tracing cores, no tensor cores, and only supports DirectX 12 (12_1), which lacks the full Ultimate feature set.

Who should pick the P104-100? The data suggests only someone who is building a system specifically for mining or compute tasks that ignore graphics output entirely. It is a dual-slot card with a 200 W suggested PSU, a 1x 8-pin power connector, and a PCIe 1.0 x4 bus interface. That PCIe interface is a major limitation: even though the card has a 256-bit memory bus and 320.3 GB/s of bandwidth, the x4 interface will bottleneck any data transfer to the host system. The RTX 3070 Ti, by contrast, uses PCIe 4.0 x16, which is the modern standard for high-throughput communication.

Who should pick the RTX 3070 Ti? Anyone who needs a graphics card for gaming, content creation, or any workload that benefits from DirectX 12 Ultimate features, ray tracing, or tensor core acceleration. The RTX 3070 Ti has 8 GB of GDDR6X memory versus the P104-100’s 4 GB of GDDR5X, and its memory bandwidth is 608.3 GB/s versus 320.3 GB/s. The RTX 3070 Ti also has a much higher FP32 throughput: 21.75 TFLOPS versus 6.655 TFLOPS. It also offers FP16 at a 1:1 ratio (21.75 TFLOPS), while the P104-100’s FP16 is a minuscule 104.0 GFLOPS at a 1:64 ratio.

Head-to-Head Benchmarks

The three head-to-head tests in the database all favor the RTX 3070 Ti by a wide margin. Starting with 3DMark Steel Nomad DX12, the RTX 3070 Ti scores 3,478, while the P104-100 manages only 1,413. The delta is -59.4%, meaning the RTX 3070 Ti is roughly 2.46 times faster in this test. This is a modern DirectX 12 workload, and the result reflects not just raw compute power but also architectural efficiency. The P104-100’s Pascal architecture lacks dedicated ray tracing hardware and uses a much older scheduling design.

In Geekbench OpenCL, the RTX 3070 Ti scores 119,718 versus 52,368 for the P104-100, a delta of -56.3%. OpenCL is a compute-oriented API, and this result shows the RTX 3070 Ti delivering over twice the compute performance. The RTX 3070 Ti has 6,144 shading units versus 1,920 on the P104-100, and 192 texture mapping units versus 120. The P104-100 does have 64 ROPs, but the RTX 3070 Ti has 96. Every major compute resource is larger on the RTX 3070 Ti.

The largest delta appears in Geekbench Vulkan, where the RTX 3070 Ti scores 139,541 and the P104-100 scores 45,165, a -67.6% difference. That is the single biggest gap in the entire comparison. Vulkan is a low-level API that can expose architectural inefficiencies, and here the Ampere architecture shows a clear advantage. The RTX 3070 Ti’s 21.75 TFLOPS of FP32 performance and its 339.8 GTexel/s texture rate dwarf the P104-100’s 6.655 TFLOPS and 208.0 GTexel/s. The pixel rates also differ: 169.9 GPixel/s for the RTX 3070 Ti versus 110.9 GPixel/s for the P104-100.

FAQ

Q: Which GPU has the higher average benchmark score in the database?

A: The NVIDIA P104-100 has a higher average benchmark score of 32,982, compared to the NVIDIA GeForce RTX 3070 Ti’s 29,945. However, the P104-100 is only tested on three benchmarks, while the RTX 3070 Ti has ten recorded results, including several low-scoring Passmark DirectX tests that pull its average down.

Q: Why does the P104-100 have a higher percentile rank (77) than the RTX 3070 Ti (75)?

A: The percentile rank is based on the average benchmark score across all tested GPUs. Because the P104-100’s limited benchmark set (3DMark Steel Nomad, Geekbench OpenCL, Geekbench Vulkan) all score relatively high, its average is higher than the RTX 3070 Ti’s, which includes Passmark DirectX 9 and DirectX 10 scores of 261 and 155 respectively.

Q: What is the biggest performance gap between the two cards?

A: The largest delta is in Geekbench Vulkan, where the RTX 3070 Ti scores 139,541 versus the P104-100’s 45,165, a difference of -67.6%. The RTX 3070 Ti is more than three times faster in that test.

Q: Does the P104-100 support ray tracing or tensor cores?

A: No. The P104-100 has no ray tracing cores and no tensor cores. The RTX 3070 Ti has 48 ray tracing cores and 192 tensor cores.

Q: What are the memory specifications for each card?

A: The P104-100 has 4 GB of GDDR5X on a 256-bit bus with 320.3 GB/s bandwidth. The RTX 3070 Ti has 8 GB of GDDR6X on a 256-bit bus with 608.3 GB/s bandwidth.

Q: Can the P104-100 be used for display output?

A: No. The P104-100 has no display outputs. The RTX 3070 Ti has 1x HDMI 2.1 and 3x DisplayPort 1.4a.

Architecture Differences

The two GPUs come from completely different architectural eras. The P104-100 uses the GP104 chip built on Pascal architecture, manufactured on a 16 nm process at TSMC. The RTX 3070 Ti uses the GA104 chip built on Ampere architecture, manufactured on an 8 nm process at Samsung. The transistor counts reflect the generational leap: the P104-100 has 7,200 million transistors on a 314 mm² die, while the RTX 3070 Ti has 17,400 million transistors on a 392 mm² die. The transistor density also differs, with the P104-100 at 22.9M per mm² and the RTX 3070 Ti at 44.4M per mm².

The feature set diverges sharply. The RTX 3070 Ti includes 48 ray tracing cores and 192 tensor cores, both of which are absent from the P104-100. This directly affects API support: the RTX 3070 Ti lists DirectX 12 Ultimate (12_2), while the P104-100 only supports DirectX 12 (12_1). The P104-100 also lacks any FP16 capability beyond a token 104.0 GFLOPS at a 1:64 ratio, whereas the RTX 3070 Ti delivers 21.75 TFLOPS of FP16 at a full 1:1 ratio. That makes the RTX 3070 Ti vastly more capable for machine learning inference and any workload using half-precision math.

The memory subsystems also differ. The P104-100 uses GDDR5X at 10 Gbps effective, while the RTX 3070 Ti uses GDDR6X at 19 Gbps effective. The result is a doubling of memory bandwidth: 320.3 GB/s versus 608.3 GB/s. Both cards have a 256-bit memory bus, but the newer memory technology gives the RTX 3070 Ti a decisive advantage in bandwidth-bound workloads.

Specification Differences

The recorded specifications show several key differences beyond the obvious performance gap. The RTX 3070 Ti has a much larger compute configuration: 6,144 shading units versus 1,920, 192 TMUs versus 120, and 96 ROPs versus 64. The clock speeds are similar, with the P104-100 boosting to 1733 MHz and the RTX 3070 Ti boosting to 1770 MHz, but the RTX 3070 Ti’s base clock is lower at 1575 MHz versus 1607 MHz. The FP32 throughput is over three times higher on the RTX 3070 Ti: 21.75 TFLOPS versus 6.655 TFLOPS.

The power requirements are drastically different. The P104-100 has no TDP listed in the database, but its suggested PSU is 200 W, and it uses a single 8-pin power connector. The RTX 3070 Ti has a TDP of 290 W, a suggested PSU of 600 W, and uses a single 12-pin power connector. The bus interfaces also differ: the P104-100 uses PCIe 1.0 x4, while the RTX 3070 Ti uses PCIe 4.0 x16. The P104-100 has no display outputs, while the RTX 3070 Ti offers 1x HDMI 2.1 and 3x DisplayPort 1.4a.

Both cards are dual-slot and share the same length of 267 mm (10.5 inches). The RTX 3070 Ti also has a recorded height of 112 mm (4.4 inches), while the P104-100’s height is not listed. The P104-100 has 4 GB of GDDR5X memory, while the RTX 3070 Ti has 8 GB of GDDR6X. The memory clocks are 1251 MHz for the P104-100 and 1188 MHz for the RTX 3070 Ti, but the effective data rates differ significantly (10 Gbps versus 19 Gbps). Both support OpenGL 4.6 and Vulkan 1.4, but the RTX 3070 Ti adds DirectX 12 Ultimate support. The P104-100 was released in December 2017, while the RTX 3070 Ti was released in May 2021. The RTX 3070 Ti has a recorded launch MSRP of 599 USD.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 3070 Ti
P104-100
Core Specs
Shading Units
6,144
1,920 -68.8%
Shaders
6,144
1,920 -68.8%
TMUs
192
120 -37.5%
ROPs
96
64 -33.3%
SM Count
48
15 -68.8%
Clocks
Base Clock
1575 MHz
1607 MHz
Boost Clock
1770 MHz
1733 MHz
Memory Clock
1188 MHz 19 Gbps effective
1251 MHz 10 Gbps effective
Memory
Memory Size
8 GB
4 GB
VRAM (MB)
8,192
4,096 -50.0%
Memory Type
GDDR6X
GDDR5X
Memory Bus
256 bit
256 bit
Bandwidth
608.3 GB/s
320.3 GB/s
Cache
L1 Cache
128 KB (per SM)
48 KB (per SM)
L2 Cache
4 MB
2 MB
Performance
Pixel Rate
169.9 GPixel/s
110.9 GPixel/s
Texture Rate
339.8 GTexel/s
208.0 GTexel/s
FP32 (TFLOPS)
21.75 TFLOPS
6.655 TFLOPS
FP64 (TFLOPS)
339.8 GFLOPS (1:64)
208.0 GFLOPS (1:32)
FP16 (TFLOPS)
21.75 TFLOPS (1:1)
104.0 GFLOPS (1:64)
AI/RT
RT Cores
48
Tensor Cores
192
Power
TDP
290 W
TDP (W)
290
Suggested PSU
600 W
200 W
Power Connectors
1x 12-pin
1x 8-pin
Architecture
Architecture
Ampere
Pascal
GPU Name
GA104
GP104
Generation
GeForce 30
Mining GPUs
Process Size
8 nm
16 nm
Transistors
17,400 million
7,200 million
Die Size
392 mm²
314 mm²
Foundry
Samsung
TSMC
Density
44.4M / mm²
22.9M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
112 mm 4.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 1.0 x4
Other
Launch Price
599 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 20
Successor
GeForce 40
View GeForce RTX 3070 Ti Details View P104-100 Details