NVIDIA GeForce RTX 4070 vs NVIDIA P104-100 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 4070

CORE STATE AD104
VRAM 12 GB
CLOCK SPEED 2475 MHz
TDP 200 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

P104-100

CORE STATE GP104
VRAM 4 GB
CLOCK SPEED 1733 MHz
TDP
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2017

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
3,854
1,413
geekbench_opencl
154,858
52,368
geekbench_vulkan
174,152
45,165
passmark_directx_10
139
N/A
passmark_directx_11
244
N/A
passmark_directx_12
103
N/A
passmark_directx_9
320
N/A
passmark_g2d
1,164
N/A
passmark_g3d
26,927
N/A
passmark_gpu_compute
14,720
N/A

Analysis: NVIDIA GeForce RTX 4070 vs NVIDIA P104-100

Head-to-Head Benchmarks

The benchmark data presents a decisive picture. The NVIDIA GeForce RTX 4070 wins all three recorded head-to-head tests against the NVIDIA P104-100, and the margins are substantial. In the 3DMark Steel Nomad DX12 test, the RTX 4070 scores 3854 against the P104-100's 1413, a 172.8% advantage. This is not a close contest; the modern architecture simply outclasses the older mining-oriented card in a demanding modern API workload.

The lead expands further in compute-oriented tests. In Geekbench OpenCL, the RTX 4070 records 154858 points versus 52368 for the P104-100, a delta of 195.7%. The Vulkan result is even more lopsided: the RTX 4070 scores 174152, while the P104-100 manages only 45165, giving the newer card a 285.6% lead. That Vulkan margin, nearly three times the score, highlights how far the older Pascal architecture has fallen behind in API efficiency and raw throughput.

Looking at the broader database averages reinforces the gap. The RTX 4070 carries an average benchmark score of 37648, placing it in the 81st percentile of all GPUs. The P104-100's average is 32982, which places it in the 77th percentile. The percentage difference in average scores is smaller than in the head-to-head tests, suggesting the P104-100 holds up better in some older or less demanding workloads. However, the recorded head-to-head data shows no test where the P104-100 wins.

The nearest rivals for each card provide additional context. The RTX 4070 sits within 1.3% of the NVIDIA GeForce RTX 4080 Mobile (which scores 38135) and is 0.1% ahead of the NVIDIA Tesla P4 (37628). It also leads the AMD Radeon RX Vega 56 (37507) by 0.4% and the AMD Radeon PRO W6400 (37157) by 1.3%. These are tight margins, meaning the RTX 4070 is positioned in a dense performance cluster. The P104-100, by contrast, is within 0.7% of the AMD Radeon Pro 570 (33207), 0.6% behind the NVIDIA GeForce RTX 3050 Mobile (33170), 0.5% behind the NVIDIA T550 Mobile (33161), and 0.4% ahead of the NVIDIA T600 Mobile (32849). Its performance neighborhood is entirely different, populated by mobile and low-power parts.

Where Each One Wins

The RTX 4070 wins every workload category represented in the head-to-head data: DX12 gaming via 3DMark Steel Nomad, OpenCL compute, and Vulkan rendering. Its 172.8% lead in DX12 suggests it is the clear choice for modern gaming at high settings, especially in titles that leverage ray tracing or DX12 Ultimate features. The compute wins, with a 195.7% margin in OpenCL and a 285.6% margin in Vulkan, indicate broad superiority in GPU-accelerated workloads such as rendering, simulation, and machine learning inference.

The P104-100, on the other hand, does not win any recorded benchmark. Its role in the database is defined by its design as a mining GPU. It has no display outputs, which means it cannot drive a monitor directly. Its strengths, if any, would lie in sustained compute tasks that do not require a display. However, the measured data does not support a performance advantage in any tested scenario. The RTX 4070 is simply faster in every metric captured.

For use-case planning, the data suggests the RTX 4070 is a versatile card. It covers gaming, general compute, and API-specific workloads with high scores. The P104-100 is a niche product that, based on the recorded benchmarks, cannot match the RTX 4070 in any tested area. Its only practical advantage is its PCIe 1.0 x4 interface, which is not a performance benefit but a compatibility note for systems with older or limited bus connectivity. Users with a P104-100 would likely see significant gains by upgrading to the RTX 4070, as the data shows a 172.8% to 285.6% improvement depending on the workload.

FAQ

Q: How much faster is the RTX 4070 in 3DMark Steel Nomad DX12?

A: The RTX 4070 scores 3854, while the P104-100 scores 1413. That is a 172.8% higher result for the RTX 4070.

Q: Which card has the higher average benchmark score in the database?

A: The RTX 4070 has an average benchmark score of 37648, compared to 32982 for the P104-100. The RTX 4070 also sits in the 81st percentile of all GPUs, versus the 77th percentile for the P104-100.

Q: What is the largest performance gap between the two cards?

A: The largest recorded gap is in Geekbench Vulkan, where the RTX 4070 scores 174152 against 45165 for the P104-100, a 285.6% difference.

Q: Does the P104-100 win any benchmark in the head-to-head data?

A: No. The RTX 4070 wins all three head-to-head tests: 3DMark Steel Nomad DX12, Geekbench OpenCL, and Geekbench Vulkan.

Q: Can the P104-100 be used for display output?

A: No. The P104-100 has no display outputs, which is consistent with its design as a mining GPU. The RTX 4070 offers 1x HDMI 2.1 and 3x DisplayPort 1.4a outputs.

Q: How does the RTX 4070 compare to its nearest rivals in the database?

A: The RTX 4070 is 0.1% ahead of the NVIDIA Tesla P4, 0.4% ahead of the AMD Radeon RX Vega 56, 1.3% ahead of the AMD Radeon PRO W6400, and 1.3% behind the NVIDIA GeForce RTX 4080 Mobile.

Specification Differences

The two cards differ in nearly every core specification. The RTX 4070 uses 5888 shading units, 184 texture mapping units, and 64 ROPs. The P104-100 has 1920 shading units, 120 texture mapping units, and 64 ROPs. The RTX 4070 also includes 46 ray tracing cores and 184 tensor cores, while the P104-100 has none of either. The FP32 compute performance is 29.15 TFLOPS for the RTX 4070 versus 6.655 TFLOPS for the P104-100. Texture rate is 455.4 GTexel/s against 208.0 GTexel/s, and pixel rate is 158.4 GPixel/s against 110.9 GPixel/s.

Memory configurations are also starkly different. The RTX 4070 has 12 GB of GDDR6X on a 192-bit bus, delivering 504.2 GB/s of bandwidth. The P104-100 has 4 GB of GDDR5X on a 256-bit bus, delivering 320.3 GB/s. The clock speeds reflect the architectural gap: the RTX 4070 runs at 1920 MHz base and 2475 MHz boost, while the P104-100 runs at 1607 MHz base and 1733 MHz boost. The RTX 4070's memory clock is 1313 MHz (21 Gbps effective), while the P104-100's is 1251 MHz (10 Gbps effective).

The bus interface differs significantly: the RTX 4070 uses PCIe 4.0 x16, while the P104-100 uses PCIe 1.0 x4. The power connectors also differ, with the RTX 4070 requiring a 1x 16-pin connector and a suggested 550 W PSU, while the P104-100 uses a 1x 8-pin connector and a suggested 200 W PSU. The RTX 4070 has a TDP of 200 W; the P104-100 has no recorded TDP. Physical dimensions vary: the RTX 4070 is 240 mm long, 110 mm tall, and 40 mm wide, while the P104-100 is 267 mm long with no recorded height or width. Both are dual-slot cards.

Architecture Differences

The architectural divide is fundamental. The RTX 4070 is built on the Ada Lovelace architecture using the AD104 chip, manufactured on a 5 nm process at TSMC. The P104-100 uses the Pascal architecture with the GP104 chip, manufactured on a 16 nm process, also at TSMC. The transistor counts reflect the process difference: the RTX 4070 packs 35,800 million transistors on a 294 mm² die, for a density of 121.8 million transistors per mm². The P104-100 has 7,200 million transistors on a 314 mm² die, for a density of 22.9 million per mm². The RTX 4070 achieves nearly 5.3 times the transistor density despite having a slightly smaller die.

The memory architecture is a direct consequence of the generational leap. The RTX 4070 uses GDDR6X with a 192-bit bus and 504.2 GB/s of bandwidth. The P104-100 uses GDDR5X with a wider 256-bit bus but only achieves 320.3 GB/s due to the lower effective clock speed. The RTX 4070's FP16 performance is 29.15 TFLOPS at a 1:1 ratio with FP32, while the P104-100 manages only 104.0 GFLOPS at a 1:64 ratio. That makes the RTX 4070 roughly 280 times faster in FP16 compute, which is critical for AI and machine learning workloads that rely on reduced precision.

API support also differs. The RTX 4070 supports DirectX 12 Ultimate (12_2), while the P104-100 is limited to DirectX 12 (12_1). Both support OpenGL 4.6 and Vulkan 1.4. The RTX 4070's ray tracing cores and tensor cores are entirely absent from the P104-100, which explains the massive gap in modern graphics workloads. The P104-100 is a product of the mining GPU generation, with no display outputs and a PCIe 1.0 x4 interface that severely limits data transfer speeds. The RTX 4070, by contrast, is a full-featured consumer GPU with modern display outputs and a PCIe 4.0 x16 connection. The release dates tell the story: the RTX 4070 launched on 2023-04-11, while the P104-100 dates to 2017-12-11. That five-year gap in release timing is reflected in every measurable performance metric.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 4070
P104-100
Core Specs
Shading Units
5,888
1,920 -67.4%
Shaders
5,888
1,920 -67.4%
TMUs
184
120 -34.8%
ROPs
64
64 0.0%
SM Count
46
15 -67.4%
Clocks
Base Clock
1920 MHz
1607 MHz
Boost Clock
2475 MHz
1733 MHz
Memory Clock
1313 MHz 21 Gbps effective
1251 MHz 10 Gbps effective
Memory
Memory Size
12 GB
4 GB
VRAM (MB)
12,288
4,096 -66.7%
Memory Type
GDDR6X
GDDR5X
Memory Bus
192 bit
256 bit
Bandwidth
504.2 GB/s
320.3 GB/s
Cache
L1 Cache
128 KB (per SM)
48 KB (per SM)
L2 Cache
36 MB
2 MB
Performance
Pixel Rate
158.4 GPixel/s
110.9 GPixel/s
Texture Rate
455.4 GTexel/s
208.0 GTexel/s
FP32 (TFLOPS)
29.15 TFLOPS
6.655 TFLOPS
FP64 (TFLOPS)
455.4 GFLOPS (1:64)
208.0 GFLOPS (1:32)
FP16 (TFLOPS)
29.15 TFLOPS (1:1)
104.0 GFLOPS (1:64)
AI/RT
RT Cores
46
Tensor Cores
184
Power
TDP
200 W
TDP (W)
200
Suggested PSU
550 W
200 W
Power Connectors
1x 16-pin
1x 8-pin
Architecture
Architecture
Ada Lovelace
Pascal
GPU Name
AD104
GP104
Generation
GeForce 40
Mining GPUs
Process Size
5 nm
16 nm
Transistors
35,800 million
7,200 million
Die Size
294 mm²
314 mm²
Foundry
TSMC
TSMC
Density
121.8M / mm²
22.9M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
240 mm 9.4 inches
267 mm 10.5 inches
Height
110 mm 4.3 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 1.0 x4
Other
Launch Price
599 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 30
Successor
GeForce 50
View GeForce RTX 4070 Details View P104-100 Details