NVIDIA GeForce RTX 4070 Ti SUPER vs NVIDIA P104-100 Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 4070 Ti SUPER

CORE STATE AD103
VRAM 16 GB
CLOCK SPEED 2610 MHz
TDP 285 W
BUS WIDTH 256 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2024
VS
NVIDIA
GEFORCE

P104-100

CORE STATE GP104
VRAM 4 GB
CLOCK SPEED 1733 MHz
TDP
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2017

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
5,569
1,413
geekbench_opencl
199,267
52,368
geekbench_vulkan
53,683
45,165
passmark_directx_10
181
N/A
passmark_directx_11
278
N/A
passmark_directx_12
119
N/A
passmark_directx_9
360
N/A
passmark_g2d
1,225
N/A
passmark_g3d
31,811
N/A
passmark_gpu_compute
18,372
N/A

Analysis: NVIDIA GeForce RTX 4070 Ti SUPER vs NVIDIA P104-100

The NVIDIA P104-100 and the NVIDIA GeForce RTX 4070 Ti SUPER represent two entirely different eras and purposes for GPUs. The P104-100 is a dedicated mining part from the Pascal generation, while the RTX 4070 Ti SUPER is a modern Ada Lovelace gaming and compute card. The benchmark data shows a decisive performance gap, but the architectural story is just as important for understanding why these cards are not interchangeable.

Head-to-Head Benchmarks

The head-to-head results are not close. Across all three shared tests, the RTX 4070 Ti SUPER wins outright, taking a clean 3-0 sweep. The most dramatic margin appears in the 3DMark Steel Nomad DX12 test, where the RTX 4070 Ti SUPER scores 5569 against the P104-100's 1413. That is a delta of -74.6% for the P104-100, meaning the newer card delivers roughly four times the raw DX12 performance. This is a generational leap that no amount of driver optimization or overclocking on the Pascal card could close.

The Geekbench OpenCL test tells a similar story, though the gap is slightly less extreme. The RTX 4070 Ti SUPER posts 199,267 points, while the P104-100 manages 52,368. The delta of -73.7% shows that even in a compute-heavy, API-agnostic workload, the Ada Lovelace architecture's massive shader count and higher clocks overwhelm the older card's 1920 shading units. The P104-100 is not a slow card by 2017 standards, but it is outclassed here by a factor of nearly four.

The closest contest comes in Geekbench Vulkan, where the RTX 4070 Ti SUPER scores 53,683 versus 45,165 for the P104-100. The delta is only -15.9%, which is still a solid win for the newer card but far narrower than the other tests. This suggests that Vulkan's lower-level overhead and the P104-100's relatively high boost clock of 1733 MHz help it stay more competitive in this specific API. Still, "closest" is relative; a 15.9% deficit is not a moral victory, and the RTX 4070 Ti SUPER remains firmly ahead in every measurable way.

Looking at the broader benchmark averages, the P104-100 actually holds a higher average score of 32,982 compared to the RTX 4070 Ti SUPER's 31,087, but this is misleading. The P104-100's average is drawn from only three tests, while the RTX 4070 Ti SUPER's average includes ten tests, several of which are Passmark DX9/DX10/DX11 legacy workloads where a modern card's drivers may not prioritize optimization. The RTX 4070 Ti SUPER's Passmark G3D score of 31,811 is actually its second-highest result, so the average is pulled down by the older DirectX tests, not by a lack of raw power.

Architecture Differences

The architectural chasm between these two cards is vast. The P104-100 uses the GP104 chip on TSMC's 16 nm process, packing 7,200 million transistors into a 314 mm² die. The RTX 4070 Ti SUPER uses the AD103 chip on a 5 nm process, fitting 45,900 million transistors into a slightly larger 379 mm² die. That is a transistor density jump from 22.9 million per mm² to 121.1 million per mm², which explains how the newer card can deliver so much more performance without a proportional increase in physical size.

The memory subsystems are also radically different. The P104-100 has 4 GB of GDDR5X on a 256-bit bus, yielding 320.3 GB/s of bandwidth. The RTX 4070 Ti SUPER has 16 GB of GDDR6X on the same 256-bit bus, but with 672.3 GB/s of bandwidth. The newer card has four times the capacity and more than double the bandwidth, which is critical for modern textures, high-resolution rendering, and large compute datasets. The P104-100's memory runs at 10 Gbps effective, while the RTX 4070 Ti SUPER's runs at 21 Gbps effective.

Compute resources differ by a similar magnitude. The P104-100 has 1920 shading units, 120 TMUs, and 64 ROPs. The RTX 4070 Ti SUPER has 8448 shading units, 264 TMUs, and 96 ROPs. The newer card also features 66 RT cores and 264 tensor cores, neither of which exist on the P104-100. This is the single most important feature gap: the P104-100 has zero hardware support for ray tracing or AI acceleration, while the RTX 4070 Ti SUPER is built around both. FP32 throughput tells the story: 6.655 TFLOPS for the P104-100 versus 44.10 TFLOPS for the RTX 4070 Ti SUPER. The FP16 ratio is also telling — the P104-100 runs FP16 at 1:64 (104.0 GFLOPS), making it effectively useless for half-precision work, while the RTX 4070 Ti SUPER runs FP16 at 1:1 (44.10 TFLOPS).

The bus interface and display outputs are equally divergent. The P104-100 uses PCIe 1.0 x4, a severe bottleneck for any modern workload, and has no display outputs at all. The RTX 4070 Ti SUPER uses PCIe 4.0 x16 and includes 1x HDMI 2.1 and 3x DisplayPort 1.4a. The P104-100 is a compute-only card by design; it cannot drive a monitor, which alone disqualifies it for most users.

FAQ

Q: Is the NVIDIA P104-100 suitable for gaming?

A: No. The P104-100 has no display outputs, making it impossible to connect to a monitor directly. It was designed for mining, and the data shows it lacks the modern feature set (RT cores, tensor cores) needed for contemporary gaming workloads.

Q: How much faster is the RTX 4070 Ti SUPER in raw compute?

A: In 3DMark Steel Nomad DX12, the RTX 4070 Ti SUPER scores 5569 versus 1413 for the P104-100, a 74.6% advantage. In Geekbench OpenCL, the margin is 73.7% (199,267 versus 52,368). The smallest gap is in Geekbench Vulkan, where the RTX 4070 Ti SUPER leads by 15.9% (53,683 versus 45,165).

Q: Which card has more memory bandwidth?

A: The RTX 4070 Ti SUPER has 672.3 GB/s of bandwidth, while the P104-100 has 320.3 GB/s. The newer card also uses 16 GB of GDDR6X memory versus 4 GB of GDDR5X.

Q: Do these cards support the same DirectX version?

A: No. The P104-100 supports DirectX 12 (12_1), while the RTX 4070 Ti SUPER supports DirectX 12 Ultimate (12_2). The newer card also adds hardware ray tracing and tensor cores, which the P104-100 lacks entirely.

Q: Which card has a higher transistor density?

A: The RTX 4070 Ti SUPER has a density of 121.1 million transistors per mm² on a 5 nm process, compared to the P104-100's 22.9 million per mm² on a 16 nm process. The newer card fits 45,900 million transistors into a 379 mm² die, versus 7,200 million in 314 mm².

Q: Can the P104-100 be used in a modern system?

A: Technically yes, but with severe limitations. It uses PCIe 1.0 x4, which is a fraction of the bandwidth available on modern motherboards. It also requires a 200 W PSU and a single 8-pin connector, but its lack of display outputs makes it useless for any visual task.

The Verdict

The data is unambiguous: the NVIDIA GeForce RTX 4070 Ti SUPER is the superior card in every benchmark where they overlap. It wins all three head-to-head tests, with margins ranging from 15.9% in Vulkan to 74.6% in DX12. It has more than four times the shading units, eight times the memory capacity, double the bandwidth, and exclusive access to RT and tensor cores. The RTX 4070 Ti SUPER also runs at higher clocks (2340 MHz base, 2610 MHz boost versus 1607 MHz base, 1733 MHz boost) and supports modern APIs like DirectX 12 Ultimate.

The P104-100 is not without merit in its original context. Its 77th percentile ranking among all GPUs and average benchmark score of 32,982 place it near the RTX 3050 Mobile and AMD Radeon Pro 570, but those are mobile or workstation parts with far lower power envelopes. The P104-100's nearest rivals include the NVIDIA T600 Mobile (delta of 0.4%) and the T550 Mobile (delta of -0.5%), which suggests it performs like a mid-range laptop GPU from a much later era. That reflects its Pascal efficiency, but it does not change the fact that it is an end-of-life mining card with no display output.

For anyone building a system today, the choice is obvious. The RTX 4070 Ti SUPER is a complete, modern GPU with 16 GB of VRAM, a triple-slot cooler, and a 285 W TDP. It is a capable card for gaming, rendering, and AI workloads, and its launch MSRP was 799 USD. The P104-100 is a historical artifact, useful only for understanding how far GPU technology has progressed in six years.

Specification Differences

The table below highlights only the fields where the two cards differ, based on the FACT PACK data.

| Specification | NVIDIA P104-100 | NVIDIA GeForce RTX 4070 Ti SUPER |

|---|---|---|

| Architecture | Pascal | Ada Lovelace |

| Generation | Mining GPUs | GeForce 40 |

| Process Node | 16 nm | 5 nm |

| Transistors | 7,200 million | 45,900 million |

| Die Size | 314 mm² | 379 mm² |

| Transistor Density | 22.9M / mm² | 121.1M / mm² |

| Base Clock | 1607 MHz | 2340 MHz |

| Boost Clock | 1733 MHz | 2610 MHz |

| Memory Size | 4 GB | 16 GB |

| Memory Type | GDDR5X | GDDR6X |

| Memory Clock | 1251 MHz / 10 Gbps effective | 1313 MHz / 21 Gbps effective |

| Memory Bandwidth | 320.3 GB/s | 672.3 GB/s |

| Shading Units | 1920 | 8448 |

| TMUs | 120 | 264 |

| ROPs | 64 | 96 |

| RT Cores | None | 66 |

| Tensor Cores | None | 264 |

| Pixel Rate | 110.9 GPixel/s | 250.6 GPixel/s |

| Texture Rate | 208.0 GTexel/s | 689.0 GTexel/s |

| FP32 | 6.655 TFLOPS | 44.10 TFLOPS |

| FP16 | 104.0 GFLOPS (1:64) | 44.10 TFLOPS (1:1) |

| TDP | Not specified | 285 W |

| Slot Width | Dual-slot | Triple-slot |

| Power Connectors | 1x 8-pin | 1x 16-pin |

| Suggested PSU | 200 W | 600 W |

| Bus Interface | PCIe 1.0 x4 | PCIe 4.0 x16 |

| Display Outputs | No outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a |

| DirectX | 12 (12_1) | 12 Ultimate (12_2) |

| Dimensions (Length) | 267 mm / 10.5 inches | 310 mm / 12.2 inches |

| Release Date | 2017-12-11 | 2024-01-23 |

| Predecessor | None | GeForce 30 |

| Successor | None | GeForce 50 |

Where Each One Wins

The RTX 4070 Ti SUPER wins in every benchmark tested, but the specific use cases where it excels are clear. For any modern DX12 gaming or ray-traced workload, the 3DMark Steel Nomad result of 5569 versus 1413 shows it is the only viable option. The 66 RT cores and 264 tensor cores make it suitable for real-time ray tracing and DLSS-style AI upscaling, features the P104-100 simply cannot offer. Its 16 GB of memory and 672.3 GB/s bandwidth also make it suitable for 4K texture packs, large 3D scenes, and compute tasks that exceed the 4 GB limit of the older card.

The P104-100 does have one niche where it can still participate: low-overhead Vulkan compute. Its Geekbench Vulkan score of 45,165 is within 15.9% of the RTX 4070 Ti SUPER, which is a respectable showing for a card from 2017. For a developer or researcher who needs a secondary compute device for Vulkan-only workloads and has a spare PCIe x4 slot, the P104-100 could still serve a purpose. Its 77th percentile ranking and average score of 32,982 also place it above the RTX 4070 Ti SUPER's 76th percentile, but that is an artifact of the differing test suites, not a sign of real-world superiority.

In practical terms, the P104-100 wins only in scenarios where its physical characteristics matter: it is shorter (267 mm versus 310 mm), dual-slot instead of triple-slot, and requires a 200 W PSU instead of 600 W. If space and power are absolute constraints and no display output is needed, the P104-100 is the easier card to install. But for anyone who wants to see results on a screen, use ray tracing, or run modern games, the RTX 4070 Ti SUPER is the only real choice. The data does not support any other conclusion.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 4070 Ti SUPER
P104-100
Core Specs
Shading Units
8,448
1,920 -77.3%
Shaders
8,448
1,920 -77.3%
TMUs
264
120 -54.5%
ROPs
96
64 -33.3%
SM Count
66
15 -77.3%
Clocks
Base Clock
2340 MHz
1607 MHz
Boost Clock
2610 MHz
1733 MHz
Memory Clock
1313 MHz 21 Gbps effective
1251 MHz 10 Gbps effective
Memory
Memory Size
16 GB
4 GB
VRAM (MB)
16,384
4,096 -75.0%
Memory Type
GDDR6X
GDDR5X
Memory Bus
256 bit
256 bit
Bandwidth
672.3 GB/s
320.3 GB/s
Cache
L1 Cache
128 KB (per SM)
48 KB (per SM)
L2 Cache
48 MB
2 MB
Performance
Pixel Rate
250.6 GPixel/s
110.9 GPixel/s
Texture Rate
689.0 GTexel/s
208.0 GTexel/s
FP32 (TFLOPS)
44.10 TFLOPS
6.655 TFLOPS
FP64 (TFLOPS)
689.0 GFLOPS (1:64)
208.0 GFLOPS (1:32)
FP16 (TFLOPS)
44.10 TFLOPS (1:1)
104.0 GFLOPS (1:64)
AI/RT
RT Cores
66
Tensor Cores
264
Power
TDP
285 W
TDP (W)
285
Suggested PSU
600 W
200 W
Power Connectors
1x 16-pin
1x 8-pin
Architecture
Architecture
Ada Lovelace
Pascal
GPU Name
AD103
GP104
Generation
GeForce 40
Mining GPUs
Process Size
5 nm
16 nm
Transistors
45,900 million
7,200 million
Die Size
379 mm²
314 mm²
Foundry
TSMC
TSMC
Density
121.1M / mm²
22.9M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
6.1
Shader Model
6.9
6.8
Physical
Slot Width
Triple-slot
Dual-slot
Length
310 mm 12.2 inches
267 mm 10.5 inches
Height
140 mm 5.5 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 1.0 x4
Other
Launch Price
799 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 30
Successor
GeForce 50
View GeForce RTX 4070 Ti SUPER Details View P104-100 Details