NVIDIA P106-100 vs NVIDIA RTX A4000 Mobile Comparison

NVIDIA
GEFORCE

NVIDIA P106-100

CORE STATE GP106
VRAM 6 GB
CLOCK SPEED 1709 MHz
TDP 120 W
BUS WIDTH 192 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2017
VS
NVIDIA
GEFORCE

RTX A4000 Mobile

CORE STATE GA104
VRAM 8 GB
CLOCK SPEED 1680 MHz
TDP 115 W
BUS WIDTH 256 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
899
N/A
geekbench_opencl
35,951
97,178
geekbench_vulkan
32,897
73,002
passmark_directx_10
N/A
105
passmark_directx_11
N/A
127
passmark_directx_12
N/A
66
passmark_directx_9
N/A
157
passmark_g2d
N/A
585
passmark_g3d
N/A
14,796
passmark_gpu_compute
N/A
6,394

Analysis: NVIDIA P106-100 vs NVIDIA RTX A4000 Mobile

The NVIDIA P106-100 and NVIDIA RTX A4000 Mobile represent two very different generations of GPU design, separated by four years of architectural evolution. The P106-100 is a Pascal-based mining card from 2017 with no display outputs, while the RTX A4000 Mobile is an Ampere-generation professional mobile GPU from 2021. Benchmark results show a decisive performance gap, with the RTX A4000 Mobile winning both head-to-head tests by substantial margins, yet the overall average benchmark scores tell a more nuanced story that rewards closer examination.

Head-to-Head Benchmarks

The direct comparison between these two GPUs is limited to two Compute benchmarks, and in both cases the RTX A4000 Mobile dominates. In Geekbench OpenCL, the RTX A4000 Mobile scores 97,178 against the P106-100’s 35,951, a delta of -63% from the perspective of the P106-100. This means the RTX A4000 Mobile is roughly 170% faster in raw compute throughput as measured by OpenCL. The Geekbench Vulkan result is similarly lopsided: the RTX A4000 Mobile posts 73,002 versus 32,897 for the P106-100, a -54.9% delta, putting the Ampere card approximately 122% ahead.

These results align with the theoretical compute specifications. The RTX A4000 Mobile delivers 17.20 TFLOPS of FP32 performance, while the P106-100 manages 4.375 TFLOPS — a nearly 4x difference in raw floating-point throughput. The RTX A4000 Mobile also shows a 1:1 FP16 ratio at 17.20 TFLOPS, whereas the P106-100’s FP16 performance is a paltry 68.36 GFLOPS (1:64), making the Pascal card effectively useless for half-precision workloads. The memory subsystem reinforces this gap: the RTX A4000 Mobile has 384.0 GB/s of bandwidth across a 256-bit bus with 8 GB of GDDR6, while the P106-100 is limited to 192.2 GB/s on a 192-bit bus with 6 GB of GDDR5.

However, the aggregate benchmark picture is less one-sided. The P106-100’s average benchmark score is 23,249, while the RTX A4000 Mobile averages 21,379. This places the P106-100 at the 68th percentile of all GPUs, slightly above the RTX A4000 Mobile’s 66th percentile. The head-to-head tests only cover two compute workloads, which heavily favor the Ampere architecture’s compute capabilities; the broader benchmark suite used to calculate average scores includes other test types where the Pascal card appears to hold its own. The nearest rival data for each card shows they sit in similar performance tiers overall: the P106-100’s closest competitor is the AMD Radeon Pro Vega 16 at a 0% delta, while the RTX A4000 Mobile’s nearest rival is the AMD Radeon HD 8970M at a 0.7% delta.

Architecture Differences

The architectural chasm between these two GPUs is vast. The P106-100 uses the GP106 chip built on TSMC’s 16 nm process, containing 4,400 million transistors on a 200 mm² die. The RTX A4000 Mobile uses the GA104 chip fabricated by Samsung on an 8 nm process, packing 17,400 million transistors into a 392 mm² die. This represents a 4x increase in transistor count and a 2x increase in transistor density, from 22.0M per mm² to 44.4M per mm² — a direct consequence of the process node shrink and architectural improvements.

The compute configuration differs dramatically. The P106-100 has 1,280 shading units, 80 TMUs, and 48 ROPs. The RTX A4000 Mobile has 5,120 shading units, 160 TMUs, and 80 ROPs — quadruple the shader count and double the texture and pixel throughput. The pixel rate jumps from 82.03 GPixel/s to 134.4 GPixel/s, and the texture rate from 136.7 GTexel/s to 268.8 GTexel/s. Most significantly, the RTX A4000 Mobile introduces dedicated hardware that the P106-100 completely lacks: 40 RT cores for ray tracing and 160 tensor cores for AI acceleration. These features enable the DirectX 12 Ultimate (12_2) API support on the Ampere card, while the P106-100 is limited to DirectX 12 (12_1).

Memory technology also separates the two. The P106-100 uses 6 GB of GDDR5 at 8 Gbps effective, while the RTX A4000 Mobile uses 8 GB of GDDR6 at 12 Gbps effective. The bus width grows from 192-bit to 256-bit, and memory bandwidth nearly doubles from 192.2 GB/s to 384.0 GB/s. Clock speeds are surprisingly similar: the P106-100 has a 1506 MHz base and 1709 MHz boost, while the RTX A4000 Mobile runs at 1140 MHz base and 1680 MHz boost. The higher transistor density and wider architecture more than compensate for the lower clocks on the Ampere part.

Power and physical characteristics differ substantially. The P106-100 is a dual-slot card drawing 120 W with a 1x 6-pin power connector and a 300 W suggested PSU. It measures 250 mm in length. The RTX A4000 Mobile has a 115 W TDP with no power connectors, as it is designed for portable devices where power delivery is integrated into the motherboard. The bus interface also evolves from PCIe 1.0 x16 on the P106-100 to PCIe 4.0 x16 on the RTX A4000 Mobile, quadrupling the available host bandwidth. The P106-100 has no display outputs at all, reflecting its mining-focused design, while the RTX A4000 Mobile’s outputs are described as "Portable Device Dependent."

Where Each One Wins

The RTX A4000 Mobile wins decisively in every head-to-head benchmark category, but its strengths are concentrated in compute-heavy workloads. The Geekbench OpenCL and Vulkan tests both measure general-purpose GPU compute, where the Ampere architecture’s massive shader count and tensor core support provide overwhelming advantages. The 17.20 TFLOPS FP32 and FP16 performance makes it suitable for scientific computing, machine learning inference, and rendering workloads that leverage half-precision arithmetic. The 40 RT cores enable hardware-accelerated ray tracing, a feature entirely absent from the P106-100, which has no RT or tensor core support.

The P106-100’s advantages are more subtle. Its average benchmark score of 23,249 exceeds the RTX A4000 Mobile’s 21,379, suggesting that in certain legacy or driver-optimized workloads, the Pascal architecture remains competitive. The 68th percentile ranking versus 66th for the RTX A4000 Mobile indicates that the P106-100 sits in a slightly higher performance tier when considering the entire GPU landscape. The P106-100 also has a higher boost clock (1709 MHz vs 1680 MHz) and a lower transistor density, which may translate to better thermal headroom in sustained workloads, though this is not directly measured in the data.

For gaming and graphics rendering, the RTX A4000 Mobile’s DirectX 12 Ultimate support and ray tracing capabilities give it a clear edge in modern titles. The P106-100’s DirectX 12 (12_1) support is sufficient for older games but lacks the feature set required for advanced effects. The RTX A4000 Mobile’s Passmark results show strong DirectX 9 performance at 157 and DirectX 11 at 127, but notably weaker DirectX 12 at 66 and DirectX 10 at 105 — suggesting some driver maturity issues or architectural inefficiencies in certain API paths. The P106-100 has no Passmark results in the data, so direct gaming comparisons are not possible.

The Verdict

The data clearly favors the RTX A4000 Mobile for any workload that leverages modern GPU features. The 63% lead in OpenCL and 54.9% lead in Vulkan are decisive, and the 4x advantage in FP32 throughput makes the Ampere card the obvious choice for compute-intensive applications. The addition of RT cores and tensor cores opens up entire categories of work — ray-traced rendering, AI inference, and mixed-precision scientific computing — that are simply impossible on the P106-100. The 8 GB of GDDR6 with 384.0 GB/s bandwidth provides twice the memory capacity and double the bandwidth, which is critical for large datasets and high-resolution textures.

However, the P106-100 is not without merit. Its higher average benchmark score and percentile ranking suggest that for legacy workloads or applications that do not utilize modern features, it can outperform the RTX A4000 Mobile. The 120 W TDP with a dual-slot cooler and 250 mm length makes it a straightforward desktop card, whereas the RTX A4000 Mobile requires a portable device platform. The P106-100’s lack of display outputs is a severe limitation for general use, but for mining or headless compute tasks, it has proven its value over many years.

For buyers who need professional-grade mobile compute with ray tracing and tensor core support, the RTX A4000 Mobile is the only choice — the P106-100 cannot compete on any modern metric. For those with existing Pascal-era systems needing compute acceleration without display output, the P106-100 remains a viable option, particularly given its higher aggregate benchmark standing. The RTX A4000 Mobile is end-of-life, as is the P106-100, so any purchase decision should factor in the availability of drivers and long-term support.

FAQ

Q: Which GPU has better raw compute performance?

A: The RTX A4000 Mobile is significantly faster in compute, delivering 17.20 TFLOPS of FP32 performance versus 4.375 TFLOPS for the P106-100. In Geekbench OpenCL, the RTX A4000 Mobile scores 97,178 against 35,951 for the P106-100, a 63% delta in its favor.

Q: Does the P106-100 support ray tracing?

A: No. The P106-100 has no RT cores and no tensor cores. The RTX A4000 Mobile includes 40 RT cores and 160 tensor cores, enabling hardware-accelerated ray tracing and AI workloads that the Pascal-based P106-100 cannot perform.

Q: Which GPU has more memory bandwidth?

A: The RTX A4000 Mobile has 384.0 GB/s of bandwidth across a 256-bit bus using 8 GB of GDDR6. The P106-100 has 192.2 GB/s on a 192-bit bus with 6 GB of GDDR5. The Ampere card offers exactly double the bandwidth.

Q: What is the process node difference between these GPUs?

A: The P106-100 is built on TSMC’s 16 nm process with 4,400 million transistors on a 200 mm² die. The RTX A4000 Mobile uses Samsung’s 8 nm process with 17,400 million transistors on a 392 mm² die, resulting in a transistor density of 44.4M per mm² versus 22.0M per mm² for the Pascal card.

Q: Are these GPUs still in production?

A: No. Both the P106-100 and the RTX A4000 Mobile are listed as end-of-life products. The P106-100 was released in June 2017, and the RTX A4000 Mobile was released in April 2021.

Q: Which GPU has a higher average benchmark score?

A: The P106-100 has a higher average benchmark score of 23,249 compared to 21,379 for the RTX A4000 Mobile. This places the P106-100 at the 68th percentile of all GPUs, while the RTX A4000 Mobile sits at the 66th percentile, despite the RTX A4000 Mobile winning both head-to-head tests.

DETAILED SPECIFICATIONS

SPECIFICATION
P106-100
RTX A4000 Mobile
Core Specs
Shading Units
1,280
5,120 +300.0%
Shaders
1,280
5,120 +300.0%
TMUs
80
160 +100.0%
ROPs
48
80 +66.7%
SM Count
10
40 +300.0%
Clocks
Base Clock
1506 MHz
1140 MHz
Boost Clock
1709 MHz
1680 MHz
Memory Clock
2002 MHz 8 Gbps effective
1500 MHz 12 Gbps effective
Memory
Memory Size
6 GB
8 GB
VRAM (MB)
6,144
8,192 +33.3%
Memory Type
GDDR5
GDDR6
Memory Bus
192 bit
256 bit
Bandwidth
192.2 GB/s
384.0 GB/s
Cache
L1 Cache
48 KB (per SM)
128 KB (per SM)
L2 Cache
1536 KB
4 MB
Performance
Pixel Rate
82.03 GPixel/s
134.4 GPixel/s
Texture Rate
136.7 GTexel/s
268.8 GTexel/s
FP32 (TFLOPS)
4.375 TFLOPS
17.20 TFLOPS
FP64 (TFLOPS)
136.7 GFLOPS (1:32)
268.8 GFLOPS (1:64)
FP16 (TFLOPS)
68.36 GFLOPS (1:64)
17.20 TFLOPS (1:1)
AI/RT
RT Cores
40
Tensor Cores
160
Power
TDP
120 W
115 W
TDP (W)
120
115 -4.2%
Suggested PSU
300 W
Power Connectors
1x 6-pin
None
Architecture
Architecture
Pascal
Ampere
GPU Name
GP106
GA104
Generation
Mining GPUs
Ampere-MW (Ax000)
Process Size
16 nm
8 nm
Transistors
4,400 million
17,400 million
Die Size
200 mm²
392 mm²
Foundry
TSMC
Samsung
Density
22.0M / mm²
44.4M / mm²
API Support
DirectX
12 (12_1)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
6.1
8.6
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Length
250 mm 9.8 inches
Outputs
No outputs
Portable Device Dependent
Bus Interface
PCIe 1.0 x16
PCIe 4.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Quadro Turing-M
Successor
Ada-MW
View P106-100 Details View RTX A4000 Mobile Details