NVIDIA A2 vs NVIDIA P104-100 Comparison

NVIDIA
GEFORCE

NVIDIA A2

CORE STATE GA107
VRAM 16 GB
CLOCK SPEED 1770 MHz
TDP 60 W
BUS WIDTH 128 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

P104-100

CORE STATE GP104
VRAM 4 GB
CLOCK SPEED 1733 MHz
TDP
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2017

PERFORMANCE BENCHMARKS

geekbench_opencl
35,357
52,368
geekbench_vulkan
34,023
45,165
3dmark_3dmark_steel_nomad_dx12
N/A
1,413

Analysis: NVIDIA A2 vs NVIDIA P104-100

NVIDIA A2 and NVIDIA P104-100 are both end-of-life NVIDIA cards with no display outputs, but they target very different eras and workloads. The A2 is a modern Ampere-generation workstation card built on an 8nm Samsung process, while the P104-100 is a Pascal-generation mining GPU from 2017 on a 16nm TSMC process. The benchmark data reveals a clear performance hierarchy, but the architectural gap means the story is more nuanced than raw scores alone.

Head-to-Head Benchmarks

The head-to-head comparison is brief, with only two shared benchmark results, and the P104-100 wins both decisively. In Geekbench OpenCL, the P104-100 scores 52,368 against the A2's 35,357, a delta of -32.5% from the A2's perspective — meaning the A2 trails by nearly a third. The Geekbench Vulkan result is closer but still favors the older card: the P104-100 hits 45,165 while the A2 manages 34,023, a 24.7% deficit. These are not marginal losses; the P104-100 outperforms the A2 by a wide margin in raw compute throughput.

The average benchmark scores tell a similar story. The A2's average is 34,690, while the P104-100's average is 32,982 — a surprising twist, because the P104-100 wins both head-to-head tests yet has a lower average overall. This discrepancy stems from the P104-100's third benchmark, a 3DMark Steel Nomad DX12 test where it scores just 1,413, which drags its average down. The A2 only has the two Geekbench results, so its average reflects only its stronger (relative) performance. The percentile rankings reinforce the parity: the A2 sits at the 79th percentile of all GPUs, while the P104-100 is at the 77th — a two-point gap that suggests they are broadly comparable in the overall GPU landscape, despite the P104-100's head-to-head dominance.

Looking at nearest rivals, the A2's average of 34,690 is just 0.4% above the NVIDIA T1000 8GB (34,561) and 0.4% above the AMD Radeon HD 7970 (34,541). It also edges out the NVIDIA TITAN V (34,355) by 1% and the RTX A1000 (34,207) by 1.4%. The P104-100's average of 32,982 sits 0.4% above the T600 Mobile (32,849), but falls 0.5% short of the T550 Mobile (33,161) and 0.6% short of the RTX 3050 Mobile (33,170). These tight deltas indicate that both cards sit in a crowded mid-range performance band, where small architectural differences translate into single-digit percentage swings against competitors.

Where Each One Wins

The P104-100 wins the only two direct comparisons, so on pure compute benchmarks it is the clear victor. Its advantage is most pronounced in OpenCL, where it leads by 32.5%, and it maintains a solid 24.7% lead in Vulkan. This suggests the P104-100's higher raw shader throughput — 1,920 shading units versus the A2's 1,280 — gives it a significant edge in general-purpose compute workloads that scale well with parallel cores. The P104-100 also has a much higher texture rate (208.0 GTexel/s vs 70.80 GTexel/s) and pixel rate (110.9 GPixel/s vs 56.64 GPixel/s), which explains its dominance in fill-rate-bound tasks.

The A2, however, is not without its own advantages. It has 16 GB of GDDR6 memory versus the P104-100's 4 GB of GDDR5X, a fourfold capacity difference that matters immensely for large datasets, deep learning models, or rendering scenes that exceed 4 GB. The A2 also supports PCIe 4.0 x8, while the P104-100 is limited to PCIe 1.0 x4 — a massive bandwidth disparity for data transfer between the CPU and GPU, even if the P104-100's internal memory bandwidth is higher (320.3 GB/s vs 200.1 GB/s). The A2's modern Ampere architecture brings hardware ray tracing cores (10) and tensor cores (40), features the Pascal-based P104-100 lacks entirely. For workloads that leverage these specialized units, the A2 wins by default, since the P104-100 cannot accelerate them at all.

The A2 also has a far lower power draw at 60 W TDP versus the P104-100's unspecified TDP, though the latter suggests a 200 W power supply compared to the A2's 250 W suggestion. The A2 is single-slot with no power connectors, while the P104-100 is dual-slot and requires an 8-pin connector. For dense server installations or low-power edge deployments, the A2's efficiency and physical footprint are compelling, even if it loses on raw compute.

Architecture Differences

The architectural chasm between these two cards is generational. The A2 uses the GA107 chip on an 8nm Samsung process, packing 8,700 million transistors into a 200 mm² die, yielding a transistor density of 43.5M per mm². The P104-100 uses the GP104 chip on a 16nm TSMC process, with 7,200 million transistors spread across a larger 314 mm² die, resulting in a density of just 22.9M per mm². The A2's modern node gives it more than double the transistor density, which explains how it achieves comparable overall performance with far fewer cores and lower power.

The A2's Ampere architecture supports DirectX 12 Ultimate (12_2), while the P104-100 is limited to DirectX 12 (12_1). Both support OpenGL 4.6 and Vulkan 1.4, but the A2's feature set includes ray tracing and tensor cores — 10 RT cores and 40 tensor cores — which are absent on the P104-100. The A2 also delivers FP16 at a 1:1 ratio with FP32 (4.531 TFLOPS both), while the P104-100's FP16 is a mere 104.0 GFLOPS, a 1:64 ratio, making it essentially useless for half-precision compute. The A2's FP32 performance is lower at 4.531 TFLOPS versus the P104-100's 6.655 TFLOPS, but the A2's AI-focused tensor cores and ray tracing hardware allow it to accelerate workloads the P104-100 simply cannot.

Memory configurations diverge sharply: the A2 has 16 GB of GDDR6 on a 128-bit bus (200.1 GB/s), while the P104-100 has 4 GB of GDDR5X on a 256-bit bus (320.3 GB/s). The P104-100's wider bus and higher bandwidth are better for traditional rasterization, but the A2's larger capacity is better for holding big models or datasets. The bus interface also differs: the A2 uses PCIe 4.0 x8, a modern high-bandwidth connection, while the P104-100 is stuck with PCIe 1.0 x4, which is dramatically slower for host-device transfers. Both cards have no display outputs, reinforcing their compute/mining focus.

FAQ

Q: Which GPU is faster in raw compute benchmarks?

A: The P104-100 wins both shared benchmarks. It leads by 32.5% in Geekbench OpenCL (52,368 vs 35,357) and by 24.7% in Geekbench Vulkan (45,165 vs 34,023).

Q: Does the A2 have any performance advantage over the P104-100?

A: In terms of raw scores, no — the A2 loses both head-to-head tests. However, the A2 has specialized hardware the P104-100 lacks: 10 RT cores and 40 tensor cores, which enable ray tracing and AI acceleration that the Pascal card cannot perform.

Q: Why does the P104-100 have a lower average benchmark score despite winning head-to-head?

A: The P104-100's average (32,982) is dragged down by a third benchmark — 3DMark Steel Nomad DX12 — where it scores only 1,413. The A2 only has the two Geekbench results, so its average (34,690) reflects only those higher-scoring tests.

Q: How much memory does each card have, and does it matter?

A: The A2 has 16 GB of GDDR6, while the P104-100 has 4 GB of GDDR5X. This is a fourfold difference in capacity. The P104-100 has higher bandwidth (320.3 GB/s vs 200.1 GB/s), but the A2's larger pool is critical for workloads that exceed 4 GB, such as large language models or high-resolution rendering.

Q: What are the power and physical differences?

A: The A2 has a 60 W TDP, is single-slot, and requires no power connectors. The P104-100 is dual-slot, requires a single 8-pin power connector, and its TDP is unspecified. The A2 suggests a 250 W power supply, while the P104-100 suggests 200 W.

Q: Which card has better connectivity to the host system?

A: The A2 uses PCIe 4.0 x8, which is a modern high-bandwidth interface. The P104-100 uses PCIe 1.0 x4, which is severely limited for data transfer, potentially bottlenecking workloads that require frequent CPU-GPU communication.

Specification Differences

| Specification | NVIDIA A2 | NVIDIA P104-100 |

|---|---|---|

| Architecture | Ampere | Pascal |

| Process Node | 8 nm (Samsung) | 16 nm (TSMC) |

| Transistors | 8,700 million | 7,200 million |

| Die Size | 200 mm² | 314 mm² |

| Transistor Density | 43.5M / mm² | 22.9M / mm² |

| Base Clock | 1440 MHz | 1607 MHz |

| Boost Clock | 1770 MHz | 1733 MHz |

| Memory Clock | 1563 MHz (12.5 Gbps effective) | 1251 MHz (10 Gbps effective) |

| Memory Size | 16 GB GDDR6 | 4 GB GDDR5X |

| Memory Bus Width | 128 bit | 256 bit |

| Memory Bandwidth | 200.1 GB/s | 320.3 GB/s |

| Shading Units | 1280 | 1920 |

| TMUs | 40 | 120 |

| ROPs | 32 | 64 |

| RT Cores | 10 | 0 (null) |

| Tensor Cores | 40 | 0 (null) |

| Pixel Rate | 56.64 GPixel/s | 110.9 GPixel/s |

| Texture Rate | 70.80 GTexel/s | 208.0 GTexel/s |

| FP32 | 4.531 TFLOPS | 6.655 TFLOPS |

| FP16 | 4.531 TFLOPS (1:1) | 104.0 GFLOPS (1:64) |

| TDP | 60 W | Unspecified |

| Slot Width | Single-slot | Dual-slot |

| Power Connectors | None | 1x 8-pin |

| Suggested PSU | 250 W | 200 W |

| Bus Interface | PCIe 4.0 x8 | PCIe 1.0 x4 |

| DirectX Support | 12 Ultimate (12_2) | 12 (12_1) |

| Length | Unspecified | 267 mm (10.5 inches) |

The Verdict

The data presents a clear choice based on workload type. If the priority is raw compute throughput in OpenCL or Vulkan, the P104-100 is the superior card — it wins both head-to-head tests by 32.5% and 24.7%, respectively, and offers higher FP32 (6.655 vs 4.531 TFLOPS), pixel rate (110.9 vs 56.64 GPixel/s), and texture rate (208.0 vs 70.80 GTexel/s). Its wider 256-bit memory bus and higher bandwidth (320.3 vs 200.1 GB/s) also favor traditional rendering or compute tasks that are bandwidth-sensitive.

However, the P104-100's limitations are severe for modern workloads. It has only 4 GB of memory, which is a hard ceiling for many current AI models or large datasets. Its PCIe 1.0 x4 interface is a bottleneck for host-device transfers. And it lacks any ray tracing or tensor core support, making it obsolete for DirectX 12 Ultimate features or accelerated AI inference.

The A2, despite losing on raw benchmarks, is the more future-proof and versatile option. Its 16 GB memory capacity is four times larger, its Ampere architecture supports the latest graphics APIs and hardware acceleration, and its 60 W TDP with no power connectors makes it far easier to deploy in dense or power-constrained environments. The A2 also holds a higher percentile ranking (79th vs 77th) and a higher average benchmark score (34,690 vs 32,982), suggesting that when the P104-100's weak 3DMark result is factored in, the A2 is the more consistent performer overall.

Choose the P104-100 for pure compute throughput with no memory constraints. Choose the A2 for modern workloads, large memory footprints, low power, and hardware-accelerated features. The benchmark results favor the older card on speed, but the architectural advantages of the A2 make it the more capable long-term investment for any task that does not fit within 4 GB or requires feature support beyond DirectX 12_1.

DETAILED SPECIFICATIONS

SPECIFICATION
A2
P104-100
Core Specs
Shading Units
1,280
1,920 +50.0%
Shaders
1,280
1,920 +50.0%
TMUs
40
120 +200.0%
ROPs
32
64 +100.0%
SM Count
10
15 +50.0%
Clocks
Base Clock
1440 MHz
1607 MHz
Boost Clock
1770 MHz
1733 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
1251 MHz 10 Gbps effective
Memory
Memory Size
16 GB
4 GB
VRAM (MB)
16,384
4,096 -75.0%
Memory Type
GDDR6
GDDR5X
Memory Bus
128 bit
256 bit
Bandwidth
200.1 GB/s
320.3 GB/s
Cache
L1 Cache
128 KB (per SM)
48 KB (per SM)
L2 Cache
2 MB
2 MB
Performance
Pixel Rate
56.64 GPixel/s
110.9 GPixel/s
Texture Rate
70.80 GTexel/s
208.0 GTexel/s
FP32 (TFLOPS)
4.531 TFLOPS
6.655 TFLOPS
FP64 (TFLOPS)
70.80 GFLOPS (1:64)
208.0 GFLOPS (1:32)
FP16 (TFLOPS)
4.531 TFLOPS (1:1)
104.0 GFLOPS (1:64)
AI/RT
RT Cores
10
Tensor Cores
40
Power
TDP
60 W
TDP (W)
60
Suggested PSU
250 W
200 W
Power Connectors
None
1x 8-pin
Architecture
Architecture
Ampere
Pascal
GPU Name
GA107
GP104
Generation
Workstation Ampere (Ax000)
Mining GPUs
Process Size
8 nm
16 nm
Transistors
8,700 million
7,200 million
Die Size
200 mm²
314 mm²
Foundry
Samsung
TSMC
Density
43.5M / mm²
22.9M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Single-slot
Dual-slot
Length
267 mm 10.5 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x8
PCIe 1.0 x4
Other
Production
End-of-life
End-of-life
Predecessor
Quadro Turing
Successor
Workstation Ada
View A2 Details View P104-100 Details