NVIDIA P104-100 vs NVIDIA RTX A1000 Comparison

NVIDIA
GEFORCE

NVIDIA P104-100

CORE STATE GP104
VRAM 4 GB
CLOCK SPEED 1733 MHz
TDP
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2017
VS
NVIDIA
GEFORCE

RTX A1000

CORE STATE GA107
VRAM 8 GB
CLOCK SPEED 1462 MHz
TDP 50 W
BUS WIDTH 128 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2024

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
1,413
969
geekbench_opencl
52,368
52,078
geekbench_vulkan
45,165
49,574

Analysis: NVIDIA P104-100 vs NVIDIA RTX A1000

The NVIDIA RTX A1000 and NVIDIA P104-100 occupy very different positions in the GPU landscape, one a modern low-power workstation card and the other a mining-focused Pascal part. Across the three shared benchmarks, the P104-100 takes two wins, but the A1000’s single victory is decisive and points to a fundamental architectural shift. The average benchmark scores are close — 34,207 for the A1000 versus 32,982 for the P104-100 — yet the per-test results tell a more nuanced story about where each card excels.

Head-to-Head Benchmarks

The most striking contrast appears in the 3DMark Steel Nomad DX12 test. The P104-100 scores 1,413, while the RTX A1000 manages only 969, a 31.4% deficit. This is a substantial gap, and it shows the P104-100’s raw rasterization throughput advantage in a modern API workload. The P104-100’s higher boost clock of 1733 MHz, compared to the A1000’s 1462 MHz, and its wider 256-bit memory bus with 320.3 GB/s of bandwidth, likely contribute to this result. The A1000’s 128-bit bus and 192.0 GB/s bandwidth are significantly narrower, which can bottleneck fill-rate-heavy scenes. In this test, the P104-100 is clearly the stronger performer, and the 31.4% delta is the largest margin in any benchmark between these two cards.

The Geekbench OpenCL test is much closer. The P104-100 edges out the A1000 with a score of 52,368 versus 52,078, a mere 0.6% difference. This near-tie is interesting because it suggests that in compute-heavy, general-purpose workloads, the two GPUs are effectively on par despite their different architectures. The A1000’s FP32 throughput is listed at 6.737 TFLOPS, while the P104-100 is at 6.655 TFLOPS, a difference of less than 1.3%. The benchmark data reflects this almost exactly, with the P104-100’s advantage being within the margin of measurement noise. The P104-100’s 6.655 TFLOPS of FP32 performance, combined with its higher 208.0 GTexel/s texture rate, helps it stay competitive, but the A1000’s newer Ampere architecture and 6.737 TFLOPS keep it close.

The third benchmark, Geekbench Vulkan, flips the script decisively. Here, the RTX A1000 scores 49,574, beating the P104-100’s 45,165 by 9.8%. This is the A1000’s only head-to-head win, but it is a meaningful one. Vulkan is a low-overhead API that can expose architectural efficiencies, and the A1000’s Ampere design, with its 18 RT cores and 72 tensor cores, appears to handle the workload more efficiently than the P104-100’s Pascal architecture, which lacks both RT and tensor cores entirely. The A1000’s 1:1 FP16 ratio (6.737 TFLOPS) versus the P104-100’s heavily reduced FP16 rate (104.0 GFLOPS, a 1:64 ratio) may also play a role in compute-oriented Vulkan tasks. The 9.8% win is not as large as the P104-100’s Steel Nomad margin, but it demonstrates that the A1000 is not simply a slower card — it has genuine strengths in specific API contexts.

Overall, the P104-100 wins two of three tests, but its victory in OpenCL is marginal. The A1000’s Vulkan win is more substantial, and its average benchmark score is actually 3.7% higher than the P104-100’s when considering all tests and the percentile rankings (79th versus 77th percentile). This suggests that while the P104-100 has a raw power edge in certain DX12 scenarios, the A1000 is the more balanced overall performer.

The Verdict

Data from the benchmark suite indicates that the NVIDIA P104-100 is the better choice for users prioritizing raw DX12 rasterization performance. Its 31.4% lead in 3DMark Steel Nomad is a decisive advantage that cannot be ignored for gaming or D3D12-heavy workloads. The card’s higher pixel rate (110.9 GPixel/s versus 46.78 GPixel/s) and texture rate (208.0 GTexel/s versus 105.3 GTexel/s) support this finding. If the primary use case involves modern DirectX 12 titles, the P104-100 is the clear winner based on this data.

However, the RTX A1000 is the more versatile card. Its 9.8% Vulkan victory and near-parity in OpenCL (within 0.6%) show that it can hold its own in compute and cross-platform APIs. The A1000 also has a higher average benchmark score (34,207 versus 32,982) and sits in a higher percentile (79th versus 77th). For users who need a card that can handle a variety of workloads — including Vulkan-based applications or those that might leverage FP16 compute — the A1000’s architectural features, such as RT cores and tensor cores, make it a more future-proof option. The P104-100’s lack of display outputs and its end-of-life production status further tilt the balance toward the A1000 for general-purpose use. The verdict is straightforward: pick the P104-100 for pure DX12 speed, and the A1000 for everything else.

Architecture Differences

The two GPUs are built on fundamentally different architectures and process nodes. The RTX A1000 uses the GA107 chip on the Ampere architecture, manufactured by Samsung on an 8 nm process. It packs 8,700 million transistors into a 200 mm² die, yielding a transistor density of 43.5 million per mm². In contrast, the P104-100 uses the GP104 chip on the older Pascal architecture, built by TSMC on a 16 nm process. It has 7,200 million transistors on a larger 314 mm² die, with a much lower density of 22.9 million per mm². The A1000’s smaller, denser die is a direct result of the newer manufacturing process, allowing it to pack more features into less space.

The A1000 features 2,304 shading units, 72 TMUs, and 32 ROPs. It also includes 18 RT cores and 72 tensor cores, which are absent entirely from the P104-100. The P104-100 has 1,920 shading units, 120 TMUs, and 64 ROPs. This means the P104-100 has more texture mapping units and ROPs, which explains its higher pixel and texture rates. The A1000 compensates with more shading units and dedicated ray tracing and tensor hardware. Memory configurations also differ significantly: the A1000 uses 8 GB of GDDR6 on a 128-bit bus, while the P104-100 uses 4 GB of GDDR5X on a 256-bit bus. The P104-100’s memory bandwidth is 320.3 GB/s versus the A1000’s 192.0 GB/s, a 66.9% advantage for the older card.

Clock speeds tell another story. The P104-100 has a base clock of 1607 MHz and a boost clock of 1733 MHz, both far higher than the A1000’s 727 MHz base and 1462 MHz boost. The A1000’s lower clocks are a trade-off for its drastically lower 50 W TDP, whereas the P104-100’s TDP is not listed but requires a 200 W suggested PSU and a single 8-pin power connector. The A1000 draws power from the PCIe slot alone and is a single-slot card, while the P104-100 is dual-slot with no display outputs. The A1000 supports PCIe 4.0 x8, while the P104-100 is limited to PCIe 1.0 x4, which could constrain data transfer in some scenarios. The A1000 also supports DirectX 12 Ultimate (12_2), while the P104-100 is limited to DirectX 12 (12_1), a distinction that matters for features like mesh shaders and variable rate shading.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA RTX A1000 has an average benchmark score of 34,207, which is 3.7% higher than the NVIDIA P104-100’s 32,982.

Q: What is the largest performance gap between the two cards in any single benchmark?

A: The P104-100 leads by 31.4% in the 3DMark Steel Nomad DX12 test, scoring 1,413 versus the A1000’s 969.

Q: Does the P104-100 support ray tracing or tensor cores?

A: No. The P104-100’s Pascal architecture has no RT cores and no tensor cores. The RTX A1000 includes 18 RT cores and 72 tensor cores.

Q: How do their memory subsystems compare?

A: The A1000 has 8 GB of GDDR6 on a 128-bit bus with 192.0 GB/s bandwidth. The P104-100 has 4 GB of GDDR5X on a 256-bit bus with 320.3 GB/s bandwidth.

Q: Which card is more power-efficient?

A: The RTX A1000 has a listed TDP of 50 W and requires no power connectors. The P104-100’s TDP is not specified, but it requires a 1x 8-pin power connector and a 200 W suggested PSU.

Q: Are there any display outputs on the P104-100?

A: No. The P104-100 has no display outputs, while the RTX A1000 offers 4x mini-DisplayPort 1.4a connections.

Where Each One Wins

The NVIDIA P104-100 wins decisively in DirectX 12 rasterization workloads. Its 31.4% lead in 3DMark Steel Nomad is the clearest indicator, supported by higher pixel and texture rates (110.9 GPixel/s and 208.0 GTexel/s). This makes it the better option for users running modern DX12 games or applications that rely heavily on fill rate and memory bandwidth. The P104-100’s wider 256-bit bus and 320.3 GB/s of bandwidth are significant assets in this context. It also edges out the A1000 in OpenCL by 0.6%, though this margin is negligible and effectively a tie.

The RTX A1000 wins in Vulkan-based workloads, taking a 9.8% lead in Geekbench Vulkan. This suggests its Ampere architecture handles low-level API compute more efficiently, possibly due to its 1:1 FP16 ratio and tensor core support. The A1000 is also the only card with display outputs, making it suitable for workstation use where visual output is required. Its 50 W TDP and lack of external power connectors mean it can fit into slim, low-power systems. The A1000’s higher average benchmark score and 79th percentile ranking versus the P104-100’s 77th percentile indicate it is the more capable all-rounder. For users who need a balance of compute performance, modern API support, and power efficiency, the A1000 is the superior choice.

DETAILED SPECIFICATIONS

SPECIFICATION
P104-100
RTX A1000
Core Specs
Shading Units
1,920
2,304 +20.0%
Shaders
1,920
2,304 +20.0%
TMUs
120
72 -40.0%
ROPs
64
32 -50.0%
SM Count
15
18 +20.0%
Clocks
Base Clock
1607 MHz
727 MHz
Boost Clock
1733 MHz
1462 MHz
Memory Clock
1251 MHz 10 Gbps effective
1500 MHz 12 Gbps effective
Memory
Memory Size
4 GB
8 GB
VRAM (MB)
4,096
8,192 +100.0%
Memory Type
GDDR5X
GDDR6
Memory Bus
256 bit
128 bit
Bandwidth
320.3 GB/s
192.0 GB/s
Cache
L1 Cache
48 KB (per SM)
128 KB (per SM)
L2 Cache
2 MB
2 MB
Performance
Pixel Rate
110.9 GPixel/s
46.78 GPixel/s
Texture Rate
208.0 GTexel/s
105.3 GTexel/s
FP32 (TFLOPS)
6.655 TFLOPS
6.737 TFLOPS
FP64 (TFLOPS)
208.0 GFLOPS (1:32)
105.3 GFLOPS (1:64)
FP16 (TFLOPS)
104.0 GFLOPS (1:64)
6.737 TFLOPS (1:1)
AI/RT
RT Cores
18
Tensor Cores
72
Power
TDP
50 W
TDP (W)
50
Suggested PSU
200 W
250 W
Power Connectors
1x 8-pin
None
Architecture
Architecture
Pascal
Ampere
GPU Name
GP104
GA107
Generation
Mining GPUs
Workstation Ampere (Ax000)
Process Size
16 nm
8 nm
Transistors
7,200 million
8,700 million
Die Size
314 mm²
200 mm²
Foundry
TSMC
Samsung
Density
22.9M / mm²
43.5M / mm²
API Support
DirectX
12 (12_1)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
6.1
8.6
Shader Model
6.8
6.9
Physical
Slot Width
Dual-slot
Single-slot
Length
267 mm 10.5 inches
163 mm 6.4 inches
Height
69 mm 2.7 inches
Outputs
No outputs
4x mini-DisplayPort 1.4a
Bus Interface
PCIe 1.0 x4
PCIe 4.0 x8
Other
Production
End-of-life
Active
Predecessor
Quadro Turing
Successor
Workstation Ada
View P104-100 Details View RTX A1000 Details