AMD Radeon Pro 580 vs NVIDIA P104-100 Comparison

AMD
RADEON

AMD Radeon Pro 580

CORE STATE Ellesmere
VRAM 8 GB
CLOCK SPEED 1200 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE GCN 4.0
nm
PROCESS 14 nm
LAUNCH DATE 2017
VS
NVIDIA
GEFORCE

P104-100

CORE STATE GP104
VRAM 4 GB
CLOCK SPEED 1733 MHz
TDP
BUS WIDTH 256 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2017

PERFORMANCE BENCHMARKS

geekbench_metal
39,213
N/A
geekbench_opencl
38,457
52,368
geekbench_vulkan
43,285
45,165
3dmark_3dmark_steel_nomad_dx12
N/A
1,413

Analysis: AMD Radeon Pro 580 vs NVIDIA P104-100

Head-to-Head Benchmarks

The head-to-head data contains two shared benchmark results, and the NVIDIA P104-100 takes both. In Geekbench OpenCL, the P104-100 scores 52,368 against the Radeon Pro 580's 38,457. That is a 26.6% lead for NVIDIA, a substantial margin that reflects the P104-100's higher raw compute throughput. The Radeon Pro 580 is not close here; the gap is nearly a third of its own score.

The second shared test, Geekbench Vulkan, is much tighter. The P104-100 scores 45,165, while the Radeon Pro 580 scores 43,285. The delta is only 4.2% in NVIDIA's favor. This suggests that in Vulkan workloads, the AMD card is far more competitive, and the gap could be attributed to driver optimization or architectural differences rather than a fundamental compute disadvantage. Still, the P104-100 wins both recorded comparisons, giving it a clean 2-0 sweep in the head-to-head.

Looking at the broader benchmark database, the Radeon Pro 580's average score across all tests is 40,318, placing it at the 82nd percentile of all GPUs. The P104-100's average is 32,982, which is the 77th percentile. This is an interesting inversion: the P104-100 wins the two shared tests, but its overall average is dragged down by the inclusion of a 3DMark Steel Nomad DX12 score of 1,413, which is a very different workload from the Geekbench compute tests. The Radeon Pro 580 has no equivalent DX12 result in the database, so its average is based solely on Geekbench Metal, OpenCL, and Vulkan scores, all of which are relatively high.

The nearest rivals for the Radeon Pro 580 include the NVIDIA GeForce RTX 5070 (average score 40,377, delta -0.1%), the AMD Radeon Pro WX 7100 (40,063, delta 0.6%), and the AMD Radeon Pro 5300 (40,870, delta -1.4%). These are all within a few percentage points, meaning the Radeon Pro 580 sits in a tightly contested performance band. The P104-100's nearest rivals are lower-tier mobile and workstation parts: the NVIDIA T600 Mobile (32,849, delta 0.4%), NVIDIA T550 Mobile (33,161, delta -0.5%), and GeForce RTX 3050 Mobile (33,170, delta -0.6%). The P104-100's average is competitive with those, but it is clearly not in the same tier as the Radeon Pro 580 when considering the full benchmark suite.

FAQ

Q: Which GPU wins the Geekbench OpenCL test?

A: The NVIDIA P104-100 wins with a score of 52,368, beating the AMD Radeon Pro 580's 38,457 by 26.6%.

Q: Is the Vulkan gap as large as the OpenCL gap?

A: No. In Geekbench Vulkan, the P104-100 scores 45,165 versus 43,285 for the Radeon Pro 580, a margin of only 4.2%.

Q: Which GPU has a higher average benchmark score across all tests?

A: The AMD Radeon Pro 580 has an average of 40,318, compared to the NVIDIA P104-100's 32,982. The Radeon Pro 580 also sits at the 82nd percentile of all GPUs, while the P104-100 sits at the 77th.

Q: Does the P104-100 have any benchmark result outside Geekbench?

A: Yes, the P104-100 has a 3DMark Steel Nomad DX12 score of 1,413. The Radeon Pro 580 has no corresponding DX12 result in the database.

Q: What does the P104-100's average score compare to among its nearest rivals?

A: The P104-100's average of 32,982 is within 0.7% of the NVIDIA T600 Mobile (32,849), T550 Mobile (33,161), GeForce RTX 3050 Mobile (33,170), and AMD Radeon Pro 570 (33,207).

Q: How does the Radeon Pro 580's average compare to its nearest rivals?

A: Its average of 40,318 is within 1.9% of the RTX 5070 (40,377), Radeon Pro WX 7100 (40,063), Radeon Pro 5300 (40,870), and RTX A500 Mobile (39,568).

Where Each One Wins

The NVIDIA P104-100 is the clear winner in raw compute benchmarks. It dominates Geekbench OpenCL with a 26.6% margin, indicating that OpenCL-heavy workloads, such as certain scientific computing or rendering tasks, will see a significant performance advantage on the P104-100. Its Vulkan lead is smaller but still consistent, making it the better choice for Vulkan-based applications. The P104-100 also has a higher pixel rate (110.9 GPixel/s versus 38.40 GPixel/s) and texture rate (208.0 GTexel/s versus 172.8 GTexel/s), which suggests it handles fill-rate-bound scenarios better.

The AMD Radeon Pro 580 wins on consistency and overall average performance. Its average benchmark score of 40,318 is 22.2% higher than the P104-100's 32,982, and it achieves this across three Geekbench tests without any low outlier. The 82nd percentile ranking versus the 77th percentile for the P104-100 shows that, in the broader GPU landscape, the Radeon Pro 580 is considered the stronger overall performer. It also offers higher FP16 throughput (5.530 TFLOPS versus 104.0 GFLOPS), which is a massive difference for workloads that utilize half-precision arithmetic. The Radeon Pro 580's memory configuration (8 GB GDDR5) is also double the capacity of the P104-100, making it more suitable for large datasets.

Specification Differences

The two cards differ on nearly every core specification. The AMD Radeon Pro 580 uses a 14 nm process from GlobalFoundries, while the NVIDIA P104-100 uses a 16 nm process from TSMC. The Radeon Pro 580 has 5,700 million transistors on a 232 mm² die, resulting in a transistor density of 24.6M per mm². The P104-100 has 7,200 million transistors on a 314 mm² die, with a density of 22.9M per mm².

Clock speeds are very different. The Radeon Pro 580 runs at a base of 1100 MHz and a boost of 1200 MHz, with memory at 1695 MHz (6.8 Gbps effective). The P104-100 runs at a base of 1607 MHz and a boost of 1733 MHz, with memory at 1251 MHz (10 Gbps effective). The P104-100's higher core clocks are a key reason for its compute wins.

Memory configurations diverge sharply. The Radeon Pro 580 has 8 GB of GDDR5 on a 256-bit bus, yielding 217.0 GB/s of bandwidth. The P104-100 has 4 GB of GDDR5X on the same 256-bit bus, but with 320.3 GB/s of bandwidth. The P104-100 has a clear bandwidth advantage, but half the capacity.

Compute unit counts also differ: the Radeon Pro 580 has 2304 shading units, 144 TMUs, and 32 ROPs. The P104-100 has 1920 shading units, 120 TMUs, and 64 ROPs. The P104-100's higher ROP count explains its much higher pixel rate. The power situation is distinct: the Radeon Pro 580 has a TDP of 185 W and no power connectors (it is an integrated GPU, or IGP, slot width), while the P104-100 has a dual-slot design, a single 8-pin power connector, and a suggested PSU of 200 W. The bus interface also differs: the Radeon Pro 580 uses PCIe 3.0 x16, while the P104-100 uses PCIe 1.0 x4, a significant bottleneck for data transfer. The P104-100 has no display outputs, while the Radeon Pro 580's outputs are portable device dependent.

Architecture Differences

The architectural split is clear: AMD uses GCN 4.0 on the Ellesmere chip, while NVIDIA uses Pascal on the GP104 chip. The Radeon Pro 580 belongs to the Radeon Pro Mac (500 Series) generation, whereas the P104-100 is part of NVIDIA's Mining GPUs generation. The Radeon Pro 580 supports DirectX 12 (12_0), OpenGL 4.6, and Vulkan 1.3. The P104-100 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The higher DirectX feature level and Vulkan version on the P104-100 are notable for API compatibility.

FP16 performance is a major architectural differentiator. The Radeon Pro 580 delivers 5.530 TFLOPS of FP16 at a 1:1 ratio with FP32, meaning it can process half-precision data at full speed. The P104-100 delivers only 104.0 GFLOPS of FP16 at a 1:64 ratio, making it effectively poor at half-precision tasks. This is a decisive factor for AI or machine learning workloads that rely on FP16.

The Radeon Pro 580's transistor density is slightly higher (24.6M / mm² versus 22.9M / mm²), despite the older 14 nm process. The P104-100's larger die (314 mm² versus 232 mm²) and higher transistor count (7,200 million versus 5,700 million) indicate a more complex design. The P104-100's memory clock of 10 Gbps effective is significantly faster than the Radeon Pro 580's 6.8 Gbps, which directly contributes to its higher bandwidth. The Radeon Pro 580 was released on June 4, 2017, while the P104-100 followed on December 11, 2017.

The Verdict

The data points to a clear split based on workload. If the priority is raw compute performance in OpenCL or Vulkan, the NVIDIA P104-100 is the stronger choice. It wins both head-to-head benchmarks and offers substantially higher bandwidth, pixel rate, and texture rate. Its higher core clocks and larger ROP count make it effective for fill-rate-bound and bandwidth-intensive tasks. However, its PCIe 1.0 x4 interface and lack of display outputs severely limit its practical use in a standard workstation or gaming PC. It is a mining-focused card with no outputs, so it cannot drive a monitor without a secondary GPU.

The AMD Radeon Pro 580 is the better all-around GPU. Its average benchmark score is 22.2% higher, it ranks in the 82nd percentile versus the 77th, and it has double the memory (8 GB versus 4 GB). It supports FP16 at full rate, which is a massive advantage for compute tasks that use half precision. It also has a standard PCIe 3.0 x16 interface and display outputs, making it a functional workstation card. The Radeon Pro 580's nearest rivals include the RTX 5070 and Radeon Pro WX 7100, placing it in a much higher performance tier than the P104-100, whose rivals are mobile and lower-end desktop parts.

For a user building a compute-focused rig that does not need display output and can tolerate a slow PCIe interface, the P104-100 delivers more performance per dollar in specific OpenCL scenarios. For anyone needing a general-purpose GPU with broad application support, memory capacity, and modern interface compatibility, the Radeon Pro 580 is the clear winner based on the recorded data. The P104-100 wins on raw speed in the shared tests, but the Radeon Pro 580 wins on versatility and overall benchmark standing. Choose the P104-100 for dedicated compute, and the Radeon Pro 580 for everything else.

DETAILED SPECIFICATIONS

SPECIFICATION
Pro 580
P104-100
Core Specs
Shading Units
2,304
1,920 -16.7%
Shaders
2,304
1,920 -16.7%
TMUs
144
120 -16.7%
ROPs
32
64 +100.0%
Compute Units
36
SM Count
15
Clocks
Base Clock
1100 MHz
1607 MHz
Boost Clock
1200 MHz
1733 MHz
Memory Clock
1695 MHz 6.8 Gbps effective
1251 MHz 10 Gbps effective
Memory
Memory Size
8 GB
4 GB
VRAM (MB)
8,192
4,096 -50.0%
Memory Type
GDDR5
GDDR5X
Memory Bus
256 bit
256 bit
Bandwidth
217.0 GB/s
320.3 GB/s
Cache
L1 Cache
16 KB (per CU)
48 KB (per SM)
L2 Cache
2 MB
2 MB
Performance
Pixel Rate
38.40 GPixel/s
110.9 GPixel/s
Texture Rate
172.8 GTexel/s
208.0 GTexel/s
FP32 (TFLOPS)
5.530 TFLOPS
6.655 TFLOPS
FP64 (TFLOPS)
345.6 GFLOPS (1:16)
208.0 GFLOPS (1:32)
FP16 (TFLOPS)
5.530 TFLOPS (1:1)
104.0 GFLOPS (1:64)
Power
TDP
185 W
TDP (W)
185
Suggested PSU
200 W
Power Connectors
None
1x 8-pin
Architecture
Architecture
GCN 4.0
Pascal
GPU Name
Ellesmere
GP104
Generation
Radeon Pro Mac (500 Series)
Mining GPUs
Process Size
14 nm
16 nm
Transistors
5,700 million
7,200 million
Die Size
232 mm²
314 mm²
Foundry
GlobalFoundries
TSMC
Density
24.6M / mm²
22.9M / mm²
API Support
DirectX
12 (12_0)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.3
1.4
OpenCL
2.1
3.0
CUDA
6.1
Shader Model
6.7
6.8
Physical
Slot Width
IGP
Dual-slot
Length
267 mm 10.5 inches
Outputs
Portable Device Dependent
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 1.0 x4
Other
Production
End-of-life
End-of-life
View Radeon Pro 580 Details View P104-100 Details