AMD Radeon Pro 580 vs NVIDIA GeForce RTX 4070 Comparison

AMD
RADEON

AMD Radeon Pro 580

CORE STATE Ellesmere
VRAM 8 GB
CLOCK SPEED 1200 MHz
TDP 185 W
BUS WIDTH 256 bit
ARCHITECTURE GCN 4.0
nm
PROCESS 14 nm
LAUNCH DATE 2017
VS
NVIDIA
GEFORCE

GeForce RTX 4070

CORE STATE AD104
VRAM 12 GB
CLOCK SPEED 2475 MHz
TDP 200 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023

PERFORMANCE BENCHMARKS

geekbench_metal
39,213
N/A
geekbench_opencl
38,457
154,858
geekbench_vulkan
43,285
174,152
3dmark_3dmark_steel_nomad_dx12
N/A
3,854
passmark_directx_10
N/A
139
passmark_directx_11
N/A
244
passmark_directx_12
N/A
103
passmark_directx_9
N/A
320
passmark_g2d
N/A
1,164
passmark_g3d
N/A
26,927
passmark_gpu_compute
N/A
14,720

Analysis: AMD Radeon Pro 580 vs NVIDIA GeForce RTX 4070

The AMD Radeon Pro 580 and NVIDIA GeForce RTX 4070 occupy vastly different positions in the hardware landscape, separated by six years of architectural evolution. The data shows a decisive performance gulf, with the RTX 4070 winning every head-to-head benchmark recorded. However, the Radeon Pro 580's legacy status and the RTX 4070's modern feature set make for a comparison that extends beyond raw frame rates, touching on compute workloads, API support, and the fundamental design philosophies of their respective generations.

Head-to-Head Benchmarks

The head-to-head results are unambiguous. In the Geekbench OpenCL test, the NVIDIA GeForce RTX 4070 scores 154,858, while the AMD Radeon Pro 580 manages 38,457. This represents a 75.2% deficit for the AMD card, meaning the RTX 4070 delivers approximately four times the compute throughput in this workload. The margin is nearly identical in Geekbench Vulkan, where the RTX 4070 scores 174,152 against the Pro 580's 43,285, again a 75.1% difference. These are not close contests; they are generational slaughters.

Interpreting these deltas requires context. The RTX 4070's OpenCL score is roughly 4.03x higher than the Pro 580's, while its Vulkan score is approximately 4.02x higher. Such a consistent ratio across two different APIs suggests the performance gap is structural rather than workload-specific. The RTX 4070's advantage stems from its newer architecture, higher clock speeds, and significantly larger shader count, which we will examine in the Architecture Differences section. For now, it is enough to state that in any compute task leveraging these APIs, the RTX 4070 will finish the job in roughly a quarter of the time required by the Radeon Pro 580.

The aggregate benchmark scores reinforce this narrative. The Radeon Pro 580's average benchmark score is 40,318, placing it at the 82nd percentile of all GPUs. The RTX 4070's average is 37,648, which sits at the 81st percentile. This near-identical percentile ranking is initially surprising given the head-to-head results, but it reflects the different benchmark suites each card was tested with. The Pro 580's average is derived from its three Geekbench scores, while the RTX 4070's includes a wider mix of PassMark and 3DMark tests. The RTX 4070's PassMark G3D score of 26,927 and its 3DMark Steel Nomad DX12 score of 3,854 are not directly comparable to the Pro 580's Geekbench numbers, but they paint a picture of a card that excels in modern, API-heavy workloads.

Looking at the nearest rivals for each card provides additional insight. The Radeon Pro 580's closest competitor is the NVIDIA GeForce RTX 5070, with an average score of 40,377, a mere 0.1% difference. This is remarkable: a 2017 workstation card is statistically tied with a next-generation consumer GPU in aggregate compute scores. The Pro 580 also edges out the AMD Radeon Pro WX 7100 by 0.6% and the NVIDIA RTX A500 Mobile by 1.9%, while trailing the AMD Radeon Pro 5300 by 1.4%. For the RTX 4070, its nearest rival is the NVIDIA Tesla P4, with a 0.1% delta, followed by the AMD Radeon RX Vega 56 at 0.4% and the AMD Radeon PRO W6400 at 1.3%. The RTX 4070 trails the NVIDIA GeForce RTX 4080 Mobile by 1.3%. These rivalries show that while the RTX 4070 dominates the Pro 580, its own standing among contemporaries is competitive rather than supreme.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The AMD Radeon Pro 580 has an average benchmark score of 40,318, which is higher than the NVIDIA GeForce RTX 4070's 37,648. However, this aggregate figure is based on different benchmark suites, and the RTX 4070 wins every direct head-to-head comparison.

Q: What is the largest performance margin in the head-to-head tests?

A: The largest margin is in the Geekbench OpenCL test, where the NVIDIA GeForce RTX 4070 leads by 75.2%. The RTX 4070 scores 154,858 versus the Radeon Pro 580's 38,457.

Q: How does the Radeon Pro 580 compare to the RTX 5070?

A: The Radeon Pro 580's average score of 40,318 is 0.1% lower than the NVIDIA GeForce RTX 5070's 40,377. This places the two cards as statistical equals in aggregate compute performance.

Q: What is the RTX 4070's best result in the PassMark suite?

A: The RTX 4070's highest PassMark score is 26,927 in the G3D test. It also achieves 14,720 in GPU Compute, 1,164 in G2D, 320 in DirectX 9, 244 in DirectX 11, 139 in DirectX 10, and 103 in DirectX 12.

Q: Which card has a higher percentile ranking among all GPUs?

A: The AMD Radeon Pro 580 sits at the 82nd percentile, while the NVIDIA GeForce RTX 4070 sits at the 81st. The Pro 580 ranks one percentile point higher despite losing all direct comparisons.

Q: Does the RTX 4070 have any benchmark scores below 100?

A: Yes, the RTX 4070 scores 103 in the PassMark DirectX 12 test, which is its lowest recorded benchmark result. This is notably lower than its DirectX 11 score of 244.

Architecture Differences

The architectural chasm between these two GPUs is the root cause of their performance disparity. The AMD Radeon Pro 580 is built on the Ellesmere chip using the GCN 4.0 architecture, fabricated on a 14 nm process at GlobalFoundries. This is a mature design from the Radeon Pro Mac (500 Series) generation, with 5,700 million transistors packed into a 232 mm² die, yielding a transistor density of 24.6M per mm². In contrast, the NVIDIA GeForce RTX 4070 uses the AD104 chip with the Ada Lovelace architecture, manufactured by TSMC on a 5 nm process. It contains 35,800 million transistors on a 294 mm² die, achieving a much higher density of 121.8M per mm². The RTX 4070 packs over six times the transistors into a die only 27% larger.

The compute resources differ dramatically. The Radeon Pro 580 has 2,304 shading units, 144 texture mapping units (TMUs), and 32 raster output units (ROPs). The RTX 4070 field 5,888 shading units, 184 TMUs, and 64 ROPs. This means the RTX 4070 has 2.55x the shading units, 1.28x the TMUs, and 2x the ROPs. The RTX 4070 also introduces dedicated hardware absent from the Pro 580: 46 ray tracing cores and 184 tensor cores. These enable hardware-accelerated ray tracing and AI-driven features like DLSS, which the GCN 4.0 architecture cannot offer.

Clock speeds and memory technology further separate the two. The Pro 580 operates at a base clock of 1100 MHz and a boost of 1200 MHz, while the RTX 4070 runs at 1920 MHz base and 2475 MHz boost, representing a 75% higher base clock and 106% higher boost clock. Memory configurations are equally divergent: the Pro 580 uses 8 GB of GDDR5 on a 256-bit bus, delivering 217.0 GB/s of bandwidth. The RTX 4070 uses 12 GB of GDDR6X on a 192-bit bus, achieving 504.2 GB/s. Despite a narrower bus, the RTX 4070's faster memory (21 Gbps effective versus 6.8 Gbps) provides 2.32x the bandwidth. The theoretical peak compute figures tell the same story: the Pro 580 delivers 5.530 TFLOPS in both FP32 and FP16 (1:1), while the RTX 4070 delivers 29.15 TFLOPS in both, a 5.27x advantage.

The Verdict

The data is unequivocal: the NVIDIA GeForce RTX 4070 is the superior performer in every direct benchmark. Its OpenCL and Vulkan scores are over four times higher than the AMD Radeon Pro 580's, and its architectural advantages are overwhelming. Anyone needing maximum compute throughput, ray tracing support, or modern API compliance should choose the RTX 4070 without hesitation. The RTX 4070 supports DirectX 12 Ultimate (12_2) and Vulkan 1.4, while the Pro 580 is limited to DirectX 12 (12_0) and Vulkan 1.3. The RTX 4070 also offers a 550 W suggested PSU requirement, 1x 16-pin power connector, and a dual-slot form factor, making it a discrete, installable card.

However, the AMD Radeon Pro 580 is not without merit. Its average benchmark score of 40,318 actually exceeds the RTX 4070's 37,648, and it ranks one percentile higher (82nd versus 81st). Its nearest rival is the RTX 5070, a much newer card, with only a 0.1% difference. This suggests that in aggregated, mixed workloads, the Pro 580 remains competitive with modern hardware. It is also an integrated GPU (IGP) with no power connectors and a portable-device-dependent display output, making it ideal for compact or mobile systems where the RTX 4070's 240 mm length and 40 mm width would not fit. The Pro 580's 185 W TDP is also lower than the RTX 4070's 200 W, though both are relatively modest for their performance classes.

The choice depends on use case. For a workstation or desktop build where space and power are available, the RTX 4070 is categorically better. Its 12 GB GDDR6X memory, 46 RT cores, and 184 tensor cores open up workloads the Pro 580 cannot touch. For an integrated, low-power system from the 2017 era, the Pro 580 remains a capable compute engine, as evidenced by its parity with the RTX 5070 in aggregate scores. The RTX 4070 wins on raw performance; the Pro 580 wins on integration and legacy compatibility.

Specification Differences

The following table highlights only the fields where the two cards differ. Identical fields, such as OpenGL 4.6 support and a PCIe x16 bus interface, are omitted.

| Specification | AMD Radeon Pro 580 | NVIDIA GeForce RTX 4070 |

|---|---|---|

| Chip | Ellesmere | AD104 |

| Architecture | GCN 4.0 | Ada Lovelace |

| Generation | Radeon Pro Mac (500 Series) | GeForce 40 |

| Process Node | 14 nm | 5 nm |

| Foundry | GlobalFoundries | TSMC |

| Transistors | 5,700 million | 35,800 million |

| Die Size | 232 mm² | 294 mm² |

| Transistor Density | 24.6M / mm² | 121.8M / mm² |

| Base Clock | 1100 MHz | 1920 MHz |

| Boost Clock | 1200 MHz | 2475 MHz |

| Memory Clock | 1695 MHz / 6.8 Gbps effective | 1313 MHz / 21 Gbps effective |

| Memory Size | 8 GB | 12 GB |

| Memory Type | GDDR5 | GDDR6X |

| Memory Bus Width | 256 bit | 192 bit |

| Memory Bandwidth | 217.0 GB/s | 504.2 GB/s |

| Shading Units | 2304 | 5888 |

| TMUs | 144 | 184 |

| ROPs | 32 | 64 |

| RT Cores | None | 46 |

| Tensor Cores | None | 184 |

| Pixel Rate | 38.40 GPixel/s | 158.4 GPixel/s |

| Texture Rate | 172.8 GTexel/s | 455.4 GTexel/s |

| FP32 | 5.530 TFLOPS | 29.15 TFLOPS |

| FP16 | 5.530 TFLOPS (1:1) | 29.15 TFLOPS (1:1) |

| TDP | 185 W | 200 W |

| Slot Width | IGP | Dual-slot |

| Power Connectors | None | 1x 16-pin |

| Suggested PSU | None | 550 W |

| Bus Interface | PCIe 3.0 x16 | PCIe 4.0 x16 |

| Display Outputs | Portable Device Dependent | 1x HDMI 2.1, 3x DisplayPort 1.4a |

| DirectX | 12 (12_0) | 12 Ultimate (12_2) |

| Vulkan | 1.3 | 1.4 |

| Dimensions | Not specified | 240 mm x 110 mm x 40 mm |

| Release Date | 2017-06-04 | 2023-04-11 |

| Predecessor | None | GeForce 30 |

| Successor | None | GeForce 50 |

| Launch MSRP | None | 599 USD |

| Production Status | End-of-life | End-of-life |

DETAILED SPECIFICATIONS

SPECIFICATION
Pro 580
RTX 4070
Core Specs
Shading Units
2,304
5,888 +155.6%
Shaders
2,304
5,888 +155.6%
TMUs
144
184 +27.8%
ROPs
32
64 +100.0%
Compute Units
36
SM Count
46
Clocks
Base Clock
1100 MHz
1920 MHz
Boost Clock
1200 MHz
2475 MHz
Memory Clock
1695 MHz 6.8 Gbps effective
1313 MHz 21 Gbps effective
Memory
Memory Size
8 GB
12 GB
VRAM (MB)
8,192
12,288 +50.0%
Memory Type
GDDR5
GDDR6X
Memory Bus
256 bit
192 bit
Bandwidth
217.0 GB/s
504.2 GB/s
Cache
L1 Cache
16 KB (per CU)
128 KB (per SM)
L2 Cache
2 MB
36 MB
Performance
Pixel Rate
38.40 GPixel/s
158.4 GPixel/s
Texture Rate
172.8 GTexel/s
455.4 GTexel/s
FP32 (TFLOPS)
5.530 TFLOPS
29.15 TFLOPS
FP64 (TFLOPS)
345.6 GFLOPS (1:16)
455.4 GFLOPS (1:64)
FP16 (TFLOPS)
5.530 TFLOPS (1:1)
29.15 TFLOPS (1:1)
AI/RT
RT Cores
46
Tensor Cores
184
Power
TDP
185 W
200 W
TDP (W)
185
200 +8.1%
Suggested PSU
550 W
Power Connectors
None
1x 16-pin
Architecture
Architecture
GCN 4.0
Ada Lovelace
GPU Name
Ellesmere
AD104
Generation
Radeon Pro Mac (500 Series)
GeForce 40
Process Size
14 nm
5 nm
Transistors
5,700 million
35,800 million
Die Size
232 mm²
294 mm²
Foundry
GlobalFoundries
TSMC
Density
24.6M / mm²
121.8M / mm²
API Support
DirectX
12 (12_0)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.3
1.4
OpenCL
2.1
3.0
CUDA
8.9
Shader Model
6.7
6.8
Physical
Slot Width
IGP
Dual-slot
Length
240 mm 9.4 inches
Height
110 mm 4.3 inches
Outputs
Portable Device Dependent
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 3.0 x16
PCIe 4.0 x16
Other
Launch Price
599 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 30
Successor
GeForce 50
View Radeon Pro 580 Details View GeForce RTX 4070 Details