AMD Radeon Pro Vega 20 vs NVIDIA P104-100 Comparison
AMD Radeon Pro Vega 20
P104-100
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon Pro Vega 20 vs NVIDIA P104-100
Head-to-Head Benchmarks
The recorded data shows a clear overall winner in direct comparisons. Across the two shared benchmark tests, the NVIDIA P104-100 takes both wins. The average benchmark score for the NVIDIA P104-100 is 32,982, while the AMD Radeon Pro Vega 20 sits at 27,839. That is a difference of roughly 18.5% in aggregate performance, placing the NVIDIA part in the 77th percentile of all GPUs, against the 73rd percentile for the AMD part.
The largest margin comes in the Geekbench OpenCL test. The NVIDIA P104-100 scores 52,368, while the AMD Radeon Pro Vega 20 scores 26,679. The delta is 96.3%, meaning the NVIDIA part nearly doubles the AMD part's output in this workload. That is not a marginal gap; it is a dominant performance differential. OpenCL compute tasks, which often scale with raw shading throughput and memory bandwidth, clearly favor the Pascal architecture here.
The second shared test, Geekbench Vulkan, also goes to the NVIDIA P104-100, but by a smaller margin. The score is 45,165 versus 26,410, a delta of 71%. Still a decisive win, but the reduced gap compared to OpenCL suggests the AMD Radeon Pro Vega 20 is relatively less disadvantaged in Vulkan workloads. The AMD part's GCN 5.0 architecture retains some efficiency in this API, even if it cannot close the overall performance chasm.
The NVIDIA P104-100 also has a benchmark that the AMD part cannot contest: 3DMark Steel Nomad DX12, where it records 1,413. There is no corresponding score for the AMD Radeon Pro Vega 20 in the database, so a direct comparison in that test is not possible. However, given the pattern in OpenCL and Vulkan, the NVIDIA part would likely maintain its lead in a DX12 workload as well.
The AMD Radeon Pro Vega 20 does have one exclusive benchmark of its own: Geekbench Metal, where it scores 30,427. The NVIDIA P104-100 has no Metal score recorded, which makes sense given its lack of display outputs and its mining-oriented design. This is not a performance comparison, but it highlights the AMD part's integration into Apple-centric environments.
Looking at the nearest rivals for each card provides additional context. The NVIDIA P104-100's closest competitors in average score are the NVIDIA T600 Mobile at 32,849 (0.4% lower), the NVIDIA T550 Mobile at 33,161 (0.5% higher), the NVIDIA GeForce RTX 3050 Mobile at 33,170 (0.6% higher), and the AMD Radeon Pro 570 at 33,207 (0.7% higher). The P104-100 sits almost exactly in the middle of this cluster, indicating its performance is competitive with modern mobile GPUs despite its older architecture and mining-specific heritage.
The AMD Radeon Pro Vega 20's nearest rivals include the AMD Radeon RX 7800M at 27,883 (0.2% higher), the AMD Radeon Pro W5500X at 27,973 (0.5% higher), the NVIDIA GeForce GTX 980 Ti at 28,020 (0.6% higher), and the NVIDIA GeForce RTX 3090 at 27,565 (1% lower). The presence of the RTX 3090 in this list is notable: a flagship desktop GPU from a later generation scores only 1% lower than the Vega 20 in this aggregate metric, which reflects how average scores can obscure workload-specific behavior.
Where Each One Wins
The NVIDIA P104-100 wins decisively in raw compute throughput. Its FP32 rating of 6.655 TFLOPS is more than double the AMD Radeon Pro Vega 20's 3.284 TFLOPS. The texture rate follows the same pattern: 208.0 GTexel/s against 102.6 GTexel/s. Pixel rate is also heavily in favor of the NVIDIA part, at 110.9 GPixel/s versus 41.06 GPixel/s. These figures explain the OpenCL result, where the NVIDIA part nearly doubles the AMD part's score. Any workload that stresses shader math, texture filtering, or rasterization will favor the P104-100.
Memory bandwidth is another area where the NVIDIA P104-100 leads, though the difference is less extreme than the compute gap. The P104-100 delivers 320.3 GB/s over a 256-bit bus using GDDR5X, while the AMD part manages 189.4 GB/s over a 1024-bit bus using HBM2. The NVIDIA part has 69% more bandwidth, which supports its advantage in memory-intensive compute tasks. The AMD part's wider bus cannot compensate for its lower memory clock and smaller effective data rate.
The AMD Radeon Pro Vega 20 wins in power efficiency. Its TDP is 100 W, while the NVIDIA P104-100 has no recorded TDP in the database, but the suggested PSU for the NVIDIA part is 200 W, which implies a substantially higher power draw. The AMD part also benefits from being an IGP (integrated graphics processor) form factor, whereas the NVIDIA part is a dual-slot card requiring a single 8-pin power connector. For mobile or power-constrained deployments, the Vega 20 is the more practical choice.
The AMD part also wins in FP16 throughput. It delivers 6.569 TFLOPS with a 2:1 ratio relative to FP32, whereas the NVIDIA P104-100 delivers only 104.0 GFLOPS with a 1:64 ratio. For workloads that can exploit half-precision math, such as certain machine learning inference tasks, the Vega 20 is vastly superior. This is a niche advantage, but it is a real one in the database's numbers.
The AMD Radeon Pro Vega 20 also has a display output path, listed as "Portable Device Dependent," whereas the NVIDIA P104-100 has no outputs at all. The P104-100 was designed for mining, not for driving displays, so any workload that requires visual output cannot use it. The Vega 20, being a Radeon Pro part for Mac systems, is intended to drive portable device displays.
Architecture Differences
The two GPUs come from different architectural generations and design philosophies. The NVIDIA P104-100 uses the GP104 chip, built on the Pascal architecture, manufactured by TSMC on a 16 nm process. The die measures 314 mm² and contains 7,200 million transistors, yielding a transistor density of 22.9 million per square millimeter. Pascal is a mature, high-clock architecture that favors raw throughput and power efficiency at moderate power envelopes.
The AMD Radeon Pro Vega 20 uses the Vega 12 chip, built on the GCN 5.0 architecture, manufactured by GlobalFoundries on a 14 nm process. The database does not record a die size or transistor count for this chip, so a direct density comparison is not possible. GCN 5.0 introduced improved geometry processing and a new memory controller for HBM2, but it retains the fundamental GCN compute-unit design that dates back earlier in AMD's lineup.
The clock behavior differs significantly. The NVIDIA P104-100 has a base clock of 1607 MHz and a boost clock of 1733 MHz, which are high for its generation. The AMD part has a base clock of 815 MHz and a boost clock of 1283 MHz, much lower. This clock gap, combined with the NVIDIA part's higher shading unit count (1920 versus 1280), explains the large FP32 difference. The AMD part compensates with its FP16 capability, which the NVIDIA part almost entirely lacks.
Memory architecture is another major divergence. The NVIDIA P104-100 uses 4 GB of GDDR5X on a 256-bit bus, with memory running at 1251 MHz (10 Gbps effective) for 320.3 GB/s. The AMD part uses 4 GB of HBM2 on a 1024-bit bus, with memory at 740 MHz (1480 Mbps effective) for 189.4 GB/s. HBM2's wide bus is a different approach to bandwidth, but in this implementation the lower clock limits its effectiveness. The bus interface also differs: the NVIDIA part uses PCIe 1.0 x4, while the AMD part uses PCIe 3.0 x16. This is unusual for the NVIDIA card, as PCIe 1.0 x4 is a severe bottleneck for data transfer from the host system, though it matters little for mining workloads that keep data on the GPU.
The API support is nearly identical. Both support DirectX 12 (12_1) and OpenGL 4.6. The difference is in Vulkan: the NVIDIA part supports Vulkan 1.4, while the AMD part supports Vulkan 1.3. This may explain why the Vulkan benchmark gap is smaller than OpenCL, as the AMD driver's Vulkan implementation is still mature despite the older API version.
Specification Differences
The NVIDIA P104-100 and AMD Radeon Pro Vega 20 differ across nearly every recorded specification. The process node is 16 nm for the NVIDIA part and 14 nm for the AMD part. The transistor count is 7,200 million for the NVIDIA part, with no recorded value for the AMD part. Die size is 314 mm² for the NVIDIA part, no recorded value for the AMD part. Transistor density is 22.9M per mm² for the NVIDIA part, no recorded value for the AMD part.
Clock speeds: the NVIDIA part has a base of 1607 MHz and boost of 1733 MHz, while the AMD part has a base of 815 MHz and boost of 1283 MHz. Memory clock is 1251 MHz (10 Gbps effective) for the NVIDIA part versus 740 MHz (1480 Mbps effective) for the AMD part. Memory type is GDDR5X for the NVIDIA part, HBM2 for the AMD part. Bus width is 256 bit versus 1024 bit. Bandwidth is 320.3 GB/s versus 189.4 GB/s.
Compute resources: shading units are 1920 versus 1280, TMUs are 120 versus 80, ROPs are 64 versus 32. Pixel rate is 110.9 GPixel/s versus 41.06 GPixel/s. Texture rate is 208.0 GTexel/s versus 102.6 GTexel/s. FP32 is 6.655 TFLOPS versus 3.284 TFLOPS. FP16 is 104.0 GFLOPS (1:64) versus 6.569 TFLOPS (2:1).
Power and form factor: the NVIDIA part has no recorded TDP, while the AMD part has a TDP of 100 W. The NVIDIA part is dual-slot with a single 8-pin power connector and a suggested PSU of 200 W. The AMD part is an IGP with no power connectors and no suggested PSU. The NVIDIA part has no display outputs; the AMD part has "Portable Device Dependent" outputs. The NVIDIA part uses PCIe 1.0 x4; the AMD part uses PCIe 3.0 x16. The NVIDIA part measures 267 mm (10.5 inches) in length; the AMD part has no recorded dimensions. The NVIDIA part supports Vulkan 1.4; the AMD part supports Vulkan 1.3.
FAQ
Q: Which GPU is faster in OpenCL compute?
A: The NVIDIA P104-100 scores 52,368 in Geekbench OpenCL, which is 96.3% higher than the AMD Radeon Pro Vega 20's 26,679. The NVIDIA part's FP32 throughput of 6.655 TFLOPS versus 3.284 TFLOPS supports this result.
Q: Does the AMD Radeon Pro Vega 20 have any performance advantage?
A: Yes, in FP16 compute. The Vega 20 delivers 6.569 TFLOPS with a 2:1 FP16 ratio, while the NVIDIA P104-100 delivers only 104.0 GFLOPS with a 1:64 ratio. The AMD part also has a lower TDP of 100 W compared to the NVIDIA part's suggested 200 W PSU.
Q: Why is the NVIDIA P104-100's memory bandwidth so much higher?
A: The P104-100 uses GDDR5X at 1251 MHz (10 Gbps effective) on a 256-bit bus, producing 320.3 GB/s. The Vega 20 uses HBM2 at 740 MHz (1480 Mbps effective) on a 1024-bit bus, producing 189.4 GB/s. The higher memory clock on the NVIDIA part overcomes the AMD part's wider bus.
Q: Can the NVIDIA P104-100 output to a display?
A: No. The database lists "No outputs" for the P104-100, which reflects its mining-focused design. The AMD Radeon Pro Vega 20 has "Portable Device Dependent" display outputs, meaning it can drive displays in portable systems.
Q: How do these GPUs compare to their nearest rivals?
A: The NVIDIA P104-100's average score of 32,982 is within 0.7% of its nearest rivals, which include the NVIDIA T600 Mobile, T550 Mobile, RTX 3050 Mobile, and AMD Radeon Pro 570. The AMD Vega 20's average score of 27,839 is within 1% of its nearest rivals, including the AMD Radeon RX 7800M, Radeon Pro W5500X, GTX 980 Ti, and RTX 3090.
Q: Which API versions do each support?
A: Both support DirectX 12 (12_1) and OpenGL 4.6. The NVIDIA P104-100 supports Vulkan 1.4, while the AMD Radeon Pro Vega 20 supports Vulkan 1.3.