AMD Radeon PRO W6400 vs NVIDIA P104-100 Comparison
AMD Radeon PRO W6400
P104-100
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon PRO W6400 vs NVIDIA P104-100
Head-to-Head Benchmarks
The recorded data contains two direct head-to-head comparisons between the AMD Radeon PRO W6400 and the NVIDIA P104-100, both of which are OpenCL and Vulkan compute workloads. The NVIDIA P104-100 wins both tests, but the margin is very different between the two APIs.
In Geekbench OpenCL, the NVIDIA P104-100 scores 52,368 against the AMD Radeon PRO W6400’s 35,027. That is a 33.1% lead for the NVIDIA card. This is a substantial gap, and it reflects the raw compute throughput difference between the two chips. The NVIDIA card’s FP32 rating is 6.655 TFLOPS, while the AMD card sits at 3.565 TFLOPS, so the OpenCL result aligns closely with that theoretical peak. The NVIDIA card also has far more shading units (1,920 vs 768) and TMUs (120 vs 48), which helps in heavily parallel workloads.
In Geekbench Vulkan, the gap narrows considerably. The NVIDIA P104-100 scores 45,165, while the AMD Radeon PRO W6400 scores 39,286. That is a 13% lead, less than half the OpenCL margin. The AMD card’s Vulkan driver and architecture appear to close some of the gap in this API, likely due to its newer RDNA 2.0 design and better modern API efficiency. Still, the NVIDIA card wins outright.
Looking at the broader database context, the AMD Radeon PRO W6400 has an average benchmark score of 37,157 across all recorded tests, while the NVIDIA P104-100 has an average of 32,982. That means the NVIDIA card’s head-to-head wins in these two tests are not representative of its overall standing. The AMD card actually sits higher in the all-GPU percentile ranking, at the 80th percentile versus the NVIDIA card’s 77th percentile. The reason is that the NVIDIA P104-100’s average is dragged down by a poor showing in 3DMark Steel Nomad DX12, where it scores only 1,413. That test is not available for the AMD card in the database, so the comparison is incomplete, but it explains why the NVIDIA card’s average is lower despite winning both Geekbench tests.
The AMD card’s nearest rivals in the database are the AMD Radeon RX Vega 56 (average 37,507, delta -0.9%), the NVIDIA Tesla P4 (37,628, -1.3%), the NVIDIA GeForce RTX 4070 (37,648, -1.3%), and the NVIDIA GeForce GTX TITAN X (36,530, +1.7%). The AMD card sits within roughly 1.7% of all of these, which shows it is a very consistent mid-pack performer. The NVIDIA P104-100’s nearest rivals are the NVIDIA T600 Mobile (32,849, +0.4%), the NVIDIA T550 Mobile (33,161, -0.5%), the NVIDIA GeForce RTX 3050 Mobile (33,170, -0.6%), and the AMD Radeon Pro 570 (33,207, -0.7%). Those deltas are all under 1%, meaning the NVIDIA card is essentially tied with those mobile and older workstation parts in average score.
So the head-to-head picture is clear: the NVIDIA P104-100 is faster in both recorded compute tests, but the AMD Radeon PRO W6400 has a better overall average benchmark score because it does not have the same weakness in DX12 workloads that drags the NVIDIA card down.
Architecture Differences
The two cards come from completely different eras and design philosophies. The AMD Radeon PRO W6400 is built on RDNA 2.0, using the Navi 24 chip, manufactured on a 6 nm process at TSMC. The NVIDIA P104-100 is built on the older Pascal architecture, using the GP104 chip, on a 16 nm process, also at TSMC. The process node difference is significant: 6 nm versus 16 nm means the AMD chip is far denser. The AMD chip packs 5,400 million transistors into a 107 mm² die, giving a transistor density of 50.5 million per mm². The NVIDIA chip has 7,200 million transistors on a 314 mm² die, which works out to only 22.9 million per mm². So the AMD chip is more than twice as dense per square millimeter.
The memory subsystems are also very different. The AMD card uses 4 GB of GDDR6 on a 64-bit bus, delivering 128.0 GB/s of bandwidth. The NVIDIA card also has 4 GB of memory, but it is GDDR5X on a 256-bit bus, delivering 320.3 GB/s. That is a 2.5x bandwidth advantage for the NVIDIA card, which is a major factor in compute workloads that are memory-bandwidth bound. The NVIDIA card also has a much wider memory bus, which helps with large data sets.
The compute resources are heavily skewed toward the NVIDIA card. It has 1,920 shading units, 120 TMUs, and 64 ROPs, versus 768 shading units, 48 TMUs, and 32 ROPs on the AMD card. The NVIDIA card’s pixel rate is 110.9 GPixel/s and its texture rate is 208.0 GTexel/s, while the AMD card manages 74.27 GPixel/s and 111.4 GTexel/s. The FP32 throughput is 6.655 TFLOPS for NVIDIA versus 3.565 TFLOPS for AMD. However, the AMD card has a major feature the NVIDIA card lacks entirely: 12 ray tracing cores. The NVIDIA P104-100 has no ray tracing hardware at all. The AMD card also supports DirectX 12 Ultimate (12_2), while the NVIDIA card only supports DirectX 12 (12_1). Both support OpenGL 4.6 and Vulkan 1.4.
The AMD card’s FP16 throughput is 7.130 TFLOPS (2:1 ratio), which is actually higher than its FP32. The NVIDIA card’s FP16 is a tiny 104.0 GFLOPS (1:64 ratio), which means it is effectively not designed for half-precision compute. That is a huge difference for AI or machine learning workloads that use FP16.
The power and physical requirements are also very different. The AMD card has a TDP of 50 W and requires no power connectors, running off the PCIe slot alone. It is single-slot. The NVIDIA card has no recorded TDP but requires a single 8-pin power connector and is dual-slot. The AMD card lists a suggested PSU of 250 W, while the NVIDIA card lists 200 W. The AMD card uses a PCIe 4.0 x4 interface, while the NVIDIA card uses PCIe 1.0 x4, which is a very old and slow bus standard. The NVIDIA card has no display outputs at all, while the AMD card has two DisplayPort 1.4a outputs. The NVIDIA card is 267 mm long (10.5 inches); the AMD card’s dimensions are not recorded.
Where Each One Wins
The NVIDIA P104-100 wins in raw compute throughput. Its OpenCL score is 33.1% higher, and its Vulkan score is 13% higher. If your workload is purely about shader math, FP32 throughput, or memory bandwidth, the NVIDIA card has a clear edge. The 320.3 GB/s bandwidth versus 128.0 GB/s is a massive advantage for large buffer operations. The NVIDIA card also has 2.5x the TMUs and 2x the ROPs, so texture-heavy and fill-rate-heavy tasks will favor it.
The AMD Radeon PRO W6400 wins in overall benchmark standing, with an 80th percentile rank versus 77th for the NVIDIA card. Its average benchmark score of 37,157 is 12.7% higher than the NVIDIA card’s 32,982. That suggests the AMD card is more consistent across a wider range of workloads, especially those that stress modern APIs. The AMD card also has ray tracing cores, which the NVIDIA card completely lacks. Any workload that uses DirectX 12 Ultimate features or ray tracing will simply not run on the NVIDIA card. The AMD card also has a much better FP16 ratio (2:1 versus 1:64), so half-precision compute is viable on the AMD card and effectively not on the NVIDIA card.
The AMD card is also far more practical from a system integration standpoint. It is a single-slot card with no power connectors, a 50 W TDP, and a 250 W suggested PSU. The NVIDIA card is dual-slot, requires an 8-pin connector, and has no display outputs. If you need to output video, the AMD card is the only choice here, as the NVIDIA card is purely a compute or mining board. The AMD card also uses a PCIe 4.0 x4 interface versus the NVIDIA card’s PCIe 1.0 x4, which is a massive difference in bus bandwidth. Even if the NVIDIA card has more memory bandwidth internally, the PCIe bus connection to the host system is far slower on the NVIDIA card.
The NVIDIA card wins on raw compute density per watt in some sense, but the AMD card wins on efficiency in terms of power draw. The AMD card’s 50 W TDP is extraordinarily low for the performance it delivers. The NVIDIA card’s TDP is not recorded, but with a dual-slot cooler and 8-pin connector, it clearly draws far more power.
The Verdict
Pick the NVIDIA P104-100 if your workload is purely about maximum compute throughput in OpenCL or Vulkan, and you do not need display output, ray tracing, or modern DirectX 12 Ultimate features. The data shows a 33.1% lead in OpenCL and a 13% lead in Vulkan. It also has 2.5x the memory bandwidth and far more shading units, TMUs, and ROPs. This is a compute-oriented board that will handle heavy FP32 workloads well.
Pick the AMD Radeon PRO W6400 if you need a card that can do more than compute. It has display outputs, supports DirectX 12 Ultimate, has ray tracing cores, and uses a modern PCIe 4.0 interface. Its overall average benchmark score is higher, and it sits at a better percentile rank. It is also a single-slot, low-power card that requires no external power connectors, making it far easier to install in a workstation or small form factor system. The 50 W TDP and 250 W suggested PSU are very modest requirements.
The data does not support a single clear winner for all use cases. The NVIDIA card is strictly faster in the two recorded head-to-head tests, but its average score suffers from a weak DX12 result. The AMD card is the more versatile and modern product, and its higher average score reflects that. If you are building a workstation that needs to output video, use modern APIs, or run ray-traced workloads, the AMD card is the obvious choice. If you are building a dedicated compute node with no display requirements and you care only about raw OpenCL or Vulkan performance, the NVIDIA card wins.
FAQ
Q: Which card has a higher Geekbench OpenCL score?
A: The NVIDIA P104-100 scores 52,368, while the AMD Radeon PRO W6400 scores 35,027. That is a 33.1% lead for the NVIDIA card.
Q: Does the AMD card have ray tracing support?
A: Yes, the AMD Radeon PRO W6400 has 12 ray tracing cores. The NVIDIA P104-100 has no ray tracing cores.
Q: Which card has more memory bandwidth?
A: The NVIDIA P104-100 has 320.3 GB/s, while the AMD Radeon PRO W6400 has 128.0 GB/s. The NVIDIA card has a 256-bit bus with GDDR5X, while the AMD card uses a 64-bit bus with GDDR6.
Q: Can the NVIDIA P104-100 output video?
A: No, it has no display outputs. The AMD Radeon PRO W6400 has two DisplayPort 1.4a outputs.
Q: Which card has the higher average benchmark score?
A: The AMD Radeon PRO W6400 has an average score of 37,157, while the NVIDIA P104-100 has an average of 32,982. The AMD card also ranks at the 80th percentile versus the NVIDIA card’s 77th percentile.
Q: What is the power connector requirement for each card?
A: The AMD Radeon PRO W6400 has no power connectors and a 50 W TDP. The NVIDIA P104-100 requires a single 8-pin connector and has no recorded TDP.
Specification Differences
| Specification | AMD Radeon PRO W6400 | NVIDIA P104-100 |
| --- | --- | --- |
| Architecture | RDNA 2.0 | Pascal |
| Process node | 6 nm | 16 nm |
| Transistors | 5,400 million | 7,200 million |
| Die size | 107 mm² | 314 mm² |
| Transistor density | 50.5M / mm² | 22.9M / mm² |
| Base clock | 2039 MHz | 1607 MHz |
| Boost clock | 2321 MHz | 1733 MHz |
| Memory clock | 2000 MHz, 16 Gbps effective | 1251 MHz, 10 Gbps effective |
| Memory size | 4 GB | 4 GB |
| Memory type | GDDR6 | GDDR5X |
| Memory bus width | 64 bit | 256 bit |
| Memory bandwidth | 128.0 GB/s | 320.3 GB/s |
| Shading units | 768 | 1920 |
| TMUs | 48 | 120 |
| ROPs | 32 | 64 |
| Ray tracing cores | 12 | None |
| Pixel rate | 74.27 GPixel/s | 110.9 GPixel/s |
| Texture rate | 111.4 GTexel/s | 208.0 GTexel/s |
| FP32 | 3.565 TFLOPS | 6.655 TFLOPS |
| FP16 | 7.130 TFLOPS (2:1) | 104.0 GFLOPS (1:64) |
| TDP | 50 W | Not recorded |
| Slot width | Single-slot | Dual-slot |
| Power connectors | None | 1x 8-pin |
| Suggested PSU | 250 W | 200 W |
| Bus interface | PCIe 4.0 x4 | PCIe 1.0 x4 |
| Display outputs | 2x DisplayPort 1.4a | No outputs |
| DirectX support | 12 Ultimate (12_2) | 12 (12_1) |
| Release date | 2022-01-18 | 2017-12-11 |