AMD Radeon Pro WX 9100 vs NVIDIA Tesla P40 Comparison
AMD Radeon Pro WX 9100
Tesla P40
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon Pro WX 9100 vs NVIDIA Tesla P40
The NVIDIA Tesla P40 and AMD Radeon Pro WX 9100 are both end-of-life workstation cards that land in the same performance tier, yet they achieve that status through radically different designs. Benchmark results indicate a split decision: the AMD card wins in OpenCL compute by 6.9%, while the NVIDIA card dominates in Vulkan by a massive 24.6%. The data shows a 1-1 tie in benchmark victories, with the Tesla P40 holding a 1.4% higher average score across all tests, placing both in the 89th percentile of all GPUs.
FAQ
Q: Which card has the higher average benchmark score?
A: The NVIDIA Tesla P40 has an average benchmark score of 65095, which is 1.4% higher than the AMD Radeon Pro WX 9100’s average of 64212.
Q: How do the two cards compare in OpenCL performance?
A: The AMD Radeon Pro WX 9100 wins the OpenCL test decisively, scoring 66605 against the Tesla P40’s 62017, a 6.9% advantage.
Q: Which card wins in Vulkan performance?
A: The NVIDIA Tesla P40 is the clear Vulkan winner, scoring 68172 compared to the WX 9100’s 54711, which translates to a 24.6% lead for NVIDIA.
Q: What are the memory configurations of these two cards?
A: The Tesla P40 features 24 GB of GDDR5 memory on a 384-bit bus, while the WX 9100 has 16 GB of HBM2 memory on a much wider 2048-bit bus.
Q: Do these cards have any display outputs?
A: No, the NVIDIA Tesla P40 has no display outputs, while the AMD Radeon Pro WX 9100 provides 6x mini-DisplayPort 1.4a connectors.
Q: What is the launch MSRP for each card?
A: The NVIDIA Tesla P40 had a launch MSRP of 5,699 USD, while the AMD Radeon Pro WX 9100 launched with a 1,599 USD price tag.
Architecture Differences
The architectural divide between these two cards is stark. The Tesla P40 is built on NVIDIA’s Pascal architecture, fabricated on a 16 nm process at TSMC. In contrast, the WX 9100 uses AMD’s GCN 5.0 architecture, built on a 14 nm process at GlobalFoundries. The process difference is minor, but the underlying designs are fundamentally different approaches to compute.
The NVIDIA GP102 chip contains 11,800 million transistors on a 471 mm² die, resulting in a transistor density of 25.1M per mm². AMD’s Vega 10 chip is slightly larger at 495 mm² and packs 12,500 million transistors, yielding a density of 25.3M per mm². The transistor counts are close, but the AMD chip allocates its resources differently, with 4096 shading units and 256 texture mapping units, versus the NVIDIA chip’s 3840 shading units and 240 TMUs. The ROP count favors NVIDIA, however, with 96 ROPs compared to AMD’s 64.
Memory architecture represents the most significant divergence. The Tesla P40 uses 24 GB of GDDR5 with a 384-bit bus, delivering 347.1 GB/s of bandwidth. The WX 9100 counters with 16 GB of HBM2 on a 2048-bit bus, achieving a substantially higher 483.8 GB/s. This bandwidth advantage is a key factor in the AMD card’s OpenCL performance. The memory clocks tell the story: NVIDIA’s memory runs at 1808 MHz (7.2 Gbps effective), while AMD’s runs at 945 MHz (1890 Mbps effective), relying on the wider bus to achieve higher throughput.
Compute capabilities differ sharply in FP16 performance. The Tesla P40 is severely limited, offering only 183.7 GFLOPS of FP16 at a 1:64 ratio, while the WX 9100 delivers 24.58 TFLOPS at a 2:1 ratio. In FP32, they are much closer: the NVIDIA card achieves 11.76 TFLOPS against the AMD card’s 12.29 TFLOPS. The AMD card also carries a higher texture rate at 384.0 GTexel/s versus 367.4 GTexel/s, but the NVIDIA card wins on pixel rate with 147.0 GPixel/s versus 96.00 GPixel/s.
Head-to-Head Benchmarks
The two benchmark tests paint contrasting pictures of these cards’ strengths. In Geekbench OpenCL, the AMD Radeon Pro WX 9100 takes the lead with a score of 66605, beating the Tesla P40’s 62017 by 6.9%. This result aligns with the AMD card’s higher FP32 throughput and significantly greater memory bandwidth. The 2048-bit HBM2 interface gives the WX 9100 a clear advantage in memory-intensive compute workloads, allowing it to sustain higher throughput where data movement is the bottleneck.
The Geekbench Vulkan test reverses the outcome dramatically. Here, the NVIDIA Tesla P40 scores 68172, while the WX 9100 manages only 54711. The 24.6% margin is the largest performance gap between the two cards in any test. This Vulkan advantage suggests that the Pascal architecture’s driver and API implementation are substantially more efficient for this workload, despite the AMD card’s theoretical hardware advantages in shading units and memory bandwidth.
The average scores across all benchmarks show the Tesla P40 at 65095, with the WX 9100 at 64212. The 1.4% average difference places these cards in the same competitive bracket, but the distribution of wins matters more than the aggregate. The NVIDIA card’s nearest rival comparison reinforces its standing: the Tesla P40 is 1.4% ahead of the WX 9100 and 2% ahead of both the NVIDIA CMP 30HX and AMD Radeon RX 9060 XT LP. The WX 9100’s closest rivals are the NVIDIA CMP 30HX at 0.6% behind and the AMD Radeon RX 7600M at 0.7% behind, showing how tightly packed this performance tier is.
Specification Differences
The two cards differ on nearly every major specification. The process node favors NVIDIA at 16 nm versus AMD’s 14 nm, though both are produced by different foundries, TSMC versus GlobalFoundries. The transistor counts are close at 11,800 million for NVIDIA and 12,500 million for AMD, with die sizes of 471 mm² and 495 mm² respectively.
Clock speeds show NVIDIA running higher: the Tesla P40 has a base clock of 1303 MHz and a boost of 1531 MHz, while the WX 9100 runs at 1200 MHz base and 1500 MHz boost. Memory capacity goes to NVIDIA with 24 GB versus 16 GB, but bandwidth goes to AMD with 483.8 GB/s versus 347.1 GB/s. The memory type is fundamentally different: GDDR5 versus HBM2, with bus widths of 384-bit versus 2048-bit.
Compute unit counts vary across the board. NVIDIA has fewer shading units at 3840 versus AMD’s 4096, and fewer TMUs at 240 versus 256, but more ROPs at 96 versus 64. The FP32 output favors AMD slightly at 12.29 TFLOPS versus 11.76 TFLOPS, and the FP16 gap is enormous at 24.58 TFLOPS versus 183.7 GFLOPS. Power consumption favors AMD at 230 W versus 250 W, and the suggested PSU is lower at 550 W versus 600 W. The power connectors differ, with NVIDIA using an 8-pin EPS and AMD using 1x 6-pin plus 1x 8-pin.
Display capabilities are completely different: the Tesla P40 has no outputs, while the WX 9100 offers 6x mini-DisplayPort 1.4a. API support is nearly identical for DirectX at 12 (12_1) and OpenGL at 4.6, but Vulkan support differs slightly at 1.4 for NVIDIA versus 1.3 for AMD. Physical dimensions are identical at 267 mm length and 111 mm height, both dual-slot designs. The release dates differ by about ten months, with the Tesla P40 launching in September 2016 and the WX 9100 in July 2017.
The Verdict
The data points to a clear split based on workload. The NVIDIA Tesla P40 is the superior choice for Vulkan-based applications, where its 24.6% lead over the WX 9100 provides a decisive performance margin. This card also offers 50% more memory capacity at 24 GB versus 16 GB, which matters for large datasets that exceed the AMD card’s capacity. The higher pixel rate of 147.0 GPixel/s versus 96.00 GPixel/s further suggests an advantage for rasterization-heavy tasks.
The AMD Radeon Pro WX 9100 claims victory in OpenCL compute with a 6.9% advantage, backed by higher FP32 throughput at 12.29 TFLOPS and dramatically better FP16 performance at 24.58 TFLOPS. The 483.8 GB/s memory bandwidth is 39% higher than the Tesla P40’s 347.1 GB/s, making the AMD card the better fit for bandwidth-bound compute workloads. Its display outputs also make it the only one of the two that can drive monitors directly.
Neither card emerges as a universal winner. The Tesla P40’s higher average score of 65095 versus 64212 is marginal, and the 1-1 split in benchmark wins confirms that the choice depends entirely on the target application. For users prioritizing Vulkan performance and maximum memory capacity, the Tesla P40 is the data-backed pick. For OpenCL compute, FP16 workloads, and any need for display connectivity, the WX 9100 is the clear choice.
Where Each One Wins
NVIDIA Tesla P40:
- Vulkan workloads, with a 24.6% benchmark advantage (68172 versus 54711)
- Memory capacity, offering 24 GB versus 16 GB for larger datasets
- Pixel fill rate at 147.0 GPixel/s versus 96.00 GPixel/s
- Higher clock speeds at 1531 MHz boost versus 1500 MHz
- DirectX and OpenGL parity is matched, but Vulkan API version is newer at 1.4 versus 1.3
AMD Radeon Pro WX 9100:
- OpenCL compute, with a 6.9% benchmark advantage (66605 versus 62017)
- Memory bandwidth at 483.8 GB/s versus 347.1 GB/s, a 39% margin
- FP32 throughput at 12.29 TFLOPS versus 11.76 TFLOPS
- FP16 performance at 24.58 TFLOPS versus 183.7 GFLOPS, an enormous 134x advantage
- Display connectivity with 6x mini-DisplayPort outputs versus none
- Lower power draw at 230 W versus 250 W, with a 550 W suggested PSU versus 600 W
- Higher shading unit count at 4096 versus 3840 and more TMUs at 256 versus 240