AMD Radeon R9 M295X vs NVIDIA P104-100 Comparison
AMD Radeon R9 M295X
P104-100
PERFORMANCE BENCHMARKS
Analysis: AMD Radeon R9 M295X vs NVIDIA P104-100
Head-to-Head Benchmarks
The database shows two comparable cross-platform tests for these GPUs: Geekbench OpenCL and Geekbench Vulkan. In both, the NVIDIA P104-100 dominates decisively.
In Geekbench OpenCL, the P104-100 scores 52,368 points against the R9 M295X’s 22,858. That is a delta of 129.1%, meaning the NVIDIA part is well over twice as fast in this compute-oriented workload. The gap is enormous and reflects the generational jump between the two architectures.
In Geekbench Vulkan, the P104-100 posts 45,165 versus 29,091 for the AMD part, a 55.3% advantage. While narrower than the OpenCL gap, it is still a commanding lead. The P104-100 wins both head-to-head tests, giving it a 2-0 record in the database.
Looking at average benchmark scores across all recorded tests, the P104-100 averages 32,982 points, placing it in the 77th percentile of all GPUs. The R9 M295X averages 28,580 points, sitting in the 74th percentile. The absolute difference is 4,402 points, roughly a 15.4% gap in the aggregate.
The P104-100’s nearest rivals are the NVIDIA T600 Mobile (32,849, delta 0.4%), the NVIDIA T550 Mobile (33,161, delta -0.5%), the NVIDIA GeForce RTX 3050 Mobile (33,170, delta -0.6%), and the AMD Radeon Pro 570 (33,207, delta -0.7%). This means the P104-100 sits in a tight cluster where performance differences among these cards are within one percent. It is effectively a peer of those mobile parts.
The R9 M295X’s nearest rivals are the NVIDIA Quadro RTX 8000 (28,421, delta 0.6%), the AMD Radeon RX 570 (28,766, delta -0.6%), the AMD Radeon RX 6800M (28,874, delta -1%), and the AMD Radeon RX 470 (28,996, delta -1.4%). Here too, the spread is narrow, but the R9 M295X is slightly below these cards, with the Quadro RTX 8000 being the only one it edges out.
The key takeaway: in any direct comparison, the P104-100 is the faster GPU by a wide margin. The OpenCL result is particularly striking, and the Vulkan result reinforces the same conclusion. The R9 M295X is not competitive in these tests.
FAQ
Q: Which GPU has the higher average benchmark score?
A: The NVIDIA P104-100 averages 32,982 points, while the AMD Radeon R9 M295X averages 28,580 points. The P104-100 also ranks higher in the overall percentile, at 77 versus 74.
Q: How large is the performance gap in the shared tests?
A: In Geekbench OpenCL, the P104-100 leads by 129.1%. In Geekbench Vulkan, it leads by 55.3%. Both are wins for the NVIDIA part.
Q: Does the R9 M295X win any head-to-head benchmark?
A: No. The database records two head-to-head tests, and the R9 M295X loses both. It has zero wins against the P104-100.
Q: What is the memory bandwidth difference?
A: The P104-100 has 320.3 GB/s of bandwidth using GDDR5X memory, while the R9 M295X has 160.0 GB/s using GDDR5. Both have 4 GB of memory on a 256-bit bus.
Q: How do the shading unit counts compare?
A: The R9 M295X has 2,048 shading units, slightly more than the P104-100’s 1,920. However, the P104-100 has higher clock speeds and a newer architecture, which explains its performance lead.
Q: What is the transistor density difference?
A: The P104-100 packs 22.9 million transistors per square millimeter, while the R9 M295X has 13.7 million per square millimeter. The P104-100 achieves this on a 16 nm process versus the R9 M295X’s 28 nm process.
The Verdict
The data is unambiguous: the NVIDIA P104-100 is the superior GPU in every recorded benchmark. It wins both head-to-head tests with deltas of 129.1% and 55.3%, and its average score is 15.4% higher. For anyone comparing these two parts for compute or graphics workloads, the P104-100 is the clear choice.
The R9 M295X, while a capable card in its own right, sits in a lower performance tier. Its average score of 28,580 places it near the NVIDIA Quadro RTX 8000 and AMD Radeon RX 570, but it trails those slightly. The P104-100, by contrast, is competitive with the NVIDIA T600 Mobile, T550 Mobile, and RTX 3050 Mobile, all within one percent of its average.
There is no scenario in the data where the R9 M295X comes out ahead. The P104-100 wins on raw compute, on graphics API performance, and on the aggregate benchmark average. The only area where the R9 M295X shows a different profile is in its feature set and power characteristics, which the next sections cover.
Specification Differences
The two cards differ in several key specifications. The NVIDIA P104-100 uses a 16 nm process from TSMC, while the R9 M295X uses a 28 nm process from the same foundry. Transistor counts are 7,200 million for the P104-100 versus 5,000 million for the R9 M295X. Die size is 314 mm² for the NVIDIA part and 366 mm² for the AMD part.
Clock speeds: the P104-100 has a base clock of 1607 MHz and a boost clock of 1733 MHz. The R9 M295X has no base or boost clock listed in the database. Memory clocks are similar in raw frequency (1251 MHz versus 1250 MHz), but the effective data rate differs: 10 Gbps for the P104-100 versus 5 Gbps for the R9 M295X.
Memory type is GDDR5X for the NVIDIA part and GDDR5 for the AMD part. Bandwidth is 320.3 GB/s versus 160.0 GB/s. Both have 4 GB capacity and 256-bit buses.
Texture units: 120 for the P104-100, 128 for the R9 M295X. ROPs: 64 versus 32. Pixel rate is 110.9 GPixel/s versus 23.14 GPixel/s. Texture rate is 208.0 GTexel/s versus 92.54 GTexel/s.
FP32 performance is 6.655 TFLOPS for the P104-100 and 2.961 TFLOPS for the R9 M295X. FP16 performance is starkly different: 104.0 GFLOPS (at a 1:64 ratio) for the NVIDIA part, but 2.961 TFLOPS (at a 1:1 ratio) for the AMD part.
Power and physical specs also differ. The P104-100 has no TDP listed, but a suggested PSU of 200 W and a single 8-pin power connector. The R9 M295X has a TDP of 250 W and no power connectors, as it is an MXM module. The P104-100 is dual-slot and 267 mm long, while the R9 M295X is an MXM module with no listed dimensions. Bus interfaces are PCIe 1.0 x4 for the NVIDIA part and MXM-B (3.0) for the AMD part.
Display outputs: the P104-100 has none, while the R9 M295X is marked as portable device dependent. API support: the P104-100 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The R9 M295X supports DirectX 12 (12_0), OpenGL 4.6, and Vulkan 1.2.170.
Architecture Differences
The two GPUs come from different architectural generations. The NVIDIA P104-100 is built on the Pascal architecture, specifically using the GP104 chip. It belongs to the Mining GPUs generation. The AMD Radeon R9 M295X uses the GCN 3.0 architecture with the Amethyst chip, part of the Gem System (R9 M200) generation.
Process node is a major differentiator: 16 nm for Pascal versus 28 nm for GCN 3.0. This directly affects transistor density, with the P104-100 achieving 22.9 million transistors per mm² versus 13.7 million for the R9 M295X. The newer process allows the P104-100 to pack more transistors into a smaller die.
The P104-100 has a 1:64 ratio for FP16 to FP32, meaning its FP16 throughput is minimal compared to FP32. The R9 M295X has a 1:1 ratio, giving it equal FP16 and FP32 performance. This makes the AMD card better suited for workloads that use FP16 precision, though the P104-100’s raw FP32 power is more than double.
The R9 M295X has more shading units (2,048 versus 1,920) and more texture units (128 versus 120), but fewer ROPs (32 versus 64). Despite the higher unit counts in some areas, the P104-100’s higher clocks and newer architecture give it superior pixel and texture rates.
Release dates differ by over three years: the P104-100 launched in December 2017, while the R9 M295X launched in November 2014. Both are end-of-life products. The R9 M295X has a predecessor listed as Solar System and a successor as Polaris Mobile; the P104-100 has no predecessor or successor in the database.
Where Each One Wins
The NVIDIA P104-100 wins in every compute workload recorded. Its Geekbench OpenCL score of 52,368 is more than double the R9 M295X’s 22,858. Its Geekbench Vulkan score of 45,165 exceeds the AMD part’s 29,091 by a healthy margin. The average benchmark score of 32,982 versus 28,580 confirms the overall trend.
Use cases for the P104-100: general GPU compute, OpenCL workloads, Vulkan rendering, and any task that benefits from high FP32 throughput. Its 6.655 TFLOPS of FP32 performance and 320.3 GB/s of memory bandwidth make it a strong candidate for compute-heavy applications. The lack of display outputs means it is meant for mining or compute-only roles, not as a primary display adapter.
Use cases for the R9 M295X: FP16 workloads, where its 1:1 FP16 to FP32 ratio gives it a relative advantage. The 2.961 TFLOPS of FP16 performance matches its FP32 figure, which could benefit certain machine learning or scientific applications that use reduced precision. It also has display outputs, making it suitable for portable devices where a display connection is needed.
However, the data does not show the R9 M295X winning any benchmark. The only areas where it has a theoretical edge are in FP16 throughput and the ability to output video. For raw performance, the P104-100 is the winner in every measured category.
In practical terms: if the workload is pure compute and does not require FP16, the P104-100 is the obvious pick. If the workload demands FP16 precision or requires display output from the GPU itself, the R9 M295X has those specific capabilities, but it will be slower in everything else. The benchmark database records no scenario where the R9 M295X outperforms the P104-100.