NVIDIA RTX A6000 vs NVIDIA Tesla P40 Comparison
NVIDIA RTX A6000
Tesla P40
PERFORMANCE BENCHMARKS
Analysis: NVIDIA RTX A6000 vs NVIDIA Tesla P40
Head-to-Head Benchmarks
The recorded head-to-head data contains two benchmark comparisons, and the NVIDIA RTX A6000 dominates both. In Geekbench OpenCL, the RTX A6000 scores 193,937 against the Tesla P40’s 62,017, a delta of -68% from the perspective of the P40. That is a decisive margin, placing the A6000 more than three times ahead of its predecessor in raw compute throughput as measured by this workload. The Geekbench Vulkan result tells a similar story: the RTX A6000 posts 164,462, while the Tesla P40 manages 68,172, a -58.5% delta. While Vulkan is not the primary API for either workstation card, the gap remains substantial.
These two tests are the only direct comparisons in the database for this pairing, and the RTX A6000 wins both, giving it a 2-0 record in head-to-head wins. The Tesla P40 has zero wins in the recorded comparisons. Looking at the broader benchmark averages, the RTX A6000’s average score is 44,075, while the Tesla P40’s average is 65,095. This is an unusual inversion: the P40’s average is higher than the A6000’s even though the A6000 wins the two common tests. The explanation lies in the benchmark sets available for each card. The P40 has only two recorded scores, both in Geekbench (OpenCL and Vulkan), and both are relatively high for that card. The A6000 has nine recorded scores, including Passmark DirectX 9, 10, 11, 12, G2D, G3D, and GPU compute, several of which are lower in absolute terms, dragging down its average.
The percentile data adds context. The Tesla P40 sits at the 89th percentile among all GPUs in the database, while the RTX A6000 sits at the 84th percentile. Despite the A6000’s massive wins in Geekbench, its percentile rank is lower because the full benchmark suite includes older DirectX tests where workstation cards do not excel. The P40, with its limited test set, benefits from having only its strongest scores counted. This demonstrates why single-metric comparisons can mislead: the A6000 is clearly superior in the tests where both cards appear, but the database’s aggregate metrics reflect different test coverage.
For the nearest rivals, the P40’s closest competitor is the AMD Radeon Pro WX 9100, which averages 64,212, a 1.4% delta from the P40’s 65,095. The AMD Radeon VII is 1.4% ahead at 66,004, and the NVIDIA CMP 30HX and AMD Radeon RX 9060 XT LP are both 2% behind at 63,842 and 63,830 respectively. The RTX A6000’s nearest rival is the NVIDIA GeForce RTX 4090 Mobile at 43,667, a 0.9% delta, followed by the NVIDIA GeForce RTX 4070 Ti at 44,795 (1.6% ahead), the NVIDIA Quadro M6000 at 43,301 (1.8% behind), and the NVIDIA GeForce RTX 5050 Mobile at 43,268 (1.9% behind). The A6000’s average score places it in a tight cluster with mobile RTX parts and the older Quadro M6000, which is notable given the A6000’s much higher peak performance in Geekbench.
The Verdict
The data is unambiguous for compute-heavy workloads: the RTX A6000 is the stronger card. In Geekbench OpenCL, it outperforms the Tesla P40 by 68%, and in Geekbench Vulkan by 58.5%. The A6000 also carries architectural advantages that align with these results: it has 10,752 shading units versus 3,840, 336 TMUs versus 240, and 112 ROPs versus 96. Its FP32 throughput is 38.71 TFLOPS versus 11.76 TFLOPS, a 3.3x difference. Memory bandwidth doubles from 347.1 GB/s to 768.0 GB/s, and capacity doubles from 24 GB to 48 GB. The A6000 also adds 84 RT cores and 336 tensor cores, features entirely absent from the Pascal-based P40.
The Tesla P40 does have one statistical advantage: its 89th percentile rank versus the A6000’s 84th, and a higher average benchmark score (65,095 versus 44,075). But this is an artifact of test coverage, not a reflection of real performance. The P40’s only recorded tests are Geekbench OpenCL and Vulkan, while the A6000’s suite includes legacy DirectX tests where its scores are modest (Passmark DirectX 9: 245, DirectX 10: 155, DirectX 11: 191, DirectX 12: 87). These low scores pull the A6000’s average down. A buyer relying solely on average scores would incorrectly conclude the P40 is faster. The head-to-head results, which use identical workloads, are the more reliable comparison.
For the user who needs a server accelerator for compute, the RTX A6000 is the appropriate choice. It wins every shared benchmark by a wide margin, offers twice the memory and bandwidth, and includes modern features like ray tracing and tensor cores. The Tesla P40, despite its higher percentile, is a Pascal-era card from 2016, end-of-life, with no display outputs and a memory architecture limited to GDDR5. Its only conceivable advantage is lower power draw at 250 W versus 300 W, but that is a modest difference relative to the performance gap. The verdict from the data: the RTX A6000 is the superior accelerator, and the P40 should only be considered if the workload is specifically tuned for Pascal-era hardware or if the lower TDP is an absolute constraint.
Where Each One Wins
The RTX A6000 wins in every scenario where both cards are measured. In Geekbench OpenCL, its score of 193,937 is more than triple the P40’s 62,017. In Geekbench Vulkan, the A6000’s 164,462 more than doubles the P40’s 68,172. These are compute-oriented workloads that stress raw FP32 throughput, memory bandwidth, and driver efficiency, all areas where the Ampere architecture holds a clear lead. The A6000 also wins on memory capacity (48 GB versus 24 GB), memory type (GDDR6 versus GDDR5), bandwidth (768.0 GB/s versus 347.1 GB/s), and feature set (RT cores, tensor cores, DirectX 12 Ultimate support). It has display outputs (4x DisplayPort 1.4a), making it usable for visualization tasks, while the P40 has no outputs at all.
The Tesla P40’s wins are narrower and statistical rather than performance-based. Its 89th percentile versus the A6000’s 84th is one such win, driven by the P40’s limited benchmark coverage. Its average score of 65,095 is higher than the A6000’s 44,075, but this is misleading as noted. The P40 also has a lower TDP at 250 W versus 300 W, and a lower suggested PSU of 600 W versus 700 W, making it easier to integrate into power-constrained systems. It uses PCIe 3.0 x16 versus the A6000’s PCIe 4.0 x16, which is a drawback for data transfer but not a performance win. In terms of raw compute, the P40 does not win any recorded benchmark against the A6000. Its only strengths are legacy compatibility (Pascal architecture, 2016 release) and lower power consumption.
For use-case splitting, the A6000 is the choice for machine learning inference and training (due to tensor cores), ray tracing workloads (due to RT cores), high-resolution rendering (due to 48 GB memory), and any task requiring display output. The P40 is only suitable for compute tasks that predate Ampere optimizations, or for systems where the 250 W TDP and 600 W PSU requirement are hard limits. Even then, the P40’s 11.76 TFLOPS FP32 is less than a third of the A6000’s 38.71 TFLOPS, so the performance penalty is severe. The data does not support recommending the P40 for any compute-heavy role where the A6000 could be installed.
FAQ
Q: Which card has a higher average benchmark score?
A: The Tesla P40 has a higher average score of 65,095 compared to the RTX A6000’s 44,075. However, this is due to different benchmark coverage: the P40 has only two Geekbench scores, while the A6000 has nine scores including legacy DirectX tests, which lower its average.
Q: How much faster is the RTX A6000 in Geekbench OpenCL?
A: The RTX A6000 scores 193,937 versus the Tesla P40’s 62,017, a 68% advantage from the P40’s perspective. In raw terms, the A6000 is roughly 3.1 times faster in this test.
Q: Does the Tesla P40 have any ray tracing or tensor cores?
A: No. The P40’s specifications list no RT cores and no tensor cores. The RTX A6000 includes 84 RT cores and 336 tensor cores.
Q: What is the memory capacity difference?
A: The Tesla P40 has 24 GB of GDDR5 memory with a 384-bit bus and 347.1 GB/s bandwidth. The RTX A6000 has 48 GB of GDDR6 memory with a 384-bit bus and 768.0 GB/s bandwidth.
Q: Which card has a higher FP32 throughput?
A: The RTX A6000 achieves 38.71 TFLOPS FP32, while the Tesla P40 achieves 11.76 TFLOPS. The A6000 is 3.3 times faster in this metric.
Q: Are both cards end-of-life?
A: Yes. The Tesla P40 was released in 2016 and is end-of-life, as is the RTX A6000, released in 2020. Both have successors in the database: Tesla Volta for the P40 and Workstation Ada for the A6000.
Q: What is the transistor density of each chip?
A: The Tesla P40’s GP102 chip has 11,800 million transistors on a 471 mm² die, yielding 25.1M transistors per mm². The RTX A6000’s GA102 chip has 28,300 million transistors on a 628 mm² die, yielding 45.1M per mm².
Architecture Differences
The Tesla P40 is built on the Pascal architecture, fabricated by TSMC on a 16 nm process. Its chip, GP102, contains 11,800 million transistors on a 471 mm² die, giving a transistor density of 25.1 million per mm². The RTX A6000 uses the Ampere architecture, fabricated by Samsung on an 8 nm process. Its chip, GA102, contains 28,300 million transistors on a 628 mm² die, for a density of 45.1 million per mm². The node shrink and denser layout allow Ampere to pack more than twice the transistors into a die that is only 33% larger by area.
The compute configuration differs dramatically. The P40 has 3,840 shading units, 240 TMUs, and 96 ROPs. The A6000 has 10,752 shading units (2.8x), 336 TMUs (1.4x), and 112 ROPs (1.17x). The A6000 also adds 84 RT cores for ray tracing and 336 tensor cores for AI workloads; the P40 has neither. Pixel rate improves from 147.0 GPixel/s to 201.6 GPixel/s, and texture rate from 367.4 GTexel/s to 604.8 GTexel/s. FP32 throughput jumps from 11.76 TFLOPS to 38.71 TFLOPS. FP16 performance is even more skewed: the P40 manages only 183.7 GFLOPS (1:64 ratio), while the A6000 achieves 38.71 TFLOPS (1:1 ratio), a 210x difference.
Memory architecture also diverges. The P40 uses 24 GB of GDDR5 on a 384-bit bus, with 347.1 GB/s bandwidth and a memory clock of 1808 MHz (7.2 Gbps effective). The A6000 uses 48 GB of GDDR6 on the same 384-bit bus, but with 768.0 GB/s bandwidth and a 2000 MHz memory clock (16 Gbps effective). The A6000 thus doubles both capacity and bandwidth. The bus interface improves from PCIe 3.0 x16 to PCIe 4.0 x16, doubling host transfer bandwidth. Power requirements rise from 250 W TDP with a 600 W suggested PSU to 300 W TDP with a 700 W suggested PSU, both using a single 8-pin EPS connector and dual-slot cooling.
API support reflects the generational gap. The P40 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The A6000 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The A6000 also provides 4x DisplayPort 1.4a outputs, while the P40 has no display outputs, making it a headless compute accelerator only. Physical dimensions are nearly identical: both are 267 mm long, with the P40 at 111 mm height and the A6000 at 112 mm. The release dates are four years apart, September 2016 for the P40 and October 2020 for the A6000, and the launch MSRP for the P40 was 5,699 USD, while the A6000 launched at 4,649 USD. The A6000 is both newer and cheaper at launch, while offering substantially higher performance in every recorded metric.