NVIDIA Quadro RTX 6000 vs NVIDIA Tesla P40 Comparison

NVIDIA
GEFORCE

NVIDIA Quadro RTX 6000

CORE STATE TU102
VRAM 24 GB
CLOCK SPEED 1770 MHz
TDP 260 W
BUS WIDTH 384 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2018
VS
NVIDIA
GEFORCE

Tesla P40

CORE STATE GP102
VRAM 24 GB
CLOCK SPEED 1531 MHz
TDP 250 W
BUS WIDTH 384 bit
ARCHITECTURE Pascal
nm
PROCESS 16 nm
LAUNCH DATE 2016

PERFORMANCE BENCHMARKS

geekbench_opencl
74,179
62,017
geekbench_vulkan
129,564
68,172

Analysis: NVIDIA Quadro RTX 6000 vs NVIDIA Tesla P40

Where Each One Wins

The benchmark data splits cleanly along workload type. The NVIDIA Quadro RTX 6000 wins both recorded tests, but the margin tells the real story. In Geekbench OpenCL, the RTX 6000 scores 74,179 against the Tesla P40's 62,017, a 19.6% advantage. That is a solid lead, but not overwhelming. In Geekbench Vulkan, the gap becomes enormous: 129,564 versus 68,172, a 90.1% difference. The RTX 6000 more than doubles the Tesla P40 in that test.

The Tesla P40's strongest position is in compute workloads that resemble OpenCL. Its 11.76 TFLOPS of FP32 throughput and 367.4 GTexel/s texture rate are respectable for a 2016-era accelerator. The data suggests the P40 remains viable for raw number crunching where API overhead is minimal. However, the RTX 6000 counters with 16.31 TFLOPS FP32 and 509.8 GTexel/s, so even in the P40's home turf, the newer card leads by roughly 38% in FP32 and 39% in texture rate.

The Vulkan result is where the architecture gap becomes decisive. The RTX 6000's Turing design includes 72 RT cores and 576 tensor cores, which the Pascal-based P40 lacks entirely. Vulkan 1.4 support exists on both cards, but the RTX 6000's feature set allows it to exploit modern graphics APIs far more effectively. The 90.1% delta in Vulkan suggests that any workload leveraging newer API features will overwhelmingly favor the RTX 6000.

For users prioritizing legacy compute or simple FP32 throughput, the P40 still holds ground. Its 183.7 GFLOPS FP16 rate (1:64 ratio) is minuscule, so half-precision work is essentially unsupported. The RTX 6000 delivers 32.62 TFLOPS FP16 (2:1 ratio), making it 177 times faster in that specific metric. Any AI inference or training task using FP16 will only consider the RTX 6000.

Head-to-Head Benchmarks

The Geekbench OpenCL test shows a 19.6% victory for the RTX 6000: 74,179 versus 62,017. This margin aligns with the raw compute gap. The RTX 6000 has 4,608 shading units versus 3,840 on the P40, and its 1,440 MHz base clock runs above the P40's 1,303 MHz. The boost clocks differ similarly: 1,770 MHz versus 1,531 MHz. These clock and core count advantages compound to produce the observed OpenCL lead.

The Vulkan test is the standout. The RTX 6000 scores 129,564, which is 90.1% higher than the P40's 68,172. This is not a modest generational improvement; it is a fundamental capability gap. Vulkan's explicit multi-threading and draw-call handling benefit from the RTX 6000's Turing architecture. The presence of RT cores and tensor cores means the RTX 6000 can offload certain operations that the P40 must handle through general-purpose shaders. The P40's 96 ROPs match the RTX 6000's 96 ROPs, so rasterization output is equal, but that is where the parity ends.

Looking at the nearest rivals for context: the RTX 6000's average benchmark score is 101,872, placing it 4.5% above the AMD Radeon RX 7900M and 4.9% above the AMD Radeon Pro VII. It sits 4.6% below the Radeon Pro Vega II Duo and 5.1% below the Radeon Pro W6600X. The P40's average is 65,095, nearly identical to the AMD Radeon VII at 66,004 (P40 trails by 1.4%) and slightly ahead of the Radeon Pro WX 9100 by 1.4%. The P40 also edges out the NVIDIA CMP 30HX and AMD Radeon RX 9060 XT LP by 2% each.

These rival comparisons reveal that the RTX 6000 competes against modern workstation flagships, while the P40 sits among older or lower-tier cards. The percentile rankings confirm this: the RTX 6000 is in the 94th percentile of all GPUs, the P40 in the 89th.

Architecture Differences

The RTX 6000 uses the TU102 chip built on Turing architecture at TSMC's 12 nm process. It packs 18,600 million transistors into a 754 mm² die, yielding a density of 24.7 million transistors per mm². The P40 uses the GP102 chip on Pascal architecture, also from TSMC but at 16 nm. It contains 11,800 million transistors on a 471 mm² die, with a slightly higher density of 25.1 million per mm². The process node difference explains the transistor count disparity: 12 nm allowed Turing to pack 57% more transistors despite being only a slightly newer node.

Turing introduces dedicated ray tracing hardware in the form of 72 RT cores. The P40 has none. Turing also includes 576 tensor cores for AI workloads; the P40 has none. These are not incremental additions but entirely new processing units that change what the GPU can do. The RTX 6000's 4,608 shading units execute at 16.31 TFLOPS FP32, while the P40's 3,840 shading units manage 11.76 TFLOPS. The RTX 6000 has 288 TMUs versus 240, and both have 96 ROPs.

Memory architecture differs substantially. The RTX 6000 uses 24 GB of GDDR6 on a 384-bit bus, delivering 672.0 GB/s bandwidth. The P40 also has 24 GB but uses GDDR5 on the same 384-bit bus, achieving only 347.1 GB/s. That is nearly half the bandwidth, and it directly impacts compute-heavy workloads that stream data. The memory clock tells the story: the RTX 6000 runs at 1,750 MHz (14 Gbps effective) versus 1,808 MHz (7.2 Gbps effective) on the P40. The GDDR6 standard doubles the per-pin data rate.

FP16 performance is an order-of-magnitude difference. The RTX 6000 delivers 32.62 TFLOPS at a 2:1 ratio, meaning two FP16 operations per FP32 operation. The P40 manages only 183.7 GFLOPS at a 1:64 ratio, meaning FP16 is heavily de-prioritized. Any modern AI framework that defaults to FP16 will run 177 times faster on the RTX 6000.

Specification Differences

The two cards share 24 GB memory capacity, a 384-bit bus width, 96 ROPs, dual-slot width, 267 mm length, 111 mm height, PCIe 3.0 x16 interface, and a 600 W suggested PSU. Beyond those, nearly everything differs.

The RTX 6000 uses TU102 on 12 nm with 18,600 million transistors and a 754 mm² die. The P40 uses GP102 on 16 nm with 11,800 million transistors and a 471 mm² die. Clock speeds: the RTX 6000 runs 1,440 MHz base and 1,770 MHz boost; the P40 runs 1,303 MHz base and 1,531 MHz boost. Memory type differs: GDDR6 at 1,750 MHz (14 Gbps effective) versus GDDR5 at 1,808 MHz (7.2 Gbps effective). Bandwidth is 672.0 GB/s versus 347.1 GB/s.

Shading units: 4,608 versus 3,840. TMUs: 288 versus 240. The RTX 6000 has 72 RT cores and 576 tensor cores; the P40 has neither. Pixel rate is 169.9 GPixel/s versus 147.0 GPixel/s. Texture rate is 509.8 GTexel/s versus 367.4 GTexel/s. FP32 throughput is 16.31 TFLOPS versus 11.76 TFLOPS. FP16 is 32.62 TFLOPS versus 183.7 GFLOPS.

TDP is 260 W versus 250 W, a negligible difference. Power connectors: the RTX 6000 uses 1x 6-pin plus 1x 8-pin; the P40 uses a single 8-pin EPS. Display outputs: the RTX 6000 offers 4x DisplayPort 1.4a and 1x USB Type-C; the P40 has no outputs. DirectX support: the RTX 6000 supports 12 Ultimate (12_2); the P40 only 12 (12_1). OpenGL is 4.6 on both, Vulkan 1.4 on both.

Release dates are two years apart: the RTX 6000 launched August 2018, the P40 in September 2016. Both are end-of-life. The RTX 6000 succeeded Quadro Volta and was succeeded by Workstation Ampere. The P40 succeeded Tesla Maxwell and was succeeded by Tesla Volta.

FAQ

Q: Which card has higher memory bandwidth?

A: The RTX 6000 offers 672.0 GB/s from GDDR6, nearly double the P40's 347.1 GB/s from GDDR5.

Q: Can the Tesla P40 handle ray tracing or AI workloads?

A: No. The P40 has no RT cores and no tensor cores. The RTX 6000 includes 72 RT cores and 576 tensor cores.

Q: How much faster is the RTX 6000 in Vulkan?

A: The RTX 6000 scores 129,564 versus 68,172, a 90.1% advantage in Geekbench Vulkan.

Q: Do both cards have the same amount of memory?

A: Yes, both have 24 GB, but the RTX 6000 uses faster GDDR6 while the P40 uses GDDR5.

Q: Which card draws more power?

A: The RTX 6000 has a 260 W TDP versus 250 W for the P40, a 10 W difference.

Q: What is the FP16 performance gap?

A: The RTX 6000 delivers 32.62 TFLOPS FP16, while the P40 manages 183.7 GFLOPS, a 177-fold difference.

The Verdict

The data directs different users to different cards. Workloads built around Vulkan or modern graphics APIs should exclusively use the RTX 6000. The 90.1% Vulkan lead is decisive, and the presence of RT and tensor cores means future-proofing for ray-traced or AI-accelerated applications. The RTX 6000 also provides superior memory bandwidth at 672.0 GB/s, which benefits any data-intensive task. Its 94th percentile ranking among all GPUs confirms it sits near the top of the performance hierarchy.

The Tesla P40 remains relevant only for legacy compute tasks where FP32 throughput is the sole concern and API overhead is minimal. Its 11.76 TFLOPS FP32 is still serviceable, and its 89th percentile shows it outperforms many current cards. However, the P40 lacks display outputs, making it unsuitable for any workstation role requiring visual output. It also lacks FP16 capability, closing the door on modern AI workflows.

For users with existing Pascal-era infrastructure, the P40's 24 GB capacity and 250 W TDP make it a drop-in replacement with similar power draw. But the RTX 6000, despite its 260 W TDP, offers roughly 38% more FP32 throughput and over double the memory bandwidth. The average benchmark scores tell the final story: 101,872 for the RTX 6000 versus 65,095 for the P40, a 56.5% overall advantage. The RTX 6000 is the clear choice for any new deployment. The P40 is a legacy option, not a competitive one.

DETAILED SPECIFICATIONS

SPECIFICATION
Quadro RTX 6000
Tesla P40
Core Specs
Shading Units
4,608
3,840 -16.7%
Shaders
4,608
3,840 -16.7%
TMUs
288
240 -16.7%
ROPs
96
96 0.0%
SM Count
72
30 -58.3%
Clocks
Base Clock
1440 MHz
1303 MHz
Boost Clock
1770 MHz
1531 MHz
Memory Clock
1750 MHz 14 Gbps effective
1808 MHz 7.2 Gbps effective
Memory
Memory Size
24 GB
24 GB
VRAM (MB)
24,576
24,576 0.0%
Memory Type
GDDR6
GDDR5
Memory Bus
384 bit
384 bit
Bandwidth
672.0 GB/s
347.1 GB/s
Cache
L1 Cache
64 KB (per SM)
48 KB (per SM)
L2 Cache
6 MB
3 MB
Performance
Pixel Rate
169.9 GPixel/s
147.0 GPixel/s
Texture Rate
509.8 GTexel/s
367.4 GTexel/s
FP32 (TFLOPS)
16.31 TFLOPS
11.76 TFLOPS
FP64 (TFLOPS)
509.8 GFLOPS (1:32)
367.4 GFLOPS (1:32)
FP16 (TFLOPS)
32.62 TFLOPS (2:1)
183.7 GFLOPS (1:64)
AI/RT
RT Cores
72
Tensor Cores
576
Power
TDP
260 W
250 W
TDP (W)
260
250 -3.8%
Suggested PSU
600 W
600 W
Power Connectors
1x 6-pin + 1x 8-pin
8-pin EPS
Architecture
Architecture
Turing
Pascal
GPU Name
TU102
GP102
Generation
Quadro Turing (Tx000)
Tesla Pascal (Pxx)
Process Size
12 nm
16 nm
Transistors
18,600 million
11,800 million
Die Size
754 mm²
471 mm²
Foundry
TSMC
TSMC
Density
24.7M / mm²
25.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
7.5
6.1
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Dual-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
111 mm 4.4 inches
111 mm 4.4 inches
Outputs
4x DisplayPort 1.4a1x USB Type-C
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 3.0 x16
Other
Launch Price
6,299 USD
5,699 USD
Production
End-of-life
End-of-life
Predecessor
Quadro Volta
Tesla Maxwell
Successor
Workstation Ampere
Tesla Volta
View Quadro RTX 6000 Details View Tesla P40 Details