GEFORCE

NVIDIA Tesla P40

NVIDIA graphics card specifications and benchmark scores

24 GB
VRAM
1531
MHz Boost
250W
TDP
384
Bus Width

At a Glance

NVIDIA
VRAM 24 GB
Boost Clock 1,531 MHz
Shaders 3,840
Bus Width 384-bit
TDP 250W
Memory Type GDDR5
Architecture Pascal
nm
Process 16 nm
Released Sep 2016

NVIDIA Tesla P40 Specifications

GPU Core

Shader units and compute resources

The NVIDIA Tesla P40 GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
3,840
Shaders
3,840
TMUs
240
ROPs
96
SM Count
30

Tesla P40 Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the Tesla P40's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Tesla P40 by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
1303 MHz
Base Clock
1,303 MHz
Boost Clock
1531 MHz
Boost Clock
1,531 MHz
Memory Clock
1808 MHz 7.2 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's Tesla P40 Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Tesla P40's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
24 GB
VRAM
24,576 MB
Memory Type
GDDR5
VRAM Type
GDDR5
Memory Bus
384 bit
Bus Width
384-bit
Bandwidth
347.1 GB/s

Tesla P40 by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the Tesla P40, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
48 KB (per SM)
L2 Cache
3 MB

Tesla P40 Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA Tesla P40 against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
11.76 TFLOPS
FP64 (Double)
367.4 GFLOPS (1:32)
FP16 (Half)
183.7 GFLOPS (1:64)
Pixel Rate
147.0 GPixel/s
Texture Rate
367.4 GTexel/s

Pascal Architecture & Process

Manufacturing and design details

The NVIDIA Tesla P40 is built on NVIDIA's Pascal architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Tesla P40 will perform in GPU benchmarks compared to previous generations.

Architecture
Pascal
GPU Name
GP102
Process Node
16 nm
Foundry
TSMC
Transistors
11,800 million
Die Size
471 mm²
Density
25.1M / mm²

Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA Tesla P40 determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Tesla P40 to maintain boost clocks without throttling.

TDP
250 W
TDP
250W
Power Connectors
8-pin EPS
Suggested PSU
600 W

Tesla P40 by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA Tesla P40 are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
Dual-slot
Length
267 mm 10.5 inches
Height
111 mm 4.4 inches
Bus Interface
PCIe 3.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA Tesla P40. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 (12_1)
DirectX
12 (12_1)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.4
Vulkan
1.4
OpenCL
3.0
CUDA
6.1
Shader Model
6.8

Tesla P40 Product Information

Release and pricing details

The NVIDIA Tesla P40 is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Tesla P40 by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Sep 2016
Launch Price
5,699 USD
Production
End-of-life
Predecessor
Tesla Maxwell
Successor
Tesla Volta

About NVIDIA Tesla P40

NVIDIA’s Tesla P40 is a Pascal-era compute accelerator built for data center workloads, not desktop gaming. With a 24 GB memory pool and a 384-bit bus, it was a high-capacity inference and rendering card for its time. Benchmark results place it just shy of much newer hardware, though the architecture’s age shows in specific workloads. The launch MSRP is 5,699 USD.

Benchmark Performance

The Tesla P40 achieves an average benchmark score of 66,127 across available tests, which places it in the 91st percentile of all GPUs tracked. That is a strong showing for a card released in 2016, and it remains competitive against several newer products in synthetic workloads. In Geekbench OpenCL, the card scores 62,017 points, while its Vulkan score reaches 70,237 points, indicating that the compute-oriented design scales well with modern API overhead.

The delta to its nearest rival is razor-thin. The GeForce RTX 4090 averages 66,473 points, which is only 0.5% ahead of the Tesla P40. That margin is negligible in real-world terms, and it reflects the fact that the P40’s raw shader throughput and memory bandwidth still hold up in compute-heavy benchmarks. The Tesla T4, a newer and more power-efficient data center card, scores 66,733 points, just 0.9% ahead. The AMD Radeon Pro Vega 56 sits 1.4% higher at 67,097 points, and the Quadro P6000—a direct workstation sibling with similar Pascal DNA—leads by 1.8% with 67,320 points.

These figures suggest the Tesla P40 is not obsolete in raw compute terms. However, the story changes when you consider that the RTX 4090 achieves its score with far higher clock speeds and newer architecture, while the P40 relies on sheer shader count and memory capacity. The gap between the P40 and the RTX 4090 is effectively a tie in these synthetic tests, but that does not translate to equal performance in modern games or AI workloads that leverage tensor cores. The data indicates the P40 punches above its age in generic compute, but it lacks the specialized hardware that pushes newer cards ahead in targeted tasks.

Memory Subsystem

The Tesla P40 carries 24 GB of GDDR5 memory on a 384-bit bus, producing 347.1 GB/s of bandwidth. Memory speed is rated at 1808 MHz, which translates to 7.2 Gbps effective. That capacity is the card’s defining feature—24 GB was enormous in 2016 and remains useful for large datasets, high-resolution textures, and multi-model inference workloads that exceed the VRAM of most consumer cards.

For high-resolution rendering, the 347.1 GB/s bandwidth is adequate but not exceptional by modern standards. Compare that to the RTX 4090, which uses faster GDDR6X on a wider 384-bit interface; the P40’s bandwidth is roughly half of what newer flagships offer. In practice, this means the P40 can hold massive scenes in memory, but moving data across the bus will be slower than on contemporary cards. The 24 GB pool also allows for 4K and 8K texture sets without swapping to system memory, which is a significant advantage over 8 GB or 12 GB cards. However, the GDDR5 type and effective speed cap the card’s ability to feed its 3840 shading units at peak load. Benchmark results show the card keeping pace with newer rivals in aggregate, but memory-bound tasks will reveal the bandwidth deficit.

Ray Tracing and Feature Set

The Tesla P40 has no dedicated ray tracing cores and no tensor cores. It relies on the Pascal architecture’s standard CUDA cores for all compute tasks. This is a critical limitation for modern workloads: ray tracing in supported applications must run on shader units, which is inefficient compared to the dedicated RT cores found in RTX-series cards. The GeForce RTX 4090, for example, leverages both RT and tensor cores to accelerate ray tracing and DLSS, giving it a massive advantage in those specific tasks.

API support is solid for the card’s age. DirectX 12 (12_1) is fully supported, as are OpenGL 4.6 and Vulkan 1.4. That means the P40 can run modern compute APIs and Vulkan-based workloads without issue. The absence of display outputs is notable—this is a headless compute card, designed for servers and workstations where video output is handled by a separate GPU or iGPU. For a builder considering this card for a desktop, that means no direct monitor connection. The feature set is purely compute-focused: FP32 performance is 11.76 TFLOPS, with FP16 throttled to 183.7 GFLOPS at a 1:64 ratio. That heavy FP16 penalty makes the card unsuitable for AI training workloads that rely on half-precision math, though it can still handle FP32 inference tasks.

How It Compares

NVIDIA GeForce RTX 4090 — The RTX 4090 leads by 0.5% in average benchmark score, but that margin is misleading. The RTX 4090 achieves its score with newer architecture, higher clocks, and dedicated RT/tensor cores. The P40 matches it in generic compute benchmarks only because of its massive 24 GB memory pool and 3840 shaders. In any ray-traced or AI-accelerated task, the RTX 4090 will vastly outperform the P40. The P40’s only advantage is memory capacity, which can matter for datasets that exceed 24 GB.

NVIDIA Tesla T4 — The T4 scores 0.9% higher on average. That is a close result, but the T4 is a much newer card with lower power draw and smaller physical footprint. The T4 also supports FP16 at a much higher rate, making it better suited for inference workloads that use half precision. The P40’s 24 GB versus the T4’s smaller memory is the key differentiator—if you need capacity over speed, the P40 wins. If you need efficiency and modern features, the T4 is the better choice.

AMD Radeon Pro Vega 56 — The Vega 56 leads by 1.4%. This is a workstation card with HBM2 memory, which gives it higher bandwidth than the P40’s GDDR5. In memory-heavy benchmarks, the Vega 56 likely pulls ahead. However, the P40 offers more VRAM (24 GB versus Vega 56’s standard 8 GB), making it the better option for large-scale rendering or data sets. The Vega 56 also has no RT or tensor cores, so neither card excels at modern ray tracing.

NVIDIA Quadro P6000 — The P6000 is 1.8% faster on average. This is the closest architectural comparison, as both cards use the GP102 chip and Pascal architecture. The P6000 has higher clock speeds and a similar memory setup, but the P40’s 24 GB matches the P6000’s capacity. The P6000 is a workstation card with display outputs, making it more flexible for desktop use. The P40 is strictly a compute card, so the P6000 wins for any workload that requires a monitor.

Who Should Consider It

The Tesla P40 is for users who need maximum VRAM capacity at a low cost, but it is not for gamers or AI researchers working with FP16. Benchmark results show the card holding its own in synthetic tests, but the lack of RT and tensor cores means it falls behind in modern games with ray tracing or DLSS. For traditional rasterization at 1440p or 4K, the 24 GB pool allows for ultra-high texture settings without stutter, but the 347.1 GB/s bandwidth and Pascal-era shading units will cap frame rates below what a modern RTX card delivers.

This card is most suitable for FP32 compute workloads where memory capacity is the bottleneck. Examples include rendering large scenes in 3D software, running multiple virtual machines with GPU passthrough, or processing big datasets in scientific computing. The 91st percentile ranking indicates it still outperforms the vast majority of GPUs ever released, so it is not a slouch. However, the 250 W TDP and 600 W suggested PSU requirement mean it is not an energy-efficient choice. If you have a workload that fits entirely within 24 GB and does not require FP16 or ray tracing, the P40 is a viable option. If you need modern features, look elsewhere.

FAQ

Q: Does the Tesla P40 support ray tracing?

A: No. The card has no dedicated ray tracing cores, and any ray tracing must be processed on the standard shading units, which is inefficient.

Q: Why does the Tesla P40 have no display outputs?

A: It is designed as a headless compute accelerator for servers and data centers. Video output would be handled by a separate GPU or integrated graphics.

Q: Can the Tesla P40 handle 4K gaming?

A: It can store 4K textures in its 24 GB memory, but the 347.1 GB/s bandwidth and lack of RT cores will limit performance in modern titles. It is not a gaming card.

Q: How does the Tesla P40 compare to the GeForce RTX 4090?

A: The RTX 4090 is only 0.5% faster in average benchmark score, but it vastly outperforms the P40 in ray tracing and AI workloads due to its dedicated hardware.

Q: Is the Tesla P40 good for AI training?

A: No. Its FP16 performance is 183.7 GFLOPS at a 1:64 ratio, which is severely limited. It is only suitable for FP32 inference tasks.

Q: What is the memory bandwidth of the Tesla P40?

A: The card has 347.1 GB/s of bandwidth, provided by 24 GB of GDDR5 on a 384-bit bus.

Detailed benchmark scores and charts for the NVIDIA Tesla P40 are below.

Benchmark Scores

geekbench_openclSource

Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA Tesla P40 handles parallel computing tasks like video encoding and scientific simulations. OpenCL is widely supported across different GPU vendors and platforms. Higher scores benefit applications that leverage GPU acceleration for non-graphics workloads.

geekbench_opencl #178 of 650
62,017
16%
Max: 388,405
Compare with other GPUs

geekbench_vulkanSource

Geekbench Vulkan tests GPU compute using the modern low-overhead Vulkan API. This shows how NVIDIA Tesla P40 performs with next-generation graphics and compute workloads.

geekbench_vulkan #136 of 446
68,172
18%
Max: 376,915

Popular NVIDIA Tesla P40 Comparisons

See how the Tesla P40 stacks up against similar graphics cards from the same generation and competing brands.

Compare with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs