NVIDIA Tesla P40
NVIDIA graphics card specifications and benchmark scores
At a Glance
NVIDIANVIDIA Tesla P40 Specifications
GPU Core
Shader units and compute resources
The NVIDIA Tesla P40 GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.
Tesla P40 Clock Speeds
GPU and memory frequencies
Clock speeds directly impact the Tesla P40's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Tesla P40 by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.
NVIDIA's Tesla P40 Memory
VRAM capacity and bandwidth
VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Tesla P40's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.
Tesla P40 by NVIDIA Cache
On-chip cache hierarchy
On-chip cache provides ultra-fast data access for the Tesla P40, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.
Tesla P40 Theoretical Performance
Compute and fill rates
Theoretical performance metrics provide a baseline for comparing the NVIDIA Tesla P40 against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.
Pascal Architecture & Process
Manufacturing and design details
The NVIDIA Tesla P40 is built on NVIDIA's Pascal architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Tesla P40 will perform in GPU benchmarks compared to previous generations.
Power & Thermal
TDP and power requirements
Power specifications for the NVIDIA Tesla P40 determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Tesla P40 to maintain boost clocks without throttling.
Tesla P40 by NVIDIA Physical & Connectivity
Dimensions and outputs
Physical dimensions of the NVIDIA Tesla P40 are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.
NVIDIA API Support
Graphics and compute APIs
API support determines which games and applications can fully utilize the NVIDIA Tesla P40. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.
Tesla P40 Product Information
Release and pricing details
The NVIDIA Tesla P40 is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Tesla P40 by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.
About NVIDIA Tesla P40
NVIDIA’s Tesla P40 is a Pascal-era compute accelerator built for data center workloads, not desktop gaming. With a 24 GB memory pool and a 384-bit bus, it was a high-capacity inference and rendering card for its time. Benchmark results place it just shy of much newer hardware, though the architecture’s age shows in specific workloads. The launch MSRP is 5,699 USD.
Benchmark Performance
The Tesla P40 achieves an average benchmark score of 66,127 across available tests, which places it in the 91st percentile of all GPUs tracked. That is a strong showing for a card released in 2016, and it remains competitive against several newer products in synthetic workloads. In Geekbench OpenCL, the card scores 62,017 points, while its Vulkan score reaches 70,237 points, indicating that the compute-oriented design scales well with modern API overhead.
The delta to its nearest rival is razor-thin. The GeForce RTX 4090 averages 66,473 points, which is only 0.5% ahead of the Tesla P40. That margin is negligible in real-world terms, and it reflects the fact that the P40’s raw shader throughput and memory bandwidth still hold up in compute-heavy benchmarks. The Tesla T4, a newer and more power-efficient data center card, scores 66,733 points, just 0.9% ahead. The AMD Radeon Pro Vega 56 sits 1.4% higher at 67,097 points, and the Quadro P6000—a direct workstation sibling with similar Pascal DNA—leads by 1.8% with 67,320 points.
These figures suggest the Tesla P40 is not obsolete in raw compute terms. However, the story changes when you consider that the RTX 4090 achieves its score with far higher clock speeds and newer architecture, while the P40 relies on sheer shader count and memory capacity. The gap between the P40 and the RTX 4090 is effectively a tie in these synthetic tests, but that does not translate to equal performance in modern games or AI workloads that leverage tensor cores. The data indicates the P40 punches above its age in generic compute, but it lacks the specialized hardware that pushes newer cards ahead in targeted tasks.
Memory Subsystem
The Tesla P40 carries 24 GB of GDDR5 memory on a 384-bit bus, producing 347.1 GB/s of bandwidth. Memory speed is rated at 1808 MHz, which translates to 7.2 Gbps effective. That capacity is the card’s defining feature—24 GB was enormous in 2016 and remains useful for large datasets, high-resolution textures, and multi-model inference workloads that exceed the VRAM of most consumer cards.
For high-resolution rendering, the 347.1 GB/s bandwidth is adequate but not exceptional by modern standards. Compare that to the RTX 4090, which uses faster GDDR6X on a wider 384-bit interface; the P40’s bandwidth is roughly half of what newer flagships offer. In practice, this means the P40 can hold massive scenes in memory, but moving data across the bus will be slower than on contemporary cards. The 24 GB pool also allows for 4K and 8K texture sets without swapping to system memory, which is a significant advantage over 8 GB or 12 GB cards. However, the GDDR5 type and effective speed cap the card’s ability to feed its 3840 shading units at peak load. Benchmark results show the card keeping pace with newer rivals in aggregate, but memory-bound tasks will reveal the bandwidth deficit.
Ray Tracing and Feature Set
The Tesla P40 has no dedicated ray tracing cores and no tensor cores. It relies on the Pascal architecture’s standard CUDA cores for all compute tasks. This is a critical limitation for modern workloads: ray tracing in supported applications must run on shader units, which is inefficient compared to the dedicated RT cores found in RTX-series cards. The GeForce RTX 4090, for example, leverages both RT and tensor cores to accelerate ray tracing and DLSS, giving it a massive advantage in those specific tasks.
API support is solid for the card’s age. DirectX 12 (12_1) is fully supported, as are OpenGL 4.6 and Vulkan 1.4. That means the P40 can run modern compute APIs and Vulkan-based workloads without issue. The absence of display outputs is notable—this is a headless compute card, designed for servers and workstations where video output is handled by a separate GPU or iGPU. For a builder considering this card for a desktop, that means no direct monitor connection. The feature set is purely compute-focused: FP32 performance is 11.76 TFLOPS, with FP16 throttled to 183.7 GFLOPS at a 1:64 ratio. That heavy FP16 penalty makes the card unsuitable for AI training workloads that rely on half-precision math, though it can still handle FP32 inference tasks.
How It Compares
NVIDIA GeForce RTX 4090 — The RTX 4090 leads by 0.5% in average benchmark score, but that margin is misleading. The RTX 4090 achieves its score with newer architecture, higher clocks, and dedicated RT/tensor cores. The P40 matches it in generic compute benchmarks only because of its massive 24 GB memory pool and 3840 shaders. In any ray-traced or AI-accelerated task, the RTX 4090 will vastly outperform the P40. The P40’s only advantage is memory capacity, which can matter for datasets that exceed 24 GB.
NVIDIA Tesla T4 — The T4 scores 0.9% higher on average. That is a close result, but the T4 is a much newer card with lower power draw and smaller physical footprint. The T4 also supports FP16 at a much higher rate, making it better suited for inference workloads that use half precision. The P40’s 24 GB versus the T4’s smaller memory is the key differentiator—if you need capacity over speed, the P40 wins. If you need efficiency and modern features, the T4 is the better choice.
AMD Radeon Pro Vega 56 — The Vega 56 leads by 1.4%. This is a workstation card with HBM2 memory, which gives it higher bandwidth than the P40’s GDDR5. In memory-heavy benchmarks, the Vega 56 likely pulls ahead. However, the P40 offers more VRAM (24 GB versus Vega 56’s standard 8 GB), making it the better option for large-scale rendering or data sets. The Vega 56 also has no RT or tensor cores, so neither card excels at modern ray tracing.
NVIDIA Quadro P6000 — The P6000 is 1.8% faster on average. This is the closest architectural comparison, as both cards use the GP102 chip and Pascal architecture. The P6000 has higher clock speeds and a similar memory setup, but the P40’s 24 GB matches the P6000’s capacity. The P6000 is a workstation card with display outputs, making it more flexible for desktop use. The P40 is strictly a compute card, so the P6000 wins for any workload that requires a monitor.
Who Should Consider It
The Tesla P40 is for users who need maximum VRAM capacity at a low cost, but it is not for gamers or AI researchers working with FP16. Benchmark results show the card holding its own in synthetic tests, but the lack of RT and tensor cores means it falls behind in modern games with ray tracing or DLSS. For traditional rasterization at 1440p or 4K, the 24 GB pool allows for ultra-high texture settings without stutter, but the 347.1 GB/s bandwidth and Pascal-era shading units will cap frame rates below what a modern RTX card delivers.
This card is most suitable for FP32 compute workloads where memory capacity is the bottleneck. Examples include rendering large scenes in 3D software, running multiple virtual machines with GPU passthrough, or processing big datasets in scientific computing. The 91st percentile ranking indicates it still outperforms the vast majority of GPUs ever released, so it is not a slouch. However, the 250 W TDP and 600 W suggested PSU requirement mean it is not an energy-efficient choice. If you have a workload that fits entirely within 24 GB and does not require FP16 or ray tracing, the P40 is a viable option. If you need modern features, look elsewhere.
FAQ
Q: Does the Tesla P40 support ray tracing?
A: No. The card has no dedicated ray tracing cores, and any ray tracing must be processed on the standard shading units, which is inefficient.
Q: Why does the Tesla P40 have no display outputs?
A: It is designed as a headless compute accelerator for servers and data centers. Video output would be handled by a separate GPU or integrated graphics.
Q: Can the Tesla P40 handle 4K gaming?
A: It can store 4K textures in its 24 GB memory, but the 347.1 GB/s bandwidth and lack of RT cores will limit performance in modern titles. It is not a gaming card.
Q: How does the Tesla P40 compare to the GeForce RTX 4090?
A: The RTX 4090 is only 0.5% faster in average benchmark score, but it vastly outperforms the P40 in ray tracing and AI workloads due to its dedicated hardware.
Q: Is the Tesla P40 good for AI training?
A: No. Its FP16 performance is 183.7 GFLOPS at a 1:64 ratio, which is severely limited. It is only suitable for FP32 inference tasks.
Q: What is the memory bandwidth of the Tesla P40?
A: The card has 347.1 GB/s of bandwidth, provided by 24 GB of GDDR5 on a 384-bit bus.
Detailed benchmark scores and charts for the NVIDIA Tesla P40 are below.
Benchmark Scores
geekbench_openclSource
Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA Tesla P40 handles parallel computing tasks like video encoding and scientific simulations. OpenCL is widely supported across different GPU vendors and platforms. Higher scores benefit applications that leverage GPU acceleration for non-graphics workloads.
geekbench_vulkanSource
Geekbench Vulkan tests GPU compute using the modern low-overhead Vulkan API. This shows how NVIDIA Tesla P40 performs with next-generation graphics and compute workloads.
Popular NVIDIA Tesla P40 Comparisons
See how the Tesla P40 stacks up against similar graphics cards from the same generation and competing brands.
Compare with Other GPUs
Select another GPU to compare specifications and benchmarks side-by-side.
Browse GPUs