GEFORCE

NVIDIA Tesla P100 PCIe 16 GB

NVIDIA graphics card specifications and benchmark scores

16 GB
VRAM
1329
MHz Boost
250W
TDP
4096
Bus Width

At a Glance

NVIDIA
VRAM 16 GB
Boost Clock 1,329 MHz
Shaders 3,584
Bus Width 4096-bit
TDP 250W
Memory Type HBM2
Architecture Pascal
nm
Process 16 nm
Released Jun 2016

NVIDIA Tesla P100 PCIe 16 GB Specifications

Tesla P100 PCIe 16 GB GPU Core

Shader units and compute resources

The NVIDIA Tesla P100 PCIe 16 GB GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
3,584
Shaders
3,584
TMUs
224
ROPs
96
SM Count
56

Tesla P100 PCIe 16 GB Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the Tesla P100 PCIe 16 GB's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Tesla P100 PCIe 16 GB by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
1190 MHz
Base Clock
1,190 MHz
Boost Clock
1329 MHz
Boost Clock
1,329 MHz
Memory Clock
715 MHz 1430 Mbps effective
GDDR GDDR 6X 6X

NVIDIA's Tesla P100 PCIe 16 GB Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Tesla P100 PCIe 16 GB's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
16 GB
VRAM
16,384 MB
Memory Type
HBM2
VRAM Type
HBM2
Memory Bus
4096 bit
Bus Width
4096-bit
Bandwidth
732.2 GB/s

Tesla P100 PCIe 16 GB by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the Tesla P100 PCIe 16 GB, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
24 KB (per SM)
L2 Cache
4 MB

Tesla P100 PCIe 16 GB Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA Tesla P100 PCIe 16 GB against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
9.526 TFLOPS
FP64 (Double)
4.763 TFLOPS (1:2)
FP16 (Half)
19.05 TFLOPS (2:1)
Pixel Rate
127.6 GPixel/s
Texture Rate
297.7 GTexel/s

Pascal Architecture & Process

Manufacturing and design details

The NVIDIA Tesla P100 PCIe 16 GB is built on NVIDIA's Pascal architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Tesla P100 PCIe 16 GB will perform in GPU benchmarks compared to previous generations.

Architecture
Pascal
GPU Name
GP100
Process Node
16 nm
Foundry
TSMC
Transistors
15,300 million
Die Size
610 mm²
Density
25.1M / mm²

NVIDIA's Tesla P100 PCIe 16 GB Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA Tesla P100 PCIe 16 GB determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Tesla P100 PCIe 16 GB to maintain boost clocks without throttling.

TDP
250 W
TDP
250W
Power Connectors
1x 8-pin
Suggested PSU
600 W

Tesla P100 PCIe 16 GB by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA Tesla P100 PCIe 16 GB are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
Dual-slot
Length
267 mm 10.5 inches
Bus Interface
PCIe 3.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA Tesla P100 PCIe 16 GB. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 (12_1)
DirectX
12 (12_1)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.3
Vulkan
1.3
OpenCL
3.0
CUDA
6.0
Shader Model
6.0

Tesla P100 PCIe 16 GB Product Information

Release and pricing details

The NVIDIA Tesla P100 PCIe 16 GB is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Tesla P100 PCIe 16 GB by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Jun 2016
Launch Price
5,699 USD
Production
End-of-life
Predecessor
Tesla Maxwell
Successor
Tesla Volta

Tesla P100 PCIe 16 GB Benchmark Scores

geekbench_openclSource

Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA Tesla P100 PCIe 16 GB handles parallel computing tasks like video encoding and scientific simulations. OpenCL is widely supported across different GPU vendors and platforms. Higher scores benefit applications that leverage GPU acceleration for non-graphics workloads.

geekbench_opencl #130 of 650
79,605
20%
Max: 388,405

About NVIDIA Tesla P100 PCIe 16 GB

Benchmark Performance

The NVIDIA Tesla P100 PCIe 16 GB occupies a peculiar position in the hardware landscape. It is a compute-oriented accelerator from the Pascal generation, yet its data sheet reveals a device that is neither a pure server part nor a gaming card. The most striking observation from the specification table is the absence of any direct benchmark scores; the `avgBenchmarkScore` field is set to zero, and the `nearestRivals` array is empty. This means the P100 cannot be positioned against contemporary competitors using quantitative deltas. Instead, one must interpret its theoretical peak throughput figures as the primary performance indicators.

The FP32 compute rating of 9.526 TFLOPS places the P100 firmly in the high-performance tier of its era. For context, this is a figure that would have been exceptional for a single-slot accelerator in mid-2016, the release date. The FP16 throughput of 19.05 TFLOPS, achieved via a 2:1 ratio, doubles the FP32 rate, which is unusual for Pascal-generation hardware. This suggests the architecture was designed with mixed-precision workloads in mind, a forward-looking decision that would later become standard in AI accelerators. The pixel rate of 127.6 GPixel/s and texture rate of 297.7 GTexel/s are derived from the 96 ROPs and 224 TMUs respectively, running at the boost clock of 1329 MHz. These are not record-breaking numbers for a 2016 flagship, but they are respectable for a compute-first product.

The `percentileVsAllGpus` field is set to 50, meaning the P100 sits exactly at the median of all GPUs ever tracked by the database. This is a curious result. A 50th percentile ranking suggests that while the P100 has substantial raw compute power, its overall benchmark aggregate is dragged down by factors such as lack of display outputs, no dedicated ray tracing hardware, and an end-of-life production status. The data implies that in purely compute-bound tasks, the P100 would outperform the majority of consumer cards, but in gaming or workstation graphics workloads, it would fall behind modern parts that integrate dedicated acceleration blocks. The empty `nearestRivals` array further complicates analysis; without named competitors or delta percentages, the P100 exists in a vacuum within this dataset.

The 16 nm TSMC process node and 15,300 million transistor count on a 610 mm² die yield a transistor density of 25.1M per mm². This is relatively low density by modern standards, but for 2016, it represented a massive chip. The large die size and high transistor count correlate directly with the 9.526 TFLOPS FP32 figure, but they also explain the 250 W TDP. The boost clock of 1329 MHz is modest compared to later Pascal consumer parts, indicating that the P100 was not optimized for raw clock speed but rather for sustained compute throughput across thousands of cores.

What the benchmark data does not show is equally important. The P100 has no entry for `rtCores` or `tensorCores`, confirming that this is a pre-Turing, pre-Volta architecture. This absence is a critical differentiator when considering the P100 against any modern accelerator. The FP16 2:1 ratio hints at early tensor-like capabilities, but without dedicated tensor cores, the P100 cannot accelerate the matrix operations that dominate contemporary AI workloads. The 50th percentile ranking, therefore, is a reflection of a device that was a compute powerhouse in its prime but has been surpassed by specialized hardware in every measurable modern benchmark category.

Memory Subsystem

The memory configuration of the Tesla P100 PCIe 16 GB is one of its most distinctive features. It employs 16 GB of HBM2 memory on a 4096-bit bus, yielding a bandwidth of 732.2 GB/s. This is an enormous memory bandwidth figure, especially when compared to the GDDR5 and GDDR5X solutions common in 2016 consumer cards. The 4096-bit bus width is the key enabler; HBM2 stacks are inherently wide, and the P100 leverages this to achieve a bandwidth that is roughly 3-4 times that of a typical 256-bit GDDR5 implementation of the same era.

For high-resolution workloads, the memory subsystem is the P100's strongest asset. The 732.2 GB/s bandwidth ensures that data movement is rarely a bottleneck, even when processing large datasets or high-resolution textures. The 16 GB capacity is also substantial, allowing for datasets that would exceed the 8 GB or 12 GB limits of contemporary consumer cards. However, the memory clock is set at 715 MHz, with an effective data rate of 1430 Mbps. This is low by modern standards, but the sheer width of the bus compensates for the modest clock speed. The result is a memory subsystem that prioritizes bandwidth over latency, which is ideal for compute kernels that stream large blocks of data.

The implications for high-resolution rendering are nuanced. In a gaming context, the P100 would handle 4K textures without capacity issues, and the bandwidth would prevent stuttering from texture streaming. However, the lack of display outputs means the P100 cannot drive a monitor directly; it must be paired with a secondary GPU for display purposes. In a compute context, the memory subsystem is unquestionably excellent. The 16 GB HBM2 pool allows for large neural network batches or scientific simulation grids to reside entirely in VRAM, avoiding the performance penalty of host-device data transfers over the PCIe 3.0 x16 bus. The bus interface itself, PCIe 3.0 x16, provides a maximum theoretical bandwidth of 16 GB/s, which is a bottleneck compared to the 732.2 GB/s internal memory bandwidth, but this is a standard limitation for all discrete GPUs.

Ray Tracing and Feature Set

The Tesla P100 PCIe 16 GB does not include any dedicated ray tracing cores or tensor cores. The `rtCores` and `tensorCores` fields are both null, which is consistent with the Pascal architecture from 2016. This is a significant omission for any modern workload that relies on real-time ray tracing or AI-accelerated denoising. The API support is listed as DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.3. The DirectX 12_1 feature level includes support for conservative rasterization and rasterizer-ordered views, but it does not include the DirectX Raytracing (DXR) API, which requires a DirectX 12 Ultimate feature set. Therefore, any ray tracing workload would have to be computed in software or via compute shaders, which would be dramatically slower than dedicated hardware.

The Vulkan 1.3 support is noteworthy because it means the P100 can access modern Vulkan extensions, including those for mesh shaders and variable rate shading, provided the driver supports them. However, without ray tracing acceleration, the P100 is fundamentally limited to rasterization and compute workloads. The FP16 2:1 ratio is the only nod to AI acceleration, and it is a primitive one compared to the tensor cores found in subsequent Volta and Turing architectures. In practice, the P100 can process FP16 data at 19.05 TFLOPS, which can be leveraged for inference tasks if the software is written to use FP16 arithmetic. But for training, the lack of tensor cores means no mixed-precision acceleration via NVIDIA's Tensor Core libraries.

The feature set is further constrained by the absence of display outputs. The `displayOutputs` field is explicitly listed as "No outputs," which means the P100 is a pure compute accelerator. It cannot be used for any visual output, and it relies on a separate GPU for rendering if used in a workstation. This is typical for Tesla-series products, which are designed for datacenter or render-farm deployments rather than desktop use. The API support, while modern on paper, is unlikely to be exercised in a gaming context due to the lack of display outputs. For compute, the Vulkan 1.3 and OpenGL 4.6 support allows for cross-platform compute workloads, but the absence of CUDA-specific details in the fact pack means the primary programming model remains a qualitative advantage that is not quantified here.

Who Should Consider It

The benchmark data, or lack thereof, paints a clear picture: the Tesla P100 PCIe 16 GB is a niche product for compute professionals, not gamers or general consumers. The 50th percentile ranking and zero average benchmark score suggest that in aggregate databases, the P100 is outclassed by many newer cards. However, the specific strengths of the memory subsystem and FP32 throughput indicate that it could still be relevant for certain workloads. Users who process large datasets that fit within 16 GB of HBM2 memory, such as genomics, molecular dynamics, or financial modeling, would benefit from the 732.2 GB/s bandwidth. The FP32 performance of 9.526 TFLOPS is sufficient for double-precision-adjacent tasks, though the fact pack does not list FP64 performance, so one must assume FP64 is either absent or severely reduced.

For gaming, the P100 is a poor choice. No display outputs, no ray tracing cores, and a 250 W TDP make it ill-suited for a desktop gaming rig. The DirectX 12_1 support is adequate for older titles, but modern games that require DirectX 12 Ultimate or ray tracing will not run properly, if at all. The Vulkan 1.3 support is a positive, but again, without display outputs, the user must have a secondary GPU, which introduces latency and complexity. For high-resolution gaming at 4K, the memory bandwidth would be excellent, but the lack of modern feature support negates this advantage.

The intended audience is clear: researchers and engineers who need a large memory pool and high bandwidth for compute tasks, and who have access to a system with a 600 W PSU and a spare PCIe 3.0 x16 slot. The production status is end-of-life, and the release date is 2016, so this is not a card for new builds. It is a legacy part that might be found in used enterprise equipment or as a stopgap for specific compute tasks. The launch MSRP is 5,699 USD, which is listed once here and reflects the enterprise pricing of the time. The 50th percentile ranking indicates that while the P100 is not a top performer, it is also not a bottom-tier part; it sits exactly in the middle of all GPUs, which is a testament to its longevity in compute benchmarks despite its age.

FAQ

Q: What is the memory bandwidth of the Tesla P100 PCIe 16 GB?

A: The memory bandwidth is 732.2 GB/s, achieved via 16 GB of HBM2 memory on a 4096-bit bus with a memory clock of 715 MHz (1430 Mbps effective).

Q: Does the Tesla P100 support real-time ray tracing?

A: No. The `rtCores` field is null, and the DirectX 12 support is limited to the 12_1 feature level, which does not include DirectX Raytracing. There are no dedicated ray tracing cores.

Q: What is the FP32 performance of the Tesla P100?

A: The FP32 compute throughput is 9.526 TFLOPS, while FP16 is 19.05 TFLOPS at a 2:1 ratio. The pixel rate is 127.6 GPixel/s and the texture rate is 297.7 GTexel/s.

Q: Can the Tesla P100 be used as a display adapter?

A: No. The display outputs field is listed as "No outputs," meaning it cannot drive a monitor. It must be paired with a separate GPU for any visual output.

Q: What power supply is recommended for the Tesla P100?

A: The suggested PSU is 600 W, and the card requires a single 8-pin power connector. The TDP is 250 W, and the card is dual-slot in width.

Q: What is the production status of the Tesla P100?

A: The production status is end-of-life. It was released on 2016-06-19, and its predecessor is Tesla Maxwell, with the successor being Tesla Volta.

Power and Cooling

The Tesla P100 PCIe 16 GB has a TDP of 250 W, which is a moderate power draw for a compute accelerator of its era. The suggested PSU is 600 W, which provides a reasonable headroom for the card's peak power consumption plus the rest of the system. The power connector requirement is a single 8-pin, which is straightforward and compatible with most standard power supplies. The slot width is dual-slot, meaning it will occupy two expansion slots in a chassis, and the card length is 267 mm (10.5 inches), which is a standard length for a dual-slot PCIe card. There are no dimensions listed for height or width, but the dual-slot design implies a substantial cooling solution, likely a blower-style cooler that exhausts air out of the chassis.

The 250 W TDP is notable because it is lower than many high-end consumer GPUs of the same generation, which often exceeded 250 W. This suggests that the P100's cooling solution is sufficient for sustained compute workloads, where the card may run at full load for hours or days. The lack of display outputs means there is no need to drive multiple monitors, which reduces the thermal load. The 16 nm process node from TSMC helps keep power consumption in check, despite the 15,300 million transistors. The boost clock of 1329 MHz is modest, which also contributes to the manageable power draw. In a datacenter environment, the dual-slot design and blower cooler are preferred for dense server configurations, as they exhaust hot air directly out of the chassis rather than recirculating it. The 600 W PSU recommendation is a conservative figure that ensures stability even under peak transient loads, which can spike above the 250 W TDP.

How It Compares

The `nearestRivals` array is empty in the fact pack, so there are no direct competitor comparisons with named GPUs or delta percentages. This absence of data is itself informative. It suggests that the Tesla P100 PCIe 16 GB does not have a close competitor within the database's tracking metrics, likely because it is a compute-specific product that does not fit neatly into the consumer or workstation categories where rival comparisons are typically drawn. Without rival names or scores, one cannot quantify how the P100 compares to, say, a Quadro or GeForce part from the same era.

The percentileVsAllGpus field of 50 provides a broad comparison: the P100 is exactly average across all GPUs in the database. This is a strange result for a card with 9.526 TFLOPS FP32 and 732.2 GB/s bandwidth, which are above-average specs. The explanation likely lies in the benchmark aggregation methodology, which may weight gaming performance heavily. Since the P100 has no display outputs, it cannot be benchmarked in gaming scenarios, and thus its overall score is dragged down by missing data. In compute-specific benchmarks, the P100 would likely rank much higher, but the database does not provide such granularity.

Given the lack of rivals, one must rely on qualitative positioning. The predecessor is Tesla Maxwell and the successor is Tesla Volta, which places the P100 between two compute generations. The Tesla Volta successor would introduce tensor cores, which the P100 lacks. This implies that the P100's primary advantage over its successor is not feature support but possibly cost or availability in the used market. The empty rivals array also means there is no deltaPct to cite, so any comparison would be speculative. The data simply does not support a quantitative comparison, and the P100 remains a solitary figure in the benchmark landscape, defined more by its specifications than by its competitive positioning.

The AMD Equivalent of Tesla P100 PCIe 16 GB

Looking for a similar graphics card from AMD? The AMD Radeon RX 480 offers comparable performance and features in the AMD lineup.

AMD Radeon RX 480

AMD • 8 GB VRAM

View Specs Compare

Popular NVIDIA Tesla P100 PCIe 16 GB Comparisons

See how the Tesla P100 PCIe 16 GB stacks up against similar graphics cards from the same generation and competing brands.

Compare Tesla P100 PCIe 16 GB with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs