GEFORCE

NVIDIA Tesla P100 SXM2

NVIDIA graphics card specifications and benchmark scores

16 GB
VRAM
1480
MHz Boost
300W
TDP
4096
Bus Width

At a Glance

NVIDIA
VRAM 16 GB
Boost Clock 1,480 MHz
Shaders 3,584
Bus Width 4096-bit
TDP 300W
Memory Type HBM2
Architecture Pascal
nm
Process 16 nm
Released Apr 2016

NVIDIA Tesla P100 SXM2 Specifications

Tesla P100 SXM2 GPU Core

Shader units and compute resources

The NVIDIA Tesla P100 SXM2 GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
3,584
Shaders
3,584
TMUs
224
ROPs
96
SM Count
56

Tesla P100 SXM2 Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the Tesla P100 SXM2's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Tesla P100 SXM2 by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
1328 MHz
Base Clock
1,328 MHz
Boost Clock
1480 MHz
Boost Clock
1,480 MHz
Memory Clock
715 MHz 1430 Mbps effective
GDDR GDDR 6X 6X

NVIDIA's Tesla P100 SXM2 Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Tesla P100 SXM2's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
16 GB
VRAM
16,384 MB
Memory Type
HBM2
VRAM Type
HBM2
Memory Bus
4096 bit
Bus Width
4096-bit
Bandwidth
732.2 GB/s

Tesla P100 SXM2 by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the Tesla P100 SXM2, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
24 KB (per SM)
L2 Cache
4 MB

Tesla P100 SXM2 Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA Tesla P100 SXM2 against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
10.61 TFLOPS
FP64 (Double)
5.304 TFLOPS (1:2)
FP16 (Half)
21.22 TFLOPS (2:1)
Pixel Rate
142.1 GPixel/s
Texture Rate
331.5 GTexel/s

Pascal Architecture & Process

Manufacturing and design details

The NVIDIA Tesla P100 SXM2 is built on NVIDIA's Pascal architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Tesla P100 SXM2 will perform in GPU benchmarks compared to previous generations.

Architecture
Pascal
GPU Name
GP100
Process Node
16 nm
Foundry
TSMC
Transistors
15,300 million
Die Size
610 mm²
Density
25.1M / mm²

NVIDIA's Tesla P100 SXM2 Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA Tesla P100 SXM2 determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Tesla P100 SXM2 to maintain boost clocks without throttling.

TDP
300 W
TDP
300W
Power Connectors
None
Suggested PSU
700 W

Tesla P100 SXM2 by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA Tesla P100 SXM2 are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
SXM Module
Bus Interface
PCIe 3.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA Tesla P100 SXM2. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 (12_1)
DirectX
12 (12_1)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.3
Vulkan
1.3
OpenCL
3.0
CUDA
6.0
Shader Model
6.0

Tesla P100 SXM2 Product Information

Release and pricing details

The NVIDIA Tesla P100 SXM2 is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Tesla P100 SXM2 by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Apr 2016
Production
End-of-life
Predecessor
Tesla Maxwell
Successor
Tesla Volta

Tesla P100 SXM2 Benchmark Scores

No benchmark data available for this GPU.

About NVIDIA Tesla P100 SXM2

# NVIDIA Tesla P100 SXM2 Analysis

The NVIDIA Tesla P100 SXM2 is a compute-oriented accelerator built on the Pascal architecture, targeting data center workloads rather than consumer gaming. With its 16 GB of HBM2 memory, 4096-bit bus, and 732.2 GB/s bandwidth, this card prioritizes memory throughput and FP32 compute over rasterization features. The data indicates a product at the 50th percentile among all GPUs, placing it exactly at the median of the performance distribution — a fitting representation of its dual role as a capable compute workhorse but an aging gaming solution. It is end-of-life, released on 2016-04-04, and succeeded the Tesla Maxwell generation before being replaced by Tesla Volta.

Who Should Consider It

The Tesla P100 SXM2 is not a card for the average desktop user. Its SXM module form factor, lack of display outputs, and absence of power connectors immediately disqualify it from standard consumer builds. Instead, this accelerator targets professionals and researchers who need high-throughput FP32 compute for simulation, scientific computing, or machine learning inference tasks that fit within 16 GB of HBM2 memory. The 10.61 TFLOPS of FP32 performance places it in a competitive range for single-precision workloads, while the FP16 capability of 21.22 TFLOPS (at a 2:1 ratio) offers a doubling of throughput for mixed-precision applications that can leverage reduced precision.

For gaming or consumer resolution-based recommendations, the data provides no direct benchmarks — the benchmark array is empty. However, the architectural facts suggest this card was never designed for real-time rendering: it has no RT cores and no tensor cores, meaning any ray-traced workloads would rely entirely on traditional shader computation. With 3584 shading units, 224 texture mapping units, and 96 ROPs, the raw rasterization throughput is present — 142.1 GPixel/s pixel rate and 331.5 GTexel/s texture rate — but the lack of display outputs means you could not connect a monitor without a secondary GPU. The 50th percentile ranking among all GPUs indicates that for any consumer scenario, this card sits in the middle of the pack, neither a high-end performer nor a low-end bargain.

Ray Tracing and Feature Set

The Pascal architecture predates NVIDIA's dedicated ray tracing hardware. The fact pack lists no RT cores and no tensor cores, confirming that this GPU has no specialized acceleration for ray-traced effects or AI-based rendering features. Any ray tracing would have to be performed through compute shaders, which is inefficient compared to dedicated hardware found in newer architectures. The API support includes DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.3, which means the card is nominally compatible with modern graphics APIs. However, the absence of RT cores means that even with these APIs, hardware-accelerated ray tracing is unavailable, and the card would fall back to software-based approaches that significantly impact performance.

The feature set is fundamentally compute-oriented. The 732.2 GB/s memory bandwidth, achieved through HBM2 across a 4096-bit bus, is the standout feature — it provides exceptional data throughput for memory-bound workloads. The 16 GB VRAM capacity allows large datasets to reside on-card without spilling to system memory. The texture rate of 331.5 GTexel/s and pixel rate of 142.1 GPixel/s show that the rasterization units are present but were not the design focus. For cloud rendering or virtual desktop scenarios, the lack of display outputs means this card must be paired with a separate GPU for any visual output.

Power and Cooling

The Tesla P100 SXM2 carries a 300 W TDP, which is substantial but not extreme for a data center accelerator. The suggested PSU rating is 700 W, which accounts for the rest of the system's power draw alongside the GPU. The power connectors are listed as "None," which is typical for SXM modules — they draw power through the motherboard or a baseboard management controller rather than through auxiliary PCIe power connectors. This means the card cannot be installed in a standard desktop PCIe slot without a custom adapter or carrier board, and the power delivery must be handled by the host system's design.

Cooling is not specified in terms of a specific cooler size or type, and the slot width is described as "SXM Module," which indicates a bare module that requires a server chassis with appropriate airflow or liquid cooling infrastructure. The 300 W TDP generates meaningful heat that must be dissipated, and in a dense server environment, this requires careful thermal management. The 16 nm process node from TSMC, with 15,300 million transistors on a 610 mm² die, gives a transistor density of 25.1M per mm² — this is relatively low by modern standards, meaning the chip runs at modest clock speeds of 1328 MHz base and 1480 MHz boost to stay within the 300 W envelope.

FAQ

Q: Does the Tesla P100 SXM2 support ray tracing?

A: No. The fact pack lists no RT cores, meaning there is no dedicated ray tracing hardware. The card relies on traditional shader-based compute for any ray tracing workloads, which is significantly less efficient than dedicated RT hardware found in newer architectures.

Q: Can I use this card for gaming?

A: The card has no display outputs, so it cannot connect to a monitor directly. While it has 3584 shading units and supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.3, its SXM form factor and lack of video outputs make it unsuitable for consumer gaming without a secondary GPU for display output.

Q: What is the memory configuration of the Tesla P100 SXM2?

A: It has 16 GB of HBM2 memory on a 4096-bit bus, providing 732.2 GB/s of memory bandwidth. The memory clock is 715 MHz, which translates to 1430 Mbps effective. This configuration is designed for high-throughput data access in compute workloads.

Q: What power supply do I need for this card?

A: The suggested PSU rating is 700 W. The card itself has a 300 W TDP, and since it uses SXM module power delivery with no power connectors, the motherboard or carrier board must supply power. A 700 W PSU is recommended to handle the full system load.

Q: Is the Tesla P100 SXM2 still in production?

A: No, the production status is end-of-life. It was released on 2016-04-04, succeeded the Tesla Maxwell generation, and was followed by the Tesla Volta generation. It is based on the Pascal architecture with the GP100 chip.

Q: What are the FP32 and FP16 compute capabilities?

A: The card delivers 10.61 TFLOPS of FP32 performance and 21.22 TFLOPS of FP16 performance (at a 2:1 ratio). This means FP16 throughput is exactly double the FP32 throughput, which is useful for mixed-precision workloads that can tolerate reduced precision.

How It Compares

The nearestRivals field is empty in the fact pack, meaning there are no direct comparison points provided for this specific card. The percentileVsAllGpus value of 50 indicates that the Tesla P100 SXM2 sits exactly at the median of all GPUs in the benchmark database — half of all GPUs perform better, and half perform worse. This positioning is consistent with a 2016-era compute card that has been superseded by multiple generations of newer hardware. Without specific rival names, scores, or deltaPct values, the comparison must rely on the percentile ranking alone.

The card's position in the product stack is defined by its generation: it is the successor to Tesla Maxwell and the predecessor to Tesla Volta. This places it in the Pascal generation, which was a significant architectural leap for compute performance at the time. However, the empty benchmark array means no empirical scores are available to quantify its performance against any specific competitor. The absence of rivalry data suggests that this card is typically evaluated in the context of data center deployments rather than head-to-head gaming benchmarks, and its 50th percentile ranking reflects a broad middle-ground performance profile across all GPU use cases.

Memory Subsystem

The memory subsystem is the defining feature of the Tesla P100 SXM2. It utilizes 16 GB of HBM2 memory, which is a high-bandwidth memory technology designed for data-intensive workloads. The bus width is an enormous 4096-bit, which is four times wider than typical GDDR5 implementations of the same era. This wide bus compensates for the relatively modest memory clock of 715 MHz (1430 Mbps effective), resulting in a total memory bandwidth of 732.2 GB/s. This bandwidth figure is critical for high-resolution compute tasks, large dataset processing, and any workload that requires frequent memory access.

For high-resolution scenarios, the 16 GB capacity allows large textures, models, or datasets to reside entirely in VRAM. The 732.2 GB/s bandwidth ensures that data can be fed to the 3584 shading units at a rate that keeps them busy. In contrast to consumer GPUs that might have smaller bus widths and lower bandwidth, this card's memory subsystem is optimized for throughput rather than latency. The HBM2 stack layout also reduces the physical footprint, which is why the card can fit into an SXM module form factor. The memory bandwidth of 732.2 GB/s is a direct enabler for the 10.61 TFLOPS FP32 throughput, as compute units require consistent data flow to reach peak performance.

Benchmark Performance

The benchmark data for the Tesla P100 SXM2 is notably sparse — the benchmarks array is empty, and the avgBenchmarkScore is 0. This means there are no empirical performance scores recorded in the fact pack to analyze. The only performance indicator is the percentileVsAllGpus value of 50, which places this card at the exact median of all GPUs. This percentile ranking suggests that the card performs better than half of all GPUs in the database and worse than the other half, making it a middle-of-the-road performer in the overall GPU landscape.

Without specific benchmark scores or nearestRivals data, it is impossible to provide exact percentage deltas or direct comparisons. The FP32 throughput of 10.61 TFLOPS and FP16 throughput of 21.22 TFLOPS are the primary compute metrics available. The FP16 figure at a 2:1 ratio indicates that the card can process half-precision data at twice the rate of full precision, which is a valuable feature for machine learning inference where reduced precision is often acceptable. The texture rate of 331.5 GTexel/s and pixel rate of 142.1 GPixel/s provide additional throughput figures, but without rival scores, these numbers exist in isolation. The 50th percentile ranking is the only comparative data point, and it indicates that this card delivers average performance relative to the entire GPU population, which is consistent with its age and compute-focused design rather than gaming optimization.

The AMD Equivalent of Tesla P100 SXM2

Looking for a similar graphics card from AMD? The AMD Radeon RX 480 offers comparable performance and features in the AMD lineup.

AMD Radeon RX 480

AMD • 8 GB VRAM

View Specs Compare

Popular NVIDIA Tesla P100 SXM2 Comparisons

See how the Tesla P100 SXM2 stacks up against similar graphics cards from the same generation and competing brands.

Compare Tesla P100 SXM2 with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs