GEFORCE

NVIDIA Tesla V100 SXM2 16 GB

NVIDIA graphics card specifications and benchmark scores

16 GB
VRAM
1530
MHz Boost
250W
TDP
4096
Bus Width
Tensor Cores

At a Glance

NVIDIA
VRAM 16 GB
Boost Clock 1,530 MHz
Shaders 5,120
Bus Width 4096-bit
TDP 250W
Memory Type HBM2
Architecture Volta
nm
Process 12 nm
Released Jun 2017

NVIDIA Tesla V100 SXM2 16 GB Specifications

GPU Core

Shader units and compute resources

The NVIDIA Tesla V100 SXM2 16 GB GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
5,120
Shaders
5,120
TMUs
320
ROPs
128
SM Count
80

Tesla V100 SXM2 16 GB Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the Tesla V100 SXM2 16 GB's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Tesla V100 SXM2 16 GB by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
1312 MHz
Base Clock
1,312 MHz
Boost Clock
1530 MHz
Boost Clock
1,530 MHz
Memory Clock
876 MHz 1752 Mbps effective
GDDR GDDR 6X 6X

NVIDIA's Tesla V100 SXM2 16 GB Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Tesla V100 SXM2 16 GB's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
16 GB
VRAM
16,384 MB
Memory Type
HBM2
VRAM Type
HBM2
Memory Bus
4096 bit
Bus Width
4096-bit
Bandwidth
897.0 GB/s

Tesla V100 SXM2 16 GB by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the Tesla V100 SXM2 16 GB, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
128 KB (per SM)
L2 Cache
6 MB

Tesla V100 SXM2 16 GB Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA Tesla V100 SXM2 16 GB against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
15.67 TFLOPS
FP64 (Double)
7.834 TFLOPS (1:2)
FP16 (Half)
31.33 TFLOPS (2:1)
Pixel Rate
195.8 GPixel/s
Texture Rate
489.6 GTexel/s

Tesla V100 SXM2 16 GB Ray Tracing & AI

Hardware acceleration features

The NVIDIA Tesla V100 SXM2 16 GB includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the Tesla V100 SXM2 16 GB capable of delivering both stunning graphics and smooth frame rates in modern titles.

Tensor Cores
640

Volta Architecture & Process

Manufacturing and design details

The NVIDIA Tesla V100 SXM2 16 GB is built on NVIDIA's Volta architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Tesla V100 SXM2 16 GB will perform in GPU benchmarks compared to previous generations.

Architecture
Volta
GPU Name
GV100
Process Node
12 nm
Foundry
TSMC
Transistors
21,100 million
Die Size
815 mm²
Density
25.9M / mm²

Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA Tesla V100 SXM2 16 GB determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Tesla V100 SXM2 16 GB to maintain boost clocks without throttling.

TDP
250 W
TDP
250W
Power Connectors
None
Suggested PSU
600 W

Tesla V100 SXM2 16 GB by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA Tesla V100 SXM2 16 GB are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
SXM Module
Bus Interface
PCIe 3.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA Tesla V100 SXM2 16 GB. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 (12_1)
DirectX
12 (12_1)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.4
Vulkan
1.4
OpenCL
3.0
CUDA
7.0
Shader Model
6.8

Tesla V100 SXM2 16 GB Product Information

Release and pricing details

The NVIDIA Tesla V100 SXM2 16 GB is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Tesla V100 SXM2 16 GB by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Jun 2017
Production
End-of-life
Predecessor
Tesla Pascal
Successor
Tesla Turing

About NVIDIA Tesla V100 SXM2 16 GB

Memory Subsystem

The NVIDIA Tesla V100 SXM2 16 GB is equipped with 16 GB of HBM2 memory, a configuration that stands out for its high-bandwidth design. The memory interface is a 4096-bit bus, which is exceptionally wide and directly contributes to the memory's ability to feed data to the GPU's compute units. This wide bus, combined with the HBM2 memory type, yields a total memory bandwidth of 897.0 GB/s. For high-resolution workloads, this bandwidth is a critical asset, as it allows the GPU to handle large datasets and complex scenes without becoming bottlenecked by data transfer speeds. In practical terms, the sheer bandwidth available here is more than sufficient for demanding scientific simulations, deep learning training on large batches, and high-resolution rendering tasks that require rapid access to substantial amounts of data.

The memory clock operates at 876 MHz, with an effective data rate of 1752 Mbps. While the clock speed itself is not notably high, the architecture's reliance on a massive bus width rather than extreme clock speeds is what defines its performance profile. The 16 GB capacity ensures that the GPU can hold large models and datasets entirely in VRAM, which is often a limiting factor for other accelerators. The combination of capacity and bandwidth makes this card particularly well-suited for compute-intensive applications where data residency and throughput are paramount.

Ray Tracing and Feature Set

The Tesla V100 SXM2 16 GB is built on the Volta architecture, which predates dedicated ray tracing hardware. It does not include RT cores, meaning that hardware-accelerated ray tracing is not a feature of this GPU. Instead, the card's compute capabilities are centered around its 5120 shading units and 640 tensor cores. These tensor cores are the standout feature, designed specifically to accelerate deep learning operations, particularly matrix multiplication used in neural network training and inference. The presence of these cores is a defining characteristic of the Volta generation, positioning the V100 as a compute-first accelerator rather than a graphics-oriented card.

In terms of API support, the GPU supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. This means it is compatible with modern graphics APIs, but its design intent is clearly for general-purpose GPU compute, as evidenced by the lack of display outputs. The card has "No outputs," indicating it is not meant to drive a display but rather to be installed in a server or compute node. The tensor cores deliver significant performance for FP16 workloads, with a rate of 31.33 TFLOPS (2:1), while FP32 performance sits at 15.67 TFLOPS. This dual-rate capability highlights the card's optimization for mixed-precision computing, a common requirement in AI and scientific workloads.

Benchmark Performance

The benchmark data for the Tesla V100 SXM2 16 GB is sparse, with no individual scores listed in the available information. However, the GPU holds a percentile rank of 50 when compared against all other GPUs, indicating it sits at the median of the performance distribution in the database. This percentile is a relative measure, suggesting that while the V100 is not the fastest accelerator ever produced, it is also far from the slowest, occupying a middle ground that reflects its age and its specialized compute focus.

Because the nearestRivals list is empty, direct comparative analysis with specific percentage deltas is not possible. The benchmark results indicate that the card's raw compute throughput, as defined by its FP32 and FP16 figures, is substantial, but without peer scores, a relative interpretation is limited to the percentile ranking. The data shows that the V100's performance is sufficient to place it in the 50th percentile of all GPUs, which implies that many newer or more powerful parts will outperform it, but it remains a competent performer for a range of compute tasks. The absence of specific benchmark scores means that conclusions must be drawn from its architectural specifications and its known position in the broader GPU landscape.

How It Compares

As there are no nearest rivals listed in the fact pack, a detailed comparison against specific competing models cannot be provided. The Tesla V100 SXM2 16 GB is positioned within the Tesla Volta generation, with its predecessor being Tesla Pascal and its successor being Tesla Turing. This lineage indicates a clear progression in NVIDIA's data center accelerator line. The Pascal generation, which came before, lacked the specialized tensor cores that the Volta architecture introduced, marking the V100 as a significant architectural leap forward. The Turing generation, which followed, would go on to introduce RT cores, something this card explicitly lacks.

In the absence of direct rival data, the comparison must be framed by the card's own specifications and its historical context. The V100's 16 GB of HBM2 memory and 4096-bit bus set it apart from many of its contemporaries at the time of release, which often relied on GDDR5 or GDDR5X with narrower buses. The 640 tensor cores were a unique selling point, making the V100 a preferred choice for AI research and development. Looking at its successor, the Tesla Turing generation would bring ray tracing capabilities, but the V100's raw FP32 and FP16 compute figures were, at the time, class-leading. The data shows that this card was a pivotal product in the evolution of GPU-accelerated computing, even if its direct comparisons to current hardware are not available in the provided facts.

Power and Cooling

The Tesla V100 SXM2 16 GB has a thermal design power (TDP) of 250 W. This is a moderate power draw for a high-performance accelerator of its era, reflecting the efficiency of the 12 nm process node manufactured by TSMC. The card is designed as an SXM Module, which means it is not a standard PCIe card with its own cooling solution and power connectors. Instead, it is meant to be installed into a proprietary SXM socket, typically found in servers like the NVIDIA DGX systems. Consequently, the card has no power connectors of its own, as power is delivered through the socket interface.

The suggested PSU rating for a system incorporating this module is 600 W. This figure is provided as a guideline for system builders, indicating the minimum recommended power supply capacity to ensure stable operation, accounting for the rest of the system's components. Because the SXM module relies on the host system's cooling infrastructure, there is no integrated fan or heatsink specification provided. The slot width is listed as "SXM Module," confirming its form factor. The lack of display outputs reinforces that this is a compute-only part, with no video connectivity. The power and cooling requirements are therefore dictated by the server chassis design, which must accommodate the 250 W TDP and provide adequate airflow to dissipate the generated heat.

Detailed benchmark scores and charts for the NVIDIA Tesla V100 SXM2 16 GB are below.

Benchmark Scores

geekbench_openclSource

Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA Tesla V100 SXM2 16 GB handles parallel computing tasks like video encoding and scientific simulations.

geekbench_opencl #118 of 650
87,454
23%
Max: 388,405
Compare with other GPUs

geekbench_vulkanSource

Geekbench Vulkan tests GPU compute using the modern low-overhead Vulkan API. This shows how NVIDIA Tesla V100 SXM2 16 GB performs with next-generation graphics and compute workloads. Vulkan offers better CPU efficiency than older APIs like OpenGL.

geekbench_vulkan #42 of 446
141,336
37%
Max: 376,915

Popular NVIDIA Tesla V100 SXM2 16 GB Comparisons

See how the Tesla V100 SXM2 16 GB stacks up against similar graphics cards from the same generation and competing brands.

Compare with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs