GEFORCE

NVIDIA Tesla V100 SXM2 32 GB

NVIDIA graphics card specifications and benchmark scores

32 GB
VRAM
1530
MHz Boost
250W
TDP
4096
Bus Width
Tensor Cores

At a Glance

NVIDIA
VRAM 32 GB
Boost Clock 1,530 MHz
Shaders 5,120
Bus Width 4096-bit
TDP 250W
Memory Type HBM2
Architecture Volta
nm
Process 12 nm
Released Mar 2018

NVIDIA Tesla V100 SXM2 32 GB Specifications

GPU Core

Shader units and compute resources

The NVIDIA Tesla V100 SXM2 32 GB GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
5,120
Shaders
5,120
TMUs
320
ROPs
128
SM Count
80

Tesla V100 SXM2 32 GB Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the Tesla V100 SXM2 32 GB's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Tesla V100 SXM2 32 GB by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
1290 MHz
Base Clock
1,290 MHz
Boost Clock
1530 MHz
Boost Clock
1,530 MHz
Memory Clock
877 MHz 1754 Mbps effective
GDDR GDDR 6X 6X

NVIDIA's Tesla V100 SXM2 32 GB Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Tesla V100 SXM2 32 GB's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
32 GB
VRAM
32,768 MB
Memory Type
HBM2
VRAM Type
HBM2
Memory Bus
4096 bit
Bus Width
4096-bit
Bandwidth
898.0 GB/s

Tesla V100 SXM2 32 GB by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the Tesla V100 SXM2 32 GB, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
128 KB (per SM)
L2 Cache
6 MB

Tesla V100 SXM2 32 GB Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA Tesla V100 SXM2 32 GB against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
15.67 TFLOPS
FP64 (Double)
7.834 TFLOPS (1:2)
FP16 (Half)
31.33 TFLOPS (2:1)
Pixel Rate
195.8 GPixel/s
Texture Rate
489.6 GTexel/s

Tesla V100 SXM2 32 GB Ray Tracing & AI

Hardware acceleration features

The NVIDIA Tesla V100 SXM2 32 GB includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the Tesla V100 SXM2 32 GB capable of delivering both stunning graphics and smooth frame rates in modern titles.

Tensor Cores
640

Volta Architecture & Process

Manufacturing and design details

The NVIDIA Tesla V100 SXM2 32 GB is built on NVIDIA's Volta architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Tesla V100 SXM2 32 GB will perform in GPU benchmarks compared to previous generations.

Architecture
Volta
GPU Name
GV100
Process Node
12 nm
Foundry
TSMC
Transistors
21,100 million
Die Size
815 mm²
Density
25.9M / mm²

Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA Tesla V100 SXM2 32 GB determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Tesla V100 SXM2 32 GB to maintain boost clocks without throttling.

TDP
250 W
TDP
250W
Power Connectors
None
Suggested PSU
600 W

Tesla V100 SXM2 32 GB by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA Tesla V100 SXM2 32 GB are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
SXM Module
Bus Interface
PCIe 3.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA Tesla V100 SXM2 32 GB. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 (12_1)
DirectX
12 (12_1)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.4
Vulkan
1.4
OpenCL
3.0
CUDA
7.0
Shader Model
6.8

Tesla V100 SXM2 32 GB Product Information

Release and pricing details

The NVIDIA Tesla V100 SXM2 32 GB is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Tesla V100 SXM2 32 GB by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Mar 2018
Production
End-of-life
Predecessor
Tesla Pascal
Successor
Tesla Turing

About NVIDIA Tesla V100 SXM2 32 GB

The NVIDIA Tesla V100 SXM2 32 GB is a Volta-architecture compute module built around the GV100 chip. TSMC fabricates it on a 12 nm process, with 21,100 million transistors in an 815 mm² die and a transistor density of 25.9M per mm². The database lists it as End-of-life, with a release date of 2018-03-26. It belongs to the Tesla Volta (Vxx) generation, sitting between Tesla Pascal as predecessor and Tesla Turing as successor.

Memory Subsystem

The memory system is the strongest single feature in the database entry. The V100 SXM2 32 GB uses 32 GB of HBM2, and that memory is attached through a 4096-bit bus. The peak bandwidth is listed as 898.0 GB/s, which is the figure that matters when a workload must constantly move large amounts of data. The memory clock is 877 MHz, with an effective data rate of 1754 Mbps, and that effective rate is exactly double the memory clock. The 4096-bit bus is what turns that effective data rate into 898.0 GB/s of bandwidth.

For high-resolution workloads, capacity and bandwidth perform different roles. The 32 GB pool can hold large framebuffers, textures, or datasets on the module itself. When a working set fits in HBM2, the system does not need to stream the same data repeatedly over the host interface. When data must move, the local memory path is 898.0 GB/s, while the host connection is PCIe 3.0 x16. The gap between those two paths affects the kind of workload this module can handle: large resident buffers and high-bandwidth access. A memory-bound kernel would be constrained by the speed at which data can be fed through the 4096-bit bus, not by a narrow memory path. The combination of 32 GB capacity and 898.0 GB/s bandwidth gives this GPU a memory subsystem built for both size and throughput.

Ray Tracing and Feature Set

The database entry does not quantify ray tracing hardware. The rtCores field is null, so there is no ray tracing core count to report from the data. What the data does quantify are 640 tensor cores. Those tensor cores sit alongside a compute layout of 5120 shading units, 320 TMUs, and 128 ROPs. The listed pixel rate is 195.8 GPixel/s, and the listed texture rate is 489.6 GTexel/s. Those rates are peak specification numbers, driven by the 128 ROPs and 320 TMUs at the boost clock.

API support in the data is DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The DirectX entry is recorded as a feature level, not as a ray tracing feature set. The module also has No display outputs, so these APIs apply to compute or off-screen rendering rather than to driving a monitor directly. The feature set is therefore not a display-oriented feature set. It is a compute-oriented set in which the 640 tensor cores and the FP16 throughput are the most distinctive items. The FP16 figure is listed as 31.33 TFLOPS at a 2:1 ratio against FP32, which makes mixed-precision work an obvious fit. Without an RT core count, any ray tracing capability cannot be established from this database entry.

Benchmark Performance

The benchmarks array is empty, and the avgBenchmarkScore field is 0. There are no sampled results to analyze. The percentileVsAllGpus field is 50, which places this GPU at the midpoint of the database's GPU list, but that percentile is not supported by a nonzero benchmark score. It is a database rank, not a measured performance result.

The quantitative identity of the V100 SXM2 32 GB therefore comes from its specification sheet. The base clock is 1290 MHz, the boost clock is 1530 MHz, and the peak FP32 rate is 15.67 TFLOPS. The peak FP16 rate is 31.33 TFLOPS, exactly double the FP32 rate, matching the 2:1 ratio recorded in the data. The pixel rate is 195.8 GPixel/s, and the texture rate is 489.6 GTexel/s. These are peak throughput figures, not application averages. Without benchmark scores, no deltaPct values can be calculated from rivals, and no relative performance statements can be made from that side of the data. The 50th percentile is the only all-GPU placement in the database, and it should be understood in the context of a missing benchmark average.

Who Should Consider It

This is not a conventional desktop graphics card. The slot width is SXM Module, the display outputs are listed as No outputs, and the power connectors are listed as None. A system using this module should still account for a 250 W TDP and a suggested PSU of 600 W, because power is delivered through the host platform rather than through on-board connectors. The intended environment is a server or compute node that can integrate an SXM module. A user who needs a display connection cannot get one from this device.

For high-resolution workloads, the relevant reasons to consider this SKU are the 32 GB HBM2 pool and the 898.0 GB/s bandwidth. Those two numbers allow large resident allocations and fast local transfers. The 640 tensor cores and 31.33 TFLOPS FP16 rate make tensor-heavy workloads a natural match. The 15.67 TFLOPS FP32 rate is also substantial for general compute. The API set, including DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4, gives software compatibility across several compute and rendering paths. Without benchmark scores, the data cannot support a settings-based or frame-rate-based recommendation. The decision should instead be driven by memory capacity, memory bandwidth, tensor throughput, and the absence of display outputs. If the workload is a high-resolution dataset that fits in 32 GB and requires high-bandwidth access, this module's specification sheet is aligned with that use. If the workload expects a standard graphics card with display outputs and on-board power connectors, the data says this is not the right device.

How It Compares

The nearestRivals array is empty in the database, so there are no rival names, scores, or deltaPct values to report. That leaves this entry without a direct head-to-head comparison. The generation context is still present: Tesla Pascal is the predecessor, Tesla Turing is the successor, and this module belongs to the Tesla Volta (Vxx) line. The production status is End-of-life, and the release date is 2018-03-26. The only database rank is percentileVsAllGpus at 50, which is a midpoint placement against the full GPU list rather than a comparative result against a specific rival. Because no nearest rivals are supplied, no relative performance conclusion can be formed from the database's own comparison fields. The available positional statements are about generation, status, and aggregate percentile, not about performance relative to named competitors.

Detailed benchmark scores and charts for the NVIDIA Tesla V100 SXM2 32 GB are below.

Benchmark Scores

geekbench_openclSource

Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA Tesla V100 SXM2 32 GB handles parallel computing tasks like video encoding and scientific simulations. OpenCL is widely supported across different GPU vendors and platforms.

geekbench_opencl #62 of 650
131,696
34%
Max: 388,405
Compare with other GPUs

Top 5 Performers

#1 NVIDIA RTX 6000D
388,405
#2 NVIDIA B300 SXM6 AC
369,831
#3 NVIDIA B200
345,482
#4 NVIDIA H200 NVL
334,891

geekbench_vulkanSource

Geekbench Vulkan tests GPU compute using the modern low-overhead Vulkan API. This shows how NVIDIA Tesla V100 SXM2 32 GB performs with next-generation graphics and compute workloads. Vulkan offers better CPU efficiency than older APIs like OpenGL. Modern games and applications increasingly use Vulkan for cross-platform GPU acceleration.

geekbench_vulkan #41 of 446
143,765
38%
Max: 376,915

Popular NVIDIA Tesla V100 SXM2 32 GB Comparisons

See how the Tesla V100 SXM2 32 GB stacks up against similar graphics cards from the same generation and competing brands.

Compare with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs