GEFORCE

NVIDIA Tesla V100S PCIe 32 GB

NVIDIA graphics card specifications and benchmark scores

32 GB
VRAM
1597
MHz Boost
250W
TDP
4096
Bus Width
Tensor Cores

At a Glance

NVIDIA
VRAM 32 GB
Boost Clock 1,597 MHz
Shaders 5,120
Bus Width 4096-bit
TDP 250W
Memory Type HBM2
Architecture Volta
nm
Process 12 nm
Released Nov 2019

NVIDIA Tesla V100S PCIe 32 GB Specifications

GPU Core

Shader units and compute resources

The NVIDIA Tesla V100S PCIe 32 GB GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
5,120
Shaders
5,120
TMUs
320
ROPs
128
SM Count
80

Tesla V100S PCIe 32 GB Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the Tesla V100S PCIe 32 GB's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Tesla V100S PCIe 32 GB by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
1245 MHz
Base Clock
1,245 MHz
Boost Clock
1597 MHz
Boost Clock
1,597 MHz
Memory Clock
1107 MHz 2.2 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's Tesla V100S PCIe 32 GB Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Tesla V100S PCIe 32 GB's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
32 GB
VRAM
32,768 MB
Memory Type
HBM2
VRAM Type
HBM2
Memory Bus
4096 bit
Bus Width
4096-bit
Bandwidth
1.13 TB/s

Tesla V100S PCIe 32 GB by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the Tesla V100S PCIe 32 GB, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
128 KB (per SM)
L2 Cache
6 MB

Tesla V100S PCIe 32 GB Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA Tesla V100S PCIe 32 GB against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
16.35 TFLOPS
FP64 (Double)
8.177 TFLOPS (1:2)
FP16 (Half)
32.71 TFLOPS (2:1)
Pixel Rate
204.4 GPixel/s
Texture Rate
511.0 GTexel/s

Tesla V100S PCIe 32 GB Ray Tracing & AI

Hardware acceleration features

The NVIDIA Tesla V100S PCIe 32 GB includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the Tesla V100S PCIe 32 GB capable of delivering both stunning graphics and smooth frame rates in modern titles.

Tensor Cores
640

Volta Architecture & Process

Manufacturing and design details

The NVIDIA Tesla V100S PCIe 32 GB is built on NVIDIA's Volta architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Tesla V100S PCIe 32 GB will perform in GPU benchmarks compared to previous generations.

Architecture
Volta
GPU Name
GV100
Process Node
12 nm
Foundry
TSMC
Transistors
21,100 million
Die Size
815 mm²
Density
25.9M / mm²

Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA Tesla V100S PCIe 32 GB determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Tesla V100S PCIe 32 GB to maintain boost clocks without throttling.

TDP
250 W
TDP
250W
Power Connectors
2x 8-pin
Suggested PSU
600 W

Tesla V100S PCIe 32 GB by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA Tesla V100S PCIe 32 GB are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
Dual-slot
Bus Interface
PCIe 3.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA Tesla V100S PCIe 32 GB. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 (12_1)
DirectX
12 (12_1)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.4
Vulkan
1.4
OpenCL
3.0
CUDA
7.0
Shader Model
6.8

Tesla V100S PCIe 32 GB Product Information

Release and pricing details

The NVIDIA Tesla V100S PCIe 32 GB is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Tesla V100S PCIe 32 GB by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Nov 2019
Production
End-of-life
Predecessor
Tesla Pascal
Successor
Tesla Turing

About NVIDIA Tesla V100S PCIe 32 GB

The NVIDIA Tesla V100S PCIe 32 GB is a data-center compute accelerator built on the Volta architecture, fabricated on TSMC's 12 nm process with 21,100 million transistors across an 815 mm² die. It pairs 5,120 shading units with 640 tensor cores and 32 GB of HBM2 memory on a 4096-bit bus, delivering 1.13 TB/s of bandwidth. The card targets high-throughput FP32 and FP16 workloads, with a base clock of 1245 MHz and a boost clock of 1597 MHz. It is now end-of-life, having launched in late 2019, and the database places it at the 50th percentile of all tracked GPUs, indicating median performance within that historical context.

Benchmark Performance

The database does not list any benchmark scores for the Tesla V100S PCIe 32 GB, so performance analysis relies on its theoretical compute metrics and its percentile ranking. The FP32 throughput is 16.35 TFLOPS, while FP16 reaches 32.71 TFLOPS via a 2:1 ratio. These figures position the card as a capable compute device, though the 50th percentile ranking suggests that, among all GPUs in the database, it sits exactly in the middle of the performance distribution. This is notable for a card that was designed for server workloads rather than consumer gaming, and the median placement reflects the fact that many modern consumer GPUs now exceed its raw FP32 output.

The pixel rate of 204.4 GPixel/s and texture rate of 511.0 GTexel/s are derived from the 128 ROPs and 320 TMUs, respectively. These rates are competitive for a 2019 accelerator, but they do not directly translate to gaming performance since the card has no display outputs. The FP16 capability is double the FP32 rate, which is typical for Volta and advantageous for mixed-precision training. The tensor cores, of which there are 640, provide dedicated matrix math acceleration for AI inference and training, but the database does not provide a separate tensor performance score.

Given the absence of benchmark entries, the percentile figure is the only comparative metric. The 50th percentile implies that half of the GPUs in the database are faster and half are slower, based on aggregated scores. For a compute-oriented card with a high memory capacity, this median standing is somewhat surprising, but it likely reflects the card's age and the rapid advancement of both consumer and professional GPUs since its release. The lack of ray tracing cores further differentiates it from modern architectures.

Ray Tracing and Feature Set

The Tesla V100S PCIe 32 GB has no dedicated ray tracing cores; the field for RT cores is null. This is a deliberate design choice for a compute card, as ray tracing hardware is primarily relevant for real-time graphics, which this card does not handle due to the absence of display outputs. Instead, the card's feature set centers on tensor cores and general compute. The 640 tensor cores enable accelerated deep learning operations, and the card supports the Volta architecture's Tensor Core instructions for mixed-precision workloads.

In terms of API support, the card lists DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. These are graphics APIs, but without display outputs, they are largely irrelevant for end-user rendering. The Vulkan 1.4 support is notable as a recent API version, but again, it would only be used in headless compute contexts. The card also lacks any display outputs, so it cannot be used for video output or direct rendering. Its bus interface is PCIe 3.0 x16, which is standard for the era.

The power delivery requires two 8-pin connectors and a 600 W suggested PSU, with a TDP of 250 W. The dual-slot design accommodates the HBM2 stacks and cooling solution. There is no mention of hardware ray tracing or DLSS-like features, as those are Turing-era innovations. The card's feature set is squarely aimed at high-performance computing, scientific simulation, and AI training, where tensor cores and memory bandwidth matter more than graphics features.

Memory Subsystem

The memory subsystem is one of the strongest aspects of the Tesla V100S PCIe 32 GB. It features 32 GB of HBM2 memory on a 4096-bit bus, yielding a bandwidth of 1.13 TB/s. This is a massive amount of memory and bandwidth, especially for its time. The 32 GB capacity allows large models and datasets to reside in GPU memory without spilling to system RAM, which is critical for deep learning training and inference on large neural networks. The 4096-bit bus is among the widest ever implemented, and the 1.13 TB/s bandwidth enables rapid data movement for compute kernels.

The memory clock is listed as 1107 MHz, with an effective data rate of 2.2 Gbps. This is relatively modest compared to later GDDR6X implementations, but the sheer bus width compensates, resulting in the high aggregate bandwidth. For high-resolution scientific simulations or AI workloads that process multi-gigabyte matrices, this memory subsystem reduces bottlenecks. The 32 GB capacity also supports larger batch sizes in training, which can improve throughput.

The card's memory type, HBM2, is known for its energy efficiency and compact footprint, though it adds manufacturing complexity. The 1.13 TB/s figure is a raw specification, and real-world effective bandwidth may vary depending on access patterns. Still, for a 2019 card, this memory configuration was top-tier, and it remains competitive even by today's standards. The combination of 32 GB and 1.13 TB/s makes the card suitable for workloads that require both high capacity and high throughput, such as computational fluid dynamics, genomic analysis, and large-scale matrix operations.

How It Compares

The nearestRivals field is empty, meaning the database does not provide any direct competitor comparisons for this card. Consequently, a rival-by-rival breakdown is not possible. The only comparative data point is the 50th percentile ranking across all GPUs. This suggests that, in aggregate, the Tesla V100S PCIe 32 GB performs at the median of the database's entire GPU population. That median status is interesting because the card is a professional accelerator, not a gaming GPU, yet it still lands in the middle when all GPUs are considered.

Without rival entries, the analysis must rely on the card's own specification sheet. Its predecessor is listed as Tesla Pascal, and its successor as Tesla Turing. These are generational markers, not performance comparisons. The card bridges the Pascal and Turing eras, incorporating Volta's tensor cores while lacking the ray tracing hardware that Turing introduced. The empty rival list may indicate that the database does not have enough comparable entries for this specific SKU, or that its unique positioning (32 GB HBM2, no display outputs) does not align with typical consumer or workstation cards.

The 50th percentile is a useful anchor. It implies that the card is neither a top performer nor a low-end part in the historical context of the database. For compute workloads, the card's strengths are its memory capacity and bandwidth, which are not fully captured by generic benchmark scores. The lack of ray tracing cores and display outputs further limits its appeal for gaming or professional visualization, but those are not the intended use cases.

Who Should Consider It

The Tesla V100S PCIe 32 GB is designed for compute-intensive applications that require large memory and high bandwidth. Its 32 GB HBM2 and 1.13 TB/s bandwidth make it well-suited for deep learning training, where model sizes often exceed 16 GB, and for inference tasks that need to hold multiple models or large batches. The 640 tensor cores accelerate matrix operations, and the FP16 throughput of 32.71 TFLOPS is double the FP32 rate, which is beneficial for mixed-precision training. Researchers and data scientists working on natural language processing, computer vision, or scientific simulations could benefit from this card's capacity and throughput.

The card has no display outputs, so it is not for desktop use or gaming. It is a headless accelerator meant for servers and workstations that handle compute remotely. Its dual-slot form factor and 250 W TDP require a standard server chassis with adequate power. The 600 W suggested PSU is a guideline for a system with this card, but actual requirements depend on other components.

Given its end-of-life status, the card is now a legacy product. Buyers looking for current support or the latest features might consider newer accelerators, but the database does not provide comparative data for those. The 50th percentile ranking suggests that its performance is not exceptional by today's standards, but for specific workloads that leverage its memory size and bandwidth, it can still be a viable option, especially in secondary markets. The lack of ray tracing cores means it is unsuitable for real-time ray-traced rendering, but that is not a concern for compute-only deployments.

In summary, the Tesla V100S PCIe 32 GB is a specialized compute card that excels in memory-heavy AI and scientific workloads. Its 32 GB capacity and 1.13 TB/s bandwidth are its defining features, while its tensor cores provide acceleration for deep learning. The card sits at the median of the database's GPU population, reflecting its age and the rapid evolution of hardware. It is not a general-purpose GPU, and its target audience is organizations that need a high-memory, high-bandwidth compute accelerator for batch processing and model training.

Detailed benchmark scores and charts for the NVIDIA Tesla V100S PCIe 32 GB are below.

Benchmark Scores

geekbench_openclSource

Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA Tesla V100S PCIe 32 GB handles parallel computing tasks like video encoding and scientific simulations. OpenCL is widely supported across different GPU vendors and platforms. Higher scores benefit applications that leverage GPU acceleration for non-graphics workloads.

geekbench_opencl #27 of 650
194,415
50%
Max: 388,405

Popular NVIDIA Tesla V100S PCIe 32 GB Comparisons

See how the Tesla V100S PCIe 32 GB stacks up against similar graphics cards from the same generation and competing brands.

Compare with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs