GEFORCE

NVIDIA Tesla V100 PCIe 16 GB

NVIDIA graphics card specifications and benchmark scores

16 GB
VRAM
1380
MHz Boost
300W
TDP
4096
Bus Width
Tensor Cores

At a Glance

NVIDIA
VRAM 16 GB
Boost Clock 1,380 MHz
Shaders 5,120
Bus Width 4096-bit
TDP 300W
Memory Type HBM2
Architecture Volta
nm
Process 12 nm
Released Jun 2017

NVIDIA Tesla V100 PCIe 16 GB Specifications

Tesla V100 PCIe 16 GB GPU Core

Shader units and compute resources

The NVIDIA Tesla V100 PCIe 16 GB GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
5,120
Shaders
5,120
TMUs
320
ROPs
128
SM Count
80

Tesla V100 PCIe 16 GB Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the Tesla V100 PCIe 16 GB's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Tesla V100 PCIe 16 GB by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
1245 MHz
Base Clock
1,245 MHz
Boost Clock
1380 MHz
Boost Clock
1,380 MHz
Memory Clock
876 MHz 1752 Mbps effective
GDDR GDDR 6X 6X

NVIDIA's Tesla V100 PCIe 16 GB Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Tesla V100 PCIe 16 GB's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
16 GB
VRAM
16,384 MB
Memory Type
HBM2
VRAM Type
HBM2
Memory Bus
4096 bit
Bus Width
4096-bit
Bandwidth
897.0 GB/s

Tesla V100 PCIe 16 GB by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the Tesla V100 PCIe 16 GB, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
128 KB (per SM)
L2 Cache
6 MB

Tesla V100 PCIe 16 GB Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA Tesla V100 PCIe 16 GB against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
14.13 TFLOPS
FP64 (Double)
7.066 TFLOPS (1:2)
FP16 (Half)
28.26 TFLOPS (2:1)
Pixel Rate
176.6 GPixel/s
Texture Rate
441.6 GTexel/s

Tesla V100 PCIe 16 GB Ray Tracing & AI

Hardware acceleration features

The NVIDIA Tesla V100 PCIe 16 GB includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the Tesla V100 PCIe 16 GB capable of delivering both stunning graphics and smooth frame rates in modern titles.

Tensor Cores
640

Volta Architecture & Process

Manufacturing and design details

The NVIDIA Tesla V100 PCIe 16 GB is built on NVIDIA's Volta architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Tesla V100 PCIe 16 GB will perform in GPU benchmarks compared to previous generations.

Architecture
Volta
GPU Name
GV100
Process Node
12 nm
Foundry
TSMC
Transistors
21,100 million
Die Size
815 mm²
Density
25.9M / mm²

NVIDIA's Tesla V100 PCIe 16 GB Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA Tesla V100 PCIe 16 GB determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Tesla V100 PCIe 16 GB to maintain boost clocks without throttling.

TDP
300 W
TDP
300W
Power Connectors
2x 8-pin
Suggested PSU
700 W

Tesla V100 PCIe 16 GB by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA Tesla V100 PCIe 16 GB are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
Dual-slot
Bus Interface
PCIe 3.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA Tesla V100 PCIe 16 GB. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 (12_1)
DirectX
12 (12_1)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.4
Vulkan
1.4
OpenCL
3.0
CUDA
7.0
Shader Model
6.8

Tesla V100 PCIe 16 GB Product Information

Release and pricing details

The NVIDIA Tesla V100 PCIe 16 GB is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Tesla V100 PCIe 16 GB by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Jun 2017
Production
End-of-life
Predecessor
Tesla Pascal
Successor
Tesla Turing

Tesla V100 PCIe 16 GB Benchmark Scores

geekbench_openclSource

Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA Tesla V100 PCIe 16 GB handles parallel computing tasks like video encoding and scientific simulations. OpenCL is widely supported across different GPU vendors and platforms. Higher scores benefit applications that leverage GPU acceleration for non-graphics workloads.

geekbench_opencl #40 of 643
163,063
42%
Max: 388,405

geekbench_vulkanSource

Geekbench Vulkan tests GPU compute using the modern low-overhead Vulkan API. This shows how NVIDIA Tesla V100 PCIe 16 GB performs with next-generation graphics and compute workloads.

geekbench_vulkan #65 of 444
113,062
30%
Max: 376,915

About NVIDIA Tesla V100 PCIe 16 GB

# NVIDIA Tesla V100 PCIe 16 GB

The NVIDIA Tesla V100 PCIe 16 GB is a data-center oriented accelerator built on the Volta architecture, fabricated on TSMC's 12 nm process. It packs 21,100 million transistors onto an 815 mm² die, yielding a transistor density of 25.9 million transistors per mm². The card operates with a base clock of 1245 MHz and a boost clock of 1380 MHz, delivering 14.13 TFLOPS of FP32 compute and 28.26 TFLOPS of FP16 compute (at a 2:1 ratio). With a 300 W TDP, dual-slot cooler, and dual 8-pin power connectors, it is positioned for server racks rather than consumer desktops. Notably, its percentile rank against all GPUs is 50, indicating median performance in the broader database, though its benchmark score is listed as zero, meaning no direct performance samples are available for this entry.

Benchmark Performance

The data for the Tesla V100 PCIe 16 GB shows no direct benchmark scores — the `benchmarks` array is empty and the average benchmark score is zero. This is typical for compute-oriented accelerators that are rarely subjected to gaming or synthetic graphics workloads. However, the raw compute specifications provide a clear picture of its capability. The FP32 throughput of 14.13 TFLOPS places it in a range that would be competitive with high-end consumer GPUs of its generation, though the percentile rank of 50 suggests that in the full database of all GPUs — which includes both gaming and professional parts — it sits exactly at the median. This is an unusual position for a card with such a large die and high transistor count, but it underscores that the V100 is optimized for specific workloads, not general-purpose graphics.

The FP16 performance of 28.26 TFLOPS is exactly double the FP32 figure, reflecting the 2:1 ratio enabled by the Volta architecture's tensor cores. This ratio is critical for deep learning training and inference, where reduced precision is acceptable. The pixel rate of 176.6 GPixel/s and texture rate of 441.6 GTexel/s, derived from 128 ROPs and 320 TMUs respectively, indicate that the card is not starved for rasterization resources, but these figures pale next to its compute capabilities. In the absence of rival data — the `nearestRivals` array is empty — no direct percentage deltas can be cited. The analysis must therefore rely on the card's own specifications and its percentile placement to infer relative standing.

How It Compares

Given that the `nearestRivals` list is empty, there are no direct competitor comparisons available from the FACT PACK. The percentile rank of 50, however, offers a baseline: half of all GPUs in the database score higher, and half score lower. This is a surprisingly middling result for a card of this caliber, but it likely reflects the fact that the database includes many consumer gaming GPUs that excel in rasterization benchmarks, while the V100's strengths lie in compute tasks that are not captured by those metrics. The predecessor is listed as Tesla Pascal, and the successor as Tesla Turing, but no specific models or scores are provided for either.

Without rival names, scores, or deltaPct values, the comparison section can only note the absence of data. The card's position in the product stack is clear: it belongs to the Volta generation, sits between Pascal and Turing in the Tesla line, and targets HPC and AI workloads. Its 5120 shading units and 640 tensor cores are the key differentiators, but without competitor figures, relative performance cannot be quantified. The lack of display outputs reinforces that this is not a card for interactive use; it is a compute accelerator designed for servers and workstations where headless operation is standard.

Who Should Consider It

The Tesla V100 PCIe 16 GB is not suited for typical gaming or consumer desktop use — it has no display outputs, and its drivers and firmware are tuned for compute workloads. For users running deep learning frameworks, scientific simulations, or data processing pipelines that leverage FP16 or FP32 compute, the card's specifications are compelling on paper. The 16 GB of HBM2 memory with 897.0 GB/s of bandwidth is particularly relevant for large models or datasets that exceed the VRAM capacity of consumer cards. At 1080p or 1440p gaming, the card's FP32 performance of 14.13 TFLOPS would theoretically handle high settings, but the absence of display outputs makes this moot.

For resolution-specific recommendations grounded in the available data: at 4K, the high memory bandwidth of 897.0 GB/s would be advantageous for texture-heavy workloads, but again, without benchmark scores, any claim about playable frame rates is unsupported. The card's 4096-bit memory bus is exceptionally wide, which helps maintain throughput in memory-bound tasks. Users who need a headless accelerator for server rooms, with a 300 W TDP and 700 W suggested PSU, will find the V100's specifications align with that use case. Cloud service providers and research institutions are the primary audience, not individual consumers.

FAQ

Q: What is the FP32 performance of the Tesla V100 PCIe 16 GB?

A: The card delivers 14.13 TFLOPS of FP32 compute, based on a base clock of 1245 MHz and a boost clock of 1380 MHz across 5120 shading units.

Q: Does this card have tensor cores?

A: Yes, it includes 640 tensor cores, which enable FP16 compute at 28.26 TFLOPS, exactly double the FP32 rate (2:1 ratio).

Q: What type of memory does it use and how much bandwidth does it offer?

A: It uses 16 GB of HBM2 memory on a 4096-bit bus, providing 897.0 GB/s of memory bandwidth.

Q: Can this card output video to a display?

A: No, the card has no display outputs; it is designed for headless compute workloads in servers or workstations.

Q: What is the power requirement for this card?

A: It has a 300 W TDP, requires two 8-pin power connectors, and the suggested PSU is 700 W.

Q: What API support is available?

A: The card supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4, though its primary use case is compute rather than graphics.

Ray Tracing and Feature Set

The Tesla V100 PCIe 16 GB does not include dedicated ray tracing cores — the `rtCores` field is null. This distinguishes it from later Turing-generation cards that introduced hardware RT acceleration. Instead, the card's feature set is centered on its 640 tensor cores, which are designed for matrix math used in AI and deep learning. The Volta architecture introduced these tensor cores, and their presence at this density (640 units) is the defining hardware feature. The card does not support real-time ray tracing via dedicated hardware, but it can theoretically handle compute-based ray tracing through general-purpose shaders, though without RT cores, performance would be limited.

The API support includes DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4, which are relevant for compute and graphics workloads. The lack of display outputs means that any graphics API usage would be offscreen rendering or compute shaders. The FP16 capability at 2:1 ratio is a major feature for mixed-precision training, and the 12 nm process node from TSMC, with 21,100 million transistors, indicates a mature but not cutting-edge manufacturing process. The card is end-of-life, with a release date of 2017-06-20, and it is part of the Tesla Volta generation, with Tesla Pascal as predecessor and Tesla Turing as successor.

Memory Subsystem

The memory subsystem is one of the V100's strongest attributes. It pairs 16 GB of HBM2 memory with a 4096-bit bus, yielding a bandwidth of 897.0 GB/s. This is significantly higher than typical GDDR5 or GDDR6 configurations of the same era, and the wide bus allows for efficient data movement in memory-bound workloads. For high-resolution compute tasks — such as training large neural networks or processing 4K video frames — the bandwidth ensures that the 5120 shading units and 640 tensor cores are not starved for data. The memory clock runs at 876 MHz, with an effective data rate of 1752 Mbps.

The 16 GB capacity is substantial for 2017, allowing models or datasets that would exceed the 8 GB or 11 GB capacities of contemporary consumer cards. At 4K resolution, the bandwidth is more than sufficient for texture streaming, though the card's lack of display outputs makes this a moot point for gaming. For scientific computing, the memory subsystem's bandwidth is a bottleneck reliever: the 897.0 GB/s figure means that large matrix operations can be fed quickly to the compute units. The combination of 4096-bit bus and HBM2 type is a clear differentiator from cards using narrower GDDR6 buses, and it explains the 300 W TDP — high-bandwidth memory requires power. The 700 W suggested PSU accounts for the card's draw plus system overhead, and the dual 8-pin connectors supply the necessary current.

The AMD Equivalent of Tesla V100 PCIe 16 GB

Looking for a similar graphics card from AMD? The AMD Radeon RX 550 Mobile offers comparable performance and features in the AMD lineup.

AMD Radeon RX 550 Mobile

AMD • 2 GB VRAM

View Specs Compare

Popular NVIDIA Tesla V100 PCIe 16 GB Comparisons

See how the Tesla V100 PCIe 16 GB stacks up against similar graphics cards from the same generation and competing brands.

Compare Tesla V100 PCIe 16 GB with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs