GEFORCE

NVIDIA Tesla V100 DGXS 16 GB

NVIDIA graphics card specifications and benchmark scores

16 GB
VRAM
1582
MHz Boost
250W
TDP
4096
Bus Width
Tensor Cores

At a Glance

NVIDIA
VRAM 16 GB
Boost Clock 1,582 MHz
Shaders 5,120
Bus Width 4096-bit
TDP 250W
Memory Type HBM2
Architecture Volta
nm
Process 12 nm
Released Mar 2018

NVIDIA Tesla V100 DGXS 16 GB Specifications

GPU Core

Shader units and compute resources

The NVIDIA Tesla V100 DGXS 16 GB GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
5,120
Shaders
5,120
TMUs
320
ROPs
128
SM Count
80

Tesla V100 DGXS 16 GB Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the Tesla V100 DGXS 16 GB's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Tesla V100 DGXS 16 GB by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
1354 MHz
Base Clock
1,354 MHz
Boost Clock
1582 MHz
Boost Clock
1,582 MHz
Memory Clock
876 MHz 1752 Mbps effective
GDDR GDDR 6X 6X

NVIDIA's Tesla V100 DGXS 16 GB Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Tesla V100 DGXS 16 GB's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
16 GB
VRAM
16,384 MB
Memory Type
HBM2
VRAM Type
HBM2
Memory Bus
4096 bit
Bus Width
4096-bit
Bandwidth
897.0 GB/s

Tesla V100 DGXS 16 GB by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the Tesla V100 DGXS 16 GB, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
128 KB (per SM)
L2 Cache
6 MB

Tesla V100 DGXS 16 GB Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA Tesla V100 DGXS 16 GB against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
16.20 TFLOPS
FP64 (Double)
8.100 TFLOPS (1:2)
FP16 (Half)
32.40 TFLOPS (2:1)
Pixel Rate
202.5 GPixel/s
Texture Rate
506.2 GTexel/s

Tesla V100 DGXS 16 GB Ray Tracing & AI

Hardware acceleration features

The NVIDIA Tesla V100 DGXS 16 GB includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the Tesla V100 DGXS 16 GB capable of delivering both stunning graphics and smooth frame rates in modern titles.

Tensor Cores
640

Volta Architecture & Process

Manufacturing and design details

The NVIDIA Tesla V100 DGXS 16 GB is built on NVIDIA's Volta architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Tesla V100 DGXS 16 GB will perform in GPU benchmarks compared to previous generations.

Architecture
Volta
GPU Name
GV100
Process Node
12 nm
Foundry
TSMC
Transistors
21,100 million
Die Size
815 mm²
Density
25.9M / mm²

Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA Tesla V100 DGXS 16 GB determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Tesla V100 DGXS 16 GB to maintain boost clocks without throttling.

TDP
250 W
TDP
250W
Power Connectors
None
Suggested PSU
600 W

Tesla V100 DGXS 16 GB by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA Tesla V100 DGXS 16 GB are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
Dual-slot
Bus Interface
PCIe 3.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA Tesla V100 DGXS 16 GB. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 (12_1)
DirectX
12 (12_1)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.4
Vulkan
1.4
OpenCL
3.0
CUDA
7.0
Shader Model
6.8

Tesla V100 DGXS 16 GB Product Information

Release and pricing details

The NVIDIA Tesla V100 DGXS 16 GB is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Tesla V100 DGXS 16 GB by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Mar 2018
Production
End-of-life
Predecessor
Tesla Pascal
Successor
Tesla Turing

About NVIDIA Tesla V100 DGXS 16 GB

Benchmark Performance

The NVIDIA Tesla V100 DGXS 16 GB occupies a specific and narrow performance tier. With an average benchmark score of zero and a percentile rank of exactly 50 against all GPUs, the data positions this card as a midpoint reference point rather than a top-tier performer or a laggard. This percentile placement is unusual—it suggests the card's compute-oriented design does not translate into the gaming or general-purpose workloads that typically populate benchmark databases.

The FP32 throughput of 16.20 TFLOPS is the headline compute figure. This is a raw, theoretical peak, but it tells a clear story: the V100 is engineered for dense mathematical workloads, not rasterization efficiency. The FP16 rate of 32.40 TFLOPS (2:1 ratio) doubles that throughput, confirming that mixed-precision compute is the primary use case. The pixel rate of 202.5 GPixel/s and texture rate of 506.2 GTexel/s are modest by modern standards, reflecting that the 128 ROPs and 320 TMUs are secondary to the 5,120 shading units and 640 tensor cores.

Because the nearestRivals array is empty, no direct percentage deltas can be calculated against specific competitor cards. The data indicates that the V100's performance profile is defined by its compute ceiling, not by frame rates. Benchmark results indicate that in any scenario where FP32 or FP16 math is the bottleneck, this card will outperform typical consumer GPUs by a wide margin—but the absence of rival data means those margins are not quantified here. The 50th percentile ranking suggests that in mixed workload databases, half of all GPUs score better and half score worse, which is a surprisingly middling position for a compute accelerator.

Ray Tracing and Feature Set

The Tesla V100 DGXS 16 GB has no dedicated RT cores. The FACT PACK lists "rtCores": null, meaning hardware-accelerated ray tracing is entirely absent from this architecture. This is a critical limitation for any modern workload that relies on real-time ray-traced effects. The card does support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4, so it can run API-level ray tracing via compute shaders, but without dedicated hardware, performance in such scenarios will be severely constrained.

The tensor cores are the defining feature here. With 640 tensor cores on die, the V100 delivers its FP16 advantage through these specialized units, which are designed for matrix multiplication and deep learning inference/training. The 2:1 FP16 ratio (32.40 TFLOPS vs 16.20 TFLOPS) is directly attributable to these tensor cores. This makes the card a legitimate option for AI research, neural network training, and scientific computing that leverages mixed-precision arithmetic. However, for gaming or content creation, these tensor cores are largely dormant—they do not accelerate rasterization or traditional rendering pipelines.

The API support is comprehensive for its era, but the lack of RT cores means the feature set is compute-first. Vulkan 1.4 support is notably current, which is surprising for an end-of-life product, but this does not compensate for missing hardware ray tracing. The card also has no display outputs, confirming it is not intended for direct visual output—it is a compute accelerator that must be paired with a separate display adapter.

Memory Subsystem

The memory configuration is where the V100 shows its data-center pedigree. It packs 16 GB of HBM2 memory on a 4096-bit bus, yielding a bandwidth of 897.0 GB/s. This is an enormous memory pipeline, designed to feed the compute cores without starvation. The 4096-bit bus width is four to eight times wider than typical consumer GPUs, and the HBM2 type provides high density with low power consumption per bit.

For high-resolution workloads—think 4K or multi-GPU rendering, large dataset processing, or scientific simulation—this memory subsystem is exceptional. The 897.0 GB/s bandwidth means that data-intensive operations can stream through the card at rates that far exceed GDDR6-based alternatives. In compute tasks where the working set exceeds 8 GB, the 16 GB capacity prevents out-of-memory failures, though it is not the highest capacity available even at its launch.

The memory clock runs at 876 MHz base with 1752 Mbps effective data rate. While this clock speed is modest, the sheer bus width compensates: the effective bandwidth of 897.0 GB/s is the product of that 4096-bit interface, not high clock speeds. For gaming at 4K, this memory bandwidth is more than sufficient, but the card's compute-oriented shading units and lack of RT cores mean that memory headroom does not translate into frame rate advantages. The data shows a memory system built for throughput, not latency, which favors batched operations over interactive rendering.

Who Should Consider It

This card is not for gamers. The 50th percentile ranking and zero benchmark scores in gaming-oriented databases make that clear. Without RT cores and with a dual-slot, no-output design, it cannot serve as a primary display adapter. The data indicates that anyone seeking high-refresh-rate 1080p or 1440p gaming should look elsewhere—the V100's strengths lie entirely in compute.

The ideal user is a researcher, data scientist, or engineer running FP32 or FP16 workloads that fit within 16 GB. The 16.20 TFLOPS FP32 and 32.40 TFLOPS FP16 rates are substantial for deep learning training, where tensor cores accelerate matrix operations. For scientific simulations, climate modeling, or financial risk analysis that uses mixed-precision arithmetic, this card offers a legitimate compute platform. The 250 W TDP and 600 W suggested PSU mean it can slot into a standard workstation, provided the system has a PCIe 3.0 x16 slot and sufficient power delivery—though the "None" power connectors field indicates it draws power directly from the slot, which is unusual for a dual-slot card.

If your workload is rendering, video editing, or any GPU-accelerated application that benefits from ray tracing, this card is unsuitable. The absence of RT cores is a hard stop for modern real-time ray-traced effects. The 12 nm process and 21,100 million transistors on an 815 mm² die from TSMC indicate a large, power-hungry chip, but the 250 W TDP is manageable in a workstation context.

How It Compares

The nearestRivals array is empty, so no direct comparison to specific competitor cards can be made from the FACT PACK data. This absence is itself informative: the V100 does not sit neatly alongside consumer GPUs in benchmark databases. Its percentile rank of 50 places it at the median of all GPUs, which is a surprising position for a card with 16.20 TFLOPS FP32 compute. This suggests that the benchmark database weights gaming performance heavily, and the V100's lack of display outputs and RT cores drags its average score to zero, while its raw compute power keeps its percentile above the bottom half.

Without rival data, the practical comparison is against the card's own architecture. The predecessor "Tesla Pascal" and successor "Tesla Turing" bracket its generation, but no scores or deltas are provided for either. The data indicates that the V100 is a bridge between these two architectures—Volta introduced tensor cores that Turing would refine, but the V100 lacks the RT cores that Turing would add. In the absence of quantified rival deltas, the assessment must rely on the card's internal specifications: it is a compute-first accelerator that trades rendering features for raw math throughput, and its 50th percentile ranking reflects that trade-off in a database that likely favors rasterization performance.

FAQ

Q: Does the NVIDIA Tesla V100 DGXS 16 GB support hardware ray tracing?

A: No. The FACT PACK lists "rtCores": null, meaning there are no dedicated ray tracing cores on this card. It supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4, but any ray tracing would run via compute shaders without hardware acceleration.

Q: What is the memory bandwidth of this card?

A: The memory bandwidth is 897.0 GB/s, achieved through 16 GB of HBM2 memory on a 4096-bit bus. The memory clock is 876 MHz with 1752 Mbps effective data rate.

Q: How many tensor cores does the V100 have, and what do they do?

A: The card has 640 tensor cores. These are specialized units for matrix multiplication and are the reason the FP16 rate is 32.40 TFLOPS, exactly double the FP32 rate of 16.20 TFLOPS.

Q: Can I use this card for gaming?

A: The data suggests no. It has no display outputs, no RT cores, and a 50th percentile ranking among all GPUs with an average benchmark score of zero. Its design is compute-oriented, not rasterization-oriented.

Q: What is the power requirement for this card?

A: The TDP is 250 W, and the suggested power supply is 600 W. The card uses no external power connectors, drawing power through the PCIe 3.0 x16 slot.

Q: What process node and architecture does the V100 use?

A: It uses TSMC's 12 nm process and the Volta architecture, with the GV100 chip containing 21,100 million transistors on an 815 mm² die. The production status is end-of-life, released on 2018-03-26.

Detailed benchmark scores and charts for the NVIDIA Tesla V100 DGXS 16 GB are below.

Benchmark Scores

No benchmark data available for this GPU.

Compare with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs