NVIDIA L4 vs NVIDIA Tesla V100 PCIe 16 GB Comparison

NVIDIA
GEFORCE

NVIDIA L4

CORE STATE AD104
VRAM 24 GB
CLOCK SPEED 2040 MHz
TDP 72 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

Tesla V100 PCIe 16 GB

CORE STATE GV100
VRAM 16 GB
CLOCK SPEED 1380 MHz
TDP 300 W
BUS WIDTH 4096 bit
ARCHITECTURE Volta
nm
PROCESS 12 nm
LAUNCH DATE 2017

PERFORMANCE BENCHMARKS

geekbench_opencl
140,838
163,063
geekbench_vulkan
121,306
113,062

Analysis: NVIDIA L4 vs NVIDIA Tesla V100 PCIe 16 GB

# Head-to-Head Benchmarks

The benchmark data presents a clear split between these two NVIDIA accelerators, with each claiming a decisive victory in one of the two recorded tests. In Geekbench OpenCL, the NVIDIA Tesla V100 PCIe 16 GB posts a score of 163,063, which is 15.8% higher than the NVIDIA L4's 140,838. This is a substantial margin that places the V100 firmly ahead in compute-heavy workloads that favor raw parallel throughput. The V100's OpenCL result also exceeds its own average benchmark score of 138,063 by a significant amount, indicating that this particular test plays to its strengths.

Conversely, the NVIDIA L4 takes the Geekbench Vulkan test with a score of 121,306, outperforming the V100's 113,062 by 6.8%. While this margin is narrower than the V100's OpenCL advantage, it demonstrates the L4's superior graphics API performance, likely benefiting from its newer architecture and higher boost clocks. The L4's Vulkan score also represents a notable improvement over its OpenCL result, whereas the V100 shows the opposite trend—its OpenCL score is vastly superior to its Vulkan score, with a gap of 50,001 points between the two tests.

Looking at the broader performance picture, the V100 achieves an average benchmark score of 138,063 across all tests, placing it at the 96th percentile of all GPUs. The L4, meanwhile, averages 131,072 and sits at the 95th percentile. These percentile rankings show that both cards are high-performing parts, but the V100 holds a slight edge in overall standing. The V100's nearest rival is the AMD Instinct MI100 with an average score of 139,035, which is 0.7% higher than the V100's average—a negligible difference. Against its direct competitors, the V100 leads the NVIDIA Tesla V100 SXM2 32 GB by 0.2%, the AMD Radeon PRO V620 by 1.2%, and the AMD Radeon Pro W6800X Duo by 1.7%. The L4's closest rival is the NVIDIA GeForce RTX 3090 Ti with an average score of 131,938, which is 0.7% higher than the L4's average. The L4 trails the NVIDIA RTX 4000 Ada Generation and NVIDIA A10M by 3.1% each, and the AMD Radeon PRO W6800 by 3.2%.

The head-to-head comparison reveals a 1-1 split in wins, but the magnitude of the victories matters. The V100's OpenCL win is more than twice as large in percentage terms as the L4's Vulkan win. This asymmetry suggests that for users prioritizing OpenCL compute, the V100 is the clear choice, while those focused on Vulkan rendering would prefer the L4. The data also shows that the V100's OpenCL score of 163,063 is significantly higher than its nearest rival's average score of 139,035, representing a 17.3% advantage over the AMD Instinct MI100. The L4's Vulkan score, in contrast, is 8.1% below the RTX 3090 Ti's average score.

# The Verdict

The benchmark results indicate two distinct performance profiles, and the choice between these cards depends entirely on the workload. The NVIDIA Tesla V100 PCIe 16 GB is the superior option for OpenCL-based compute tasks, delivering a 15.8% performance advantage over the L4 in that specific test. Its average benchmark score of 138,063 also exceeds the L4's 131,072 by 5.3%, and its 96th percentile ranking versus the L4's 95th confirms its higher overall standing. The V100's OpenCL performance is particularly impressive when compared to its rivals—it leads the AMD Instinct MI100 by 0.7% in average score, despite the MI100 having a higher peak average among the V100's nearest rivals.

However, the NVIDIA L4 is the better choice for Vulkan-based workloads, where it outperforms the V100 by 6.8%. The L4's Vulkan score of 121,306 also represents a 14.9% improvement over its own OpenCL score, suggesting that its architecture is better optimized for modern graphics APIs. The L4's newer Ada Lovelace architecture, with its support for DirectX 12 Ultimate (12_2), also gives it a feature advantage in graphics-oriented applications. The V100, by contrast, supports DirectX 12 (12_1), which lacks some of the newer features found in the Ultimate specification.

For users who need a single card for mixed workloads, the V100's higher average benchmark score and better OpenCL performance make it the safer recommendation. The V100's 15.8% OpenCL advantage is a larger margin than the L4's 6.8% Vulkan advantage, meaning the V100's strengths are more pronounced than the L4's. Additionally, the V100's 96th percentile ranking versus the L4's 95th indicates that the V100 is marginally better positioned among all GPUs. The data does not support choosing the L4 for general-purpose compute, but it is the clear winner for Vulkan-specific applications.

# Where Each One Wins

The NVIDIA Tesla V100 PCIe 16 GB wins in scenarios that rely on OpenCL compute performance. Its score of 163,063 in Geekbench OpenCL is 15.8% higher than the L4's, making it the dominant choice for OpenCL-accelerated scientific computing, machine learning inference, and data processing tasks. The V100 also holds the advantage in overall average benchmark score, with 138,063 versus the L4's 131,072, a 5.3% lead. This translates to better performance across a broad range of compute workloads, not just OpenCL. The V100's 640 tensor cores, compared to the L4's 240, further reinforce its compute credentials, though tensor core performance is not directly measured in the available benchmarks.

The NVIDIA L4 wins in Vulkan-based graphics workloads. Its Geekbench Vulkan score of 121,306 is 6.8% higher than the V100's 113,062, demonstrating superior performance in Vulkan rendering and compute tasks. The L4's 60 ray tracing cores, which the V100 lacks entirely, provide hardware acceleration for ray-traced graphics workloads. The L4 also supports DirectX 12 Ultimate (12_2), while the V100 is limited to DirectX 12 (12_1), making the L4 the better choice for applications that leverage the latest graphics features. The L4's higher boost clock of 2040 MHz, compared to the V100's 1380 MHz, contributes to its Vulkan advantage despite having fewer shading units (7424 versus 5120). The L4's FP32 throughput of 30.29 TFLOPS also doubles the V100's 14.13 TFLOPS, which benefits floating-point compute tasks, though this advantage is not reflected in the OpenCL benchmark where the V100 wins.

# FAQ

Q: Which card has the higher average benchmark score?

A: The NVIDIA Tesla V100 PCIe 16 GB has an average benchmark score of 138,063, which is 5.3% higher than the NVIDIA L4's average of 131,072.

Q: How do the two cards compare in Geekbench OpenCL?

A: The V100 scores 163,063 in Geekbench OpenCL, which is 15.8% higher than the L4's 140,838. The V100 wins this test.

Q: Which card performs better in Geekbench Vulkan?

A: The L4 scores 121,306 in Geekbench Vulkan, which is 6.8% higher than the V100's 113,062. The L4 wins this test.

Q: What are the percentile rankings for each card?

A: The V100 is at the 96th percentile of all GPUs, while the L4 is at the 95th percentile.

Q: How do the cards compare to their nearest rivals?

A: The V100's nearest rival is the AMD Instinct MI100, which has an average score of 139,035, 0.7% higher than the V100. The L4's nearest rival is the NVIDIA GeForce RTX 3090 Ti, which has an average score of 131,938, 0.7% higher than the L4.

Q: Which card has more tensor cores?

A: The V100 has 640 tensor cores, while the L4 has 240 tensor cores.

# Architecture Differences

The two cards represent fundamentally different generations of NVIDIA architecture. The V100 is built on the Volta architecture using the GV100 chip, while the L4 uses the Ada Lovelace architecture with the AD104 chip. This generational gap manifests in several key specifications. The V100 is fabricated on a 12 nm process at TSMC, whereas the L4 uses a 5 nm process, also at TSMC. The L4's newer process node allows for significantly higher transistor density: 121.8 million transistors per square millimeter compared to the V100's 25.9 million. The L4 packs 35,800 million transistors onto a 294 mm² die, while the V100 has 21,100 million transistors on a much larger 815 mm² die.

Memory configurations differ substantially. The V100 uses 16 GB of HBM2 memory with a 4096-bit bus and 897.0 GB/s bandwidth. The L4 uses 24 GB of GDDR6 memory with a 192-bit bus and 300.1 GB/s bandwidth. The V100's HBM2 memory provides nearly three times the bandwidth of the L4's GDDR6, which is crucial for memory-intensive compute workloads. The L4 compensates with 50% more memory capacity, which benefits large datasets that exceed the V100's 16 GB limit.

Compute resources also vary. The V100 has 5120 shading units, 320 texture mapping units, and 128 render output units. The L4 has 7424 shading units, 240 TMUs, and 80 ROPs. Despite having fewer shading units, the V100's higher memory bandwidth and older architecture result in different performance characteristics. The V100's FP32 throughput is 14.13 TFLOPS, while the L4 achieves 30.29 TFLOPS—more than double. The V100's FP16 performance is 28.26 TFLOPS (2:1 ratio), while the L4 delivers 30.29 TFLOPS (1:1 ratio), meaning the L4 does not gain a throughput advantage when switching to FP16. The V100 includes 640 tensor cores, while the L4 has 240 tensor cores and adds 60 ray tracing cores, which the V100 lacks entirely.

Clock speeds and power characteristics differ markedly. The V100 has a base clock of 1245 MHz and a boost clock of 1380 MHz, with a TDP of 300 W requiring a dual-slot cooler and two 8-pin power connectors. The L4 has a base clock of 795 MHz but a much higher boost clock of 2040 MHz, with a TDP of just 72 W, a single-slot design, and no power connectors required. The L4's suggested PSU is 250 W, compared to the V100's 700 W, making the L4 dramatically more power-efficient. The L4 also supports PCIe 4.0 x16, while the V100 uses PCIe 3.0 x16. Neither card has display outputs, as both are designed for server deployment. The V100 was released on June 20, 2017, and is end-of-life, while the L4 was released on March 20, 2023, and remains active in production.

DETAILED SPECIFICATIONS

SPECIFICATION
L4
Tesla V100 PCIe 16 GB
Core Specs
Shading Units
7,424
5,120 -31.0%
Shaders
7,424
5,120 -31.0%
TMUs
240
320 +33.3%
ROPs
80
128 +60.0%
SM Count
60
80 +33.3%
Clocks
Base Clock
795 MHz
1245 MHz
Boost Clock
2040 MHz
1380 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
876 MHz 1752 Mbps effective
Memory
Memory Size
24 GB
16 GB
VRAM (MB)
24,576
16,384 -33.3%
Memory Type
GDDR6
HBM2
Memory Bus
192 bit
4096 bit
Bandwidth
300.1 GB/s
897.0 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
48 MB
6 MB
Performance
Pixel Rate
163.2 GPixel/s
176.6 GPixel/s
Texture Rate
489.6 GTexel/s
441.6 GTexel/s
FP32 (TFLOPS)
30.29 TFLOPS
14.13 TFLOPS
FP64 (TFLOPS)
473.3 GFLOPS (1:64)
7.066 TFLOPS (1:2)
FP16 (TFLOPS)
30.29 TFLOPS (1:1)
28.26 TFLOPS (2:1)
AI/RT
RT Cores
60
—
Tensor Cores
240
640 +166.7%
Power
TDP
72 W
300 W
TDP (W)
72
300 +316.7%
Suggested PSU
250 W
700 W
Power Connectors
None
2x 8-pin
Architecture
Architecture
Ada Lovelace
Volta
GPU Name
AD104
GV100
Generation
Server Ada (Lxx)
Tesla Volta (Vxx)
Process Size
5 nm
12 nm
Transistors
35,800 million
21,100 million
Die Size
294 mm²
815 mm²
Foundry
TSMC
TSMC
Density
121.8M / mm²
25.9M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
7.0
Shader Model
6.8
6.8
Physical
Slot Width
Single-slot
Dual-slot
Length
169 mm 6.7 inches
—
Height
56 mm 2.2 inches
—
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 3.0 x16
Other
Production
Active
End-of-life
Predecessor
Server Ampere
Tesla Pascal
Successor
Server Hopper
Tesla Turing
View L4 Details View Tesla V100 PCIe 16 GB Details