NVIDIA L4 vs NVIDIA Tesla T4 Comparison

NVIDIA
GEFORCE

NVIDIA L4

CORE STATE AD104
VRAM 24 GB
CLOCK SPEED 2040 MHz
TDP 72 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

Tesla T4

CORE STATE TU104
VRAM 16 GB
CLOCK SPEED 1590 MHz
TDP 70 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2018

PERFORMANCE BENCHMARKS

geekbench_opencl
140,838
61,276
geekbench_vulkan
121,306
72,190

Analysis: NVIDIA L4 vs NVIDIA Tesla T4

The NVIDIA L4 and NVIDIA Tesla T4 are both single-slot, passively cooled server accelerators from NVIDIA, but they represent very different generations of GPU architecture. The L4, built on the Ada Lovelace architecture and a 5 nm process, is the current active product, while the Tesla T4, based on the older Turing architecture and a 12 nm process, has reached end-of-life status. This comparison examines the recorded benchmark data and specification differences between these two cards, focusing on what the numbers reveal about their relative performance and positioning.

Head-to-Head Benchmarks

The database contains two direct benchmark comparisons between the NVIDIA L4 and the NVIDIA Tesla T4: Geekbench OpenCL and Geekbench Vulkan. In both tests, the L4 emerges as the clear winner, and the margins are substantial.

In the Geekbench OpenCL test, the L4 scores 140,838 points, while the Tesla T4 scores 61,276 points. This represents a delta of 129.8% in favor of the L4. To put this in perspective, the L4 is more than double the score of the T4 in this compute-oriented workload. The OpenCL test typically stresses general-purpose GPU compute, including floating-point operations and memory throughput, and the L4’s advantage here is decisive.

The Geekbench Vulkan test shows a narrower but still significant gap. The L4 scores 121,306, while the Tesla T4 scores 72,190, resulting in a 68% delta for the L4. Vulkan is a lower-level graphics and compute API, and while the L4 still wins by a wide margin, the relative difference is smaller than in OpenCL. This suggests that the T4’s Turing architecture holds up better in certain graphics-oriented workloads than in pure compute tasks, though it still trails considerably.

The overall win count is 2 for the L4 and 0 for the Tesla T4. The average benchmark score for the L4 is 131,072, placing it in the 95th percentile of all GPUs in the database. The Tesla T4’s average score is 66,733, which sits in the 90th percentile. While both are high-performing cards relative to the entire GPU landscape, the L4’s average score is roughly double that of the T4, confirming the head-to-head results.

Looking at the nearest rivals for each card provides additional context. The L4’s closest competitors include the NVIDIA GeForce RTX 3090 Ti (average score 131,938, a delta of -0.7% relative to the L4), the NVIDIA RTX 4000 Ada Generation (135,218, -3.1%), the NVIDIA A10M (135,230, -3.1%), and the AMD Radeon PRO W6800 (135,396, -3.2%). These are all high-end cards, and the L4 is essentially trading blows with them, sitting slightly below each in average score but within 3.2% of all of them. This indicates that the L4 is a very capable accelerator that competes with top-tier consumer and workstation GPUs.

The Tesla T4’s nearest rivals are a different class of hardware. The AMD Radeon VII (66,004, a delta of 1.1% relative to the T4) is slightly below the T4, while the NVIDIA Tesla P40 (65,095, 2.5%) is also below. The AMD Radeon Instinct MI25 (68,562, -2.7%) and the Intel Arc A770 (68,809, -3%) both sit above the T4. This places the T4 in a mid-range tier, competitive with older or lower-end accelerators but far from the performance levels of the L4.

The Verdict

The benchmark data is unambiguous: the NVIDIA L4 is the superior performer in every recorded test. In OpenCL, the L4 is 129.8% faster, and in Vulkan, it is 68% faster. The average benchmark score of 131,072 for the L4 compared to 66,733 for the Tesla T4 represents a doubling of overall performance.

For users or systems requiring maximum compute throughput from a single-slot card, the L4 is the clear choice. Its performance places it in the 95th percentile of all GPUs, and its nearest rivals are top-tier cards like the RTX 3090 Ti and RTX 4000 Ada Generation. The L4 is also an active product, meaning it is currently in production and supported, whereas the Tesla T4 is end-of-life.

However, the Tesla T4 is not without merit, though its advantages are not reflected in the raw performance numbers. The T4’s average score of 66,733 still places it in the 90th percentile, which is respectable. Its nearest rivals include the AMD Radeon VII and the NVIDIA Tesla P40, and it slightly outperforms both. For legacy systems or applications built around the Turing architecture, the T4 could still be a functional option, especially if its lower absolute performance is acceptable for the workload.

The data suggests that anyone choosing between these two should pick the L4 unless there is a specific compatibility or software constraint that requires the older Turing architecture. The performance gap is too large to justify the T4 on speed alone, and the L4’s active production status is a further point in its favor.

Where Each One Wins

Based strictly on the recorded data, the NVIDIA L4 wins in every benchmark category. There are no tests in which the Tesla T4 outperforms the L4. The wins are 2 for the L4 and 0 for the T4.

The L4’s largest advantage is in OpenCL, where its 129.8% delta indicates that it excels at general-purpose compute tasks. This is consistent with its higher FP32 throughput of 30.29 TFLOPS compared to the T4’s 8.141 TFLOPS, a difference of roughly 3.7 times. The L4 also has a significantly higher pixel rate (163.2 GPixel/s vs. 101.8 GPixel/s) and texture rate (489.6 GTexel/s vs. 254.4 GTexel/s), which contribute to its performance in both compute and graphics workloads.

The Vulkan test shows a smaller margin, at 68%. This suggests that while the L4 is still much faster, the T4’s Turing architecture with its dedicated RT and Tensor cores may handle certain graphics or ray-tracing tasks relatively better than its raw compute numbers would imply. The T4 has 320 Tensor Cores compared to the L4’s 240, and it also has 40 RT cores versus the L4’s 60. The T4’s FP16 performance of 16.28 TFLOPS (at a 2:1 ratio) is actually higher relative to its FP32 than the L4’s, which offers 30.29 TFLOPS FP16 at a 1:1 ratio. This means the T4’s Tensor Core performance may be more balanced relative to its overall compute, though the L4’s absolute numbers are still far higher.

For users focused on FP32 compute, the L4 is the clear winner. For workloads that heavily utilize FP16 or rely on Tensor Cores, the T4’s architecture offers a different balance, but the L4’s raw power still gives it an edge in most scenarios. The T4’s only potential niche is in applications that are specifically optimized for its Turing-era feature set or that require its 16 GB of memory in a particular configuration, though the L4 offers 24 GB, which is larger.

FAQ

Q: Which GPU is faster in Geekbench OpenCL?

A: The NVIDIA L4 is faster, scoring 140,838 compared to the Tesla T4’s 61,276. This is a 129.8% difference in favor of the L4.

Q: How do the two cards compare in Geekbench Vulkan?

A: The L4 also wins this test, with a score of 121,306 versus the T4’s 72,190, a delta of 68%.

Q: What is the average benchmark score for each card?

A: The L4 has an average benchmark score of 131,072, while the Tesla T4 has an average of 66,733. The L4’s score is nearly double that of the T4.

Q: Are both cards still in production?

A: No. The NVIDIA L4 has an active production status, while the NVIDIA Tesla T4 is listed as end-of-life.

Q: Which GPU has a higher memory capacity?

A: The L4 has 24 GB of GDDR6 memory, while the Tesla T4 has 16 GB of GDDR6 memory. The L4 also has a wider memory bus at 192 bit versus the T4’s 256 bit, but the T4’s bandwidth is slightly higher at 320.0 GB/s compared to the L4’s 300.1 GB/s.

Q: What are the nearest rivals for each card?

A: The L4’s nearest rivals include the NVIDIA GeForce RTX 3090 Ti (delta -0.7%), NVIDIA RTX 4000 Ada Generation (-3.1%), NVIDIA A10M (-3.1%), and AMD Radeon PRO W6800 (-3.2%). The Tesla T4’s nearest rivals are the AMD Radeon VII (delta 1.1%), NVIDIA Tesla P40 (2.5%), AMD Radeon Instinct MI25 (-2.7%), and Intel Arc A770 (-3%).

Architecture Differences

The NVIDIA L4 and NVIDIA Tesla T4 are built on fundamentally different architectures. The L4 uses the Ada Lovelace architecture with the AD104 chip, manufactured on a 5 nm process by TSMC. The T4 uses the Turing architecture with the TU104 chip, on a 12 nm process also by TSMC. This process difference is significant: the L4 packs 35,800 million transistors into a 294 mm² die, resulting in a transistor density of 121.8 million per mm². The T4 has 13,600 million transistors on a much larger 545 mm² die, giving a density of just 25.0 million per mm². The L4’s newer process allows for far greater integration and efficiency.

The chip designs also differ in core counts. The L4 has 7,424 shading units, 240 texture mapping units (TMUs), and 80 raster operations units (ROPs). The T4 has 2,560 shading units, 160 TMUs, and 64 ROPs. The L4 also features 60 RT cores and 240 Tensor Cores, while the T4 has 40 RT cores and 320 Tensor Cores. Despite having fewer Tensor Cores, the L4’s newer architecture delivers higher FP16 performance in absolute terms (30.29 TFLOPS vs. 16.28 TFLOPS), though the T4 offers a 2:1 FP16 to FP32 ratio while the L4 is 1:1.

Both cards support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, so API compatibility is identical. However, the L4 uses a PCIe 4.0 x16 interface, while the T4 uses PCIe 3.0 x16. The L4 also has no display outputs, and the same is true for the T4; both are compute-focused accelerators.

The L4 is part of the Server Ada generation (Lxx) and has a predecessor of Server Ampere and a successor of Server Hopper. The T4 is from the Tesla Turing generation (Txx), with a predecessor of Tesla Volta and a successor of Server Ampere. The release dates differ as well: the L4 was released on 2023-03-20, while the T4 was released on 2018-09-12.

Specification Differences

The specification differences between the two cards are extensive. The most notable difference is in the process node: the L4 uses 5 nm, while the T4 uses 12 nm. The L4 has 35,800 million transistors compared to the T4’s 13,600 million, and the die sizes are 294 mm² for the L4 versus 545 mm² for the T4. The transistor density is 121.8 million per mm² for the L4 and 25.0 million per mm² for the T4.

Clock speeds differ, with the L4 having a base clock of 795 MHz and a boost clock of 2040 MHz, while the T4 has a base of 585 MHz and a boost of 1590 MHz. Memory clocks also differ: the L4 runs at 1563 MHz with 12.5 Gbps effective, while the T4 runs at 1250 MHz with 10 Gbps effective.

Memory capacity is 24 GB for the L4 versus 16 GB for the T4, and the bus widths are 192 bit for the L4 and 256 bit for the T4. The memory bandwidth is 300.1 GB/s for the L4 and 320.0 GB/s for the T4, meaning the T4 actually has a slight bandwidth advantage due to its wider bus.

Core counts are higher on the L4: 7,424 shading units versus 2,560, 240 TMUs versus 160, and 80 ROPs versus 64. The L4 has 60 RT cores and 240 Tensor Cores, while the T4 has 40 RT cores and 320 Tensor Cores. The pixel rate is 163.2 GPixel/s for the L4 and 101.8 GPixel/s for the T4, and the texture rate is 489.6 GTexel/s versus 254.4 GTexel/s.

FP32 performance is 30.29 TFLOPS for the L4 and 8.141 TFLOPS for the T4, while FP16 is 30.29 TFLOPS (1:1) for the L4 and 16.28 TFLOPS (2:1) for the T4. The TDP is 72 W for the L4 and 70 W for the T4, nearly identical. Both are single-slot with no power connectors and a suggested PSU of 250 W.

The bus interface is PCIe 4.0 x16 for the L4 and PCIe 3.0 x16 for the T4. Dimensions are similar: the L4 is 169 mm long and 56 mm high, while the T4 is 168 mm long with no recorded height. The production status is Active for the L4 and End-of-life for the T4.

The release dates are also different, with the L4 launching on 2023-03-20 and the T4 on 2018-09-12. The T4’s predecessor is Tesla Volta and its successor is Server Ampere, while the L4’s predecessor is Server Ampere and its successor is Server Hopper.

DETAILED SPECIFICATIONS

SPECIFICATION
L4
Tesla T4
Core Specs
Shading Units
7,424
2,560 -65.5%
Shaders
7,424
2,560 -65.5%
TMUs
240
160 -33.3%
ROPs
80
64 -20.0%
SM Count
60
40 -33.3%
Clocks
Base Clock
795 MHz
585 MHz
Boost Clock
2040 MHz
1590 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
1250 MHz 10 Gbps effective
Memory
Memory Size
24 GB
16 GB
VRAM (MB)
24,576
16,384 -33.3%
Memory Type
GDDR6
GDDR6
Memory Bus
192 bit
256 bit
Bandwidth
300.1 GB/s
320.0 GB/s
Cache
L1 Cache
128 KB (per SM)
64 KB (per SM)
L2 Cache
48 MB
4 MB
Performance
Pixel Rate
163.2 GPixel/s
101.8 GPixel/s
Texture Rate
489.6 GTexel/s
254.4 GTexel/s
FP32 (TFLOPS)
30.29 TFLOPS
8.141 TFLOPS
FP64 (TFLOPS)
473.3 GFLOPS (1:64)
254.4 GFLOPS (1:32)
FP16 (TFLOPS)
30.29 TFLOPS (1:1)
16.28 TFLOPS (2:1)
AI/RT
RT Cores
60
40 -33.3%
Tensor Cores
240
320 +33.3%
Power
TDP
72 W
70 W
TDP (W)
72
70 -2.8%
Suggested PSU
250 W
250 W
Power Connectors
None
None
Architecture
Architecture
Ada Lovelace
Turing
GPU Name
AD104
TU104
Generation
Server Ada (Lxx)
Tesla Turing (Txx)
Process Size
5 nm
12 nm
Transistors
35,800 million
13,600 million
Die Size
294 mm²
545 mm²
Foundry
TSMC
TSMC
Density
121.8M / mm²
25.0M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
7.5
Shader Model
6.8
6.9
Physical
Slot Width
Single-slot
Single-slot
Length
169 mm 6.7 inches
168 mm 6.6 inches
Height
56 mm 2.2 inches
Outputs
No outputs
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 3.0 x16
Other
Production
Active
End-of-life
Predecessor
Server Ampere
Tesla Volta
Successor
Server Hopper
Server Ampere
View L4 Details View Tesla T4 Details