NVIDIA Tesla T4
NVIDIA graphics card specifications and benchmark scores
At a Glance
NVIDIANVIDIA Tesla T4 Specifications
GPU Core
Shader units and compute resources
The NVIDIA Tesla T4 GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.
Tesla T4 Clock Speeds
GPU and memory frequencies
Clock speeds directly impact the Tesla T4's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Tesla T4 by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.
NVIDIA's Tesla T4 Memory
VRAM capacity and bandwidth
VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Tesla T4's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.
Tesla T4 by NVIDIA Cache
On-chip cache hierarchy
On-chip cache provides ultra-fast data access for the Tesla T4, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.
Tesla T4 Theoretical Performance
Compute and fill rates
Theoretical performance metrics provide a baseline for comparing the NVIDIA Tesla T4 against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.
Tesla T4 Ray Tracing & AI
Hardware acceleration features
The NVIDIA Tesla T4 includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the Tesla T4 capable of delivering both stunning graphics and smooth frame rates in modern titles.
Turing Architecture & Process
Manufacturing and design details
The NVIDIA Tesla T4 is built on NVIDIA's Turing architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Tesla T4 will perform in GPU benchmarks compared to previous generations.
Power & Thermal
TDP and power requirements
Power specifications for the NVIDIA Tesla T4 determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Tesla T4 to maintain boost clocks without throttling.
Tesla T4 by NVIDIA Physical & Connectivity
Dimensions and outputs
Physical dimensions of the NVIDIA Tesla T4 are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.
NVIDIA API Support
Graphics and compute APIs
API support determines which games and applications can fully utilize the NVIDIA Tesla T4. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.
Tesla T4 Product Information
Release and pricing details
The NVIDIA Tesla T4 is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Tesla T4 by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.
About NVIDIA Tesla T4
The NVIDIA Tesla T4 occupies a distinctive position in the hardware landscape: it is an end-of-life data center accelerator built on the 12 nm Turing architecture, yet its benchmark scores place it within a razor-thin margin of much newer and more power-hungry consumer and workstation flagships. With an average benchmark score of 66733 and a 91st percentile ranking among all GPUs, the T4 demonstrates that its legacy status does not equate to obsolescence. The data shows a product that traded raw clock speed for efficiency and feature completeness, delivering competitive compute results in a 70 W package that requires no external power connectors.
How It Compares
Against the NVIDIA GeForce RTX 4090, the Tesla T4 trails by a negligible 0.4% in average benchmark score (66733 vs 66473). This is a remarkable outcome given the generational gap and the vast difference in power envelopes. The data indicates that for the specific workloads captured by the Geekbench OpenCL and Vulkan tests, the T4’s 40 RT cores and 320 Tensor cores provide enough compute throughput to nearly match the flagship consumer card, suggesting that the T4’s Turing architecture is exceptionally well-optimized for these synthetic benchmarks.
The AMD Radeon Pro Vega 56 edges out the Tesla T4 by just 0.5% (67097 vs 66733). This workstation-class AMD card offers a slightly higher average score, but the margin is within noise for most real-world applications. The T4 counters with a substantially lower TDP of 70 W versus the Vega 56’s much higher power draw, and the T4’s GDDR6 memory at 320.0 GB/s bandwidth provides a modern memory subsystem that the older Vega architecture lacks. The benchmark results indicate a statistical tie, with the T4 winning on efficiency and feature set.
The NVIDIA Quadro P6000 leads the Tesla T4 by 0.9% (67320 vs 66733), the largest deficit among the listed rivals. The P6000, based on the older Pascal architecture, achieves its score through sheer shader count and high clock speeds, but it lacks the T4’s dedicated RT cores, Tensor cores, and DirectX 12 Ultimate support. The data suggests that the T4’s architectural advantages in ray tracing and AI acceleration are not fully reflected in these particular Geekbench tests, which favor raw compute throughput.
The NVIDIA Tesla P40 is the closest competitor, sitting 0.9% behind the T4 (66127 vs 66733). Both are data center cards, but the P40 uses the older Pascal architecture with no RT or Tensor cores. The T4’s 320 Tensor cores and 40 RT cores provide capabilities that the P40 simply cannot match, and the T4 also offers a higher boost clock of 1590 MHz versus the P40’s more modest specifications. Benchmark results show the T4 winning this head-to-head, cementing its position as the more advanced accelerator.
Ray Tracing and Feature Set
The Tesla T4 is built on the Turing architecture and incorporates 40 dedicated RT cores, which are specifically designed for hardware-accelerated ray tracing. This makes the T4 capable of real-time ray-traced workloads, a feature absent from its immediate predecessor, Tesla Volta. The presence of these cores means that the T4 can handle ray-traced rendering tasks that would otherwise be relegated to software or would be impossible on older architectures.
Complementing the RT cores are 320 Tensor cores, which are specialized for matrix math and deep learning inference. These Tensor cores enable the T4 to accelerate AI workloads, including neural network training and inference, far more efficiently than traditional shader cores. The combination of RT and Tensor cores positions the T4 as a dual-purpose accelerator for both graphics and compute-intensive AI tasks, a rarity in the server segment at its launch.
On the API front, the T4 supports DirectX 12 Ultimate (12_2), which is the latest feature level, ensuring compatibility with modern games and applications that leverage advanced rendering techniques. It also supports OpenGL 4.6 and Vulkan 1.4, providing broad cross-platform compatibility. The T4 features no display outputs, confirming its role as a headless compute or rendering accelerator rather than a consumer graphics card. Its PCIe 3.0 x16 interface provides adequate bandwidth for most workloads, though newer generations offer higher throughput.
Benchmark Performance
The Tesla T4’s benchmark results are remarkably consistent across the two Geekbench tests, with a Geekbench OpenCL score of 61276 and a Geekbench Vulkan score of 72190. The Vulkan score is notably higher, indicating that the T4’s architecture scales well with modern low-level APIs that better utilize its parallel compute resources. The average benchmark score of 66733 places the T4 in the 91st percentile of all GPUs, a strong showing for a card released in 2018.
The deltaPct values against nearest rivals reveal a tightly contested field. The T4 is 0.4% ahead of the RTX 4090 (66473), which is statistically insignificant but noteworthy given the RTX 4090’s much larger die and power budget. Against the Radeon Pro Vega 56, the T4 is 0.5% behind (67097), and against the Quadro P6000, it is 0.9% behind (67320). The only rival the T4 clearly beats is the Tesla P40, where it holds a 0.9% advantage (66127). These margins are all under 1%, meaning the T4 performs at the same tier as these competitors in synthetic benchmarks, despite being a lower-power, end-of-life product.
The FP32 compute rating of 8.141 TFLOPS is modest by modern standards, but the FP16 rating of 65.13 TFLOPS (8:1 ratio) is exceptionally high, thanks to the Tensor cores. This suggests that the T4 is disproportionately faster in FP16 workloads, which are common in AI inference and training, compared to its raw FP32 throughput. The pixel rate of 101.8 GPixel/s and texture rate of 254.4 GTexel/s are adequate for the card’s intended server workloads, but they are not competitive with high-end consumer gaming cards.
FAQ
Q: How does the Tesla T4 compare to the RTX 4090 in benchmark scores?
A: The Tesla T4 has an average benchmark score of 66733, which is 0.4% higher than the RTX 4090’s 66473, indicating near-identical performance in the Geekbench OpenCL and Vulkan tests.
Q: Does the Tesla T4 support ray tracing?
A: Yes, the T4 includes 40 dedicated RT cores based on the Turing architecture, enabling hardware-accelerated ray tracing.
Q: What is the memory bandwidth of the Tesla T4?
A: The T4 has a 256-bit memory bus with 16 GB of GDDR6 memory, providing a bandwidth of 320.0 GB/s.
Q: Is the Tesla T4 still in production?
A: No, the production status is listed as end-of-life, with a release date of September 12, 2018.
Q: What API features does the Tesla T4 support?
A: The T4 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.
Q: How many Tensor cores does the Tesla T4 have?
A: The T4 contains 320 Tensor cores, which are used for AI and deep learning acceleration.
Power and Cooling
The Tesla T4 has a TDP of just 70 W, which is exceptionally low for a GPU with 16 GB of GDDR6 memory and 2560 shading units. This low power draw is achieved through a base clock of 585 MHz and a boost clock of 1590 MHz, running on a 12 nm TSMC process with 13,600 million transistors. The T4 requires no external power connectors, drawing all its power from the PCIe 3.0 x16 slot, and a suggested power supply of 250 W is sufficient for a system containing this card.
The card is single-slot in design, measuring 168 mm (6.6 inches) in length, which makes it suitable for dense server configurations where space is at a premium. With no display outputs and no power connectors, the T4 is designed for rack-mounted compute nodes rather than desktop towers. The low TDP also means that cooling requirements are minimal, and a simple passive heatsink or low-profile active cooler is adequate, though the fact pack does not specify a cooler type.
The 70 W TDP is a fraction of what rival cards like the RTX 4090 consume, yet the benchmark scores are nearly identical. This efficiency is the T4’s primary advantage in power-constrained environments, where multiple cards can be installed without overwhelming thermal or power budgets. The data indicates that the T4 achieves its performance through architectural efficiency rather than brute-force clock speeds.
Who Should Consider It
The Tesla T4 is best suited for server and data center environments where AI inference and ray-traced rendering are required, but power and space are limited. With its 70 W TDP and single-slot form factor, the T4 allows for high-density deployments where multiple accelerators are needed. The benchmark scores show that it performs at the level of the RTX 4090 in synthetic tests, making it a viable option for compute workloads that do not require display output.
For users working with FP16 workloads, the T4’s 65.13 TFLOPS (8:1) rating is particularly compelling, as it offers over 8 times the throughput of its FP32 rating of 8.141 TFLOPS. This makes the T4 an excellent choice for deep learning inference tasks, where Tensor cores can be fully utilized. The 16 GB of GDDR6 memory provides sufficient capacity for large models, and the 320.0 GB/s bandwidth ensures data can be fed to the compute units quickly.
At high resolutions, such as 4K or multi-monitor setups, the T4’s memory subsystem is adequate, but the card’s lack of display outputs means it is not intended for direct gaming or workstation use. Instead, it is a rendering or compute accelerator that would be paired with a separate display adapter. The 91st percentile ranking indicates that the T4 outperforms the vast majority of GPUs, making it a strong choice for any server workload that aligns with its feature set.
Memory Subsystem
The Tesla T4 is equipped with 16 GB of GDDR6 memory on a 256-bit bus, yielding a bandwidth of 320.0 GB/s. The memory operates at 1250 MHz, with an effective data rate of 10 Gbps. This configuration provides a balanced memory subsystem that is well-suited for the card’s compute and AI workloads, where large datasets and intermediate results need to be stored on-chip.
The 16 GB capacity is generous for a card of this class, allowing it to handle large neural network models or high-resolution render buffers without spilling to system memory. The 256-bit bus width is narrower than some high-end cards, but the GDDR6 technology ensures that the 320.0 GB/s bandwidth is sufficient to keep the 2560 shading units and 40 RT cores fed with data. For FP16 workloads, the memory bandwidth becomes even more critical, as the 65.13 TFLOPS compute rate requires substantial data throughput to avoid stalling.
The memory subsystem’s performance is reflected in the benchmark scores, where the T4 matches or exceeds rivals with similar or larger memory configurations. The Radeon Pro Vega 56 and Quadro P6000 both have comparable bandwidth, but the T4’s GDDR6 is more modern and efficient than the older HBM2 and GDDR5X used by those cards. The 320.0 GB/s bandwidth is adequate for 4K rendering and large-scale compute, though it is not a standout feature compared to newer cards with 1 TB/s or higher bandwidth. For the T4’s intended server role, the memory configuration strikes a practical balance between capacity, speed, and power consumption.
Detailed benchmark scores and charts for the NVIDIA Tesla T4 are below.
Benchmark Scores
geekbench_openclSource
Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA Tesla T4 handles parallel computing tasks like video encoding and scientific simulations.
geekbench_vulkanSource
Geekbench Vulkan tests GPU compute using the modern low-overhead Vulkan API. This shows how NVIDIA Tesla T4 performs with next-generation graphics and compute workloads. Vulkan offers better CPU efficiency than older APIs like OpenGL.
Popular NVIDIA Tesla T4 Comparisons
See how the Tesla T4 stacks up against similar graphics cards from the same generation and competing brands.
Compare with Other GPUs
Select another GPU to compare specifications and benchmarks side-by-side.
Browse GPUs