GEFORCE

NVIDIA Tesla T4G

NVIDIA graphics card specifications and benchmark scores

16 GB
VRAM
1590
MHz Boost
70W
TDP
256
Bus Width
Ray Tracing Tensor Cores

At a Glance

NVIDIA
VRAM 16 GB
Boost Clock 1,590 MHz
Shaders 2,560
Bus Width 256-bit
TDP 70W
Memory Type GDDR6
RT Cores 40
Architecture Turing
nm
Process 12 nm
Released Sep 2018

NVIDIA Tesla T4G Specifications

GPU Core

Shader units and compute resources

The NVIDIA Tesla T4G GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
2,560
Shaders
2,560
TMUs
160
ROPs
64
SM Count
40

Tesla T4G Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the Tesla T4G's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Tesla T4G by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
585 MHz
Base Clock
585 MHz
Boost Clock
1590 MHz
Boost Clock
1,590 MHz
Memory Clock
1250 MHz 10 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's Tesla T4G Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Tesla T4G's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
16 GB
VRAM
16,384 MB
Memory Type
GDDR6
VRAM Type
GDDR6
Memory Bus
256 bit
Bus Width
256-bit
Bandwidth
320.0 GB/s

Tesla T4G by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the Tesla T4G, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
64 KB (per SM)
L2 Cache
4 MB

Tesla T4G Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA Tesla T4G against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
8.141 TFLOPS
FP64 (Double)
254.4 GFLOPS (1:32)
FP16 (Half)
65.13 TFLOPS (8:1)
Pixel Rate
101.8 GPixel/s
Texture Rate
254.4 GTexel/s

Tesla T4G Ray Tracing & AI

Hardware acceleration features

The NVIDIA Tesla T4G includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the Tesla T4G capable of delivering both stunning graphics and smooth frame rates in modern titles.

RT Cores
40
Tensor Cores
320

Turing Architecture & Process

Manufacturing and design details

The NVIDIA Tesla T4G is built on NVIDIA's Turing architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Tesla T4G will perform in GPU benchmarks compared to previous generations.

Architecture
Turing
GPU Name
TU104
Process Node
12 nm
Foundry
TSMC
Transistors
13,600 million
Die Size
545 mm²
Density
25.0M / mm²

Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA Tesla T4G determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Tesla T4G to maintain boost clocks without throttling.

TDP
70 W
TDP
70W
Power Connectors
None
Suggested PSU
250 W

Tesla T4G by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA Tesla T4G are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
Single-slot
Length
168 mm 6.6 inches
Bus Interface
PCIe 3.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA Tesla T4G. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 Ultimate (12_2)
DirectX
12 Ultimate (12_2)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.4
Vulkan
1.4
OpenCL
3.0
CUDA
7.5
Shader Model
6.8

Tesla T4G Product Information

Release and pricing details

The NVIDIA Tesla T4G is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Tesla T4G by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Sep 2018
Production
End-of-life
Predecessor
Tesla Volta
Successor
Server Ampere

About NVIDIA Tesla T4G

The NVIDIA Tesla T4G is a Turing-architecture accelerator designed for server deployment, built on a 12 nm process at TSMC with 13,600 million transistors on a 545 mm² die. It is a single-slot, 70 W part with no display outputs and no power connectors, relying entirely on its PCIe 3.0 x16 slot for power, with a suggested system power supply of 250 W. Its production status is end-of-life, having been released on September 12, 2018, positioned between the Tesla Volta and Server Ampere generations. This analysis uses the available data to interpret its benchmark standing, memory configuration, competitive positioning, and feature set.

Benchmark Performance

The Tesla T4G holds a percentile rank of 50 among all GPUs, placing it exactly at the median of the database's performance distribution. Its average benchmark score is recorded as 0, which provides no direct numerical reference point; consequently, its standing is best understood through its percentile placement and architectural characteristics rather than absolute scores. The data shows that this is a mid-pack performer, neither a top-tier compute monster nor a low-end entry point, which aligns with its design as a power-efficient inference and virtualized workstation accelerator.

The compute capabilities are defined by its 2,560 shading units, 160 texture mapping units, and 64 raster output units. The FP32 throughput is 8.141 TFLOPS, a figure that establishes its general-purpose compute ceiling. In contrast, the FP16 performance is 65.13 TFLOPS with an 8:1 ratio, meaning the tensor cores are heavily leveraged for mixed-precision workloads, but the raw FP32 rate is modest by modern standards. The pixel rate is 101.8 GPixel/s, and the texture rate is 254.4 GTexel/s, which are moderate figures that reflect the card's orientation toward throughput in AI inference rather than raw rasterization.

The base clock is 585 MHz, which is notably low, but the boost clock reaches 1,590 MHz, indicating a wide dynamic range that allows the card to scale up significantly under load while maintaining idle efficiency. This clock behavior, combined with the 70 W TDP, suggests that the T4G is designed for sustained operation in dense server environments where thermal and power budgets are constrained. The benchmark results indicate that while it does not lead in raw FLOPS, its efficiency profile is a key attribute, making it a viable option for specific workloads rather than a general-purpose gaming or rendering card.

Memory Subsystem

The memory configuration is a critical strength for the Tesla T4G. It is equipped with 16 GB of GDDR6 memory on a 256-bit bus, yielding a bandwidth of 320.0 GB/s. The memory clock operates at 1,250 MHz, translating to 10 Gbps effective data rate. This capacity and bandwidth combination is substantial for its power envelope, allowing the card to handle large model footprints in inference tasks or substantial frame buffers in virtualized environments.

For high-resolution workloads, the 16 GB capacity is more decisive than the raw bandwidth. At 4K or higher resolutions in a virtual desktop infrastructure (VDI) context, the memory bus width of 256 bits is adequate, but the 320.0 GB/s bandwidth may become a limiting factor compared to higher-end accelerators with wider buses. However, the data suggests that for the T4G's intended use cases—such as AI inference and cloud rendering—the capacity to hold larger datasets in VRAM reduces the need for frequent host-device transfers, which is often more important than peak bandwidth. The lack of display outputs means this memory is not used for direct frame presentation but rather for compute buffers, making the capacity and bandwidth trade-off aligned with server-side processing.

The 10 Gbps effective memory speed is a conservative figure for GDDR6, prioritizing stability and lower power draw over peak throughput. This aligns with the 70 W TDP, as faster memory would increase thermal output. Consequently, the memory subsystem is balanced for sustained, high-density deployment rather than burst performance, and the 16 GB size ensures that multi-tenant workloads can coexist without exhausting VRAM.

How It Compares

The nearest rivals list is empty in the data, which means no direct comparative scores or deltaPct values are available. Therefore, formal comparisons against specific competitor GPUs cannot be quantified. However, its position in the product stack can be inferred from its generation and specifications.

Relative to its predecessor, the Tesla Volta, the T4G introduces Turing architecture features, notably hardware ray tracing cores and enhanced tensor cores, though the data does not provide performance deltas. The shift to GDDR6 from Volta's HBM2 indicates a cost and power optimization, trading bandwidth for capacity and availability.

Compared to its successor, Server Ampere, the T4G is older and likely slower in raw compute, but the data does not offer metrics to substantiate this. The T4G's 50th percentile standing suggests it is roughly average among all GPUs in the database, which positions it below enthusiast-class cards but above entry-level integrated solutions.

Without rival scores, the assessment must rely on architectural merits: the 40 RT cores and 320 tensor cores are present, but their effectiveness is unquantified here. The FP16 ratio of 8:1 indicates that tensor core performance is heavily weighted toward AI workloads, which is a differentiator from general-purpose cards that may have lower FP16 throughput. The 12 nm process, while mature, is less dense than newer nodes, but the 70 W TDP remains a compelling advantage for server density.

Who Should Consider It

The Tesla T4G is not a consumer-oriented card, given its lack of display outputs and server form factor. Its target audience is organizations deploying AI inference at scale, particularly where power efficiency is paramount. The 70 W TDP allows for high-density configurations in servers without extensive cooling or power infrastructure, and the 16 GB memory capacity is sufficient for many natural language processing or computer vision models that do not exceed that footprint.

For high-resolution rendering tasks, the 8.141 TFLOPS FP32 performance is modest, so it is not recommended for heavy 3D rendering or simulation workloads where compute throughput is the primary bottleneck. The data suggests it is better suited for batch inference or light virtualized workloads rather than real-time, high-fidelity graphics. The 101.8 GPixel/s pixel rate and 254.4 GTexel/s texture rate confirm that it is not designed for high-refresh-rate gaming or professional visualization.

The 50th percentile rank indicates that it offers average performance relative to all GPUs, which means it will not excel in compute-heavy tasks but will provide acceptable results for its power draw. Organizations with existing Turing-based software stacks for AI inference will find the T4G compatible, but those seeking leading-edge compute should look elsewhere. It is also an end-of-life product, so procurement should consider long-term availability and support.

Ray Tracing and Feature Set

The Tesla T4G includes 40 ray tracing cores and 320 tensor cores, reflecting its Turing architecture heritage. The presence of RT cores enables hardware-accelerated ray tracing, but the data does not provide specific ray tracing performance metrics, so its efficacy remains qualitative. Given the 8.141 TFLOPS FP32 rate, ray tracing workloads would likely be limited compared to higher-end Turing or newer cards, but the capability is present for applications that require it.

The tensor cores are a major feature, with FP16 performance reaching 65.13 TFLOPS at an 8:1 ratio. This indicates that the card is optimized for AI inference and training acceleration, where tensor core utilization is high. The API support includes DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, ensuring compatibility with modern graphics and compute APIs. DirectX 12 Ultimate support confirms feature-level 12_2, which includes variable rate shading and mesh shaders, though these are more relevant for gaming than server workloads.

The feature set is rounded out by PCIe 3.0 x16 connectivity, which is sufficient for data transfer but not as fast as newer PCIe 4.0 or 5.0 interfaces. The card has no display outputs, reinforcing its role as a compute-only accelerator. The 12 nm process and TSMC foundry are standard for the 2018 era, and the transistor density of 25.0M per mm² is a reflection of that node's capabilities. Overall, the T4G's feature set is targeted at AI and virtualized environments, with ray tracing as a secondary capability, and its API support ensures broad software compatibility.

Detailed benchmark scores and charts for the NVIDIA Tesla T4G are below.

Benchmark Scores

No benchmark data available for this GPU.

Compare with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs