GPU Comparison

NVIDIA
GEFORCE

NVIDIA GeForce RTX 4090

CORE STATE AD102
VRAM 24 GB
CLOCK SPEED 2520 MHz
TDP 450 W
BUS WIDTH 384 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2022
VS
NVIDIA
GEFORCE

Tesla T4

CORE STATE TU104
VRAM 16 GB
CLOCK SPEED 1590 MHz
TDP 70 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2018

PERFORMANCE BENCHMARKS

3dmark_3dmark_steel_nomad_dx12
9,223
N/A
geekbench_opencl
255,416
61,276
geekbench_vulkan
271,631
72,190
passmark_directx_10
224
N/A
passmark_directx_11
326
N/A
passmark_directx_12
150
N/A
passmark_directx_9
397
N/A
passmark_g2d
1,299
N/A
passmark_g3d
38,194
N/A
passmark_gpu_compute
26,613
N/A

Analysis: NVIDIA GeForce RTX 4090 vs NVIDIA Tesla T4

The NVIDIA Tesla T4 and NVIDIA GeForce RTX 4090 represent two vastly different design philosophies from the same manufacturer. The T4 is a low-power, server-oriented accelerator built for broad compatibility and efficiency, while the RTX 4090 is a consumer flagship engineered for maximum throughput. Benchmark data places both at the 91st percentile among all GPUs, yet their average scores are nearly identical: the Tesla T4 averages 66,733, while the RTX 4090 averages 66,473, a difference of just 0.4%. This statistical tie masks profound architectural and performance disparities that surface in specific workloads.

FAQ

Q: How do the average benchmark scores compare between the Tesla T4 and RTX 4090?

A: The Tesla T4 has an average benchmark score of 66,733, while the RTX 4090 scores 66,473. The RTX 4090 trails by 0.4% in this aggregate metric, placing both cards at the 91st percentile of all GPUs.

Q: Which card wins in Geekbench OpenCL performance, and by how much?

A: The RTX 4090 wins decisively with a score of 317,684 versus the Tesla T4's 61,276. This represents an 80.7% advantage for the RTX 4090 in that test.

Q: What is the difference in memory bandwidth between the two cards?

A: The Tesla T4 provides 320.0 GB/s of bandwidth, while the RTX 4090 delivers 1.01 TB/s. The RTX 4090's bandwidth is more than three times higher.

Q: Are both cards based on the same transistor count?

A: No. The Tesla T4 contains 13,600 million transistors on a 545 mm² die, whereas the RTX 4090 packs 76,300 million transistors into a 609 mm² die. The RTX 4090 achieves a transistor density of 125.3M / mm² versus 25.0M / mm² for the T4.

Q: What are the thermal design power ratings for each card?

A: The Tesla T4 has a TDP of 70 W, while the RTX 4090 has a TDP of 450 W. The T4 also requires no power connectors and a 250 W suggested PSU, compared to the RTX 4090's 16-pin connector and 850 W suggested PSU.

Q: Do both GPUs support the same DirectX and Vulkan versions?

A: Yes, both support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, despite their different architectures and release timelines.

Architecture Differences

The Tesla T4 is built on the Turing architecture using a TU104 chip fabricated on TSMC's 12 nm process. The RTX 4090 employs the Ada Lovelace architecture with an AD102 chip on a 5 nm process, also from TSMC. This node transition enables the RTX 4090 to house 76,300 million transistors versus the T4's 13,600 million, despite only a modest die size increase from 545 mm² to 609 mm². Transistor density jumps from 25.0M / mm² on the T4 to 125.3M / mm² on the RTX 4090, a fivefold improvement that underpins the latter's massive compute advantage.

The compute resources differ by an order of magnitude. The Tesla T4 has 2,560 shading units, 160 texture mapping units, and 64 ROPs, while the RTX 4090 scales these to 16,384 shading units, 512 TMUs, and 176 ROPs. Ray tracing hardware follows a similar pattern: the T4 includes 40 RT cores, whereas the RTX 4090 features 128. Tensor core counts are 320 and 512, respectively. Clock speeds also diverge significantly, with the T4's base clock at 585 MHz and boost at 1590 MHz, against the RTX 4090's 2235 MHz base and 2520 MHz boost.

Memory subsystems are equally distinct. The T4 uses 16 GB of GDDR6 on a 256-bit bus, achieving 320.0 GB/s bandwidth. The RTX 4090 pairs 24 GB of GDDR6X with a 384-bit bus for 1.01 TB/s throughput. The T4's memory operates at 1250 MHz (10 Gbps effective), while the RTX 4090's runs at 1313 MHz (21 Gbps effective). The T4's FP16 performance of 65.13 TFLOPS is achieved through an 8:1 ratio relative to FP32, whereas the RTX 4090 delivers 82.58 TFLOPS in both FP16 and FP32 with a 1:1 ratio.

Head-to-Head Benchmarks

The shared benchmark suite between these two cards is limited to two Geekbench tests, and the RTX 4090 dominates both. In Geekbench OpenCL, the RTX 4090 scores 317,684 against the Tesla T4's 61,276, yielding an 80.7% deficit for the T4. This gap reflects the RTX 4090's 82.58 TFLOPS FP32 throughput versus the T4's 8.141 TFLOPS, a tenfold raw compute advantage that translates directly into the benchmark result.

Geekbench Vulkan tells a similar story, with the RTX 4090 posting 270,615 versus the T4's 72,190. The delta is 73.3% in favor of the RTX 4090. While slightly narrower than the OpenCL margin, this still represents a substantial performance gap. The RTX 4090 wins both head-to-head tests, giving it a 2-0 record in the shared benchmark set.

The aggregate average scores, however, tell a more nuanced story. The Tesla T4's average of 66,733 edges out the RTX 4090's 66,473 by 0.4%, a margin within normal run-to-run variance. This near-parity in the overall metric arises because the T4's benchmark portfolio includes only Geekbench results, while the RTX 4090's includes a broader set of tests such as Passmark and 3DMark, which pull its average downward. The RTX 4090's Passmark G3D score of 38,194 and GPU compute score of 26,613 are high in absolute terms but dilute the average when combined with lower DirectX-specific scores like the Passmark DirectX 9 result of 397.

Specification Differences

| Specification | NVIDIA Tesla T4 | NVIDIA GeForce RTX 4090 |

|---|---|---|

| Architecture | Turing | Ada Lovelace |

| Process Node | 12 nm | 5 nm |

| Transistors | 13,600 million | 76,300 million |

| Die Size | 545 mm² | 609 mm² |

| Transistor Density | 25.0M / mm² | 125.3M / mm² |

| Base Clock | 585 MHz | 2235 MHz |

| Boost Clock | 1590 MHz | 2520 MHz |

| Memory Size | 16 GB | 24 GB |

| Memory Type | GDDR6 | GDDR6X |

| Memory Bus | 256 bit | 384 bit |

| Memory Bandwidth | 320.0 GB/s | 1.01 TB/s |

| Memory Clock | 1250 MHz (10 Gbps effective) | 1313 MHz (21 Gbps effective) |

| Shading Units | 2560 | 16384 |

| TMUs | 160 | 512 |

| ROPs | 64 | 176 |

| RT Cores | 40 | 128 |

| Tensor Cores | 320 | 512 |

| Pixel Rate | 101.8 GPixel/s | 443.5 GPixel/s |

| Texture Rate | 254.4 GTexel/s | 1,290.2 GTexel/s |

| FP32 | 8.141 TFLOPS | 82.58 TFLOPS |

| FP16 | 65.13 TFLOPS (8:1) | 82.58 TFLOPS (1:1) |

| TDP | 70 W | 450 W |

| Slot Width | Single-slot | Triple-slot |

| Power Connectors | None | 1x 16-pin |

| Suggested PSU | 250 W | 850 W |

| Bus Interface | PCIe 3.0 x16 | PCIe 4.0 x16 |

| Display Outputs | No outputs | 1x HDMI 2.13x DisplayPort 1.4a |

| Length | 168 mm (6.6 inches) | 304 mm (12 inches) |

| Height | N/A | 137 mm (5.4 inches) |

| Width | N/A | 61 mm (2.4 inches) |

| Release Date | 2018-09-12 | 2022-09-19 |

| Predecessor | Tesla Volta | GeForce 30 |

| Successor | Server Ampere | GeForce 50 |

| Launch MSRP | N/A | 1,599 USD |

Where Each One Wins

The RTX 4090 wins in every measurable performance category where both cards have data. Its Geekbench OpenCL score of 317,684 is 80.7% higher than the T4's 61,276, and its Vulkan score of 270,615 beats the T4's 72,190 by 73.3%. The RTX 4090 also leads in theoretical throughput metrics: 82.58 TFLOPS FP32 versus 8.141 TFLOPS, 1.01 TB/s bandwidth versus 320.0 GB/s, and 443.5 GPixel/s pixel rate versus 101.8 GPixel/s. Its 24 GB memory capacity and 512 tensor cores make it the clear choice for compute-heavy workloads that fit within its 450 W power envelope.

The Tesla T4's advantages lie outside raw performance. Its 70 W TDP allows operation without any power connectors, and its single-slot, 168 mm length makes it suitable for dense server installations where space and power are constrained. The T4's 16 GB GDDR6 memory and 320 tensor cores still provide meaningful compute capability, and its support for DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4 matches the RTX 4090's API coverage. The T4 also carries a higher average benchmark score of 66,733 versus 66,473, though this 0.4% edge stems from differing test portfolios rather than superior performance.

For inference workloads in power-constrained environments, the T4's efficiency profile is compelling: 8.141 TFLOPS FP32 from 70 W. For scenarios demanding maximum throughput, the RTX 4090's tenfold FP32 advantage and greater memory bandwidth make it the superior instrument. The data does not present a single winner; it presents two tools optimized for different constraints. The RTX 4090 is the performance king, while the T4 wins on operational flexibility and power economy.

DETAILED SPECIFICATIONS

SPECIFICATION
RTX 4090
Tesla T4
Core Specs
Shading Units
16,384
2,560 -84.4%
Shaders
16,384
2,560 -84.4%
TMUs
512
160 -68.8%
ROPs
176
64 -63.6%
SM Count
128
40 -68.8%
Clocks
Base Clock
2235 MHz
585 MHz
Boost Clock
2520 MHz
1590 MHz
Memory Clock
1313 MHz 21 Gbps effective
1250 MHz 10 Gbps effective
Memory
Memory Size
24 GB
16 GB
VRAM (MB)
24,576
16,384 -33.3%
Memory Type
GDDR6X
GDDR6
Memory Bus
384 bit
256 bit
Bandwidth
1.01 TB/s
320.0 GB/s
Cache
L1 Cache
128 KB (per SM)
64 KB (per SM)
L2 Cache
72 MB
4 MB
Performance
Pixel Rate
443.5 GPixel/s
101.8 GPixel/s
Texture Rate
1,290.2 GTexel/s
254.4 GTexel/s
FP32 (TFLOPS)
82.58 TFLOPS
8.141 TFLOPS
FP64 (TFLOPS)
1,290.2 GFLOPS (1:64)
254.4 GFLOPS (1:32)
FP16 (TFLOPS)
82.58 TFLOPS (1:1)
16.28 TFLOPS (2:1)
AI/RT
RT Cores
128
40 -68.8%
Tensor Cores
512
320 -37.5%
Power
TDP
450 W
70 W
TDP (W)
450
70 -84.4%
Suggested PSU
850 W
250 W
Power Connectors
1x 16-pin
None
Architecture
Architecture
Ada Lovelace
Turing
GPU Name
AD102
TU104
Generation
GeForce 40
Tesla Turing (Txx)
Process Size
5 nm
12 nm
Transistors
76,300 million
13,600 million
Die Size
609 mm²
545 mm²
Foundry
TSMC
TSMC
Density
125.3M / mm²
25.0M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.9
7.5
Shader Model
6.8
6.9
Physical
Slot Width
Triple-slot
Single-slot
Length
304 mm 12 inches
168 mm 6.6 inches
Height
137 mm 5.4 inches
Outputs
1x HDMI 2.13x DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 3.0 x16
Other
Launch Price
1,599 USD
Production
End-of-life
End-of-life
Predecessor
GeForce 30
Tesla Volta
Successor
GeForce 50
Server Ampere
View GeForce RTX 4090 Details View Tesla T4 Details