AMD Radeon RX 9060 XT LP vs NVIDIA Tesla T4 Comparison

AMD
RADEON

AMD Radeon RX 9060 XT LP

CORE STATE Navi 44
VRAM 16 GB
CLOCK SPEED 3050 MHz
TDP 140 W
BUS WIDTH 128 bit
ARCHITECTURE RDNA 4.0
nm
PROCESS 4 nm
LAUNCH DATE 2025
VS
NVIDIA
GEFORCE

Tesla T4

CORE STATE TU104
VRAM 16 GB
CLOCK SPEED 1590 MHz
TDP 70 W
BUS WIDTH 256 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2018

PERFORMANCE BENCHMARKS

geekbench_opencl
88,183
61,276
geekbench_vulkan
39,476
72,190

Analysis: AMD Radeon RX 9060 XT LP vs NVIDIA Tesla T4

NVIDIA Tesla T4 and AMD Radeon RX 9060 XT LP are two GPUs with almost nothing in common except their 16 GB memory capacity. The Tesla T4 is an end-of-life, single-slot server accelerator from 2018, while the RX 9060 XT LP is an active, dual-slot consumer graphics card from 2025. The benchmark data shows a perfect split: each card wins one of the two available tests decisively. The T4 takes the Geekbench Vulkan test by 82.9%, while the RX 9060 XT LP crushes the OpenCL test by 30.5%. The average benchmark scores, however, tell a different story, with the T4 averaging 66,733 versus 63,830 for the AMD card, placing them at the 90th and 89th percentiles of all GPUs respectively.

Where Each One Wins

The NVIDIA Tesla T4 is the clear winner in Vulkan workloads. Its score of 72,190 in Geekbench Vulkan is not just a small margin — it is 82.9% higher than the RX 9060 XT LP’s score of 39,476. This is a massive gap that suggests the T4’s Turing architecture has significantly better driver optimization or hardware scheduling for Vulkan’s low-level API. The T4 also holds a higher average benchmark score (66,733) than its rival (63,830), which pushes it to the 90th percentile versus the AMD card’s 89th. In terms of nearest rivals, the T4 sits within a tight band: it is 1.1% above the AMD Radeon VII, 2.5% above the NVIDIA Tesla P40, and only 2.7% below the AMD Radeon Instinct MI25.

The AMD Radeon RX 9060 XT LP dominates in OpenCL compute. Its Geekbench OpenCL score of 88,183 dwarfs the T4’s 61,276, representing a 30.5% advantage. This is consistent with the card’s raw specifications: it delivers 24.99 TFLOPS of FP32 performance versus the T4’s 8.141 TFLOPS, a three-fold difference in theoretical compute throughput. The RX 9060 XT LP also wins on texture and pixel rates, with 390.4 GTexel/s and 195.2 GPixel/s respectively, compared to the T4’s 254.4 GTexel/s and 101.8 GPixel/s. However, the AMD card’s average benchmark score is dragged down by its poor Vulkan showing, landing it at the 89th percentile with an average of 63,830 — nearly identical to its closest rival, the NVIDIA CMP 30HX (63,842, 0% delta).

Architecture Differences

The two cards represent completely different design philosophies separated by seven years of silicon evolution. The NVIDIA Tesla T4 uses the TU104 chip on a 12 nm TSMC process, packing 13,600 million transistors into a 545 mm² die. This yields a transistor density of just 25.0 million per mm². The T4 is built on the Turing architecture, which includes 40 RT cores for ray tracing and 320 tensor cores for AI acceleration. Its FP16 performance of 16.28 TFLOPS is exactly double its FP32 rate (2:1 ratio), indicating dedicated tensor core support for mixed-precision workloads.

The AMD Radeon RX 9060 XT LP uses the Navi 44 chip on a 4 nm TSMC process — a much more modern node. It packs 29,700 million transistors into a 199 mm² die, achieving a stellar transistor density of 149.2 million per mm², nearly six times higher than the T4. The RDNA 4.0 architecture dispenses with tensor cores entirely, instead focusing on raw shader throughput. Its FP16 performance of 24.99 TFLOPS equals its FP32 rate (1:1 ratio), meaning no dedicated mixed-precision hardware. The AMD card has 32 RT cores, 2,048 shading units, and 128 TMUs, while the T4 has 2,560 shading units, 160 TMUs, and 64 ROPs. Both have 64 ROPs, but the RX 9060 XT LP’s higher clocks (3050 MHz boost versus 1590 MHz) give it the edge in fill rates.

Head-to-Head Benchmarks

The Geekbench results reveal two entirely different performance profiles. In OpenCL, the RX 9060 XT LP wins with a score of 88,183 against the T4’s 61,276, a delta of -30.5% from the AMD card’s perspective. This is a straightforward compute test, and the AMD card’s 24.99 TFLOPS FP32 simply overwhelms the T4’s 8.141 TFLOPS. The RX 9060 XT LP also has nearly identical memory bandwidth (322.3 GB/s versus 320.0 GB/s) despite a narrower 128-bit bus, thanks to faster GDDR6 memory running at 20.1 Gbps effective versus 10 Gbps on the T4.

In Vulkan, the tables turn completely. The Tesla T4 scores 72,190 versus the RX 9060 XT LP’s 39,476, a 82.9% advantage for NVIDIA. This is surprising given the T4’s much lower raw compute figures. The gap suggests that either the T4’s Turing architecture handles Vulkan’s command buffers far more efficiently, or the RX 9060 XT LP’s drivers are not yet optimized for this API. The T4’s 40 RT cores and 320 tensor cores may also play a role in graphics workloads that leverage these units, though the benchmark does not specify. Notably, the T4’s Vulkan score is higher than its OpenCL score (72,190 versus 61,276), while the AMD card shows the opposite pattern (39,476 versus 88,183), indicating a fundamental architectural preference.

Specification Differences

| Specification | NVIDIA Tesla T4 | AMD Radeon RX 9060 XT LP |

|---|---|---|

| Process Node | 12 nm | 4 nm |

| Transistors | 13,600 million | 29,700 million |

| Die Size | 545 mm² | 199 mm² |

| Transistor Density | 25.0M / mm² | 149.2M / mm² |

| Base Clock | 585 MHz | 1380 MHz |

| Boost Clock | 1590 MHz | 3050 MHz |

| Memory Clock | 1250 MHz (10 Gbps) | 2518 MHz (20.1 Gbps) |

| Memory Bus Width | 256 bit | 128 bit |

| Shading Units | 2560 | 2048 |

| TMUs | 160 | 128 |

| RT Cores | 40 | 32 |

| Tensor Cores | 320 | None |

| FP32 Performance | 8.141 TFLOPS | 24.99 TFLOPS |

| FP16 Performance | 16.28 TFLOPS (2:1) | 24.99 TFLOPS (1:1) |

| Pixel Rate | 101.8 GPixel/s | 195.2 GPixel/s |

| Texture Rate | 254.4 GTexel/s | 390.4 GTexel/s |

| TDP | 70 W | 140 W |

| Slot Width | Single-slot | Dual-slot |

| Power Connectors | None | 1x 8-pin |

| Suggested PSU | 250 W | 300 W |

| Bus Interface | PCIe 3.0 x16 | PCIe 5.0 x16 |

| Display Outputs | No outputs | 1x HDMI 2.1b, 2x DisplayPort 2.1a |

| Release Date | 2018-09-12 | 2025-12-16 |

| Production Status | End-of-life | Active |

FAQ

Q: Which card is faster in OpenCL compute?

A: The AMD Radeon RX 9060 XT LP is significantly faster, scoring 88,183 versus 61,276 in Geekbench OpenCL, a 30.5% advantage. This aligns with its 24.99 TFLOPS FP32 performance versus the T4’s 8.141 TFLOPS.

Q: Why does the NVIDIA Tesla T4 win the Vulkan test by such a large margin?

A: The T4 scores 72,190 in Geekbench Vulkan versus 39,476 for the RX 9060 XT LP, an 82.9% difference. This is despite the T4 having lower raw compute, suggesting a significant architectural or driver advantage for NVIDIA in Vulkan workloads. The T4’s 320 tensor cores and 40 RT cores may contribute to this result.

Q: What is the average benchmark score for each card?

A: The NVIDIA Tesla T4 averages 66,733 across its benchmarks, placing it in the 90th percentile of all GPUs. The AMD Radeon RX 9060 XT LP averages 63,830, placing it in the 89th percentile.

Q: How do the memory systems compare?

A: Both cards have 16 GB of GDDR6 memory. The T4 uses a 256-bit bus with 320.0 GB/s bandwidth, while the RX 9060 XT LP uses a 128-bit bus but achieves 322.3 GB/s due to faster 20.1 Gbps effective memory clocks versus 10 Gbps on the T4.

Q: Which card has better raw fill rates?

A: The AMD Radeon RX 9060 XT LP has nearly double the pixel rate (195.2 GPixel/s versus 101.8 GPixel/s) and a 53% higher texture rate (390.4 GTexel/s versus 254.4 GTexel/s), driven by its much higher boost clock of 3050 MHz versus 1590 MHz.

Q: What are the power and physical differences?

A: The T4 is a 70 W single-slot card with no power connectors and no display outputs, designed for server use. The RX 9060 XT LP is a 140 W dual-slot card with a single 8-pin connector and full display outputs (1x HDMI 2.1b, 2x DisplayPort 2.1a). The T4 suggests a 250 W PSU, while the AMD card suggests 300 W.

The Verdict

The data presents a clear use-case split. The NVIDIA Tesla T4 is the choice for Vulkan-based graphics workloads and applications that can leverage its tensor cores and RT cores. Its 82.9% Vulkan advantage over the RX 9060 XT LP is decisive, and its higher average score (66,733 versus 63,830) and 90th percentile ranking make it the better overall performer in mixed workloads. Its single-slot, 70 W design with no power connectors is ideal for dense server deployments where space and power are constrained. The card is end-of-life, but its performance profile remains competitive in specific niches.

The AMD Radeon RX 9060 XT LP is the clear winner for OpenCL compute tasks, delivering a 30.5% higher score (88,183 versus 61,276) thanks to its 24.99 TFLOPS of FP32 throughput. It is also a much more modern part with a 4 nm process, PCIe 5.0 interface, and active production status. Its dual-slot design with display outputs makes it suitable for desktop use, and its 322.3 GB/s memory bandwidth slightly exceeds the T4’s. However, its poor Vulkan performance (39,476) pulls its average down to 63,830, just below the T4’s average, and its nearest rival is the NVIDIA CMP 30HX with essentially identical performance.

Choose the Tesla T4 if your workload is Vulkan-heavy or requires low-power, single-slot server acceleration with ray tracing and tensor core support. Choose the RX 9060 XT LP if you need maximum OpenCL compute throughput, modern connectivity, and an active product with display outputs. The benchmark data shows no universal winner — the correct choice depends entirely on whether your application prefers Vulkan or OpenCL.

DETAILED SPECIFICATIONS

SPECIFICATION
RX 9060 XT LP
Tesla T4
Core Specs
Shading Units
2,048
2,560 +25.0%
Shaders
2,048
2,560 +25.0%
TMUs
128
160 +25.0%
ROPs
64
64 0.0%
Compute Units
32
—
SM Count
—
40
Clocks
Base Clock
1380 MHz
585 MHz
Boost Clock
3050 MHz
1590 MHz
Game Clock
2450 MHz
—
Memory Clock
2518 MHz 20.1 Gbps effective
1250 MHz 10 Gbps effective
Memory
Memory Size
16 GB
16 GB
VRAM (MB)
16,384
16,384 0.0%
Memory Type
GDDR6
GDDR6
Memory Bus
128 bit
256 bit
Bandwidth
322.3 GB/s
320.0 GB/s
Cache
L1 Cache
—
64 KB (per SM)
L2 Cache
4 MB
4 MB
L3 Cache
32 MB
—
L0 Cache
32 KB per WGP
—
Performance
Pixel Rate
195.2 GPixel/s
101.8 GPixel/s
Texture Rate
390.4 GTexel/s
254.4 GTexel/s
FP32 (TFLOPS)
24.99 TFLOPS
8.141 TFLOPS
FP64 (TFLOPS)
780.8 GFLOPS (1:32)
254.4 GFLOPS (1:32)
FP16 (TFLOPS)
24.99 TFLOPS (1:1)
16.28 TFLOPS (2:1)
AI/RT
RT Cores
32
40 +25.0%
Tensor Cores
—
320
Matrix Cores
64
—
Power
TDP
140 W
70 W
TDP (W)
140
70 -50.0%
Suggested PSU
300 W
250 W
Power Connectors
1x 8-pin
None
Architecture
Architecture
RDNA 4.0
Turing
GPU Name
Navi 44
TU104
Generation
Navi IV (RX 9000)
Tesla Turing (Txx)
Process Size
4 nm
12 nm
Transistors
29,700 million
13,600 million
Die Size
199 mm²
545 mm²
Foundry
TSMC
TSMC
Density
149.2M / mm²
25.0M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.2
3.0
CUDA
—
7.5
Shader Model
6.9
6.9
Physical
Slot Width
Dual-slot
Single-slot
Length
—
168 mm 6.6 inches
Outputs
1x HDMI 2.1b2x DisplayPort 2.1a
No outputs
Bus Interface
PCIe 5.0 x16
PCIe 3.0 x16
Other
Production
Active
End-of-life
Predecessor
Navi III
Tesla Volta
Successor
—
Server Ampere
View Radeon RX 9060 XT LP Details View Tesla T4 Details