NVIDIA T400 4 GB vs NVIDIA Tesla K40c Comparison

NVIDIA
GEFORCE

NVIDIA T400 4 GB

CORE STATE TU117
VRAM 4 GB
CLOCK SPEED 1425 MHz
TDP 30 W
BUS WIDTH 64 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Tesla K40c

CORE STATE GK180
VRAM 12 GB
CLOCK SPEED 876 MHz
TDP 245 W
BUS WIDTH 384 bit
ARCHITECTURE Kepler
nm
PROCESS 28 nm
LAUNCH DATE 2013

PERFORMANCE BENCHMARKS

geekbench_opencl
17,320
17,468
geekbench_vulkan
16,263
N/A

Analysis: NVIDIA T400 4 GB vs NVIDIA Tesla K40c

The NVIDIA Tesla K40c and NVIDIA T400 4 GB are both end-of-life workstation cards, but they represent radically different eras of GPU design. The K40c is a 2013-era compute behemoth built on Kepler, while the T400 is a 2021 Turing-based entry-level card. Benchmark data shows they are surprisingly close in raw compute scores, yet their specifications and intended use cases could not be more different.

FAQ

Q: Which card has the higher average benchmark score?

A: The NVIDIA Tesla K40c has an average benchmark score of 17,468, while the NVIDIA T400 4 GB averages 16,792. The K40c is about 4% higher on average, though the T400's average is pulled down by its second Vulkan test score of 16,263.

Q: How do the two cards compare in the Geekbench OpenCL test?

A: The Tesla K40c scores 17,468, edging out the T400's 17,320 by a margin of 0.9%. This is the only head-to-head benchmark recorded, and it goes to the K40c.

Q: What is the most significant architectural difference between the two?

A: The K40c uses the Kepler architecture on a 28 nm process with the GK180 chip, while the T400 uses the Turing architecture on a 12 nm process with the TU117 chip. The T400 also supports a higher DirectX version (12_1 vs 11_0 for the K40c).

Q: Which card has more memory bandwidth?

A: The Tesla K40c offers 288.4 GB/s of bandwidth from its 384-bit bus and 12 GB of GDDR5 memory. The T400 provides 80.00 GB/s over a 64-bit bus with 4 GB of GDDR6 memory.

Q: Are these cards comparable in power requirements?

A: No. The K40c has a 245 W TDP, requires dual-slot cooling, and needs both a 6-pin and 8-pin power connector. The T400 has a 30 W TDP, fits in a single slot, and requires no power connectors at all.

Q: Which card has better API support?

A: The T400 supports newer APIs, including DirectX 12 (12_1) and Vulkan 1.4, while the K40c supports DirectX 12 (11_0) and Vulkan 1.2.175. Both support OpenGL 4.6.

The Verdict

The data paints a clear picture of two cards built for different jobs. The Tesla K40c is the compute-oriented option: it wins the only head-to-head benchmark, offers 12 GB of memory, and delivers 5.046 TFLOPS of FP32 performance. Its 61st percentile ranking among all GPUs places it just ahead of the T400's 60th percentile.

The T400 4 GB is the practical workstation card. It consumes only 30 W versus 245 W, requires no external power connectors, and fits in a single slot. It also provides display outputs (3x mini-DisplayPort 1.4a) where the K40c has none. For users who need a simple, low-power card with modern API support, the T400 is the logical choice.

For raw compute throughput, the K40c still holds a slight edge. Its 5.046 TFLOPS FP32 figure dwarfs the T400's 1,094.4 GFLOPS, and its 288.4 GB/s memory bandwidth is 3.6x higher. However, the T400 counters with FP16 support at 2.189 TFLOPS (2:1), a feature the K40c lacks entirely. Users needing modern display outputs or low power draw should pick the T400; those prioritizing raw FP32 compute and large memory pools should pick the K40c.

Head-to-Head Benchmarks

The only recorded head-to-head benchmark is the Geekbench OpenCL test, and it is remarkably close. The Tesla K40c scores 17,468 against the T400's 17,320, a delta of 0.9% in favor of the older card. This narrow margin is surprising given the massive specification differences between the two.

Looking at the nearest rivals provides context. The K40c's closest competitor is the AMD Radeon Pro 460 at 17,509 (0.2% faster), followed by the AMD Radeon Pro 560 at 17,551 (0.5% faster), the AMD Radeon 780M at 17,588 (0.7% faster), and the NVIDIA GeForce RTX 4060 at 17,639 (1% faster). The K40c trails all of these by less than a single percentage point.

The T400's OpenCL score of 17,320 places it near the AMD Radeon RX 7600S, which scores 16,696 (the T400 is 0.6% ahead). The NVIDIA Tesla M4 scores 16,932 (0.8% behind the T400), the AMD Radeon HD 7970M scores 17,019 (1.3% behind), and the NVIDIA GeForce GTX 690 scores 17,037 (1.4% behind). The T400 also has a Vulkan score of 16,263, which is not directly compared against the K40c since the K40c has no recorded Vulkan benchmark.

The K40c wins the head-to-head tally with 1 win to 0. That said, the 0.9% delta is within the margin of benchmark noise. Real-world differences in this score range are unlikely to be perceptible in most workloads.

Specification Differences

The memory subsystems are the most dramatic differentiator. The K40c packs 12 GB of GDDR5 on a 384-bit bus, delivering 288.4 GB/s of bandwidth. The T400 offers 4 GB of GDDR6 on a 64-bit bus, with 80.00 GB/s bandwidth. The K40c's bandwidth advantage is more than 3.6x, and its memory capacity is 3x larger.

Core counts follow a similar pattern. The K40c features 2,880 shading units, 240 texture mapping units, and 48 ROPs. The T400 has 384 shading units, 24 TMUs, and 16 ROPs. Pixel and texture rates reflect this: the K40c outputs 52.56 GPixel/s and 210.2 GTexel/s, while the T400 manages 22.80 GPixel/s and 34.20 GTexel/s.

Power and physical design diverge sharply. The K40c draws 245 W, spans dual slots, and requires both a 6-pin and 8-pin power connector, with a suggested 550 W PSU. The T400 draws just 30 W, occupies a single slot, needs no power connectors, and works with a 200 W suggested PSU. The K40c measures 267 mm (10.5 inches) in length; the T400's dimensions are not listed.

The K40c has no display outputs, making it a pure compute accelerator. The T400 includes 3x mini-DisplayPort 1.4a outputs. Release dates are eight years apart: the K40c launched on October 7, 2013, while the T400 launched on May 5, 2021. The K40c has a launch MSRP of 7,699 USD; the T400 has no listed launch MSRP.

Architecture Differences

The two cards come from completely different architectural generations. The K40c is built on Kepler, using the GK180 chip fabricated by TSMC on a 28 nm process. It contains 7,080 million transistors on a 561 mm² die, yielding a transistor density of 12.6M per mm². The architecture supports DirectX 12 (11_0) and Vulkan 1.2.175.

The T400 uses the Turing architecture with the TU117 chip, also from TSMC but on a 12 nm process. It contains 4,700 million transistors on a 200 mm² die, achieving a higher density of 23.5M per mm². Turing brings newer API support: DirectX 12 (12_1) and Vulkan 1.4.

The K40c's generational context is Tesla Kepler (Kxx), succeeding Tesla Fermi and preceding Tesla Maxwell. The T400 belongs to the Quadro Turing (Tx000) family, succeeding Quadro Volta and preceding Workstation Ampere.

Neither card features ray tracing cores or tensor cores. The K40c's FP32 throughput is 5.046 TFLOPS, while the T400 delivers 1,094.4 GFLOPS. The T400 adds FP16 capability at 2.189 TFLOPS (2:1), which the K40c does not offer. Clock behavior also differs: the K40c runs at a 745 MHz base with an 876 MHz boost, while the T400 has a much lower 420 MHz base but a 1425 MHz boost. Memory clocks show 1502 MHz (6 Gbps effective) for the K40c and 1250 MHz (10 Gbps effective) for the T400.

DETAILED SPECIFICATIONS

SPECIFICATION
T400 4 GB
Tesla K40c
Core Specs
Shading Units
384
2,880 +650.0%
Shaders
384
2,880 +650.0%
TMUs
24
240 +900.0%
ROPs
16
48 +200.0%
SM Count
6
—
Clocks
Base Clock
420 MHz
745 MHz
Boost Clock
1425 MHz
876 MHz
Memory Clock
1250 MHz 10 Gbps effective
1502 MHz 6 Gbps effective
Memory
Memory Size
4 GB
12 GB
VRAM (MB)
4,096
12,288 +200.0%
Memory Type
GDDR6
GDDR5
Memory Bus
64 bit
384 bit
Bandwidth
80.00 GB/s
288.4 GB/s
Cache
L1 Cache
64 KB (per SM)
16 KB (per SMX)
L2 Cache
1024 KB
1536 KB
Performance
Pixel Rate
22.80 GPixel/s
52.56 GPixel/s
Texture Rate
34.20 GTexel/s
210.2 GTexel/s
FP32 (TFLOPS)
1,094.4 GFLOPS
5.046 TFLOPS
FP64 (TFLOPS)
34.20 GFLOPS (1:32)
1.682 TFLOPS (1:3)
FP16 (TFLOPS)
2.189 TFLOPS (2:1)
—
Power
TDP
30 W
245 W
TDP (W)
30
245 +716.7%
Suggested PSU
200 W
550 W
Power Connectors
None
1x 6-pin + 1x 8-pin
Architecture
Architecture
Turing
Kepler
GPU Name
TU117
GK180
Generation
Quadro Turing (Tx000)
Tesla Kepler (Kxx)
Process Size
12 nm
28 nm
Transistors
4,700 million
7,080 million
Die Size
200 mm²
561 mm²
Foundry
TSMC
TSMC
Density
23.5M / mm²
12.6M / mm²
API Support
DirectX
12 (12_1)
12 (11_0)
OpenGL
4.6
4.6
Vulkan
1.4
1.2.175
OpenCL
3.0
3.0
CUDA
7.5
3.5
Shader Model
6.8
5.1
Physical
Slot Width
Single-slot
Dual-slot
Length
—
267 mm 10.5 inches
Outputs
3x mini-DisplayPort 1.4a
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 3.0 x16
Other
Launch Price
—
7,699 USD
Production
End-of-life
End-of-life
Predecessor
Quadro Volta
Tesla Fermi
Successor
Workstation Ampere
Tesla Maxwell
View T400 4 GB Details View Tesla K40c Details