NVIDIA T400 vs NVIDIA Tesla K40m Comparison

NVIDIA
GEFORCE

NVIDIA T400

CORE STATE TU117
VRAM 2 GB
CLOCK SPEED 1425 MHz
TDP 30 W
BUS WIDTH 64 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Tesla K40m

CORE STATE GK110B
VRAM 12 GB
CLOCK SPEED 876 MHz
TDP 245 W
BUS WIDTH 384 bit
ARCHITECTURE Kepler
nm
PROCESS 28 nm
LAUNCH DATE 2013

PERFORMANCE BENCHMARKS

geekbench_opencl
17,039
19,885
geekbench_vulkan
15,976
N/A

Analysis: NVIDIA T400 vs NVIDIA Tesla K40m

Head-to-Head Benchmarks

The only shared benchmark in the database is Geekbench OpenCL, and the recorded results show a clear edge for the NVIDIA Tesla K40m. The Tesla K40m scores 19,885 points against 17,039 for the NVIDIA T400, a delta of 16.7% in favor of the older card. This is a substantial margin for a compute-oriented workload, indicating that the Tesla K40m's raw execution resources still translate into a measurable advantage in OpenCL throughput even though it belongs to an earlier architecture generation.

Looking at the broader database context, the Tesla K40m's 19,885 score places it just 0.1% behind the AMD FirePro W7000, which averages 19,905. It also edges out the AMD Radeon RX 6650 XT (19,765) by 0.6%, the AMD FirePro D300 (19,637) by 1.3%, and the NVIDIA Quadro K5200 (19,602) by 1.4%. In percentile terms, the Tesla K40m sits at the 65th percentile among all GPUs in the database, which is a respectable standing for a product released much earlier.

The NVIDIA T400, by contrast, posts an average benchmark score of 16,508 across its two recorded tests: 17,039 in Geekbench OpenCL and 15,976 in Geekbench Vulkan. Its OpenCL result trails the Tesla K40m by the 16.7% margin noted above. In the percentile ranking, the T400 lands at the 60th percentile, five points below the Tesla K40m. Its nearest rivals in the database are remarkably close: the NVIDIA GeForce RTX 5090 D V2 (16,504, 0% delta), the AMD Radeon PRO W7500 (16,415, 0.6% higher), the NVIDIA RTX PRO 6000 Blackwell (16,408, 0.6% higher), and the AMD Radeon RX 5700 XT (16,361, 0.9% higher). This cluster of scores suggests the T400 is positioned in a dense performance band where small variations separate competitors.

The head-to-head data records one win for the Tesla K40m and none for the T400. That asymmetry is meaningful: in the only test where both cards were measured under identical conditions, the Tesla K40m outperformed the T400 by a double-digit percentage. The T400 does not win any recorded benchmark against the Tesla K40m, so any decision favoring the T400 must rest on other factors, such as feature support, power efficiency, or output capabilities, rather than raw compute performance in this specific workload.

Architecture Differences

The two cards come from fundamentally different architectural generations. The Tesla K40m uses the GK110B chip, built on the Kepler architecture, while the T400 uses the TU117 chip, based on the Turing architecture. The manufacturing process reflects this gap: the Tesla K40m is fabricated on a 28 nm process at TSMC, whereas the T400 uses a 12 nm process, also from TSMC. The newer node gives the T400 a much higher transistor density, 23.5 million transistors per square millimeter versus 12.6 million for the Tesla K40m, despite the Tesla chip having more total transistors.

The die sizes tell a similar story of generational scaling. The Tesla K40m's GK110B packs 7,080 million transistors on a 561 mm² die, making it a large, power-hungry chip. The T400's TU117 contains 4,700 million transistors on a 200 mm² die, which is far smaller while still holding a substantial transistor count. This difference in physical scale directly influences the cards' thermal and power profiles, which are starkly different.

The compute resources are heavily skewed toward the Tesla K40m. It features 2,880 shading units, 240 texture mapping units, and 48 raster output units. The T400 has only 384 shading units, 24 texture mapping units, and 16 raster output units. These are not close numbers: the Tesla K40m has 7.5 times the shading units and 10 times the texture units of the T400. Consequently, the theoretical pixel rate is 52.56 GPixel/s for the Tesla K40m versus 22.80 GPixel/s for the T400, and the texture rate is 210.2 GTexel/s versus 34.20 GTexel/s. The floating-point throughput follows the same pattern: the Tesla K40m delivers 5.046 TFLOPS of FP32, while the T400 manages 1,094.4 GFLOPS, which is roughly one-fifth of the older card's output.

Memory is another area of major divergence. The Tesla K40m has 12 GB of GDDR5 on a 384-bit bus, yielding 288.4 GB/s of bandwidth. The T400 has 2 GB of GDDR6 on a 64-bit bus, providing 80.00 GB/s. The memory type is newer on the T400 (GDDR6 versus GDDR5), and its effective speed is higher at 10 Gbps versus 6 Gbps, but the narrow bus and smaller capacity result in a bandwidth figure that is less than a third of the Tesla K40m's. The T400 does support FP16 at 2.189 TFLOPS with a 2:1 ratio, a capability the Tesla K40m lacks entirely, as its FP16 field is null in the database.

Clock behavior also differs markedly. The Tesla K40m runs at a base of 745 MHz and a boost of 876 MHz, while the T400 runs at a base of 420 MHz and a boost of 1425 MHz. The T400's boost clock is over 60% higher than the Tesla K40m's, which partially compensates for its far lower core count in some workloads, but not enough to overcome the massive shader advantage of the older card in the OpenCL test.

The Verdict

The data points to a clear split in intended use cases. For raw compute throughput in OpenCL, the NVIDIA Tesla K40m is the stronger card, with a 16.7% lead over the T400 in the only head-to-head benchmark recorded. Its much larger shader array, wider memory bus, and higher bandwidth make it the logical pick for workloads that stress parallel floating-point execution and large data sets. The fact that it sits at the 65th percentile versus the T400's 60th percentile reinforces this, even though both cards are near the middle of the database distribution.

The T400, however, has advantages that are not captured by the single shared benchmark. It is a Turing-generation card with a much newer feature set, including DirectX 12 (12_1) support, Vulkan 1.4, and three mini-DisplayPort 1.4a outputs. The Tesla K40m has no display outputs at all, making the T400 the only viable choice for any task that requires driving a monitor. The T400 also has a dramatically lower thermal design power of 30 W versus 245 W for the Tesla K40m, and it requires no power connectors, while the Tesla K40m's power requirements are not specified but its suggested PSU is 550 W versus 200 W for the T400.

For users who prioritize compute density and memory capacity, the Tesla K40m is the clear winner from the data. For users who need a low-power, display-capable card with modern API support and a smaller physical footprint, the T400 is the only sensible option. The T400's wins are qualitative: it is single-slot, passively compatible with a 200 W PSU, and offers outputs. The Tesla K40m's win is quantitative: it scores higher in the only benchmark where both were measured. Neither card is a universal recommendation, but the decision criteria are unambiguous from the recorded facts.

Specification Differences

  • Process node: 28 nm (Tesla K40m) versus 12 nm (T400)
  • Transistor count: 7,080 million (Tesla K40m) versus 4,700 million (T400)
  • Die size: 561 mm² (Tesla K40m) versus 200 mm² (T400)
  • Transistor density: 12.6M / mm² (Tesla K40m) versus 23.5M / mm² (T400)
  • Base clock: 745 MHz (Tesla K40m) versus 420 MHz (T400)
  • Boost clock: 876 MHz (Tesla K40m) versus 1425 MHz (T400)
  • Memory size: 12 GB (Tesla K40m) versus 2 GB (T400)
  • Memory type: GDDR5 (Tesla K40m) versus GDDR6 (T400)
  • Memory bus width: 384 bit (Tesla K40m) versus 64 bit (T400)
  • Memory bandwidth: 288.4 GB/s (Tesla K40m) versus 80.00 GB/s (T400)
  • Shading units: 2880 (Tesla K40m) versus 384 (T400)
  • TMUs: 240 (Tesla K40m) versus 24 (T400)
  • ROPs: 48 (Tesla K40m) versus 16 (T400)
  • Pixel rate: 52.56 GPixel/s (Tesla K40m) versus 22.80 GPixel/s (T400)
  • Texture rate: 210.2 GTexel/s (Tesla K40m) versus 34.20 GTexel/s (T400)
  • FP32: 5.046 TFLOPS (Tesla K40m) versus 1,094.4 GFLOPS (T400)
  • FP16: null (Tesla K40m) versus 2.189 TFLOPS (T400)
  • TDP: 245 W (Tesla K40m) versus 30 W (T400)
  • Slot width: Dual-slot (Tesla K40m) versus Single-slot (T400)
  • Power connectors: null (Tesla K40m) versus None (T400)
  • Suggested PSU: 550 W (Tesla K40m) versus 200 W (T400)
  • Display outputs: No outputs (Tesla K40m) versus 3x mini-DisplayPort 1.4a (T400)
  • DirectX support: 12 (11_1) (Tesla K40m) versus 12 (12_1) (T400)
  • Vulkan support: 1.2.175 (Tesla K40m) versus 1.4 (T400)
  • Release date: 2013-11-21 (Tesla K40m) versus 2021-05-05 (T400)

FAQ

Q: Which card has higher OpenCL performance?

A: The NVIDIA Tesla K40m scores 19,885 in Geekbench OpenCL, which is 16.7% higher than the NVIDIA T400's 17,039.

Q: Does the T400 support display outputs?

A: Yes, the T400 has 3x mini-DisplayPort 1.4a outputs. The Tesla K40m has no display outputs.

Q: What is the memory capacity difference?

A: The Tesla K40m has 12 GB of GDDR5 memory, while the T400 has 2 GB of GDDR6 memory.

Q: Which card has a lower power requirement?

A: The T400 has a TDP of 30 W and a suggested PSU of 200 W, whereas the Tesla K40m has a TDP of 245 W and a suggested PSU of 550 W.

Q: How do the cards compare in percentile ranking?

A: The Tesla K40m sits at the 65th percentile among all GPUs, while the T400 sits at the 60th percentile.

Q: What is the release date gap between them?

A: The Tesla K40m was released on 2013-11-21, and the T400 was released on 2021-05-05, a gap of roughly seven and a half years.

DETAILED SPECIFICATIONS

SPECIFICATION
T400
Tesla K40m
Core Specs
Shading Units
384
2,880 +650.0%
Shaders
384
2,880 +650.0%
TMUs
24
240 +900.0%
ROPs
16
48 +200.0%
SM Count
6
Clocks
Base Clock
420 MHz
745 MHz
Boost Clock
1425 MHz
876 MHz
Memory Clock
1250 MHz 10 Gbps effective
1502 MHz 6 Gbps effective
Memory
Memory Size
2 GB
12 GB
VRAM (MB)
2,048
12,288 +500.0%
Memory Type
GDDR6
GDDR5
Memory Bus
64 bit
384 bit
Bandwidth
80.00 GB/s
288.4 GB/s
Cache
L1 Cache
64 KB (per SM)
16 KB (per SMX)
L2 Cache
1024 KB
1536 KB
Performance
Pixel Rate
22.80 GPixel/s
52.56 GPixel/s
Texture Rate
34.20 GTexel/s
210.2 GTexel/s
FP32 (TFLOPS)
1,094.4 GFLOPS
5.046 TFLOPS
FP64 (TFLOPS)
34.20 GFLOPS (1:32)
1.682 TFLOPS (1:3)
FP16 (TFLOPS)
2.189 TFLOPS (2:1)
Power
TDP
30 W
245 W
TDP (W)
30
245 +716.7%
Suggested PSU
200 W
550 W
Power Connectors
None
Architecture
Architecture
Turing
Kepler
GPU Name
TU117
GK110B
Generation
Quadro Turing (Tx000)
Tesla Kepler (Kxx)
Process Size
12 nm
28 nm
Transistors
4,700 million
7,080 million
Die Size
200 mm²
561 mm²
Foundry
TSMC
TSMC
Density
23.5M / mm²
12.6M / mm²
API Support
DirectX
12 (12_1)
12 (11_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.2.175
OpenCL
3.0
3.0
CUDA
7.5
3.5
Shader Model
6.8
6.5 (5.1)
Physical
Slot Width
Single-slot
Dual-slot
Length
267 mm 10.5 inches
Outputs
3x mini-DisplayPort 1.4a
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 3.0 x16
Other
Launch Price
7,699 USD
Production
End-of-life
End-of-life
Predecessor
Quadro Volta
Tesla Fermi
Successor
Workstation Ampere
Tesla Maxwell
View T400 Details View Tesla K40m Details