NVIDIA T400 4 GB vs NVIDIA Tesla K40m Comparison

NVIDIA
GEFORCE

NVIDIA T400 4 GB

CORE STATE TU117
VRAM 4 GB
CLOCK SPEED 1425 MHz
TDP 30 W
BUS WIDTH 64 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Tesla K40m

CORE STATE GK110B
VRAM 12 GB
CLOCK SPEED 876 MHz
TDP 245 W
BUS WIDTH 384 bit
ARCHITECTURE Kepler
nm
PROCESS 28 nm
LAUNCH DATE 2013

PERFORMANCE BENCHMARKS

geekbench_opencl
17,320
19,885
geekbench_vulkan
16,263
N/A

Analysis: NVIDIA T400 4 GB vs NVIDIA Tesla K40m

Head-to-Head Benchmarks

The recorded head-to-head data contains a single OpenCL benchmark, and the NVIDIA Tesla K40m wins that test outright. In Geekbench OpenCL, the Tesla K40m scores 19,885 points against 17,320 points for the NVIDIA T400 4 GB, a delta of 14.8% in favor of the older card. This is the only direct comparison available, so the Tesla K40m holds a 1-0 win record in the head-to-head section. The T400 4 GB does not register any wins in the direct comparison data.

The OpenCL result aligns with the overall average benchmark scores in the database. The Tesla K40m carries an average benchmark score of 19,885, while the T400 4 GB averages 16,792 across its two recorded tests, which include a Vulkan score of 16,263 alongside the OpenCL result. The gap in the single shared test is substantial, and it reflects the raw compute resources each card brings to the table.

Context from the nearest rivals section reinforces the positioning. The Tesla K40m sits at the 65th percentile among all GPUs, with its closest competitor being the AMD FirePro W7000 at a score of 19,905, a delta of -0.1%. That means the K40m is essentially tied with that rival. The T400 4 GB sits at the 60th percentile, and its nearest rival is the NVIDIA Tesla M4 at 16,932, a delta of -0.8%. The T400 trails its own peer group slightly, whereas the K40m effectively matches its top rival.

Where Each One Wins

The Tesla K40m wins the compute-heavy OpenCL test by a wide margin, and that is its clear strength. The data shows a 14.8% advantage over the T400 4 GB in that specific workload. The K40m also has a much higher theoretical FP32 output at 5.046 TFLOPS versus 1,094.4 GFLOPS for the T400, which explains why the older card leads in raw number crunching. The K40m also offers 12 GB of GDDR5 memory with a 384-bit bus, delivering 288.4 GB/s of bandwidth, compared to the T400's 4 GB of GDDR6 on a 64-bit bus at 80.00 GB/s. Those are the numbers behind the OpenCL win.

The T400 4 GB has no recorded wins in the head-to-head section, but it does have strengths in other areas that the database records. The T400 supports Vulkan 1.4, while the K40m supports Vulkan 1.2.175. The T400 also has display outputs, specifically 3x mini-DisplayPort 1.4a, whereas the K40m has no outputs at all. The T400 is a single-slot card with no power connectors and a 30 W thermal design power, while the K40m is dual-slot with a 245 W TDP and requires a 550 W suggested PSU. The T400 also has a higher boost clock at 1425 MHz versus 876 MHz for the K40m, and it supports FP16 at 2.189 TFLOPS with a 2:1 ratio, a feature the K40m does not list.

For use cases, the K40m is the choice for pure OpenCL compute tasks where memory bandwidth and FP32 throughput dominate. The T400 is the choice for a low-power, display-capable workstation card that fits in a single slot and can drive multiple monitors. The benchmark data does not include a Vulkan score for the K40m, so the T400's Vulkan advantage is not directly measured, but the API version support suggests the T400 is more current in that respect.

Architecture Differences

The two cards come from different architectural generations. The Tesla K40m uses the GK110B chip on the Kepler architecture, manufactured on a 28 nm process at TSMC. The T400 4 GB uses the TU117 chip on the Turing architecture, also built by TSMC but on a 12 nm process. The transistor counts differ significantly: the K40m packs 7,080 million transistors on a 561 mm² die, while the T400 has 4,700 million transistors on a 200 mm² die. The transistor density reflects the process difference, with the K40m at 12.6 million transistors per square millimeter and the T400 at 23.5 million per square millimeter.

The compute units also differ. The K40m has 2,880 shading units, 240 texture mapping units, and 48 raster output units. The T400 has 384 shading units, 24 TMUs, and 16 ROPs. Neither card lists any ray tracing cores or tensor cores, so both rely on traditional shader-based compute. The pixel rate for the K40m is 52.56 GPixel/s, and its texture rate is 210.2 GTexel/s. The T400 posts a pixel rate of 22.80 GPixel/s and a texture rate of 34.20 GTexel/s.

The memory architecture is also fundamentally different. The K40m uses 12 GB of GDDR5 with a 384-bit bus, while the T400 uses 4 GB of GDDR6 with a 64-bit bus. The effective memory clock differs as well: the K40m runs at 6 Gbps effective, and the T400 runs at 10 Gbps effective. Despite the faster memory clock on the T400, the narrow bus limits its bandwidth to 80.00 GB/s, far below the K40m's 288.4 GB/s.

The API support shows generational differences. The K40m supports DirectX 12 (11_1), OpenGL 4.6, and Vulkan 1.2.175. The T400 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. Both are end-of-life products, but they come from different lineage: the K40m follows Tesla Fermi and precedes Tesla Maxwell, while the T400 follows Quadro Volta and precedes Workstation Ampere.

Specification Differences

The two cards differ on nearly every recorded specification. The process node is 28 nm for the K40m and 12 nm for the T400. The base clock is 745 MHz versus 420 MHz, and the boost clock is 876 MHz versus 1425 MHz. The memory size is 12 GB versus 4 GB, the memory type is GDDR5 versus GDDR6, the bus width is 384 bit versus 64 bit, and the bandwidth is 288.4 GB/s versus 80.00 GB/s.

The shading units are 2,880 versus 384, TMUs are 240 versus 24, and ROPs are 48 versus 16. The pixel rate is 52.56 GPixel/s versus 22.80 GPixel/s, and the texture rate is 210.2 GTexel/s versus 34.20 GTexel/s. The FP32 output is 5.046 TFLOPS versus 1,094.4 GFLOPS. The T400 lists FP16 at 2.189 TFLOPS with a 2:1 ratio, and the K40m has no FP16 listing.

The TDP is 245 W versus 30 W. The slot width is dual-slot versus single-slot. The K40m has no power connector listing, while the T400 explicitly has none. The suggested PSU is 550 W versus 200 W. The display outputs are "No outputs" for the K40m versus 3x mini-DisplayPort 1.4a for the T400. The dimensions are recorded for the K40m at 267 mm or 10.5 inches in length, while the T400 has no dimension data.

The release dates are far apart: the K40m launched on 2013-11-21, and the T400 launched on 2021-05-05. The launch MSRP for the K40m is 7,699 USD, and the T400 has no recorded launch MSRP. The production status for both is end-of-life.

FAQ

Q: Which card wins the only direct benchmark comparison?

A: The NVIDIA Tesla K40m wins the Geekbench OpenCL test with a score of 19,885 against 17,320 for the NVIDIA T400 4 GB, a delta of 14.8%.

Q: What is the average benchmark score for each card?

A: The Tesla K40m has an average benchmark score of 19,885, while the T400 4 GB has an average of 16,792. The T400's average comes from two tests: 17,320 in OpenCL and 16,263 in Vulkan.

Q: How do the two cards compare in terms of memory bandwidth?

A: The Tesla K40m delivers 288.4 GB/s of bandwidth using 12 GB of GDDR5 on a 384-bit bus. The T400 4 GB delivers 80.00 GB/s using 4 GB of GDDR6 on a 64-bit bus.

Q: What are the FP32 compute outputs for both cards?

A: The Tesla K40m outputs 5.046 TFLOPS of FP32 performance. The T400 4 GB outputs 1,094.4 GFLOPS, which is roughly one-fifth of the K40m's figure.

Q: Which card supports a newer Vulkan version?

A: The NVIDIA T400 4 GB supports Vulkan 1.4, while the NVIDIA Tesla K40m supports Vulkan 1.2.175. The T400 also supports DirectX 12 (12_1), whereas the K40m supports DirectX 12 (11_1).

Q: Do either of these cards have display outputs?

A: The NVIDIA T400 4 GB has 3x mini-DisplayPort 1.4a outputs. The NVIDIA Tesla K40m has no display outputs at all.

The Verdict

The data clearly separates these two cards by purpose. The NVIDIA Tesla K40m is the compute winner. It leads the only head-to-head OpenCL test by 14.8%, provides 5.046 TFLOPS of FP32, and offers 288.4 GB/s of memory bandwidth. Its 12 GB frame buffer and 384-bit bus make it suitable for large data sets in compute workloads. The K40m sits at the 65th percentile among all GPUs, and its nearest rival, the AMD FirePro W7000, is only 0.1% behind, confirming the K40m is competitive with its direct peers.

The NVIDIA T400 4 GB is the practical workstation card. It has no wins in the head-to-head data, but it offers a 30 W TDP, single-slot design, no power connectors, and three mini-DisplayPort outputs. Its Vulkan 1.4 support and FP16 capability at 2.189 TFLOPS give it modern API coverage that the K40m lacks. The T400 sits at the 60th percentile, and its nearest rival, the NVIDIA Tesla M4, is 0.8% ahead, meaning the T400 slightly trails its closest competitor.

For a user who prioritizes raw OpenCL compute and memory bandwidth, the Tesla K40m is the clear choice. For a user who needs a low-power, display-capable card with modern API support, the T400 4 GB is the only option that fits, since the K40m cannot output video. The benchmark data does not include a Vulkan test for the K40m, so the T400's Vulkan advantage is supported only by API version listings, not by a measured score. The verdict from the recorded numbers is straightforward: the K40m wins on compute, and the T400 wins on versatility and efficiency.

DETAILED SPECIFICATIONS

SPECIFICATION
T400 4 GB
Tesla K40m
Core Specs
Shading Units
384
2,880 +650.0%
Shaders
384
2,880 +650.0%
TMUs
24
240 +900.0%
ROPs
16
48 +200.0%
SM Count
6
Clocks
Base Clock
420 MHz
745 MHz
Boost Clock
1425 MHz
876 MHz
Memory Clock
1250 MHz 10 Gbps effective
1502 MHz 6 Gbps effective
Memory
Memory Size
4 GB
12 GB
VRAM (MB)
4,096
12,288 +200.0%
Memory Type
GDDR6
GDDR5
Memory Bus
64 bit
384 bit
Bandwidth
80.00 GB/s
288.4 GB/s
Cache
L1 Cache
64 KB (per SM)
16 KB (per SMX)
L2 Cache
1024 KB
1536 KB
Performance
Pixel Rate
22.80 GPixel/s
52.56 GPixel/s
Texture Rate
34.20 GTexel/s
210.2 GTexel/s
FP32 (TFLOPS)
1,094.4 GFLOPS
5.046 TFLOPS
FP64 (TFLOPS)
34.20 GFLOPS (1:32)
1.682 TFLOPS (1:3)
FP16 (TFLOPS)
2.189 TFLOPS (2:1)
Power
TDP
30 W
245 W
TDP (W)
30
245 +716.7%
Suggested PSU
200 W
550 W
Power Connectors
None
Architecture
Architecture
Turing
Kepler
GPU Name
TU117
GK110B
Generation
Quadro Turing (Tx000)
Tesla Kepler (Kxx)
Process Size
12 nm
28 nm
Transistors
4,700 million
7,080 million
Die Size
200 mm²
561 mm²
Foundry
TSMC
TSMC
Density
23.5M / mm²
12.6M / mm²
API Support
DirectX
12 (12_1)
12 (11_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.2.175
OpenCL
3.0
3.0
CUDA
7.5
3.5
Shader Model
6.8
6.5 (5.1)
Physical
Slot Width
Single-slot
Dual-slot
Length
267 mm 10.5 inches
Outputs
3x mini-DisplayPort 1.4a
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 3.0 x16
Other
Launch Price
7,699 USD
Production
End-of-life
End-of-life
Predecessor
Quadro Volta
Tesla Fermi
Successor
Workstation Ampere
Tesla Maxwell
View T400 4 GB Details View Tesla K40m Details