GPU Comparison

NVIDIA
GEFORCE

NVIDIA T400 4 GB

CORE STATE TU117
VRAM 4 GB
CLOCK SPEED 1425 MHz
TDP 30 W
BUS WIDTH 64 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Tesla M4

CORE STATE GM206
VRAM 4 GB
CLOCK SPEED 1072 MHz
TDP 50 W
BUS WIDTH 128 bit
ARCHITECTURE Maxwell 2.0
nm
PROCESS 28 nm
LAUNCH DATE 2015

PERFORMANCE BENCHMARKS

geekbench_opencl
17,320
16,932
geekbench_vulkan
16,263
N/A

Analysis: NVIDIA T400 4 GB vs NVIDIA Tesla M4

The NVIDIA Tesla M4 and NVIDIA T400 4 GB are both end-of-life, single-slot workstation cards, but they represent two very different generations of NVIDIA’s GPU design. The Tesla M4 is a Maxwell 2.0-era compute accelerator with no display outputs, while the T400 is a Turing-based entry-level workstation card with three mini-DisplayPort outputs. The benchmark data shows a narrow overall victory for the T400, but the underlying architecture and feature set tell a more nuanced story.

Where Each One Wins

The T400 4 GB claims the only head-to-head benchmark victory in the provided data. In the Geekbench OpenCL test, the T400 scores 17,320 points against the Tesla M4’s 16,932, a delta of -2.2% from the perspective of the M4. This means the T400 is roughly 2.3% faster in raw compute workloads measured by OpenCL. The T400 also has a second benchmark result, a Geekbench Vulkan score of 16,263, which the Tesla M4 lacks entirely, indicating the Turing card has additional graphics API capabilities that the Maxwell part cannot match.

The Tesla M4, by contrast, does not win any head-to-head benchmark. However, its single OpenCL score of 16,932 places it in the 60th percentile of all GPUs, identical to the T400’s percentile ranking. Looking at the rival lists, the M4’s average score of 16,932 is actually higher than the T400’s average of 16,792, despite losing the direct comparison. This discrepancy arises because the T400’s average includes its lower Vulkan score, while the M4’s average is based solely on its OpenCL result. In terms of pure compute density, the M4’s 2.195 TFLOPS FP32 performance exceeds the T400’s 1,094.4 GFLOPS, making the M4 the stronger choice for raw floating-point throughput.

Architecture Differences

The architectural gap between these two cards is substantial. The Tesla M4 uses the GM206 chip built on TSMC’s 28 nm process, packing 2,940 million transistors into a 228 mm² die. The T400 uses the TU117 chip on a 12 nm process, with 4,700 million transistors in a slightly smaller 200 mm² die. This means the T400 achieves a transistor density of 23.5M per mm², nearly double the M4’s 12.9M per mm², reflecting the newer manufacturing node.

The M4 is built on Maxwell 2.0 architecture, while the T400 uses Turing. This generational leap brings several key differences. The M4 has 1,024 shading units, 64 texture mapping units, and 32 ROPs. The T400 is far leaner in these counts: 384 shading units, 24 TMUs, and 16 ROPs. Despite having fewer than half the shaders, the T400’s higher boost clock of 1,425 MHz (versus the M4’s 1,072 MHz boost) helps it stay competitive in compute tasks. The T400 also supports FP16 at 2.189 TFLOPS with a 2:1 ratio, a feature entirely absent from the M4’s specification sheet.

Memory configurations differ as well. Both cards have 4 GB of VRAM, but the M4 uses GDDR5 on a 128-bit bus, delivering 88.00 GB/s of bandwidth. The T400 uses GDDR6 on a 64-bit bus, which yields 80.00 GB/s, slightly lower bandwidth despite the faster memory technology. The M4’s memory runs at 1,375 MHz (5.5 Gbps effective), while the T400’s runs at 1,250 MHz (10 Gbps effective). The T400 also has display outputs (3x mini-DisplayPort 1.4a), whereas the M4 has none, making the T400 suitable for visual output while the M4 is strictly a compute accelerator.

FAQ

Q: Which GPU has higher raw FP32 compute performance?

A: The Tesla M4 is clearly ahead, with 2.195 TFLOPS compared to the T400’s 1,094.4 GFLOPS. This is more than double the FP32 throughput.

Q: Does the T400 support any API features the M4 lacks?

A: Yes, the T400 has FP16 support at 2.189 TFLOPS with a 2:1 ratio. The M4 has no FP16 specification listed. Both cards support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4.

Q: How do their memory bandwidths compare?

A: The M4 offers 88.00 GB/s over a 128-bit GDDR5 bus. The T400 provides 80.00 GB/s over a 64-bit GDDR6 bus. The M4 has a 10% bandwidth advantage.

Q: What are the power requirements for each card?

A: The Tesla M4 has a 50 W TDP with a suggested 250 W PSU. The T400 has a 30 W TDP with a suggested 200 W PSU. The T400 also requires no power connectors, while the M4’s connector situation is unspecified.

Q: Which card can drive displays?

A: Only the T400 can output video, featuring 3x mini-DisplayPort 1.4a. The Tesla M4 has no display outputs and is intended for compute-only workloads.

Q: How do their average benchmark scores rank them against other GPUs?

A: Both sit in the 60th percentile of all GPUs. The M4’s average score is 16,932, slightly above the T400’s 16,792. In the direct OpenCL head-to-head, however, the T400 wins 17,320 to 16,932.

Specification Differences

The two cards diverge on nearly every major specification. The chip and process differ entirely: GM206 (Maxwell 2.0) on 28 nm versus TU117 (Turing) on 12 nm. Transistor counts are 2,940 million versus 4,700 million, and die sizes are 228 mm² versus 200 mm². The T400’s transistor density of 23.5M/mm² more than doubles the M4’s 12.9M/mm².

Clock speeds show a mixed picture. The M4 has a higher base clock (872 MHz versus 420 MHz), but the T400 has a much higher boost clock (1,425 MHz versus 1,072 MHz). Memory clocks are 1,375 MHz (5.5 Gbps effective) for the M4 and 1,250 MHz (10 Gbps effective) for the T400, reflecting the GDDR5-to-GDDR6 transition.

Core counts heavily favor the M4: 1,024 shading units versus 384, 64 TMUs versus 24, and 32 ROPs versus 16. This translates to pixel rates of 34.30 GPixel/s versus 22.80 GPixel/s and texture rates of 68.61 GTexel/s versus 34.20 GTexel/s. FP32 output is 2.195 TFLOPS versus 1,094.4 GFLOPS. Power consumption is lower on the T400 at 30 W versus 50 W, with suggested PSUs of 200 W versus 250 W. The T400 is the only one with display outputs (3x mini-DisplayPort 1.4a) and power connectors (none). Their bus interfaces are identical: PCIe 3.0 x16.

Head-to-Head Benchmarks

The only direct benchmark comparison available is Geekbench OpenCL. The T400 scores 17,320, while the Tesla M4 scores 16,932. The delta is -2.2% relative to the M4, meaning the T400 outperforms the M4 by approximately 388 points. This is a narrow margin, roughly 2.3% faster, which is notable given the M4’s significantly higher shader count and FP32 throughput. The T400’s advantage likely stems from its higher boost clock and newer architecture, which compensates for its fewer cores.

The T400 also posts a Geekbench Vulkan score of 16,263, a result that has no M4 counterpart. This absence suggests the M4 either was not tested under Vulkan or lacks the same level of driver optimization. The Vulkan score is lower than the T400’s OpenCL result by 1,057 points, indicating that even within the same GPU, API choice affects performance. The M4’s average benchmark score of 16,932 is higher than the T400’s average of 16,792, but this is because the T400’s two scores are averaged together, pulling its mean down.

Looking at nearest rivals provides additional context. The M4’s closest competitor is the AMD Radeon HD 7970M, which scores 17,019, just 0.5% above the M4. The T400’s nearest rival is the AMD Radeon RX 7600S at 16,696, which the T400 beats by 0.6%. Both cards sit in a tight performance cluster around 17,000 points, meaning real-world application differences may matter more than the raw benchmark deltas.

The Verdict

The data presents a clear split. For pure compute throughput, the Tesla M4 is the superior card: it delivers 2.195 TFLOPS FP32, 88.00 GB/s of memory bandwidth, and a higher texture and pixel rate. Its single OpenCL score of 16,932 is respectable and places it in the 60th percentile. However, the M4 is a compute-only device with no display outputs, a 50 W TDP, and a 2015 release date, making it a legacy accelerator suited for headless compute tasks.

For anyone needing a workstation card with visual output, the T400 4 GB is the only option between the two. It wins the direct OpenCL benchmark (17,320 versus 16,932), offers Vulkan support, and includes three mini-DisplayPort 1.4a connectors. Its 30 W TDP and lack of power connectors make it far easier to integrate into low-power systems. The T400’s FP16 capability at 2.189 TFLOPS is a bonus for workloads that leverage half-precision math, something the M4 cannot do.

The verdict hinges on use case. If the workload is pure compute on a server without display needs, the M4’s higher FP32 and bandwidth make it the logical pick. If the task involves any visualization, desktop output, or modern API flexibility, the T400 is the better choice despite its lower raw compute. The benchmark scores are close, within 2.2%, so the architectural differences and feature sets should drive the decision. Neither card is a clear winner; they are tools for different jobs, and the data reflects that divergence.

DETAILED SPECIFICATIONS

SPECIFICATION
T400 4 GB
Tesla M4
Core Specs
Shading Units
384
1,024 +166.7%
Shaders
384
1,024 +166.7%
TMUs
24
64 +166.7%
ROPs
16
32 +100.0%
SM Count
6
Clocks
Base Clock
420 MHz
872 MHz
Boost Clock
1425 MHz
1072 MHz
Memory Clock
1250 MHz 10 Gbps effective
1375 MHz 5.5 Gbps effective
Memory
Memory Size
4 GB
4 GB
VRAM (MB)
4,096
4,096 0.0%
Memory Type
GDDR6
GDDR5
Memory Bus
64 bit
128 bit
Bandwidth
80.00 GB/s
88.00 GB/s
Cache
L1 Cache
64 KB (per SM)
48 KB (per SMM)
L2 Cache
1024 KB
1024 KB
Performance
Pixel Rate
22.80 GPixel/s
34.30 GPixel/s
Texture Rate
34.20 GTexel/s
68.61 GTexel/s
FP32 (TFLOPS)
1,094.4 GFLOPS
2.195 TFLOPS
FP64 (TFLOPS)
34.20 GFLOPS (1:32)
68.61 GFLOPS (1:32)
FP16 (TFLOPS)
2.189 TFLOPS (2:1)
Power
TDP
30 W
50 W
TDP (W)
30
50 +66.7%
Suggested PSU
200 W
250 W
Power Connectors
None
Architecture
Architecture
Turing
Maxwell 2.0
GPU Name
TU117
GM206
Generation
Quadro Turing (Tx000)
Tesla Maxwell (Mxx)
Process Size
12 nm
28 nm
Transistors
4,700 million
2,940 million
Die Size
200 mm²
228 mm²
Foundry
TSMC
TSMC
Density
23.5M / mm²
12.9M / mm²
API Support
DirectX
12 (12_1)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
7.5
5.2
Shader Model
6.8
6.8
Physical
Slot Width
Single-slot
Single-slot
Outputs
3x mini-DisplayPort 1.4a
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 3.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Quadro Volta
Tesla Kepler
Successor
Workstation Ampere
Tesla Pascal
View T400 4 GB Details View Tesla M4 Details