NVIDIA T400 vs NVIDIA Tesla M4 Comparison

NVIDIA
GEFORCE

NVIDIA T400

CORE STATE TU117
VRAM 2 GB
CLOCK SPEED 1425 MHz
TDP 30 W
BUS WIDTH 64 bit
ARCHITECTURE Turing
nm
PROCESS 12 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

Tesla M4

CORE STATE GM206
VRAM 4 GB
CLOCK SPEED 1072 MHz
TDP 50 W
BUS WIDTH 128 bit
ARCHITECTURE Maxwell 2.0
nm
PROCESS 28 nm
LAUNCH DATE 2015

PERFORMANCE BENCHMARKS

geekbench_opencl
17,039
16,932
geekbench_vulkan
15,976
N/A

Analysis: NVIDIA T400 vs NVIDIA Tesla M4

FAQ

Q: Which GPU wins in the Geekbench OpenCL benchmark?

A: The NVIDIA T400 edges out the Tesla M4, scoring 17039 versus 16932. The performance difference is a marginal 0.6% in favor of the T400.

Q: How does the Tesla M4 compare to its nearest rivals?

A: The Tesla M4 scores 16932, placing it 0.5% behind the AMD Radeon HD 7970M (17019) and 0.6% behind the NVIDIA GeForce GTX 690 (17037). It also trails the AMD Radeon RX 7600 XT by 0.9%, while leading the NVIDIA T400 4 GB by 0.8%.

Q: What is the performance context for the NVIDIA T400?

A: The T400's average benchmark score is 16508, derived from its OpenCL score of 17039 and Vulkan score of 15976. It sits at the 60th percentile among all GPUs, exactly matching the Tesla M4's percentile ranking.

Q: What are the memory specifications for each card?

A: The Tesla M4 comes with 4 GB of GDDR5 memory on a 128-bit bus, delivering 88.00 GB/s bandwidth. The T400 offers 2 GB of GDDR6 memory on a 64-bit bus, providing 80.00 GB/s bandwidth.

Q: Which card has a higher boost clock?

A: The T400's boost clock reaches 1425 MHz, substantially higher than the Tesla M4's 1072 MHz boost. However, the M4 has a higher base clock at 872 MHz versus the T400's 420 MHz.

Q: Do both cards support the same API levels?

A: Yes, both the Tesla M4 and T400 support DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. This means API-level feature compatibility is identical between the two.

Architecture Differences

The NVIDIA Tesla M4 and NVIDIA T400 represent two distinct architectural generations from NVIDIA, built on different process nodes and targeting different workloads. The Tesla M4 utilizes the GM206 chip based on Maxwell 2.0 architecture, manufactured on TSMC's 28 nm process. In contrast, the T400 uses the TU117 chip based on Turing architecture, also from TSMC but on a more advanced 12 nm node. This process difference is significant: the M4 packs 2,940 million transistors on a 228 mm² die, yielding a transistor density of 12.9 million per mm², while the T400 integrates 4,700 million transistors on a smaller 200 mm² die, achieving a much higher density of 23.5 million per mm².

The compute configurations differ markedly between the two. The Tesla M4 fields 1,024 shading units, 64 texture mapping units, and 32 render output units. The T400, despite being a newer architecture, is more modest in raw unit counts, featuring 384 shading units, 24 TMUs, and 16 ROPs. Neither card includes dedicated ray tracing or tensor cores, keeping both focused on traditional rasterization and compute workloads. The FP32 throughput tells the story: the M4 delivers 2.195 TFLOPS, while the T400 provides 1,094.4 GFLOPS — roughly half the single-precision compute of the older card. Interestingly, the T400 supports FP16 at 2.189 TFLOPS with a 2:1 ratio, whereas the M4 has no listed FP16 capability.

Clock behavior reveals a generational split in design philosophy. The Tesla M4 runs at a modest 872 MHz base with a 1072 MHz boost, while the T400 starts at a very low 420 MHz base but boosts aggressively to 1425 MHz. Memory technology also advances: the M4 uses GDDR5 at 1375 MHz (5.5 Gbps effective), while the T400 employs GDDR6 at 1250 MHz (10 Gbps effective). The M4's 128-bit memory bus gives it a bandwidth advantage at 88.00 GB/s versus the T400's 64-bit bus at 80.00 GB/s.

Power and physical characteristics follow the architectural trends. The Tesla M4 has a 50 W TDP and requires a suggested 250 W power supply, while the T400 draws only 30 W with a suggested 200 W PSU. Both are single-slot cards, but the M4 has no display outputs — it is a compute-oriented accelerator — whereas the T400 provides three mini-DisplayPort 1.4a connectors for visual output. The M4 belongs to the Tesla Maxwell generation with a 2015 release date, while the T400 is part of the Quadro Turing lineup released in 2021.

Head-to-Head Benchmarks

The only direct benchmark comparison available is Geekbench OpenCL, and the results are remarkably close. The NVIDIA T400 scores 17039, while the NVIDIA Tesla M4 scores 16932, giving the T400 a win with a delta of -0.6% from the M4's perspective. This translates to a 107-point difference — a margin so thin that it falls well within typical run-to-run variance for OpenCL workloads. In practical terms, the two cards deliver statistically equivalent OpenCL performance despite their vastly different architectures and specifications.

Contextualizing these scores against their nearest rivals reinforces the closeness of this matchup. The Tesla M4's 16932 score places it just 0.5% behind the AMD Radeon HD 7970M and 0.6% behind the NVIDIA GeForce GTX 690, while leading the T400's own listing as a rival by 0.8% (based on the T400's 16792 average in that particular comparison context). The T400's broader benchmark profile, which includes a Geekbench Vulkan score of 15976, pulls its average down to 16508 — a figure that sits 0.6% ahead of the AMD Radeon PRO W7500 and exactly matches the NVIDIA GeForce RTX 5090 D V2's average of 16504.

The performance parity is striking when you consider the underlying hardware differences. The Tesla M4 achieves its 16932 OpenCL score with double the shading units (1,024 versus 384) and more than double the FP32 throughput (2.195 TFLOPS versus 1,094.4 GFLOPS) compared to the T400. Yet the T400's higher boost clock of 1425 MHz and newer GDDR6 memory help it compensate, delivering a score that is essentially indistinguishable from the older, more compute-heavy card. Both GPUs land at the 60th percentile among all GPUs, confirming that in the aggregate benchmark landscape, they occupy the same performance tier.

Neither card demonstrates a decisive advantage in this head-to-head. The T400's single win in the OpenCL test is by a margin that would not be perceptible in real-world use. The data suggests that for OpenCL compute tasks, system configuration, driver optimization, and thermal behavior would likely have a greater impact on results than the choice between these two specific GPUs.

The Verdict

Based strictly on the benchmark data, the NVIDIA T400 is the nominal winner in this comparison, taking the Geekbench OpenCL test with a 0.6% advantage over the Tesla M4. However, this victory is so narrow that it carries little practical significance — a score difference of 107 points out of approximately 17,000 is within the noise floor of most benchmarking methodologies. Users selecting between these cards should not base their decision on raw performance, as both deliver statistically equivalent OpenCL results.

The decision should instead hinge on workload requirements and system constraints. The Tesla M4 offers twice the memory capacity at 4 GB versus 2 GB, along with higher memory bandwidth (88.00 GB/s versus 80.00 GB/s) thanks to its wider 128-bit bus. This makes the M4 the better choice for memory-intensive compute tasks or datasets that exceed the T400's 2 GB capacity. Its higher FP32 throughput of 2.195 TFLOPS also gives it an edge for raw single-precision compute workloads, despite the benchmark parity.

The T400, meanwhile, brings significant advantages in power efficiency and flexibility. Its 30 W TDP is 40% lower than the M4's 50 W, and it requires a less robust power supply (200 W suggested versus 250 W). The T400 also includes three mini-DisplayPort 1.4a outputs, making it suitable for display-driven workloads, whereas the M4 has no outputs at all. The T400's FP16 capability at 2.189 TFLOPS could be relevant for workloads that leverage reduced precision, a feature the M4 lacks entirely.

For compute-focused deployments where memory capacity and FP32 throughput matter most, the Tesla M4 is the defensible pick. For power-constrained environments, display output requirements, or workloads that can leverage FP16, the T400 is the more versatile option. Given the performance parity, the T400's newer architecture and lower power draw make it the more sensible default for most users, but the M4's memory advantage should not be dismissed for specific use cases.

Specification Differences

The table below highlights only the key specification fields where the NVIDIA Tesla M4 and NVIDIA T400 differ.

| Specification | NVIDIA Tesla M4 | NVIDIA T400 |

| --- | --- | --- |

| Chip | GM206 | TU117 |

| Architecture | Maxwell 2.0 | Turing |

| Generation | Tesla Maxwell (Mxx) | Quadro Turing (Tx000) |

| Process Node | 28 nm | 12 nm |

| Transistors | 2,940 million | 4,700 million |

| Die Size | 228 mm² | 200 mm² |

| Transistor Density | 12.9M / mm² | 23.5M / mm² |

| Base Clock | 872 MHz | 420 MHz |

| Boost Clock | 1072 MHz | 1425 MHz |

| Memory Clock | 1375 MHz (5.5 Gbps effective) | 1250 MHz (10 Gbps effective) |

| Memory Size | 4 GB | 2 GB |

| Memory Type | GDDR5 | GDDR6 |

| Memory Bus Width | 128 bit | 64 bit |

| Memory Bandwidth | 88.00 GB/s | 80.00 GB/s |

| Shading Units | 1024 | 384 |

| TMUs | 64 | 24 |

| ROPs | 32 | 16 |

| Pixel Rate | 34.30 GPixel/s | 22.80 GPixel/s |

| Texture Rate | 68.61 GTexel/s | 34.20 GTexel/s |

| FP32 | 2.195 TFLOPS | 1,094.4 GFLOPS |

| FP16 | — | 2.189 TFLOPS (2:1) |

| TDP | 50 W | 30 W |

| Power Connectors | — | None |

| Suggested PSU | 250 W | 200 W |

| Display Outputs | No outputs | 3x mini-DisplayPort 1.4a |

| Release Date | 2015-11-09 | 2021-05-05 |

| Predecessor | Tesla Kepler | Quadro Volta |

| Successor | Tesla Pascal | Workstation Ampere |

| Geekbench OpenCL | 16932 | 17039 |

| Geekbench Vulkan | — | 15976 |

| Average Benchmark Score | 16932 | 16508 |

DETAILED SPECIFICATIONS

SPECIFICATION
T400
Tesla M4
Core Specs
Shading Units
384
1,024 +166.7%
Shaders
384
1,024 +166.7%
TMUs
24
64 +166.7%
ROPs
16
32 +100.0%
SM Count
6
Clocks
Base Clock
420 MHz
872 MHz
Boost Clock
1425 MHz
1072 MHz
Memory Clock
1250 MHz 10 Gbps effective
1375 MHz 5.5 Gbps effective
Memory
Memory Size
2 GB
4 GB
VRAM (MB)
2,048
4,096 +100.0%
Memory Type
GDDR6
GDDR5
Memory Bus
64 bit
128 bit
Bandwidth
80.00 GB/s
88.00 GB/s
Cache
L1 Cache
64 KB (per SM)
48 KB (per SMM)
L2 Cache
1024 KB
1024 KB
Performance
Pixel Rate
22.80 GPixel/s
34.30 GPixel/s
Texture Rate
34.20 GTexel/s
68.61 GTexel/s
FP32 (TFLOPS)
1,094.4 GFLOPS
2.195 TFLOPS
FP64 (TFLOPS)
34.20 GFLOPS (1:32)
68.61 GFLOPS (1:32)
FP16 (TFLOPS)
2.189 TFLOPS (2:1)
Power
TDP
30 W
50 W
TDP (W)
30
50 +66.7%
Suggested PSU
200 W
250 W
Power Connectors
None
Architecture
Architecture
Turing
Maxwell 2.0
GPU Name
TU117
GM206
Generation
Quadro Turing (Tx000)
Tesla Maxwell (Mxx)
Process Size
12 nm
28 nm
Transistors
4,700 million
2,940 million
Die Size
200 mm²
228 mm²
Foundry
TSMC
TSMC
Density
23.5M / mm²
12.9M / mm²
API Support
DirectX
12 (12_1)
12 (12_1)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
7.5
5.2
Shader Model
6.8
6.8
Physical
Slot Width
Single-slot
Single-slot
Outputs
3x mini-DisplayPort 1.4a
No outputs
Bus Interface
PCIe 3.0 x16
PCIe 3.0 x16
Other
Production
End-of-life
End-of-life
Predecessor
Quadro Volta
Tesla Kepler
Successor
Workstation Ampere
Tesla Pascal
View T400 Details View Tesla M4 Details