GEFORCE

NVIDIA Tesla M4

NVIDIA graphics card specifications and benchmark scores

4 GB
VRAM
1072
MHz Boost
50W
TDP
128
Bus Width

At a Glance

NVIDIA
VRAM 4 GB
Boost Clock 1,072 MHz
Shaders 1,024
Bus Width 128-bit
TDP 50W
Memory Type GDDR5
Architecture Maxwell 2.0
nm
Process 28 nm
Released Nov 2015

NVIDIA Tesla M4 Specifications

GPU Core

Shader units and compute resources

The NVIDIA Tesla M4 GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
1,024
Shaders
1,024
TMUs
64
ROPs
32

Tesla M4 Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the Tesla M4's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Tesla M4 by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
872 MHz
Base Clock
872 MHz
Boost Clock
1072 MHz
Boost Clock
1,072 MHz
Memory Clock
1375 MHz 5.5 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's Tesla M4 Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Tesla M4's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
4 GB
VRAM
4,096 MB
Memory Type
GDDR5
VRAM Type
GDDR5
Memory Bus
128 bit
Bus Width
128-bit
Bandwidth
88.00 GB/s

Tesla M4 by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the Tesla M4, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
48 KB (per SMM)
L2 Cache
1024 KB

Tesla M4 Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA Tesla M4 against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
2.195 TFLOPS
FP64 (Double)
68.61 GFLOPS (1:32)
Pixel Rate
34.30 GPixel/s
Texture Rate
68.61 GTexel/s

Maxwell 2.0 Architecture & Process

Manufacturing and design details

The NVIDIA Tesla M4 is built on NVIDIA's Maxwell 2.0 architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Tesla M4 will perform in GPU benchmarks compared to previous generations.

Architecture
Maxwell 2.0
GPU Name
GM206
Process Node
28 nm
Foundry
TSMC
Transistors
2,940 million
Die Size
228 mm²
Density
12.9M / mm²

Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA Tesla M4 determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Tesla M4 to maintain boost clocks without throttling.

TDP
50 W
TDP
50W
Suggested PSU
250 W

Tesla M4 by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA Tesla M4 are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
Single-slot
Bus Interface
PCIe 3.0 x16
Display Outputs
No outputs
Display Outputs
No outputs

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA Tesla M4. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 (12_1)
DirectX
12 (12_1)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.4
Vulkan
1.4
OpenCL
3.0
CUDA
5.2
Shader Model
6.8

Tesla M4 Product Information

Release and pricing details

The NVIDIA Tesla M4 is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Tesla M4 by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Nov 2015
Production
End-of-life
Predecessor
Tesla Kepler
Successor
Tesla Pascal

About NVIDIA Tesla M4

The NVIDIA Tesla M4 is a compact, single-slot accelerator built on the Maxwell 2.0 architecture, designed for specific server and datacenter workloads rather than consumer gaming. Its 28 nm TSMC process and 2,940 million transistors in a 228 mm² die place it in a particular performance tier, as evidenced by its benchmark results. The data reveals a GPU that sits near the middle of the pack, with a 59th percentile ranking among all GPUs, indicating that while it is not a high-end part, it holds its own against several more modern and power-hungry alternatives.

Memory Subsystem

The Tesla M4 is equipped with 4 GB of GDDR5 memory, which is paired with a 128-bit memory bus. This configuration yields a memory bandwidth of 88.00 GB/s, a figure that is modest by contemporary standards but was adequate for the card's intended server-side tasks at its launch. The memory operates at an effective data rate of 5.5 Gbps, translating to a base memory clock of 1375 MHz.

For high-resolution workloads, the 4 GB capacity is a significant limiting factor. While the bandwidth of 88.00 GB/s is sufficient for many compute tasks, it is not designed for the massive framebuffers and texture data associated with modern 4K gaming or rendering. The data suggests that the M4 is more suited to tasks where memory size is less critical than compute throughput. The 128-bit bus width, when compared to the 88.00 GB/s bandwidth, indicates a design balanced for efficiency rather than raw data movement. The pixel rate of 34.30 GPixel/s and texture rate of 68.61 GTexel/s further reinforce that this is a compute-oriented part, not a graphics powerhouse. In essence, the memory subsystem is adequate for its era and purpose, but it would become a bottleneck when pushing large datasets or high-resolution textures.

Power and Cooling

The Tesla M4 has a notably low thermal design power (TDP) of just 50 W. This is a defining characteristic of the card, allowing it to be passively cooled in a single-slot form factor, which is ideal for dense server environments. The suggested power supply unit (PSU) for a system containing this card is 250 W, reflecting the card's frugal nature. There are no power connectors listed, which is consistent with a 50 W TDP; the card can draw all its power from the PCIe 3.0 x16 slot itself.

This low power envelope is a significant advantage in multi-GPU server configurations, as it reduces overall system heat and power draw. The 50 W TDP also means that cooling requirements are minimal, allowing for a simpler and more reliable chassis design. The data indicates that the M4 was engineered with a clear focus on efficiency, making it a low-impact addition to any server. This is in stark contrast to many high-performance accelerators that require substantial cooling and power delivery infrastructure.

Benchmark Performance

The primary benchmark data available for the Tesla M4 is a single Geekbench OpenCL score of 16970. This score is the average benchmark score and places the card in the 59th percentile of all GPUs. To understand what this means, it is essential to compare it against the nearest rivals. The data shows the M4 is 2.9% ahead of the NVIDIA T400, which scores 16486. This is a slim lead, indicating near-parity in raw compute performance.

Conversely, the M4 is 2.9% behind the NVIDIA Tesla K40c, which scores 17468. This is a crucial comparison, as the K40c is a much larger and more power-hungry card from the previous Kepler generation. The M4 achieving near-parity with a 2.9% deficit confirms the efficiency of the Maxwell architecture. Similarly, the M4 is 3.0% behind the AMD Radeon Pro 560, which scores 17497. The margin is again minimal, showing that the M4 competes effectively with a much newer mobile workstation GPU. Finally, the M4 is 3.4% ahead of the AMD Radeon 680M, an integrated graphics solution, which scores 16407. This shows that the dedicated M4 still holds a measurable performance advantage over a modern high-end iGPU.

The 2.195 TFLOPS of FP32 compute power is the theoretical peak, but the OpenCL score provides a more realistic picture of performance. The benchmark results indicate that the M4's performance is tightly clustered with its rivals, all within a 3.4% band. This suggests that for compute workloads, the architectural differences and clock speeds (base 872 MHz, boost 1072 MHz) result in very similar real-world performance.

How It Compares

NVIDIA T400: The M4 is 2.9% faster than the T400 in the OpenCL benchmark. This is a notable result, as the T400 is a newer, more modern card. The M4's lead, while small, indicates that its Maxwell architecture with 1024 shading units can still outperform a more recent entry-level card. The performance parity suggests that for basic compute tasks, the M4 remains relevant despite being end-of-life.

NVIDIA Tesla K40c: The M4 is 2.9% slower than the K40c. This is a remarkable data point. The K40c is a flagship-class compute card from the Kepler era, with a much higher TDP and larger die. The M4, with its 50 W TDP, nearly matches it. This reinforces the architectural efficiency of Maxwell 2.0 and shows that the M4 was a significant step forward in performance-per-watt.

AMD Radeon Pro 560: The M4 trails the Radeon Pro 560 by 3.0%. The Pro 560 is a mobile workstation GPU from a later generation. The M4's ability to stay within 3% of this card demonstrates that its compute capabilities were well-aligned with contemporary mid-range solutions. The data implies that for many compute tasks, the M4 is functionally equivalent to this newer card.

AMD Radeon 680M: The M4 is 3.4% ahead of the Radeon 680M. The 680M is a high-performance integrated GPU found in modern APUs. The M4's lead, while modest, is significant because it shows that a dedicated, older GPU can still outperform a high-end iGPU. This positions the M4 as a viable low-cost compute option where dedicated VRAM and a discrete GPU are preferred.

Ray Tracing and Feature Set

The Tesla M4 does not have dedicated ray tracing (RT) cores or tensor cores, as these are absent from the FACT PACK data. This is consistent with its Maxwell 2.0 architecture, which predates the introduction of hardware-accelerated ray tracing in NVIDIA's consumer and professional lines. The card's feature set is therefore focused on traditional rasterization and compute workloads.

In terms of API support, the M4 supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. This is a robust set of modern APIs, ensuring broad software compatibility. The DirectX 12_1 support means it can handle feature level 12_1, which is essential for many modern games and applications. The Vulkan 1.4 support is particularly relevant for compute and cross-platform workloads. The absence of RT and tensor cores means that the M4 is not suitable for modern ray-traced rendering or AI-accelerated tasks that rely on these dedicated hardware units. Its strength lies in general-purpose compute and legacy graphics workloads.

Who Should Consider It

Based on the benchmark scores and memory configuration, the Tesla M4 is a candidate for specific, low-power compute scenarios. The data shows it performs at a level comparable to the NVIDIA T400 and AMD Radeon 680M, making it suitable for tasks like basic GPU compute, virtual desktop infrastructure, or as a dedicated PhysX or compute offload card. The 4 GB VRAM is sufficient for moderate datasets, but not for large-scale deep learning or high-resolution texture-heavy workloads.

For users with a 250 W PSU and a need for a single-slot, low-power accelerator, the M4's 50 W TDP is a compelling attribute. Its performance in the 59th percentile suggests it can handle a wide range of tasks, but it will not excel in demanding applications. The 2.9% lead over the T400 and 3.4% lead over the Radeon 680M indicate that it is not obsolete, but rather a viable option for budget-conscious compute environments where power efficiency is paramount. It is not recommended for high-resolution gaming or professional 3D rendering, as the 88.00 GB/s bandwidth and 4 GB capacity are limiting factors. The card is best suited for headless servers where its display outputs (none) are not a drawback.

FAQ

Q: What is the performance of the NVIDIA Tesla M4 in compute benchmarks?

A: The Tesla M4 achieves an average Geekbench OpenCL score of 16970, placing it in the 59th percentile of all GPUs.

Q: How does the Tesla M4 compare to the NVIDIA T400?

A: The Tesla M4 is 2.9% faster than the NVIDIA T400, which has an average score of 16486.

Q: Does the Tesla M4 have hardware ray tracing capabilities?

A: No, the Tesla M4 does not have dedicated RT cores or tensor cores, as it is based on the Maxwell 2.0 architecture.

Q: What is the power consumption of the Tesla M4?

A: The Tesla M4 has a TDP of 50 W, and the suggested system PSU is 250 W.

Q: What is the memory bandwidth of the Tesla M4?

A: The Tesla M4 has a memory bandwidth of 88.00 GB/s, utilizing 4 GB of GDDR5 memory on a 128-bit bus.

Q: Which GPU is slower than the Tesla M4?

A: The AMD Radeon 680M is 3.4% slower than the Tesla M4, with an average score of 16407.

Detailed benchmark scores and charts for the NVIDIA Tesla M4 are below.

Benchmark Scores

geekbench_openclSource

Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA Tesla M4 handles parallel computing tasks like video encoding and scientific simulations. OpenCL is widely supported across different GPU vendors and platforms. Higher scores benefit applications that leverage GPU acceleration for non-graphics workloads.

geekbench_opencl #332 of 650
16,932
4%
Max: 388,405
Compare with other GPUs

Top 5 Performers

#1 NVIDIA RTX 6000D
388,405
#2 NVIDIA B300 SXM6 AC
369,831
#3 NVIDIA B200
345,482
#4 NVIDIA H200 NVL
334,891

Popular NVIDIA Tesla M4 Comparisons

See how the Tesla M4 stacks up against similar graphics cards from the same generation and competing brands.

Compare with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs