NVIDIA Tesla M40
NVIDIA graphics card specifications and benchmark scores
At a Glance
NVIDIANVIDIA Tesla M40 Specifications
GPU Core
Shader units and compute resources
The NVIDIA Tesla M40 GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.
Tesla M40 Clock Speeds
GPU and memory frequencies
Clock speeds directly impact the Tesla M40's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The Tesla M40 by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.
NVIDIA's Tesla M40 Memory
VRAM capacity and bandwidth
VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The Tesla M40's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.
Tesla M40 by NVIDIA Cache
On-chip cache hierarchy
On-chip cache provides ultra-fast data access for the Tesla M40, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.
Tesla M40 Theoretical Performance
Compute and fill rates
Theoretical performance metrics provide a baseline for comparing the NVIDIA Tesla M40 against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.
Maxwell 2.0 Architecture & Process
Manufacturing and design details
The NVIDIA Tesla M40 is built on NVIDIA's Maxwell 2.0 architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the Tesla M40 will perform in GPU benchmarks compared to previous generations.
Power & Thermal
TDP and power requirements
Power specifications for the NVIDIA Tesla M40 determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the Tesla M40 to maintain boost clocks without throttling.
Tesla M40 by NVIDIA Physical & Connectivity
Dimensions and outputs
Physical dimensions of the NVIDIA Tesla M40 are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.
NVIDIA API Support
Graphics and compute APIs
API support determines which games and applications can fully utilize the NVIDIA Tesla M40. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.
Tesla M40 Product Information
Release and pricing details
The NVIDIA Tesla M40 is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the Tesla M40 by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.
About NVIDIA Tesla M40
The NVIDIA Tesla M40 is a Tesla Maxwell-generation compute accelerator (an “Mxx” part) built around the GM200 chip on TSMC’s 28 nm process. The die measures 601 mm² and contains 8,000 million transistors, giving a transistor density of 13.3M per square millimeter. The GPU is configured with 3072 shading units, 192 texture mapping units, and 96 ROPs. Core clocks are 948 MHz base and 1112 MHz boost; memory is clocked at 1502 MHz, listed as 6 Gbps effective. Single-precision throughput is 6.832 TFLOPS. The M40 was released on November 9, 2015, is currently end-of-life, and sits between the Tesla Kepler and Tesla Pascal generations in the product line.
Power and Cooling — TDP, PSU recommendation, connector requirements
The M40 carries a TDP of 250 W. The listed power supply recommendation is 600 W, and the card draws auxiliary power through a single 8-pin EPS connector. The use of an EPS connector is a notable requirement for system integrators, since it is not a standard consumer PCIe power input in the same way. The card is dual-slot in width and has a listed length of 267 mm / 10.5 inches, so chassis clearance needs to accommodate both dimensions. The bus interface is PCIe 3.0 x16, and the card has no display outputs, meaning the power envelope is dedicated to compute work rather than display pipeline output. For a board with a 250 W TDP, the 600 W suggested PSU is the reference figure that should be used when planning host power delivery. The 8-pin EPS connector is the only power input listed in the specification record.
Ray Tracing and Feature Set — RT/tensor cores, API support from facts
There is no RT core count and no tensor core count listed for the M40. The feature set is therefore defined by API support rather than dedicated ray tracing or tensor accelerator blocks. The card supports DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. It is built on the Maxwell 2.0 architecture with the GM200 chip. The lack of display outputs makes the M40 a headless accelerator; no monitor or display connection is available directly from the board. For compute workloads, the supported APIs matter more than display capability. The DirectX 12 feature level 12_1 is the relevant compatibility target for applications using that API. OpenGL 4.6 and Vulkan 1.4 provide the compute and rendering entry points on their respective platforms. Since no RT or tensor core counts are present in the data, any comparison involving those blocks cannot be made from this fact pack.
Benchmark Performance — analyze scores vs rivals with exact % deltas
The available benchmark entries for the M40 are Geekbench OpenCL and Geekbench Vulkan. The OpenCL score is 39192, and the Vulkan score is 44602. The reported average benchmark score is 41897. That average places the M40 at the 84th percentile of all GPUs in the database, meaning it scores higher than 84% of tracked GPUs on aggregate. The nearest rival list is tightly packed. The AMD Radeon Pro 580X has an average score of 41991 and a deltaPct of -0.2, meaning the M40 is 0.2% behind the 580X. The NVIDIA GeForce RTX 5070 has an average score of 41687 and a deltaPct of 0.5, meaning the M40 is 0.5% ahead. The AMD Radeon Pro 5300 has an average score of 41610 and a deltaPct of 0.7, meaning the M40 is 0.7% ahead. The NVIDIA GeForce RTX 3090 has an average score of 41441 and a deltaPct of 1.1, meaning the M40 is 1.1% ahead. The delta values run from -0.2 to 1.1, which places all four nearest rivals inside a narrow band around the M40’s 41897 average. The Vulkan score is higher than the OpenCL score, indicating that the M40 delivers a higher measured result in Vulkan compute than in OpenCL compute. The aggregate score is what drives the 84th percentile ranking and the nearest-rival comparisons.
Who Should Consider It
The M40 is suited to headless compute environments. Because it has no display outputs, it is not a card for connecting a monitor directly. Its 84th percentile standing means it outperforms the majority of GPUs in the database, so it can handle substantial OpenCL or Vulkan compute workloads. The 12 GB GDDR5 frame buffer is the key resource for large working sets at high resolutions; the 288.4 GB/s bandwidth keeps those large buffers moving. Since the nearest rivals are all within the delta range of -0.2 to 1.1, average-score differences should not be the primary decision factor. Instead, system requirements matter: the 250 W TDP, the 600 W PSU guidance, the 8-pin EPS connector, the dual-slot width, and the 267 mm / 10.5 inches length all determine whether the M40 fits a given build. For workloads that depend more on memory capacity than on raw average score, the M40’s 12 GB GDDR5 allocation is a meaningful asset. For workloads that need a direct display output, the M40 is not an option because no outputs are listed.
How It Compares — position vs each nearest rival, one short paragraph per rival
Against the AMD Radeon Pro 580X, the M40 is effectively even. The 580X averages 41991, while the M40 averages 41897, and the listed deltaPct is -0.2. This is the only nearest rival that edges ahead of the M40, and the gap is the smallest in the immediate rivalry set.
Against the NVIDIA GeForce RTX 5070, the M40 is 0.5% ahead. The RTX 5070 averages 41687, which is below the M40’s 41897 but by a very small margin. Despite the different product naming and generation positions, the aggregate benchmark scores put the two cards in the same performance tier.
Against the AMD Radeon Pro 5300, the M40 is 0.7% ahead. The Pro 5300 averages 41610, so the M40’s lead is slightly larger than its lead over the RTX 5070. The listed deltaPct is 0.7, positioning the M40 clearly but narrowly above this rival.
Against the NVIDIA GeForce RTX 3090, the M40 is 1.1% ahead. The RTX 3090 averages 41441, making this the largest delta among the four nearest rivals. Even so, the M40’s 41897 average is only that stated margin above the RTX 3090 in this comparison data.
Memory Subsystem — VRAM size/type, bus width, bandwidth and what it means for high resolutions
The M40 pairs 12 GB of GDDR5 memory with a 384-bit memory bus. The memory clock is 1502 MHz, which the fact pack lists as 6 Gbps effective, and the resulting bandwidth is 288.4 GB/s. For high-resolution work, the 12 GB capacity is the first resource to consider; large textures or large compute buffers must fit within that space. The 384-bit bus and 288.4 GB/s bandwidth determine how quickly that data can be read and written. The listed pixel rate is 106.8 GPixel/s, and the texture rate is 213.5 GTexel/s, both of which rely on the memory subsystem staying fed. A wide 384-bit bus is particularly relevant at high resolutions because more pixels and texels are in flight at any given time. The 12 GB allocation gives the M40 a larger memory pool than a purely low-capacity compute card would have, and the 288.4 GB/s figure is the rated throughput of that pool. The memory subsystem is therefore a central component of the M40’s overall compute profile, especially when working with high-resolution data sets.
Detailed benchmark scores and charts for the NVIDIA Tesla M40 are below.
Benchmark Scores
geekbench_openclSource
Geekbench OpenCL tests GPU compute performance using the cross-platform OpenCL API. This shows how NVIDIA Tesla M40 handles parallel computing tasks like video encoding and scientific simulations. OpenCL is widely supported across different GPU vendors and platforms. Higher scores benefit applications that leverage GPU acceleration for non-graphics workloads.
geekbench_vulkanSource
Geekbench Vulkan tests GPU compute using the modern low-overhead Vulkan API. This shows how NVIDIA Tesla M40 performs with next-generation graphics and compute workloads.
The AMD Equivalent of Tesla M40
Looking for a similar graphics card from AMD? The AMD Radeon RX 480 offers comparable performance and features in the AMD lineup.
Popular NVIDIA Tesla M40 Comparisons
See how the Tesla M40 stacks up against similar graphics cards from the same generation and competing brands.
Compare with Other GPUs
Select another GPU to compare specifications and benchmarks side-by-side.
Browse GPUs