NVIDIA A2 vs NVIDIA Tesla M40 Comparison
NVIDIA A2
Tesla M40
PERFORMANCE BENCHMARKS
Analysis: NVIDIA A2 vs NVIDIA Tesla M40
FAQ
Q: Which GPU wins the Geekbench OpenCL benchmark?
A: The NVIDIA Tesla M40 scores 39,192 in Geekbench OpenCL, which is 10.8% higher than the NVIDIA A2’s score of 35,357.
Q: How large is the gap in the Geekbench Vulkan test?
A: The Tesla M40 scores 44,602, while the A2 scores 34,023. That gives the M40 a 31.1% lead in Vulkan, its largest margin in the head-to-head data.
Q: Which card has a higher average benchmark score?
A: The Tesla M40 has an average benchmark score of 41,897, compared to the A2’s 34,690. The M40 also sits at the 83rd percentile among all GPUs, while the A2 sits at the 79th.
Q: What are the node sizes for these two architectures?
A: The Tesla M40 is built on a 28 nm process at TSMC, while the A2 uses an 8 nm process at Samsung. The A2’s die is 200 mm² versus the M40’s 601 mm².
Q: Do both cards support the same DirectX version?
A: No. The Tesla M40 supports DirectX 12 (12_1), while the A2 supports DirectX 12 Ultimate (12_2). Both support OpenGL 4.6 and Vulkan 1.4.
Q: Which card has more memory and what type?
A: The A2 has 16 GB of GDDR6 memory on a 128-bit bus, while the Tesla M40 has 12 GB of GDDR5 on a 384-bit bus. The M40 still delivers higher memory bandwidth at 288.4 GB/s versus 200.1 GB/s for the A2.
Architecture Differences
The Tesla M40 and NVIDIA A2 come from completely different design eras. The M40 uses the GM200 chip with Maxwell 2.0 architecture, a 28 nm TSMC process, and a massive 601 mm² die holding 8,000 million transistors. The A2 uses the GA107 chip with Ampere architecture, an 8 nm Samsung process, and a much smaller 200 mm² die that packs 8,700 million transistors. The transistor density tells the story: the A2 reaches 43.5M transistors per mm², more than three times the M40’s 13.3M per mm².
Compute resources differ sharply. The M40 fields 3,072 shading units, 192 TMUs, and 96 ROPs. The A2 has 1,280 shading units, 40 TMUs, and 32 ROPs. Yet the A2 adds hardware that the M40 lacks entirely: 10 ray tracing cores and 40 tensor cores. The A2 also supports FP16 at a 1:1 ratio with FP32, offering 4.531 TFLOPS in each precision, while the M40 lists only FP32 at 6.832 TFLOPS with no FP16 figure recorded.
Clock speeds favor the A2. Its base clock is 1,440 MHz with a boost of 1,770 MHz, versus the M40’s 948 MHz base and 1,112 MHz boost. The A2’s memory runs at 1,563 MHz (12.5 Gbps effective) compared to the M40’s 1,502 MHz (6 Gbps effective). The M40 compensates with a 384-bit memory bus and 12 GB of GDDR5, while the A2 uses a 128-bit bus and 16 GB of GDDR6. The result: the M40’s 288.4 GB/s bandwidth exceeds the A2’s 200.1 GB/s despite the newer memory type on the A2.
Power and physical design also diverge. The M40 draws 250 W, requires a dual-slot cooler with an 8-pin EPS connector, and a 600 W suggested PSU. The A2 draws 60 W, fits in a single slot, needs no power connectors, and only suggests a 250 W PSU. The M40 measures 267 mm (10.5 inches) in length; the A2’s dimensions are not recorded. Both cards have no display outputs, but the A2 uses PCIe 4.0 x8 while the M40 uses PCIe 3.0 x16.
Release timing and lineage also separate them. The M40 launched on November 9, 2015, with the Tesla Kepler as its predecessor and Tesla Pascal as its successor. The A2 launched exactly six years later on November 9, 2021, following Quadro Turing and leading to Workstation Ada. Both are now end-of-life products.
Head-to-Head Benchmarks
The recorded head-to-head data covers two tests, and the Tesla M40 wins both. In Geekbench OpenCL, the M40 scores 39,192 against the A2’s 35,357, a 10.8% advantage. In Geekbench Vulkan, the M40 scores 44,602 against 34,023, a 31.1% advantage. The Vulkan result is notable because it nearly triples the OpenCL gap, suggesting the M40’s Maxwell architecture performs disproportionately well under the Vulkan API compared to the A2’s Ampere implementation.
The average benchmark scores reinforce the pattern. The M40 averages 41,897 and sits in the 83rd percentile of all GPUs. The A2 averages 34,690 and sits in the 79th percentile. The M40’s nearest rivals in the database include the Tesla M40 24 GB at 41,707 (only 0.5% behind), the GeForce RTX 3080 Ti at 41,187 (1.7% behind), and the Radeon Pro 5300 at 40,870 (2.5% behind). The only rival ahead of it is the Radeon RX 7650 GRE at 42,723, which leads by 1.9%. The A2’s nearest rivals cluster tightly around it: the T1000 8 GB trails by 0.4%, the Radeon HD 7970 trails by 0.4%, the TITAN V trails by 1%, and the RTX A1000 trails by 1.4%. This indicates the A2 sits in a competitive mid-range bracket where small score differences separate cards, whereas the M40 sits near the top of its own tier.
Looking at the two benchmark wins, the M40’s Vulkan advantage is particularly striking. A 31.1% delta is a substantial margin for two NVIDIA cards, and it suggests that software workloads using Vulkan will see a much larger performance gap than those using OpenCL. For OpenCL-oriented tasks, the M40 still leads, but the 10.8% edge is more modest and could be influenced by driver optimizations or workload characteristics.
Specification Differences
The two cards differ across nearly every major specification category. Process node: 28 nm TSMC for the M40 versus 8 nm Samsung for the A2. Transistor count: 8,000 million versus 8,700 million. Die size: 601 mm² versus 200 mm². Transistor density: 13.3M per mm² versus 43.5M per mm².
Compute units: 3,072 shading units, 192 TMUs, and 96 ROPs on the M40; 1,280 shading units, 40 TMUs, and 32 ROPs on the A2. The A2 adds 10 RT cores and 40 tensor cores; the M40 has none of either. Pixel rate: 106.8 GPixel/s versus 56.64 GPixel/s. Texture rate: 213.5 GTexel/s versus 70.80 GTexel/s. FP32: 6.832 TFLOPS versus 4.531 TFLOPS. FP16 is only listed for the A2 at 4.531 TFLOPS (1:1).
Memory: 12 GB GDDR5 on a 384-bit bus with 288.4 GB/s bandwidth versus 16 GB GDDR6 on a 128-bit bus with 200.1 GB/s bandwidth. Clocks: 948/1112 MHz base/boost for the M40 versus 1440/1770 MHz for the A2. Memory clocks: 1502 MHz (6 Gbps effective) versus 1563 MHz (12.5 Gbps effective).
Power and cooling: 250 W TDP, dual-slot, 8-pin EPS, 600 W suggested PSU for the M40; 60 W TDP, single-slot, no power connectors, 250 W suggested PSU for the A2. Bus interface: PCIe 3.0 x16 versus PCIe 4.0 x8. Dimensions: the M40 is 267 mm long; the A2 has no recorded length, height, or width. DirectX: 12 (12_1) versus 12 Ultimate (12_2). Both share OpenGL 4.6 and Vulkan 1.4. Display outputs: none on either. Release dates: November 9, 2015 versus November 9, 2021. Both are end-of-life.
Where Each One Wins
The Tesla M40 wins on raw compute throughput. Its FP32 rate of 6.832 TFLOPS is about 51% higher than the A2’s 4.531 TFLOPS. Its pixel rate of 106.8 GPixel/s nearly doubles the A2’s 56.64 GPixel/s, and its texture rate of 213.5 GTexel/s is roughly three times the A2’s 70.80 GTexel/s. Memory bandwidth also favors the M40: 288.4 GB/s versus 200.1 GB/s. For workloads that saturate bandwidth, such as large data transfers or high-resolution texture streaming, the M40’s wider 384-bit bus gives it a structural advantage. The benchmark data agrees: the M40 wins both recorded tests, with the Vulkan margin at 31.1% and the OpenCL margin at 10.8%.
The A2 wins on efficiency and modern features. Its 60 W TDP is a fraction of the M40’s 250 W, and it requires no external power connectors. It also carries ray tracing cores and tensor cores, enabling workloads that the M40 cannot accelerate at all. The A2’s 16 GB of GDDR6 memory doubles the capacity advantage in VRAM-heavy scenarios, even though its bandwidth is lower. Its FP16 support at 1:1 ratio means mixed-precision workloads can run at the same throughput as FP32, a capability absent from the recorded M40 data. The A2 also uses PCIe 4.0, which doubles the per-lane bandwidth of the M40’s PCIe 3.0 interface, though only at x8 width.
The TDP difference is dramatic: 250 W versus 60 W. For dense server deployments or systems with tight power budgets, the A2 is the only viable choice between the two. The M40’s dual-slot footprint and 8-pin EPS connector also impose physical and cabling requirements that the A2 avoids entirely.
The Verdict
The data paints a clear picture. The Tesla M40 is the faster card in both recorded benchmarks. It wins Geekbench OpenCL by 10.8% and Geekbench Vulkan by 31.1%, and its average score of 41,897 places it at the 83rd percentile versus the A2’s 34,690 and 79th percentile. Its higher FP32 throughput, pixel rate, texture rate, and memory bandwidth all support the benchmark results. Anyone prioritizing raw compute performance should choose the M40.
The A2 appeals to a different set of priorities. Its 60 W power draw, single-slot design, and lack of power connectors make it far easier to deploy in constrained environments. Its 16 GB of GDDR6 memory provides more capacity than the M40’s 12 GB, and its tensor cores and RT cores open up acceleration paths that the M40 cannot match. Its FP16 capability at 1:1 with FP32 further extends its usefulness in mixed-precision workloads.
For compute-bound tasks like rendering, simulation, or large-scale FP32 math, the M40’s 6.832 TFLOPS and 31.1% Vulkan lead make it the clear winner. For power-sensitive inference, ray tracing, or any workload that can exploit tensor cores or FP16, the A2 is the sensible pick despite losing both benchmarks. The two cards are not direct substitutes; they are solutions to different problems. The M40 is a high-throughput compute card from the Maxwell era, while the A2 is a low-power Ampere accelerator with modern feature support. The choice depends entirely on whether the workload prioritizes raw speed or efficiency and feature breadth.