NVIDIA Quadro GV100 vs NVIDIA Tesla M40 24 GB Comparison
NVIDIA Quadro GV100
Tesla M40 24 GB
PERFORMANCE BENCHMARKS
Analysis: NVIDIA Quadro GV100 vs NVIDIA Tesla M40 24 GB
Where Each One Wins
The benchmark data splits these two cards into completely different performance tiers. The NVIDIA Quadro GV100 wins every recorded head-to-head comparison, and by substantial margins. In Geekbench OpenCL, the GV100 scores 150,004 against the Tesla M40 24 GB's 37,439, a 75% gap. In Geekbench Vulkan, the GV100 posts 139,526 versus 45,975, a 67% advantage. The Tesla M40 24 GB does not win a single benchmark in the database.
This is not a close contest. The GV100's average benchmark score of 35,520 across all recorded tests places it in the 80th percentile of all GPUs, while the Tesla M40 24 GB sits slightly higher at the 83rd percentile with an average of 41,707. The percentile discrepancy is interesting: the M40's average is pulled up by its two strong Geekbench results, but the GV100 has a much wider benchmark portfolio including Passmark tests where it scores lower, dragging its average down. The GV100 still holds the decisive edge in compute workloads.
For users prioritizing raw compute throughput, the GV100 is the clear choice. For users who only need OpenCL or Vulkan performance, the GV100 again dominates. The M40 24 GB has no recorded category where it leads.
Architecture Differences
The two cards come from different NVIDIA architectures and generations. The Tesla M40 24 GB uses the GM200 chip on Maxwell 2.0 architecture, built on a 28 nm process at TSMC. It packs 8,000 million transistors on a 601 mm² die, for a transistor density of 13.3 million per mm². The Quadro GV100 uses the GV100 chip on Volta architecture, built on a 12 nm process, also at TSMC. It houses 21,100 million transistors on an 815 mm² die, for a density of 25.9 million per mm².
The GV100 is the more modern and more complex chip. Its transistor count is roughly 2.6 times that of the M40, and its die is 214 mm² larger. The Volta architecture brings tensor cores, 640 of them, which the Maxwell-based M40 lacks entirely. The M40 has no tensor cores and no RT cores; the GV100 also has no RT cores, but the tensor core presence is a major architectural differentiator for compute workloads.
Memory architecture diverges sharply. The M40 uses 24 GB of GDDR5 on a 384-bit bus, delivering 288.4 GB/s of bandwidth. The GV100 uses 32 GB of HBM2 on a 4096-bit bus, delivering 868.4 GB/s. That is roughly three times the memory bandwidth, a massive advantage for memory-bound workloads.
Shader resources also favor the GV100. The M40 has 3,072 shading units, 192 TMUs, and 96 ROPs. The GV100 has 5,120 shading units, 320 TMUs, and 128 ROPs. Pixel rate is 106.8 GPixel/s on the M40 versus 208.3 GPixel/s on the GV100. Texture rate is 213.5 GTexel/s versus 520.6 GTexel/s. FP32 compute is 6.832 TFLOPS on the M40 versus 16.66 TFLOPS on the GV100. The GV100 also supports FP16 at 33.32 TFLOPS with a 2:1 ratio; the M40 has no recorded FP16 capability.
Both cards share the same TDP of 250 W, the same dual-slot form factor, the same 267 mm length, the same PCIe 3.0 x16 interface, and the same API support for DirectX 12 (12_1), OpenGL 4.6, and Vulkan 1.4. The GV100 adds 4x DisplayPort 1.4a outputs; the M40 has no display outputs at all.
FAQ
Q: Which card has more memory bandwidth?
A: The Quadro GV100 has 868.4 GB/s of bandwidth from 32 GB of HBM2 on a 4096-bit bus. The Tesla M40 24 GB has 288.4 GB/s from 24 GB of GDDR5 on a 384-bit bus.
Q: Does either card support ray tracing?
A: No. Neither the Tesla M40 24 GB nor the Quadro GV100 has any RT cores recorded in the database.
Q: What is the FP32 performance difference?
A: The Quadro GV100 delivers 16.66 TFLOPS of FP32 compute, while the Tesla M40 24 GB delivers 6.832 TFLOPS. The GV100 is roughly 2.4 times as fast in FP32.
Q: Do both cards have the same power draw?
A: Yes. Both the Tesla M40 24 GB and the Quadro GV100 have a TDP of 250 W and a suggested PSU of 600 W.
Q: Which card has tensor cores?
A: Only the Quadro GV100. It has 640 tensor cores. The Tesla M40 24 GB has no tensor cores.
Q: What display outputs does each card offer?
A: The Quadro GV100 has 4x DisplayPort 1.4a. The Tesla M40 24 GB has no display outputs.
Specification Differences
The recorded specifications show a clear generational leap. The key differences are as follows.
The chip and process differ: GM200 on 28 nm for the M40 versus GV100 on 12 nm for the Quadro. Transistor count is 8,000 million versus 21,100 million. Die size is 601 mm² versus 815 mm². Transistor density is 13.3M per mm² versus 25.9M per mm².
Clock speeds differ. The M40 has a base clock of 948 MHz and a boost of 1112 MHz. The GV100 has a base of 1132 MHz and a boost of 1627 MHz. Memory clocks are 1502 MHz (6 Gbps effective) on the M40 versus 848 MHz (1696 Mbps effective) on the GV100.
Memory configuration differs completely. The M40 uses 24 GB GDDR5 with a 384-bit bus and 288.4 GB/s bandwidth. The GV100 uses 32 GB HBM2 with a 4096-bit bus and 868.4 GB/s bandwidth.
Compute resources differ: 3,072 shading units, 192 TMUs, 96 ROPs on the M40 versus 5,120 shading units, 320 TMUs, 128 ROPs on the GV100. The GV100 adds 640 tensor cores; the M40 has none.
Output rates differ. Pixel rate is 106.8 GPixel/s versus 208.3 GPixel/s. Texture rate is 213.5 GTexel/s versus 520.6 GTexel/s. FP32 is 6.832 TFLOPS versus 16.66 TFLOPS. FP16 is absent on the M40 and 33.32 TFLOPS (2:1) on the GV100.
Power connectors differ: 8-pin EPS on the M40 versus 1x 8-pin on the GV100. Display outputs differ: none on the M40 versus 4x DisplayPort 1.4a on the GV100. The GV100 has a recorded height of 111 mm; the M40 has no height recorded. The GV100 has a launch MSRP of 8,999 USD; the M40 has no launch MSRP recorded.
Head-to-Head Benchmarks
Only two head-to-head benchmarks exist in the database, and the Quadro GV100 wins both decisively.
In Geekbench OpenCL, the GV100 scores 150,004 against the M40's 37,439. The delta is -75% from the GV100's perspective, meaning the M40 trails by three-quarters. That is a four-fold difference in raw compute throughput. OpenCL workloads such as general-purpose GPU compute, physics simulation, and data processing will run dramatically faster on the GV100.
In Geekbench Vulkan, the GV100 scores 139,526 against the M40's 45,975. The delta is -67%, meaning the M40 is about one-third of the GV100's performance. Vulkan is a modern graphics and compute API, and the gap reflects both the architectural advancement of Volta and the substantially higher shader and memory resources of the GV100.
The M40's own benchmark results, taken alone, are respectable: 37,439 in OpenCL and 45,975 in Vulkan place it in the 83rd percentile of all GPUs. Its nearest rivals include the NVIDIA Tesla M40 (average score 41,897, delta -0.5%), the GeForce RTX 3080 Ti (41,187, delta 1.3%), and the AMD Radeon Pro 5300 (40,870, delta 2%). These are all close scores, indicating the M40 performs around the level of a mid-range modern card in its recorded tests.
The GV100's nearest rivals are very different. It sits near the GeForce RTX 5070 Ti Mobile (35,435, delta 0.2%), the AMD Radeon Pro Duo (35,860, delta -0.9%), and the NVIDIA T1000 (36,289, delta -2.1%). These comparisons reflect the GV100's average across a wider set of benchmarks, including Passmark tests where it records lower scores such as 19650 in G3D and 9069 in GPU compute. Its Passmark DirectX scores are modest: 140 in DirectX 10, 168 in DirectX 11, 84 in DirectX 12, and 207 in DirectX 9. The Passmark G2D score is 836.
The overall picture is unambiguous. In every recorded comparison, the GV100 is faster, often by a wide margin. The M40 24 GB has no benchmark win in the database.
The Verdict
The data supports one clear conclusion: the NVIDIA Quadro GV100 is the superior card in every recorded benchmark. It wins both head-to-head tests, offers more than double the FP32 compute, nearly three times the memory bandwidth, and adds tensor cores that the Tesla M40 24 GB completely lacks. It is also a newer design, released in 2018 versus 2015 for the M40, and it carries a launch MSRP of 8,999 USD.
Who should pick the Tesla M40 24 GB? The M40 is the lower performer by every metric in the database, but it does have advantages that matter in specific contexts. It is an end-of-life product with no display outputs, so it is strictly a compute accelerator. Its 24 GB of GDDR5 memory is still substantial, and its 83rd percentile ranking shows it remains a capable card relative to the broader GPU landscape. For workloads that only need OpenCL or Vulkan compute and do not require tensor cores or FP16, the M40 can handle the job, but it will do so at a fraction of the GV100's speed.
Who should pick the Quadro GV100? Anyone whose workloads benefit from the recorded advantages: high FP32 throughput, FP16 compute, tensor cores, 32 GB of HBM2 memory, and 868.4 GB/s of bandwidth. The GV100's display outputs make it usable in workstation configurations where the M40 cannot drive a monitor at all. The GV100's 80th percentile ranking, despite a broader and harsher benchmark suite, confirms its strength.
The verdict is straightforward. The Quadro GV100 is the faster, more capable, and more modern card. The Tesla M40 24 GB is an older, slower accelerator that only makes sense when its specific limitations, such as the lack of tensor cores or the lower memory bandwidth, are acceptable for the task at hand. For any performance-sensitive workload, the GV100 is the only rational choice from the recorded data.