NVIDIA GB10 vs NVIDIA L40 Comparison
NVIDIA GB10
L40
PERFORMANCE BENCHMARKS
Analysis: NVIDIA GB10 vs NVIDIA L40
Where Each One Wins
The benchmark data splits these two NVIDIA server parts cleanly by workload class. The NVIDIA L40 wins both recorded tests outright, with no test in the database where the GB10 takes the lead. That gives the L40 a 2-0 win record in head-to-head comparisons.
The L40’s wins come in OpenCL and Vulkan compute workloads. Its Geekbench OpenCL score of 330,926 dwarfs the GB10’s 120,137, a 175.5% advantage. In Vulkan, the L40 posts 237,295 against 114,648, a 107% lead. These are not narrow margins; the L40 essentially doubles the GB10’s output in both API environments.
The GB10, however, occupies a different tier entirely. Its average benchmark score of 117,393 places it at the 95th percentile of all GPUs, while the L40 sits at the 99th percentile. The GB10’s nearest rivals are the NVIDIA RTX 4000 SFF Ada Generation at 117,088 (0.3% behind), the AMD Radeon PRO W7700 at 118,976 (1.3% ahead), the NVIDIA Tesla V100 SXM2 16 GB at 114,395 (2.6% behind), and the NVIDIA RTX A5500 Mobile at 113,944 (3% behind). This is a workstation-class performance envelope.
The L40’s rivals sit much higher. The NVIDIA RTX 6000 Ada Generation averages 287,237 (1.1% behind the L40), the NVIDIA L40S averages 295,763 (3.9% ahead), the NVIDIA L20 averages 251,147 (13.1% behind), and the AMD Instinct MI300X averages 317,994 (10.7% ahead). The L40 competes at the top of the server accelerator stack.
So the use-case split is stark: the L40 is for heavy compute, rendering, and AI inference where raw throughput matters, while the GB10 is a compact, low-power entry point for lighter workloads that still demand respectable performance.
Architecture Differences
The two chips come from different NVIDIA generations. The L40 uses the AD102 chip built on the Ada Lovelace architecture, while the GB10 uses the GB20B chip on the Blackwell 2.0 architecture. Both are fabricated by TSMC on a 5 nm process, but the design philosophies diverge sharply.
The L40’s AD102 die measures 609 mm² and packs 76,300 million transistors, giving a transistor density of 125.3 million per square millimeter. The GB10’s GB20B die is much smaller at 382 mm², and its transistor count is listed as unknown in the database.
The compute resources tell the story. The L40 has 18,176 shading units, 568 texture mapping units, 192 raster output units, 142 ray tracing cores, and 568 tensor cores. The GB10 has 6,144 shading units, 384 TMUs, 48 ROPs, 48 ray tracing cores, and 384 tensor cores. That is roughly one-third the shading units, two-thirds the TMUs, one-quarter the ROPs, and about one-third the ray tracing cores.
Clock behavior also differs. The L40’s base clock is 735 MHz with a boost of 2,490 MHz. The GB10 starts at 1,665 MHz base and boosts to 2,418 MHz. The GB10 runs a much higher base clock, but the boost ceiling is nearly identical, within 72 MHz.
Memory architecture diverges completely. The L40 uses 48 GB of GDDR6 on a 384-bit bus with 864.0 GB/s of bandwidth. The GB10 uses 128 GB of LPDDR5X on a 256-bit bus with 273.2 GB/s. The GB10 has nearly three times the capacity but the L40 has more than three times the bandwidth.
The GB10’s API support is listed as unavailable for DirectX, OpenGL, and Vulkan, which suggests it is not intended for traditional graphics workloads. The L40 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.
Head-to-Head Benchmarks
The Geekbench OpenCL result is the single largest gap in the comparison. The L40 scores 330,926 against the GB10’s 120,137, a 175.5% delta. This means the L40 delivers roughly 2.75 times the OpenCL compute throughput of the GB10. For context, the L40’s nearest rival, the RTX 6000 Ada Generation, is only 1.1% behind, so this is not a case of the L40 being an outlier; it is simply a much higher-tier part.
The Vulkan test shows a smaller but still commanding gap. The L40 posts 237,295, while the GB10 manages 114,648, a 107% delta. The GB10’s Vulkan score is actually lower than its OpenCL score, while the L40’s Vulkan score is about 28% lower than its OpenCL result. This suggests the L40’s advantage grows in compute-heavy OpenCL workloads.
Looking at the GB10’s position relative to its own rivals helps contextualize its scores. Its average benchmark score of 117,393 is within 3% of the RTX 4000 SFF Ada Generation, the Radeon PRO W7700, the Tesla V100 SXM2 16 GB, and the RTX A5500 Mobile. The GB10 is not a weak performer; it is just in a completely different performance class from the L40.
The L40’s average benchmark score of 284,111 puts it 13.1% ahead of the NVIDIA L20 and 10.7% behind the AMD Instinct MI300X. The L40S leads it by 3.9%. These are the margins at the top of the stack, and the L40 holds its own.
Specification Differences
The two cards differ across nearly every major specification field. Process node and foundry are the same (5 nm, TSMC), but almost everything else diverges.
The L40 uses the Ada Lovelace architecture with the AD102 chip; the GB10 uses Blackwell 2.0 with the GB20B chip. The L40’s die is 609 mm² with 76,300 million transistors; the GB10’s die is 382 mm² with unknown transistor count. Transistor density is 125.3M per mm² for the L40, not listed for the GB10.
Clock speeds differ in base but converge at boost: 735 MHz base and 2,490 MHz boost for the L40, versus 1,665 MHz base and 2,418 MHz boost for the GB10. Memory clocks are 2,250 MHz (18 Gbps effective) for the L40, versus 1,067 MHz (8.5 Gbps effective) for the GB10.
Memory configuration is a major split: 48 GB GDDR6 on a 384-bit bus with 864.0 GB/s bandwidth for the L40, versus 128 GB LPDDR5X on a 256-bit bus with 273.2 GB/s for the GB10.
Compute units: the L40 has 18,176 shading units, 568 TMUs, 192 ROPs, 142 RT cores, and 568 tensor cores. The GB10 has 6,144 shading units, 384 TMUs, 48 ROPs, 48 RT cores, and 384 tensor cores.
Pixel rate is 478.1 GPixel/s for the L40 versus 116.1 GPixel/s for the GB10. Texture rate is 1,414.3 GTexel/s versus 928.5 GTexel/s. FP32 and FP16 performance are both 90.52 TFLOPS for the L40, both 29.71 TFLOPS for the GB10, with a 1:1 ratio on both.
Power and physical design differ substantially. The L40 has a 300 W TDP, dual-slot form factor, 1x 16-pin power connector, and requires a 700 W suggested PSU. The GB10 has a 140 W TDP, IGP form factor, no power connectors, and a 300 W suggested PSU. The L40 is 267 mm long and 111 mm tall; the GB10 is 150 mm by 51 mm.
Interface and outputs also differ. The L40 uses PCIe 4.0 x16 with 4x DisplayPort 1.4a. The GB10 uses PCIe 5.0 x16 with 1x HDMI. API support: the L40 has DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4; the GB10 lists none.
Production status and release dates differ. The L40 is end-of-life, released 2022-10-12. The GB10 is active, released 2025-10-14. The L40’s predecessor is Server Ampere and successor is Server Hopper. The GB10’s predecessor is Server Hopper and successor is Server Rubin. The GB10 has a launch MSRP of 3,999 USD.
FAQ
Q: Which GPU has higher raw compute throughput?
A: The NVIDIA L40. Its FP32 and FP16 performance are both 90.52 TFLOPS, while the GB10 delivers 29.71 TFLOPS in both. The L40’s OpenCL score is 175.5% higher than the GB10’s.
Q: How do the memory configurations compare?
A: The GB10 has 128 GB of LPDDR5X on a 256-bit bus with 273.2 GB/s bandwidth. The L40 has 48 GB of GDDR6 on a 384-bit bus with 864.0 GB/s bandwidth. The GB10 has more capacity, but the L40 has more than three times the bandwidth.
Q: What is the performance gap in the recorded benchmarks?
A: In OpenCL, the L40 scores 330,926 versus 120,137 for the GB10, a 175.5% delta. In Vulkan, the L40 scores 237,295 versus 114,648, a 107% delta. The L40 wins both tests.
Q: Which GPU is positioned closer to its direct competitors?
A: The GB10. Its average score of 117,393 is within 3% of its nearest rivals, including the RTX 4000 SFF Ada Generation (0.3% behind) and the Radeon PRO W7700 (1.3% ahead). The L40’s nearest rival, the RTX 6000 Ada Generation, is only 1.1% behind, but the L40S is 3.9% ahead of it.
Q: Do these GPUs support the same graphics APIs?
A: No. The L40 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The GB10 lists no API support for DirectX, OpenGL, or Vulkan, indicating it is not designed for traditional graphics rendering.
Q: What are the power requirements?
A: The L40 has a 300 W TDP and requires a 700 W suggested PSU. The GB10 has a 140 W TDP and requires a 300 W suggested PSU. The GB10 uses no power connectors, while the L40 uses a single 16-pin connector.
The Verdict
The data points to a clear division of purpose. The NVIDIA L40 is the choice for workloads that demand maximum compute throughput, high memory bandwidth, and full graphics API support. Its 90.52 TFLOPS FP32 performance, 864.0 GB/s bandwidth, and 99th percentile standing among all GPUs make it a serious server accelerator. It wins both head-to-head benchmarks by margins of 175.5% and 107%, and it stands within 3.9% of the stronger L40S while beating the L20 by 13.1%. Users who need to process large datasets, run heavy CUDA or OpenCL kernels, or drive multiple displays will find the L40’s 4x DisplayPort outputs and 48 GB GDDR6 memory suited to the task.
The NVIDIA GB10, by contrast, is a different class of product. Its 128 GB of LPDDR5X memory is double-plus the capacity of the L40, and its 140 W TDP with no power connectors makes it far easier to integrate into compact systems. Its 29.71 TFLOPS FP32 and 273.2 GB/s bandwidth are modest next to the L40, but its average benchmark score of 117,393 places it at the 95th percentile overall, within 3% of the RTX 4000 SFF Ada Generation and the Radeon PRO W7700. The GB10’s PCIe 5.0 x16 interface is a generation ahead of the L40’s PCIe 4.0.
Who should pick which? The L40 suits environments where absolute performance is the priority and power draw, physical size, and end-of-life status are acceptable trade-offs. The GB10 suits deployments where memory capacity, low power, compact size, and an active production status matter more than raw speed, and where the lack of graphics API support is not a limitation. The benchmark record shows the L40 is the compute king, but the GB10’s 95th percentile standing proves it is far from a weak option. The choice hinges not on which is better overall, but on which fits the workload. For raw throughput, the L40. For capacity and efficiency, the GB10.