NVIDIA GB10 vs NVIDIA N1X 40SM Comparison
NVIDIA GB10
N1X 40SM
PERFORMANCE BENCHMARKS
Analysis: NVIDIA GB10 vs NVIDIA N1X 40SM
Head-to-Head Benchmarks
The recorded database contains no direct head-to-head benchmark results between the NVIDIA GB10 and the NVIDIA N1X 40SM. The GB10 has completed two standardized tests, while the N1X 40SM has no benchmark entries at all. This absence of direct comparison data means the analysis must rely on the available specifications and the GB10's measured performance against its nearest rivals.
The GB10 posts an OpenCL score of 120,137 and a Vulkan score of 114,648 in Geekbench testing. Its average benchmark score across all recorded tests is 117,393. This average places the GB10 at the 95th percentile among all GPUs in the database, indicating that its measured performance sits well above the majority of recorded graphics processors.
The nearest rival data for the GB10 provides context for interpreting these numbers. The AMD Radeon PRO W7700 holds an average score of 118,976, which is 1.3% higher than the GB10's average. The NVIDIA RTX 4000 SFF Ada Generation scores 117,088 on average, which is 0.3% lower than the GB10. The NVIDIA Tesla V100 SXM2 16 GB averages 114,395, placing it 2.6% behind the GB10. The NVIDIA RTX A5500 Mobile averages 113,944, which is 3% lower than the GB10's average.
These deltas show that the GB10 sits in a tight competitive cluster. Its performance is within a few percentage points of several established workstation and mobile GPUs, but it does not dominate any of them decisively. The largest advantage the GB10 holds over its nearest rivals is the 3% gap against the RTX A5500 Mobile, while its largest deficit is the 1.3% shortfall against the Radeon PRO W7700.
The N1X 40SM has no benchmark scores recorded in the database. Its average benchmark score is listed as zero, and its percentile ranking sits at the 50th percentile, which is the median position. This does not indicate that the N1X 40SM is an average performer; rather, it reflects the complete lack of measured data for this part. The GB10's 95th percentile ranking is based on actual test results, whereas the N1X 40SM's 50th percentile is a placeholder value assigned when no scores exist.
Given the absence of head-to-head tests, the specification sheet becomes the primary basis for comparing these two NVIDIA parts. Both chips share the same silicon: the GB20B die, fabricated on TSMC's 5 nm process, with a die size of 382 mm². The architectural foundation is identical, with both using Blackwell 2.0 architecture. The GB10 belongs to the Server Blackwell (Bxx) generation, while the N1X 40SM belongs to the Blackwell IGP (N1x) generation.
The GB10 carries 6,144 shading units, 384 texture mapping units, 48 raster output units, 48 ray tracing cores, and 384 tensor cores. The N1X 40SM reduces these counts to 5,120 shading units, 320 TMUs, 40 ROPs, 40 ray tracing cores, and 160 tensor cores. The GB10 therefore has 20% more shading units, 20% more TMUs, 20% more ROPs, and 20% more ray tracing cores than the N1X 40SM. The tensor core difference is larger: the GB10 has 384 tensor cores versus 160 on the N1X 40SM, which is a 140% advantage for the GB10.
Clock speeds also differ substantially. The GB10 has a base clock of 1,665 MHz and a boost clock of 2,418 MHz. The N1X 40SM runs at 741 MHz base and 2,346 MHz boost. The GB10's base clock is more than double that of the N1X 40SM, while its boost clock is 72 MHz higher. This clock advantage compounds the shading unit advantage, resulting in a significant throughput gap.
The pixel rate for the GB10 is 116.1 GPixel/s, compared to 93.84 GPixel/s for the N1X 40SM. This is a 23.7% higher pixel throughput for the GB10. The texture rate shows a similar pattern: 928.5 GTexel/s for the GB10 versus 750.7 GTexel/s for the N1X 40SM, a 23.7% difference. In FP32 compute, the GB10 delivers 29.71 TFLOPS, while the N1X 40SM delivers 24.02 TFLOPS, a 23.7% gap. The FP16 figures match the FP32 figures on both parts, with each listed at a 1:1 ratio, so the same 23.7% difference applies.
The Verdict
The data supports a clear separation between these two parts. The GB10 is the higher-performing option by every measurable compute metric in the specification sheet. Its shading unit count, texture rate, pixel rate, and FP32 throughput all exceed those of the N1X 40SM by roughly 20% to 24%. The tensor core count advantage is far larger, with the GB10 carrying more than double the tensor cores of the N1X 40SM.
The N1X 40SM's lower base clock of 741 MHz versus 1,665 MHz on the GB10 suggests that the N1X 40SM is designed for a different operating envelope, likely one that prioritizes power efficiency or thermal constraints over raw throughput. The GB10 has a recorded TDP of 140 W, while the N1X 40SM's TDP is listed as unknown in the database. Both parts use no power connectors and are classified as IGP (integrated graphics processor) form factors with the same PCIe 5.0 x16 bus interface.
For the GB10, the benchmark data confirms its position relative to the wider GPU landscape. Its 95th percentile ranking and average score of 117,393 place it among the upper tier of recorded GPUs. The nearest rival analysis shows that it trades the top spot with the Radeon PRO W7700, which leads by 1.3%, and the RTX 4000 SFF Ada Generation, which trails by only 0.3%. The GB10 is a competitive part in its segment, not an outlier in either direction.
For the N1X 40SM, the lack of benchmark data means no performance verdict can be drawn from measurements. The specification sheet shows a part that is clearly configured with fewer compute resources than the GB10. Every throughput metric is lower, and the tensor core count is dramatically lower. The N1X 40SM appears to be a more restrained implementation of the same GB20B die, likely intended for a narrower set of workloads or a lower thermal budget.
Users who require the highest available compute throughput from this silicon should select the GB10. Users who are evaluating the N1X 40SM should note that no measured performance data exists in the database, and its specification-derived performance is lower across all recorded metrics. The GB10 is the only one of the two with actual benchmark scores, and those scores place it at the 95th percentile. The N1X 40SM's 50th percentile placeholder is not a performance claim; it is a missing-data marker.
Architecture Differences
Both the NVIDIA GB10 and the NVIDIA N1X 40SM are built on the same GB20B chip, manufactured by TSMC on a 5 nm process, with a die size of 382 mm². The transistor count is listed as unknown for both parts. The architecture is identical: Blackwell 2.0. The generation labels differ, with the GB10 classified as Server Blackwell (Bxx) and the N1X 40SM as Blackwell IGP (N1x), but the underlying silicon is the same.
The memory subsystem is identical on both parts. Each uses 128 GB of LPDDR5X memory on a 256-bit bus, with a memory clock of 1,067 MHz running at 8.5 Gbps effective. This yields a memory bandwidth of 273.2 GB/s for both. The identical memory configuration means that memory-bound workloads will not differentiate between the two parts based on bandwidth alone; the difference will come from the compute resources that feed that memory.
The compute resources are where the two parts diverge. The GB10 has 6,144 shading units, 384 TMUs, 48 ROPs, 48 ray tracing cores, and 384 tensor cores. The N1X 40SM has 5,120 shading units, 320 TMUs, 40 ROPs, 40 ray tracing cores, and 160 tensor cores. The GB10 leads in every category, with the tensor core difference being the most pronounced.
Clock behavior also differs. The GB10's base clock is 1,665 MHz, which is unusually high for a base clock, suggesting that the part is designed to sustain high frequencies even under load. The N1X 40SM's base clock is 741 MHz, less than half of the GB10's. Boost clocks are closer: 2,418 MHz for the GB10 and 2,346 MHz for the N1X 40SM, a 72 MHz gap. The much lower base clock on the N1X 40SM implies that its sustained operating frequency is significantly below that of the GB10, which will affect any workload that runs for extended periods.
The pixel rate of the GB10 is 116.1 GPixel/s, derived from its 48 ROPs and boost clock. The N1X 40SM's pixel rate is 93.84 GPixel/s from 40 ROPs. Texture rates follow the same proportion: 928.5 GTexel/s versus 750.7 GTexel/s. FP32 compute is 29.71 TFLOPS for the GB10 and 24.02 TFLOPS for the N1X 40SM. Both parts list FP16 at the same figures as FP32, with a 1:1 ratio, meaning neither part offers a dedicated FP16 acceleration path beyond what the FP32 units provide.
The GB10 has a recorded TDP of 140 W and a suggested PSU rating of 300 W. The N1X 40SM has no TDP listed and no suggested PSU. Both are classified as IGP with no power connectors. The GB10 has physical dimensions of 150 mm length, 51 mm height, and 150 mm width. The N1X 40SM has no dimensions recorded. Both use the same PCIe 5.0 x16 bus interface and feature a single HDMI display output. Neither part supports DirectX, OpenGL, or Vulkan, as their API listings are all N/A.
The GB10 has a release date of October 14, 2025, and a launch MSRP of 3,999 USD. The N1X 40SM has a release date of May 31, 2026, and no launch MSRP. The GB10 lists its predecessor as Server Hopper and its successor as Server Rubin. The N1X 40SM has neither a predecessor nor a successor listed.
FAQ
Q: Which GPU has the higher FP32 compute performance?
A: The NVIDIA GB10 delivers 29.71 TFLOPS of FP32 compute, while the NVIDIA N1X 40SM delivers 24.02 TFLOPS. The GB10 leads by 23.7%.
Q: Do the two GPUs use the same memory configuration?
A: Yes. Both use 128 GB of LPDDR5X memory on a 256-bit bus with a memory clock of 1,067 MHz (8.5 Gbps effective), yielding 273.2 GB/s of bandwidth for each.
Q: What are the tensor core counts for each part?
A: The GB10 has 384 tensor cores, while the N1X 40SM has 160 tensor cores. The GB10 has more than double the tensor core count of the N1X 40SM.
Q: Does the N1X 40SM have any recorded benchmark scores?
A: No. The database lists no benchmark entries for the N1X 40SM, and its average benchmark score is recorded as zero. The GB10, by contrast, has OpenCL and Vulkan scores of 120,137 and 114,648, respectively.
Q: How does the GB10 compare to its nearest rivals in average score?
A: The GB10's average score of 117,393 is 0.3% higher than the RTX 4000 SFF Ada Generation (117,088), 1.3% lower than the AMD Radeon PRO W7700 (118,976), 2.6% higher than the Tesla V100 SXM2 16 GB (114,395), and 3% higher than the RTX A5500 Mobile (113,944).
Q: Which GPU has a higher base clock speed?
A: The GB10 has a base clock of 1,665 MHz, while the N1X 40SM has a base clock of 741 MHz. The GB10's base clock is more than twice that of the N1X 40SM.
Where Each One Wins
The GB10 wins in every compute category that has recorded data. Its shading unit count is 20% higher, its texture rate is 23.7% higher, its pixel rate is 23.7% higher, and its FP32 throughput is 23.7% higher. The tensor core gap is the single largest difference, with the GB10 holding 384 tensor cores versus 160 on the N1X 40SM, a 140% advantage. For any workload that depends on tensor operations, the GB10 is the clear choice based on the specification data.
The GB10 also wins on clock behavior. Its base clock of 1,665 MHz is far above the N1X 40SM's 741 MHz, which means the GB10 can sustain higher performance in steady-state workloads without relying on boost behavior. The boost clocks are closer, at 2,418 MHz versus 2,346 MHz, but the sustained frequency advantage of the GB10 is substantial.
The GB10 has a recorded TDP of 140 W and a suggested PSU of 300 W, indicating that it has a defined power envelope. The N1X 40SM has no TDP recorded, no suggested PSU, and no dimensions listed, which limits the characterization of its operating requirements. The GB10's earlier release date of October 2025 versus May 2026 for the N1X 40SM means the GB10 has a longer availability window in the database.
The N1X 40SM does not win in any measured benchmark category, because it has no measured benchmarks. Its only advantage in the recorded data is its lower base clock, which could indicate a lower power draw, but no TDP figure exists to confirm this. The N1X 40SM's median percentile placeholder of 50 is not a performance result, and it cannot be interpreted as a competitive position.
For rasterization workloads, the GB10's higher ROP count and pixel rate give it the advantage. For texture-heavy workloads, the GB10's higher TMU count and texture rate dominate. For compute workloads, the GB10's higher shading unit count and FP32 throughput lead. For tensor workloads, the GB10's tensor core count is decisively higher. In every scenario where the database provides data, the GB10 is the stronger part. The N1X 40SM remains an uncharacterized variant of the same die with fewer allocated resources.