NVIDIA L4 vs NVIDIA N1X 40SM Comparison

NVIDIA
GEFORCE

NVIDIA L4

CORE STATE AD104
VRAM 24 GB
CLOCK SPEED 2040 MHz
TDP 72 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

N1X 40SM

CORE STATE GB20B
VRAM 128 GB
CLOCK SPEED 2346 MHz
TDP unknown
BUS WIDTH 256 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2026

PERFORMANCE BENCHMARKS

geekbench_opencl
140,838
N/A
geekbench_vulkan
121,306
N/A

Analysis: NVIDIA L4 vs NVIDIA N1X 40SM

Head-to-Head Benchmarks

The recorded data presents an unusual comparison. The NVIDIA L4 has two benchmark entries in the database, while the NVIDIA N1X 40SM has none. This means there are no direct head-to-head results to walk through. The L4 delivers a Geekbench OpenCL score of 140,838 and a Geekbench Vulkan score of 121,306. Its average benchmark score across all tests is 131,072. The N1X 40SM, by contrast, has an average benchmark score of 0, indicating that no performance measurements have been recorded for it yet.

The L4's nearest rivals in the database provide context for its standing. The NVIDIA GeForce RTX 3090 Ti averages 131,938, which is 0.7% higher than the L4. The NVIDIA RTX 4000 Ada Generation averages 135,218, putting it 3.1% above the L4. The NVIDIA A10M also averages 135,230, again 3.1% higher. The AMD Radeon PRO W6800 averages 135,396, which is 3.2% above the L4. These deltas show the L4 sits slightly below a cluster of established workstation and prosumer cards, with the closest rival being the RTX 3090 Ti at less than one percent difference.

The N1X 40SM has no nearest rivals listed, no benchmark scores, and no delta percentages. The percentile data further separates the two. The L4 sits in the 95th percentile of all GPUs in the database, meaning it outperforms the vast majority of recorded graphics hardware. The N1X 40SM sits in the 50th percentile, which with a zero average score reflects an absence of recorded data rather than a measured performance level. The data indicates the L4 is a measurable, competitive product, while the N1X 40SM is essentially unquantified in this database.

The Verdict

The data strongly favors the NVIDIA L4 for any workload where performance has been recorded. Its average benchmark score of 131,072 places it in the 95th percentile of all GPUs. The OpenCL result of 140,838 and Vulkan result of 121,306 confirm that the L4 delivers substantial compute capability in both general-purpose and graphics API workloads. The nearest rival data shows it within 3.2% of several high-end cards, meaning it trades blows with the RTX 3090 Ti, RTX 4000 Ada Generation, A10M, and Radeon PRO W6800.

The NVIDIA N1X 40SM cannot be recommended for performance-sensitive tasks based on this database. With an average benchmark score of 0 and no recorded tests, there is no evidence of its compute capability. The 50th percentile ranking appears to be a placeholder rather than a measured outcome. Anyone selecting between these two for a workload that depends on recorded performance data should choose the L4 without hesitation.

For specific workloads, the L4's memory configuration matters. It carries 24 GB of GDDR6 memory on a 192-bit bus, delivering 300.1 GB/s of bandwidth. The N1X 40SM carries 128 GB of LPDDR5X memory on a 256-bit bus, delivering 273.2 GB/s of bandwidth. The N1X has more total memory but lower bandwidth. The L4 also boasts higher raw compute: 30.29 TFLOPS FP32 versus 24.02 TFLOPS FP32 for the N1X. The L4's shading units number 7,424 versus 5,120 for the N1X, and its RT cores at 60 exceed the N1X's 40. The L4 has 240 tensor cores versus 160 for the N1X. The texture rate favors the N1X at 750.7 GTexel/s versus 489.6 GTexel/s for the L4, but the pixel rate favors the L4 at 163.2 GPixel/s versus 93.84 GPixel/s.

The N1X is an IGP (integrated graphics processor) with a single HDMI output, while the L4 has no display outputs and is designed as a single-slot server accelerator. The N1X uses the PCIe 5.0 x16 interface, while the L4 uses PCIe 4.0 x16. The L4 has a known TDP of 72 W, while the N1X's TDP is unknown. The L4 is an active production product released in March 2023. The N1X is also active in production with a release date of May 2026.

FAQ

Q: Which GPU has the higher average benchmark score?

A: The NVIDIA L4 has an average benchmark score of 131,072. The NVIDIA N1X 40SM has an average benchmark score of 0, meaning no benchmark results have been recorded for it.

Q: How does the L4 compare to its closest rival in the database?

A: The L4's closest rival is the NVIDIA GeForce RTX 3090 Ti, which has an average score of 131,938. That is 0.7% higher than the L4's average score of 131,072.

Q: What are the memory capacities of the two GPUs?

A: The NVIDIA L4 has 24 GB of GDDR6 memory with a 192-bit bus. The NVIDIA N1X 40SM has 128 GB of LPDDR5X memory with a 256-bit bus.

Q: Which GPU delivers higher FP32 compute throughput?

A: The NVIDIA L4 delivers 30.29 TFLOPS FP32. The NVIDIA N1X 40SM delivers 24.02 TFLOPS FP32, which is lower.

Q: Do the two GPUs support the same graphics APIs?

A: No. The NVIDIA L4 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The NVIDIA N1X 40SM lists N/A for DirectX, OpenGL, and Vulkan.

Q: What are the pixel and texture rates for each GPU?

A: The NVIDIA L4 has a pixel rate of 163.2 GPixel/s and a texture rate of 489.6 GTexel/s. The NVIDIA N1X 40SM has a pixel rate of 93.84 GPixel/s and a texture rate of 750.7 GTexel/s.

Specification Differences

The two GPUs differ across nearly every measured specification. The process node is the same, both use 5 nm TSMC fabrication, but the chip designs diverge. The L4 uses the AD104 chip with 35,800 million transistors on a 294 mm² die, giving a transistor density of 121.8M per mm². The N1X 40SM uses the GB20B chip with an unknown transistor count on a 382 mm² die, and no transistor density is recorded.

Clock speeds differ substantially. The L4 has a base clock of 795 MHz and a boost clock of 2040 MHz. The N1X 40SM has a lower base clock of 741 MHz but a higher boost clock of 2346 MHz. Memory clocks also differ: the L4 runs at 1563 MHz with 12.5 Gbps effective data rate, while the N1X runs at 1067 MHz with 8.5 Gbps effective.

The memory subsystem presents a clear trade-off. The L4 has 24 GB of GDDR6 on a 192-bit bus with 300.1 GB/s bandwidth. The N1X has 128 GB of LPDDR5X on a 256-bit bus with 273.2 GB/s bandwidth. The N1X offers over five times the capacity but lower bandwidth.

Compute unit counts favor the L4 in most categories. The L4 has 7,424 shading units, 240 TMUs, 80 ROPs, 60 RT cores, and 240 tensor cores. The N1X has 5,120 shading units, 320 TMUs, 40 ROPs, 40 RT cores, and 160 tensor cores. The N1X has more TMUs but fewer of everything else.

Rates and throughput differ as well. The L4 achieves 163.2 GPixel/s pixel rate and 489.6 GTexel/s texture rate. The N1X achieves 93.84 GPixel/s and 750.7 GTexel/s. FP32 and FP16 performance are identical within each GPU at a 1:1 ratio, with the L4 at 30.29 TFLOPS and the N1X at 24.02 TFLOPS.

Power and physical characteristics are distinct. The L4 has a TDP of 72 W and a suggested PSU of 250 W, fits in a single slot, and measures 169 mm in length and 56 mm in height. The N1X has an unknown TDP, no suggested PSU, is classified as an IGP, and has no recorded dimensions. The L4 uses PCIe 4.0 x16, while the N1X uses PCIe 5.0 x16. The L4 has no display outputs; the N1X has one HDMI output. Neither uses power connectors.

Architecture Differences

The architectural split is fundamental. The L4 is built on Ada Lovelace architecture under the Server Ada generation, while the N1X 40SM uses Blackwell 2.0 under the Blackwell IGP generation. The L4's predecessor is listed as Server Ampere and its successor as Server Hopper. The N1X has no predecessor or successor recorded.

The L4 is a discrete server accelerator. Its Ada Lovelace architecture is designed for data center inference and rendering workloads, with no display outputs and a single-slot form factor. The N1X is an integrated graphics processor, meaning it is designed to be built into a system rather than installed as a separate card. It has one HDMI output, confirming its role as a display-capable IGP.

The API support reveals another architectural difference. The L4 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, making it suitable for graphics applications that rely on modern rendering APIs. The N1X lists N/A for all three APIs, suggesting it may not expose traditional graphics APIs or that its drivers are not yet validated for them in the database.

The chip designs reflect different priorities. The L4 uses AD104, a chip with a known transistor count of 35,800 million and a die size of 294 mm². The N1X uses GB20B with an unknown transistor count and a larger die size of 382 mm². The larger die with fewer shading units (5,120 versus 7,424) suggests the N1X dedicates more area to memory or other integrated functions. The N1X's 128 GB LPDDR5X memory is far beyond what a typical discrete GPU carries, indicating the memory is likely shared with a host system rather than exclusively dedicated to graphics.

The N1X's higher boost clock of 2346 MHz versus the L4's 2040 MHz does not compensate for its lower FP32 throughput. The L4's 30.29 TFLOPS versus the N1X's 24.02 TFLOPS confirms that the Ada Lovelace chip packs more compute per clock. The RT core counts, 60 versus 40, and tensor core counts, 240 versus 160, both favor the L4, indicating stronger ray tracing and AI acceleration capabilities.

The texture rate discrepancy stands out. The N1X's 320 TMUs produce 750.7 GTexel/s, while the L4's 240 TMUs produce 489.6 GTexel/s. This means the N1X is more efficient at texture mapping despite having fewer shading units. The pixel rate goes the other way: the L4's 80 ROPs deliver 163.2 GPixel/s versus the N1X's 40 ROPs at 93.84 GPixel/s. The N1X's ROP count is half the L4's, which limits its fill-rate performance.

Where Each One Wins

The L4 wins in compute throughput. Its FP32 performance of 30.29 TFLOPS is 26% higher than the N1X's 24.02 TFLOPS. The shading unit advantage, 7,424 versus 5,120, gives it more parallel execution lanes. The tensor core advantage, 240 versus 160, makes it better suited for AI and deep learning inference. The RT core advantage, 60 versus 40, gives it stronger ray tracing capability.

The L4 also wins in memory bandwidth. Its 300.1 GB/s exceeds the N1X's 273.2 GB/s, meaning it can feed its compute units faster. The L4's pixel rate of 163.2 GPixel/s is 74% higher than the N1X's 93.84 GPixel/s, making it better for rasterization-heavy workloads. The L4's 95th percentile ranking and measured benchmark scores confirm its real-world performance advantage.

The N1X 40SM wins in memory capacity. Its 128 GB of LPDDR5X is more than five times the L4's 24 GB. For workloads that require holding very large datasets on the GPU, such as certain inference or simulation tasks, this capacity advantage matters. The N1X also wins in texture throughput, with 750.7 GTexel/s versus 489.6 GTexel/s, and in texture mapping units, 320 versus 240.

The N1X's higher boost clock of 2346 MHz versus 2040 MHz indicates it can reach higher peak frequencies, but the benchmark data does not show how this translates into real performance. The N1X's PCIe 5.0 x16 interface is a newer standard than the L4's PCIe 4.0 x16, potentially offering higher host transfer rates. The N1X has a display output while the L4 has none, making it the only one of the two that can drive a monitor.

The L4 wins on power efficiency as a measured quantity. Its TDP of 72 W is known and extremely low for its performance level. The N1X's TDP is unknown, so no efficiency comparison can be made from recorded data. The L4's suggested PSU of 250 W indicates it can run in systems with modest power supplies.

For inference workloads, the L4's tensor cores and higher FP16 throughput make it the stronger choice. For memory-bound tasks requiring over 24 GB of capacity, the N1X's 128 GB is unmatched. For general compute, the L4's recorded scores and 95th percentile ranking provide concrete evidence of capability, while the N1X's zero benchmark score leaves its performance entirely undocumented. The data points to the L4 as the measured performer and the N1X as an unverified quantity with specific capacity advantages.

DETAILED SPECIFICATIONS

SPECIFICATION
L4
N1X 40SM
Core Specs
Shading Units
7,424
5,120 -31.0%
Shaders
7,424
5,120 -31.0%
TMUs
240
320 +33.3%
ROPs
80
40 -50.0%
SM Count
60
40 -33.3%
Clocks
Base Clock
795 MHz
741 MHz
Boost Clock
2040 MHz
2346 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
1067 MHz 8.5 Gbps effective
Memory
Memory Size
24 GB
128 GB
VRAM (MB)
24,576
131,072 +433.3%
Memory Type
GDDR6
LPDDR5X
Memory Bus
192 bit
256 bit
Bandwidth
300.1 GB/s
273.2 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
48 MB
50 MB
Performance
Pixel Rate
163.2 GPixel/s
93.84 GPixel/s
Texture Rate
489.6 GTexel/s
750.7 GTexel/s
FP32 (TFLOPS)
30.29 TFLOPS
24.02 TFLOPS
FP64 (TFLOPS)
473.3 GFLOPS (1:64)
375.4 GFLOPS (1:64)
FP16 (TFLOPS)
30.29 TFLOPS (1:1)
24.02 TFLOPS (1:1)
AI/RT
RT Cores
60
40 -33.3%
Tensor Cores
240
160 -33.3%
Power
TDP
72 W
unknown
TDP (W)
72
Suggested PSU
250 W
Power Connectors
None
None
Architecture
Architecture
Ada Lovelace
Blackwell 2.0
GPU Name
AD104
GB20B
Generation
Server Ada (Lxx)
Blackwell IGP (N1x)
Process Size
5 nm
5 nm
Transistors
35,800 million
unknown
Die Size
294 mm²
382 mm²
Foundry
TSMC
TSMC
Density
121.8M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
8.9
12.1
Shader Model
6.8
Physical
Slot Width
Single-slot
IGP
Length
169 mm 6.7 inches
Height
56 mm 2.2 inches
Outputs
No outputs
1x HDMI
Bus Interface
PCIe 4.0 x16
PCIe 5.0 x16
Other
Production
Active
Active
Predecessor
Server Ampere
Successor
Server Hopper
View L4 Details View N1X 40SM Details