NVIDIA L4 vs NVIDIA N1X 48SM Comparison

NVIDIA
GEFORCE

NVIDIA L4

CORE STATE AD104
VRAM 24 GB
CLOCK SPEED 2040 MHz
TDP 72 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

N1X 48SM

CORE STATE GB20B
VRAM 128 GB
CLOCK SPEED 2346 MHz
TDP unknown
BUS WIDTH 256 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2026

PERFORMANCE BENCHMARKS

geekbench_opencl
140,838
N/A
geekbench_vulkan
121,306
N/A

Analysis: NVIDIA L4 vs NVIDIA N1X 48SM

The NVIDIA L4 and the NVIDIA N1X 48SM are both active server-oriented GPUs, but they occupy different positions in the database. The L4 has a full set of recorded benchmark scores, ranking in the 95th percentile of all GPUs, while the N1X 48SM has no recorded benchmarks yet and sits at the 50th percentile with an average score of zero. The data shows a clear split in capability and purpose, with the L4 being a proven, measurable compute device and the N1X 48SM representing a larger, less-tested design.

Where Each One Wins

The L4 wins in every measured benchmark category because it is the only one with recorded scores. Its Geekbench OpenCL score is 140,838 and its Vulkan score is 121,306. These results place it just 0.7% behind the NVIDIA GeForce RTX 3090 Ti, which has an average score of 131,938, and 3.1% behind both the NVIDIA RTX 4000 Ada Generation and the NVIDIA A10M. The L4’s average benchmark score of 131,072 confirms it as a high-performing part, especially for its 72 W TDP and single-slot design.

The N1X 48SM has no benchmark scores, so it cannot win any direct performance comparisons. Instead, its advantages are structural. It carries 128 GB of LPDDR5X memory on a 256-bit bus, which is far larger than the L4’s 24 GB of GDDR6 on a 192-bit bus. The N1X 48SM also supports PCIe 5.0 x16, double the bandwidth generation of the L4’s PCIe 4.0 x16. This suggests the N1X 48SM is designed for capacity-heavy workloads, such as large model residency or high-bandwidth data access, rather than raw shader throughput.

The use-case split is therefore straightforward: the L4 wins in compute speed and efficiency metrics that are already validated, while the N1X 48SM wins in memory capacity, interface generation, and the ability to handle larger datasets. The L4 is a tested accelerator for immediate deployment, whereas the N1X 48SM is a forward-looking part with untested performance but substantially more memory.

Architecture Differences

The two GPUs come from different architectural generations. The L4 is built on Ada Lovelace, using the AD104 chip, fabricated on a 5 nm process at TSMC with 35,800 million transistors on a 294 mm² die. The N1X 48SM uses the Blackwell 2.0 architecture, based on the GB20B chip, also on a 5 nm process at TSMC, but with a larger 382 mm² die and an unknown transistor count. The L4’s transistor density is 121.8 million per mm², while the N1X 48SM has no recorded density figure.

The L4 is described as a server Ada part, while the N1X 48SM is a Blackwell IGP (integrated graphics processor) part, which aligns with its lack of a dedicated power connector and its IGP slot width. The L4 has no display outputs, while the N1X 48SM includes a single HDMI output, indicating that the latter can serve as a display-capable solution despite its server classification.

In terms of compute resources, the L4 has more shading units (7,424 vs. 6,144) and more tensor cores (240 vs. 192), but fewer texture mapping units (240 vs. 384) and fewer ROPs (80 vs. 48). The L4 also has more RT cores (60 vs. 48). The N1X 48SM’s higher texture rate (900.9 GTexel/s vs. 489.6 GTexel/s) and lower pixel rate (112.6 GPixel/s vs. 163.2 GPixel/s) reflect this different balance of resources.

The memory subsystem differs fundamentally: the L4 uses GDDR6 with a 300.1 GB/s bandwidth, while the N1X 48SM uses LPDDR5X with a 273.2 GB/s bandwidth. The N1X 48SM’s memory clock is lower (1067 MHz, 8.5 Gbps effective) compared to the L4’s 1563 MHz, 12.5 Gbps effective, but the N1X 48SM’s wider 256-bit bus compensates partially, though not fully, for the bandwidth gap.

Head-to-Head Benchmarks

There are no direct head-to-head benchmark results recorded between these two parts. The database lists zero wins for each side and no head-to-head tests. This means the only comparative numeric evidence comes from the L4’s own scores and its nearest rivals, which do not include the N1X 48SM. The L4’s Geekbench OpenCL score of 140,838 and Vulkan score of 121,306 stand as the sole measured performance indicators.

Relative to its nearest rivals, the L4 trails the GeForce RTX 3090 Ti by 0.7%, the RTX 4000 Ada Generation by 3.1%, and the A10M by 3.1%, and it is 3.2% behind the AMD Radeon PRO W6800. These narrow margins suggest the L4 is competitive within its immediate performance class, despite its much lower 72 W power draw compared to typical high-end parts.

For the N1X 48SM, the absence of benchmark data means its performance can only be inferred from its specifications. Its FP32 throughput is 28.83 TFLOPS, which is close to the L4’s 30.29 TFLOPS, and its FP16 throughput is identical at 28.83 TFLOPS, also matching in a 1:1 ratio. The L4’s higher shading unit count and boost clock (2040 MHz vs. 2346 MHz for the N1X 48SM) contribute to its slightly higher peak FP32 figure, but the difference is under 5%.

The N1X 48SM’s texture rate is nearly double that of the L4 (900.9 vs. 489.6 GTexel/s), which indicates a design optimized for texturing workloads. Its lower pixel rate (112.6 vs. 163.2 GPixel/s) shows that fill-rate-heavy tasks would favor the L4. Without recorded scores, these specification-derived differences are the only way to compare them.

Specification Differences

The two GPUs differ across nearly every major specification field. The L4 uses the AD104 chip with Ada Lovelace architecture, while the N1X 48SM uses the GB20B chip with Blackwell 2.0 architecture. The L4 has a transistor count of 35,800 million, whereas the N1X 48SM’s transistor count is unknown. The die size is 294 mm² for the L4 and 382 mm² for the N1X 48SM.

The L4’s base clock is 795 MHz and boost clock is 2040 MHz, compared to the N1X 48SM’s 741 MHz base and 2346 MHz boost. Memory specifications diverge sharply: the L4 has 24 GB of GDDR6 on a 192-bit bus with 300.1 GB/s bandwidth, while the N1X 48SM has 128 GB of LPDDR5X on a 256-bit bus with 273.2 GB/s bandwidth.

Compute unit counts favor the L4 in shading units (7,424 vs. 6,144), ROPs (80 vs. 48), RT cores (60 vs. 48), and tensor cores (240 vs. 192). The N1X 48SM has more TMUs (384 vs. 240). Pixel rates are 163.2 GPixel/s for the L4 and 112.6 GPixel/s for the N1X 48SM. Texture rates are 489.6 GTexel/s for the L4 and 900.9 GTexel/s for the N1X 48SM.

Power and physical design differ as well. The L4 has a TDP of 72 W, a single-slot form factor, and a suggested PSU of 250 W. The N1X 48SM has an unknown TDP, an IGP slot width, and no suggested PSU. The L4 has no display outputs, while the N1X 48SM has one HDMI output. The bus interface is PCIe 4.0 x16 for the L4 and PCIe 5.0 x16 for the N1X 48SM.

API support is another differentiator: the L4 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, while the N1X 48SM records N/A for all three. The L4’s dimensions are 169 mm in length and 56 mm in height, while the N1X 48SM has no recorded dimensions. The L4 was released on March 20, 2023, and the N1X 48SM on May 31, 2026.

FAQ

Q: Which GPU has more memory?

A: The NVIDIA N1X 48SM has 128 GB of LPDDR5X memory, which is substantially more than the NVIDIA L4’s 24 GB of GDDR6.

Q: What is the difference in shader throughput?

A: The L4 has a higher FP32 throughput of 30.29 TFLOPS compared to the N1X 48SM’s 28.83 TFLOPS, though the N1X 48SM has a higher boost clock.

Q: Which GPU has better texture processing capability?

A: The N1X 48SM has a texture rate of 900.9 GTexel/s, nearly double the L4’s 489.6 GTexel/s, due to its larger number of texture mapping units.

Q: Do both GPUs support the same APIs?

A: No. The L4 supports DirectX 12 Ultimate, OpenGL 4.6, and Vulkan 1.4, while the N1X 48SM records N/A for all of these APIs.

Q: How does the L4 compare to its nearest rivals?

A: The L4 is 0.7% slower than the GeForce RTX 3090 Ti and 3.1% slower than both the RTX 4000 Ada Generation and the A10M, based on average benchmark scores.

Q: What is the memory bandwidth difference?

A: The L4 has a higher memory bandwidth of 300.1 GB/s, while the N1X 48SM has 273.2 GB/s, despite the latter’s wider 256-bit bus.

The Verdict

The data supports a clear division of roles. The NVIDIA L4 is the tested and validated performer, with an average benchmark score of 131,072, a 95th percentile ranking, and competitive results against the GeForce RTX 3090 Ti and other near-peers. Its 72 W power draw and single-slot design make it a practical choice for dense server deployments where compute speed and efficiency are known quantities.

The NVIDIA N1X 48SM is a different kind of product. With no recorded benchmarks and a 50th percentile ranking, it offers no measured performance evidence. Its advantages are structural: 128 GB of memory, PCIe 5.0 x16 support, and a higher texture rate. These features point toward workloads that require large memory footprints or heavy texturing, but the lack of API support and benchmark data leaves its real-world behavior unverified.

For users who need a working, high-percentile accelerator with proven compute scores, the L4 is the only option with recorded data. For users who prioritize memory capacity and interface generation, the N1X 48SM presents a larger, unproven alternative. The choice hinges entirely on whether measured performance or raw capacity is the primary requirement.

DETAILED SPECIFICATIONS

SPECIFICATION
L4
N1X 48SM
Core Specs
Shading Units
7,424
6,144 -17.2%
Shaders
7,424
6,144 -17.2%
TMUs
240
384 +60.0%
ROPs
80
48 -40.0%
SM Count
60
48 -20.0%
Clocks
Base Clock
795 MHz
741 MHz
Boost Clock
2040 MHz
2346 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
1067 MHz 8.5 Gbps effective
Memory
Memory Size
24 GB
128 GB
VRAM (MB)
24,576
131,072 +433.3%
Memory Type
GDDR6
LPDDR5X
Memory Bus
192 bit
256 bit
Bandwidth
300.1 GB/s
273.2 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
48 MB
50 MB
Performance
Pixel Rate
163.2 GPixel/s
112.6 GPixel/s
Texture Rate
489.6 GTexel/s
900.9 GTexel/s
FP32 (TFLOPS)
30.29 TFLOPS
28.83 TFLOPS
FP64 (TFLOPS)
473.3 GFLOPS (1:64)
450.4 GFLOPS (1:64)
FP16 (TFLOPS)
30.29 TFLOPS (1:1)
28.83 TFLOPS (1:1)
AI/RT
RT Cores
60
48 -20.0%
Tensor Cores
240
192 -20.0%
Power
TDP
72 W
unknown
TDP (W)
72
Suggested PSU
250 W
Power Connectors
None
None
Architecture
Architecture
Ada Lovelace
Blackwell 2.0
GPU Name
AD104
GB20B
Generation
Server Ada (Lxx)
Blackwell IGP (N1x)
Process Size
5 nm
5 nm
Transistors
35,800 million
unknown
Die Size
294 mm²
382 mm²
Foundry
TSMC
TSMC
Density
121.8M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
8.9
12.1
Shader Model
6.8
Physical
Slot Width
Single-slot
IGP
Length
169 mm 6.7 inches
Height
56 mm 2.2 inches
Outputs
No outputs
1x HDMI
Bus Interface
PCIe 4.0 x16
PCIe 5.0 x16
Other
Production
Active
Active
Predecessor
Server Ampere
Successor
Server Hopper
View L4 Details View N1X 48SM Details