NVIDIA L4 vs NVIDIA N1 16SM Comparison

NVIDIA
GEFORCE

NVIDIA L4

CORE STATE AD104
VRAM 24 GB
CLOCK SPEED 2040 MHz
TDP 72 W
BUS WIDTH 192 bit
ARCHITECTURE Ada Lovelace
nm
PROCESS 5 nm
LAUNCH DATE 2023
VS
NVIDIA
GEFORCE

N1 16SM

CORE STATE GB20B
VRAM 128 GB
CLOCK SPEED 2346 MHz
TDP unknown
BUS WIDTH 256 bit
ARCHITECTURE Blackwell 2.0
nm
PROCESS 5 nm
LAUNCH DATE 2026

PERFORMANCE BENCHMARKS

geekbench_opencl
140,838
N/A
geekbench_vulkan
121,306
N/A

Analysis: NVIDIA L4 vs NVIDIA N1 16SM

Head-to-Head Benchmarks

The recorded data shows a stark contrast between these two NVIDIA parts. The L4 has measured benchmark scores in two tests, while the N1 16SM has no recorded benchmark entries in the database. The L4 achieves an average benchmark score of 131,072, placing it in the 95th percentile of all GPUs tracked. The N1 16SM, by contrast, sits in the 50th percentile with an average benchmark score of 0, meaning no test data has been captured for it.

The L4 posts a Geekbench OpenCL score of 140,838 and a Geekbench Vulkan score of 121,306. These figures place it narrowly behind several established workstation and server cards. The closest rival is the NVIDIA GeForce RTX 3090 Ti with an average score of 131,938, which edges out the L4 by 0.7 percent. The NVIDIA RTX 4000 Ada Generation and NVIDIA A10M both average 135,218 and 135,230 respectively, each sitting 3.1 percent ahead of the L4. The AMD Radeon PRO W6800 leads this group with a 135,396 average, a 3.2 percent advantage over the L4. These margins are small, indicating that the L4 clusters tightly with these comparable accelerators despite its lower thermal envelope.

The N1 16SM has no head-to-head benchmark data available. The head-to-head table is empty, and the wins counter shows zero victories for either side. The absence of Geekbench entries for the N1 16SM means the database cannot directly compare its compute performance against the L4. The L4's recorded scores stand alone as the only quantitative performance evidence in this matchup. Its Vulkan result trails its OpenCL result by roughly 14 percent, a pattern consistent with drivers optimized for compute workloads rather than graphics rendering. The L4's nearest rivals all cluster within a 3.2 percent band, which indicates the card's performance is well positioned among its peers in the database.

Architecture Differences

The two GPUs diverge sharply in their underlying designs. The L4 uses the AD104 chip built on the Ada Lovelace architecture, manufactured on a 5 nm process by TSMC. The N1 16SM uses the GB20B chip with the Blackwell 2.0 architecture, also on a 5 nm TSMC process. The L4 belongs to the Server Ada (Lxx) generation, while the N1 16SM belongs to the Blackwell IGP (N1x) generation. The L4 relies on 35,800 million transistors spread across a 294 mm² die, giving it a transistor density of 121.8 million per square millimeter. The N1 16SM's transistor count is listed as unknown, but its die measures 382 mm², which is roughly 30 percent larger than the L4's die.

The compute resources differ substantially. The L4 carries 7,424 shading units, 240 texture mapping units, and 80 raster output units. The N1 16SM has 2,048 shading units, 128 TMUs, and 24 ROPs. That means the L4 has 3.6 times more shading units and 3.3 times more ROPs. Ray tracing hardware follows the same pattern: the L4 has 60 RT cores versus 16 on the N1 16SM. Tensor core counts also favor the L4, with 240 versus 64. The L4's FP32 throughput is 30.29 TFLOPS, and its FP16 throughput matches at 30.29 TFLOPS with a 1:1 ratio. The N1 16SM delivers 9.609 TFLOPS for both FP32 and FP16, also at 1:1. The L4 therefore offers roughly 3.2 times the raw floating-point performance.

Clock behavior differs as well. The L4 has a base clock of 795 MHz and a boost clock of 2040 MHz. The N1 16SM runs a base of 741 MHz and boosts to 2346 MHz. The N1 16SM's higher boost clock does not compensate for its fewer cores. The L4 achieves a pixel rate of 163.2 GPixel/s and a texture rate of 489.6 GTexel/s, whereas the N1 16SM reaches 56.30 GPixel/s and 300.3 GTexel/s. The L4 leads in both fill-rate metrics by wide margins.

Memory configurations also separate the two. The L4 uses 24 GB of GDDR6 on a 192-bit bus, producing 300.1 GB/s of bandwidth. The N1 16SM uses 128 GB of LPDDR5X on a 256-bit bus, delivering 273.2 GB/s. The N1 16SM has more capacity but slightly lower bandwidth. The L4's memory runs at 1563 MHz with 12.5 Gbps effective speed, while the N1 16SM's memory runs at 1067 MHz with 8.5 Gbps effective. The L4 consumes 72 W according to its TDP rating, and it needs no external power connectors, relying on the PCIe slot. The N1 16SM's TDP is unknown, but it also uses no power connectors and is an integrated graphics processor (IGP) with a single HDMI output. The L4 has no display outputs at all.

The bus interfaces differ: the L4 uses PCIe 4.0 x16, while the N1 16SM uses PCIe 5.0 x16. The L4 measures 169 mm in length and 56 mm in height, fitting a single-slot form factor. The N1 16SM has no listed dimensions. API support diverges completely. The L4 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The N1 16SM lists N/A for DirectX, OpenGL, and Vulkan, which indicates it is not intended for conventional graphics API workloads. The L4's release date is recorded as March 20, 2023, while the N1 16SM's release date is May 31, 2026. The L4 has a predecessor in Server Ampere and a successor in Server Hopper. The N1 16SM lists neither a predecessor nor a successor.

The Verdict

The database presents a clear choice based on measured evidence. The L4 is the only one of the two with recorded benchmark scores, and those scores place it in the 95th percentile of all GPUs tracked. The N1 16SM has no benchmarks, a 50th percentile ranking, and an average score of zero. Anyone seeking a discrete accelerator with proven compute performance should select the L4. Its 30.29 TFLOPS of FP32 throughput, 60 RT cores, and 240 tensor cores make it a capable server-side compute card, and its 72 W TDP with no external power connectors allows deployment in power-constrained environments.

The N1 16SM, despite its larger die and newer Blackwell 2.0 architecture, lacks any performance validation in the database. Its 128 GB of LPDDR5X memory is its most distinctive asset, along with its PCIe 5.0 x16 interface and integrated form factor. The data indicates it is positioned as an IGP with a single HDMI output, which suggests a different use case entirely: embedded or integrated systems where graphics output matters. Its 9.609 TFLOPS of FP32 is roughly one-third of the L4's throughput. The N1 16SM's 50th percentile ranking places it at the median of all GPUs, while the L4 sits near the top.

The L4's release date of 2023 predates the N1 16SM's 2026 release, so the newer part has had less time to accumulate benchmark entries. However, the absence of any entries means no evidence exists to overturn the L4's advantage. The L4's nearest rivals all sit within a 3.2 percent band, which shows it competes directly with high-end workstation cards like the RTX 4000 Ada Generation and the Radeon PRO W6800. The N1 16SM has no rivals listed at all.

FAQ

Q: Which GPU has better raw compute throughput?

A: The L4 delivers 30.29 TFLOPS of FP32 and FP16, while the N1 16SM delivers 9.609 TFLOPS of both. The L4 offers roughly 3.2 times the floating-point performance.

Q: How much memory does each GPU have?

A: The L4 has 24 GB of GDDR6 on a 192-bit bus with 300.1 GB/s of bandwidth. The N1 16SM has 128 GB of LPDDR5X on a 256-bit bus with 273.2 GB/s of bandwidth.

Q: Are there any benchmark scores for the N1 16SM?

A: No. The N1 16SM has an empty benchmark list, an average benchmark score of 0, and no nearest rivals recorded in the database.

Q: What is the performance percentile of each GPU?

A: The L4 sits in the 95th percentile of all GPUs, while the N1 16SM sits in the 50th percentile.

Q: How do the two GPUs compare on ray tracing hardware?

A: The L4 has 60 RT cores, while the N1 16SM has 16 RT cores. The L4 also has 240 tensor cores versus 64 on the N1 16SM.

Q: What are the power requirements for each?

A: The L4 has a TDP of 72 W and uses no external power connectors. The N1 16SM's TDP is unknown, but it also uses no external power connectors.

Q: Do both GPUs support standard graphics APIs?

A: The L4 supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The N1 16SM lists N/A for all three APIs.

Where Each One Wins

The L4 wins on every measured performance metric in the database. Its FP32 throughput of 30.29 TFLOPS dwarfs the N1 16SM's 9.609 TFLOPS. Its pixel rate of 163.2 GPixel/s versus 56.30 GPixel/s gives it a 2.9 times advantage in rasterization throughput. Its texture rate of 489.6 GTexel/s versus 300.3 GTexel/s is a 1.6 times lead. The L4 also wins on memory bandwidth at 300.1 GB/s versus 273.2 GB/s, and it has significantly more shading units, TMUs, ROPs, RT cores, and tensor cores. The L4 is the clear choice for compute-heavy server workloads, as its benchmark scores confirm.

The N1 16SM wins on memory capacity with 128 GB versus 24 GB, a 5.3 times advantage. Its 256-bit memory bus is wider than the L4's 192-bit bus, though its lower effective memory speed reduces the bandwidth advantage to a deficit. The N1 16SM also wins on bus interface with PCIe 5.0 x16 versus PCIe 4.0 x16, and it offers a display output via HDMI, which the L4 lacks entirely. The N1 16SM's integrated form factor and larger die size of 382 mm² suggest it targets a different market segment, likely embedded or integrated systems where the 128 GB unified memory pool and display output are priorities. The L4's single-slot 169 mm length and no-output design indicate it belongs in server racks where compute density matters more than connectivity.

The L4 also wins on release timing in the database: its March 2023 release date has allowed benchmark data to accumulate, whereas the N1 16SM's May 2026 release date has not. The L4's position among its nearest rivals, all within 3.2 percent of its average score, confirms it competes at a high level. The N1 16SM, with no rivals and no scores, cannot claim any measured performance advantage. For users who need proven compute capability, the L4 is the only option with evidence. For users who need massive memory capacity and an integrated display output, the N1 16SM offers those features, but without any performance validation in the database.

DETAILED SPECIFICATIONS

SPECIFICATION
L4
N1 16SM
Core Specs
Shading Units
7,424
2,048 -72.4%
Shaders
7,424
2,048 -72.4%
TMUs
240
128 -46.7%
ROPs
80
24 -70.0%
SM Count
60
16 -73.3%
Clocks
Base Clock
795 MHz
741 MHz
Boost Clock
2040 MHz
2346 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
1067 MHz 8.5 Gbps effective
Memory
Memory Size
24 GB
128 GB
VRAM (MB)
24,576
131,072 +433.3%
Memory Type
GDDR6
LPDDR5X
Memory Bus
192 bit
256 bit
Bandwidth
300.1 GB/s
273.2 GB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
48 MB
50 MB
Performance
Pixel Rate
163.2 GPixel/s
56.30 GPixel/s
Texture Rate
489.6 GTexel/s
300.3 GTexel/s
FP32 (TFLOPS)
30.29 TFLOPS
9.609 TFLOPS
FP64 (TFLOPS)
473.3 GFLOPS (1:64)
150.1 GFLOPS (1:64)
FP16 (TFLOPS)
30.29 TFLOPS (1:1)
9.609 TFLOPS (1:1)
AI/RT
RT Cores
60
16 -73.3%
Tensor Cores
240
64 -73.3%
Power
TDP
72 W
unknown
TDP (W)
72
Suggested PSU
250 W
Power Connectors
None
None
Architecture
Architecture
Ada Lovelace
Blackwell 2.0
GPU Name
AD104
GB20B
Generation
Server Ada (Lxx)
Blackwell IGP (N1x)
Process Size
5 nm
5 nm
Transistors
35,800 million
unknown
Die Size
294 mm²
382 mm²
Foundry
TSMC
TSMC
Density
121.8M / mm²
API Support
DirectX
12 Ultimate (12_2)
OpenGL
4.6
Vulkan
1.4
OpenCL
3.0
3.0
CUDA
8.9
12.1
Shader Model
6.8
Physical
Slot Width
Single-slot
IGP
Length
169 mm 6.7 inches
Height
56 mm 2.2 inches
Outputs
No outputs
1x HDMI
Bus Interface
PCIe 4.0 x16
PCIe 5.0 x16
Other
Production
Active
Active
Predecessor
Server Ampere
Successor
Server Hopper
View L4 Details View N1 16SM Details