NVIDIA Jetson Orin Nano Super vs NVIDIA L4 Comparison
NVIDIA Jetson Orin Nano Super
L4
PERFORMANCE BENCHMARKS
Analysis: NVIDIA Jetson Orin Nano Super vs NVIDIA L4
FAQ
Q: What is the key architectural difference between the NVIDIA Jetson Orin Nano Super and the NVIDIA L4?
A: The Jetson Orin Nano Super uses the GA10B chip on the Ampere architecture, built on an 8 nm process by Samsung. The L4 uses the AD104 chip on the Ada Lovelace architecture, built on a 5 nm process by TSMC.
Q: How do their memory subsystems compare?
A: The Jetson Orin Nano Super has 8 GB of LPDDR5 on a 128-bit bus with 102.4 GB/s bandwidth. The L4 has 24 GB of GDDR6 on a 192-bit bus with 300.1 GB/s bandwidth.
Q: Which GPU has higher raw compute throughput?
A: The L4 delivers 30.29 TFLOPS FP32 and 30.29 TFLOPS FP16 (1:1 ratio). The Jetson Orin Nano Super delivers 2.089 TFLOPS FP32 and 4.178 TFLOPS FP16 (2:1 ratio).
Q: What are the power requirements?
A: The Jetson Orin Nano Super has a TDP of 25 W. The L4 has a TDP of 72 W and a suggested PSU of 250 W.
Q: What is the physical form factor difference?
A: The Jetson Orin Nano Super is an IGP (integrated graphics processor) sized at 70 mm by 45 mm. The L4 is a single-slot card measuring 169 mm by 56 mm.
Q: What is the production status and release timing?
A: Both are Active. The Jetson Orin Nano Super was released on 2024-12-16, while the L4 was released on 2023-03-20.
Architecture Differences
The Jetson Orin Nano Super and the L4 represent two distinct NVIDIA architectures and design philosophies. The Jetson Orin Nano Super uses the GA10B chip, which belongs to the Ampere architecture, manufactured on Samsung's 8 nm process. Its die size is 200 mm². This is a Tegra-generation part, indicating an embedded or edge-oriented design. The L4 uses the AD104 chip, part of the Ada Lovelace architecture, built on TSMC's 5 nm process. Its die size is 294 mm² with 35,800 million transistors and a transistor density of 121.8M per mm². The process node difference is significant: 8 nm versus 5 nm, which directly impacts power efficiency and density.
The compute resources differ sharply. The Jetson Orin Nano Super has 1024 shading units, 32 TMUs, 16 ROPs, and 32 tensor cores. It has no dedicated RT cores. The L4 has 7424 shading units, 240 TMUs, 80 ROPs, 60 RT cores, and 240 tensor cores. That is a 7.25x difference in shading units and a 7.5x difference in tensor cores. The L4 also supports ray tracing hardware, which the Jetson Orin Nano Super lacks entirely.
Memory architecture also differs fundamentally. The Jetson Orin Nano Super uses 8 GB of LPDDR5 with a 128-bit bus, running at 800 MHz with 6.4 Gbps effective speed, yielding 102.4 GB/s bandwidth. The L4 uses 24 GB of GDDR6 with a 192-bit bus, running at 1563 MHz with 12.5 Gbps effective speed, yielding 300.1 GB/s bandwidth. The L4 has three times the capacity and nearly three times the bandwidth. FP16 processing differs in ratio: the Jetson Orin Nano Super runs FP16 at 4.178 TFLOPS with a 2:1 ratio relative to FP32, while the L4 runs FP16 at 30.29 TFLOPS with a 1:1 ratio. This means the L4 does not sacrifice FP16 throughput relative to FP32, while the Jetson Orin Nano Super effectively halves its FP32 rate when using FP16.
The L4 is a server-class Ada part with a PCIe 4.0 x16 interface, while the Jetson Orin Nano Super uses PCIe 4.0 x4. The L4 has no display outputs, while the Jetson Orin Nano Super's outputs are portable device dependent. The L4 is a single-slot card with no power connectors (relying on the slot), and the Jetson Orin Nano Super is an IGP.
Where Each One Wins
The Jetson Orin Nano Super wins in situations requiring minimal power draw and compact physical integration. Its 25 W TDP makes it suitable for embedded deployments, portable devices, or powered-by-battery systems where the L4's 72 W TDP and 250 W suggested PSU would be prohibitive. The 70 mm by 45 mm dimensions allow placement in small chassis or custom carrier boards. Its PCIe 4.0 x4 interface is sufficient for edge inference workloads that do not demand massive data movement. The FP16 2:1 ratio indicates a design optimized for mixed-precision workloads where FP16 acceleration matters more than raw FP32.
The L4 wins on almost every raw performance metric. Its 30.29 TFLOPS FP32 is roughly 14.5x higher than the Jetson Orin Nano Super's 2.089 TFLOPS. It has 24 GB of GDDR6 versus 8 GB of LPDDR5, which matters for large model footprints in AI inference or rendering tasks. The 300.1 GB/s bandwidth enables faster data feeding for compute-heavy operations. Its 60 RT cores provide hardware ray tracing, which the Jetson Orin Nano Super cannot do. The 240 tensor cores deliver 7.5x the tensor throughput of the 32 tensor cores on the Jetson Orin Nano Super. The L4 is also a single-slot card that fits standard server racks, with a PCIe 4.0 x16 interface for maximum host bandwidth.
In benchmark percentile terms, the L4 sits at the 95th percentile of all GPUs, while the Jetson Orin Nano Super sits at the 50th percentile. The L4 has recorded benchmark scores (Geekbench OpenCL 140838, Geekbench Vulkan 121306, average 131072), while the Jetson Orin Nano Super has no recorded benchmark scores in the database. The L4's nearest rivals include the NVIDIA GeForce RTX 3090 Ti (avg score 131938, delta -0.7%), the NVIDIA RTX 4000 Ada Generation (avg score 135218, delta -3.1%), the NVIDIA A10M (avg score 135230, delta -3.1%), and the AMD Radeon PRO W6800 (avg score 135396, delta -3.2%). These deltas indicate the L4 is within 3.2% of these high-end cards, with the RTX 3090 Ti being slightly ahead by 0.7%.
Specification Differences
The two GPUs differ across nearly every specification field. The chip is GA10B on the Jetson Orin Nano Super versus AD104 on the L4. Architecture is Ampere versus Ada Lovelace. Generation is Tegra (Ampere) versus Server Ada (Lxx). Process node is 8 nm (Samsung) versus 5 nm (TSMC). The L4 has known transistor count (35,800 million) and density (121.8M / mm²), while the Jetson Orin Nano Super's transistor count is unknown. Die size is 200 mm² versus 294 mm².
Clocks differ: the L4 has a base clock of 795 MHz and a boost clock of 2040 MHz, while the Jetson Orin Nano Super has no base or boost clock listed. Memory clock is 800 MHz with 6.4 Gbps effective on the Jetson Orin Nano Super versus 1563 MHz with 12.5 Gbps effective on the L4. Memory size is 8 GB versus 24 GB. Memory type is LPDDR5 versus GDDR6. Bus width is 128 bit versus 192 bit. Bandwidth is 102.4 GB/s versus 300.1 GB/s.
Compute resources differ: shading units 1024 versus 7424, TMUs 32 versus 240, ROPs 16 versus 80. The L4 has 60 RT cores while the Jetson Orin Nano Super has none. Tensor cores are 32 versus 240. Pixel rate is 16.32 GPixel/s versus 163.2 GPixel/s. Texture rate is 32.64 GTexel/s versus 489.6 GTexel/s. FP32 is 2.089 TFLOPS versus 30.29 TFLOPS. FP16 is 4.178 TFLOPS (2:1) versus 30.29 TFLOPS (1:1).
Power and physical specs: TDP is 25 W versus 72 W. Slot width is IGP versus single-slot. Power connectors are none on both, but the L4 has a suggested PSU of 250 W while the Jetson Orin Nano Super has none listed. Bus interface is PCIe 4.0 x4 versus PCIe 4.0 x16. Display outputs are portable device dependent versus no outputs. Dimensions are 70 mm by 45 mm versus 169 mm by 56 mm. Release dates are 2024-12-16 versus 2023-03-20. The L4 has a predecessor (Server Ampere) and successor (Server Hopper), while the Jetson Orin Nano Super has neither listed. The Jetson Orin Nano Super has a launch MSRP of 249 USD, while the L4 has no launch MSRP listed.
Head-to-Head Benchmarks
The database contains no direct head-to-head benchmark results between these two GPUs, and the Jetson Orin Nano Super has no individual benchmark scores recorded. The L4, however, has two recorded benchmark results. In Geekbench OpenCL, the L4 scores 140838. In Geekbench Vulkan, it scores 121306. Its average benchmark score is 131072.
The L4's percentile ranking of 95 places it firmly in the upper tier of all GPUs. Its nearest rival, the NVIDIA GeForce RTX 3090 Ti, has an average score of 131938, which is 0.7% higher than the L4's average. The NVIDIA RTX 4000 Ada Generation scores 135218, which is 3.1% higher. The NVIDIA A10M scores 135230, also 3.1% higher. The AMD Radeon PRO W6800 scores 135396, which is 3.2% higher. These numbers show the L4 is competitive with, but slightly behind, these established high-end workstation and server cards.
Given the Jetson Orin Nano Super has no benchmark scores, the comparison must rely on raw specification differences. The L4's FP32 throughput of 30.29 TFLOPS is approximately 14.5 times the Jetson Orin Nano Super's 2.089 TFLOPS. The L4's FP16 throughput of 30.29 TFLOPS is approximately 7.25 times the Jetson Orin Nano Super's 4.178 TFLOPS. The L4's memory bandwidth of 300.1 GB/s is approximately 2.93 times the Jetson Orin Nano Super's 102.4 GB/s. The L4's texture rate of 489.6 GTexel/s is 15 times the Jetson Orin Nano Super's 32.64 GTexel/s. The L4's pixel rate of 163.2 GPixel/s is 10 times the Jetson Orin Nano Super's 16.32 GPixel/s.
These ratios indicate that in compute-bound, memory-bound, or texture-bound workloads, the L4 holds a decisive advantage. The Jetson Orin Nano Super's only clear wins are in power consumption (25 W versus 72 W) and physical footprint (70 mm by 45 mm versus 169 mm by 56 mm). The L4 also has the advantage of RT cores, which the Jetson Orin Nano Super does not have at all.
The Verdict
The data supports a clear separation of use cases. The NVIDIA Jetson Orin Nano Super is an embedded or edge device. Its 25 W TDP, IGP form factor, and 70 mm by 45 mm dimensions make it suitable for compact, power-constrained deployments. Its 1024 shading units, 32 tensor cores, and 8 GB LPDDR5 are adequate for lightweight inference or low-resolution processing. The 2:1 FP16 ratio suggests it is intended for mixed-precision edge AI workloads, where the 4.178 TFLOPS FP16 throughput can be leveraged without the power draw of a full-size GPU. The 249 USD launch MSRP positions it as an accessible entry point for developers building small-scale systems.
The NVIDIA L4 is a server-class accelerator. Its 7424 shading units, 240 tensor cores, 60 RT cores, and 24 GB GDDR6 place it in a different performance tier. The 30.29 TFLOPS FP32 and FP16 (1:1) throughput, combined with 300.1 GB/s bandwidth, make it suitable for larger models, higher-resolution workloads, and ray-traced rendering. The 95th percentile ranking and average benchmark score of 131072 confirm its competitive position against the RTX 3090 Ti, RTX 4000 Ada Generation, A10M, and Radeon PRO W6800. The 72 W TDP and single-slot form factor with PCIe 4.0 x16 allow deployment in standard server environments.
The choice between them depends on the deployment context. For a portable, low-power device that must operate on battery or within a small enclosure, the Jetson Orin Nano Super is the only viable option given its power and size constraints. For a rack-mounted server handling compute-heavy inference, rendering, or ray tracing workloads, the L4 delivers roughly 14.5 times the FP32 throughput, 7.25 times the FP16 throughput, and nearly three times the memory bandwidth.
The L4's lack of display outputs means it is not intended for direct video output, while the Jetson Orin Nano Super's display outputs are portable device dependent, indicating flexibility for embedded displays. The L4's predecessor and successor designations (Server Ampere and Server Hopper) show a clear lineage in NVIDIA's server GPU roadmap, while the Jetson Orin Nano Super's Tegra generation indicates a different product family aimed at system-on-chip applications. The recorded benchmark data for the L4, including the 140838 OpenCL and 121306 Vulkan scores, provides measurable evidence of its capability. The Jetson Orin Nano Super has no such data, so its performance must be inferred from its specification sheet alone.