NVIDIA Jetson Orin Nano Super vs NVIDIA Rubin GPU Comparison

NVIDIA
GEFORCE

NVIDIA Jetson Orin Nano Super

CORE STATE GA10B
VRAM 8 GB
CLOCK SPEED —
TDP 25 W
BUS WIDTH 128 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2024
VS
NVIDIA
GEFORCE

Rubin GPU

CORE STATE GR100
VRAM 288 GB
CLOCK SPEED 2267 MHz
TDP 2300 W
BUS WIDTH 16384 bit
ARCHITECTURE Rubin
nm
PROCESS 3 nm
LAUNCH DATE 2026

Analysis: NVIDIA Jetson Orin Nano Super vs NVIDIA Rubin GPU

Head-to-Head Benchmarks

The recorded database contains no direct head-to-head benchmark results for these two parts. Neither the NVIDIA Jetson Orin Nano Super nor the NVIDIA Rubin GPU has any submitted benchmark scores, and the wins counter shows zero for both sides. This absence of measured data is itself informative: the two products occupy entirely different operational domains, and no common workload has been recorded against both.

The Jetson Orin Nano Super carries a percentile ranking of 50 among all GPUs in the database, which places it at the median of the recorded distribution. The Rubin GPU also shows a percentile ranking of 50. Without actual benchmark submissions, these percentile values reflect only the absence of data rather than comparative performance. No average benchmark score exists for either unit; both show a value of 0 in the database.

What can be stated from the recorded figures is the theoretical peak throughput for each design. The Jetson Orin Nano Super delivers 2.089 TFLOPS of FP32 compute and 4.178 TFLOPS of FP16 compute at a 2:1 ratio. The Rubin GPU delivers 130.0 TFLOPS of FP32 compute and 260.0 TFLOPS of FP16 compute at the same 2:1 ratio. The Rubin GPU therefore shows roughly 62 times the FP32 throughput of the Jetson part based on the recorded specification values, and roughly 62 times the FP16 throughput as well. These are architectural ceilings, not measured workload results, but they establish the scale of the performance gap.

The pixel throughput figures follow the same pattern. The Jetson Orin Nano Super achieves a pixel rate of 16.32 GPixel/s and a texture rate of 32.64 GTexel/s. The Rubin GPU achieves 54.41 GPixel/s and 2,031.2 GTexel/s. The texture rate gap is particularly large: the Rubin part processes textures at over 62 times the rate of the Jetson unit. The pixel rate gap is more modest at roughly 3.3 times, reflecting the Rubin's relatively small count of 24 ROPs compared to its massive shader and texture arrays.

Memory bandwidth shows the most extreme divergence. The Jetson Orin Nano Super has 102.4 GB/s of bandwidth from 8 GB of LPDDR5 on a 128-bit bus. The Rubin GPU has 22.1 TB/s of bandwidth from 288 GB of HBM4 on a 16,384-bit bus. The Rubin part offers roughly 215 times the memory bandwidth of the Jetson unit. This difference is central to understanding where each part can be deployed.

Architecture Differences

The two chips come from different architectural generations and different foundries. The Jetson Orin Nano Super uses the GA10B chip built on the Ampere architecture, fabricated on an 8 nm process at Samsung. The Rubin GPU uses the GR100 chip built on the Rubin architecture, fabricated on a 3 nm process at TSMC. The process node difference of 8 nm versus 3 nm is substantial, and the foundry choice differs as well.

The die sizes reflect the gulf in design intent. The Jetson's GA10B measures 200 mm². The Rubin's GR100 measures 1,456 mm², which is more than seven times the area. Transistor counts tell a similar story: the Jetson part has no recorded transistor count, while the Rubin part lists 336,000 million transistors, which is 336 billion. The transistor density for the Rubin part is recorded at 230.8 million transistors per square millimeter. No density figure exists for the Jetson chip in the database.

Shading unit counts differ by a factor of 28. The Jetson Orin Nano Super has 1,024 shading units, 32 TMUs, and 16 ROPs. The Rubin GPU has 28,672 shading units, 896 TMUs, and 24 ROPs. The TMU count scales with the shader count almost exactly at the same 28-to-1 ratio, but the ROP count scales far less, only 1.5 times. This imbalance suggests the Rubin design prioritizes compute and texture throughput over final pixel output, which is typical for a server-oriented accelerator.

Tensor core counts also differ by a factor of 28: the Jetson has 32 tensor cores, while the Rubin has 896. Neither part lists dedicated ray tracing cores in the database, so no comparison can be made on that front. Both parts support the same FP16 compute ratio of 2:1 relative to FP32, indicating that both use a similar paired-issue scheme for reduced precision.

The memory architectures are completely different. The Jetson uses LPDDR5 in an 8 GB configuration with a 128-bit bus. The Rubin uses HBM4 in a 288 GB configuration with a 16,384-bit bus. The bus width difference is a factor of 128, which explains the enormous bandwidth gap. The Jetson memory clock is recorded as 800 MHz with 6.4 Gbps effective. The Rubin memory clock is 2,695 MHz with 10.8 Gbps effective. The higher clock and vastly wider bus combine to produce the 22.1 TB/s figure.

Power envelopes separate the two designs entirely. The Jetson Orin Nano Super has a TDP of 25 W and is classified as an IGP with a mobile form factor. The Rubin GPU has a TDP of 2,300 W, requires a suggested PSU of 2,700 W, and is packaged as an SXM Module. The Rubin's power draw is 92 times the Jetson's total board power. This single figure dictates deployment scenarios more than any other specification.

Where Each One Wins

The Jetson Orin Nano Super wins in scenarios where power is constrained and physical size matters. Its 25 W TDP allows operation in portable or embedded contexts. Its dimensions of 70 mm in length and 45 mm in height fit compact carrier boards. The display outputs are listed as "Portable Device Dependent," which means the Jetson can drive a screen in a handheld or embedded product. The Rubin GPU lists "No outputs," so it cannot drive displays at all. For edge devices, robotics, or any system requiring on-device graphics output, the Jetson is the only viable choice between the two.

The Jetson also wins on the software interface front. It supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The Rubin GPU lists N/A for all three APIs. This means the Jetson can run conventional graphics workloads and interactive applications, while the Rubin part is a pure compute accelerator with no graphics API support. The bus interface also favors the Jetson in some respects: it uses PCIe 4.0 x4, which is broadly compatible with existing platforms, while the Rubin uses PCIe 6.0 x16, which requires newer server infrastructure.

The Rubin GPU wins in every raw compute metric. Its FP32 throughput of 130.0 TFLOPS exceeds the Jetson's 2.089 TFLOPS by a factor of roughly 62. Its FP16 throughput of 260.0 TFLOPS exceeds the Jetson's 4.178 TFLOPS by the same factor. The Rubin's texture rate of 2,031.2 GTexel/s dwarfs the Jetson's 32.64 GTexel/s. The Rubin's memory bandwidth of 22.1 TB/s is in a different class entirely from the Jetson's 102.4 GB/s. For any workload that scales with compute throughput or memory bandwidth, the Rubin part is categorically superior.

The Rubin also wins on capacity. Its 288 GB of HBM4 memory allows large models and datasets to reside on-chip, whereas the Jetson's 8 GB of LPDDR5 limits working set size. The Rubin's 896 tensor cores provide massive matrix multiplication throughput for neural network training and inference at scale. The Jetson's 32 tensor cores are sufficient for lightweight inference at the edge but cannot approach the Rubin's throughput.

The release dates reflect different product cycles. The Jetson Orin Nano Super was released on 2024-12-16. The Rubin GPU is scheduled for release on 2025-12-31. The Rubin's predecessor is recorded as "Server Blackwell," while the Jetson has no predecessor listed. Both parts are listed as Active in production status.

Specification Differences

The recorded specifications differ across nearly every field. The process node differs: 8 nm for the Jetson versus 3 nm for the Rubin. The foundry differs: Samsung versus TSMC. The die size differs: 200 mm² versus 1,456 mm². The transistor count is unknown for the Jetson and 336,000 million for the Rubin. The transistor density is unlisted for the Jetson and 230.8 million per mm² for the Rubin.

Base clock is unlisted for the Jetson, while the Rubin has a base clock of 700 MHz. Boost clock is unlisted for the Jetson, while the Rubin has a boost clock of 2,267 MHz. Memory clock differs: 800 MHz with 6.4 Gbps effective for the Jetson versus 2,695 MHz with 10.8 Gbps effective for the Rubin. Memory size differs: 8 GB versus 288 GB. Memory type differs: LPDDR5 versus HBM4. Bus width differs: 128-bit versus 16,384-bit. Bandwidth differs: 102.4 GB/s versus 22.1 TB/s.

Shading units differ: 1,024 versus 28,672. TMUs differ: 32 versus 896. ROPs differ: 16 versus 24. Tensor cores differ: 32 versus 896. Pixel rate differs: 16.32 GPixel/s versus 54.41 GPixel/s. Texture rate differs: 32.64 GTexel/s versus 2,031.2 GTexel/s. FP32 throughput differs: 2.089 TFLOPS versus 130.0 TFLOPS. FP16 throughput differs: 4.178 TFLOPS versus 260.0 TFLOPS.

TDP differs: 25 W versus 2,300 W. Slot width differs: IGP versus SXM Module. Suggested PSU is unlisted for the Jetson and 2,700 W for the Rubin. Bus interface differs: PCIe 4.0 x4 versus PCIe 6.0 x16. Display outputs differ: "Portable Device Dependent" versus "No outputs." API support differs: the Jetson supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4; the Rubin lists N/A for all three.

Dimensions differ: the Jetson measures 70 mm in length and 45 mm in height, while the Rubin has no recorded dimensions. Release dates differ: 2024-12-16 versus 2025-12-31. The Rubin has a predecessor ("Server Blackwell") and a launch MSRP of 249 USD, while the Rubin has no launch MSRP. The Jetson's launch MSRP is 249 USD, stated here as recorded in the database.

FAQ

Q: Which part has higher FP32 compute throughput?

A: The NVIDIA Rubin GPU delivers 130.0 TFLOPS of FP32 compute, compared to 2.089 TFLOPS for the NVIDIA Jetson Orin Nano Super. This is roughly 62 times higher on the Rubin side.

Q: What is the memory bandwidth difference?

A: The Jetson Orin Nano Super has 102.4 GB/s of bandwidth from 8 GB of LPDDR5 on a 128-bit bus. The Rubin GPU has 22.1 TB/s of bandwidth from 288 GB of HBM4 on a 16,384-bit bus, which is approximately 215 times higher.

Q: Which part supports graphics APIs?

A: The Jetson Orin Nano Super supports DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4. The Rubin GPU lists N/A for DirectX, OpenGL, and Vulkan, indicating no graphics API support.

Q: What are the power requirements for each?

A: The Jetson Orin Nano Super has a TDP of 25 W and is an IGP. The Rubin GPU has a TDP of 2,300 W, requires a suggested PSU of 2,700 W, and is packaged as an SXM Module.

Q: Which part has tensor cores, and how many?

A: Both parts have tensor cores. The Jetson Orin Nano Super has 32 tensor cores, while the Rubin GPU has 896 tensor cores.

Q: What is the process node and die size for each?

A: The Jetson Orin Nano Super is built on an 8 nm process at Samsung with a die size of 200 mm². The Rubin GPU is built on a 3 nm process at TSMC with a die size of 1,456 mm².

DETAILED SPECIFICATIONS

SPECIFICATION
Jetson Orin Nano Super
Rubin GPU
Core Specs
Shading Units
1,024
28,672 +2700.0%
Shaders
1,024
28,672 +2700.0%
TMUs
32
896 +2700.0%
ROPs
16
24 +50.0%
SM Count
8
224 +2700.0%
Clocks
Base Clock
—
700 MHz
Boost Clock
—
2267 MHz
GPU Clock
1020 MHz
—
Memory Clock
800 MHz 6.4 Gbps effective
2695 MHz 10.8 Gbps effective
Memory
Memory Size
8 GB
288 GB
VRAM (MB)
8,192
294,912 +3500.0%
Memory Type
LPDDR5
HBM4
Memory Bus
128 bit
16384 bit
Bandwidth
102.4 GB/s
22.1 TB/s
Cache
L1 Cache
128 KB (per SM)
256 KB (per SM)
L2 Cache
2 MB
128 MB
Performance
Pixel Rate
16.32 GPixel/s
54.41 GPixel/s
Texture Rate
32.64 GTexel/s
2,031.2 GTexel/s
FP32 (TFLOPS)
2.089 TFLOPS
130.0 TFLOPS
FP64 (TFLOPS)
—
32.50 TFLOPS (1:4)
FP16 (TFLOPS)
4.178 TFLOPS (2:1)
260.0 TFLOPS (2:1)
AI/RT
Tensor Cores
32
896 +2700.0%
Power
TDP
25 W
2300 W
TDP (W)
25
2,300 +9100.0%
Suggested PSU
—
2700 W
Architecture
Architecture
Ampere
Rubin
GPU Name
GA10B
GR100
Generation
Tegra (Ampere)
Server Rubin (Rxx)
Process Size
8 nm
3 nm
Transistors
unknown
336,000 million
Die Size
200 mm²
1456 mm²
Foundry
Samsung
TSMC
Density
—
230.8M / mm²
API Support
DirectX
12 Ultimate (12_2)
—
OpenGL
4.6
—
Vulkan
1.4
—
OpenCL
3.0
3.0
CUDA
8.7
10.7
Shader Model
6.8
—
Physical
Slot Width
IGP
SXM Module
Length
70 mm 2.8 inches
—
Height
45 mm 1.8 inches
—
Outputs
Portable Device Dependent
No outputs
Bus Interface
PCIe 4.0 x4
PCIe 6.0 x16
Other
Launch Price
249 USD
—
Production
Active
Active
Predecessor
—
Server Blackwell
View Jetson Orin Nano Super Details View Rubin GPU Details