GEFORCE

NVIDIA N1X 40SM

NVIDIA graphics card specifications and benchmark scores

128 GB
VRAM
2346
MHz Boost
TDP
256
Bus Width
Ray Tracing Tensor Cores

At a Glance

NVIDIA
VRAM 128 GB
Boost Clock 2,346 MHz
Shaders 5,120
Bus Width 256-bit
TDP unknown
Memory Type LPDDR5X
RT Cores 40
Architecture Blackwell 2.0
nm
Process 5 nm
Released Jun 2026

NVIDIA N1X 40SM Specifications

N1X 40SM GPU Core

Shader units and compute resources

The NVIDIA N1X 40SM GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
5,120
Shaders
5,120
TMUs
320
ROPs
40
SM Count
40

N1X 40SM Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the N1X 40SM's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The N1X 40SM by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
741 MHz
Base Clock
741 MHz
Boost Clock
2346 MHz
Boost Clock
2,346 MHz
Memory Clock
1067 MHz 8.5 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's N1X 40SM Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The N1X 40SM's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
128 GB
VRAM
131,072 MB
Memory Type
LPDDR5X
VRAM Type
LPDDR5X
Memory Bus
256 bit
Bus Width
256-bit
Bandwidth
273.2 GB/s

N1X 40SM by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the N1X 40SM, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
128 KB (per SM)
L2 Cache
50 MB

N1X 40SM Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA N1X 40SM against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
24.02 TFLOPS
FP64 (Double)
375.4 GFLOPS (1:64)
FP16 (Half)
24.02 TFLOPS (1:1)
Pixel Rate
93.84 GPixel/s
Texture Rate
750.7 GTexel/s

N1X 40SM Ray Tracing & AI

Hardware acceleration features

The NVIDIA N1X 40SM includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the N1X 40SM capable of delivering both stunning graphics and smooth frame rates in modern titles.

RT Cores
40
Tensor Cores
160

Blackwell 2.0 Architecture & Process

Manufacturing and design details

The NVIDIA N1X 40SM is built on NVIDIA's Blackwell 2.0 architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the N1X 40SM will perform in GPU benchmarks compared to previous generations.

Architecture
Blackwell 2.0
GPU Name
GB20B
Process Node
5 nm
Foundry
TSMC
Transistors
unknown
Die Size
382 mm²

NVIDIA's N1X 40SM Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA N1X 40SM determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the N1X 40SM to maintain boost clocks without throttling.

TDP
unknown
Power Connectors
None

N1X 40SM by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA N1X 40SM are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
IGP
Bus Interface
PCIe 5.0 x16
Display Outputs
1x HDMI
Display Outputs
1x HDMI

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA N1X 40SM. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
N/A
DirectX
N/A
OpenGL
N/A
OpenGL
N/A
Vulkan
N/A
Vulkan
N/A
OpenCL
3.0
CUDA
12.1
Shader Model
N/A

N1X 40SM Product Information

Release and pricing details

The NVIDIA N1X 40SM is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the N1X 40SM by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Jun 2026
Production
Active

N1X 40SM Benchmark Scores

No benchmark data available for this GPU.

About NVIDIA N1X 40SM

# NVIDIA N1X 40SM: A 128 GB IGP with Blackwell 2.0 Architecture

The NVIDIA N1X 40SM is an integrated graphics processor built on the Blackwell 2.0 architecture, fabricated on TSMC's 5 nm process with a 382 mm² die. It targets users who need an exceptionally large memory pool in a compact, integrated form factor, offering 128 GB of LPDDR5X VRAM alongside 5120 shading units and 40 ray tracing cores. The data shows a GPU that prioritizes memory capacity and modern feature support over raw compute throughput, positioning it as a specialized solution within the broader graphics landscape.

Who Should Consider It

The N1X 40SM is best suited for users whose primary workload involves massive datasets that must reside in GPU memory, rather than those seeking maximum frame rates in traditional gaming scenarios. With 128 GB of LPDDR5X memory on a 256-bit bus delivering 273.2 GB/s of bandwidth, this IGP can hold entire large-scale models, high-resolution texture packs, or complex scientific visualizations without spilling to system memory. The 50th percentile ranking against all GPUs indicates it sits squarely in the mid-pack for overall performance, meaning users should not expect flagship-level compute speeds, but the memory capacity is a differentiator that few discrete cards can match.

For resolution and settings recommendations, the pixel rate of 93.84 GPixel/s and texture rate of 750.7 GTexel/s suggest that at 1080p with moderate settings, the N1X 40SM can handle many titles comfortably, though the 40 ROPs may limit performance in fill-rate-bound scenarios at higher resolutions. At 1440p, users should plan for reduced texture quality or rely on the large VRAM to offset potential bottlenecks, as the 24.02 TFLOPS of FP32 compute is adequate for modern games but not exceptional. For 4K gaming, the data indicates this is not the primary use case; the memory bandwidth of 273.2 GB/s, while respectable, is lower than what high-end discrete solutions offer, and the 40 ROPs will struggle with heavy pixel work at that resolution.

The integrated nature of the GPU (slot width is IGP) means it is ideal for small-form-factor or workstation builds where a discrete card is impractical. Users running AI inference, large-scale data visualization, or virtualized GPU workloads will find the 128 GB capacity transformative, even if the compute throughput is mid-pack. Conversely, competitive gamers seeking high refresh rates at 1440p or above should look elsewhere, as the benchmark percentile data does not support such expectations. The 160 tensor cores provide acceleration for AI-assisted features, making this a compelling choice for creative professionals who need both memory and neural network support.

How It Compares

The N1X 40SM has no direct rivals listed in the benchmark database, which is unsurprising given its unique combination of integrated design and 128 GB memory capacity. The nearestRivals array is empty, indicating that the data currently contains no comparable GPUs for direct percentile-based comparison. This absence of rivals suggests that the N1X 40SM occupies a niche that other products do not fill, making positional analysis qualitative rather than quantitative. Benchmark results alone cannot place it against specific competitors, so users must extrapolate from its raw specifications.

Given the lack of rival data, the 50th percentile rank against all GPUs serves as the primary reference point. This median position implies that half of all GPUs in the database outperform it in aggregate benchmark scores, while half fall behind. The 24.02 TFLOPS FP32 figure places it in the range of mid-generation discrete graphics cards, but the 128 GB memory capacity is an outlier that skews its utility profile. Without rival names or deltaPct values, the analysis must rely on architectural characteristics: the GB20B chip with Blackwell 2.0 features suggests a modern feature set, while the 741 MHz base and 2346 MHz boost clocks indicate a power-efficient design that prioritizes sustained throughput over peak performance.

The absence of comparison data means potential buyers should weigh the N1X 40SM's memory capacity against the unknown performance of alternatives. The 50th percentile ranking provides a baseline expectation, but the specific workloads that benefit from 128 GB VRAM may see outsized value relative to the raw score. In essence, this GPU competes on capacity rather than speed, and the lack of rivals in the database reinforces that positioning.

Ray Tracing and Feature Set

The N1X 40SM includes 40 ray tracing cores, enabling hardware-accelerated ray tracing in supported applications. The Blackwell 2.0 architecture underpins this feature set, providing the foundation for both RT and tensor operations. The 160 tensor cores deliver AI acceleration for features like deep learning super sampling and other neural network-based enhancements, which can partially offset the mid-pack compute performance by improving effective frame rates in compatible titles.

However, the API support table shows DirectX, OpenGL, and Vulkan all listed as "N/A," which is a significant caveat for gaming applications. This suggests that the N1X 40SM may not expose standard graphics APIs to software, potentially limiting it to proprietary or compute-focused workloads rather than mainstream games. The single HDMI display output reinforces this interpretation, indicating a headless or single-display configuration typical of server or compute environments.

For ray tracing specifically, the 40 RT cores are present, but without DirectX or Vulkan API support, traditional game-based ray tracing may not be accessible. Instead, the RT cores likely serve compute workloads that utilize ray tracing algorithms, such as physical simulations or rendering pipelines that bypass standard graphics APIs. The 24.02 TFLOPS of FP16 (1:1 ratio) suggests that the GPU can handle mixed-precision workloads efficiently, which is valuable for AI training or inference tasks that leverage the tensor cores. The 5 nm process from TSMC and 382 mm² die size indicate a modern, density-optimized design that balances power efficiency with transistor budget.

Power and Cooling

The N1X 40SM has a TDP listed as unknown, but the integrated form factor (slot width IGP) and lack of power connectors (powerConnectors: "None") imply that it draws power from the host system rather than requiring external supply. The suggested PSU is null, meaning no specific power supply recommendation is provided in the data. The absence of power connectors is notable, as it suggests the GPU is designed to be powered entirely through the PCIe 5.0 x16 slot, which can deliver up to 75W under standard specifications.

For cooling, the IGP designation means it uses the system's cooling solution rather than a dedicated cooler. The 741 MHz base clock and 2346 MHz boost clock are relatively moderate, which likely keeps thermal output manageable within an integrated design. The 5 nm process node contributes to efficiency, but without a TDP figure, users cannot precisely plan for thermal loads. The 382 mm² die size is substantial for an integrated part, suggesting that adequate airflow over the chip is necessary for sustained boost clocks.

Given the unknown TDP and no PSU recommendation, buyers should ensure their system provides robust cooling for the chip, especially if they plan to sustain boost clocks of 2346 MHz under load. The lack of power connectors simplifies installation, but the 273.2 GB/s memory bandwidth and 128 GB capacity mean the memory subsystem may draw additional power through the system board. The PCIe 5.0 x16 interface ensures sufficient data transfer rates between the CPU and GPU, but the integrated nature means system-level thermal design is critical.

FAQ

Q: How much VRAM does the N1X 40SM have, and what type is it?

A: The N1X 40SM features 128 GB of LPDDR5X memory on a 256-bit bus, providing 273.2 GB/s of bandwidth. This is an exceptionally large memory pool for an integrated GPU, suited for data-intensive workloads.

Q: What is the FP32 compute performance of this GPU?

A: The FP32 performance is 24.02 TFLOPS, with FP16 also at 24.02 TFLOPS (1:1 ratio). This places it in the mid-range of GPUs, as evidenced by its 50th percentile rank against all GPUs.

Q: Does the N1X 40SM support DirectX or Vulkan?

A: The API support is listed as N/A for DirectX, OpenGL, and Vulkan. This suggests the GPU is not designed for traditional gaming APIs, but rather for compute or proprietary workloads.

Q: How many ray tracing and tensor cores does it have?

A: It has 40 ray tracing cores and 160 tensor cores, enabling hardware-accelerated ray tracing and AI-based compute tasks. The tensor cores support neural network workloads that can benefit from mixed-precision processing.

Q: What is the clock speed of the N1X 40SM?

A: The base clock is 741 MHz, and the boost clock is 2346 MHz. Memory runs at 1067 MHz, effective 8.5 Gbps. These clocks are moderate, reflecting the integrated design's power constraints.

Q: What power connectors does it require?

A: The N1X 40SM requires no power connectors (listed as "None"), as it is an integrated GPU that draws power from the system. The TDP is unknown, and no PSU recommendation is provided.

Q: What display outputs are available?

A: The GPU offers a single HDMI output, indicating a minimal display configuration. This aligns with its focus on compute workloads rather than multi-monitor gaming setups.

Memory Subsystem

The N1X 40SM's memory subsystem is its defining feature: 128 GB of LPDDR5X memory on a 256-bit bus, delivering 273.2 GB/s of bandwidth. This capacity is extraordinary for any GPU, let alone an integrated one, and it fundamentally shapes the product's use case. The 256-bit bus width is wide enough to support the bandwidth figure, but the 273.2 GB/s throughput is moderate compared to high-end discrete GPUs with GDDR6X or HBM memory. However, for workloads that require holding large datasets entirely in VRAM, the 128 GB capacity far outweighs the bandwidth limitation.

At high resolutions, the memory subsystem's impact is nuanced. The 273.2 GB/s bandwidth is sufficient for 1080p and 1440p gaming with high texture quality, as long as the 40 ROPs and 24.02 TFLOPS compute do not become the bottleneck. For 4K, the bandwidth may be insufficient for the most demanding textures, but the large capacity allows for massive texture caches that could reduce repeated loading. The 8.5 Gbps effective memory speed is modest, but LPDDR5X is optimized for power efficiency, which aligns with the integrated design. The pixel rate of 93.84 GPixel/s and texture rate of 750.7 GTexel/s suggest that the memory can feed the shader units adequately at lower resolutions, but the 40 ROPs will cap fill-rate performance.

The 128 GB capacity also enables unique workflows: users can load entire large language models, scientific datasets, or multi-layer image stacks into GPU memory without partitioning. The 256-bit bus ensures that data movement does not become a severe bottleneck, though the 273.2 GB/s bandwidth means that massive sequential transfers will take time. For compute tasks that repeatedly access the same data, the capacity is a clear advantage. For streaming workloads that demand high bandwidth, the N1X 40SM may lag behind narrower-bus but faster-memory solutions. The 160 tensor cores can leverage the large memory for AI workloads, but the overall performance envelope remains constrained by the mid-pack compute and memory speed.

Benchmark Performance

The N1X 40SM has an average benchmark score of 0 and a 50th percentile rank against all GPUs, with no benchmarks listed and no nearest rivals. This data absence means quantitative comparison is impossible, but the percentile provides a critical anchor: the GPU is exactly at the median of the database. The 24.02 TFLOPS FP32 compute is a concrete figure that places it in the range of mid-generation discrete GPUs, but the 50th percentile suggests that many alternatives offer similar or better raw performance.

Without rival names or deltaPct values, the analysis must rely on internal consistency. The 24.02 TFLOPS at a boost clock of 2346 MHz with 5120 shading units implies an efficiency of roughly 10.2 TFLOPS per thousand shading units, which is reasonable for a 5 nm process. The 40 RT cores and 160 tensor cores are present, but their performance is not quantified in benchmarks. The 273.2 GB/s bandwidth, combined with the 128 GB capacity, suggests that memory-bound workloads may see better-than-expected results due to the large cache-like pool, even if the raw bandwidth is not class-leading.

The 50th percentile rank is the single most informative benchmark metric. It indicates that in aggregate scoring, the N1X 40SM outperforms half of all GPUs and underperforms the other half. For users coming from older or lower-end GPUs, this represents a significant upgrade; for those with high-end discrete cards, it will be a step down. The 24.02 TFLOPS FP32 figure means that compute-heavy tasks like rendering or physics simulations will run at a moderate pace, but the 128 GB memory allows for problem sizes that would cripple GPUs with 16 GB or 24 GB VRAM. The FP16 1:1 ratio doubles the effective throughput for mixed-precision workloads, which can benefit AI inference tasks that use the tensor cores.

Ultimately, the N1X 40SM's benchmark performance is defined by its memory capacity rather than its compute speed. The 50th percentile rank confirms it is not a performance leader, but the 128 GB VRAM is a feature that no amount of TFLOPS can substitute for in capacity-bound workloads. The data shows a GPU that trades raw speed for memory size, making it a niche but potentially invaluable tool for specific professional and scientific applications.

The AMD Equivalent of N1X 40SM

Looking for a similar graphics card from AMD? The AMD Radeon RX 9050 offers comparable performance and features in the AMD lineup.

AMD Radeon RX 9050

AMD • 8 GB VRAM

View Specs Compare

Popular NVIDIA N1X 40SM Comparisons

See how the N1X 40SM stacks up against similar graphics cards from the same generation and competing brands.

Compare N1X 40SM with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs