GEFORCE

NVIDIA A40 PCIe

NVIDIA graphics card specifications and benchmark scores

48 GB
VRAM
1740
MHz Boost
300W
TDP
384
Bus Width
Ray Tracing Tensor Cores

At a Glance

NVIDIA
VRAM 48 GB
Boost Clock 1,740 MHz
Shaders 10,752
Bus Width 384-bit
TDP 300W
Memory Type GDDR6
RT Cores 84
Architecture Ampere
nm
Process 8 nm
Released Oct 2020

NVIDIA A40 PCIe Specifications

GPU Core

Shader units and compute resources

The NVIDIA A40 PCIe GPU core specifications define its raw processing power for graphics and compute workloads. Shading units (also called CUDA cores, stream processors, or execution units depending on manufacturer) handle the parallel calculations required for rendering. TMUs (Texture Mapping Units) process texture data, while ROPs (Render Output Units) handle final pixel output. Higher shader counts generally translate to better GPU benchmark performance, especially in demanding games and 3D applications.

Shading Units
10,752
Shaders
10,752
TMUs
336
ROPs
112
SM Count
84

A40 PCIe Clock Speeds

GPU and memory frequencies

Clock speeds directly impact the A40 PCIe's performance in GPU benchmarks and real-world gaming. The base clock represents the minimum guaranteed frequency, while the boost clock indicates peak performance under optimal thermal conditions. Memory clock speed affects texture loading and frame buffer operations. The A40 PCIe by NVIDIA dynamically adjusts frequencies based on workload, temperature, and power limits to maximize performance while maintaining stability.

Base Clock
1305 MHz
Base Clock
1,305 MHz
Boost Clock
1740 MHz
Boost Clock
1,740 MHz
Memory Clock
1812 MHz 14.5 Gbps effective
GDDR GDDR 6X 6X

NVIDIA's A40 PCIe Memory

VRAM capacity and bandwidth

VRAM (Video RAM) is dedicated memory for storing textures, frame buffers, and shader data. The A40 PCIe's memory capacity determines how well it handles high-resolution textures and multiple displays. Memory bandwidth, measured in GB/s, affects how quickly data moves between the GPU and VRAM. Higher bandwidth improves performance in memory-intensive scenarios like 4K gaming. The memory bus width and type (GDDR6, GDDR6X, HBM) significantly influence overall GPU benchmark scores.

Memory Size
48 GB
VRAM
49,152 MB
Memory Type
GDDR6
VRAM Type
GDDR6
Memory Bus
384 bit
Bus Width
384-bit
Bandwidth
695.8 GB/s

A40 PCIe by NVIDIA Cache

On-chip cache hierarchy

On-chip cache provides ultra-fast data access for the A40 PCIe, reducing the need to fetch data from slower VRAM. L1 and L2 caches store frequently accessed data close to the compute units. AMD's Infinity Cache (L3) dramatically increases effective bandwidth, improving GPU benchmark performance without requiring wider memory buses. Larger cache sizes help maintain high frame rates in memory-bound scenarios and reduce power consumption by minimizing VRAM accesses.

L1 Cache
128 KB (per SM)
L2 Cache
6 MB

A40 PCIe Theoretical Performance

Compute and fill rates

Theoretical performance metrics provide a baseline for comparing the NVIDIA A40 PCIe against other graphics cards. FP32 (single-precision) performance, measured in TFLOPS, indicates compute capability for gaming and general GPU workloads. FP64 (double-precision) matters for scientific computing. Pixel and texture fill rates determine how quickly the GPU can render complex scenes. While real-world GPU benchmark results depend on many factors, these specifications help predict relative performance levels.

FP32 (Float)
37.42 TFLOPS
FP64 (Double)
584.6 GFLOPS (1:64)
FP16 (Half)
37.42 TFLOPS (1:1)
Pixel Rate
194.9 GPixel/s
Texture Rate
584.6 GTexel/s

A40 PCIe Ray Tracing & AI

Hardware acceleration features

The NVIDIA A40 PCIe includes dedicated hardware for ray tracing and AI acceleration. RT cores handle real-time ray tracing calculations for realistic lighting, reflections, and shadows in supported games. Tensor cores (NVIDIA) or XMX cores (Intel) accelerate AI workloads including DLSS, FSR, and XeSS upscaling technologies. These features enable higher visual quality without proportional performance costs, making the A40 PCIe capable of delivering both stunning graphics and smooth frame rates in modern titles.

RT Cores
84
Tensor Cores
336

Ampere Architecture & Process

Manufacturing and design details

The NVIDIA A40 PCIe is built on NVIDIA's Ampere architecture, which defines how the GPU processes graphics and compute workloads. The manufacturing process node affects power efficiency, thermal characteristics, and maximum clock speeds. Smaller process nodes pack more transistors into the same die area, enabling higher performance per watt. Understanding the architecture helps predict how the A40 PCIe will perform in GPU benchmarks compared to previous generations.

Architecture
Ampere
GPU Name
GA102
Process Node
8 nm
Foundry
Samsung
Transistors
28,300 million
Die Size
628 mm²
Density
45.1M / mm²

Power & Thermal

TDP and power requirements

Power specifications for the NVIDIA A40 PCIe determine PSU requirements and thermal management needs. TDP (Thermal Design Power) indicates the heat output under typical loads, guiding cooler selection. Power connector requirements ensure adequate power delivery for stable operation during demanding GPU benchmarks. The suggested PSU wattage accounts for the entire system, not just the graphics card. Efficient power delivery enables the A40 PCIe to maintain boost clocks without throttling.

TDP
300 W
TDP
300W
Power Connectors
8-pin EPS
Suggested PSU
700 W

A40 PCIe by NVIDIA Physical & Connectivity

Dimensions and outputs

Physical dimensions of the NVIDIA A40 PCIe are critical for case compatibility. Card length, height, and slot width determine whether it fits in your chassis. The PCIe interface version affects bandwidth for communication with the CPU. Display outputs define monitor connectivity options, with modern cards supporting multiple high-resolution displays simultaneously. Verify these specifications against your case and motherboard before purchasing to ensure a proper fit.

Slot Width
Dual-slot
Length
267 mm 10.5 inches
Height
111 mm 4.4 inches
Bus Interface
PCIe 4.0 x16
Display Outputs
3x DisplayPort 1.4a
Display Outputs
3x DisplayPort 1.4a

NVIDIA API Support

Graphics and compute APIs

API support determines which games and applications can fully utilize the NVIDIA A40 PCIe. DirectX 12 Ultimate enables advanced features like ray tracing and variable rate shading. Vulkan provides cross-platform graphics capabilities with low-level hardware access. OpenGL remains important for professional applications and older games. CUDA (NVIDIA) and OpenCL enable GPU compute for video editing, 3D rendering, and scientific applications. Higher API versions unlock newer graphical features in GPU benchmarks and games.

DirectX
12 Ultimate (12_2)
DirectX
12 Ultimate (12_2)
OpenGL
4.6
OpenGL
4.6
Vulkan
1.4
Vulkan
1.4
OpenCL
3.0
CUDA
8.6
Shader Model
6.8

A40 PCIe Product Information

Release and pricing details

The NVIDIA A40 PCIe is manufactured by NVIDIA as part of their graphics card lineup. Release date and launch pricing provide context for comparing GPU benchmark results with competing products from the same era. Understanding the product lifecycle helps evaluate whether the A40 PCIe by NVIDIA represents good value at current market prices. Predecessor and successor information aids in tracking generational improvements and planning future upgrades.

Manufacturer
NVIDIA
Release Date
Oct 2020
Production
End-of-life
Predecessor
Tesla Turing
Successor
Server Ada

About NVIDIA A40 PCIe

How It Compares

The NVIDIA A40 PCIe occupies a unique position in the database: it is an end-of-life server accelerator based on the GA102 chip from the Ampere architecture. Its percentile rank of 50 against all GPUs places it squarely in the middle of the field, but that figure is misleading because the A40 is designed for professional and datacenter workloads, not consumer gaming. The benchmark data shows no nearest rivals are listed, meaning the A40 stands alone in this dataset without direct comparative scores from other cards. Consequently, its positioning must be inferred from its architectural traits rather than head-to-head measurements.

The A40's closest conceptual predecessor is the Tesla Turing generation, and its successor is the Server Ada line. This generational placement indicates NVIDIA intended the A40 as a bridge between the Turing-era compute accelerators and the newer Ada-based server parts. Without rival scores, the data suggests the A40 is a specialized tool rather than a mainstream product, and its 50th percentile ranking reflects a broad distribution where it neither dominates nor lags. The absence of benchmark entries in the fact pack means no average score exists, so relative performance can only be discussed through its raw specifications.

Ray Tracing and Feature Set

The A40 integrates 84 RT cores and 336 tensor cores, making it a fully capable ray tracing accelerator. The RT cores are dedicated hardware for bounding volume hierarchy traversal and ray intersection tests, which offload this work from the shading units. The tensor cores, numbering 336, are present for AI-accelerated workloads such as denoising, deep learning super sampling, and inference tasks. Both core types are standard for the Ampere generation, and their counts align with the GA102 die's full configuration.

API support is comprehensive for the card's era: DirectX 12 Ultimate with feature level 12_2 is listed, which includes hardware ray tracing, variable rate shading, and mesh shaders. Vulkan 1.4 and OpenGL 4.6 round out the modern graphics APIs. The card supports three DisplayPort 1.4a outputs, which is notable for a server-focused product, as it allows direct display output for visualization workloads. The feature set is oriented toward professional rendering and compute, not just headless datacenter operation.

Power and Cooling

The A40 carries a thermal design power of 300 W, which dictates its cooling and power delivery requirements. NVIDIA recommends a 700 W power supply for systems using this card, a figure that accounts for the rest of the system's draw beyond the GPU itself. The power connector is an 8-pin EPS, which is a departure from the more common PCIe 8-pin connectors used on consumer cards. This EPS connector is typically found on server motherboards and power supplies, reinforcing the A40's datacenter positioning.

Cooling is handled by a dual-slot design, which is standard for high-TDP accelerators in rack-mounted servers. The card's physical dimensions are 267 mm in length (10.5 inches) and 111 mm in height (4.4 inches), making it longer than many consumer flagship cards but shorter than some triple-slot server accelerators. The dual-slot cooler is a capable air-based solution for the 300 W envelope, and the 8-pin EPS connector ensures stable power delivery under sustained load. The card uses PCIe 4.0 x16 for host connectivity, which provides sufficient bandwidth for its 48 GB memory pool.

FAQ

Q: What is the memory configuration of the A40?

A: The A40 features 48 GB of GDDR6 memory on a 384-bit bus, delivering 695.8 GB/s of bandwidth. The memory operates at 1812 MHz, which translates to 14.5 Gbps effective.

Q: Does the A40 support ray tracing?

A: Yes, the A40 includes 84 RT cores dedicated to ray tracing acceleration, and it supports DirectX 12 Ultimate (12_2), which mandates hardware ray tracing support.

Q: What is the card's compute performance in FP32?

A: The A40 delivers 37.42 TFLOPS of FP32 compute and an identical 37.42 TFLOPS in FP16, indicating a 1:1 ratio for those precision formats.

Q: What power supply is recommended for the A40?

A: NVIDIA recommends a 700 W power supply. The card itself has a TDP of 300 W and requires a single 8-pin EPS power connector.

Q: What is the production status of the A40?

A: The A40 is listed as end-of-life. It was released on October 4, 2020, and its predecessor is the Tesla Turing generation, with the Server Ada line as its successor.

Q: How many display outputs does the A40 have?

A: It has three DisplayPort 1.4a outputs, allowing direct video output despite its server-oriented design.

Benchmark Performance

No benchmark scores are available for the A40 in the fact pack; the benchmarks array is empty, and the average benchmark score is 0. This absence of data means the card cannot be compared numerically to any rivals, as the nearestRivals list is also empty. The only quantitative performance indicator is the 50th percentile rank against all GPUs, but without a distribution of scores, this percentile cannot be interpreted as a specific performance level. The data simply records that the A40 sits at the median of the database's GPU population.

Given this lack of direct measurements, performance analysis must rest on architectural specifications. The A40's 10,752 shading units, 336 TMUs, and 112 ROPs are the highest counts for the GA102 die, matching the full configuration. Its pixel rate is 194.9 GPixel/s, and its texture rate is 584.6 GTexel/s. These figures indicate the A40 is a fully enabled GA102 part, with no disabled subunits. In FP32 compute, the 37.42 TFLOPS figure is substantial for a 300 W card, and the 1:1 FP16 ratio is unusual for Ampere, as many consumer cards halve FP16 throughput.

The memory subsystem is also noteworthy: 48 GB of GDDR6 on a 384-bit bus is the maximum capacity for this memory interface, and the 695.8 GB/s bandwidth is sufficient for large datasets in scientific computing or AI inference. The card's 8 nm process node, manufactured by Samsung, houses 28,300 million transistors on a 628 mm² die, resulting in a transistor density of 45.1 million per square millimeter. These physical characteristics are identical to the consumer RTX 3090, but the A40's drivers and memory capacity target professional workloads.

Who Should Consider It

The A40 PCIe is tailored for professionals who need large memory capacity and high compute throughput in a dual-slot form factor. The 48 GB VRAM is the primary differentiator; it allows loading large neural network models, high-resolution 3D scenes, or massive scientific datasets entirely into GPU memory without spilling to system RAM. Users working with datasets that exceed 24 GB—the typical capacity of consumer flagships—would find the A40's memory pool essential.

For rendering and visualization workloads, the 84 RT cores and 37.42 TFLOPS FP32 performance handle ray-traced frames and complex shading. The three DisplayPort 1.4a outputs mean the card can drive multiple 4K displays for interactive visualization or digital signage, which is uncommon for server accelerators. The dual-slot design and 267 mm length make it compatible with most server chassis and many workstation towers, though the 8-pin EPS connector requires a motherboard or power supply that provisions this server-standard plug.

The card is not aimed at gamers or consumer enthusiasts. Its 300 W TDP and 700 W PSU recommendation are modest for its compute power, but the lack of consumer display features like HDMI and the server-oriented EPS connector make it an awkward fit for desktop builds. The end-of-life status means it is likely being phased out in favor of Server Ada parts, but for existing infrastructure or secondhand acquisitions, the A40 remains a capable compute accelerator. Its 50th percentile ranking across all GPUs suggests it is neither a performance outlier nor a weak link; it is a balanced, high-memory workhorse for datacenter and professional visualization tasks.

Detailed benchmark scores and charts for the NVIDIA A40 PCIe are below.

Benchmark Scores

No benchmark data available for this GPU.

Compare with Other GPUs

Select another GPU to compare specifications and benchmarks side-by-side.

Browse GPUs