AMD Radeon PRO W6800 vs NVIDIA A10M Comparison

AMD
RADEON

AMD Radeon PRO W6800

CORE STATE Navi 21
VRAM 32 GB
CLOCK SPEED 2322 MHz
TDP 250 W
BUS WIDTH 256 bit
ARCHITECTURE RDNA 2.0
nm
PROCESS 7 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

A10M

CORE STATE GA102
VRAM 20 GB
CLOCK SPEED 1635 MHz
TDP 150 W
BUS WIDTH 320 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE —

PERFORMANCE BENCHMARKS

geekbench_metal
174,420
N/A
geekbench_opencl
121,808
135,230
geekbench_vulkan
109,961
N/A

Analysis: AMD Radeon PRO W6800 vs NVIDIA A10M

# FAQ

Q: How does the NVIDIA A10M compare to the AMD Radeon PRO W6800 in the only shared benchmark?

A: The A10M wins the Geekbench OpenCL test with a score of 135,230, while the W6800 scores 120,399. That gives the NVIDIA card a 12.3% performance advantage in this particular workload.

Q: Which GPU has higher average benchmark scores across all tests?

A: The NVIDIA A10M edges out the AMD Radeon PRO W6800 with an average score of 135,230 versus 133,588. The delta between the two is just 1.2% in favor of the A10M.

Q: What are the memory specifications for each card?

A: The A10M features 20 GB of GDDR6 memory on a 320-bit bus, delivering 500.2 GB/s of bandwidth. The W6800 offers 32 GB of GDDR6 on a 256-bit bus, with slightly higher bandwidth at 512.0 GB/s.

Q: Which GPU has higher clock speeds?

A: The AMD Radeon PRO W6800 runs with a base clock of 1575 MHz and a boost clock of 2322 MHz. The NVIDIA A10M is significantly lower, with a 975 MHz base and 1635 MHz boost.

Q: What is the difference in power consumption between the two cards?

A: The A10M has a TDP of just 150 W and requires a suggested 450 W power supply. The W6800 draws 250 W and needs a 600 W supply, making the NVIDIA card noticeably more power-efficient.

Q: Does the AMD card support more display outputs?

A: Yes. The W6800 provides six mini-DisplayPort 1.4a outputs, whereas the A10M has no display outputs at all — it is designed purely for compute/server workloads.

# Architecture Differences

The NVIDIA A10M and AMD Radeon PRO W6800 represent fundamentally different architectural approaches. The A10M is built on NVIDIA's Ampere architecture using the GA102 chip, fabricated on Samsung's 8 nm process. It packs 28,300 million transistors into a 628 mm² die, yielding a transistor density of 45.1 million per square millimeter. In contrast, the W6800 uses AMD's RDNA 2.0 architecture with the Navi 21 chip, produced on TSMC's 7 nm process. AMD's chip contains 26,800 million transistors on a smaller 520 mm² die, achieving a higher density of 51.5 million per square millimeter.

The compute layouts diverge significantly. The A10M deploys 7,168 shading units, 224 texture mapping units (TMUs), and 80 render output units (ROPs). It also includes 56 ray tracing cores and 224 tensor cores, which are critical for AI and machine learning workloads. The W6800, by contrast, has 3,840 shading units, 240 TMUs, and 96 ROPs, along with 60 ray tracing cores but no tensor core equivalent. This means the NVIDIA card is explicitly designed for accelerated AI inference and training, while the AMD card focuses on conventional graphics and compute.

Clock behavior reveals another key difference. The A10M operates at 975 MHz base and 1635 MHz boost, whereas the W6800 runs much higher at 1575 MHz base and 2322 MHz boost. Despite the higher clocks, the W6800's FP32 throughput is lower at 17.83 TFLOPS versus the A10M's 23.44 TFLOPS. However, the AMD card excels in FP16 performance, delivering 35.67 TFLOPS (2:1 ratio), while the A10M matches its FP32 figure at 23.44 TFLOPS (1:1 ratio). This makes the W6800 potentially better suited for workloads that leverage FP16 precision.

Both GPUs support DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4, ensuring compatibility with modern graphics APIs. The A10M uses an 8-pin EPS power connector and occupies a single slot, while the W6800 requires dual-slot spacing with a 6-pin plus 8-pin configuration. Physically, both cards share the same 267 mm length, but the W6800 is slightly taller at 120 mm versus 112 mm for the A10M, and it has a 50 mm width compared to unknown dimensions for the NVIDIA card.

# Head-to-Head Benchmarks

The only directly comparable benchmark between these two GPUs is Geekbench OpenCL, and the results clearly favor the NVIDIA A10M. The A10M scores 135,230, while the AMD Radeon PRO W6800 posts 120,399. This represents a 12.3% lead for the NVIDIA card — a substantial margin in a compute-focused test like OpenCL. The A10M's advantage likely stems from its higher FP32 throughput (23.44 TFLOPS versus 17.83 TFLOPS) and its large complement of shading units (7,168 versus 3,840).

However, the average benchmark score across all available tests narrows the gap considerably. The A10M's average matches its single OpenCL score at 135,230, while the W6800's average of 133,588 includes results from three different benchmarks: Geekbench Metal, OpenCL, and Vulkan. The W6800's Metal score of 171,137 is particularly strong, and its Vulkan score of 109,228 shows respectable performance. When considering only the shared OpenCL test, the A10M wins outright, but the W6800's broader benchmark portfolio demonstrates strengths in other API environments.

Looking at the nearest rivals provides context for these scores. The A10M sits just 0% behind the NVIDIA RTX 4000 Ada Generation (135,218) and 0.6% ahead of the AMD Radeon RX 9070 GRE (134,417). It also surpasses the NVIDIA GeForce RTX 3090 Ti (131,911) by 2.5%. The W6800, on the other hand, trails the RX 9070 GRE by 0.6%, the RTX 4000 Ada by 1.2%, and the A10M by 1.2%, while leading the RTX 3090 Ti by 1.3%. Both cards rank in the 97th percentile among all GPUs, placing them in the top tier of performance.

The deltaPct values in the nearestRivals data tell a compelling story. The A10M and W6800 are separated by just 1.2% in average score, making them effectively peer performers in overall benchmark terms. Yet the 12.3% OpenCL gap shows that the A10M has a decisive edge in specific compute scenarios. This suggests that the choice between these two cards depends heavily on the workload and API used.

# Specification Differences

| Specification | NVIDIA A10M | AMD Radeon PRO W6800 |

|---|---|---|

| Architecture | Ampere | RDNA 2.0 |

| Process Node | 8 nm (Samsung) | 7 nm (TSMC) |

| Transistors | 28,300 million | 26,800 million |

| Die Size | 628 mm² | 520 mm² |

| Base Clock | 975 MHz | 1575 MHz |

| Boost Clock | 1635 MHz | 2322 MHz |

| Memory Size | 20 GB GDDR6 | 32 GB GDDR6 |

| Memory Bus Width | 320 bit | 256 bit |

| Memory Bandwidth | 500.2 GB/s | 512.0 GB/s |

| Shading Units | 7,168 | 3,840 |

| TMUs | 224 | 240 |

| ROPs | 80 | 96 |

| Ray Tracing Cores | 56 | 60 |

| Tensor Cores | 224 | None |

| FP32 Performance | 23.44 TFLOPS | 17.83 TFLOPS |

| FP16 Performance | 23.44 TFLOPS (1:1) | 35.67 TFLOPS (2:1) |

| Pixel Rate | 130.8 GPixel/s | 222.9 GPixel/s |

| Texture Rate | 366.2 GTexel/s | 557.3 GTexel/s |

| TDP | 150 W | 250 W |

| Slot Width | Single-slot | Dual-slot |

| Power Connectors | 8-pin EPS | 1x 6-pin + 1x 8-pin |

| Suggested PSU | 450 W | 600 W |

| Display Outputs | None | 6x mini-DisplayPort 1.4a |

| Height | 112 mm | 120 mm |

| Width | Not specified | 50 mm |

| Release Date | Not specified | June 7, 2021 |

# Where Each One Wins

NVIDIA A10M wins for raw compute throughput and AI workloads. Its 23.44 TFLOPS FP32 performance is 31.5% higher than the W6800's 17.83 TFLOPS, and the inclusion of 224 tensor cores makes it a purpose-built solution for machine learning tasks. The 12.3% OpenCL benchmark victory confirms this advantage in real-world compute scenarios. Additionally, its 150 W TDP is 40% lower than the W6800's 250 W, making it a far more energy-efficient option for dense server deployments where power density matters. The single-slot form factor and 8-pin EPS connector also suit high-density rack environments.

AMD Radeon PRO W6800 wins for graphics workstations and memory capacity. The 32 GB of GDDR6 memory is 60% larger than the A10M's 20 GB, which is critical for massive datasets, complex 3D scenes, or GPU-accelerated rendering with large texture sets. The six mini-DisplayPort 1.4a outputs enable multi-monitor setups that the A10M simply cannot support. The higher pixel rate (222.9 GPixel/s versus 130.8 GPixel/s) and texture rate (557.3 GTexel/s versus 366.2 GTexel/s) indicate stronger rasterization performance for traditional graphics workloads. FP16 performance is also double that of the A10M, which benefits certain compute tasks that use reduced precision.

The W6800 also wins on clock speed and pixel throughput. Its 2322 MHz boost clock is 42% higher than the A10M's 1635 MHz, and the higher ROP count (96 versus 80) contributes to the substantial pixel rate advantage. For users working with the Metal API, the W6800's 171,137 Geekbench Metal score demonstrates strong Apple ecosystem compatibility. Its Vulkan score of 109,228 provides an alternative compute path, though it trails the OpenCL results of both cards.

In summary, the A10M is the compute specialist, while the W6800 is the graphics generalist. The data shows a clear split: NVIDIA's card dominates in OpenCL compute and AI acceleration, while AMD's card offers more memory, better display flexibility, and superior rasterization rates. The 1.2% difference in average benchmark scores suggests they are closely matched overall, but the specific strengths of each make them suitable for different professional environments. Both sit at the 97th percentile among all GPUs, indicating top-tier performance regardless of the choice.

DETAILED SPECIFICATIONS

SPECIFICATION
PRO W6800
A10M
Core Specs
Shading Units
3,840
7,168 +86.7%
Shaders
3,840
7,168 +86.7%
TMUs
240
224 -6.7%
ROPs
96
80 -16.7%
Compute Units
60
—
SM Count
—
56
Clocks
Base Clock
1575 MHz
975 MHz
Boost Clock
2322 MHz
1635 MHz
Memory Clock
2000 MHz 16 Gbps effective
1563 MHz 12.5 Gbps effective
Memory
Memory Size
32 GB
20 GB
VRAM (MB)
32,768
20,480 -37.5%
Memory Type
GDDR6
GDDR6
Memory Bus
256 bit
320 bit
Bandwidth
512.0 GB/s
500.2 GB/s
Cache
L1 Cache
128 KB per Array
128 KB (per SM)
L2 Cache
4 MB
6 MB
L3 Cache
128 MB
—
L0 Cache
32 KB per WGP
—
Performance
Pixel Rate
222.9 GPixel/s
130.8 GPixel/s
Texture Rate
557.3 GTexel/s
366.2 GTexel/s
FP32 (TFLOPS)
17.83 TFLOPS
23.44 TFLOPS
FP64 (TFLOPS)
1,114.6 GFLOPS (1:16)
732.5 GFLOPS (1:32)
FP16 (TFLOPS)
35.67 TFLOPS (2:1)
23.44 TFLOPS (1:1)
AI/RT
RT Cores
60
56 -6.7%
Tensor Cores
—
224
Power
TDP
250 W
150 W
TDP (W)
250
150 -40.0%
Suggested PSU
600 W
450 W
Power Connectors
1x 6-pin + 1x 8-pin
8-pin EPS
Architecture
Architecture
RDNA 2.0
Ampere
GPU Name
Navi 21
GA102
Generation
Radeon Pro Navi (Navi II Series)
Server Ampere (Axx)
Process Size
7 nm
8 nm
Transistors
26,800 million
28,300 million
Die Size
520 mm²
628 mm²
Foundry
TSMC
Samsung
Density
51.5M / mm²
45.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
2.1
3.0
CUDA
—
8.6
Shader Model
6.8
6.8
Physical
Slot Width
Dual-slot
Single-slot
Length
267 mm 10.5 inches
267 mm 10.5 inches
Height
120 mm 4.7 inches
112 mm 4.4 inches
Outputs
6x mini-DisplayPort 1.4a
No outputs
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Launch Price
2,249 USD
—
Production
End-of-life
End-of-life
Predecessor
Radeon Pro Vega
Tesla Turing
Successor
—
Server Ada
View Radeon PRO W6800 Details View A10M Details