NVIDIA A10G vs NVIDIA GeForce RTX 3090 Ti Comparison

NVIDIA
GEFORCE

NVIDIA A10G

CORE STATE GA102
VRAM 24 GB
CLOCK SPEED 1710 MHz
TDP 150 W
BUS WIDTH 384 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2021
VS
NVIDIA
GEFORCE

GeForce RTX 3090 Ti

CORE STATE GA102
VRAM 24 GB
CLOCK SPEED 1860 MHz
TDP 450 W
BUS WIDTH 384 bit
ARCHITECTURE Ampere
nm
PROCESS 8 nm
LAUNCH DATE 2022

PERFORMANCE BENCHMARKS

geekbench_opencl
158,063
174,441
geekbench_vulkan
145,863
215,633
3dmark_3dmark_steel_nomad_dx12
N/A
5,741

Analysis: NVIDIA A10G vs NVIDIA GeForce RTX 3090 Ti

FAQ

Q: Which GPU has a higher average benchmark score?

A: The NVIDIA A10G has an average benchmark score of 151,963, while the NVIDIA GeForce RTX 3090 Ti has an average score of 131,938. The A10G sits in the 97th percentile of all GPUs, whereas the 3090 Ti sits in the 95th percentile.

Q: How do the two compare in Vulkan compute performance?

A: The RTX 3090 Ti wins decisively in the Geekbench Vulkan test. It scores 215,633, which is 32.4% higher than the A10G’s score of 145,863.

Q: Which card offers more memory bandwidth?

A: The RTX 3090 Ti provides significantly more memory bandwidth. It uses GDDR6X memory running at 21 Gbps effective, delivering 1.01 TB/s, compared to the A10G’s GDDR6 memory at 12.5 Gbps effective, which provides 600.2 GB/s.

Q: Are there differences in the number of shading units?

A: Yes. The RTX 3090 Ti has 10,752 shading units, while the A10G has 9,216 shading units. This gives the 3090 Ti a higher theoretical FP32 throughput of 40.00 TFLOPS versus the A10G’s 31.52 TFLOPS.

Q: What is the power draw difference between the two cards?

A: The TDP differs substantially. The A10G is rated at 150 W with a suggested PSU of 450 W, while the RTX 3090 Ti is rated at 450 W with a suggested PSU of 850 W.

Q: Which GPU has a higher pixel fill rate?

A: The RTX 3090 Ti has a higher pixel rate. It achieves 208.3 GPixel/s, whereas the A10G achieves 164.2 GPixel/s.

---

Architecture Differences

Both GPUs are built on the same fundamental architecture, but they target different environments. The A10G and the RTX 3090 Ti both use the GA102 chip, the Ampere architecture, an 8 nm process from Samsung, and share identical transistor counts of 28,300 million on a 628 mm² die. The transistor density is also identical at 45.1M / mm².

The core configurations diverge. The A10G is configured with 9,216 shading units, 288 TMUs, 96 ROPs, 72 RT cores, and 288 tensor cores. The RTX 3090 Ti is a fuller implementation, with 10,752 shading units, 336 TMUs, 112 ROPs, 84 RT cores, and 336 tensor cores. This results in the 3090 Ti having a higher FP32 and FP16 compute throughput of 40.00 TFLOPS (1:1 ratio) compared to the A10G’s 31.52 TFLOPS (1:1 ratio).

Memory subsystems are a major point of difference. While both cards feature 24 GB of memory on a 384-bit bus, the A10G uses GDDR6 at 12.5 Gbps effective, yielding 600.2 GB/s of bandwidth. The RTX 3090 Ti uses faster GDDR6X at 21 Gbps effective, yielding 1.01 TB/s of bandwidth — a substantial 68% advantage in raw bandwidth.

Clock speeds also favor the consumer card. The RTX 3090 Ti has a base clock of 1560 MHz and a boost clock of 1860 MHz. The A10G is clocked lower, at 1320 MHz base and 1710 MHz boost. The memory clock is also higher on the 3090 Ti at 1313 MHz versus the A10G’s 1563 MHz, though the effective transfer rate is what differs most due to the memory type change.

Form factor and power delivery tell the story of their intended use cases. The A10G is a single-slot card, 267 mm long, with an 8-pin EPS power connector and a 150 W TDP. The RTX 3090 Ti is a triple-slot card, 336 mm long, requiring a single 16-pin connector and a 450 W TDP. The A10G has no display outputs, while the 3090 Ti includes 1x HDMI 2.1 and 3x DisplayPort 1.4a outputs. Both share the same API support: DirectX 12 Ultimate (12_2), OpenGL 4.6, and Vulkan 1.4.

The production status for both is end-of-life. The A10G was released on 2021-04-11, while the RTX 3090 Ti was released on 2022-01-26. The A10G’s predecessor is Tesla Turing, and its successor is Server Ada. The 3090 Ti’s predecessor is GeForce 20, and its successor is GeForce 40.

---

The Verdict

The data presents a clear split between compute capacity and power efficiency. The RTX 3090 Ti is the outright performance leader in every head-to-head benchmark recorded. It wins the Geekbench OpenCL test with a score of 174,441 versus the A10G’s 158,063, a 9.4% margin. In Vulkan, the margin is much larger: 215,633 versus 145,863, a 32.4% advantage.

However, the A10G should not be dismissed as a weaker product. Its average benchmark score of 151,963 is actually higher than the RTX 3090 Ti’s average of 131,938. This apparent contradiction is explained by the benchmark set. The A10G’s nearest rivals include the NVIDIA Tesla V100 PCIe 32 GB (average score 150,305, only 1.1% behind) and the AMD Instinct MI100 (average score 139,035, which the A10G beats by 9.3%). The 3090 Ti’s rivals are different, including the NVIDIA L4 (131,072, only 0.7% behind) and the AMD Radeon PRO W6800 (135,396, which beats the 3090 Ti by 2.6%).

The choice depends on the workload. For raw compute throughput, especially in Vulkan, the RTX 3090 Ti is the superior part. For a lower-power, single-slot solution that still delivers competitive average performance, the A10G is the better fit. The 3090 Ti’s 450 W TDP and triple-slot footprint are significant physical constraints that the A10G avoids entirely with its 150 W TDP and single-slot design. The A10G is a server-oriented card with no display outputs, while the 3090 Ti can drive displays directly.

---

Specification Differences

The following table lists only the fields where the two GPUs differ:

| Field | NVIDIA A10G | NVIDIA GeForce RTX 3090 Ti |

|:--- |:--- |:--- |

| Generation | Server Ampere (Axx) | GeForce 30 |

| Base Clock | 1320 MHz | 1560 MHz |

| Boost Clock | 1710 MHz | 1860 MHz |

| Memory Clock | 1563 MHz (12.5 Gbps effective) | 1313 MHz (21 Gbps effective) |

| Memory Type | GDDR6 | GDDR6X |

| Memory Bandwidth | 600.2 GB/s | 1.01 TB/s |

| Shading Units | 9216 | 10752 |

| TMUs | 288 | 336 |

| ROPs | 96 | 112 |

| RT Cores | 72 | 84 |

| Tensor Cores | 288 | 336 |

| Pixel Rate | 164.2 GPixel/s | 208.3 GPixel/s |

| Texture Rate | 492.5 GTexel/s | 625.0 GTexel/s |

| FP32 / FP16 | 31.52 TFLOPS (1:1) | 40.00 TFLOPS (1:1) |

| TDP | 150 W | 450 W |

| Slot Width | Single-slot | Triple-slot |

| Power Connectors | 8-pin EPS | 1x 16-pin |

| Suggested PSU | 450 W | 850 W |

| Display Outputs | No outputs | 1x HDMI 2.1, 3x DisplayPort 1.4a |

| Dimensions | 267 mm (10.5 in) length, 112 mm (4.4 in) height | 336 mm (13.2 in) length, 140 mm (5.5 in) height, 61 mm (2.4 in) width |

| Release Date | 2021-04-11 | 2022-01-26 |

| Predecessor | Tesla Turing | GeForce 20 |

| Successor | Server Ada | GeForce 40 |

| Launch MSRP | Not listed | 1,999 USD |

---

Head-to-Head Benchmarks

The head-to-head data is limited to two tests, both of which the RTX 3090 Ti wins. The first is Geekbench OpenCL. Here, the 3090 Ti scores 174,441 against the A10G’s 158,063. The delta is 9.4% in favor of the 3090 Ti. This is a moderate lead, reflecting the 3090 Ti’s higher core count and clock speeds.

The second test is Geekbench Vulkan. The outcome is far more lopsided. The RTX 3090 Ti scores 215,633, while the A10G scores 145,863. This represents a 32.4% advantage for the 3090 Ti. This large gap suggests that the 3090 Ti’s architecture is better optimized for the Vulkan API, or that its higher memory bandwidth (1.01 TB/s versus 600.2 GB/s) plays a more significant role in Vulkan workloads than in OpenCL.

It is worth remembering the A10G has zero wins in this head-to-head comparison. However, the average benchmark scores present a different perspective. The A10G’s average of 151,963 is 15.2% higher than the 3090 Ti’s average of 131,938. This is because the 3090 Ti’s average includes its scores from the 3DMark Steel Nomad DX12 test (5,741), which is a different workload that appears to lower its overall average. The A10G does not have a 3DMark score in the data, so its average is based on its two Geekbench results.

---

Where Each One Wins

The RTX 3090 Ti wins in raw compute and graphics performance. Its victories in both OpenCL and Vulkan are clear. The 32.4% lead in Vulkan is particularly strong, making it the better choice for applications that leverage this API. Its higher pixel rate (208.3 GPixel/s) and texture rate (625.0 GTexel/s) also indicate superior rasterization capability. With display outputs included, this card is suited for direct video output and high-end gaming or workstation tasks.

The A10G wins in efficiency and deployment flexibility. Its 150 W TDP is one-third of the 3090 Ti’s 450 W TDP. This drastically reduces power supply requirements (450 W suggested PSU versus 850 W) and cooling demands, allowing for dense server installations. The single-slot design, at 267 mm in length, is much easier to fit into chassis than the 3090 Ti’s triple-slot, 336 mm length. The A10G also has a higher average benchmark score overall (151,963 versus 131,938), indicating that in mixed workloads, its average performance is superior. Its nearest rival data shows it outperforms the AMD Instinct MI100 by 9.3%, while the 3090 Ti is beaten by the AMD Radeon PRO W6800 by 2.6%.

In summary, the RTX 3090 Ti is the performance pick for tasks that demand maximum throughput in compute or graphics, while the A10G is the pick for server environments where power, space, and average workload performance are the primary concerns.

DETAILED SPECIFICATIONS

SPECIFICATION
A10G
RTX 3090 Ti
Core Specs
Shading Units
9,216
10,752 +16.7%
Shaders
9,216
10,752 +16.7%
TMUs
288
336 +16.7%
ROPs
96
112 +16.7%
SM Count
72
84 +16.7%
Clocks
Base Clock
1320 MHz
1560 MHz
Boost Clock
1710 MHz
1860 MHz
Memory Clock
1563 MHz 12.5 Gbps effective
1313 MHz 21 Gbps effective
Memory
Memory Size
24 GB
24 GB
VRAM (MB)
24,576
24,576 0.0%
Memory Type
GDDR6
GDDR6X
Memory Bus
384 bit
384 bit
Bandwidth
600.2 GB/s
1.01 TB/s
Cache
L1 Cache
128 KB (per SM)
128 KB (per SM)
L2 Cache
6 MB
6 MB
Performance
Pixel Rate
164.2 GPixel/s
208.3 GPixel/s
Texture Rate
492.5 GTexel/s
625.0 GTexel/s
FP32 (TFLOPS)
31.52 TFLOPS
40.00 TFLOPS
FP64 (TFLOPS)
985.0 GFLOPS (1:32)
625.0 GFLOPS (1:64)
FP16 (TFLOPS)
31.52 TFLOPS (1:1)
40.00 TFLOPS (1:1)
AI/RT
RT Cores
72
84 +16.7%
Tensor Cores
288
336 +16.7%
Power
TDP
150 W
450 W
TDP (W)
150
450 +200.0%
Suggested PSU
450 W
850 W
Power Connectors
8-pin EPS
1x 16-pin
Architecture
Architecture
Ampere
Ampere
GPU Name
GA102
GA102
Generation
Server Ampere (Axx)
GeForce 30
Process Size
8 nm
8 nm
Transistors
28,300 million
28,300 million
Die Size
628 mm²
628 mm²
Foundry
Samsung
Samsung
Density
45.1M / mm²
45.1M / mm²
API Support
DirectX
12 Ultimate (12_2)
12 Ultimate (12_2)
OpenGL
4.6
4.6
Vulkan
1.4
1.4
OpenCL
3.0
3.0
CUDA
8.6
8.6
Shader Model
6.8
6.8
Physical
Slot Width
Single-slot
Triple-slot
Length
267 mm 10.5 inches
336 mm 13.2 inches
Height
112 mm 4.4 inches
140 mm 5.5 inches
Outputs
No outputs
1x HDMI 2.13x DisplayPort 1.4a
Bus Interface
PCIe 4.0 x16
PCIe 4.0 x16
Other
Launch Price
1,999 USD
Production
End-of-life
End-of-life
Predecessor
Tesla Turing
GeForce 20
Successor
Server Ada
GeForce 40
View A10G Details View GeForce RTX 3090 Ti Details