AMD Radeon HD 8790M vs NVIDIA Quadro K4000M Comparison

AMD
RADEON

AMD Radeon HD 8790M

CORE STATE Mars
VRAM 2 GB
CLOCK SPEED 900 MHz
TDP —
BUS WIDTH 128 bit
ARCHITECTURE GCN 1.0
nm
PROCESS 28 nm
LAUNCH DATE 2013
VS
NVIDIA
GEFORCE

Quadro K4000M

CORE STATE GK104
VRAM 4 GB
CLOCK SPEED 601 MHz
TDP 100 W
BUS WIDTH 256 bit
ARCHITECTURE Kepler
nm
PROCESS 28 nm
LAUNCH DATE 2012

PERFORMANCE BENCHMARKS

geekbench_opencl
5,017
5,986
geekbench_vulkan
6,365
N/A

Analysis: AMD Radeon HD 8790M vs NVIDIA Quadro K4000M

The NVIDIA Quadro K4000M and AMD Radeon HD 8790M are both end-of-life mobile workstation graphics solutions, but they approach the task from fundamentally different architectural directions. The data shows a single head-to-head benchmark result, yet the specification sheets reveal two distinct design philosophies. This analysis breaks down the measurable performance gap, the underlying silicon differences, and the practical implications for a buyer choosing between these two legacy parts.

Head-to-Head Benchmarks

The only direct benchmark comparison available is Geekbench OpenCL, and it decisively favors the NVIDIA Quadro K4000M. The K4000M scores 5986 points, while the Radeon HD 8790M manages 5017 points. That works out to a 19.3% advantage for the NVIDIA part. This is not a marginal win; it is a substantial margin that places the two GPUs in different performance tiers despite their similar market positioning.

To contextualize this score, the K4000M sits at the 34th percentile of all GPUs, with an average benchmark score of 5986. Its nearest rivals in the database are the AMD FirePro W4100 at 5987 (a 0% delta), the NVIDIA Quadro K4000 at 5982 (0.1% ahead), and even the NVIDIA RTX PRO 6000 Blackwell Server at 5996, which the K4000M trails by only 0.2%. This clustering indicates that the K4000M delivers a level of compute performance that is remarkably consistent with a broad range of other professional and even high-end consumer parts.

The Radeon HD 8790M, by contrast, produces an average benchmark score of 5691 across its two recorded tests, placing it at the 33rd percentile. Its nearest rival is the Intel Iris Pro Graphics P6300 at 5712, a 0.4% deficit, and the NVIDIA GeForce GTX 670MX at 5721, which it trails by 0.5%. The AMD part does hold a 1.6% lead over the NVIDIA Quadro M500M, which scores 5604. The deltaPct figures show that the HD 8790M trades blows with integrated and older discrete solutions, whereas the K4000M operates in a higher performance band.

When interpreting the head-to-head result, it is important to note that the K4000M’s 19.3% lead in OpenCL is consistent with its overall hardware profile. The NVIDIA chip has more raw compute resources, which directly translates to this benchmark outcome. The Radeon’s sole consolation is that its Geekbench Vulkan score of 6365 is higher than its OpenCL score, but since there is no Vulkan result for the K4000M, no direct comparison can be made there.

Architecture Differences

The two GPUs come from different architectural generations and design philosophies. The NVIDIA Quadro K4000M uses the GK104 chip, built on the Kepler architecture, and belongs to the Quadro Kepler-M (Kx000M) generation. The AMD Radeon HD 8790M uses the Mars chip, based on GCN 1.0, and belongs to the Solar System (HD 8700M) family. Both are fabricated by TSMC on a 28 nm process node, but the similarities end there.

The most striking difference is in die size and transistor count. The GK104 is a massive 294 mm² die containing 3,540 million transistors, yielding a transistor density of 12.0 million per mm². The Mars chip is far smaller at 77 mm², with 950 million transistors, which actually gives it a slightly higher density of 12.3 million per mm². This means AMD packed transistors more tightly, but NVIDIA used its larger budget to deploy far more execution units.

The K4000M features 960 shading units, 80 texture mapping units, and 32 ROPs. The HD 8790M counters with 384 shading units, 24 TMUs, and only 8 ROPs. This is a 2.5x advantage for NVIDIA in shading units and a 3.3x advantage in TMUs, with a 4x advantage in ROPs. The pixel rate tells the story clearly: the K4000M delivers 12.02 GPixel/s versus the Radeon’s 7.200 GPixel/s. Similarly, texture rate is 48.08 GTexel/s for NVIDIA versus 21.60 GTexel/s for AMD. The FP32 compute throughput is 1,153.9 GFLOPS versus 691.2 GFLOPS, a 67% advantage for the Quadro.

Clock speeds tell a different story. The Radeon runs at a base clock of 850 MHz and boosts to 900 MHz, while the Quadro is locked at 601 MHz for both base and boost. AMD’s higher clocks help close the gap, but cannot overcome the massive difference in execution resources. Memory also differs: the K4000M has 4 GB of GDDR5 on a 256-bit bus, delivering 89.60 GB/s of bandwidth, while the HD 8790M has 2 GB on a 128-bit bus, providing 64.00 GB/s. The NVIDIA part’s memory clocks at 700 MHz (2.8 Gbps effective), while AMD’s runs at 1000 MHz (4 Gbps effective).

FAQ

Q: Which GPU has a higher average benchmark score?

A: The NVIDIA Quadro K4000M has an average benchmark score of 5986, while the AMD Radeon HD 8790M has an average of 5691. The K4000M is ahead by roughly 5.2%.

Q: How much faster is the K4000M in the only direct head-to-head test?

A: In the Geekbench OpenCL test, the K4000M scores 5986 against the HD 8790M’s 5017, a 19.3% delta in favor of NVIDIA.

Q: Does the Radeon have any benchmark where it scores higher than its OpenCL result?

A: Yes, the Radeon HD 8790M records a Geekbench Vulkan score of 6365, which is higher than its OpenCL score of 5017. However, there is no Vulkan result for the K4000M to compare against.

Q: What are the respective process nodes for these GPUs?

A: Both the NVIDIA Quadro K4000M and the AMD Radeon HD 8790M are fabricated by TSMC on a 28 nm process node.

Q: Which GPU has more memory bandwidth?

A: The NVIDIA Quadro K4000M has 89.60 GB/s of bandwidth from its 256-bit bus and 4 GB of GDDR5, while the Radeon HD 8790M has 64.00 GB/s from a 128-bit bus and 2 GB of GDDR5.

Q: What is the transistor density difference between the two chips?

A: The AMD Mars chip has a slightly higher transistor density at 12.3 million per mm², compared to the NVIDIA GK104’s 12.0 million per mm², despite the NVIDIA chip having over three times the total transistors.

Specification Differences

The two GPUs differ in nearly every measurable hardware specification. The NVIDIA Quadro K4000M uses the GK104 chip with Kepler architecture, while the AMD Radeon HD 8790M uses the Mars chip with GCN 1.0. The K4000M has a 294 mm² die with 3,540 million transistors, versus the HD 8790M’s 77 mm² die with 950 million transistors. Transistor density is nearly identical at 12.0M versus 12.3M per mm².

Clock speeds favor AMD: the Radeon runs at 850 MHz base and 900 MHz boost, while the Quadro is fixed at 601 MHz for both. Memory clocks also favor AMD at 1000 MHz (4 Gbps effective) versus 700 MHz (2.8 Gbps effective). However, the NVIDIA part has superior memory capacity and bandwidth: 4 GB on a 256-bit bus delivering 89.60 GB/s, versus 2 GB on a 128-bit bus delivering 64.00 GB/s.

Compute resources heavily favor NVIDIA. The K4000M has 960 shading units, 80 TMUs, and 32 ROPs, while the HD 8790M has 384 shading units, 24 TMUs, and 8 ROPs. Pixel rate is 12.02 GPixel/s versus 7.200 GPixel/s, texture rate is 48.08 GTexel/s versus 21.60 GTexel/s, and FP32 performance is 1,153.9 GFLOPS versus 691.2 GFLOPS.

The TDP is listed as 100 W for the NVIDIA part, while no TDP is specified for the AMD part. Both use MXM modules, but the bus interfaces differ: the K4000M uses MXM-B (3.0), while the HD 8790M uses MXM-A (3.0). The NVIDIA part supports DirectX 12 (11_0), while the AMD part supports DirectX 12 (11_1). OpenGL support is identical at 4.6, but Vulkan versions differ: 1.2.175 for NVIDIA versus 1.2.170 for AMD. The K4000M was released on 2012-05-31, while the HD 8790M came later on 2013-03-31.

The Verdict

The benchmark data is unambiguous: the NVIDIA Quadro K4000M is the faster GPU in raw compute performance. Its 19.3% lead in Geekbench OpenCL is backed by a hardware specification sheet that shows overwhelming advantages in shading units, TMUs, ROPs, memory bandwidth, and FP32 throughput. The K4000M’s 960 shading units and 89.60 GB/s bandwidth are simply in a different class from the Radeon’s 384 units and 64.00 GB/s.

The Radeon HD 8790M does have a higher boost clock of 900 MHz versus 601 MHz, which indicates higher per-clock efficiency, but this cannot compensate for the 2.5x deficit in shading units. The AMD part’s smaller die and lower transistor count suggest it was designed for lower power consumption, though no TDP is listed for it. The K4000M’s 100 W TDP is explicitly stated.

For a buyer choosing between these two end-of-life mobile GPUs, the data points firmly toward the NVIDIA part if compute performance is the priority. The K4000M also offers double the memory capacity, which is significant for large datasets. However, the Radeon’s Vulkan score of 6365 hints at stronger performance in that specific API, though without a comparative K4000M result, this remains speculative.

Where Each One Wins

The NVIDIA Quadro K4000M wins in the only direct benchmark comparison, the Geekbench OpenCL test, by a margin of 19.3%. It also wins on every relevant hardware specification: compute units, texture units, ROPs, memory size, memory bus width, memory bandwidth, pixel rate, texture rate, and FP32 GFLOPS. Its 4 GB frame buffer is double that of the Radeon, which is a practical advantage for memory-intensive workloads. The K4000M’s higher percentile ranking (34th versus 33rd) and higher average score (5986 versus 5691) confirm its overall superiority.

The AMD Radeon HD 8790M wins on clock speed, with a 900 MHz boost versus 601 MHz, and on memory clock at 1000 MHz versus 700 MHz. It also has a slightly higher transistor density and supports a marginally newer DirectX version (12 (11_1) versus 12 (11_0)). The Radeon’s Vulkan score of 6365 is its one benchmark highlight, suggesting that in Vulkan-specific workloads, it may perform better than its OpenCL result implies. Its smaller die size and lower transistor count could also indicate lower power draw, though no TDP figure is provided to confirm this.

In practical terms, the K4000M is the choice for OpenCL compute tasks, memory-heavy applications, and any workload that benefits from its 4 GB frame buffer. The HD 8790M is a more modest part that might suffice for lighter duties, but the data shows it is outclassed by the NVIDIA option in nearly every measurable category.

DETAILED SPECIFICATIONS

SPECIFICATION
HD 8790M
Quadro K4000M
Core Specs
Shading Units
384
960 +150.0%
Shaders
384
960 +150.0%
TMUs
24
80 +233.3%
ROPs
8
32 +300.0%
Compute Units
6
—
Clocks
Base Clock
850 MHz
601 MHz
Boost Clock
900 MHz
601 MHz
Memory Clock
1000 MHz 4 Gbps effective
700 MHz 2.8 Gbps effective
Memory
Memory Size
2 GB
4 GB
VRAM (MB)
2,048
4,096 +100.0%
Memory Type
GDDR5
GDDR5
Memory Bus
128 bit
256 bit
Bandwidth
64.00 GB/s
89.60 GB/s
Cache
L1 Cache
16 KB (per CU)
16 KB (per SMX)
L2 Cache
256 KB
512 KB
Performance
Pixel Rate
7.200 GPixel/s
12.02 GPixel/s
Texture Rate
21.60 GTexel/s
48.08 GTexel/s
FP32 (TFLOPS)
691.2 GFLOPS
1,153.9 GFLOPS
FP64 (TFLOPS)
43.20 GFLOPS (1:16)
48.08 GFLOPS (1:24)
Power
TDP
—
100 W
TDP (W)
—
100
Power Connectors
None
None
Architecture
Architecture
GCN 1.0
Kepler
GPU Name
Mars
GK104
Generation
Solar System (HD 8700M)
Quadro Kepler-M (Kx000M)
Process Size
28 nm
28 nm
Transistors
950 million
3,540 million
Die Size
77 mm²
294 mm²
Foundry
TSMC
TSMC
Density
12.3M / mm²
12.0M / mm²
API Support
DirectX
12 (11_1)
12 (11_0)
OpenGL
4.6
4.6
Vulkan
1.2.170
1.2.175
OpenCL
2.1 (1.2)
3.0
CUDA
—
3.0
Shader Model
6.5 (5.1)
6.5 (5.1)
Physical
Slot Width
MXM Module
MXM Module
Outputs
Portable Device Dependent
Portable Device Dependent
Bus Interface
MXM-A (3.0)
MXM-B (3.0)
Other
Production
End-of-life
End-of-life
Predecessor
London
Quadro Fermi-M
Successor
Gem System
Quadro Maxwell-M
View Radeon HD 8790M Details View Quadro K4000M Details