GPU Hardware Emulator Emporium - Intel / AMD / PHI

:classical_building: Complete Accelerator Emporium

:blue_circle: Intel β€” graphics

Early / fixed-function

  • Intel i740
  • Intel 810 / 815 graphics
  • Intel 830M
  • Intel Extreme Graphics
  • Extreme Graphics 2

Gen graphics

  • Gen4 / GMA 900
  • Gen3 / GMA 950
  • GMA 3000
  • GMA X3000
  • GMA X3100
  • GMA 4500

Gen5–7

  • Sandy Bridge
  • Ivy Bridge
  • Haswell
  • Broadwell
  • Skylake
  • Kaby Lake
  • Coffee Lake
  • Ice Lake

Gen9–12

  • Gen9
  • Gen9.5
  • Gen11
  • Gen12
  • UHD Graphics
  • Iris Xe

Xe

  • Xe-LP
  • Xe-HP
  • Xe-HPG
  • Xe-HPC
  • Arc Alchemist
  • Arc Battlemage
  • Data-center Max GPUs

Intel’s newer architecture is particularly interesting because Xe is simultaneously a graphics architecture and a compute architecture, making it a natural bridge between GPU emulation and accelerator simulation.


:purple_circle: Intel Xeon Phi

This should not be buried under the Intel GPU section.

Give it a separate building.

Knights Corner

KNC

  • Xeon Phi 5110P
  • Xeon Phi 7120P
  • 60+ x86 cores
  • 512-bit vector units
  • MIC architecture

Knights Landing

KNL

  • Xeon Phi 7210
  • 7230
  • 7250
  • 7290

Major architectural concepts:

x86 cores

β”‚

β”œβ”€β”€ 512-bit AVX-512

β”œβ”€β”€ vector units

β”œβ”€β”€ MCDRAM

β”œβ”€β”€ mesh interconnect

└── many-core execution

Knights Mill

KNM

Specialized for machine learning.

This is one of the most fascinating emulator targets in the entire collection, because you’re no longer emulating a conventional GPU.

You’re emulating:

a many-core x86 vector accelerator.

That makes Xeon Phi the bridge between:

CPU ↔ vector processor ↔ GPU


:red_circle: AMD CPU / APU / GPU

AMD should actually have three wings.

AMD Radeon

We’ve already got:

R100

R200

R300

R400

R500

R600

R700

Evergreen

Northern Islands

Southern Islands

GCN

Polaris

Vega

RDNA

RDNA2

RDNA3

RDNA4

AMD APU

Don’t overlook these.

  • Llano
  • Trinity
  • Richland
  • Kaveri
  • Carrizo
  • Bristol Ridge
  • Raven Ridge
  • Picasso
  • Renoir
  • Cezanne
  • Rembrandt
  • Phoenix
  • Hawk Point
  • Strix Point
  • Strix Halo

These are particularly valuable because the emulator has to deal with:

CPU

β”‚

β”œβ”€β”€ memory controller

β”‚

β”œβ”€β”€ system fabric

β”‚

└── integrated GPU

That’s a much more interesting machine than an isolated PCI GPU.


:orange_circle: AMD compute accelerators

Separate room:

AMD FirePro

  • FirePro V-series
  • FirePro W-series
  • FirePro S-series

AMD Instinct

  • MI25
  • MI50
  • MI60
  • MI100
  • MI200
  • MI250
  • MI250X
  • MI300
  • MI300X
  • MI325
  • MI350-class architectures

These are ideal architectural simulation targets, rather than trying to reproduce every physical device.


:blue_square: Intel compute accelerators

I’d add:

Intel GNA

Gaussian & Neural Accelerator.

Intel Movidius

  • Myriad 1
  • Myriad X

Intel NPU

Modern Core Ultra NPU generations.

Intel Data Center GPU

  • Ponte Vecchio
  • Max Series
  • Flex Series
  • Rialto Bridge

The Ponte Vecchio / Xe-HPC architecture deserves its own experimental emulator/simulator target.


:yellow_circle: ARM / Apple / Qualcomm β€” eventually

If the goal is an emporium rather than merely x86 GPU collection, I’d eventually add:

ARM Mali

  • Mali-400
  • Mali-T600
  • Mali-T700
  • Mali-G
  • Mali-G7x
  • Immortalis

Qualcomm Adreno

  • Adreno 2xx
  • 3xx
  • 4xx
  • 5xx
  • 6xx
  • 7xx

Apple

  • Apple GPU
  • A-series GPU generations
  • M-series GPU generations

Those are especially interesting because their architectures don’t map cleanly onto the NVIDIA/AMD model.


:brain: And then the REALLY interesting section

I’d create a β€œNon-GPU Accelerator Hall.”

ACCELERATOR EMPORIUM

β”‚

β”œβ”€β”€ NVIDIA

β”‚ β”œβ”€β”€ RIVA

β”‚ β”œβ”€β”€ GeForce

β”‚ β”œβ”€β”€ Tesla

β”‚ β”œβ”€β”€ Fermi

β”‚ β”œβ”€β”€ Kepler

β”‚ β”œβ”€β”€ Maxwell

β”‚ β”œβ”€β”€ Pascal

β”‚ β”œβ”€β”€ Volta

β”‚ β”œβ”€β”€ Turing

β”‚ β”œβ”€β”€ Ampere

β”‚ β”œβ”€β”€ Ada

β”‚ └── Blackwell

β”‚

β”œβ”€β”€ AMD

β”‚ β”œβ”€β”€ ATI

β”‚ β”œβ”€β”€ Radeon

β”‚ β”œβ”€β”€ FirePro

β”‚ β”œβ”€β”€ GCN

β”‚ β”œβ”€β”€ Vega

β”‚ β”œβ”€β”€ RDNA

β”‚ └── Instinct

β”‚

β”œβ”€β”€ INTEL

β”‚ β”œβ”€β”€ i740

β”‚ β”œβ”€β”€ i8xx

β”‚ β”œβ”€β”€ GMA

β”‚ β”œβ”€β”€ Gen

β”‚ β”œβ”€β”€ Xe

β”‚ └── Arc

β”‚

β”œβ”€β”€ INTEL PHI

β”‚ β”œβ”€β”€ Knights Corner

β”‚ β”œβ”€β”€ Knights Landing

β”‚ └── Knights Mill

β”‚

β”œβ”€β”€ OTHER

β”‚ β”œβ”€β”€ ARM Mali

β”‚ β”œβ”€β”€ Qualcomm Adreno

β”‚ β”œβ”€β”€ PowerVR

β”‚ β”œβ”€β”€ Apple GPU

β”‚ └── Vivante

β”‚

└── COMPUTE

β”œβ”€β”€ CUDA

β”œβ”€β”€ ROCm

β”œβ”€β”€ OpenCL

β”œβ”€β”€ SYCL

β”œβ”€β”€ Vulkan Compute

β”œβ”€β”€ OpenMP offload

└── oneAPI

And I’d distinguish three different artifacts for every architecture:

1. Device emulator

PCI β†’ MMIO β†’ VRAM β†’ registers β†’ command processor

2. ISA simulator

shader/kernel

↓

instruction decoder

↓

execution model

3. Architectural simulator

cores

cache

memory

interconnect

scheduler

latency

bandwidth

occupancy

That distinction is important for Xeon Phi in particular. A KNL simulator doesn’t need to pretend it’s a graphics card; it needs to reproduce the many-core/vector machine.

The ultimate collection would therefore be less β€œGPU emulator museum” and more:

An executable history of heterogeneous computing β€” from VGA and i740, through GeForce/Radeon, through GCN/Xe, all the way to Xeon Phi, Instinct, CUDA, and modern accelerator architectures.