AI Accelerator Trends 2022: Navigating the Post-Moore Era of Silicon

AI and ML accelerator survey and trends

2022-01-01
Albert Reuther, Peter Michaleas, Michael Jones, Vijay Gadepally, Siddharth Samsi, Jeremy Kepner
Summary
Problem
Method
Results
Takeaways
Abstract

This paper provides a comprehensive 2022 update on the AI accelerator landscape, categorizing commercial processors by peak performance (GOps/s) and power consumption. It highlights the continued dominance of Int8 for inference and the emergence of specialized architectures like Google TPUv4, NVIDIA Hopper (H100), and wafer-scale systems that achieve SOTA efficiency in both data center and embedded environments.

TL;DR

The MIT Lincoln Laboratory’s latest survey confirms that AI hardware is no longer just about "faster chips"—it’s a race of architectural efficiency. As Moore’s Law pales, the industry is pivoting to Dataflow architectures, Processor-in-Memory (PIM), and aggressive quantization (FP8/Int8) to keep up with the explosive demand of Transformers and edge AI.

Background: The End of General-Purpose Dominance

For decades, we relied on Dennard scaling and clock frequency boosts. Those days are gone. Today’s AI workloads are hitting a "Power Wall." This survey explores how vendors from NVIDIA to startups like Mythic are bypassing this wall by trading off functional flexibility for specialized AI kernels.

The Performance-Power Landscape

The core of the research is visualized in a massive scatter plot mapping Peak Performance (GOps/s) against Power (Watts).

AI Accelerator Performance vs Power

Key Observations:

  • The "Crowded" Data Center: The region between 100W and 700W is becoming extremely dense. NVIDIA’s Hopper (H100) and Graphcore’s Bow represent the high-performance frontier, while PCIe v5 is anticipated to break the current 300W limit.
  • Ultra-Low Power Frontier: At the <1W end, chips like the Maxim MAX78000 (30mW) enable battery-powered keyword spotting and vision, often using 1-bit or 2-bit operations.

Methodology: Why These Chips Work

The survey highlights two major architectural shifts:

  1. Dataflow Processing: Unlike CPUs that fetch instructions, dataflow processors (Cerebras, Groq) "place and route" the neural network graph directly onto the hardware. This eliminates cache-line overhead and makes power consumption deterministic.
  2. Processor-in-Memory (PIM): Startups like Mythic and Syntiant use analog compute or flash-memory circuits to perform multiply-accumulate (MAC) operations inside the memory itself, effectively "solving" the data movement bottleneck.

Architecture Comparison Zoomed

Precision Trends: The Rise of FP8

A critical takeaway is the evolution of numerical formats. While Int8 remains the standard for inference, the survey notes that FP8 is the new battleground for training. Platforms from NVIDIA, Intel (Habana Gaudi2), and Graphcore are all adopting 8-bit floating point to reduce memory bandwidth requirements by half compared to FP16, without significant accuracy loss.

Beyond AI: Accelerators as the New HPC

In a fascinating turn, these "AI chips" are being hijacked for traditional scientific computing. The survey highlights how Google TPUs and Cerebras wafers are now running:

  • Computational Fluid Dynamics (CFD)
  • Fast Fourier Transforms (FFTs)
  • Molecular Dynamics

Because these chips handle massive parallel arithmetic so well, they are becoming the default engines for any math-heavy simulation, not just neural networks.

Critical Insight & Conclusion

This survey proves that we have entered the Age of Domain-Specific Architectures (DSA). The winners in the next five years won't just have the most transistors; they will have the best "Software-Hardware Co-design," allowing developers to map complex graphs onto these heterogeneous dataflow fabrics.

Limitations

The report relies on "Peak Performance" numbers provided by vendors. Real-world "Sustainable Performance" often differs due to thermal throttling and software stack inefficiencies. Future surveys must look deeper into the Model-to-Silicon utilization ratio.

Takeaway: If you are building AI software, your next performance 10x won't come from code optimization alone—it will come from targeting the specific numerical and dataflow strengths of these new-age accelerators.

Find Similar Papers

Try Our Examples

  • Find the most recent survey papers from 2024 or 2025 that update the AI accelerator performance-to-power scatter plots for Blackwell-generation GPUs and dedicated NPU startups.
  • Which paper first formally defined the "Dataflow" architecture as used by Cerebras and Groq, and how does it fundamentally differ from von Neumann architectures in energy efficiency?
  • Search for research articles where Google TPUs or Graphcore IPUs were applied to non-AI High-Performance Computing (HPC) tasks such as fluid dynamics or financial Monte Carlo simulations.
Contents
AI Accelerator Trends 2022: Navigating the Post-Moore Era of Silicon
1. TL;DR
2. Background: The End of General-Purpose Dominance
3. The Performance-Power Landscape
3.1. Key Observations:
4. Methodology: Why These Chips Work
5. Precision Trends: The Rise of FP8
6. Beyond AI: Accelerators as the New HPC
7. Critical Insight & Conclusion
7.1. Limitations