AI Accelerator Trends 2022: Navigating the Post-Moore Era of Silicon
AI and ML accelerator survey and trends
This paper provides a comprehensive 2022 update on the AI accelerator landscape, categorizing commercial processors by peak performance (GOps/s) and power consumption. It highlights the continued dominance of Int8 for inference and the emergence of specialized architectures like Google TPUv4, NVIDIA Hopper (H100), and wafer-scale systems that achieve SOTA efficiency in both data center and embedded environments.
TL;DR
The MIT Lincoln Laboratory’s latest survey confirms that AI hardware is no longer just about "faster chips"—it’s a race of architectural efficiency. As Moore’s Law pales, the industry is pivoting to Dataflow architectures, Processor-in-Memory (PIM), and aggressive quantization (FP8/Int8) to keep up with the explosive demand of Transformers and edge AI.
Background: The End of General-Purpose Dominance
For decades, we relied on Dennard scaling and clock frequency boosts. Those days are gone. Today’s AI workloads are hitting a "Power Wall." This survey explores how vendors from NVIDIA to startups like Mythic are bypassing this wall by trading off functional flexibility for specialized AI kernels.
The Performance-Power Landscape
The core of the research is visualized in a massive scatter plot mapping Peak Performance (GOps/s) against Power (Watts).

Key Observations:
- The "Crowded" Data Center: The region between 100W and 700W is becoming extremely dense. NVIDIA’s Hopper (H100) and Graphcore’s Bow represent the high-performance frontier, while PCIe v5 is anticipated to break the current 300W limit.
- Ultra-Low Power Frontier: At the <1W end, chips like the Maxim MAX78000 (30mW) enable battery-powered keyword spotting and vision, often using 1-bit or 2-bit operations.
Methodology: Why These Chips Work
The survey highlights two major architectural shifts:
- Dataflow Processing: Unlike CPUs that fetch instructions, dataflow processors (Cerebras, Groq) "place and route" the neural network graph directly onto the hardware. This eliminates cache-line overhead and makes power consumption deterministic.
- Processor-in-Memory (PIM): Startups like Mythic and Syntiant use analog compute or flash-memory circuits to perform multiply-accumulate (MAC) operations inside the memory itself, effectively "solving" the data movement bottleneck.

Precision Trends: The Rise of FP8
A critical takeaway is the evolution of numerical formats. While Int8 remains the standard for inference, the survey notes that FP8 is the new battleground for training. Platforms from NVIDIA, Intel (Habana Gaudi2), and Graphcore are all adopting 8-bit floating point to reduce memory bandwidth requirements by half compared to FP16, without significant accuracy loss.
Beyond AI: Accelerators as the New HPC
In a fascinating turn, these "AI chips" are being hijacked for traditional scientific computing. The survey highlights how Google TPUs and Cerebras wafers are now running:
- Computational Fluid Dynamics (CFD)
- Fast Fourier Transforms (FFTs)
- Molecular Dynamics
Because these chips handle massive parallel arithmetic so well, they are becoming the default engines for any math-heavy simulation, not just neural networks.
Critical Insight & Conclusion
This survey proves that we have entered the Age of Domain-Specific Architectures (DSA). The winners in the next five years won't just have the most transistors; they will have the best "Software-Hardware Co-design," allowing developers to map complex graphs onto these heterogeneous dataflow fabrics.
Limitations
The report relies on "Peak Performance" numbers provided by vendors. Real-world "Sustainable Performance" often differs due to thermal throttling and software stack inefficiencies. Future surveys must look deeper into the Model-to-Silicon utilization ratio.
Takeaway: If you are building AI software, your next performance 10x won't come from code optimization alone—it will come from targeting the specific numerical and dataflow strengths of these new-age accelerators.
