DeepBurning: Bridging the Gap Between Neural Network Innovation and FPGA Deployment
10996_DeepBurning automatic generation of FPGA-based learning accelerators for the neural network family.
DeepBurning is an automated toolchain designed to generate customized FPGA-based hardware accelerators for a wide range of Neural Networks (NNs). By providing a library of modular functional building blocks and a specialized compiler, it maps high-level descriptors (like Caffe models) directly into RTL code, achieving high power efficiency and hardware utilization across diverse architectures including CNNs, MLPs, and Recurrent Networks.
TL;DR
DeepBurning is an automated framework that transforms high-level neural network descriptions (e.g., Caffe) into optimized FPGA hardware. By utilizing a library of "Functional Building Blocks" and a smart compiler, it eliminates the need for manual RTL coding, supporting everything from simple MLPs to complex architectures like AlexNet and GoogleNet with high efficiency.
Background & Motivation: The Hardware Bottleneck
The rapid evolution of neural networks (NNs) has created a significant challenge: while software frameworks allow for quick iteration, deploying these models on specialized hardware like FPGAs remains a grueling manual process. Traditional methods require hardware engineers to hand-craft RTL code for every new layer or architecture change.
The authors identify a critical "productivity gap." Existing automated tools are often too rigid, optimized for only one type of network (like CNNs), and fail when confronted with the diverse requirements of the broader "Neural Network Family" (including Recurrent and Associative networks).
Methodology: The "Lego" Approach to Hardware
DeepBurning's core innovation lies in its modularity and its compiler-driven synthesis.
1. Functional Building Blocks (FBBs)
Instead of synthesizing the entire network from scratch, DeepBurning uses a library of highly optimized, parameterized components. These include:
- Computation Units: Convolution, Fully-Connected, and Pooling modules.
- Non-linear Activation: Support for ReLU, Sigmoid, and Tanh.
- Data Management: Specialized Address Generation Units (AGUs) that handle the "tiling" and "folding" of data.
2. The Compiler and Resource Mapping
The DeepBurning compiler takes a model description and performs Temporal and Spatial Folding. This ensures that even if a model is too large for the FPGA's physical gates, the framework can reuse hardware components over time to process the entire workload.

Experimental Validation
The authors tested DeepBurning across a spectrum of networks to prove its versatility.
Broad Support Across the NN Family
As shown in the table below, DeepBurning successfully mapped diverse layers (Conv, FC, LRN, Dropout) across multiple SOTA models:
| Feature | MLP | CMAC | AlexNet | GoogleNet |
|---|---|---|---|---|
| Convolution | No | No | Yes | Yes |
| FC Layer | Yes | Yes | Yes | Yes |
| Pooling | No | No | Yes | Yes |
Performance and Accuracy
One of the primary concerns with automated generation is the loss of precision due to fixed-point arithmetic. DeepBurning utilizes a customizable data width, achieving nearly software-level accuracy (99%+) while maintaining the power-efficiency benefits of FPGAs.
The table above highlights how DeepBurning scales resource usage (DSP, LUT, FF) across different models from MNIST to the resource-intensive NiN (Network in Network).
Critical Insight: Why it Matters
DeepBurning represents a shift from "Hardware Design" to "Hardware Compilation." By abstracting the hardware details through a modular library, it allows AI researchers to treat FPGAs as a transparent execution target, similar to how they use GPUs.
Limitations: While revolutionary for its time (2016), the framework's reliance on static building blocks may struggle with the most recent dynamic "Attention" mechanisms found in Transformers without further updates to its FBB library.
Conclusion
DeepBurning provides a robust blueprint for the "Neural Network to Silicon" pipeline. Its emphasis on a flexible, modular architecture ensures that it isn't just a "CNN accelerator," but a comprehensive generation platform for the ever-expanding family of learning algorithms.
