{"product_id":"accelerators-for-convolutional-neural-networks-isbn-9781394171880","title":"Accelerators for Convolutional Neural Networks","description":"\u003cb\u003eAccelerators for Convolutional Neural Networks\u003c\/b\u003e \u003cp\u003e\u003cb\u003eComprehensive and thorough resource exploring different types of convolutional neural networks and complementary accelerators\u003c\/b\u003e \u003c\/p\u003e\u003cp\u003e\u003ci\u003eAccelerators for Convolutional Neural Networks\u003c\/i\u003e provides basic deep learning knowledge and instructive content to build up convolutional neural network (CNN) accelerators for the Internet of things (IoT) and edge computing practitioners, elucidating compressive coding for CNNs, presenting a two-step lossless input feature maps compression method, discussing arithmetic coding -based lossless weights compression method and the design of an associated decoding method, describing contemporary sparse CNNs that consider sparsity in both weights and activation maps, and discussing hardware\/software co-design and co-scheduling techniques that can lead to better optimization and utilization of the available hardware resources for CNN acceleration. \u003c\/p\u003e\u003cp\u003eThe first part of the book provides an overview of CNNs along with the composition and parameters of different contemporary CNN models. Later chapters focus on compressive coding for CNNs and the design of dense CNN accelerators. The book also provides directions for future research and development for CNN accelerators. \u003c\/p\u003e\u003cp\u003eOther sample topics covered in \u003ci\u003eAccelerators for Convolutional Neural Networks\u003c\/i\u003e include: \u003c\/p\u003e\u003cul\u003e\n\u003cli\u003e How to apply arithmetic coding and decoding with range scaling for lossless weight compression for 5-bit CNN weights to deploy CNNs in extremely resource-constrained systems\u003c\/li\u003e \u003cli\u003eState-of-the-art research surrounding dense CNN accelerators, which are mostly based on systolic arrays or parallel multiply-accumulate (MAC) arrays\u003c\/li\u003e \u003cli\u003eiMAC dense CNN accelerator, which combines image-to-column (im2col) and general matrix multiplication (GEMM) hardware acceleration\u003c\/li\u003e \u003cli\u003eMulti-threaded, low-cost, log-based processing element (PE) core, instances of which are stacked in a spatial grid to engender NeuroMAX dense accelerator\u003c\/li\u003e \u003cli\u003eSparse-PE, a multi-threaded and flexible CNN PE core that exploits sparsity in both weights and activation maps, instances of which can be stacked in a spatial grid for engendering sparse CNN accelerators\u003c\/li\u003e\n\u003c\/ul\u003e \u003cp\u003eFor researchers in AI, computer vision, computer architecture, and embedded systems, along with graduate and senior undergraduate students in related programs of study, \u003ci\u003eAccelerators for Convolutional Neural Networks\u003c\/i\u003e is an essential resource to understanding the many facets of the subject and relevant applications. \u003c\/p\u003e\u003cp\u003eAbout the Authors xiii\u003c\/p\u003e \u003cp\u003ePreface xv\u003c\/p\u003e \u003cp\u003e\u003cb\u003ePart I Overview 1\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e\u003cb\u003e1 Introduction 3\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e1.1 History and Applications 5\u003c\/p\u003e \u003cp\u003e1.2 Pitfalls of High-Accuracy DNNs\/CNNs 6\u003c\/p\u003e \u003cp\u003e1.2.1 Compute and Energy Bottleneck 6\u003c\/p\u003e \u003cp\u003e1.2.2 Sparsity Considerations 9\u003c\/p\u003e \u003cp\u003e1.3 Chapter Summary 11\u003c\/p\u003e \u003cp\u003e\u003cb\u003e2 Overview of Convolutional Neural Networks 13\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e2.1 Deep Neural Network Architecture 13\u003c\/p\u003e \u003cp\u003e2.2 Convolutional Neural Network Architecture 15\u003c\/p\u003e \u003cp\u003e2.3 Popular CNN Models 26\u003c\/p\u003e \u003cp\u003e2.4 Popular CNN Datasets 30\u003c\/p\u003e \u003cp\u003e2.5 CNN Processing Hardware 31\u003c\/p\u003e \u003cp\u003e2.6 Chapter Summary 37\u003c\/p\u003e \u003cp\u003e\u003cb\u003ePart II Compressive Coding for CNNs 39\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e\u003cb\u003e3 Contemporary Advances in Compressive Coding for CNNs 41\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e3.1 Background of Compressive Coding 41\u003c\/p\u003e \u003cp\u003e3.2 Compressive Coding for CNNs 43\u003c\/p\u003e \u003cp\u003e3.3 Lossy Compression for CNNs 43\u003c\/p\u003e \u003cp\u003e3.4 Lossless Compression for CNNs 44\u003c\/p\u003e \u003cp\u003e3.5 Recent Advancements in Compressive Coding for CNNs 48\u003c\/p\u003e \u003cp\u003e3.6 Chapter Summary 50\u003c\/p\u003e \u003cp\u003e\u003cb\u003e4 Lossless Input Feature Map Compression 51\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e4.1 Two-Step Input Feature Map Compression Technique 52\u003c\/p\u003e \u003cp\u003e4.2 Evaluation 55\u003c\/p\u003e \u003cp\u003e4.3 Chapter Summary 57\u003c\/p\u003e \u003cp\u003e\u003cb\u003e5 Arithmetic Coding and Decoding for 5-Bit CNN Weights 59\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e5.1 Architecture and Design Overview 60\u003c\/p\u003e \u003cp\u003e5.2 Algorithm Overview 63\u003c\/p\u003e \u003cp\u003e5.3 Weight Decoding Algorithm 67\u003c\/p\u003e \u003cp\u003e5.4 Encoding and Decoding Examples 69\u003c\/p\u003e \u003cp\u003e5.5 Evaluation Methodology 74\u003c\/p\u003e \u003cp\u003e5.6 Evaluation Results 75\u003c\/p\u003e \u003cp\u003e5.7 Chapter Summary 84\u003c\/p\u003e \u003cp\u003e\u003cb\u003ePart III Dense CNN Accelerators 85\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e\u003cb\u003e6 Contemporary Dense CNN Accelerators 87\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e6.1 Background on Dense CNN Accelerators 87\u003c\/p\u003e \u003cp\u003e6.2 Representation of the CNNWeights and Feature Maps in Dense Format 87\u003c\/p\u003e \u003cp\u003e6.3 Popular Architectures for Dense CNN Accelerators 89\u003c\/p\u003e \u003cp\u003e6.4 Recent Advancements in Dense CNN Accelerators 92\u003c\/p\u003e \u003cp\u003e6.5 Chapter Summary 93\u003c\/p\u003e \u003cp\u003e\u003cb\u003e7 iMAC: Image-to-Column and General Matrix Multiplication-Based Dense CNN Accelerator 95\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e7.1 Background and Motivation 95\u003c\/p\u003e \u003cp\u003e7.2 Architecture 97\u003c\/p\u003e \u003cp\u003e7.3 Implementation 99\u003c\/p\u003e \u003cp\u003e7.4 Chapter Summary 100\u003c\/p\u003e \u003cp\u003e\u003cb\u003e8 NeuroMAX: A Dense CNN Accelerator 101\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e8.1 RelatedWork 102\u003c\/p\u003e \u003cp\u003e8.2 Log Mapping 103\u003c\/p\u003e \u003cp\u003e8.3 Hardware Architecture 105\u003c\/p\u003e \u003cp\u003e8.4 Data Flow and Processing 108\u003c\/p\u003e \u003cp\u003e8.5 Implementation and Results 118\u003c\/p\u003e \u003cp\u003e8.6 Chapter Summary 124\u003c\/p\u003e \u003cp\u003e\u003cb\u003ePart IV Sparse CNN Accelerators 125\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e9 Contemporary Sparse CNN Accelerators 127\u003c\/p\u003e \u003cp\u003e\u003cb\u003e9.1 Background of Sparsity in CNN Models 127\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e9.2 Background of Sparse CNN Accelerators 128\u003c\/p\u003e \u003cp\u003e9.3 Recent Advancements in Sparse CNN Accelerators 131\u003c\/p\u003e \u003cp\u003e9.4 Chapter Summary 133\u003c\/p\u003e \u003cp\u003e\u003cb\u003e10 CNN Accelerator for In Situ Decompression and Convolution of Sparse Input Feature Maps 135\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e10.1 Overview 135\u003c\/p\u003e \u003cp\u003e10.2 Hardware Design Overview 135\u003c\/p\u003e \u003cp\u003e10.3 Design Optimization Techniques Utilized in the Hardware Accelerator 140\u003c\/p\u003e \u003cp\u003e10.4 FPGA Implementation 141\u003c\/p\u003e \u003cp\u003e10.5 Evaluation Results 143\u003c\/p\u003e \u003cp\u003e10.6 Chapter Summary 149\u003c\/p\u003e \u003cp\u003e\u003cb\u003e11 Sparse-PE: A Sparse CNN Accelerator 151\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e11.1 RelatedWork 155\u003c\/p\u003e \u003cp\u003e11.2 Sparse-PE 156\u003c\/p\u003e \u003cp\u003e11.3 Implementation and Results 174\u003c\/p\u003e \u003cp\u003e11.4 Chapter Summary 184\u003c\/p\u003e \u003cp\u003e\u003cb\u003e12 Phantom: A High-Performance Computational Core for Sparse CNNs 185\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e12.1 RelatedWork 189\u003c\/p\u003e \u003cp\u003e12.2 Phantom 190\u003c\/p\u003e \u003cp\u003e12.3 Phantom-2D 201\u003c\/p\u003e \u003cp\u003e12.4 Experiments and Results 209\u003c\/p\u003e \u003cp\u003e12.5 Chapter Summary 218\u003c\/p\u003e \u003cp\u003e\u003cb\u003ePart V HW\/SW Co-Design and Co-Scheduling for CNN Acceleration 221\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e\u003cb\u003e13 State-of-the-Art in HW\/SW Co-Design and Co-Scheduling for CNN Acceleration 223\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e13.1 HW\/SW Co-Design 223\u003c\/p\u003e \u003cp\u003e13.2 HW\/SW Co-Scheduling 228\u003c\/p\u003e \u003cp\u003e13.3 Chapter Summary 230\u003c\/p\u003e \u003cp\u003e\u003cb\u003e14 Hardware\/Software Co-Design for CNN Acceleration 231\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e14.1 Background of iMAC Accelerator 231\u003c\/p\u003e \u003cp\u003e14.2 Software Partition for iMAC Accelerator 232\u003c\/p\u003e \u003cp\u003e14.3 Experimental Evaluations 235\u003c\/p\u003e \u003cp\u003e14.4 Chapter Summary 237\u003c\/p\u003e \u003cp\u003e\u003cb\u003e15 CPU-Accelerator Co-Scheduling for CNN Acceleration 239\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003e15.1 Background and Preliminaries 240\u003c\/p\u003e \u003cp\u003e15.2 CNN Acceleration with CPU-Accelerator Co-Scheduling 242\u003c\/p\u003e \u003cp\u003e15.3 Experimental Results 251\u003c\/p\u003e \u003cp\u003e15.4 Chapter Summary 257\u003c\/p\u003e \u003cp\u003e\u003cb\u003e16 Conclusions 259\u003c\/b\u003e\u003c\/p\u003e \u003cp\u003eReferences 265\u003c\/p\u003e \u003cp\u003eIndex 285\u003c\/p\u003e  \u003cp\u003e\u003cb\u003eARSLAN MUNIR, PhD,\u003c\/b\u003e is an Associate Professor in the Department of Computer Science of Kansas State University. He is also the Director of the Intelligent Systems, Computer Architecture, Analytics, and Security (ISCAAS) Laboratory at the university. \u003c\/p\u003e\u003cp\u003e\u003cb\u003eJOONHO KONG, PhD,\u003c\/b\u003e is an Associate Professor in the School of Electronics Engineering College of IT Engineering at Kyungpook National University, South Korea. \u003c\/p\u003e\u003cp\u003e\u003cb\u003eMAHMOOD AZHAR QURESHI, PhD,\u003c\/b\u003e is a Senior IP Logic Design Engineer at Intel Corporation in Santa Clara, California.   \u003c\/p\u003e\u003cp\u003e\u003cb\u003eComprehensive and thorough resource exploring different types of convolutional neural networks and complementary accelerators\u003c\/b\u003e \u003c\/p\u003e\u003cp\u003e\u003ci\u003eAccelerators for Convolutional Neural Networks\u003c\/i\u003e provides basic deep learning knowledge and instructive content to build up convolutional neural network (CNN) accelerators for the Internet of things (IoT) and edge computing practitioners, elucidating compressive coding for CNNs, presenting a two-step lossless input feature maps compression method, discussing arithmetic coding -based lossless weights compression method and the design of an associated decoding method, describing contemporary sparse CNNs that consider sparsity in both weights and activation maps, and discussing hardware\/software co-design and co-scheduling techniques that can lead to better optimization and utilization of the available hardware resources for CNN acceleration. \u003c\/p\u003e\u003cp\u003eThe first part of the book provides an overview of CNNs along with the composition and parameters of different contemporary CNN models. Later chapters focus on compressive coding for CNNs and the design of dense CNN accelerators. The book also provides directions for future research and development for CNN accelerators. \u003c\/p\u003e\u003cp\u003eOther sample topics covered in \u003ci\u003eAccelerators for Convolutional Neural Networks\u003c\/i\u003e include: \u003c\/p\u003e\u003cul\u003e\n\u003cli\u003e How to apply arithmetic coding and decoding with range scaling for lossless weight compression for 5-bit CNN weights to deploy CNNs in extremely resource-constrained systems\u003c\/li\u003e \u003cli\u003eState-of-the-art research surrounding dense CNN accelerators, which are mostly based on systolic arrays or parallel multiply-accumulate (MAC) arrays\u003c\/li\u003e \u003cli\u003eiMAC dense CNN accelerator, which combines image-to-column (im2col) and general matrix multiplication (GEMM) hardware acceleration\u003c\/li\u003e \u003cli\u003eMulti-threaded, low-cost, log-based processing element (PE) core, instances of which are stacked in a spatial grid to engender NeuroMAX dense accelerator\u003c\/li\u003e \u003cli\u003eSparse-PE, a multi-threaded and flexible CNN PE core that exploits sparsity in both weights and activation maps, instances of which can be stacked in a spatial grid for engendering sparse CNN accelerators\u003c\/li\u003e\n\u003c\/ul\u003e \u003cp\u003eFor researchers in AI, computer vision, computer architecture, and embedded systems, along with graduate and senior undergraduate students in related programs of study, \u003ci\u003eAccelerators for Convolutional Neural Networks\u003c\/i\u003e is an essential resource to understanding the many facets of the subject and relevant applications.\u003c\/p\u003e","brand":"Wiley-IEEE Press","offers":[{"title":"Default Title","offer_id":47988652441829,"sku":"NP9781394171880","price":145.0,"currency_code":"USD","in_stock":false}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/1842\/7735\/files\/9781394171880.jpg?v=1761781125","url":"https:\/\/k12savings.com\/products\/accelerators-for-convolutional-neural-networks-isbn-9781394171880","provider":"K12savings","version":"1.0","type":"link"}