請用此 Handle URI 來引用此文件:
http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103280完整後設資料紀錄
| DC 欄位 | 值 | 語言 |
|---|---|---|
| dc.contributor.advisor | 闕志達 | zh_TW |
| dc.contributor.advisor | Tzi-Dar Chiueh | en |
| dc.contributor.author | 黎慶玲 | zh_TW |
| dc.contributor.author | Khanh-Linh Le | en |
| dc.date.accessioned | 2026-08-10T16:21:08Z | - |
| dc.date.available | 2026-08-11 | - |
| dc.date.copyright | 2026-08-10 | - |
| dc.date.issued | 2026 | - |
| dc.date.submitted | 2026-07-30 00:00:00 | - |
| dc.identifier.citation | [1] S. Ali, T. Abuhmed, S. El-Sappagh, K. Muhammad, J. M. Alonso-Moral, R. Confalonieri, R. Guidotti, J. Del Ser, J. Díaz-Rodríguez, and F. Herrera, "Explainable artificial intelligence (XAI): What we know and what is left to attain trustworthy artificial intelligence," Information Fusion, vol. 99, Art. no. 101805, 2023.
[2] X. Liu, D. Huang, J. Yao, J. Dong, L. Song, H. Wang, C. Yao, and W. Chu, "From black box to glass box: A practical review of explainable artificial intelligence (XAI)," AI, vol. 6, no. 11, Art. no. 285, 2025. [3] F. Basheer, "Explainable AI in edge and IoT devices: A review of lightweight techniques," in Proc. IEEE 4th Int. Conf. Technology, Engineering, Management for Societal Impact Using Marketing, Entrepreneurship and Talent (TEMSMET), 2025. [4] H. Sun, Y. Liu, A. Al-Tahmeesschi, A. Nag, M. Soleimanpour, B. Canberk, H. Arslan, and H. Ahmadi, "Advancing 6G: Survey for explainable AI on communications and network slicing," IEEE Open J. Commun. Soc., vol. 6, pp. 1372–1412, Jan. 2025. [5] A. Bhat, A. S. Assoa, and A. Raychowdhury, "Gradient backpropagation based feature attribution to enable explainable AI on the edge," in Proc. IFIP/IEEE 30th Int. Conf. Very Large Scale Integr. (VLSI-SoC), 2022, pp. 1–6. [6] Z. Pan and P. Mishra, "Hardware acceleration of explainable machine learning," in Proc. Design, Autom. Test Eur. Conf. Exhib. (DATE), 2022, pp. 1127–1130. [7] J. Kim, S. Han, G. Ko, J. Kim, C. Le, T. Kim, C. Youn, and J. Kim, "EPU: An energy-efficient explainable AI accelerator with sparsity-free computation and heat map compression/pruning," IEEE J. Solid-State Circuits, vol. 59, no. 3, pp. 830–841, Mar. 2024. [8] A. Siddique, K. Khalil, and K. A. Hoque, "ApproXAI: Energy-efficient hardware acceleration of explainable AI using approximate computing," in Proc. Int. Joint Conf. Neural Netw. (IJCNN), Rome, Italy, Jun.–Jul. 2025, doi: 10.1109/IJCNN64981.2025.11227243. [9] P. Fantozzi and M. Naldi, "The explainability of Transformers: Current status and directions," Computers, vol. 13, no. 4, Art. no. 92, 2024. [10] S. Jain and B. C. Wallace, "Attention is not explanation," in Proc. Conf. North Amer. Chapter Assoc. Comput. Linguistics: Hum. Lang. Technol. (NAACL-HLT), vol. 1, Jun. 2019, pp. 3543–3556. [11] M. Sundararajan, A. Taly, and Q. Yan, "Axiomatic attribution for deep networks," in Proc. Int. Conf. Mach. Learn. (ICML), Jul. 2017, pp. 3319–3328. [12] C. Walker, S. Jha, K. Chen, and R. Ewetz, "Integrated decision gradients: Compute your attributions where the model makes its decision," in Proc. AAAI Conf. Artif. Intell., vol. 38, no. 6, pp. 5289–5297, Mar. 2024. [13] H.-S. Wang, D.-J. Jwo, and Y.-H. Liu, "Transformer-based ionospheric prediction and explainability analysis for enhanced GNSS positioning," Remote Sens., vol. 17, no. 1, Art. no. 81, 2025. [14] A. Kök, F. Y. Okay, Ö. Muyanlı, and S. Özdemir, "Explainable artificial intelligence (XAI) for internet of things: A survey," IEEE Internet Things J., vol. 10, no. 16, pp. 14764–14779, 2023. [15] M. T. Ribeiro, S. Singh, and C. Guestrin, "'Why should I trust you?' Explaining the predictions of any classifier," in Proc. ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining (KDD), Aug. 2016, pp. 1135–1144. [16] S. M. Lundberg and S.-I. Lee, "A unified approach to interpreting model predictions," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 30, 2017. [17] R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, "Grad-CAM: Visual explanations from deep networks via gradient-based localization," in Proc. IEEE Int. Conf. Comput. Vis. (ICCV), 2017, pp. 618–626. [18] J. Li, C. Zhang, J. T. Zhou, H. Fu, S. Xia, and Q. Hu, "Deep-LIFT: Deep label-specific feature learning for image annotation," IEEE Trans. Cybern., vol. 52, no. 8, pp. 7732–7741, Aug. 2022. [19] V. Miglani, N. Kokhlikyan, B. Alsallakh, M. Martin, and O. Reblitz-Richardson, "Investigating saturation effects in integrated gradients," arXiv preprint arXiv:2010.12697, 2020. [20] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, "Attention is all you need," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 30, 2017. [21] H. Shao, J. Lu, M. Wang, and Z. Wang, "An efficient training accelerator for Transformer with hardware-algorithm co-optimization," IEEE Trans. Very Large Scale Integr. (VLSI) Syst., vol. 31, no. 11, pp. 1788–1801, 2023. [22] N. Shazeer, "GLU variants improve Transformer," arXiv preprint arXiv:2002.05202, 2020. [23] D. Hendrycks and K. Gimpel, "Gaussian error linear units (GELUs)," arXiv preprint arXiv:1606.08415, 2016. [24] M. Lee, "GELU activation function in deep learning: A comprehensive mathematical analysis and performance," arXiv preprint arXiv:2305.12073, 2023. [25] J. L. Ba, J. R. Kiros, and G. E. Hinton, "Layer normalization," arXiv preprint arXiv:1607.06450, 2016. [26] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, "ImageNet: A large-scale hierarchical image database," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2009, pp. 248–255. [27] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, and J. Uszkoreit, "An image is worth 16×16 words: Transformers for image recognition at scale," arXiv preprint arXiv:2010.11929, 2020. [28] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, "Swin Transformer: Hierarchical vision transformer using shifted windows," in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2021, pp. 10012–10022. [29] H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, "Training data-efficient image transformers & distillation through attention," in Proc. Int. Conf. Mach. Learn. (ICML), Jul. 2021, pp. 10347–10357. [30] M. Ding, B. Xiao, N. Codella, P. Luo, J. Wang, and L. Yuan, "DaViT: Dual attention vision transformers," in Proc. Eur. Conf. Comput. Vis. (ECCV), Oct. 2022, pp. 74–92. [31] A. Kapishnikov, T. Bolukbasi, F. Viégas, and M. Terry, "XRAI: Better attributions through regions," in Proc. IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 2019, pp. 4948–4957. [32] V. Petsiuk, A. Das, and K. Saenko, "RISE: Randomized input sampling for explanation of black-box models," arXiv preprint arXiv:1806.07421, 2018. [33] N. Kokhlikyan et al., "Captum: A unified and generic model interpretability library for PyTorch," arXiv preprint arXiv:2009.07896, 2020. [34] A. Kapishnikov, S. Venugopalan, B. Avci, B. Wedin, M. Terry, and T. Bolukbasi, "Guided integrated gradients: An adaptive path method for removing noise," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2021, pp. 5050–5058. [35] R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts, "Recursive deep models for semantic compositionality over a sentiment treebank," in Proc. Conf. Empirical Methods Natural Lang. Process. (EMNLP), Seattle, WA, USA, 2013, pp. 1631–1642. [36] A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts, "Learning word vectors for sentiment analysis," in Proc. 49th Annu. Meeting Assoc. Comput. Linguistics: Hum. Lang. Technol. (ACL-HLT), Portland, OR, USA, 2011, pp. 142–150. [37] B. Pang and L. Lee, "Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales," in Proc. 43rd Annu. Meeting Assoc. Comput. Linguistics (ACL), Ann Arbor, MI, USA, 2005, pp. 115–124. [38] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," in Proc. Conf. North Amer. Chapter Assoc. Comput. Linguistics: Hum. Lang. Technol. (NAACL-HLT), 2019, pp. 4171–4186. [39] V. Sanh, L. Debut, J. Chaumond, and T. Wolf, "DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter," arXiv preprint arXiv:1910.01108, 2019. [40] J. Enguehard, S. O'Hagan, M. Hollifield, A. Hock, and N. Jacobs, "Discretized integrated gradients for explaining language models," in Proc. Conf. Empirical Methods Natural Lang. Process. (EMNLP), 2023. [41] A. Shrikumar, P. Greenside, and A. Kundaje, "Learning important features through propagating activation differences," in Proc. Int. Conf. Mach. Learn. (ICML), Jul. 2017, pp. 3145–3153. [42] J. DeYoung, S. Jain, N. F. Rajani, E. Lehman, C. Xiong, R. Socher, and B. C. Wallace, "ERASER: A benchmark to evaluate rationalized NLP models," in Proc. Annu. Meeting Assoc. Comput. Linguistics (ACL), Jul. 2020, pp. 4443–4458. [43] F. Chen, Ternary-Weight Transformer Acceleration Circuit Design and Chip Implementation, M.S. thesis, National Taiwan University, Taipei, Taiwan, 2025. [44] M. Hsieh, Y. Liu, and T. Chiueh, "A multiplier-less convolutional neural network inference accelerator for intelligent edge devices," IEEE J. Emerg. Sel. Topics Circuits Syst., vol. 11, no. 4, pp. 739–750, 2021. [45] F. Li, B. Liu, X. Wang, B. Zhang, and J. Yan, "Ternary weight networks," arXiv preprint arXiv:1605.04711, 2016. [46] Y. Bhalgat, J. Lee, M. Nagel, T. Blankevoort, and N. Kwak, "LSQ+: Improving low-bit quantization through learnable offsets and better initialization," in Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), 2020, pp. 696–697. [47] K. Prabhu, R. Radway, J. Yu, K. Bartolone, M. Giordano, F. Peddinghaus, Y. Urman, W. Khwa, Y. Chih, M. Chang, S. Mitra, and P. Raina, "MINOTAUR: An edge transformer inference and training accelerator with 12 Mbytes on-chip resistive RAM and fine-grained spatiotemporal power gating," in Proc. IEEE Symp. VLSI Technol. Circuits, 2024, pp. 1–2. [48] Y. Qin, Y. Wang, D. Deng, X. Yang, Z. Zhao, Y. Zhou, Y. Fan, J. Wei, T. Chen, L. Liu, S. Wei, Y. Hu, and S. Yin, "Ayaka: A versatile transformer accelerator with low-rank estimation and heterogeneous dataflow," IEEE J. Solid-State Circuits, vol. 59, no. 10, pp. 3342–3356, 2024. [49] S. Lim, J.-H. Kim, S. Moon, J. Cha, D. Seo, J. Kim, H. Lee, J. Lee, and J.-Y. Kim, "Adelia: A 4-nm LLM processing unit with streamlined dataflow and dual-mode parallelism for maximizing hardware efficiency," IEEE J. Solid-State Circuits, vol. 61, no. 4, pp. 1513–1525, 2026, doi: 10.1109/JSSC.2026.3663603. [50] S. Moon, M. Li, G. K. Chen, P. C. Knag, R. K. Krishnamurthy, and M. Seok, "T-REX: Hardware–software co-optimized transformer accelerator with reduced external memory access and enhanced hardware utilization," IEEE J. Solid-State Circuits, vol. 61, no. 1, pp. 154–167, 2026, doi: 10.1109/JSSC.2025.3603187. | - |
| dc.identifier.uri | http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103280 | - |
| dc.description.abstract | 隨著人工智慧(AI)模型在現代社會各個領域逐漸成為關鍵技術,建立具備可信度與可問責性的 AI 模型已成為重要需求。然而,隨著 AI 模型持續發展,其準確率與可解釋性之間存在著權衡。傳統機器學習方法因具有白盒特性而易於解釋,但如今已逐漸被具備高準確率、卻屬於黑盒模型的深度學習所取代。可解釋人工智慧(Explainable Artificial Intelligence, XAI)旨在透過事後解釋方法,說明黑盒模型輸入與輸出之間的關聯,以彌補此一缺口。近年來,Transformer 已成為最成功的深度學習模型之一,在影像分類與自然語言處理等多種應用中展現卓越效能,並逐漸取代傳統深度神經網路。與其他現代深度學習模型相同,Transformer 的決策過程難以解釋。然而,不同於其他模型的是,Transformer 本身具備注意力機制,因此已有許多研究嘗試利用注意力權重來解釋模型決策。然而,注意力機制是否真正能反映模型決策依據仍存在爭議,因此仍需透過其他事後解釋式 XAI 方法來說明 Transformer 模型輸入與輸出之間的關係。
XAI 領域至今已提出許多不同的方法,其中整合決策梯度(Integrated Decision Gradient, IDG)因能同時在視覺與自然語言處理任務中提供高品質解釋而受到關注。然而,IDG 為了提升解釋品質,需要對基準輸入與原始輸入之間進行多次插值,並重複執行前向傳播與反向傳播運算,因此具有較高的計算成本。基於此,有必要設計一套兼具高效能與低運算成本的可解釋性加速器。現有大多數 Transformer 加速器主要針對推論設計,尚無法有效支援以硬體方式實現梯度式解釋方法,例如 IDG。為解決此問題,本論文提出一款基於IDG的高硬體效率可解釋 Transformer 加速器晶片。 首先,本研究針對不同 Transformer 模型,包括用於影像分類的 Vision Transformer以及用於情感分類的 BERT,驗證 IDG 解釋結果的穩健性,並與其他 XAI 方法進行量化比較。由於解釋品質本身難以量化評估,目前學術界亦缺乏統一的評估基準,因此本研究參考相關文獻,選擇多項具代表性的評估指標,以建立較完整的量化評估方法。實驗結果顯示,相較於多種歸因方法,IDG 在視覺與自然語言 Transformer 模型上皆能提供較佳的解釋品質。 接著,本研究針對前向傳播與反向傳播運算進行硬體加速設計。前向傳播採用先前提出之三值權重四位元活化(Ternary-Weight 4-Bit Activation, TW4A)Transformer 模型,並搭配近似 Softmax 與 GELU 非線性運算,同時利用量化感知訓練維持近似後的模型準確率。BP 則採用 Signed-Digit-4(SD4)梯度量化,以及適合硬體實作的 Softmax 與 GELU 反向運算近似方法。此外,為提升模型效能,本研究針對不同校正與範圍映射百分位數進行分析,以選擇各層最佳的百分位設定。我們利用前述量化評估指標分析近似運算後的解釋品質變化,並以量化前之結果進行正規化比較。雖然正規化後的 SIC/AIC 相較於 FP32 IDG 最多下降約 3–5%,但由於評估結果係以 FP32 為基準進行正規化,因此實際的絕對差異仍相當有限。更重要的是,在採用 SD4 梯度量化與非線性近似運算後,整體解釋品質及其變化趨勢仍能獲得良好保留。 本研究提出之加速器支援 Vision Transformer、Swin Transformer 與 BERT 等模型,可同時應用於電腦視覺與自然語言處理任務,並透過 TSRI MPW 平台採用台積電TSMC28 奈米製程完成晶片實作。本設計採用資料重用策略,以降低片上記憶體使用量與記憶體存取成本,使完整 Transformer 運算僅需 82 KB 的片上記憶體即可完成。此外,考量 TSRI ADVANTEST V93000 測試平台的限制,本研究設計 Self-Test 模式,提供晶片內部驗證機制,使晶片可於高於測試設備限制的操作頻率下驗證其功能正確性。實作完成之晶片總面積為 6.16 mm²,其中核心面積為 4.33 mm²;正常模式操作頻率為 200 MHz,Self-Test 模式則可達 400 MHz。在 Self-Test 模式下,本設計可達到 13.1 TOPS 的運算效能,能源效率為 24.95 TOPS/W,面積效率則為 2.13 TOPS/mm²。後佈局模擬結果證實,本加速器於正常模式與 Self-Test 模式下皆能正確完成功能運作。 | zh_TW |
| dc.description.abstract | As AI models become a critical part of the modern world in many aspects, there is a need for trustworthy and accountable AI models. However, as AI models develop, there is a trade-off between accuracy and interpretability. The traditional machine learning techniques, which are easy to interpret thanks to their white-box characteristic, have now been replaced by robust deep learning models that achieve high accuracy but are black-box. XAI aims to fill this gap by introducing a post hoc method to explain the relationship between the input and output of black-box models. One of the most successful models in recent years is Transformer, which has demonstrated robust performance and has replaced traditional deep neural networks in different applications, including image classification and natural language processing. Like many other modern deep learning models, the decision-making process of the Transformer is difficult to interpret. But unlike other models, the Transformer itself has an attention mechanism, which several researchers have proposed using to explain its decisions. However, the attention mechanism is under dispute for its relation to explanation. Therefore, another post hoc XAI is required to explain the relationship between the Transformer model's inputs and outputs.
Many XAI methods have been proposed throughout the field's history. Among them, Integrated Decision Gradient (IDG) stands out for its high-quality explanations for both visual and NLP tasks. However, IDG requires multiple FP and BP computations to interpolate the baseline to the original input to increase its explainability, thereby requiring greater computational effort. Therefore, it was necessary to develop an explainable accelerator that could deliver high-performance explainability while reducing computational effort. Most of the existing Transformer accelerators are primarily designed for inference and do not efficiently support hardware-based gradient-based explanation generation, such as IDG. To address this gap, this thesis presents a hardware-efficient, explainable Transformer accelerator chip based on Integrated Decision Gradients (IDG). First, we validate the robustness of the IDG explanation across different Transformer models, including Vision Transformers and BERTs, for image and sentiment classification, respectively, and compare the quantitative results with those of other XAI methods. A quantitative explanation is difficult to evaluate and lacks a unified benchmark across all research. We sought to select several reliable metrics, informed by previous work, to provide a comprehensive evaluation approach. A quantitative comparison shows that IDG achieves better explanation results than multiple attribution methods across both vision and NLP Transformer models. Next, we consider the computing acceleration for both the Forward Pass (FP) and the Backward Pass (BP). The FP leverages previously developed ternary-weight-4-bit-activation (TW4A) Transformer models with approximate Softmax and GELU nonlinear activations, which employ QAT to maintain high accuracy after approximation. The BP employs Signed-Digit-4 (SD4) gradient quantization and hardware-friendly approximations for Softmax and GELU backward computation. In addition, to achieve better performance, multiple calibration and range-mapping percentiles are explored to determine the best percentile for each layer. We used the previous quantitative metrics to evaluate the trend after approximation and normalized them with the metric before quantization. Although the normalized SIC/AIC shows up to 3–5% degradation relative to FP32 IDG, the absolute metric difference remains small because the values are normalized against the FP32 baseline. More importantly, the overall explanation behavior and trend are still preserved under SD4 and approximate nonlinear computation. The accelerator IC proposed in this study, which supports Vision Transformer (ViT), Swin Transformer, and BERT models for both computer vision and natural language processing tasks, was implemented using the TSMC 28 nm process through the TSRI MPW platform. A data reuse strategy is employed to reduce on-chip memory usage and memory access overhead, enabling complete Transformer computation with only 82 KB of on-chip memory. The proposed accelerator, considering the TSRI ADVANTEST V93000 testing environment, includes a Self-Test mode that provides an internal validation method for testing the IC's robustness at higher frequencies than the environment's limitations allow. The implemented chip occupies a die area of 6.16 mm², with a core area of 4.33 mm², operating in normal mode at 200 MHz and in Self-Test mode at 400 MHz. In Self-Test mode, the design achieves a throughput of 13.1 TOPS, with an energy efficiency of 24.95 TOPS/W and an area efficiency of 2.13 TOPS/mm². Post-layout simulation results verify the accelerator's functional correctness in both normal and Self-Test modes. | en |
| dc.description.provenance | Submitted by admin ntu (admin@lib.ntu.edu.tw) on 2026-08-10T16:21:08Z No. of bitstreams: 0 | en |
| dc.description.provenance | Made available in DSpace on 2026-08-10T16:21:08Z (GMT). No. of bitstreams: 0 | en |
| dc.description.tableofcontents | ACKNOWLEDGEMENT I
摘要 III ABSTRACT V TABLE OF CONTENTS VIII TABLE OF FIGURES XIII LIST OF TABLES XVII CHAPTER 1 INTRODUCTION 1 1.1 RESEARCH BACKGROUND 1 1.2 RELATED WORK 4 1.3 RESEARCH MOTIVATION AND OBJECTIVES 5 1.4 THESIS ORGANIZATION AND CONTRIBUTION 7 CHAPTER 2 FOUNDATIONS OF EXPLAINABLE AI AND INTEGRATED DECISION GRADIENT 9 2.1 OVERVIEW OF XAI 9 2.1.1 Fundamental concepts of XAI 9 2.1.2 Taxonomy of XAI methods 11 2.1.2.1 Stage-based classification 11 2.1.2.2 Scope-based classification 11 2.1.2.3 Methods of application classification 12 2.1.3 Application of XAI 12 2.2 STATE-OF-THE-ART XAI METHODS 13 2.3 GRADIENT-BASED ATTRIBUTION METHODS 15 2.3.1 Backpropagation mechanism 15 2.3.2 Gradient saturation problem 15 2.3.3 Integrated Gradients 16 2.3.3.1 Integrated Gradients definition 16 2.3.3.2 IG limitation – Saturation effects within path-integrals 17 2.4 INTEGRATED DECISION GRADIENT 19 2.4.1 Important factor 19 2.4.2 Integrated Decision Gradients definition 20 2.4.3 Observation on integral path 21 2.4.4 Different Sampling Approaches 22 2.4.4.1 Uniform Sampling 22 2.4.4.2 Adaptive Sampling 22 2.4.4.3 Semi-uniform Sampling 23 2.5 CHAPTER 2 SUMMARY 24 CHAPTER 3 BACKGROUND OF TRANSFORMER NEURAL NETWORKS 26 3.1 TRANSFORMER ARCHITECTURE 26 3.1.1 Encoder-decoder structure 26 3.1.2 Multi-Head Attention 27 3.1.3 Positional encoding 29 3.2 FEED-FORWARD NETWORK 29 3.2.1 GELU 29 3.3 LAYER NORMALIZATION 30 3.4 BACKWARD COMPUTATION IN TRANSFORMER MODELS 31 3.4.1 Linear layer 33 3.4.2 Softmax layer 33 3.4.3 GELU layer 34 3.5 CHAPTER 3 SUMMARY 35 CHAPTER 4 QUANTIZATION AND APPROXIMATION FOR EXPLAINABLE TRANSFORMER ACCELERATION 36 4.1 OVERVIEW 36 4.2 EVALUATION OF IDG ON THE IMAGE CLASSIFICATION TASK 36 4.2.1 ImageNet dataset 37 4.2.2 Vision Transformer architectures 37 4.2.2.1 ViT 38 4.2.2.2 Swin 39 4.2.3 Experimental setting 42 4.2.4 Evaluation metrics 42 4.2.5 Experimental results 44 4.3 EVALUATION OF IDG ON SENTIMENT CLASSIFICATION TASKS 45 4.3.1 Datasets 45 4.3.2 BERT-based models 46 4.3.2.1 BERT 46 4.3.2.2 DistilBERT 46 4.3.3 Experimental setting 47 4.3.4 Evaluation metrics 47 4.3.5 Experimental results 49 4.4 QUANTIZATION STRATEGY FOR XAI ACCELERATION 49 4.4.1 Neural network quantization 49 4.4.2 Ternary weight quantization 50 4.4.3 Activation quantization 51 4.4.4 Gradient quantization 54 4.5 APPROXIMATION OF NON-LINEAR FUNCTIONS 58 4.5.1 Softmax-FP approximation (Multimax) 58 4.5.2 Softmax-BP approximation 59 4.5.3 GELU-FP approximation (QGELU) 60 4.5.4 GELU-BP approximation (DGELU) 61 4.6 EXPERIMENTAL EVALUATION 62 4.6.1 Image classification explanation evaluation 63 4.6.2 Sentiment classification explanation evaluation 65 4.7 CHAPTER 4 SUMMARY 67 CHAPTER 5 PROPOSED XAI ACCELERATOR ARCHITECTURE 69 5.1 CHIP ARCHITECTURE OVERVIEW 69 5.1.1 System architecture overview 69 5.1.2 Memory configuration 70 5.2 PROCESSING ELEMENT CUBE 70 5.2.1 Partial Product Generator architecture 70 5.2.2 PE unit 71 5.2.3 PE Cube 72 5.3 POST-PROCESSING UNIT (PPU) 74 5.3.1 PPU architecture overview 74 5.3.2 Accumulator unit 75 5.3.3 Requantizer unit 76 5.3.4 GELU-FP approximation (QGELU) 77 5.3.5 Derivative of GELU approximation (DGELU) 78 5.3.6 Pre-Sort 79 5.3.7 Softmax-BP 80 5.4 MULTIMAX CALCULATION UNIT 81 5.5 SELF-TEST CIRCUIT 82 5.6 CHIP ARCHITECTURE 83 5.7 CHIP COMPUTATION FLOW 84 5.8 SUMMARY OF CHIP-SUPPORTED COMPUTATIONS 87 5.9 CHAPTER 5 SUMMARY 88 CHAPTER 6 CHIP IMPLEMENTATION AND MEASUREMENT 89 6.1 DESIGN AND IMPLEMENTATION FLOW 89 6.2 CHIP VERIFICATION METHODOLOGY 92 6.3 CHIP IMPLEMENTATION RESULTS 93 6.3.1 Chip layout 93 6.3.2 Packaging and wire bonding 93 6.3.3 Chip specification 94 6.3.4 Comparison with State-of-the-Art designs 96 6.4 PERFORMANCE AND POWER CONSUMPTION ANALYSIS 97 6.5 SILICON MEASUREMENT AND VALIDATION 102 6.6 CHAPTER 6 SUMMARY 103 CHAPTER 7 RESEARCH CONCLUSIONS AND FUTURE DIRECTIONS 104 REFERENCES 107 | - |
| dc.language.iso | en | - |
| dc.subject | 加速器 | - |
| dc.subject | 深度學習 | - |
| dc.subject | XAI | - |
| dc.subject | IC | - |
| dc.subject | Transformer | - |
| dc.subject | IDG | - |
| dc.subject | SD4 | - |
| dc.subject | accelerator | - |
| dc.subject | deep learning | - |
| dc.subject | XAI | - |
| dc.subject | IC | - |
| dc.subject | Transformer | - |
| dc.subject | IDG | - |
| dc.subject | SD4 | - |
| dc.title | 基於整合決策梯度之可解釋 Transformer 加速晶片設計與實作 | zh_TW |
| dc.title | Design and Implementation of an Explainable Transformer Acceleration Chip Based on Integrated Decision Gradients | en |
| dc.type | Thesis | - |
| dc.date.schoolyear | 114-2 | - |
| dc.description.degree | 碩士 | - |
| dc.contributor.oralexamcommittee | 蔡佩芸;馬席彬 | zh_TW |
| dc.contributor.oralexamcommittee | Pei-Yun Tsai;Hsi-Pin Ma | en |
| dc.subject.keyword | 加速器; 深度學習; XAI; IC; Transformer; IDG; SD4 | zh_TW |
| dc.subject.keyword | accelerator; deep learning; XAI; IC; Transformer; IDG; SD4 | en |
| dc.relation.page | 112 | - |
| dc.identifier.doi | 10.6342/NTU202602785 | - |
| dc.rights.note | 同意授權(限校園內公開) | - |
| dc.date.accepted | 2026-08-03 | - |
| dc.contributor.author-college | 重點科技研究學院 | - |
| dc.contributor.author-dept | 積體電路設計與自動化學位學程 | - |
| dc.date.embargo-lift | 2026-08-11 | - |
| 顯示於系所單位: | 積體電路設計與自動化學位學程 | |
文件中的檔案:
| 檔案 | 大小 | 格式 | |
|---|---|---|---|
| ntu-114-2.pdf 授權僅限NTU校內IP使用(校園外請利用VPN校外連線服務) | 8.78 MB | Adobe PDF |
系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。
