請用此 Handle URI 來引用此文件:
http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103280| 標題: | 基於整合決策梯度之可解釋 Transformer 加速晶片設計與實作 Design and Implementation of an Explainable Transformer Acceleration Chip Based on Integrated Decision Gradients |
| 作者: | 黎慶玲 Khanh-Linh Le |
| 指導教授: | 闕志達 Tzi-Dar Chiueh |
| 關鍵字: | 加速器; 深度學習; XAI; IC; Transformer; IDG; SD4 accelerator; deep learning; XAI; IC; Transformer; IDG; SD4 |
| 出版年 : | 2026 |
| 學位: | 碩士 |
| 摘要: | 隨著人工智慧(AI)模型在現代社會各個領域逐漸成為關鍵技術,建立具備可信度與可問責性的 AI 模型已成為重要需求。然而,隨著 AI 模型持續發展,其準確率與可解釋性之間存在著權衡。傳統機器學習方法因具有白盒特性而易於解釋,但如今已逐漸被具備高準確率、卻屬於黑盒模型的深度學習所取代。可解釋人工智慧(Explainable Artificial Intelligence, XAI)旨在透過事後解釋方法,說明黑盒模型輸入與輸出之間的關聯,以彌補此一缺口。近年來,Transformer 已成為最成功的深度學習模型之一,在影像分類與自然語言處理等多種應用中展現卓越效能,並逐漸取代傳統深度神經網路。與其他現代深度學習模型相同,Transformer 的決策過程難以解釋。然而,不同於其他模型的是,Transformer 本身具備注意力機制,因此已有許多研究嘗試利用注意力權重來解釋模型決策。然而,注意力機制是否真正能反映模型決策依據仍存在爭議,因此仍需透過其他事後解釋式 XAI 方法來說明 Transformer 模型輸入與輸出之間的關係。
XAI 領域至今已提出許多不同的方法,其中整合決策梯度(Integrated Decision Gradient, IDG)因能同時在視覺與自然語言處理任務中提供高品質解釋而受到關注。然而,IDG 為了提升解釋品質,需要對基準輸入與原始輸入之間進行多次插值,並重複執行前向傳播與反向傳播運算,因此具有較高的計算成本。基於此,有必要設計一套兼具高效能與低運算成本的可解釋性加速器。現有大多數 Transformer 加速器主要針對推論設計,尚無法有效支援以硬體方式實現梯度式解釋方法,例如 IDG。為解決此問題,本論文提出一款基於IDG的高硬體效率可解釋 Transformer 加速器晶片。 首先,本研究針對不同 Transformer 模型,包括用於影像分類的 Vision Transformer以及用於情感分類的 BERT,驗證 IDG 解釋結果的穩健性,並與其他 XAI 方法進行量化比較。由於解釋品質本身難以量化評估,目前學術界亦缺乏統一的評估基準,因此本研究參考相關文獻,選擇多項具代表性的評估指標,以建立較完整的量化評估方法。實驗結果顯示,相較於多種歸因方法,IDG 在視覺與自然語言 Transformer 模型上皆能提供較佳的解釋品質。 接著,本研究針對前向傳播與反向傳播運算進行硬體加速設計。前向傳播採用先前提出之三值權重四位元活化(Ternary-Weight 4-Bit Activation, TW4A)Transformer 模型,並搭配近似 Softmax 與 GELU 非線性運算,同時利用量化感知訓練維持近似後的模型準確率。BP 則採用 Signed-Digit-4(SD4)梯度量化,以及適合硬體實作的 Softmax 與 GELU 反向運算近似方法。此外,為提升模型效能,本研究針對不同校正與範圍映射百分位數進行分析,以選擇各層最佳的百分位設定。我們利用前述量化評估指標分析近似運算後的解釋品質變化,並以量化前之結果進行正規化比較。雖然正規化後的 SIC/AIC 相較於 FP32 IDG 最多下降約 3–5%,但由於評估結果係以 FP32 為基準進行正規化,因此實際的絕對差異仍相當有限。更重要的是,在採用 SD4 梯度量化與非線性近似運算後,整體解釋品質及其變化趨勢仍能獲得良好保留。 本研究提出之加速器支援 Vision Transformer、Swin Transformer 與 BERT 等模型,可同時應用於電腦視覺與自然語言處理任務,並透過 TSRI MPW 平台採用台積電TSMC28 奈米製程完成晶片實作。本設計採用資料重用策略,以降低片上記憶體使用量與記憶體存取成本,使完整 Transformer 運算僅需 82 KB 的片上記憶體即可完成。此外,考量 TSRI ADVANTEST V93000 測試平台的限制,本研究設計 Self-Test 模式,提供晶片內部驗證機制,使晶片可於高於測試設備限制的操作頻率下驗證其功能正確性。實作完成之晶片總面積為 6.16 mm²,其中核心面積為 4.33 mm²;正常模式操作頻率為 200 MHz,Self-Test 模式則可達 400 MHz。在 Self-Test 模式下,本設計可達到 13.1 TOPS 的運算效能,能源效率為 24.95 TOPS/W,面積效率則為 2.13 TOPS/mm²。後佈局模擬結果證實,本加速器於正常模式與 Self-Test 模式下皆能正確完成功能運作。 As AI models become a critical part of the modern world in many aspects, there is a need for trustworthy and accountable AI models. However, as AI models develop, there is a trade-off between accuracy and interpretability. The traditional machine learning techniques, which are easy to interpret thanks to their white-box characteristic, have now been replaced by robust deep learning models that achieve high accuracy but are black-box. XAI aims to fill this gap by introducing a post hoc method to explain the relationship between the input and output of black-box models. One of the most successful models in recent years is Transformer, which has demonstrated robust performance and has replaced traditional deep neural networks in different applications, including image classification and natural language processing. Like many other modern deep learning models, the decision-making process of the Transformer is difficult to interpret. But unlike other models, the Transformer itself has an attention mechanism, which several researchers have proposed using to explain its decisions. However, the attention mechanism is under dispute for its relation to explanation. Therefore, another post hoc XAI is required to explain the relationship between the Transformer model's inputs and outputs. Many XAI methods have been proposed throughout the field's history. Among them, Integrated Decision Gradient (IDG) stands out for its high-quality explanations for both visual and NLP tasks. However, IDG requires multiple FP and BP computations to interpolate the baseline to the original input to increase its explainability, thereby requiring greater computational effort. Therefore, it was necessary to develop an explainable accelerator that could deliver high-performance explainability while reducing computational effort. Most of the existing Transformer accelerators are primarily designed for inference and do not efficiently support hardware-based gradient-based explanation generation, such as IDG. To address this gap, this thesis presents a hardware-efficient, explainable Transformer accelerator chip based on Integrated Decision Gradients (IDG). First, we validate the robustness of the IDG explanation across different Transformer models, including Vision Transformers and BERTs, for image and sentiment classification, respectively, and compare the quantitative results with those of other XAI methods. A quantitative explanation is difficult to evaluate and lacks a unified benchmark across all research. We sought to select several reliable metrics, informed by previous work, to provide a comprehensive evaluation approach. A quantitative comparison shows that IDG achieves better explanation results than multiple attribution methods across both vision and NLP Transformer models. Next, we consider the computing acceleration for both the Forward Pass (FP) and the Backward Pass (BP). The FP leverages previously developed ternary-weight-4-bit-activation (TW4A) Transformer models with approximate Softmax and GELU nonlinear activations, which employ QAT to maintain high accuracy after approximation. The BP employs Signed-Digit-4 (SD4) gradient quantization and hardware-friendly approximations for Softmax and GELU backward computation. In addition, to achieve better performance, multiple calibration and range-mapping percentiles are explored to determine the best percentile for each layer. We used the previous quantitative metrics to evaluate the trend after approximation and normalized them with the metric before quantization. Although the normalized SIC/AIC shows up to 3–5% degradation relative to FP32 IDG, the absolute metric difference remains small because the values are normalized against the FP32 baseline. More importantly, the overall explanation behavior and trend are still preserved under SD4 and approximate nonlinear computation. The accelerator IC proposed in this study, which supports Vision Transformer (ViT), Swin Transformer, and BERT models for both computer vision and natural language processing tasks, was implemented using the TSMC 28 nm process through the TSRI MPW platform. A data reuse strategy is employed to reduce on-chip memory usage and memory access overhead, enabling complete Transformer computation with only 82 KB of on-chip memory. The proposed accelerator, considering the TSRI ADVANTEST V93000 testing environment, includes a Self-Test mode that provides an internal validation method for testing the IC's robustness at higher frequencies than the environment's limitations allow. The implemented chip occupies a die area of 6.16 mm², with a core area of 4.33 mm², operating in normal mode at 200 MHz and in Self-Test mode at 400 MHz. In Self-Test mode, the design achieves a throughput of 13.1 TOPS, with an energy efficiency of 24.95 TOPS/W and an area efficiency of 2.13 TOPS/mm². Post-layout simulation results verify the accelerator's functional correctness in both normal and Self-Test modes. |
| URI: | http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103280 |
| DOI: | 10.6342/NTU202602785 |
| 全文授權: | 同意授權(限校園內公開) |
| 電子全文公開日期: | 2026-08-11 |
| 顯示於系所單位: | 積體電路設計與自動化學位學程 |
文件中的檔案:
| 檔案 | 大小 | 格式 | |
|---|---|---|---|
| ntu-114-2.pdf 授權僅限NTU校內IP使用(校園外請利用VPN校外連線服務) | 8.78 MB | Adobe PDF |
系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。
