Skip navigation

DSpace

機構典藏 DSpace 系統致力於保存各式數位資料(如:文字、圖片、PDF)並使其易於取用。

點此認識 DSpace
DSpace logo
English
中文
  • 瀏覽論文
    • 校院系所
    • 出版年
    • 作者
    • 標題
    • 關鍵字
    • 指導教授
  • 搜尋 TDR
  • 授權 Q&A
    • 我的頁面
    • 接受 E-mail 通知
    • 編輯個人資料
  1. NTU Theses and Dissertations Repository
  2. 重點科技研究學院
  3. 精準健康學位學程
請用此 Handle URI 來引用此文件: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103810
標題: 基於注視引導與知識蒸餾進行具有資源效率的胸部X光報告生成
GAZE-GUIDED KNOWLEDGE DISTILLATION FOR RESOURCE-EFFICIENT CHEST X-RAY REPORT GENERATION
作者: Nandhitha Suruliandi
Nandhitha Suruliandi
指導教授: 蕭輔仁
Furen Xiao
關鍵字: 胸部X光報告生成; 醫學視覺語言模型; 知識蒸餾; 注視引導; 邊緣部署
Chest X-ray report generation; medical vision-language models; knowledge distillation; gaze guidance; edge deployment
出版年 : 2026
學位: 碩士
摘要: 儘管胸部X光報告生成技術已有長足進展,效能最佳的系統多半體積龐大、運算資源需求高,在放射科醫師人力最為短缺的地區反而最難取得。本研究旨在探討能否訓練一個輕量模型,使其在較大模型的監督下學習,並在不需專用圖形處理器的情況下產生具臨床價值的報告,同時釐清此種臨床價值應如何衡量。
為此,本論文提出一套以師生架構為基礎的注視引導知識蒸餾框架。教師模型採用具四十億參數的醫學視覺語言模型 MedGemma。其內部特徵事先由胸部X光影像萃取並儲存,因此教師模型本身不更新參數,於學生模型訓練過程中亦無需載入。學生模型參數量約五千萬,並非教師模型的縮減複本,而是一個獨立訓練、學習以與教師相同方式編碼影像的網路。除影像特徵之外,學生模型同時學習放射科醫師閱片時所注視的區域(由眼動追蹤記錄取得)以及參考報告內容。訓練完成後,學生模型可在不依賴教師模型的情況下獨立產生報告,其規模適用於硬體受限的部署環境。
實驗結果顯示,學生模型雖遠小於教師模型,但在本領域常用的文字重疊指標(BLEU、METEOR、ROUGE-L 與 CIDEr)以及臨床準確度上皆優於其教師模型。其中注視引導為關鍵因素,納入注視引導後臨床準確度約提升一倍。本研究並指出,僅以自然語言生成指標並不足以判斷所產生的報告是否具臨床可靠性,因此另採用逐項發現層級的指標作為互補評估。
Although the generation of chest X-ray reports has come a long way, the most sophisticated systems are large, resource intensive and least accessible in the regions where there are few radiologists. The goal of this work is to evaluate the feasibility of developing a compact model which can be trained under the supervision of a much larger model and produce clinically useful reports without the need for a dedicated GPU, and to determine how that clinical usefulness should be measured. In this respect, this thesis introduces a gaze-guided knowledge distillation framework based on the teacher-student design. MedGemma is a 4 billion parameter medical vision-language model that serves as the teacher. Its internal features are extracted from the chest X-rays once in advance and stored, so the teacher never gets updated and is not required during student training. The student is a smaller network of the order of 50 million parameters; the student isn't a reduced copy of the teacher, but a separate model trained to encode each image in the same manner as the teacher. It also trains on the areas a radiologist would examine when reading a chest X-ray, as recorded by eye tracking, and on the reference reports. Once trained, the student writes reports without the teacher, at a size suitable for hardware-limited situations. Although it is much smaller, the student is better than its own teacher on all the standard text-overlap metrics in this field (BLEU, METEOR, ROUGE-L and CIDEr), and on clinical accuracy. Experimental results show that gaze is the key, with clinical accuracy doubling when gaze guidance is included. It is also reported that NLG metrics alone are not sufficient for judging whether a generated report is clinically reliable, and hence per-finding metrics are used in this work as a complementary measure.
URI: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103810
DOI: 10.6342/NTU202603739
全文授權: 同意授權(全球公開)
電子全文公開日期: 2026-08-20
顯示於系所單位:精準健康學位學程

文件中的檔案:
檔案 大小格式 
ntu-114-2.pdf6.63 MBAdobe PDF檢視/開啟
顯示文件完整紀錄


系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。

社群連結
聯絡資訊
10617臺北市大安區羅斯福路四段1號
No.1 Sec.4, Roosevelt Rd., Taipei, Taiwan, R.O.C. 106
Tel: (02)33662353
Email: ntuetds@ntu.edu.tw
意見箱
相關連結
館藏目錄
國內圖書館整合查詢 MetaCat
臺大學術典藏 NTU Scholars
臺大圖書館數位典藏館
本站聲明
© NTU Library All Rights Reserved