Skip navigation

DSpace

機構典藏 DSpace 系統致力於保存各式數位資料(如:文字、圖片、PDF)並使其易於取用。

點此認識 DSpace
DSpace logo
English
中文
  • 瀏覽論文
    • 校院系所
    • 出版年
    • 作者
    • 標題
    • 關鍵字
    • 指導教授
  • 搜尋 TDR
  • 授權 Q&A
    • 我的頁面
    • 接受 E-mail 通知
    • 編輯個人資料
  1. NTU Theses and Dissertations Repository
  2. 重點科技研究學院
  3. 精準健康學位學程
請用此 Handle URI 來引用此文件: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103810
完整後設資料紀錄
DC 欄位值語言
dc.contributor.advisor蕭輔仁zh_TW
dc.contributor.advisorFuren Xiaoen
dc.contributor.authorNandhitha Suruliandizh_TW
dc.contributor.authorNandhitha Suruliandien
dc.date.accessioned2026-08-19T16:48:51Z-
dc.date.available2026-08-20-
dc.date.copyright2026-08-19-
dc.date.issued2026-
dc.date.submitted2026-08-11 15:29:34-
dc.identifier.citationReference
[1] P. Singh and S. Singh, "ChestX-Transcribe: a multimodal transformer for automated radiology report generation from chest x-rays," Frontiers in Digital Health, vol. 7, 1535168, 2025.
[2] M. Y. Ouis and M. A. Akhloufi, "Deep learning for report generation on chest X-ray images," Computerized Medical Imaging and Graphics, vol. 111, 102320, 2024.
[3] A. M. Fahmy and M. M. El Mashaad, "Automated X-Ray Chest Report Generation," in 2025 Twelfth International Conference on Intelligent Computing and Information Systems (ICICIS), 2025, pp. 651-658.
[4] K. Konstantinidis, "The shortage of radiographers: A global crisis in healthcare," Journal of Medical Imaging and Radiation Sciences, vol. 55, no. 4, 101333, 2024.
[5] T. Wang, J. Guo, B. Zhang, G. Yang, and D. Li, "Deploying AI on edge: Advancement and challenges in edge intelligence," Mathematics, vol. 13, no. 11, 1878, 2025.
[6] L. Huang, Y. Cao, X. Zhao, C. Li, and J. Tang, "Medical report generation via knowledge distill and medical keywords," Neurocomputing, 133823, 2026.
[7] A. Sellergren et al., "MedGemma technical report," arXiv preprint arXiv:2507.05201, 2025.
[8] D. Khatri and S. T. P. Gupta, "Towards comprehensive benchmarking of medical vision language models," Briefings in Bioinformatics, vol. 26, no. Supplement_1, pp. i24-i25, 2025.
[9] B. C. Kalpélbé, A. G. Adaambiik, and W. Peng, "Vision language models in medicine," arXiv preprint arXiv:2503.01863, 2025.
[10] A. Moslemi, A. Briskina, Z. Dang, and J. Li, "A survey on knowledge distillation: Recent advancements," Machine Learning with Applications, vol. 18, 100605, 2024.
[11] G. Hinton, O. Vinyals, and J. Dean, "Distilling the knowledge in a neural network," arXiv preprint arXiv:1503.02531, 2015.
[12] A. Karargyris et al., "Creation and validation of a chest X-ray dataset with eye-tracking and report dictation for AI development," Scientific Data, vol. 8, no. 1, 92, 2021.
[13] T.-T. Pham, A. Nguyen, Z. Deng, C. C. Wu, H. Nguyen, and N. Le, "Interpreting Radiologist's Intention from Eye Movements in Chest X-ray Diagnosis," in Proceedings of the 33rd ACM International Conference on Multimedia, 2025, pp. 257-266.
[14] G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, "Densely Connected Convolutional Networks," in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2017, pp. 4700-4708.
[15] A. Vaswani et al., "Attention Is All You Need," in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017, pp. 5998-6008.
[16] A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio, "FitNets: Hints for Thin Deep Nets," in Proc. Int. Conf. Learning Representations (ICLR), 2015.
[17] H.-C. Shin, K. Roberts, L. Lu, D. Demner-Fushman, J. Yao, and R. M. Summers, "Learning to read chest x-rays: Recurrent neural cascade model for automated image annotation," in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2016, pp. 2497-2506.
[18] X. Wang, Y. Peng, L. Lu, Z. Lu, and R. M. Summers, "TieNet: Text-image embedding network for common thorax disease classification and reporting in chest x-rays," in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2018, pp. 9049-9058.
[19] B. Jing, P. Xie, and E. Xing, "On the automatic generation of medical imaging reports," in Proc. 56th Annu. Meeting Assoc. for Computational Linguistics (ACL), 2018, pp. 2577-2586.
[20] C. Yin, B. Qian, J. Wei, X. Li, X. Zhang, Y. Li, and Q. Zheng, "Automatic generation of medical imaging diagnostic report with hierarchical recurrent neural network," in Proc. IEEE Int. Conf. Data Mining (ICDM), 2019, pp. 728-737.
[21] G. Liu, T.-M. H. Hsu, M. McDermott, W. Boag, W.-H. Weng, P. Szolovits, and M. Ghassemi, "Clinically accurate chest x-ray report generation," in Proc. Machine Learning for Healthcare Conf. (MLHC), 2019, pp. 249-269.
[22] Z. Chen, Y. Song, T.-H. Chang, and X. Wan, "Generating radiology reports via memory-driven transformer," in Proc. 2020 Conf. Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 1439-1449.
[23] N. Aksoy, N. Ravikumar, and A. F. Frangi, "Radiology report generation using transformers conditioned with non-imaging data," in Medical Imaging 2023: Imaging Informatics for Healthcare, Research, and Applications, vol. 12469, SPIE, 2023, pp. 146-153.
[24] Z. Wang, L. Liu, L. Wang, and L. Zhou, "METransformer: Radiology report generation by transformer with multiple learnable expert tokens," in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2023, pp. 11558-11567.
[25] I. Sîrbu, I.-R. Sîrbu, J. Bogojeska, and T. Rebedea, "GIT-CXR: End-to-end transformer for chest X-ray report generation," Information, vol. 16, no. 7, 524, 2025.
[26] F. Liu, X. Wu, S. Ge, W. Fan, and Y. Zou, "Exploring and distilling posterior and prior knowledge for radiology report generation," in Proc. IEEE/CVF Conf. Computer Vision and Pattern Recognition (CVPR), 2021, pp. 13753-13762.
[27] C. Pellegrini, E. Özsoy, B. Busam, N. Navab, and M. Keicher, "RaDialog: A large vision-language model for radiology report generation and conversational assistance," arXiv preprint arXiv:2311.18681, 2023.
[28] D. Edirisinghe, W. Nimalsiri, M. Hennayake, D. Meedeniya, and G. Lim, "Chest X-ray report generation using abnormality guided vision language model," IEEE Access, 2025.
[29] Z. Wang, S. Yan, K. Yin, X. Zhang, and W. K. Cheung, "CURV: Coherent uncertainty-aware reasoning in vision-language models for X-ray report generation," Advances in Neural Information Processing Systems, vol. 38, pp. 93686-93713, 2025.
[30] S. Wang, Z. Yan, D. Zhang, H. Wei, Z. Li, and R. Li, "Prototype knowledge distillation for medical segmentation with missing modality," in Proc. IEEE Int. Conf. Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1-5.
[31] J. Chen, X. Huang, M. Jiang, Y. Li, Z. Zou, and D. Qian, "Graph-driven medical report generation with adaptive knowledge distillation," Applied Sciences, vol. 15, no. 20, 10974, 2025.
[32] H. Tsaniya, C. Fatichah, N. Suciati, T. Obi, and J. Lee, "Medical report generation with knowledge distillation and multi-stage hierarchical attention in vision transformer encoder and GPT-2 decoder," IEEE Access, 2025.
[33] J. Duan, M. Zhang, M. Song, X. Xu, and H. Lu, "Eye tracking-enhanced deep learning for medical image analysis: A systematic review on data efficiency, interpretability, and multimodal integration," Bioengineering, vol. 12, no. 9, 954, 2025.
[34] J. Neves, C. Hsieh, I. B. Nobre, S. C. Sousa, C. Ouyang, A. Maciel, A. Duchowski, J. Jorge, and C. Moreira, "Shedding light on AI in radiology: A systematic review and taxonomy of eye gaze-driven interpretability in deep learning," European Journal of Radiology, vol. 172, 111341, 2024.
[35] M. Bhattacharya, G. Singh, S. Jain, and P. Prasanna, "GazeLT: Visual attention-guided long-tailed disease classification in chest radiographs," arXiv preprint arXiv:2508.09478, 2025.
[36] B. Liu, G. Li, Y. Zou, R. Zhang, X. Zhou, H. Zhang, J. Lv, G. Luo, and D. Ji, "Towards clinician-like reasoning: Stage-aware gaze-guided heterogeneous network for medical image recognition," Displays, 103550, 2026.
[37] C. Ma et al., "Eye-gaze guided multi-modal alignment for medical representation learning," Advances in Neural Information Processing Systems, vol. 37, pp. 6126-6153, 2024.
[38] Z. Zhang, K. Qian, B. Zhou, F. Fang, and X. Ma, "Gaze-assisted visual grounding via knowledge distillation for referred object grasping with under-specified object referring," Engineering Applications of Artificial Intelligence, vol. 133, 108493, 2024.
[39] D. Demner-Fushman, M. D. Kohli, M. B. Rosenman, S. E. Shooshan, L. Rodriguez, S. Antani, G. R. Thoma, and C. J. McDonald, "Preparing a collection of radiology examinations for distribution and retrieval," Journal of the American Medical Informatics Association, vol. 23, no. 2, pp. 304-310, 2016.
[40] K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, "BLEU: A method for automatic evaluation of machine translation," in Proc. 40th Annu. Meeting Assoc. Computational Linguistics (ACL), 2002, pp. 311-318.
[41] S. Banerjee and A. Lavie, "METEOR: An automatic metric for MT evaluation with improved correlation with human judgments," in Proc. ACL Workshop Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, 2005, pp. 65-72.
[42] C.-Y. Lin, "ROUGE: A package for automatic evaluation of summaries," in Text Summarization Branches Out, 2004, pp. 74-81.
[43] R. Vedantam, C. L. Zitnick, and D. Parikh, "CIDEr: Consensus-based image description evaluation," in Proc. IEEE Conf. Computer Vision and Pattern Recognition (CVPR), 2015, pp. 4566-4575.
[44] Z. Chen, Y. Shen, Y. Song, and X. Wan, "Cross-modal memory networks for radiology report generation," in Proc. 59th Annu. Meeting Assoc. Comput. Linguistics and 11th Int. Joint Conf. Natural Language Processing (Vol. 1: Long Papers), 2021, pp. 5904–5914.
[45] Z. Wang, L. Liu, L. Wang, and L. Zhou, "R2GenGPT: Radiology report generation with frozen LLMs," Meta-Radiology, vol. 1, no. 3, art. no. 100033, 2023.
[46] P. Peng et al., "Eye gaze guided cross-modal alignment network for radiology report generation," IEEE J. Biomed. Health Informat., vol. 28, no. 12, pp. 7406–7419, 2024.
[47] A. Konwer, M. Bhattacharya, and P. Prasanna, "Gaze2Report: Radiology report generation via visual-gaze prompt tuning of LLMs," in Proc. IEEE 23rd Int. Symp. Biomedical Imaging (ISBI), 2026, pp. 1–5.
[48] Z. Wang, M. Tang, L. Wang, X. Li, and L. Zhou, "A medical semantic-assisted transformer for radiographic report generation," in Proc. Int. Conf. Medical Image Computing and Computer-Assisted Intervention (MICCAI), 2022, pp. 655–664.
-
dc.identifier.urihttp://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103810-
dc.description.abstract儘管胸部X光報告生成技術已有長足進展,效能最佳的系統多半體積龐大、運算資源需求高,在放射科醫師人力最為短缺的地區反而最難取得。本研究旨在探討能否訓練一個輕量模型,使其在較大模型的監督下學習,並在不需專用圖形處理器的情況下產生具臨床價值的報告,同時釐清此種臨床價值應如何衡量。
為此,本論文提出一套以師生架構為基礎的注視引導知識蒸餾框架。教師模型採用具四十億參數的醫學視覺語言模型 MedGemma。其內部特徵事先由胸部X光影像萃取並儲存,因此教師模型本身不更新參數,於學生模型訓練過程中亦無需載入。學生模型參數量約五千萬,並非教師模型的縮減複本,而是一個獨立訓練、學習以與教師相同方式編碼影像的網路。除影像特徵之外,學生模型同時學習放射科醫師閱片時所注視的區域(由眼動追蹤記錄取得)以及參考報告內容。訓練完成後,學生模型可在不依賴教師模型的情況下獨立產生報告,其規模適用於硬體受限的部署環境。
實驗結果顯示,學生模型雖遠小於教師模型,但在本領域常用的文字重疊指標(BLEU、METEOR、ROUGE-L 與 CIDEr)以及臨床準確度上皆優於其教師模型。其中注視引導為關鍵因素,納入注視引導後臨床準確度約提升一倍。本研究並指出,僅以自然語言生成指標並不足以判斷所產生的報告是否具臨床可靠性,因此另採用逐項發現層級的指標作為互補評估。
zh_TW
dc.description.abstractAlthough the generation of chest X-ray reports has come a long way, the most sophisticated systems are large, resource intensive and least accessible in the regions where there are few radiologists. The goal of this work is to evaluate the feasibility of developing a compact model which can be trained under the supervision of a much larger model and produce clinically useful reports without the need for a dedicated GPU, and to determine how that clinical usefulness should be measured. In this respect, this thesis introduces a gaze-guided knowledge distillation framework based on the teacher-student design. MedGemma is a 4 billion parameter medical vision-language model that serves as the teacher. Its internal features are extracted from the chest X-rays once in advance and stored, so the teacher never gets updated and is not required during student training. The student is a smaller network of the order of 50 million parameters; the student isn't a reduced copy of the teacher, but a separate model trained to encode each image in the same manner as the teacher. It also trains on the areas a radiologist would examine when reading a chest X-ray, as recorded by eye tracking, and on the reference reports. Once trained, the student writes reports without the teacher, at a size suitable for hardware-limited situations. Although it is much smaller, the student is better than its own teacher on all the standard text-overlap metrics in this field (BLEU, METEOR, ROUGE-L and CIDEr), and on clinical accuracy. Experimental results show that gaze is the key, with clinical accuracy doubling when gaze guidance is included. It is also reported that NLG metrics alone are not sufficient for judging whether a generated report is clinically reliable, and hence per-finding metrics are used in this work as a complementary measure.en
dc.description.provenanceSubmitted by admin ntu (admin@lib.ntu.edu.tw) on 2026-08-19T16:48:51Z
No. of bitstreams: 0
en
dc.description.provenanceMade available in DSpace on 2026-08-19T16:48:51Z (GMT). No. of bitstreams: 0en
dc.description.tableofcontentsCONTENTS
ACKNOWLEDGEMENTS i
中文摘要 ii
ABSTRACT iii
CONTENTS iv
LIST OF FIGURES viii
LIST OF TABLES ix
ABBREVIATIONS AND ACRONYMS x
NOTATIONS xii
Chapter 1 Introduction 1
1.1 Background and Motivation 1
1.2 Problem Statement 2
1.3 Research Objectives 3
1.4 Proposed Solution 4
Chapter 2 Theoretical Background 7
2.1 Chest X-Ray Imaging and Radiology Reporting 7
2.1.1 Chest Radiography 7
2.1.2 Radiology Reports 8
2.1.3 Radiologist Visual Search and Gaze 9
2.2 Deep Learning Foundations 11
2.2.1 Convolutional Neural Networks and DenseNet 11
2.2.2 Transformer Architecture and Attention 12
2.2.3 Vision-Language Models 14
2.2.4 Knowledge Distillation 15
Chapter 3 Literature Review 17
3.1 Chest X-Ray Report Generation 17
3.1.1 Early CNN-RNN Approaches 18
3.1.2 Transformer-Based Approaches 20
3.1.3 Recent Vision-Language Model Approaches 22
3.2 Knowledge Distillation for Medical Imaging and Report Generation 24
3.2.1 Knowledge Distillation in Medical Imaging 24
3.2.2 Knowledge Distillation for Report Generation 25
3.3 Gaze and Eye-Tracking in Medical AI 26
3.3.1 Radiologist Gaze as a Source of Expert Knowledge 27
3.3.2 Gaze-Guided Deep Learning Models 28
3.4 Research Gap 29
Chapter 4 Methodology 31
4.1 Dataset 33
4.1.1 IU Chest X-Ray Dataset 33
4.1.2 Radiologist Gaze Data 35
4.2 Data Preprocessing 38
4.2.1 Image Preprocessing 38
4.2.2 Report Preprocessing 38
4.3 Model Architecture 39
4.3.1 Spatial Image Encoder 41
4.3.2 Auxiliary Clinical Heads and Measurement Encoder 41
4.3.3 Teacher Projection Modules 42
4.3.4 Report Decoder 42
4.4 Knowledge Distillation from MedGemma 44
4.4.1 Selection of Teacher Model 45
4.4.2 Teacher Feature Extraction 46
4.4.3 Distillation Objectives 46
4.5 Gaze-Guided Spatial Supervision 47
4.5.1 Forming the Gaze-Derived Heat Map 47
4.5.2 Applying the Gaze-Derived Heat Map 50
Chapter 5 Experimental Setup and Results 53
5.1 Training Strategy 53
5.1.1 Training Objective to Minimize the Loss Function 53
5.1.2 Optimization Setup 55
5.1.3 Training Workflow 55
5.1.4 Inference and Report Generation 56
5.2 Evaluation Metrics 56
5.2.1 Natural Language Generation Metrics 57
5.2.2 Clinical Metrics 58
5.3 Results 60
5.3.1 Baselines and Comparative Methods 61
5.3.2 Quantitative Analysis 62
5.3.2.1 Comparison with State-of-the-Art Methods 62
5.3.2.2 Teacher and Student Model Comparison 64
5.3.2.3 Ablation Study 65
5.3.2.4 Per-Finding Clinical Analysis 68
5.3.2.5 Hallucination and Report Diversity 74
5.3.2.6 Report Vocabulary and Finding Coverage 77
5.3.2.7 Model Efficiency 80
5.3.3 Qualitative Analysis 81
Chapter 6 Discussion 86
6.1 Quantitative Performance and Model Efficiency 86
6.2 The Contribution of Gaze Guidance and Distillation 87
6.3 Clinical Accuracy at the Level of Individual Findings 88
6.4 Report Diversity and Expressive Range 89
6.5 Qualitative Assessment and the Limits of Aggregate Metrics 90
6.6 Configurations Explored and Rejected 91
Chapter 7 Limitations 95
7.1 Dataset and labels 95
7.2 Evaluation methodology 96
7.3 Model behaviour 97
7.4 Inference conditions 98
Chapter 8 Conclusion 99
8.1 Summary of Findings 99
8.2 Future Directions 100
8.3 Concluding Remarks 102
Reference 103
-
dc.language.isoen-
dc.subject胸部X光報告生成-
dc.subject醫學視覺語言模型-
dc.subject知識蒸餾-
dc.subject注視引導-
dc.subject邊緣部署-
dc.subjectChest X-ray report generation-
dc.subjectmedical vision-language models-
dc.subjectknowledge distillation-
dc.subjectgaze guidance-
dc.subjectedge deployment-
dc.title基於注視引導與知識蒸餾進行具有資源效率的胸部X光報告生成zh_TW
dc.titleGAZE-GUIDED KNOWLEDGE DISTILLATION FOR RESOURCE-EFFICIENT CHEST X-RAY REPORT GENERATIONen
dc.typeThesis-
dc.date.schoolyear114-2-
dc.description.degree碩士-
dc.contributor.oralexamcommittee廖俊智;黃國彥zh_TW
dc.contributor.oralexamcommitteeJun-Zhi Liao;Kuo-Yen Huangen
dc.subject.keyword胸部X光報告生成; 醫學視覺語言模型; 知識蒸餾; 注視引導; 邊緣部署zh_TW
dc.subject.keywordChest X-ray report generation; medical vision-language models; knowledge distillation; gaze guidance; edge deploymenten
dc.relation.page109-
dc.identifier.doi10.6342/NTU202603739-
dc.rights.note同意授權(全球公開)-
dc.date.accepted2026-08-13-
dc.contributor.author-college重點科技研究學院-
dc.contributor.author-dept精準健康學位學程-
dc.date.embargo-lift2026-08-20-
顯示於系所單位:精準健康學位學程

文件中的檔案:
檔案 大小格式 
ntu-114-2.pdf6.63 MBAdobe PDF檢視/開啟
顯示文件簡單紀錄


系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。

社群連結
聯絡資訊
10617臺北市大安區羅斯福路四段1號
No.1 Sec.4, Roosevelt Rd., Taipei, Taiwan, R.O.C. 106
Tel: (02)33662353
Email: ntuetds@ntu.edu.tw
意見箱
相關連結
館藏目錄
國內圖書館整合查詢 MetaCat
臺大學術典藏 NTU Scholars
臺大圖書館數位典藏館
本站聲明
© NTU Library All Rights Reserved