Skip navigation

DSpace

機構典藏 DSpace 系統致力於保存各式數位資料(如:文字、圖片、PDF)並使其易於取用。

點此認識 DSpace
DSpace logo
English
中文
  • 瀏覽論文
    • 校院系所
    • 出版年
    • 作者
    • 標題
    • 關鍵字
    • 指導教授
  • 搜尋 TDR
  • 授權 Q&A
    • 我的頁面
    • 接受 E-mail 通知
    • 編輯個人資料
  1. NTU Theses and Dissertations Repository
  2. 共同教育中心
  3. 智慧醫療與健康資訊碩士學位學程
請用此 Handle URI 來引用此文件: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103607
標題: 遲發性運動障礙辨識之多模態機器學習框架:基於語 音與臉部動作分析
A Multimodal Machine Learning Framework for Distinguishing Tardive Dyskinesia via Voice and Facial Movement Analysis
作者: 許漢強
Dwayne Reinaldy
指導教授: 曾宇鳳
Yufeng Jane Tseng
關鍵字: 遲發性運動障礙; 自動化篩查; 語音生物標記; 多模態機器學習; 三維臉部分析; 說話者解纏; 可解釋人工智慧
Tardive Dyskinesia; Automated Screening; Speech Biomarker; Multimodal Machine Learning; 3D Facial Analysis; Speaker Disentanglement; Explainable AI
出版年 : 2026
學位: 碩士
摘要: 遲發性運動障礙(Tardive Dyskinesia, TD)是一種由藥物引起的運動障礙,估計影響20%至30%接受長期抗精神病藥物治療的患者,但在臨床實踐中仍普遍存在診斷不足的問題。TD患者的言語障礙已被臨床記錄超過三十年,但目前尚無自動化系統用於分析TD患者的言語特徵。本論文提出了一種多模態機器學習框架,透過整合來自簡短、真實場景視訊錄製中的語音和面部分析,實現TD的自動化篩檢。該框架結合了基於解耦機制的音訊模組與3D面部運動學分析模組,並透過跨模態注意力機制將兩者連接,使兩種模態能夠動態互動。我們在包含80名受試者的資料集上評估了該框架,受試者分為三類:TD、帕金森病(Parkinson's Disease)和健康對照組,並採用受試者獨立的交叉驗證方法。該多模態模型在健康對照組與TD的二分類任務中取得了0.93的AUROC值,在三分類任務中取得了0.89的AUROC值,在所有任務中均持續優於僅使用面部或僅使用音訊的基線模型。我們還採用了可解釋人工智慧(Explainable AI, XAI)方法,以提供可解釋的預測結果,將模型的注意力機制與已知TD言語障礙相關的語音片段聯繫起來。總體而言,我們的研究填補了一個持續三十餘年的診斷空白,並朝著在真實臨床環境中實現客觀TD篩檢邁出了重要一步。
Tardive Dyskinesia (TD) is a medication-induced movement disorder affecting an estimated 20-30% of patients on long-term antipsychotic treatment, yet remains widely underdiagnosed in clinical practice. Speech impairment in TD has been clinically documented for over three decades, yet no automated system has been used for analyzing speech in TD. This thesis presents a multimodal machine learning framework for automated TD screening by integrating voice and facial analysis from brief, in-the-wild video recordings. The proposed framework combines a disentanglement mechanism-based audio module with a 3D facial kinematic analysis module connected via cross-modal attention, enabling the two modalities to interact dynamically. We evaluated this framework on a dataset of 80 subjects across three classes: TD, Parkinson's Disease, and healthy controls under subject-independent cross-validation. The multimodal model achieved an AUROC of 0.93 for healthy control versus TD discrimination and 0.89 for the three-class classification task, consistently outperforming the face and audio only baseline across all tasks. We also employed Explainable AI (XAI) to provide interpretable predictions, linking model attention to speech segments consistent with known speech impairments in TD. Overall, our work addresses a diagnostic gap that has persisted for over three decades and represents a step toward objective TD screening in real-world clinical settings.
URI: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103607
DOI: 10.6342/NTU202602120
全文授權: 同意授權(全球公開)
電子全文公開日期: 2026-08-19
顯示於系所單位:智慧醫療與健康資訊碩士學位學程

文件中的檔案:
檔案 大小格式 
ntu-114-2.pdf5.17 MBAdobe PDF檢視/開啟
顯示文件完整紀錄


系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。

社群連結
聯絡資訊
10617臺北市大安區羅斯福路四段1號
No.1 Sec.4, Roosevelt Rd., Taipei, Taiwan, R.O.C. 106
Tel: (02)33662353
Email: ntuetds@ntu.edu.tw
意見箱
相關連結
館藏目錄
國內圖書館整合查詢 MetaCat
臺大學術典藏 NTU Scholars
臺大圖書館數位典藏館
本站聲明
© NTU Library All Rights Reserved