Skip navigation

DSpace

機構典藏 DSpace 系統致力於保存各式數位資料(如:文字、圖片、PDF)並使其易於取用。

點此認識 DSpace
DSpace logo
English
中文
  • 瀏覽論文
    • 校院系所
    • 出版年
    • 作者
    • 標題
    • 關鍵字
    • 指導教授
  • 搜尋 TDR
  • 授權 Q&A
    • 我的頁面
    • 接受 E-mail 通知
    • 編輯個人資料
  1. NTU Theses and Dissertations Repository
  2. 公共衛生學院
  3. 健康數據拓析統計研究所
請用此 Handle URI 來引用此文件: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/104645
標題: 大腸癌腺瘤及癌症進展之機器學習模式分析
Progression of Colorectal Adenoma and Cancer: A Machine Learning Analysis
作者: 簡瑞伶
Ruei-Ling Chien
指導教授: 陳秀熙
Hsiu-Hsi Chen
關鍵字: 大腸直腸癌; 疾病自然史; 連續時間隱藏狀態馬可夫模型; 生成矩陣; 機器學習; 隨機森林; SHAP; 個人化疾病進展; 精準篩檢
Colorectal cancer; disease natural history; continuous-time hidden-state Markov model; generator matrix; machine learning; Random Forest; SHAP; personalized disease progression; precision screening
出版年 : 2026
學位: 碩士
摘要: 摘要
研究背景
大腸直腸癌 (colorectal cancer, CRC) 腺瘤-癌症發生序列 (adenoma–carcinoma sequence) 由正常黏膜歷經小腺瘤、進展型腺瘤、臨床症前可偵測期 (preclinical detectable phase, PCDP) 逐步演變為臨床期癌症,疾病進展屬連續且不可直接觀測之潛在過程。傳統疾病自然史模型多以族群平均 (population average) 描述疾病進展,難以反映個體間之疾病異質性;而機器學習方法雖廣泛應用於風險預測,卻多侷限於疾病分類,缺乏與疾病自然史之動態連結。本研究旨在整合群體層級疾病自然史推論與個人化風險預測,建立一套混合式疾病進展架構 (Hybrid Disease Progression Framework)。

研究方法
群體層級分析採用臺灣全國糞便免疫化學檢查 (fecal immunochemical test, FIT) 篩檢資料 (14,478,478 人次篩檢紀錄、5,640,906 名受檢者),建立含九個潛在疾病狀態之連續時間隱藏狀態馬可夫模型 (continuous-time hidden-state Markov model),並透過觀測映射機制 (observation mapping) 結合 FIT 篩檢敏感度連結潛在狀態與觀測結果,以最大概似估計法 (MLE) 推估生成矩陣 (generator matrix) 之各轉移強度,再以矩陣指數法模擬潛在疾病分布。個體層級分析則以彰化地區社區篩檢資料建立「正常 vs. 腺瘤」與「腺瘤 vs. 篩檢偵測癌」兩項二元分類模型,比較人工神經網路 (ANN) 與隨機森林 (Random Forest, RF),並以 SHapley Additive exPlanations (SHAP) 辨識重要危險因子。最後將機器學習預測機率轉換為相對風險倍數,回饋調節群體生成矩陣以建立個人化生成矩陣 (personalized generator matrix) ,推估個人化疾病進展光譜。

研究結果
生成矩陣估計顯示,正常進入腺瘤之轉移強度 (q₁₂ = 0.00112) 遠低於後續各階段之進展速度 (如 q₂₃ = 0.5896、q₃₄ = 0.5889),且觀測結果與真實疾病狀態並非一對一對應,FIT 陰性個案之間隔癌反映篩檢敏感度之限制。在機器學習分類中,「正常 vs. 腺瘤」以 RF 表現最佳 (測試集 AUC = 0.935),優於 ANN (AUC = 0.896)。然而,「腺瘤 vs. 篩檢偵測癌」因兩階段特徵重疊、分類難度較高,測試集 AUC 分別降至 0.716(ANN)與 0.694(RF)。SHAP 分析一致指出誘導時間(induction time)、糞便血紅素濃度 (f-Hb)、年齡及代謝症候群相關指標為最重要之預測因子。個人化進展模擬顯示,q₁₂ (是否容易形成腺瘤)對整體癌症風險之影響遠大於 q₃₄ (腺瘤惡化速度):高 q₁₂/低q₃₄ 個案之十年臨床期癌症風險 (2.46%) 明顯高於低 q₁₂/高 q₃₄ 個案,顯示「正常→腺瘤」階段為整體風險之上游主閘門。

結論
本研究首次整合九狀態連續時間隱藏狀態疾病自然史模型與人工智慧個人化風險預測,建立由群體疾病進展延伸至個人化疾病進展之分析架構,使疾病自然史研究由「描述平均進展」提升至「推估個人進展」。此一兼具機制解釋能力與個人化預測能力之架構,可作為風險導向篩檢 (risk-adaptive screening)、個人化篩檢間隔設計與臨床共享決策之理論基礎,並有望成為 AI 數位雙生 (AI-enabled Digital Twin)及精準公共衛生之方法學平台。辨識「容易形成腺瘤」之高風險個體較辨識「腺瘤惡化速度快」者更能於上游有效攔截大腸直腸癌之發生。
Abstract
Background
Colorectal cancer (CRC) follows the adenoma–carcinoma sequence, progressing from normal mucosa through small and advanced adenomas and the preclinical detectable phase (PCDP) to clinical cancer. This progression is a continuous and latent process that cannot be directly observed. Conventional disease natural history models characterize progression at the population-average level and fail to capture individual heterogeneity, whereas machine learning approaches, though widely applied to risk prediction, are largely confined to disease classification and lack a dynamic link to the disease natural history. This study aimed to integrate population-level disease natural history inference with personalized risk prediction into a Hybrid Disease Progression Framework.

Methods
At the population level, nationwide fecal immunochemical test (FIT) screening data from Taiwan (14,478,478 screening records; 5,640,906 participants) were used to construct a nine-state continuous-time hidden-state Markov model. An observation mapping mechanism incorporating FIT screening sensitivity linked latent states to observed screening outcomes, and the transition intensities of the generator matrix were estimated by maximum likelihood estimation (MLE) , with the latent disease distribution simulated via the matrix exponential. At the individual level, community screening data from Changhua were used to build two binary classifiers—Normal vs. Adenoma and Adenoma vs. Screen-detected CRC—comparing an artificial neural network (ANN) and a Random Forest (RF), with SHapley Additive exPlanations (SHAP) applied to identify important risk factors. Finally, RF-predicted probabilities were converted into relative risk multipliers to modulate the population generator matrix, yielding a personalized generator matrix and a personalized disease progression spectrum.

Results
The estimated generator matrix showed that the transition from normal to small adenoma (q₁₂ = 0.00112) was markedly slower than subsequent progression steps (e.g., q₂₃ = 0.5896, q₃₄ = 0.5889), and observed outcomes did not correspond one-to-one with true disease states—interval cancers among FIT-negative cases reflected the limits of screening sensitivity. For the Normal vs. Adenoma classification, RF performed best (test AUC = 0.935), outperforming ANN (AUC = 0.896). The Adenoma vs. Screen-detected CRC task was more challenging owing to overlapping features, yielding test AUCs of 0.716 and 0.694 for ANN and RF, respectively. SHAP analyses consistently identified induction time, fecal hemoglobin concentration (f-Hb), age, and metabolic-syndrome-related indicators as the most important predictors. Personalized simulations indicated that the propensity to form an adenoma (q₁₂) influenced overall cancer risk far more than the speed of adenoma malignant progression (q₃₄): a high-q₁₂/low-q₃₄ case carried a 10-year clinical-cancer risk of 2.46%, substantially higher than a low-q₁₂/high-q₃₄ case, identifying the Normal→Adenoma step as the upstream gatekeeper of overall risk.

Conclusion
This study is the first to integrate a nine-state continuous-time hidden-state disease natural history model with AI-based personalized risk prediction, establishing an analytical framework that extends population-level progression to personalized disease progression and elevates natural history research from describing average progression to estimating individual progression. Combining mechanistic interpretability with personalized prediction, the framework can serve as a theoretical basis for risk-adaptive screening, personalized screening-interval design, and shared clinical decision-making, and as a methodological platform for AI-enabled Digital Twins and precision public health. Identifying individuals prone to forming adenomas appears more effective for intercepting CRC upstream than identifying those with fast-progressing adenomas. Limitations include reliance on single-region data for the personalized model and the static nature of the generator matrix; external validation and integration of longitudinal and multi-omics data are warranted.
URI: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/104645
DOI: 10.6342/NTU202602503
全文授權: 同意授權(全球公開)
電子全文公開日期: 2026-08-29
顯示於系所單位:健康數據拓析統計研究所

文件中的檔案:
檔案 描述 大小格式 
ntu-114-2.pdf3.91 MBAdobe PDF檢視/開啟
ntu-114-2.pdf3.91 MBAdobe PDF檢視/開啟
顯示文件完整紀錄


系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。

社群連結
聯絡資訊
10617臺北市大安區羅斯福路四段1號
No.1 Sec.4, Roosevelt Rd., Taipei, Taiwan, R.O.C. 106
Tel: (02)33662353
Email: ntuetds@ntu.edu.tw
意見箱
相關連結
館藏目錄
國內圖書館整合查詢 MetaCat
臺大學術典藏 NTU Scholars
臺大圖書館數位典藏館
本站聲明
© NTU Library All Rights Reserved