Skip navigation

DSpace

機構典藏 DSpace 系統致力於保存各式數位資料(如:文字、圖片、PDF)並使其易於取用。

點此認識 DSpace
DSpace logo
English
中文
  • 瀏覽論文
    • 校院系所
    • 出版年
    • 作者
    • 標題
    • 關鍵字
    • 指導教授
  • 搜尋 TDR
  • 授權 Q&A
    • 我的頁面
    • 接受 E-mail 通知
    • 編輯個人資料
  1. NTU Theses and Dissertations Repository
  2. 公共衛生學院
  3. 健康數據拓析統計研究所
請用此 Handle URI 來引用此文件: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/104645
完整後設資料紀錄
DC 欄位值語言
dc.contributor.advisor陳秀熙zh_TW
dc.contributor.advisorHsiu-Hsi Chenen
dc.contributor.author簡瑞伶zh_TW
dc.contributor.authorRuei-Ling Chienen
dc.date.accessioned2026-08-28T16:54:06Z-
dc.date.available2026-08-29-
dc.date.copyright2026-08-28-
dc.date.issued2026-
dc.date.submitted2026-07-27 00:00:00-
dc.identifier.citation1.Fearon, E.R. and B. Vogelstein, A genetic model for colorectal tumorigenesis. Cell, 1990. 61(5): p. 759–67.
2.Feinleib, M. and M. Zelen, Some pitfalls in the evaluation of screening programs. Arch Environ Health, 1969. 19(3): p. 412–5.
3.Cole, P. and A.S. Morrison, Basic issues in population screening for cancer. J Natl Cancer Inst, 1980. 64(5): p. 1263–72.
4.Chen, T.H.-H., et al., Evaluation of a selective screening for colorectal carcinoma. Cancer, 1999. 86(7): p. 1116–1128.
5.Chiu, Y.-H., et al., Health information system for community-based multiple screening in Keelung, Taiwan (Keelung Community-based Integrated Screening No. 3). International Journal of Medical Informatics, 2006. 75(5): p. 369–383.
6.Yang, K.C., et al., Colorectal cancer screening with faecal occult blood test within a multiple disease screening programme: an experience from Keelung, Taiwan. J Med Screen, 2006. 13 Suppl 1: p. S8–13.
7.Chiu, H.-M., et al., Long-Term Effectiveness Associated With Fecal Immunochemical Testing for Early-Age Screening. JAMA Oncology, 2025. 11(8): p. 846–854.
8.Hsu, W.-F., et al., Classifying interval cancers as false negatives or newly occurring in fecal immunochemical testing. Journal of Medical Screening, 2021. 28(3): p. 286–294.
9.Chiu, H.M., et al., Long-term effectiveness of faecal immunochemical test screening for proximal and distal colorectal cancers. Gut, 2021. 70(12): p. 2321–2329.
10.Yen, A.M.-F. and H.-H. Chen, Modeling the overdetection of screen-identified cancers in population-based cancer screening with the Coxian phase-type Markov process. Statistics in Medicine, 2020. 39(5): p. 660–673.
11.Lin, T.-Y., et al., Assessing overdiagnosis of fecal immunological test screening for colorectal cancer with a digital twin approach. npj Digital Medicine, 2023. 6(1): p. 24.
12.Yen, A.M.-F. and H.-H. Chen, Bayesian measurement-error-driven hidden Markov regression model for calibrating the effect of covariates on multistate outcomes: Application to androgenetic alopecia. Statistics in Medicine, 2018. 37(21): p. 3125–3146.
13.Johnson, C.M., et al., Meta-analyses of colorectal cancer risk factors. Cancer Causes Control, 2013. 24(6): p. 1207–22.
14.Yen, A.M.-F., et al., Precision Colorectal Cancer Fecal Immunological Test Screening With Fecal-Hemoglobin-Concentration–Guided Interscreening Intervals. JAMA Oncology, 2024. 10(6): p. 765–772.
15.Chen, L.S., et al., Baseline faecal occult blood concentration as a predictor of incident colorectal neoplasia: longitudinal follow-up of a Taiwanese population-based colorectal cancer screening cohort. Lancet Oncol, 2011. 12(6): p. 551–8.
16.Rajkomar, A., J. Dean, and I. Kohane, Machine Learning in Medicine. N Engl J Med, 2019. 380(14): p. 1347–1358.
17.Nartowt, B.J., et al., Robust Machine Learning for Colorectal Cancer Risk Prediction and Stratification. Frontiers in Big Data, 2020. Volume 3 - 2020.
18.Burnett, B., et al., Machine Learning in Colorectal Cancer Risk Prediction from Routinely Collected Data: A Review. Diagnostics (Basel), 2023. 13(2).
19.Lundberg, S.M. and S.-I. Lee, A unified approach to interpreting model predictions, in Proceedings of the 31st International Conference on Neural Information Processing Systems. 2017, Curran Associates Inc.: Long Beach, California, USA. p. 4768–4777.
20.Allman, E.S., C. Matias, and J.A. Rhodes, Identifiability of parameters in latent structure models with many observed variables. The Annals of Statistics, 2009. 37(6A): p. 3099–3132, 34.
21.Andersen, P.K., et al., Introduction, in Statistical Models Based on Counting Processes, P.K. Andersen, et al., Editors. 1993, Springer US: New York, NY. p. 1–44.
-
dc.identifier.urihttp://tdr.lib.ntu.edu.tw/jspui/handle/123456789/104645-
dc.description.abstract摘要
研究背景
大腸直腸癌 (colorectal cancer, CRC) 腺瘤-癌症發生序列 (adenoma–carcinoma sequence) 由正常黏膜歷經小腺瘤、進展型腺瘤、臨床症前可偵測期 (preclinical detectable phase, PCDP) 逐步演變為臨床期癌症,疾病進展屬連續且不可直接觀測之潛在過程。傳統疾病自然史模型多以族群平均 (population average) 描述疾病進展,難以反映個體間之疾病異質性;而機器學習方法雖廣泛應用於風險預測,卻多侷限於疾病分類,缺乏與疾病自然史之動態連結。本研究旨在整合群體層級疾病自然史推論與個人化風險預測,建立一套混合式疾病進展架構 (Hybrid Disease Progression Framework)。

研究方法
群體層級分析採用臺灣全國糞便免疫化學檢查 (fecal immunochemical test, FIT) 篩檢資料 (14,478,478 人次篩檢紀錄、5,640,906 名受檢者),建立含九個潛在疾病狀態之連續時間隱藏狀態馬可夫模型 (continuous-time hidden-state Markov model),並透過觀測映射機制 (observation mapping) 結合 FIT 篩檢敏感度連結潛在狀態與觀測結果,以最大概似估計法 (MLE) 推估生成矩陣 (generator matrix) 之各轉移強度,再以矩陣指數法模擬潛在疾病分布。個體層級分析則以彰化地區社區篩檢資料建立「正常 vs. 腺瘤」與「腺瘤 vs. 篩檢偵測癌」兩項二元分類模型,比較人工神經網路 (ANN) 與隨機森林 (Random Forest, RF),並以 SHapley Additive exPlanations (SHAP) 辨識重要危險因子。最後將機器學習預測機率轉換為相對風險倍數,回饋調節群體生成矩陣以建立個人化生成矩陣 (personalized generator matrix) ,推估個人化疾病進展光譜。

研究結果
生成矩陣估計顯示,正常進入腺瘤之轉移強度 (q₁₂ = 0.00112) 遠低於後續各階段之進展速度 (如 q₂₃ = 0.5896、q₃₄ = 0.5889),且觀測結果與真實疾病狀態並非一對一對應,FIT 陰性個案之間隔癌反映篩檢敏感度之限制。在機器學習分類中,「正常 vs. 腺瘤」以 RF 表現最佳 (測試集 AUC = 0.935),優於 ANN (AUC = 0.896)。然而,「腺瘤 vs. 篩檢偵測癌」因兩階段特徵重疊、分類難度較高,測試集 AUC 分別降至 0.716(ANN)與 0.694(RF)。SHAP 分析一致指出誘導時間(induction time)、糞便血紅素濃度 (f-Hb)、年齡及代謝症候群相關指標為最重要之預測因子。個人化進展模擬顯示,q₁₂ (是否容易形成腺瘤)對整體癌症風險之影響遠大於 q₃₄ (腺瘤惡化速度):高 q₁₂/低q₃₄ 個案之十年臨床期癌症風險 (2.46%) 明顯高於低 q₁₂/高 q₃₄ 個案,顯示「正常→腺瘤」階段為整體風險之上游主閘門。

結論
本研究首次整合九狀態連續時間隱藏狀態疾病自然史模型與人工智慧個人化風險預測,建立由群體疾病進展延伸至個人化疾病進展之分析架構,使疾病自然史研究由「描述平均進展」提升至「推估個人進展」。此一兼具機制解釋能力與個人化預測能力之架構,可作為風險導向篩檢 (risk-adaptive screening)、個人化篩檢間隔設計與臨床共享決策之理論基礎,並有望成為 AI 數位雙生 (AI-enabled Digital Twin)及精準公共衛生之方法學平台。辨識「容易形成腺瘤」之高風險個體較辨識「腺瘤惡化速度快」者更能於上游有效攔截大腸直腸癌之發生。
zh_TW
dc.description.abstractAbstract
Background
Colorectal cancer (CRC) follows the adenoma–carcinoma sequence, progressing from normal mucosa through small and advanced adenomas and the preclinical detectable phase (PCDP) to clinical cancer. This progression is a continuous and latent process that cannot be directly observed. Conventional disease natural history models characterize progression at the population-average level and fail to capture individual heterogeneity, whereas machine learning approaches, though widely applied to risk prediction, are largely confined to disease classification and lack a dynamic link to the disease natural history. This study aimed to integrate population-level disease natural history inference with personalized risk prediction into a Hybrid Disease Progression Framework.

Methods
At the population level, nationwide fecal immunochemical test (FIT) screening data from Taiwan (14,478,478 screening records; 5,640,906 participants) were used to construct a nine-state continuous-time hidden-state Markov model. An observation mapping mechanism incorporating FIT screening sensitivity linked latent states to observed screening outcomes, and the transition intensities of the generator matrix were estimated by maximum likelihood estimation (MLE) , with the latent disease distribution simulated via the matrix exponential. At the individual level, community screening data from Changhua were used to build two binary classifiers—Normal vs. Adenoma and Adenoma vs. Screen-detected CRC—comparing an artificial neural network (ANN) and a Random Forest (RF), with SHapley Additive exPlanations (SHAP) applied to identify important risk factors. Finally, RF-predicted probabilities were converted into relative risk multipliers to modulate the population generator matrix, yielding a personalized generator matrix and a personalized disease progression spectrum.

Results
The estimated generator matrix showed that the transition from normal to small adenoma (q₁₂ = 0.00112) was markedly slower than subsequent progression steps (e.g., q₂₃ = 0.5896, q₃₄ = 0.5889), and observed outcomes did not correspond one-to-one with true disease states—interval cancers among FIT-negative cases reflected the limits of screening sensitivity. For the Normal vs. Adenoma classification, RF performed best (test AUC = 0.935), outperforming ANN (AUC = 0.896). The Adenoma vs. Screen-detected CRC task was more challenging owing to overlapping features, yielding test AUCs of 0.716 and 0.694 for ANN and RF, respectively. SHAP analyses consistently identified induction time, fecal hemoglobin concentration (f-Hb), age, and metabolic-syndrome-related indicators as the most important predictors. Personalized simulations indicated that the propensity to form an adenoma (q₁₂) influenced overall cancer risk far more than the speed of adenoma malignant progression (q₃₄): a high-q₁₂/low-q₃₄ case carried a 10-year clinical-cancer risk of 2.46%, substantially higher than a low-q₁₂/high-q₃₄ case, identifying the Normal→Adenoma step as the upstream gatekeeper of overall risk.

Conclusion
This study is the first to integrate a nine-state continuous-time hidden-state disease natural history model with AI-based personalized risk prediction, establishing an analytical framework that extends population-level progression to personalized disease progression and elevates natural history research from describing average progression to estimating individual progression. Combining mechanistic interpretability with personalized prediction, the framework can serve as a theoretical basis for risk-adaptive screening, personalized screening-interval design, and shared clinical decision-making, and as a methodological platform for AI-enabled Digital Twins and precision public health. Identifying individuals prone to forming adenomas appears more effective for intercepting CRC upstream than identifying those with fast-progressing adenomas. Limitations include reliance on single-region data for the personalized model and the static nature of the generator matrix; external validation and integration of longitudinal and multi-omics data are warranted.
en
dc.description.provenanceSubmitted by admin ntu (admin@lib.ntu.edu.tw) on 2026-08-28T16:54:06Z
No. of bitstreams: 0
en
dc.description.provenanceMade available in DSpace on 2026-08-28T16:54:06Z (GMT). No. of bitstreams: 0en
dc.description.tableofcontents目次
致謝 i
摘要 ii
Abstract iv
目次 vii
圖次 x
表次 xi
第一章 引言 1
1.1 研究背景 1
1.2 研究目的 3
1.3 研究流程 3
1.4 論文架構 4
第二章 文獻回顧 5
2.1 大腸直腸癌篩檢與疾病自然史 5
2.2 臺灣大腸直腸癌篩檢計畫與糞便免疫化學檢查 6
2.3 大腸直腸癌自然史模型 7
2.4 個人化因子與大腸直腸癌進展風險 9
2.5 機器學習於大腸直腸癌個人化風險預測之角色 11
2.6 現有研究限制與研究動機 13
第三章 研究架構 15
3.1 大腸直腸癌疾病自然史模型 15
3.2 連續時間隱藏狀態疾病進展模型 18
3.3 模型可識別性 21
3.3.1 基本假設 22
3.3.2 轉移矩陣之可識別性 24
3.3.3 生成矩陣之唯一性 26
3.3.4 全域可識別性 29
3.4 大樣本概似理論 30
3.4.1一致性 32
3.4.2漸近常態性 33
3.5 機器學習預測模型 35
3.5.1研究對象與研究變項 35
3.5.2 資料處理與模型建立 38
3.5.3 人工神經網路模型 42
3.5.4 隨機森林模型 45
3.5.5 模型效能評估 47
3.5.6 模型解釋性分析 50
3.6 個人化風險分層之疾病進展模型 54
3.7小結 56
第四章 資料分析結果 58
4.1 全國FIT篩檢資料庫之人口特徵與觀測結果分布 58
4.2 生成矩陣參數估計結果 64
4.3 發射矩陣估計結果 65
4.4 潛在疾病狀態分布模擬結果 66
4.5機器學習模型分析 69
4.5.1 彰化社區篩檢資料之人口學與個人化特徵描述 69
4.5.2 正常與腺瘤分類器 81
4.5.3 腺瘤與篩檢癌分類器 89
4.5.4綜合比較與研究意涵 97
4.6 個人化風險分層之疾病進展 98
4.6.1 風險分層相對風險倍數 98
4.6.2 各風險層之累積進展機率 101
4.6.3 四種代表性風險情境之十年疾病進展模擬 103
第五章 討論與結論 117
5.1 主要研究發現 117
5.2 群體層級疾病自然史模型之方法學意義 118
5.3 個人化風險預測與機器學習 118
5.4 群體疾病自然史與個人化疾病進展之整合 119
5.5 臨床與公共衛生意涵 121
5.6 研究限制與未來發展 122
5.7 結論 124
-
dc.language.isozh_TW-
dc.subject大腸直腸癌-
dc.subject疾病自然史-
dc.subject連續時間隱藏狀態馬可夫模型-
dc.subject生成矩陣-
dc.subject機器學習-
dc.subject隨機森林-
dc.subjectSHAP-
dc.subject個人化疾病進展-
dc.subject精準篩檢-
dc.subjectColorectal cancer-
dc.subjectdisease natural history-
dc.subjectcontinuous-time hidden-state Markov model-
dc.subjectgenerator matrix-
dc.subjectmachine learning-
dc.subjectRandom Forest-
dc.subjectSHAP-
dc.subjectpersonalized disease progression-
dc.subjectprecision screening-
dc.title大腸癌腺瘤及癌症進展之機器學習模式分析zh_TW
dc.titleProgression of Colorectal Adenoma and Cancer: A Machine Learning Analysisen
dc.typeThesis-
dc.date.schoolyear114-2-
dc.description.degree碩士-
dc.contributor.oralexamcommittee嚴明芳;江濬如;許文峰zh_TW
dc.contributor.oralexamcommitteeMing-Fang Yen;Chun-Ju Chiang;Wen-Feng Hsuen
dc.subject.keyword大腸直腸癌; 疾病自然史; 連續時間隱藏狀態馬可夫模型; 生成矩陣; 機器學習; 隨機森林; SHAP; 個人化疾病進展; 精準篩檢zh_TW
dc.subject.keywordColorectal cancer; disease natural history; continuous-time hidden-state Markov model; generator matrix; machine learning; Random Forest; SHAP; personalized disease progression; precision screeningen
dc.relation.page126-
dc.identifier.doi10.6342/NTU202602503-
dc.rights.note同意授權(全球公開)-
dc.date.accepted2026-07-28-
dc.contributor.author-college公共衛生學院-
dc.contributor.author-dept健康數據拓析統計研究所-
dc.date.embargo-lift2026-08-29-
顯示於系所單位:健康數據拓析統計研究所

文件中的檔案:
檔案 描述 大小格式 
ntu-114-2.pdf3.91 MBAdobe PDF檢視/開啟
ntu-114-2.pdf3.91 MBAdobe PDF檢視/開啟
顯示文件簡單紀錄


系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。

社群連結
聯絡資訊
10617臺北市大安區羅斯福路四段1號
No.1 Sec.4, Roosevelt Rd., Taipei, Taiwan, R.O.C. 106
Tel: (02)33662353
Email: ntuetds@ntu.edu.tw
意見箱
相關連結
館藏目錄
國內圖書館整合查詢 MetaCat
臺大學術典藏 NTU Scholars
臺大圖書館數位典藏館
本站聲明
© NTU Library All Rights Reserved