請用此 Handle URI 來引用此文件:
http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/104528完整後設資料紀錄
| DC 欄位 | 值 | 語言 |
|---|---|---|
| dc.contributor.advisor | 洪弘 | zh_TW |
| dc.contributor.advisor | Hung Hung | en |
| dc.contributor.author | 王筱媛 | zh_TW |
| dc.contributor.author | Hsiao-Yuan Wang | en |
| dc.date.accessioned | 2026-08-28T16:08:34Z | - |
| dc.date.available | 2026-08-29 | - |
| dc.date.copyright | 2026-08-28 | - |
| dc.date.issued | 2026 | - |
| dc.date.submitted | 2026-06-29 00:00:00 | - |
| dc.identifier.citation | Frank Critchley. Influence in principal components analysis. Biometrika, 72(3):627–636, 1985.
Hung Hung and Su-Yun Huang. On the efficiency-loss free ordering-robustness of product-pca. arXiv preprint arXiv:2302.11124, 2023. HungHung,Su-YunHuang, andChing-KangIng. Ageneralized information criterion for high-dimensional pca rank selection. Statistical Papers, 63(4):1295–1321, 2022. Sadanori Konishi. Normalizing transformatins and bootstrap confidence intervals. The Annals of Statistics, 19(4):2209–2225, 1991. Sadanori Konishi and Genshiro Kitagawa. Generalised information criteria in model selection. Biometrika, 83(4):875–890, 1996. Solomon Kullback and Richard A Leibler. On information and sufficiency. The annals of mathematical statistics, 22(1):79–86, 1951. Kazuyoshi YataandMakotoAoshima. Effectivepcaforhigh-dimension, low-sample-size data with singular value decomposition of cross data matrix. Journal of multivariate analysis, 101(9):2060–2077, 2010 | - |
| dc.identifier.uri | http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/104528 | - |
| dc.description.abstract | 在巨量資料時代,主成分分析(PCA)已成為降維與特徵提取的核心技術。然而,實務應用中仍面臨一項關鍵挑戰:如何在維持高保真度還原的前提下,縮減至最精簡的秩。儘管廣義資訊準則(GIC)提供了一個穩健的評估框架,但傳統 PCA 對非結構化雜訊與異常值極為敏感,易導致估計的訊號子空間產生偏誤。為彌補此理論與實務間的缺憾,本文將乘積主成分分析(Product-PCA, PPCA)整合至 GIC 框架中,利用其資料分割的特性來克服異常值的干擾。
本研究的核心貢獻為根據一階影響函數,推導出 PPCA 估計量的偏差校正項,為乘積型協方差結構的模型選擇奠定理論基礎。數值模擬與 ImageNet 影像重建實驗均證實,PPCA-GIC 的性能顯著優於傳統方法。在資料受汙染的情況下,PPCA 展現出卓越的排序穩健性,僅需較低的秩即可達到與 PCA 相同的子空間相似度,實現更高的模型簡約性與穩健性。此外,在未受汙染的理想環境下,兩者仍保持漸進等價性,不損失統計效率。研究結果證實,PPCA-GIC 是高維噪音環境下進行秩辨識與訊號恢復的高效工具。 | zh_TW |
| dc.description.abstract | In the era of massive data, Principal Component Analysis (PCA) serves as a critical gateway for dimensionality reduction. However, a fundamental trade-off persists: how to achieve high-fidelity reconstruction using the most parsimonious rank. While the Generalized Information Criterion (GIC) offers a robust evaluative framework, classical PCA remains inherently sensitive to unstructured noise and outliers, which can distort the estimated signal subspace. To bridge this gap, this thesis integrates Product-PCA (PPCA)—which leverages data partitioning to filter out idiosyncratic interference—into the GIC framework.
The primary contribution of this research is the formal derivation of bias correction terms for PPCA based on first-order influence functions. Numerical simulations and ImageNet experiments validate that PPCA-GIC outperforms traditional methods. In contaminated scenarios, PPCA achieves equivalent subspace similarity to PCA while utilizing a significantly lower rank, demonstrating superior model parsimony and robustness. Conversely, it maintains asymptotic equivalence to PCA in uncontaminated settings. These findings confirm that PPCA-GIC is a highly efficient tool for rank identification in high-dimensional, noisy environments. | en |
| dc.description.provenance | Submitted by admin ntu (admin@lib.ntu.edu.tw) on 2026-08-28T16:08:34Z No. of bitstreams: 0 | en |
| dc.description.provenance | Made available in DSpace on 2026-08-28T16:08:34Z (GMT). No. of bitstreams: 0 | en |
| dc.description.tableofcontents | Chapter 1 Introduction 1
1.1 Review of PCA and GIC for Rank Selection 1 1.2 Transition to Product-PCA (PPCA) 7 1.3 Main Contributions and Thesis Structure 9 Chapter 2 Methodology 11 2.1 Derivation of the Generalized Information Criterion 12 2.2 Derivation of Influence Functions for PPCA 14 2.3 Formulation of the GIC-PPCA Criterion 20 Chapter 3 Numerical Simulations 25 3.1 Simulation Settings 25 3.2 Evaluation Metrics 27 3.3 Simulation Results 28 3.3.1 Performance under Null Perturbation 28 3.3.2 Performance under Heterogeneous Contamination for Multivariate T Distribution Outlier 29 Chapter 4 Application 39 4.1 Image Reconstruction Methodology 39 4.2 Experimental Results and Analysis 40 4.2.1 U-shaped Error Profile: Sharp Signal Identification 41 4.2.2 Asymptotic Profile: Robustness in Cluttered Scenarios 42 Chapter 5 Discussion 47 References 51 | - |
| dc.language.iso | en | - |
| dc.subject | 乘積主成分分析 | - |
| dc.subject | 廣義資訊準則 | - |
| dc.subject | 影響函數對等性 | - |
| dc.subject | 高維模型選擇 | - |
| dc.subject | 自動化秩選取 | - |
| dc.subject | Product-Principal Component Analysis | - |
| dc.subject | Generalized Information Criterion | - |
| dc.subject | Influence Function Equivalence | - |
| dc.subject | High-dimensional Model Selection | - |
| dc.subject | Automatic Rank Selection | - |
| dc.title | 應用廣義資訊準則於乘積主成分分析之模型選擇 | zh_TW |
| dc.title | A Generalized Information Model Selection Criterion for Product-Principal Component Analysis | en |
| dc.type | Thesis | - |
| dc.date.schoolyear | 114-2 | - |
| dc.description.degree | 碩士 | - |
| dc.contributor.oralexamcommittee | 盧子彬;陳素雲 | zh_TW |
| dc.contributor.oralexamcommittee | TZU-PIN LU;Su-Yun Huang | en |
| dc.subject.keyword | 乘積主成分分析; 廣義資訊準則; 影響函數對等性; 高維模型選擇; 自動化秩選取 | zh_TW |
| dc.subject.keyword | Product-Principal Component Analysis; Generalized Information Criterion; Influence Function Equivalence; High-dimensional Model Selection; Automatic Rank Selection | en |
| dc.relation.page | 81 | - |
| dc.identifier.doi | 10.6342/NTU202600923 | - |
| dc.rights.note | 未授權 | - |
| dc.date.accepted | 2026-06-30 | - |
| dc.contributor.author-college | 公共衛生學院 | - |
| dc.contributor.author-dept | 健康數據拓析統計研究所 | - |
| dc.date.embargo-lift | N/A | - |
| 顯示於系所單位: | 健康數據拓析統計研究所 | |
文件中的檔案:
| 檔案 | 大小 | 格式 | |
|---|---|---|---|
| ntu-114-2.pdf 未授權公開取用 | 3.71 MB | Adobe PDF |
系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。
