請用此 Handle URI 來引用此文件:
http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103981| 標題: | ORBIT:應用於開源 AI 資料集的 IPFS 區域協作儲存架構 ORBIT: On-Demand Regional Binding of IPFS Teams for Open-Source AI Dataset Storage |
| 作者: | 李奕杰 Yi-Jie Li |
| 指導教授: | 莊裕澤 Yuh-Jzer Joung |
| 關鍵字: | IPFS; 去中心化儲存; 協同快取; 地理局部性; AI 資料集; 點對點系統 IPFS; decentralized storage; cooperative caching; geographic locality; AI datasets; peer-to-peer systems |
| 出版年 : | 2026 |
| 學位: | 碩士 |
| 摘要: | 開源人工智慧(Open-Source AI)所依賴的資料集,普遍具有體積龐大、內容不可變、且每逢訓練任務便需被完整下載的特性。然而,這類資料集目前幾乎全數託管於中心化的基礎設施之上,使開放生態系賴以協作的資料,其可得性與主權皆繫於少數業者的裁量。星際檔案系統(InterPlanetary File System, IPFS)以內容定址與去中心化的特性,為此提供了具抗審查性的替代方案;惟原生 IPFS 的檢索須歷經分散式雜湊表(DHT)查找與 Bitswap 的循序區塊交換,若請求者周遭找不到鄰近檔案副本,對於資料集規模的應用場景而言仍嫌緩慢。
快取是一個可行的解決方案,但現有的被動式快取並不足以化解此一瓶頸:各節點僅依自身觀察到的熱度獨立快取,既無法保證任一區域能湊齊完整的資料集,也可能因重複快取相同內容而虛耗容量。有鑑於此,本研究提出 ORBIT(On-Demand Regional Binding of IPFS Teams),一套將資料集綁定至其請求來源區域的協同快取架構。本研究首先透過儲存者與計算者的單位效益分析,論證地理鄰近的資料配置對雙方均為有利,並據此設計兩項核心機制:其一為優先級競標,將每筆任務指派予最適合服務的區域;其二為按需組隊,於資料集首次被請求時將其決定性地分片至一組節點,在近乎最小的儲存成本下保證整個資料集於區域內的完整覆蓋。此外,額外快取機制使團隊的持有內容得以隨請求樣態的漂移而動態演進。 本研究於三種環境中驗證此架構:橫跨三個真實雲端區域的 30 節點實機部署、注入跨區延遲的 100 節點三區叢集,以及擴展至十區的 300 節點叢集。實驗結果顯示,ORBIT 的下載速度相對於原生 IPFS 達約 3.2 至 4.1 倍,相對於表現最佳的 Edge CDN 對照組亦達約 3.0 至 3.3 倍,且各項差異均具統計顯著性(p < 0.001),其下載時間的變異與長尾延遲亦為所有方法中最低。同時,ORBIT 在每節點與全系統的資源用量皆逼近原生 IPFS;當部署規模由三區擴展至十區時,其效能依然穩定。整體而言,本研究為下一代去中心化的開源 AI 生態系,奠定了一套高效、穩定且可擴展的資料儲存基礎。 Open-source AI depends on large, immutable datasets repeatedly downloaded in full, yet they sit almost exclusively on centralized infrastructure. IPFS offers a content-addressed, censorship-resistant alternative, but DHT lookups and sequential cross-continent block exchange make retrieval slow, and Edge CDN, in which each node independently pins popular content, guarantees neither complete regional coverage nor unique copies, and can slow retrieval at scale. We present ORBIT (On-Demand Regional Binding of IPFS Teams), a cooperative caching architecture that binds datasets to requesting regions. Its key enabler is a decoupling of development from computation: compute providers train on developers' behalf and return only megabyte-scale weights, and we prove that geographically local placement then maximizes the benefit of rational storage and compute providers. ORBIT exploits this alignment through a coverage-driven priority auction dispatching each task to the best-positioned region, on-demand team formation partitioning datasets across regional nodes exactly once, and an extra-pin cache tracking drifting requests. On a 30-node three-region deployment and emulated clusters up to 300 nodes and ten regions, ORBIT retrieves datasets 3.0 to 3.3 times faster than the best Edge CDN deployment and 3.2 to 4.1 times faster than vanilla IPFS, with the lowest variance and tail latency, and near-vanilla resource usage. |
| URI: | http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103981 |
| DOI: | 10.6342/NTU202602927 |
| 全文授權: | 同意授權(全球公開) |
| 電子全文公開日期: | 2026-08-21 |
| 顯示於系所單位: | 資訊管理學系 |
文件中的檔案:
| 檔案 | 大小 | 格式 | |
|---|---|---|---|
| ntu-114-2.pdf | 5.43 MB | Adobe PDF | 檢視/開啟 |
系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。
