Skip navigation

DSpace

機構典藏 DSpace 系統致力於保存各式數位資料(如:文字、圖片、PDF)並使其易於取用。

點此認識 DSpace
DSpace logo
English
中文
  • 瀏覽論文
    • 校院系所
    • 出版年
    • 作者
    • 標題
    • 關鍵字
    • 指導教授
  • 搜尋 TDR
  • 授權 Q&A
    • 我的頁面
    • 接受 E-mail 通知
    • 編輯個人資料
  1. NTU Theses and Dissertations Repository
  2. 電機資訊學院
  3. 電信工程學研究所
請用此 Handle URI 來引用此文件: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/102965
標題: BOCCHI 資料集與 MSDCT-UNet:局部動態模糊偵測
The BOCCHI Benchmark and MSDCT-UNet for Local Motion Blur Detection
作者: 陳冠霖
Kuan-Lin Chen
指導教授: 丁建均
Jian-Jiun Ding
關鍵字: 局部動態模糊偵測; BOCCHI 資料集; 離散餘弦轉換; 頻域學習; 語義分割
Local motion blur detection; BOCCHI dataset; Discrete Cosine Transform; Frequency domain learning; Semantic segmentation
出版年 : 2026
學位: 碩士
摘要: 局部動態模糊偵測旨在以像素級精度定位影像中受運動模糊影響的區域。當模糊與清晰區域可由梯度幅值簡單區分時,現有的模糊偵測基準容易讓模型依賴低梯度的代理特徵,而非真正學到運動模糊背後的物理性質。
為填補此空白,本研究建立 BOCCHI(Blurred Objects Captured across Cameras with Human-annotated Imagery)資料集:共 633 張以五台分屬不同廠牌的消費級相機在真實場景拍攝、並完成像素級人工標註的影像。BOCCHI 的清晰區域呈現雙峰梯度分布,其中低梯度模態(清晰 PR25 = 20.15)落在模糊區平均梯度(µblur = 29.62)之下,使清晰 PR25/µblur 比值(0.68)為受評五個資料集中最高,梯度分布重疊使單純依賴梯度幅值的判別捷徑失效。
我們進一步提出 MSDCT-UNet(Multi-Scale Discrete Cosine Transform UNet),一個對影像梯度進行多尺度 DCT 變換以取得頻域先驗、以多頭 DCT Attention 學習各頻率成分的重要性、並以 FiLM 機制融合頻域與空間特徵的編碼器-解碼器網路。在 BOCCHI 上,MSDCT-UNet 較第二名 baseline 提升 1.25 個百分點的 mIoU 並在邊界偵測上領先。跨資料集評估進一步揭示一個資料集層級的訓練效率優勢:以 BOCCHI 訓練的模型在 mIoU / Dice / Recall 平均值上勝過其他三個訓練來源(ReLoBlur、CUHKmotion、OMoBlur),且僅使用 633 張訓練影像,遠少於 ReLoBlur(1,200)與 OMoBlur(994)。此優勢與模型架構無關,顯示主導跨資料集可遷移性的關鍵是場景組成,而非訓練資料規模。
Local motion blur detection requires pixel-level localization of blurred regions. Existing benchmarks can permit gradient-based shortcuts when blurred and sharp regions are easily separable, letting models rely on low-gradient proxies rather than the underlying physics of motion blur.
To address this gap, we introduce BOCCHI (Blurred Objects Captured across Cameras with Human-annotated Imagery), 633 real-captured, pixel-annotated images shot with five consumer-grade cameras spanning five different manufacturers. BOCCHI’s sharp regions exhibit a bimodal gradient distribution including a low-gradient mode (sharp PR25 = 20.15) that falls below the blur-region mean (µblur = 29.62), giving the largest sharp PR25/µblur ratio (0.68) among five evaluated datasets and creating gradient overlap that defeats simple gradient shortcuts.
We further propose MSDCT-UNet (Multi-Scale Discrete Cosine Transform UNet), a frequency-aware encoder-decoder that extracts multi-scale DCT priors from gradients, weights them via multi-head DCT Attention, and fuses frequency and spatial cues through FiLM modulation. On BOCCHI, MSDCT-UNet surpasses the second-best baseline by 1.25 percentage points in mIoU and leads in boundary localization. Cross-dataset evaluation further reveals a dataset-level data-efficiency advantage: BOCCHI-trained models outperform every other training source (ReLoBlur, CUHKmotion, OMoBlur) on mIoU/Dice/Recall averages using only 633 training images, substantially fewer than ReLoBlur (1,200) and OMoBlur (994). This advantage is independent of architecture, suggesting that scene composition, rather than dataset scale, drives cross-dataset transferability for local motion blur detection.
URI: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/102965
DOI: 10.6342/NTU202602037
全文授權: 同意授權(限校園內公開)
電子全文公開日期: 2031-07-16
顯示於系所單位:電信工程學研究所

文件中的檔案:
檔案 大小格式 
ntu-114-2.pdf
  未授權公開取用
14.27 MBAdobe PDF檢視/開啟
顯示文件完整紀錄


系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。

社群連結
聯絡資訊
10617臺北市大安區羅斯福路四段1號
No.1 Sec.4, Roosevelt Rd., Taipei, Taiwan, R.O.C. 106
Tel: (02)33662353
Email: ntuetds@ntu.edu.tw
意見箱
相關連結
館藏目錄
國內圖書館整合查詢 MetaCat
臺大學術典藏 NTU Scholars
臺大圖書館數位典藏館
本站聲明
© NTU Library All Rights Reserved