Skip navigation

DSpace

機構典藏 DSpace 系統致力於保存各式數位資料(如:文字、圖片、PDF)並使其易於取用。

點此認識 DSpace
DSpace logo
English
中文
  • 瀏覽論文
    • 校院系所
    • 出版年
    • 作者
    • 標題
    • 關鍵字
    • 指導教授
  • 搜尋 TDR
  • 授權 Q&A
    • 我的頁面
    • 接受 E-mail 通知
    • 編輯個人資料
  1. NTU Theses and Dissertations Repository
  2. 電機資訊學院
  3. 資訊工程學系
請用此 Handle URI 來引用此文件: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103502
標題: CoMoM:透過單一影片運動先驗學習協調行動操作
CoMoM: Learning Coordinated Mobile Manipulation via Single-Video Motion Priors
作者: 楊育勝
Yu-Sheng Yang
指導教授: 徐宏民
Winston H. Hsu
關鍵字: 協調行動操作; 從人類影片學習; 強化學習; 策略蒸餾; 模擬到真實遷移; 視覺運動策略
Coordinated Mobile Manipulation; Learning from Human Videos; Reinforcement Learning; Policy Distillation; Sim-to-Real Transfer; Visuomotor Policy
出版年 : 2026
學位: 碩士
摘要: 協調行動操作要求機器人同步控制底盤移動與手臂操作,藉由動態擴展工作空間,避免底盤固定時在關節式物體操作任務(例如開門)上遭遇的運動學限制。模仿學習雖能達成協調控制,卻高度仰賴耗費人力的機器人遙控操作;從人類影片學習提供了可擴展的替代方案,但現有方法侷限於桌面型操作或解耦式執行;強化學習雖不需示範資料,卻需人工針對個別任務設計高成本的獎勵函數。
為了解決上述限制,我們提出 CoMoM,僅需一支單目人類影片即可學習協調行動操作策略。CoMoM 先從影片擷取三維的手部與底盤軌跡,再以基於語意錨點的空間扭曲將其對齊至隨機化的模擬環境,產生可適應不同場景配置的軌跡先驗。這些先驗作為參考路徑,構造出軌跡引導式獎勵函數,為強化學習提供訓練特權專家所需的訊號。訓練完成的專家接著生成加入雜訊擾動的合成資料集,蒸餾出具備三維感知能力的視覺運動策略以供部署。
在五項協調式任務上的評估顯示,相較於具代表性的基準方法,CoMoM 在任務進度上平均提升 40%、任務執行效率上平均提升 56%。我們將學習到的策略部署於實體 Stretch 機器人,驗證了零樣本模擬到真實遷移的可行性。這些結果顯示,單一人類影片足以作為協調行動操作的低成本運動先驗來源。
Coordinated mobile manipulation synchronizes base locomotion and arm manipulation. This coordination avoids the kinematic limitations of a stationary base during articulated tasks, such as door opening, by dynamically expanding the workspace. Imitation learning achieves coordinated control but remains constrained by a reliance on labor-intensive robot teleoperation. While learning from a human video offers a scalable alternative, existing methods are restricted to tabletop manipulation or decoupled execution. Alternatively, reinforcement learning (RL) learns without demonstrations but relies on manual task-specific reward engineering.
To address these limitations, we present CoMoM, a framework that learns coordinated mobile manipulation policies from a single monocular human video. CoMoM extracts 3D hand and base trajectories from the video and aligns them to randomized simulations via semantic anchor-based spatial warping. The warping yields layout-adaptive trajectory priors, which serve as reference paths to formulate a task-agnostic trajectory-guided reward function, providing RL signals to train a privileged expert. The trained RL expert then generates a noise-augmented synthetic dataset to distill a 3D-aware visuomotor policy for deployment.
Evaluated across five coordinated tasks, CoMoM achieves an average gain of 40% in task progress and 56% in execution efficiency relative to representative baselines. Finally, we deploy the learned policy on a physical Stretch robot, demonstrating the feasibility of zero-shot sim-to-real transfer. Together, these results show that a single human video can serve as a low-cost source of motion priors for coordinated mobile manipulation.
URI: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103502
DOI: 10.6342/NTU202603803
全文授權: 未授權
電子全文公開日期: N/A
顯示於系所單位:資訊工程學系

文件中的檔案:
檔案 大小格式 
ntu-114-2.pdf
  未授權公開取用
5.65 MBAdobe PDF
顯示文件完整紀錄


系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。

社群連結
聯絡資訊
10617臺北市大安區羅斯福路四段1號
No.1 Sec.4, Roosevelt Rd., Taipei, Taiwan, R.O.C. 106
Tel: (02)33662353
Email: ntuetds@ntu.edu.tw
意見箱
相關連結
館藏目錄
國內圖書館整合查詢 MetaCat
臺大學術典藏 NTU Scholars
臺大圖書館數位典藏館
本站聲明
© NTU Library All Rights Reserved