請用此 Handle URI 來引用此文件:
http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103502| 標題: | CoMoM:透過單一影片運動先驗學習協調行動操作 CoMoM: Learning Coordinated Mobile Manipulation via Single-Video Motion Priors |
| 作者: | 楊育勝 Yu-Sheng Yang |
| 指導教授: | 徐宏民 Winston H. Hsu |
| 關鍵字: | 協調行動操作; 從人類影片學習; 強化學習; 策略蒸餾; 模擬到真實遷移; 視覺運動策略 Coordinated Mobile Manipulation; Learning from Human Videos; Reinforcement Learning; Policy Distillation; Sim-to-Real Transfer; Visuomotor Policy |
| 出版年 : | 2026 |
| 學位: | 碩士 |
| 摘要: | 協調行動操作要求機器人同步控制底盤移動與手臂操作,藉由動態擴展工作空間,避免底盤固定時在關節式物體操作任務(例如開門)上遭遇的運動學限制。模仿學習雖能達成協調控制,卻高度仰賴耗費人力的機器人遙控操作;從人類影片學習提供了可擴展的替代方案,但現有方法侷限於桌面型操作或解耦式執行;強化學習雖不需示範資料,卻需人工針對個別任務設計高成本的獎勵函數。
為了解決上述限制,我們提出 CoMoM,僅需一支單目人類影片即可學習協調行動操作策略。CoMoM 先從影片擷取三維的手部與底盤軌跡,再以基於語意錨點的空間扭曲將其對齊至隨機化的模擬環境,產生可適應不同場景配置的軌跡先驗。這些先驗作為參考路徑,構造出軌跡引導式獎勵函數,為強化學習提供訓練特權專家所需的訊號。訓練完成的專家接著生成加入雜訊擾動的合成資料集,蒸餾出具備三維感知能力的視覺運動策略以供部署。 在五項協調式任務上的評估顯示,相較於具代表性的基準方法,CoMoM 在任務進度上平均提升 40%、任務執行效率上平均提升 56%。我們將學習到的策略部署於實體 Stretch 機器人,驗證了零樣本模擬到真實遷移的可行性。這些結果顯示,單一人類影片足以作為協調行動操作的低成本運動先驗來源。 Coordinated mobile manipulation synchronizes base locomotion and arm manipulation. This coordination avoids the kinematic limitations of a stationary base during articulated tasks, such as door opening, by dynamically expanding the workspace. Imitation learning achieves coordinated control but remains constrained by a reliance on labor-intensive robot teleoperation. While learning from a human video offers a scalable alternative, existing methods are restricted to tabletop manipulation or decoupled execution. Alternatively, reinforcement learning (RL) learns without demonstrations but relies on manual task-specific reward engineering. To address these limitations, we present CoMoM, a framework that learns coordinated mobile manipulation policies from a single monocular human video. CoMoM extracts 3D hand and base trajectories from the video and aligns them to randomized simulations via semantic anchor-based spatial warping. The warping yields layout-adaptive trajectory priors, which serve as reference paths to formulate a task-agnostic trajectory-guided reward function, providing RL signals to train a privileged expert. The trained RL expert then generates a noise-augmented synthetic dataset to distill a 3D-aware visuomotor policy for deployment. Evaluated across five coordinated tasks, CoMoM achieves an average gain of 40% in task progress and 56% in execution efficiency relative to representative baselines. Finally, we deploy the learned policy on a physical Stretch robot, demonstrating the feasibility of zero-shot sim-to-real transfer. Together, these results show that a single human video can serve as a low-cost source of motion priors for coordinated mobile manipulation. |
| URI: | http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103502 |
| DOI: | 10.6342/NTU202603803 |
| 全文授權: | 未授權 |
| 電子全文公開日期: | N/A |
| 顯示於系所單位: | 資訊工程學系 |
文件中的檔案:
| 檔案 | 大小 | 格式 | |
|---|---|---|---|
| ntu-114-2.pdf 未授權公開取用 | 5.65 MB | Adobe PDF |
系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。
