Skip navigation

DSpace

機構典藏 DSpace 系統致力於保存各式數位資料(如:文字、圖片、PDF)並使其易於取用。

點此認識 DSpace
DSpace logo
English
中文
  • 瀏覽論文
    • 校院系所
    • 出版年
    • 作者
    • 標題
    • 關鍵字
    • 指導教授
  • 搜尋 TDR
  • 授權 Q&A
    • 我的頁面
    • 接受 E-mail 通知
    • 編輯個人資料
  1. NTU Theses and Dissertations Repository
  2. 共同教育中心
  3. 統計碩士學位學程
請用此 Handle URI 來引用此文件: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103717
標題: 多智慧體在強化學習中湧現出攻擊行為: 統計建模與 AI 安全啟示
Emergent Attack in Multi-Agent Reinforcement Learning: Statistical Modeling and Implications for AI Safety
作者: 魏逸豪
I-Hao Wei
指導教授: 黃從仁
Tsung-Ren Huang
關鍵字: 多智慧體強化學習; 深度 Q 網路; 湧現行為; 有害湧現; 統計建模; 人工智慧安全
Multi-Agent Reinforcement Learning; Deep Q-Network; Emergent Behavior; Harmful Emergence; Statistical Modeling; AI Safety
出版年 : 2026
學位: 碩士
摘要: 強化學習已成為現代人工智慧的核心方法之一,多智慧體強化學習(MARL)的相關研究也隨之蓬勃發展。然而,既有研究多聚焦於合作型任務的效率提升或顯式對抗賽局,在非預先設計為對抗的環境中,智慧體是否會自發產生傷害性行為,則仍缺乏系統性探討。本研究透過模擬實驗,以掃地機器人共享充電座為情境,模擬兩種 MARL 設定:其一為雙智慧體環境,觀察主動攻擊行為的湧現;其二為三智慧體環境,觀察團體主動攻擊行為的湧現。
本研究識別出主動攻擊湧現的三項必要條件:(1)自我保護動機、(2)復原時間、(3)資源稀缺。其中自我保護由獎勵設計(對能量損耗與死亡的懲罰)內建;復原時間與資源稀缺則經對照實驗逐一驗證。這兩者只要缺一,主動攻擊就不再穩定湧現。在雙智慧體環境中,不論雙方攻擊力是否對稱,只要兩項條件成立,個體間的主動攻擊都會湧現,顯示攻擊力的物理不對稱並不是它的前提。在三智慧體的聯盟環境中,兩個共享能量的智慧體共同面對一個資源較充裕的目標,團體攻擊也自發湧現,並展現共同攻擊與供電分工兩種型態,同樣以這兩項條件為必要前提。換言之,相同的條件不需額外條件,就從個體延伸到群體層次。
研究結果指出,即使沒有明顯的對抗性獎勵設計,無論在個體或群體層次,只要部署環境同時具備足夠長的復原時間與資源稀缺,主動攻擊行為仍可能湧現。這項發現對 MARL 系統從模擬走向真實部署提出一項人工智慧安全的警示:要抑制這類危害行為,只需破除其中任一條件。縮短復原時間,使撞擊不再構成生存威脅;或解除資源稀缺,讓雙方各有充電去處——兩者都能使主動攻擊不再湧現。本研究也為這類行為的環境與獎勵設計提供初步指引。
Reinforcement learning has become one of the central methods of modern artificial intelligence, and research on multi-agent reinforcement learning (MARL) has flourished accordingly. However, existing studies have largely focused either on improving efficiency in cooperative tasks or on explicitly adversarial games; the question of whether agents can spontaneously develop harmful behaviors in environments not designed to be adversarial remains under-explored. This study addresses this gap through simulation experiments, situated in a scenario in which robot vacuum cleaners share a charging station. Two MARL settings are examined: a two-agent environment, in which the emergence of proactive physical attack is observed; and a three-agent environment, in which the emergence of proactive coalition-based group attack is observed.
This study identifies three necessary conditions for the emergence of proactive attack: (1) a self-preservation motive, (2) post-collision recovery time, and (3) resource scarcity. The self-preservation motive is built in through the reward design (penalties on energy loss and death) and serves as a premise underlying all experiments; recovery time and resource scarcity are each validated through controlled removal— remove either, and proactive attack no longer emerges in a stable manner. In the two-agent environment, individual proactive attack emerges whenever both conditions hold, regardless of whether the agents’ attack powers are symmetric, indicating that physical asymmetry is not a prerequisite. In the three-agent coalition environment, where two energy-sharing agents jointly face a single target, group attack likewise emerges—taking the forms of coordinated co-attack and a supply-based division of labor— and again requires these same two conditions; the same conditions thus extend from the individual to the group level without any additional requirement.
The results indicate that, even without any explicit adversarial reward design, at both the individual and group levels, proactive attack behavior can still emerge whenever the deployment environment provides both sufficient post-collision recovery time and resource scarcity. These findings raise an AI-safety warning for the future transition of MARL systems from simulation to real-world deployment: to suppress such harmful behavior, removing either condition alone suffices— shortening recovery time so that a collision no longer poses a survival threat, or relieving resource scarcity so that each agent has its own charging option, is enough to keep proactive attack from emerging. This study further provides preliminary guidelines for designing environments and reward structures that avoid such harmful behaviors.
URI: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/103717
DOI: 10.6342/NTU202601689
全文授權: 同意授權(全球公開)
電子全文公開日期: 2026-08-20
顯示於系所單位:統計碩士學位學程

文件中的檔案:
檔案 大小格式 
ntu-114-2.pdf2.38 MBAdobe PDF檢視/開啟
顯示文件完整紀錄


系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。

社群連結
聯絡資訊
10617臺北市大安區羅斯福路四段1號
No.1 Sec.4, Roosevelt Rd., Taipei, Taiwan, R.O.C. 106
Tel: (02)33662353
Email: ntuetds@ntu.edu.tw
意見箱
相關連結
館藏目錄
國內圖書館整合查詢 MetaCat
臺大學術典藏 NTU Scholars
臺大圖書館數位典藏館
本站聲明
© NTU Library All Rights Reserved