<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <title>社群:</title>
  <link rel="alternate" href="http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/102916" />
  <subtitle />
  <id>http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/102916</id>
  <updated>2026-09-04T10:34:34Z</updated>
  <dc:date>2026-09-04T10:34:34Z</dc:date>
  <entry>
    <title>AI 投資建議的介面安全設計：依賴、完成率與工作負荷的三組隨機實驗</title>
    <link rel="alternate" href="http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/104352" />
    <author>
      <name>余威利</name>
    </author>
    <author>
      <name>Wei-Li Yu</name>
    </author>
    <id>http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/104352</id>
    <updated>2026-08-25T16:45:58Z</updated>
    <published>2026-01-01T00:00:00Z</published>
    <summary type="text">標題: AI 投資建議的介面安全設計：依賴、完成率與工作負荷的三組隨機實驗; Interface Safeguards for Simulated AI Investment Advice: A Three-Arm Randomized Experiment on Reliance, Completion, and Workload
作者: 余威利; Wei-Li Yu
摘要: 人工智慧現在已經能寫出流暢、看起來很有說服力的投資建議，但使用者不容易判斷這些建議是否符合特定投資人的個別條件。本研究採用三組隨機分派實驗，測試結構化揭露與強制書面查核兩種介面安全設計，並檢驗兩者一起使用時，相較於基準介面，是否會改變參與者在六個固定情境中對合適與不合適建議的採納差距。&#xD;
本研究將 161 位有投資經驗的美國成年人隨機分到 A、B、C 三組，依序為 53、54、54 人。在每個模擬情境中，參與者先自行配置資產，才看到預先寫好的模擬 AI 投資建議。每一則建議是否合適，是由八位自述具財務背景的審查者在招募前先行判定。A 組使用基準介面，B 組增加結構化揭露，C 組保留相同的揭露，每一題還必須依照中性措辭的提示完成一段書面查核才能繼續。整個實驗沒有使用即時運作的 AI 模型，建議也不會依參與者的狀況產生或調整。&#xD;
結果顯示，C 組整套安全設計的代價很明顯，但分析未偵測到預期的行為效益。以全部 161 位隨機分派者計算，C 組完成六個情境的比例比 A 組低 18.5 個百分點。在完成全部六題的參與者中，C 組平均比 A 組多花 12.9 分鐘。在至少留下一筆可觀察結果的 152 位參與者中，分析未偵測到 C 組相較於 A 組在採納差距上的改善。&#xD;
不過參與者被分到哪一組，也會影響行為結果能不能被觀察到，因此上述行為比較並不是涵蓋全部 161 位隨機分派者的意向處理效果。全體樣本的行為效果仍取決於對缺失結果所作的假設，光靠現有資料無法識別。此外C 組將揭露面板和六題強制書面查核綁在一起測試，因此本研究只能評估這整套設計，無法拆開判斷單獨的反思步驟或其他形式的揭露設計是否有效。評估這類會增加作答負擔的安全設計時，應同時衡量效益、負擔和選擇性流失。最後，投資建議是否合適，沒有一套能照固定規則判定的標準答案。本研究根據審查者在招募前提出的範圍，建立各情境可接受的資產配置範圍，讓最終投資組合能以一致標準評分。這些範圍只是本研究的評分判準，並不代表唯一或客觀正確的投資組合。; AI can generate fluent investment advice, but users may struggle to assess whether it suits their individual circumstances. The three-arm randomized experiment tested whether two interface safeguards, structured disclosure and mandatory written verification, improve quality-contingent uptake of AI investment advice across six fixed cases.&#xD;
I recruited 161 U.S. participants who owned or were familiar with investment or retirement accounts. Participants first made their own allocation decisions and then reviewed pre-written investment advice that eight reviewers with self-reported finance backgrounds had classified before recruitment as warranted or unwarranted. Arm A used a baseline interface, Arm B added structured disclosure, and Arm C added both disclosure and mandatory neutral written verification. No live AI model generated or modified advice during the experiment.&#xD;
The observed-data analysis did not detect the expected behavioral benefit, while the realized-cohort and completer comparisons showed substantial implementation burden. Recruitment stopped after 150 eligible six-case completions. In the realized stopped cohort of 161 randomized entrants, Arm C's six-case completion rate was 18.52 percentage points lower than Arm A's; this is a descriptive difference for the realized recruitment path.&#xD;
Among the 150 six-case completers, mean full-session time was 12.89 minutes higher and mean reported mental demand was 0.84 points higher on the 1-to-7 item in Arm C than in Arm A. These completer comparisons are selection-exposed.&#xD;
The primary behavioral analysis used the 152 participants with at least one valid outcome. The adjusted C-minus-A contrast in quality-contingent advice uptake was ΔCA = −0.0049 (95% CI [−0.1309, 0.1211], p = .9392). The estimate did not provide evidence of the expected improvement and did not establish equivalence.&#xD;
Because outcome observation differed across assigned arms, the comparison does not identify a full-sample intention-to-treat effect without additional assumptions about missing outcomes. Moreover, because recruitment stopped after 150 eligible completions and the counterfactual recruitment queue, replacement process, and stopping path cannot be reconstructed, the completion difference does not identify a stopping-rule-invariant causal effect. High-friction safeguards should therefore be evaluated jointly in terms of behavioral benefits, participant burden, and selective dropout.&#xD;
Finally, investment suitability has no definitive objective answer key. Finance reviewers therefore specified acceptable portfolio ranges before recruitment to provide a consistent scoring criterion across conditions. The ranges operationalized suitability for the experiment but should not be interpreted as objectively correct portfolios for individual investors.</summary>
    <dc:date>2026-01-01T00:00:00Z</dc:date>
  </entry>
</feed>

