Skip navigation

DSpace

機構典藏 DSpace 系統致力於保存各式數位資料(如:文字、圖片、PDF)並使其易於取用。

點此認識 DSpace
DSpace logo
English
中文
  • 瀏覽論文
    • 校院系所
    • 出版年
    • 作者
    • 標題
    • 關鍵字
    • 指導教授
  • 搜尋 TDR
  • 授權 Q&A
    • 我的頁面
    • 接受 E-mail 通知
    • 編輯個人資料
  1. NTU Theses and Dissertations Repository
  2. 國際政經學院
  3. 財金碩士學位學程
請用此 Handle URI 來引用此文件: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/104352
完整後設資料紀錄
DC 欄位值語言
dc.contributor.advisor謝志昇zh_TW
dc.contributor.advisorChih-Sheng Hsiehen
dc.contributor.author余威利zh_TW
dc.contributor.authorWei-Li Yuen
dc.date.accessioned2026-08-25T16:45:54Z-
dc.date.available2026-08-26-
dc.date.copyright2026-08-25-
dc.date.issued2026-
dc.date.submitted2026-08-21 15:00:40-
dc.identifier.citationBansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., & Weld, D. S. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, Article 81. Association for Computing Machinery. https://doi.org/10.1145/3411764.3445717
Bianchi, M., & Brière, M. (2026). Human-robot interactions in investment decisions. Management Science, 72(1), 14–31. https://doi.org/10.1287/mnsc.2022.03886
Bo, J. Y., Wan, S., & Anderson, A. (2025). To rely or not to rely? Evaluating interventions for appropriate reliance on large language models. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Article 905, 1–23. Association for Computing Machinery. https://doi.org/10.1145/3706598.3714097
Bonaccio, S., & Dalal, R. S. (2006). Advice taking and decision-making: An integrative literature review, and implications for the organizational sciences. Organizational Behavior and Human Decision Processes, 101(2), 127–151. https://doi.org/10.1016/j.obhdp.2006.07.001
Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 188. https://doi.org/10.1145/3449287
D’Acunto, F., Prabhala, N., & Rossi, A. G. (2019). The promises and pitfalls of robo-advising. The Review of Financial Studies, 32(5), 1983–2020. https://doi.org/10.1093/rfs/hhz014
Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114–126. https://doi.org/10.1037/xge0000033
Fieberg, C., Hornuf, L., Meiler, M., & Streich, D. J. (2025, January). Using large language models for financial advice (CESifo Working Paper No. 11666). CESifo. https://www.ifo.de/DocDL/cesifo1_wp11666.pdf
Financial Industry Regulatory Authority. (n.d.). 2111. Suitability. Retrieved August 20, 2026, from https://www.finra.org/rules-guidance/rulebooks/finra-rules/2111
Fok, R., & Weld, D. S. (2024). In search of verifiability: Explanations rarely enable complementary performance in AI-advised decision making. AI Magazine, 45(3), 317–332. https://doi.org/10.1002/aaai.12182
Green, P., & MacLeod, C. J. (2016). SIMR: An R package for power analysis of generalized linear mixed models by simulation. Methods in Ecology and Evolution, 7(4), 493–498. https://doi.org/10.1111/2041-210X.12504
Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65–70.
Kim, S. S. Y., Liao, Q. V., Vorvoreanu, M., Ballard, S., & Wortman Vaughan, J. (2024). “I’m not sure, but…”: Examining the impact of large language models’ uncertainty expression on user reliance and trust. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, 822–835. Association for Computing Machinery. https://doi.org/10.1145/3630106.3658941
Klingbeil, A., Grützner, C., & Schreck, P. (2024). Trust and reliance on AI: An experimental study on the extent and costs of overreliance on AI. Computers in Human Behavior, 160, Article 108352. https://doi.org/10.1016/j.chb.2024.108352
Kohn, S. C., de Visser, E. J., Wiese, E., Lee, Y.-C., & Shaw, T. H. (2021). Measurement of trust in automation: A narrative review and reference guide. Frontiers in Psychology, 12, Article 604977. https://doi.org/10.3389/fpsyg.2021.604977
Lee, D. S. (2009). Training, wages, and sample selection: Estimating sharp bounds on treatment effects. The Review of Economic Studies, 76(3), 1071–1102. https://doi.org/10.1111/j.1467-937X.2009.00536.x
Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50–80. https://doi.org/10.1518/hfes.46.1.50_30392
Li, S., Li, J., Li, H., Zhu, H., & Li, X. (2026). Guided reflection in AI-assisted decision-making: Effects on AI overreliance and decision accuracy. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Article 807, 1–19. Association for Computing Machinery. https://doi.org/10.1145/3772318.3790632
Logg, J. M., Minson, J. A., & Moore, D. A. (2019). Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes, 151, 90–103. https://doi.org/10.1016/j.obhdp.2018.12.005
Meade, A. W., & Craig, S. B. (2012). Identifying careless responses in survey data. Psychological Methods, 17(3), 437–455. https://doi.org/10.1037/a0028085
Naiseh, M., Cemiloglu, D., Dogan, H., & Jiang, N. (2026). How is reliance on AI measured? Mapping and evaluating measurement approaches in human-AI interaction [Working paper]. SSRN. https://doi.org/10.2139/ssrn.6845387
Natali, C., Naiseh, M., Cabitza, F., & Frischmann, B. M. (2025). Better AI with designed friction: Theories, applications and research agenda. In HHAI 2025: Proceedings of the 4th International Conference on Hybrid Human-Artificial Intelligence (Frontiers in Artificial Intelligence and Applications, Vol. 408, pp. 518–521). IOS Press. https://doi.org/10.3233/FAIA250680
National Research Council. (2010). The prevention and treatment of missing data in clinical trials. National Academies Press. https://doi.org/10.17226/12955
National Taiwan University. (2026, June). 國立臺灣大學生成式 AI 使用與研究誠信指引 [Guidelines for the Use of Generative AI and Research Integrity]. https://ori.ntu.edu.tw/news/download/fileSn/59
National Taiwan University Office of Research Integrity. (2026, July 3). 本校已訂定「國立臺灣大學生成式 AI 使用與研究誠信指引」 [NTU has established the Guidelines for the Use of Generative AI and Research Integrity]. National Taiwan University. https://ori.ntu.edu.tw/news/detail/news_sn/108
Newcombe, R. G. (1998). Interval estimation for the difference between independent proportions: Comparison of eleven methods. Statistics in Medicine, 17(8), 873–890. https://doi.org/10.1002/(SICI)1097-0258(19980430)17:8<873::AID-SIM779>3.0.CO;2-I
Niszczota, P., & Abbas, S. (2023). GPT has become financially literate: Insights from financial literacy tests of GPT and a preliminary test of how people use it as a source of advice. Finance Research Letters, 58, Article 104333. https://doi.org/10.1016/j.frl.2023.104333
Oehler, A., & Horn, M. (2024). Does ChatGPT provide better advice than robo-advisors? Finance Research Letters, 60, Article 104898. https://doi.org/10.1016/j.frl.2023.104898
Ospina, R., & Ferrari, S. L. P. (2012). A general class of zero-or-one inflated beta regression models. Computational Statistics & Data Analysis, 56(6), 1609–1623. https://doi.org/10.1016/j.csda.2011.10.005
Parasuraman, R., & Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230–253. https://doi.org/10.1518/001872097778543886
Poursabzi-Sangdeh, F., Goldstein, D. G., Hofman, J. M., Wortman Vaughan, J., & Wallach, H. (2021). Manipulating and measuring model interpretability. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, Article 237. Association for Computing Machinery. https://doi.org/10.1145/3411764.3445315
Raees, M., Khan, V.-J., Lykourentzou, I., & Papangelis, K. (2026). Do people appropriately rely on AI-advice? An analytical review of HCI research on human-AI decision-making. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Article 805, 1–24. Association for Computing Machinery. https://doi.org/10.1145/3772318.3791467
Schemmer, M., Kühl, N., Benz, C., Bartos, A., & Satzger, G. (2023). Appropriate reliance on AI advice: Conceptualization and the effect of explanations. Proceedings of the 28th International Conference on Intelligent User Interfaces, 410–422. Association for Computing Machinery. https://doi.org/10.1145/3581641.3584066
Schoeffer, J., Jakubik, J., Vössing, M., Kühl, N., & Satzger, G. (2025). AI reliance and decision quality: Fundamentals, interdependence, and the effects of interventions. Journal of Artificial Intelligence Research, 82, 471–501. https://doi.org/10.1613/jair.1.15873
Spiegelhalter, D. (2017). Risk and uncertainty communication. Annual Review of Statistics and Its Application, 4(1), 31–60. https://doi.org/10.1146/annurev-statistics-010814-020148
Takayanagi, T., Izumi, K., Sanz-Cruzado, J., McCreadie, R., & Ounis, I. (2025). Are generative AI agents effective personalized financial advisors? Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 286–295. Association for Computing Machinery. https://doi.org/10.1145/3726302.3729897
Tian, H. Y., Amin, H., & Yin, M. (2026). Understanding the effects of AI-assisted critical thinking on human-AI decision making. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Article 810, 1–30. Association for Computing Machinery. https://doi.org/10.1145/3772318.3790785
U.S. Securities and Exchange Commission. (2019, June 5). Regulation Best Interest: The broker-dealer standard of conduct (Release No. 34-86031; codified at 17 C.F.R. § 240.15l-1). https://www.sec.gov/files/rules/final/2019/34-86031.pdf
van der Bles, A. M., van der Linden, S., Freeman, A. L. J., Mitchell, J., Galvao, A. B., Zaval, L., & Spiegelhalter, D. J. (2019). Communicating uncertainty about facts, numbers and science. Royal Society Open Science, 6(5), Article 181870. https://doi.org/10.1098/rsos.181870
van der Bles, A. M., van der Linden, S., Freeman, A. L. J., & Spiegelhalter, D. J. (2020). The effects of communicating uncertainty on public trust in facts and numbers. Proceedings of the National Academy of Sciences, 117(14), 7672–7683. https://doi.org/10.1073/pnas.1913678117
Westfall, J., Kenny, D. A., & Judd, C. M. (2014). Statistical power and optimal design in experiments in which samples of participants respond to samples of stimuli. Journal of Experimental Psychology: General, 143(5), 2020–2045. https://doi.org/10.1037/xge0000014
Winder, P., Hildebrand, C., & Hartmann, J. (2025). Biased echoes: Large language models reinforce investment biases and increase portfolio risks of private investors. PLOS ONE, 20(6), Article e0325459. https://doi.org/10.1371/journal.pone.0325459
Wischnewski, M., Krämer, N., & Müller, E. (2023). Measuring and understanding trust calibrations for automated systems: A survey of the state-of-the-art and future directions. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, Article 755. Association for Computing Machinery. https://doi.org/10.1145/3544548.3581197
Xu, Z., Song, T., & Lee, Y.-C. (2025). Confronting verbalized uncertainty: Understanding how LLM’s verbalized uncertainty influences users in AI-assisted decision-making. International Journal of Human-Computer Studies, 197, Article 103455. https://doi.org/10.1016/j.ijhcs.2025.103455
Yang, C. (L.), Bauer, K., Li, X., & Hinz, O. (2026). My advisor, her AI, and me: Evidence from a field experiment on human-AI collaboration and investment decisions. Management Science, 72(1), 242–264. https://doi.org/10.1287/mnsc.2022.03918
Zhang, Y., Liao, Q. V., & Bellamy, R. K. E. (2020). Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 295–305. Association for Computing Machinery. https://doi.org/10.1145/3351095.3372852
Zhou, K., Hwang, J. D., Ren, X., Dziri, N., Jurafsky, D., & Sap, M. (2025). REL-A.I.: An interaction-centered approach to measuring human-LM reliance. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 11148–11167. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.naacl-long.556
Zhou, K., Hwang, J. D., Ren, X., & Sap, M. (2024). Relying on the unreliable: The impact of language models’ reluctance to express uncertainty. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3623–3643. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.198
-
dc.identifier.urihttp://tdr.lib.ntu.edu.tw/jspui/handle/123456789/104352-
dc.description.abstract人工智慧現在已經能寫出流暢、看起來很有說服力的投資建議,但使用者不容易判斷這些建議是否符合特定投資人的個別條件。本研究採用三組隨機分派實驗,測試結構化揭露與強制書面查核兩種介面安全設計,並檢驗兩者一起使用時,相較於基準介面,是否會改變參與者在六個固定情境中對合適與不合適建議的採納差距。
本研究將 161 位有投資經驗的美國成年人隨機分到 A、B、C 三組,依序為 53、54、54 人。在每個模擬情境中,參與者先自行配置資產,才看到預先寫好的模擬 AI 投資建議。每一則建議是否合適,是由八位自述具財務背景的審查者在招募前先行判定。A 組使用基準介面,B 組增加結構化揭露,C 組保留相同的揭露,每一題還必須依照中性措辭的提示完成一段書面查核才能繼續。整個實驗沒有使用即時運作的 AI 模型,建議也不會依參與者的狀況產生或調整。
結果顯示,C 組整套安全設計的代價很明顯,但分析未偵測到預期的行為效益。以全部 161 位隨機分派者計算,C 組完成六個情境的比例比 A 組低 18.5 個百分點。在完成全部六題的參與者中,C 組平均比 A 組多花 12.9 分鐘。在至少留下一筆可觀察結果的 152 位參與者中,分析未偵測到 C 組相較於 A 組在採納差距上的改善。
不過參與者被分到哪一組,也會影響行為結果能不能被觀察到,因此上述行為比較並不是涵蓋全部 161 位隨機分派者的意向處理效果。全體樣本的行為效果仍取決於對缺失結果所作的假設,光靠現有資料無法識別。此外C 組將揭露面板和六題強制書面查核綁在一起測試,因此本研究只能評估這整套設計,無法拆開判斷單獨的反思步驟或其他形式的揭露設計是否有效。評估這類會增加作答負擔的安全設計時,應同時衡量效益、負擔和選擇性流失。最後,投資建議是否合適,沒有一套能照固定規則判定的標準答案。本研究根據審查者在招募前提出的範圍,建立各情境可接受的資產配置範圍,讓最終投資組合能以一致標準評分。這些範圍只是本研究的評分判準,並不代表唯一或客觀正確的投資組合。
zh_TW
dc.description.abstractAI can generate fluent investment advice, but users may struggle to assess whether it suits their individual circumstances. The three-arm randomized experiment tested whether two interface safeguards, structured disclosure and mandatory written verification, improve quality-contingent uptake of AI investment advice across six fixed cases.
I recruited 161 U.S. participants who owned or were familiar with investment or retirement accounts. Participants first made their own allocation decisions and then reviewed pre-written investment advice that eight reviewers with self-reported finance backgrounds had classified before recruitment as warranted or unwarranted. Arm A used a baseline interface, Arm B added structured disclosure, and Arm C added both disclosure and mandatory neutral written verification. No live AI model generated or modified advice during the experiment.
The observed-data analysis did not detect the expected behavioral benefit, while the realized-cohort and completer comparisons showed substantial implementation burden. Recruitment stopped after 150 eligible six-case completions. In the realized stopped cohort of 161 randomized entrants, Arm C's six-case completion rate was 18.52 percentage points lower than Arm A's; this is a descriptive difference for the realized recruitment path.
Among the 150 six-case completers, mean full-session time was 12.89 minutes higher and mean reported mental demand was 0.84 points higher on the 1-to-7 item in Arm C than in Arm A. These completer comparisons are selection-exposed.
The primary behavioral analysis used the 152 participants with at least one valid outcome. The adjusted C-minus-A contrast in quality-contingent advice uptake was ΔCA = −0.0049 (95% CI [−0.1309, 0.1211], p = .9392). The estimate did not provide evidence of the expected improvement and did not establish equivalence.
Because outcome observation differed across assigned arms, the comparison does not identify a full-sample intention-to-treat effect without additional assumptions about missing outcomes. Moreover, because recruitment stopped after 150 eligible completions and the counterfactual recruitment queue, replacement process, and stopping path cannot be reconstructed, the completion difference does not identify a stopping-rule-invariant causal effect. High-friction safeguards should therefore be evaluated jointly in terms of behavioral benefits, participant burden, and selective dropout.
Finally, investment suitability has no definitive objective answer key. Finance reviewers therefore specified acceptable portfolio ranges before recruitment to provide a consistent scoring criterion across conditions. The ranges operationalized suitability for the experiment but should not be interpreted as objectively correct portfolios for individual investors.
en
dc.description.provenanceSubmitted by admin ntu (admin@lib.ntu.edu.tw) on 2026-08-25T16:45:54Z
No. of bitstreams: 0
en
dc.description.provenanceMade available in DSpace on 2026-08-25T16:45:54Z (GMT). No. of bitstreams: 0en
dc.description.tableofcontentsMaster's Thesis Acceptance Certificate i
Acknowledgements ii
摘要 iii
Abstract iv
Table of Contents vi
List of Figures viii
List of Tables ix
Chapter 1 Introduction 1
1.1 Research Problem 1
1.2 Contributions 3
1.3 Scope and Structure 3
Chapter 2 Literature Review 5
2.1 Trust and Reliance 5
2.2 Measuring Reliance 5
2.3 Interface Safeguards 6
2.4 Automated Investment Advice 8
2.5 Study Positioning 8
Chapter 3 Conceptual Framework and Hypotheses 11
3.1 Core Constructs 11
3.2 Treatment Contrasts 11
3.3 Hypotheses 13
3.4 Alternative Explanations 14
Chapter 4 Method as Implemented 16
4.1 Design and Platform 16
4.2 Participants and Recruitment 17
4.3 Randomization and Repair 18
4.4 Scenarios and Interface Treatments 18
4.5 Procedure 22
4.6 Measures 23
4.7 Reviewer Benchmark 25
4.8 Pilots, Governance, and Analysis Implementation 26
Chapter 5 Analysis Plan and Estimands 29
5.1 Estimands 29
5.2 Primary Model 30
5.3 Secondary and Sensitivity Analyses 32
5.4 Planning and Post-Outcome Analyses 32
Chapter 6 Results 34
6.1 Participant Flow 34
6.2 Sample Characteristics 36
6.3 Data Quality 37
6.4 Primary Results 37
6.5 Behavioral Outcomes 41
6.6 Subjective Outcomes 43
6.7 Implementation Costs 45
6.8 Sensitivity Analyses 48
Chapter 7 Discussion, Limitations, and Conclusion 52
7.1 Burden and Selection 52
7.2 Behavioral Findings 53
7.3 Limitations 54
7.4 Implications and Future Research 55
7.5 Conclusion 56
Declarations 57
References 58
Appendix A. Case Materials 66
A.1 Scoring Basis 66
A.2 Instructions 66
A.3 Case S1: Warranted Drawdown 67
A.4 Case S2: Unwarranted Aggression 69
A.5 Case S3: Warranted Recovery 71
A.6 Case S4: Unwarranted Concentration 73
A.7 Case S5: Warranted Goal Constraints 75
A.8 Case S6: Unwarranted Conservatism 77
Appendix B. Reviewer Benchmark 80
B.1 Reviewer Panel 80
B.2 Review Instrument 81
B.3 Decision Rules and Results 81
B.4 Range Aggregation 82
B.5 Presentation Changes 83
B.6 Documentation 83
Appendix C. Participant Materials 84
C.1 Consent 84
C.2 Debrief 84
Appendix D. Governance and AI Disclosure 85
D.1 Register 85
D.2 Author Contributions and AI-Assistance Disclosure 85
Appendix E. Sensitivity Analyses 86
E.1 Tipping-Point Grid 86
E.2 Lee Bounds 88
Appendix F. Interface Screens 90
-
dc.language.isoen-
dc.subject人工智慧投資建議-
dc.subject依建議品質而異的採納行為-
dc.subject建議依賴-
dc.subject隨機實驗-
dc.subject差異性流失-
dc.subject缺失結果敏感度分析-
dc.subjectAI investment advice-
dc.subjectquality-contingent advice uptake-
dc.subjectadvice reliance-
dc.subjectrandomized experiment-
dc.subjectdifferential attrition-
dc.subjectmissing-outcome sensitivity analysis-
dc.titleAI 投資建議的介面安全設計:依賴、完成率與工作負荷的三組隨機實驗zh_TW
dc.titleInterface Safeguards for Simulated AI Investment Advice: A Three-Arm Randomized Experiment on Reliance, Completion, and Workloaden
dc.typeThesis-
dc.date.schoolyear114-2-
dc.description.degree碩士-
dc.contributor.oralexamcommittee王道一;陳暐zh_TW
dc.contributor.oralexamcommitteeJoseph Tao-yi Wang;Wei Chenen
dc.subject.keyword人工智慧投資建議; 依建議品質而異的採納行為; 建議依賴; 隨機實驗; 差異性流失; 缺失結果敏感度分析zh_TW
dc.subject.keywordAI investment advice; quality-contingent advice uptake; advice reliance; randomized experiment; differential attrition; missing-outcome sensitivity analysisen
dc.relation.page101-
dc.identifier.doi10.6342/NTU202604571-
dc.rights.note未授權-
dc.date.accepted2026-08-21-
dc.contributor.author-college國際政經學院-
dc.contributor.author-dept財金碩士學位學程-
dc.date.embargo-liftN/A-
顯示於系所單位:財金碩士學位學程

文件中的檔案:
檔案 大小格式 
ntu-114-2.pdf
  未授權公開取用
2.5 MBAdobe PDF
顯示文件簡單紀錄


系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。

社群連結
聯絡資訊
10617臺北市大安區羅斯福路四段1號
No.1 Sec.4, Roosevelt Rd., Taipei, Taiwan, R.O.C. 106
Tel: (02)33662353
Email: ntuetds@ntu.edu.tw
意見箱
相關連結
館藏目錄
國內圖書館整合查詢 MetaCat
臺大學術典藏 NTU Scholars
臺大圖書館數位典藏館
本站聲明
© NTU Library All Rights Reserved