請用此 Handle URI 來引用此文件:
http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/104352完整後設資料紀錄
| DC 欄位 | 值 | 語言 |
|---|---|---|
| dc.contributor.advisor | 謝志昇 | zh_TW |
| dc.contributor.advisor | Chih-Sheng Hsieh | en |
| dc.contributor.author | 余威利 | zh_TW |
| dc.contributor.author | Wei-Li Yu | en |
| dc.date.accessioned | 2026-08-25T16:45:54Z | - |
| dc.date.available | 2026-08-26 | - |
| dc.date.copyright | 2026-08-25 | - |
| dc.date.issued | 2026 | - |
| dc.date.submitted | 2026-08-21 15:00:40 | - |
| dc.identifier.citation | Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., & Weld, D. S. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, Article 81. Association for Computing Machinery. https://doi.org/10.1145/3411764.3445717
Bianchi, M., & Brière, M. (2026). Human-robot interactions in investment decisions. Management Science, 72(1), 14–31. https://doi.org/10.1287/mnsc.2022.03886 Bo, J. Y., Wan, S., & Anderson, A. (2025). To rely or not to rely? Evaluating interventions for appropriate reliance on large language models. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Article 905, 1–23. Association for Computing Machinery. https://doi.org/10.1145/3706598.3714097 Bonaccio, S., & Dalal, R. S. (2006). Advice taking and decision-making: An integrative literature review, and implications for the organizational sciences. Organizational Behavior and Human Decision Processes, 101(2), 127–151. https://doi.org/10.1016/j.obhdp.2006.07.001 Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 188. https://doi.org/10.1145/3449287 D’Acunto, F., Prabhala, N., & Rossi, A. G. (2019). The promises and pitfalls of robo-advising. The Review of Financial Studies, 32(5), 1983–2020. https://doi.org/10.1093/rfs/hhz014 Dietvorst, B. J., Simmons, J. P., & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114–126. https://doi.org/10.1037/xge0000033 Fieberg, C., Hornuf, L., Meiler, M., & Streich, D. J. (2025, January). Using large language models for financial advice (CESifo Working Paper No. 11666). CESifo. https://www.ifo.de/DocDL/cesifo1_wp11666.pdf Financial Industry Regulatory Authority. (n.d.). 2111. Suitability. Retrieved August 20, 2026, from https://www.finra.org/rules-guidance/rulebooks/finra-rules/2111 Fok, R., & Weld, D. S. (2024). In search of verifiability: Explanations rarely enable complementary performance in AI-advised decision making. AI Magazine, 45(3), 317–332. https://doi.org/10.1002/aaai.12182 Green, P., & MacLeod, C. J. (2016). SIMR: An R package for power analysis of generalized linear mixed models by simulation. Methods in Ecology and Evolution, 7(4), 493–498. https://doi.org/10.1111/2041-210X.12504 Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65–70. Kim, S. S. Y., Liao, Q. V., Vorvoreanu, M., Ballard, S., & Wortman Vaughan, J. (2024). “I’m not sure, but…”: Examining the impact of large language models’ uncertainty expression on user reliance and trust. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, 822–835. Association for Computing Machinery. https://doi.org/10.1145/3630106.3658941 Klingbeil, A., Grützner, C., & Schreck, P. (2024). Trust and reliance on AI: An experimental study on the extent and costs of overreliance on AI. Computers in Human Behavior, 160, Article 108352. https://doi.org/10.1016/j.chb.2024.108352 Kohn, S. C., de Visser, E. J., Wiese, E., Lee, Y.-C., & Shaw, T. H. (2021). Measurement of trust in automation: A narrative review and reference guide. Frontiers in Psychology, 12, Article 604977. https://doi.org/10.3389/fpsyg.2021.604977 Lee, D. S. (2009). Training, wages, and sample selection: Estimating sharp bounds on treatment effects. The Review of Economic Studies, 76(3), 1071–1102. https://doi.org/10.1111/j.1467-937X.2009.00536.x Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50–80. https://doi.org/10.1518/hfes.46.1.50_30392 Li, S., Li, J., Li, H., Zhu, H., & Li, X. (2026). Guided reflection in AI-assisted decision-making: Effects on AI overreliance and decision accuracy. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Article 807, 1–19. Association for Computing Machinery. https://doi.org/10.1145/3772318.3790632 Logg, J. M., Minson, J. A., & Moore, D. A. (2019). Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes, 151, 90–103. https://doi.org/10.1016/j.obhdp.2018.12.005 Meade, A. W., & Craig, S. B. (2012). Identifying careless responses in survey data. Psychological Methods, 17(3), 437–455. https://doi.org/10.1037/a0028085 Naiseh, M., Cemiloglu, D., Dogan, H., & Jiang, N. (2026). How is reliance on AI measured? Mapping and evaluating measurement approaches in human-AI interaction [Working paper]. SSRN. https://doi.org/10.2139/ssrn.6845387 Natali, C., Naiseh, M., Cabitza, F., & Frischmann, B. M. (2025). Better AI with designed friction: Theories, applications and research agenda. In HHAI 2025: Proceedings of the 4th International Conference on Hybrid Human-Artificial Intelligence (Frontiers in Artificial Intelligence and Applications, Vol. 408, pp. 518–521). IOS Press. https://doi.org/10.3233/FAIA250680 National Research Council. (2010). The prevention and treatment of missing data in clinical trials. National Academies Press. https://doi.org/10.17226/12955 National Taiwan University. (2026, June). 國立臺灣大學生成式 AI 使用與研究誠信指引 [Guidelines for the Use of Generative AI and Research Integrity]. https://ori.ntu.edu.tw/news/download/fileSn/59 National Taiwan University Office of Research Integrity. (2026, July 3). 本校已訂定「國立臺灣大學生成式 AI 使用與研究誠信指引」 [NTU has established the Guidelines for the Use of Generative AI and Research Integrity]. National Taiwan University. https://ori.ntu.edu.tw/news/detail/news_sn/108 Newcombe, R. G. (1998). Interval estimation for the difference between independent proportions: Comparison of eleven methods. Statistics in Medicine, 17(8), 873–890. https://doi.org/10.1002/(SICI)1097-0258(19980430)17:8<873::AID-SIM779>3.0.CO;2-I Niszczota, P., & Abbas, S. (2023). GPT has become financially literate: Insights from financial literacy tests of GPT and a preliminary test of how people use it as a source of advice. Finance Research Letters, 58, Article 104333. https://doi.org/10.1016/j.frl.2023.104333 Oehler, A., & Horn, M. (2024). Does ChatGPT provide better advice than robo-advisors? Finance Research Letters, 60, Article 104898. https://doi.org/10.1016/j.frl.2023.104898 Ospina, R., & Ferrari, S. L. P. (2012). A general class of zero-or-one inflated beta regression models. Computational Statistics & Data Analysis, 56(6), 1609–1623. https://doi.org/10.1016/j.csda.2011.10.005 Parasuraman, R., & Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230–253. https://doi.org/10.1518/001872097778543886 Poursabzi-Sangdeh, F., Goldstein, D. G., Hofman, J. M., Wortman Vaughan, J., & Wallach, H. (2021). Manipulating and measuring model interpretability. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, Article 237. Association for Computing Machinery. https://doi.org/10.1145/3411764.3445315 Raees, M., Khan, V.-J., Lykourentzou, I., & Papangelis, K. (2026). Do people appropriately rely on AI-advice? An analytical review of HCI research on human-AI decision-making. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Article 805, 1–24. Association for Computing Machinery. https://doi.org/10.1145/3772318.3791467 Schemmer, M., Kühl, N., Benz, C., Bartos, A., & Satzger, G. (2023). Appropriate reliance on AI advice: Conceptualization and the effect of explanations. Proceedings of the 28th International Conference on Intelligent User Interfaces, 410–422. Association for Computing Machinery. https://doi.org/10.1145/3581641.3584066 Schoeffer, J., Jakubik, J., Vössing, M., Kühl, N., & Satzger, G. (2025). AI reliance and decision quality: Fundamentals, interdependence, and the effects of interventions. Journal of Artificial Intelligence Research, 82, 471–501. https://doi.org/10.1613/jair.1.15873 Spiegelhalter, D. (2017). Risk and uncertainty communication. Annual Review of Statistics and Its Application, 4(1), 31–60. https://doi.org/10.1146/annurev-statistics-010814-020148 Takayanagi, T., Izumi, K., Sanz-Cruzado, J., McCreadie, R., & Ounis, I. (2025). Are generative AI agents effective personalized financial advisors? Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 286–295. Association for Computing Machinery. https://doi.org/10.1145/3726302.3729897 Tian, H. Y., Amin, H., & Yin, M. (2026). Understanding the effects of AI-assisted critical thinking on human-AI decision making. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Article 810, 1–30. Association for Computing Machinery. https://doi.org/10.1145/3772318.3790785 U.S. Securities and Exchange Commission. (2019, June 5). Regulation Best Interest: The broker-dealer standard of conduct (Release No. 34-86031; codified at 17 C.F.R. § 240.15l-1). https://www.sec.gov/files/rules/final/2019/34-86031.pdf van der Bles, A. M., van der Linden, S., Freeman, A. L. J., Mitchell, J., Galvao, A. B., Zaval, L., & Spiegelhalter, D. J. (2019). Communicating uncertainty about facts, numbers and science. Royal Society Open Science, 6(5), Article 181870. https://doi.org/10.1098/rsos.181870 van der Bles, A. M., van der Linden, S., Freeman, A. L. J., & Spiegelhalter, D. J. (2020). The effects of communicating uncertainty on public trust in facts and numbers. Proceedings of the National Academy of Sciences, 117(14), 7672–7683. https://doi.org/10.1073/pnas.1913678117 Westfall, J., Kenny, D. A., & Judd, C. M. (2014). Statistical power and optimal design in experiments in which samples of participants respond to samples of stimuli. Journal of Experimental Psychology: General, 143(5), 2020–2045. https://doi.org/10.1037/xge0000014 Winder, P., Hildebrand, C., & Hartmann, J. (2025). Biased echoes: Large language models reinforce investment biases and increase portfolio risks of private investors. PLOS ONE, 20(6), Article e0325459. https://doi.org/10.1371/journal.pone.0325459 Wischnewski, M., Krämer, N., & Müller, E. (2023). Measuring and understanding trust calibrations for automated systems: A survey of the state-of-the-art and future directions. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, Article 755. Association for Computing Machinery. https://doi.org/10.1145/3544548.3581197 Xu, Z., Song, T., & Lee, Y.-C. (2025). Confronting verbalized uncertainty: Understanding how LLM’s verbalized uncertainty influences users in AI-assisted decision-making. International Journal of Human-Computer Studies, 197, Article 103455. https://doi.org/10.1016/j.ijhcs.2025.103455 Yang, C. (L.), Bauer, K., Li, X., & Hinz, O. (2026). My advisor, her AI, and me: Evidence from a field experiment on human-AI collaboration and investment decisions. Management Science, 72(1), 242–264. https://doi.org/10.1287/mnsc.2022.03918 Zhang, Y., Liao, Q. V., & Bellamy, R. K. E. (2020). Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, 295–305. Association for Computing Machinery. https://doi.org/10.1145/3351095.3372852 Zhou, K., Hwang, J. D., Ren, X., Dziri, N., Jurafsky, D., & Sap, M. (2025). REL-A.I.: An interaction-centered approach to measuring human-LM reliance. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 11148–11167. Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.naacl-long.556 Zhou, K., Hwang, J. D., Ren, X., & Sap, M. (2024). Relying on the unreliable: The impact of language models’ reluctance to express uncertainty. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3623–3643. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.198 | - |
| dc.identifier.uri | http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/104352 | - |
| dc.description.abstract | 人工智慧現在已經能寫出流暢、看起來很有說服力的投資建議,但使用者不容易判斷這些建議是否符合特定投資人的個別條件。本研究採用三組隨機分派實驗,測試結構化揭露與強制書面查核兩種介面安全設計,並檢驗兩者一起使用時,相較於基準介面,是否會改變參與者在六個固定情境中對合適與不合適建議的採納差距。
本研究將 161 位有投資經驗的美國成年人隨機分到 A、B、C 三組,依序為 53、54、54 人。在每個模擬情境中,參與者先自行配置資產,才看到預先寫好的模擬 AI 投資建議。每一則建議是否合適,是由八位自述具財務背景的審查者在招募前先行判定。A 組使用基準介面,B 組增加結構化揭露,C 組保留相同的揭露,每一題還必須依照中性措辭的提示完成一段書面查核才能繼續。整個實驗沒有使用即時運作的 AI 模型,建議也不會依參與者的狀況產生或調整。 結果顯示,C 組整套安全設計的代價很明顯,但分析未偵測到預期的行為效益。以全部 161 位隨機分派者計算,C 組完成六個情境的比例比 A 組低 18.5 個百分點。在完成全部六題的參與者中,C 組平均比 A 組多花 12.9 分鐘。在至少留下一筆可觀察結果的 152 位參與者中,分析未偵測到 C 組相較於 A 組在採納差距上的改善。 不過參與者被分到哪一組,也會影響行為結果能不能被觀察到,因此上述行為比較並不是涵蓋全部 161 位隨機分派者的意向處理效果。全體樣本的行為效果仍取決於對缺失結果所作的假設,光靠現有資料無法識別。此外C 組將揭露面板和六題強制書面查核綁在一起測試,因此本研究只能評估這整套設計,無法拆開判斷單獨的反思步驟或其他形式的揭露設計是否有效。評估這類會增加作答負擔的安全設計時,應同時衡量效益、負擔和選擇性流失。最後,投資建議是否合適,沒有一套能照固定規則判定的標準答案。本研究根據審查者在招募前提出的範圍,建立各情境可接受的資產配置範圍,讓最終投資組合能以一致標準評分。這些範圍只是本研究的評分判準,並不代表唯一或客觀正確的投資組合。 | zh_TW |
| dc.description.abstract | AI can generate fluent investment advice, but users may struggle to assess whether it suits their individual circumstances. The three-arm randomized experiment tested whether two interface safeguards, structured disclosure and mandatory written verification, improve quality-contingent uptake of AI investment advice across six fixed cases.
I recruited 161 U.S. participants who owned or were familiar with investment or retirement accounts. Participants first made their own allocation decisions and then reviewed pre-written investment advice that eight reviewers with self-reported finance backgrounds had classified before recruitment as warranted or unwarranted. Arm A used a baseline interface, Arm B added structured disclosure, and Arm C added both disclosure and mandatory neutral written verification. No live AI model generated or modified advice during the experiment. The observed-data analysis did not detect the expected behavioral benefit, while the realized-cohort and completer comparisons showed substantial implementation burden. Recruitment stopped after 150 eligible six-case completions. In the realized stopped cohort of 161 randomized entrants, Arm C's six-case completion rate was 18.52 percentage points lower than Arm A's; this is a descriptive difference for the realized recruitment path. Among the 150 six-case completers, mean full-session time was 12.89 minutes higher and mean reported mental demand was 0.84 points higher on the 1-to-7 item in Arm C than in Arm A. These completer comparisons are selection-exposed. The primary behavioral analysis used the 152 participants with at least one valid outcome. The adjusted C-minus-A contrast in quality-contingent advice uptake was ΔCA = −0.0049 (95% CI [−0.1309, 0.1211], p = .9392). The estimate did not provide evidence of the expected improvement and did not establish equivalence. Because outcome observation differed across assigned arms, the comparison does not identify a full-sample intention-to-treat effect without additional assumptions about missing outcomes. Moreover, because recruitment stopped after 150 eligible completions and the counterfactual recruitment queue, replacement process, and stopping path cannot be reconstructed, the completion difference does not identify a stopping-rule-invariant causal effect. High-friction safeguards should therefore be evaluated jointly in terms of behavioral benefits, participant burden, and selective dropout. Finally, investment suitability has no definitive objective answer key. Finance reviewers therefore specified acceptable portfolio ranges before recruitment to provide a consistent scoring criterion across conditions. The ranges operationalized suitability for the experiment but should not be interpreted as objectively correct portfolios for individual investors. | en |
| dc.description.provenance | Submitted by admin ntu (admin@lib.ntu.edu.tw) on 2026-08-25T16:45:54Z No. of bitstreams: 0 | en |
| dc.description.provenance | Made available in DSpace on 2026-08-25T16:45:54Z (GMT). No. of bitstreams: 0 | en |
| dc.description.tableofcontents | Master's Thesis Acceptance Certificate i
Acknowledgements ii 摘要 iii Abstract iv Table of Contents vi List of Figures viii List of Tables ix Chapter 1 Introduction 1 1.1 Research Problem 1 1.2 Contributions 3 1.3 Scope and Structure 3 Chapter 2 Literature Review 5 2.1 Trust and Reliance 5 2.2 Measuring Reliance 5 2.3 Interface Safeguards 6 2.4 Automated Investment Advice 8 2.5 Study Positioning 8 Chapter 3 Conceptual Framework and Hypotheses 11 3.1 Core Constructs 11 3.2 Treatment Contrasts 11 3.3 Hypotheses 13 3.4 Alternative Explanations 14 Chapter 4 Method as Implemented 16 4.1 Design and Platform 16 4.2 Participants and Recruitment 17 4.3 Randomization and Repair 18 4.4 Scenarios and Interface Treatments 18 4.5 Procedure 22 4.6 Measures 23 4.7 Reviewer Benchmark 25 4.8 Pilots, Governance, and Analysis Implementation 26 Chapter 5 Analysis Plan and Estimands 29 5.1 Estimands 29 5.2 Primary Model 30 5.3 Secondary and Sensitivity Analyses 32 5.4 Planning and Post-Outcome Analyses 32 Chapter 6 Results 34 6.1 Participant Flow 34 6.2 Sample Characteristics 36 6.3 Data Quality 37 6.4 Primary Results 37 6.5 Behavioral Outcomes 41 6.6 Subjective Outcomes 43 6.7 Implementation Costs 45 6.8 Sensitivity Analyses 48 Chapter 7 Discussion, Limitations, and Conclusion 52 7.1 Burden and Selection 52 7.2 Behavioral Findings 53 7.3 Limitations 54 7.4 Implications and Future Research 55 7.5 Conclusion 56 Declarations 57 References 58 Appendix A. Case Materials 66 A.1 Scoring Basis 66 A.2 Instructions 66 A.3 Case S1: Warranted Drawdown 67 A.4 Case S2: Unwarranted Aggression 69 A.5 Case S3: Warranted Recovery 71 A.6 Case S4: Unwarranted Concentration 73 A.7 Case S5: Warranted Goal Constraints 75 A.8 Case S6: Unwarranted Conservatism 77 Appendix B. Reviewer Benchmark 80 B.1 Reviewer Panel 80 B.2 Review Instrument 81 B.3 Decision Rules and Results 81 B.4 Range Aggregation 82 B.5 Presentation Changes 83 B.6 Documentation 83 Appendix C. Participant Materials 84 C.1 Consent 84 C.2 Debrief 84 Appendix D. Governance and AI Disclosure 85 D.1 Register 85 D.2 Author Contributions and AI-Assistance Disclosure 85 Appendix E. Sensitivity Analyses 86 E.1 Tipping-Point Grid 86 E.2 Lee Bounds 88 Appendix F. Interface Screens 90 | - |
| dc.language.iso | en | - |
| dc.subject | 人工智慧投資建議 | - |
| dc.subject | 依建議品質而異的採納行為 | - |
| dc.subject | 建議依賴 | - |
| dc.subject | 隨機實驗 | - |
| dc.subject | 差異性流失 | - |
| dc.subject | 缺失結果敏感度分析 | - |
| dc.subject | AI investment advice | - |
| dc.subject | quality-contingent advice uptake | - |
| dc.subject | advice reliance | - |
| dc.subject | randomized experiment | - |
| dc.subject | differential attrition | - |
| dc.subject | missing-outcome sensitivity analysis | - |
| dc.title | AI 投資建議的介面安全設計:依賴、完成率與工作負荷的三組隨機實驗 | zh_TW |
| dc.title | Interface Safeguards for Simulated AI Investment Advice: A Three-Arm Randomized Experiment on Reliance, Completion, and Workload | en |
| dc.type | Thesis | - |
| dc.date.schoolyear | 114-2 | - |
| dc.description.degree | 碩士 | - |
| dc.contributor.oralexamcommittee | 王道一;陳暐 | zh_TW |
| dc.contributor.oralexamcommittee | Joseph Tao-yi Wang;Wei Chen | en |
| dc.subject.keyword | 人工智慧投資建議; 依建議品質而異的採納行為; 建議依賴; 隨機實驗; 差異性流失; 缺失結果敏感度分析 | zh_TW |
| dc.subject.keyword | AI investment advice; quality-contingent advice uptake; advice reliance; randomized experiment; differential attrition; missing-outcome sensitivity analysis | en |
| dc.relation.page | 101 | - |
| dc.identifier.doi | 10.6342/NTU202604571 | - |
| dc.rights.note | 未授權 | - |
| dc.date.accepted | 2026-08-21 | - |
| dc.contributor.author-college | 國際政經學院 | - |
| dc.contributor.author-dept | 財金碩士學位學程 | - |
| dc.date.embargo-lift | N/A | - |
| 顯示於系所單位: | 財金碩士學位學程 | |
文件中的檔案:
| 檔案 | 大小 | 格式 | |
|---|---|---|---|
| ntu-114-2.pdf 未授權公開取用 | 2.5 MB | Adobe PDF |
系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。
