Skip navigation

DSpace

機構典藏 DSpace 系統致力於保存各式數位資料(如:文字、圖片、PDF)並使其易於取用。

點此認識 DSpace
DSpace logo
English
中文
  • 瀏覽論文
    • 校院系所
    • 出版年
    • 作者
    • 標題
    • 關鍵字
    • 指導教授
  • 搜尋 TDR
  • 授權 Q&A
    • 我的頁面
    • 接受 E-mail 通知
    • 編輯個人資料
  1. NTU Theses and Dissertations Repository
  2. 文學院
  3. 翻譯碩士學位學程
請用此 Handle URI 來引用此文件: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/104027
標題: 從「訊息轉換」到「話語行動」:以戲劇節拍探討人類口譯與 AI 即時翻譯
From Information to Action: Exploring Human Interpreting and AI Live Translation through Dramatic Beats
作者: 王子瑄
Tzu-Hsuan Wang
指導教授: 吳茵茵
Yin-Yin Wu
關鍵字: 戲劇節拍; 行動; 人類口譯; AI 即時字幕; 口譯評估
dramatic beats; action; human interpreting; AI live translation; interpreting assessment
出版年 : 2026
學位: 碩士
摘要: 現有比較人類口譯與 AI 即時翻譯字幕(AI live translation,以下簡稱 AI 即時字幕)表現的研究,多以內容準確性為主要評估標準。此一取向有助於衡量資訊傳遞的完整程度,卻未必足以呈現口譯所涉及的溝通行動。另一方面,有關人類口譯價值的論述常指出,口譯員較能掌握現場氛圍、理解互動脈絡、預判話語走向,並傳達言外之意。這些說法雖具說服力,往往仍停留在描述層次,較難轉化為可操作且可檢視的分析架構。

為補充準確性導向評估所能觀察的面向,本研究引入史坦尼斯拉夫斯基(Stanislavski)的戲劇節拍(dramatic beats)概念,作為比較人類口譯與 AI 即時字幕表現的分析視角。節拍在本研究中作為切分語篇與定位話語行動轉折的操作單位。研究聚焦講者於各節拍中所執行的話語行動,檢視這些行動能否在不同輸出版本中被辨識出來。

研究語料取自 CAPRI 2025 年度論壇的座談場次,包含原文演說、專業人類口譯輸出與 AI 即時字幕輸出三種對應版本。分析首先依話語行動的轉折將原文演說切分為節拍,再以原文節拍為基準,逐一檢視兩種輸出是否保留各節拍中的話語行動。若某一輸出仍能呈現該節拍的行動功能,即判定該話語行動獲得保留。研究並依據此一以原文節拍為基準的分析程序,選取代表性時刻進行細部分析。

本研究首先設定一項方法論目的,即將戲劇節拍轉化為可操作的分析單位,用以追蹤原文演說中的話語行動在不同即時翻譯輸出中是否仍可辨識。研究另提出兩項問題:第一,專業人類口譯與 AI 即時字幕保留原文各節拍話語行動的程度為何;第二,兩種輸出在何種條件下能保留原文節拍中的話語行動,又在何種條件下雖保留字面內容,行動功能卻未保留。

整體而言,專業人類口譯比 AI 即時字幕更穩定地保留原文各節拍中的話語行動,但兩者的差距會隨原文行動的呈現方式而變化。當原文的行動鋪陳明確,且主要由字面形式承載,兩種輸出的表現較為接近;當行動仰賴語氣、發話角色、隱喻結構、聽眾導向,或節拍之間的推進關係時,兩者差異最為明顯。細部分析進一步顯示,兩種輸出呈現不同的行動保留模式,因此不宜簡化為單一的人機優劣之分。人類口譯較能透過重組、壓縮與改寫,保留節拍之間的關係及具有表演性的行動;AI 即時字幕則在字面形式明確且系統正確解析原文時,較能保留由表層形式直接呈現的行動。

透過在準確性之外加入話語行動的分析層次,本研究為人類口譯與 AI 即時字幕的比較提供另一個觀察角度。研究結果顯示,人類口譯的價值之一,可具體呈現在其取捨與重組話語行動的方式上,使講者的話在傳遞資訊的同時,仍能持續推進話語行動。節拍分析因此為既有評估架構較難掌握的人類口譯價值,提供一套更具體且有分析依據的描述方式。
Current comparisons between human interpreting and AI live translation are largely based on accuracy, especially the degree of propositional matching between source and target output. While this approach is useful for assessing informational transfer, it may not fully capture interpreting as a form of communicative action. Arguments for the value of human interpreters point instead to abilities such as reading the room, sensing interactional dynamics, and conveying implied intent. Although these claims are persuasive, they remain difficult to operationalize and test empirically.

This study therefore explores an alternative evaluative perspective by introducing the concept of dramatic beats from theatrical analysis. Drawing on Stanislavski's beat analysis, it treats a speech as a sequence of beats, each representing a stretch of discourse in which the speaker is doing one thing in one way, pursuing a single communicative action by a single means; a new beat begins where that action or means changes. Beats serve here as segmentation units, marking the points where action turns become perceptible so that those turns can be traced across output versions.

The study draws on panel discussion data from the CAPRI 2025 Annual Forum and compares source speeches, professional human interpreting output, and AI live translation output. The source speech is first segmented into beats, and each output is then read against that structure, so that a beat is marked in an output only where the output preserves the action of the corresponding source beat. Representative moments are selected from within that beat structure because they show especially clearly how action turns are preserved or lost across outputs.
This study pursues a methodological purpose and two empirical questions. The methodological purpose is to operationalize dramatic beats as a workable analytical unit for tracing the action of a source speech across interpreting outputs. Building on that unit, the two empirical questions ask to what extent professional human interpreting and AI live translation preserve the action each source beat performs, and under what conditions each output re-performs that action rather than retaining only its wording.

Across the corpus, the professional human interpreting output preserved the action of source beats more consistently than the AI live translation output did. Action preservation is treated here as one dimension of performance, not as a comprehensive measure of translation quality. The size of this difference varied with how the source action was realized. When the action was explicitly signposted and carried largely by surface form, the two outputs were more similar. The difference widened when recoverability depended on tone, voicing, metaphorical framing, audience orientation, or the progression across adjacent beats. Close readings further revealed distinct preservation patterns. Human interpreting more often used restructuring, compression, and reformulation to preserve relations among beats and actions that depended on performance. AI live translation was more likely to preserve actions that were directly encoded in surface form, provided that the source was correctly parsed.

By adding communicative action to accuracy-based evaluation, this study offers a more concrete way of describing one dimension of human interpreting value. The findings show that this value can be observed in how interpreters selectively reorganize discourse so that a speaker’s words continue to advance as action while information is being conveyed.
URI: http://tdr.lib.ntu.edu.tw/jspui/handle/123456789/104027
DOI: 10.6342/NTU202603626
全文授權: 同意授權(全球公開)
電子全文公開日期: 2026-08-22
顯示於系所單位:翻譯碩士學位學程

文件中的檔案:
檔案 大小格式 
ntu-114-2.pdf6.73 MBAdobe PDF檢視/開啟
顯示文件完整紀錄


系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。

社群連結
聯絡資訊
10617臺北市大安區羅斯福路四段1號
No.1 Sec.4, Roosevelt Rd., Taipei, Taiwan, R.O.C. 106
Tel: (02)33662353
Email: ntuetds@ntu.edu.tw
意見箱
相關連結
館藏目錄
國內圖書館整合查詢 MetaCat
臺大學術典藏 NTU Scholars
臺大圖書館數位典藏館
本站聲明
© NTU Library All Rights Reserved