发表机构
Lehigh University(里海大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对AI翻译中可读性与源保留分歧导致的可评估性差距,通过TransLingo开展2×2对比实验,揭示了源文本条件与输出呈现方式对感知质量等的影响,明确了任务绩效信任与披露意愿的关联。
AI 中文摘要
可读的AI输出会留下可评估性差距:即使展示了源文本,整体质量判断也可能无法反映输出保留了什么内容。我们研究了源文本条件和输出呈现方式如何与感知翻译质量相关,以及输出和系统评估如何与纯文本界面中的信任和声明披露意愿相关。我们使用TransLingo开展了一项重点2×2比较研究(N=306),考察了简单生成叙事和复杂文学哲学散文,以及LLM生成的面向可读性的输出和研究者修订的面向保真度的输出。描述性刺激审计显示,在两种源文本条件下,面向保真度的输出都保留了更多源内容。因子分析显示,感知质量中存在显著的呈现方式与源文本条件的交互作用。对于简单叙事,参与者将面向保真度的输出评为高于面向可读性的输出;而对于复杂散文,未出现可靠的呈现方式差异。在感知智力、面向能动性的拟人化归因和任务绩效信任方面,观察到了相应的源条件依赖模式。一项独立的理论有序评估结构SEM表征了六个领域中感知质量、感知智力、面向能动性的拟人化归因、任务绩效信任与声明披露意愿之间的并发关联,其中任务绩效信任是声明意愿的近端关联因素。观察到的评分模式将源访问与源可评估性区分开来:对于复杂刺激,展示源文本并不能确保整体质量评分反映保留内容的差异。它还将对翻译输出评估的支持与对个人文本委托给系统的决策的数据处理支持区分开来。
英文摘要
Readable AI output can leave an evaluability gap: even when the source is shown, an overall-quality judgment may not reflect what an output preserves. We investigated how source-text condition and output rendering relate to perceived translation quality, and how output and system appraisals relate to trust and stated disclosure willingness in a plain-text interface. A focal 2 * 2 comparison (N=306) using TransLingo examined simple generated narratives and complex literary-philosophical prose alongside LLM-generated readability-oriented outputs and researcher-revised fidelity-oriented outputs. A descriptive stimulus audit indicated greater source retention in fidelity-oriented outputs in both source-text conditions. Factorial analyses showed a significant rendering-by-source-text-condition interaction in perceived quality. Participants rated fidelity-oriented outputs higher than readability-oriented outputs for the simple narratives, whereas no reliable rendering difference emerged for the complex prose. A corresponding source-condition-dependent pattern was observed for perceived intelligence, agency-oriented anthropomorphic attribution, and task-performance trust. A separate theory-ordered appraisal-structure SEM characterized concurrent associations among perceived quality, perceived intelligence, agency-oriented anthropomorphic attribution, task-performance trust, and stated disclosure willingness across six domains, with task-performance trust as the proximal correlate of stated willingness. The observed rating pattern distinguishes source access from source evaluability: for the complex stimuli, displaying the source did not ensure that one overall-quality rating reflected differences in retained content. It also separates support for evaluating translation output from data-handling support for decisions about what personal text to entrust to a system.