arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33691cs.CL

智能体何时提供帮助?经典文本及其翻译的嵌入、LLM与智能体对齐

When Do Agents Help? Embedding, LLM and Agentic Alignment of Classical Texts and Their Translations

Máté Metzger

首次发表
浏览论文内容

中文总结 AI 辅助

本研究比较经典文本翻译对齐的七种方法,发现生成式智能体在结构可靠性和缺陷减少上优于嵌入方法,但在长文档分块后优势消失。

中文摘要 AI 辅助

经典文本与其翻译的对齐支持机器翻译、检索和计算研究,但比较对齐工作流的证据较为分散。本研究在巴利语、梵语、密释纳希伯来语和藏语的452篇文本上比较了七个系统,共包含9,833个人类对齐单元:四个嵌入流水线、一个直接LLM调用、一个自主智能体,以及由独立审计员修订的智能体。生成式工作流恢复了参考对应关系的93-94%,而嵌入方法最多仅为77%。上限分析表明,句子边界使得某些参考无法由嵌入流水线表示。各生成式工作流的参考恢复率相似:智能体的优势为0.5个百分点(95%置信区间-0.02至1.17),审计未增加已确立的益处。然而,智能体对所有452篇文本均生成了结构有效的输出,而直接调用仅为437篇。一个盲法三LLM专家组对每个生成式不匹配与源文本和人工参考进行了评估。大多数不匹配被标记为可辩护的编辑性变异;共识重大错误标签仅覆盖0.06-0.14%的单元。专家组标记的智能体残余缺陷显著少于直接调用(0.7%对比1.4%),表明仅参考恢复低估了对齐质量。在在线发布的十篇长巴利语经文中,智能体和经审计的智能体将恢复率从直接调用的71%提高到84%和92%。相同的参考定位块使三者均达到93%。因此,智能体在短段落上提高了结构可靠性并减少了判定缺陷,而其在长文档上的显著恢复优势在分块后消失。在此设置中,独立审计对准备好的段落提供的可测量额外益处甚微。

英文摘要

Classical texts aligned with their translations support machine translation, retrieval and computational research, but evidence comparing alignment workflows is scattered. This study compares seven systems on 452 texts in Pali, Sanskrit, Mishnaic Hebrew and Tibetan, comprising 9,833 human-aligned units: four embedding pipelines, a direct LLM call, an autonomous agent, and the agent revised by an independent auditor. Generative workflows recover 93-94% of reference correspondences, against at most 77% for embeddings. A ceiling analysis shows that sentence boundaries make some references unrepresentable by the embedding pipelines. Reference recovery is similar across generative workflows: the agent's advantage is 0.5 percentage points (95% CI -0.02 to 1.17), and auditing adds no established benefit. Agents nevertheless produce structurally valid output for all 452 texts, against 437 for direct calls. A blinded three-LLM panel assesses every generative mismatch against the source and human reference. Most mismatches are labelled defensible editorial variation; consensus major-error labels cover only 0.06-0.14% of units. The panel labels significantly fewer residual defects for agents than direct calls (0.7% versus 1.4%), suggesting that reference recovery alone understates alignment quality. On ten long Pali discourses taken as published online, agents and audited agents raise recovery from the direct call's 71% to 84% and 92%. Identical reference-located chunks bring all three to 93%. Agents thus improve structural reliability and reduce judged defects on short passages, while their large recovery advantage on long documents disappears after chunking. In this setting, independent auditing offers little measurable additional benefit on prepared passages.

补充信息

↑