arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

行动胜于言语:测量使用工具的智能体的跨语言策略保留率

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram

arXiv 2608.11110首次发表:更新:

发表机构

Microsoft Research India(微软研究院印度分部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究测量使用工具的智能体的跨语言行动策略保留率,发现前沿模型在贪心解码下保留71%-73%策略,参数低于100亿时保留率崩溃,且存在影响测量的混淆因素。

AI 中文摘要

当使用工具的智能体在不同语言中被赋予相同任务时,它是否仍会采取相同步骤?多语言评估很少提出这个问题:它们比较最终答案,却忽略了行动。然而这些行动是成果的体现:它们决定成本和延迟,决定系统如何失败,也是其行为唯一可审计的部分。我们以行动策略为测量对象,涉及8个模型、6个并行基准和41种语言(共238万次 rollout)。朴素测量方法失效:原始轨迹相似度与任何可辩护结论之间存在5个混淆因素,每个因素都能翻转结论:短轨迹得分更高,空轨迹得分完美,无关轨迹有超过一半的时间偶然一致,差距受每个模型的可复现性限制,且同一模型在一种语言中被问两次相同问题会给出不同答案,没有基线。我们消除了所有5个混淆因素,每一次修正都使效应更大。差异证明是结构性的,而非采样噪声:它在每个单元中都能在贪心解码下存活,且随着温度升高保持稳定,即使模型的自一致性降低。以自身可复现性归一化后,4个不同的前沿模型在贪心解码下收敛,每个模型在跨语言中保留71%-73%的行动策略,模型身份仅解释5.7%的方差。在参数规模约100亿以下时,保留率会崩溃,而较小模型之间的排序在很大程度上是偶然基线的产物,我们通过排列而非假设来测量该基线。智能体通过英语路由非英语任务;这种转换具有因果重要性,已通过4个模型的预注册预测得到确认,且模型在被要求时不会放弃这种转换。最后,单个轨迹提取正则表达式(而非模型)造成了多语言失败:两个工作示例使一个模型的测量准确率提升了26倍,而其在可读输出上的准确率几乎没有变化。

英文摘要

When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product: they fix cost and latency, decide how the system fails, and are the only auditable part of its behaviour. We make the action policy the measured object across 8 models, 6 parallel benchmarks and 41 languages (2.38M rollouts). The naive measurement fails: five confounds sit between raw trace similarity and any defensible claim, each able to flip a conclusion. Short traces score higher, empty traces score perfectly, unrelated traces agree by chance over half the time, the gap is capped by each model's reproducibility, and a model asked the same question twice in one language answers differently, leaving no baseline. We remove all five, and every correction makes the effect larger. Divergence proves structural, not sampling noise: it survives greedy decoding in every cell and stays flat as temperature rises, even as models grow less self-consistent. Normalised by their own reproducibility, four very different frontier models converge under greedy decoding, each keeping 71-73% of its action policy across languages, with model identity explaining only 5.7% of the variance. Below roughly 10B parameters it breaks down, and the ordering among smaller models is largely an artifact of a chance floor we measure by permutation rather than assume. Agents route non-English tasks through English; this pivot is causally load-bearing, confirmed by a pre-registered prediction across four models, and models will not abandon it when told to. Finally, a single trace-extraction regex, not the model, manufactured a multilingual failure: two worked examples raise one model's measured accuracy twenty-sixfold while its accuracy on readable outputs barely moves.

CommentsAccepted in COLM 26

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑