arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35749cs.CL

迈向语言代理中的通信高效社交智能

Towards Communication-Efficient Social Intelligence in Language Agents

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • University of Science and Technology of China(中国科学技术大学)
  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Linxiao Gong, Yijie Xu, Tianfu Wang, Yin Wu, Yili Wang, Xingbo Yao, Huizai Yao, Xilin Xia, Haowen Yang, Hui Xiong

AI总结:

本文提出教师辅助通信训练(TACT),通过表达与策略专家修订及伙伴反馈蒸馏,在减少通信成本的同时提升语言代理的社交目标达成率,在SOTOPIA和AgentSense上验证了有效性。

AI中文摘要:

社交智能语言代理必须进行谈判、协调并解决冲突偏好,同时尊重参与者的时间和注意力。平衡这些需求具有挑战性,因为代理必须传达足够的信息以解决合作伙伴的约束并推进其目标,同时不添加对交互无帮助的词语。在本文中,我们提出了教师辅助通信训练(Teacher-Assisted Communication Training, TACT),以提高社交目标达成率,同时降低通信成本,使与代理的交互更加高效且要求更低。我们首先从行动策略和表达两方面刻画通信效率,其影响不仅限于当前话语,还延伸至合作伙伴的回应及后续交流。我们设计TACT来修订学生生成的行动,通过合作伙伴的回应测试修订,并将有用的反馈提炼给学生。表达专家去除不必要的细节,同时保留预期的行动,而策略专家提出可能更好地解决合作伙伴约束的替代方案。为确定哪种修订有帮助,TACT为每个候选采样合作伙伴回应,并通过平衡局部目标支持与行动令牌成本来选择教师参考。该参考指导学生自身生成前缀上的策略内蒸馏,使学生能够在部署时独立行动。我们在SOTOPIA和AgentSense上评估TACT。在SOTOPIA上,它在All和Hard子集上取得了评估方法中最高的目标达成率,同时使用的目标令牌远少于SFT+SDPO。在AgentSense上,它相比初始学生提高了目标成功率,同时减少了目标令牌和交互消息。

英文摘要:

Socially intelligent language agents must negotiate, coordinate, and resolve conflicting preferences while respecting the time and attention of both participants. Balancing these demands is challenging because agents must convey enough to address a partner's constraints and advance their goals without adding words that do not help the interaction. In this paper, we propose Teacher-Assisted Communication Training (TACT) to improve social goal attainment while reducing communication cost, making interactions with agents more productive and less demanding. We first characterize communication efficiency in terms of action strategy and expression, whose effects extend beyond the current utterance to the partner's response and subsequent exchanges. We design TACT to revise student-generated actions, test the revisions through partner responses, and distill useful feedback into the student. An expression specialist removes unnecessary detail while preserving the intended action, while a strategy specialist proposes alternatives that may better address the partner's constraints. To determine which revision helps, TACT samples a partner response for each candidate and selects a teacher reference by balancing local goal support against action-token cost. That reference guides on-policy distillation on the student's own generation prefixes, allowing the student to act independently at deployment. We evaluate TACT on SOTOPIA and AgentSense. On SOTOPIA, it achieves the highest Goal among the evaluated methods on All and Hard while using substantially fewer target tokens than SFT+SDPO. On AgentSense, it improves goal success over the initial student while reducing target tokens and interaction messages.

↑