arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HoosierHelp:面向社会服务导航的大语言模型智能体基准测试

HoosierHelp: Benchmarking LLM Agents for Social Service Navigation

Yiyang Li, Weixiang Sun, Tianyi Ma, Kaiwen Shi, Zheyuan Zhang, Yanfang Ye

arXiv 2608.09946首次发表:更新:

发表机构

University of Notre Dame(圣母大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出基于印第安纳州3971项公共社会服务资源的交互式基准HoosierHelp,测试发现现有LLM智能体在社会服务导航中对复杂非理想用户交互鲁棒性不足,需更可靠的智能体。

AI 中文摘要

社会服务导航需要将寻求帮助的个体与满足其需求及特定约束的资源相连接。尽管大语言模型(LLM)智能体为对话式资源导航提供了有前景的界面,但现有基准测试未能捕捉该场景的交互复杂性与约束落地需求。我们推出HoosierHelp,这一交互式基准测试基于印第安纳州3971项公共社会服务资源构建。智能体与模拟用户交互,发出结构化资源搜索请求,处理非理想交互,并选择工具返回的最终资源。HoosierHelp通过改变模拟用户的需求结构、约束可满足性及行为模式(包括不耐烦、表述杂乱、无依据请求与自我矛盾)提升模拟用户的真实性。对7种大语言模型的240个样本开展的实验表明,当前LLM智能体在社会服务导航方面仍存在显著不可靠性,在需回退及自我矛盾的对话中性能急剧下降,凸显了对能更鲁棒应对复杂非理想用户交互的智能体的需求。

英文摘要

Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Although LLM agents offer a promising interface for conversational resource navigation, existing benchmarks do not capture the interaction complexity and constraint-grounding demands of this setting. We introduce HoosierHelp, an interactive benchmark grounded in 3,971 Indiana public social service resources. Agents interact with simulated users, issue structured resource-search calls, handle non-ideal interactions, and select the final resources returned by the tool. HoosierHelp enhances the realism of simulated users by varying their need structure, constraint satisfiability, and behavior patterns, including impatience, rambling, unsupported requests, and self-contradiction. Experiments on 240 samples across seven LLMs show that current LLM agents remain substantially unreliable for social service navigation. Performance drops sharply on fallback-required and self-contradictory conversations, highlighting the need for agents that are more robust to complex and non-ideal user interactions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑