发表机构
Zhejiang University; Ant Group(浙江大学; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对债务催收场景中现有基准无法体现用户行为异质性的问题,提出DebtBench基准并开发DebtGPT代理,用16个先进LLMs实验,发现多数模型表现不佳,DebtGPT优于开源基线且性能与GPT-4o相当。
AI 中文摘要
债务催收是金融行业中的一项关键协商任务,作为一个行为丰富、风险高的以用户为中心的对话系统测试平台,具有很强的实际相关性和特殊的学术价值。虽然大语言模型(LLMs)在对话和协商中显示出了潜力,但在这种复杂场景中有效评估它们的性能仍然是一个重大挑战:现有基准统一假设用户是具有固定偏好的静态、理性主体,未能捕捉到现实世界债务催收中固有的丰富行为异质性。为了弥合这一差距,我们提出了DebtBench,第一个丰富了角色的债务催收基准,突出了协商中的行为异质性。此外,我们开发了DebtGPT,一个经过训练以共同优化财务回收和交互体验的债务催收代理。我们使用16个最先进的大语言模型的实验结果发现,大多数现有模型在这种复杂但现实的场景中表现不佳,而DebtGPT优于所有开源基线,并且达到了与GPT-4o相当的性能。代码和数据可在该https网址获取。
英文摘要
Debt collection is a critical negotiation task in the financial industry, with strong practical relevance and exceptional academic value as a behaviorally rich, high-stakes testbed for human-centered dialogue systems. While large language models (LLMs) have shown promise in dialogue and negotiation, effectively evaluating their performance in this complex scenarios remains a major challenge: existing benchmarks uniformly assume users to be static, rational agents with fixed preferences, failing to capture the rich behavioral heterogeneity inherent in real-world debt collection. To bridge this gap, we propose DebtBench, the first public persona-enriched debt collection benchmark, that highlights behavioral heterogeneity in negotiation. Moreover, we develop DebtGPT, a debt collection agent trained to jointly optimize financial recovery and interaction experience. Our experimental results, using 16 state-of-the-art LLMs, find that most existing models struggle in this complex but realistic scenarios, whereas DebtGPT outperforms all open-source baselines and achieves performance on par with GPT-4o. The code and data are available at https://github.com/YYuHhhh/DebtNegotiation.