发表机构
AmpUp Research(AmpUp Research)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究发现语言模型智能体在处理CRM记录时,会轻信有激励偏差的销售代表断言,即使与公司政策矛盾也批准交易;该问题在七个模型中普遍存在,规模与推理无法抵御,本质是说服而非信息缺失;作者提出诊断方法而非架构改进,通过桶分析、同信息对照和计算步骤对照分离说服与信息缺口,并发布评估工件。
AI 中文摘要
语言模型智能体越来越多地基于客户关系管理(CRM)记录回答问题,例如是否应认定一个销售线索为合格。我们识别出一种更强大的模型也无法解决的失败模式:当上下文中包含来自具有乐观倾向一方的断言时——此处为销售代表,即CRM中记录的证人——模型将该断言视为证据,并批准了公司自身记录认为不可接受的交易。在来自CRMArena-Pro的100个线索资格认定任务中,代表在每次通话中都断言了可接受的时间表,并在76次通话中断言了可接受的预算;在31个此类断言与价格表和安装政策相矛盾的任务中,仅阅读通话记录的模型在31个案例中的29个中批准了交易。这一特征在来自四个提供商的七个模型上一致(被误导的比例为87-97%);规模和显式推理并未提供任何抵抗力。在35个真实失败中,只有3个不涉及断言:失败在于说服,而非信息缺失。我们贡献的是一种诊断方法而非架构:(i)一种将说服与信息缺口分离的桶分析,(ii)一种相同信息对照,表明向模型提供记录会将严格准确率从41降至18,同时提高召回率——精确率崩溃,(iii)一种计算步骤对照,保持提取固定,仅改变谁计算预算和时间表。差异范围从廉价模型上的42个百分点到已正确计算的模型上的2-5个百分点;在最强大的模型上,各分支在置信区间内,因此该模式是一致的方向和健全性属性,而非已证明的性能下限。我们预先指定了一项返回阴性结果的泛化测试,描述了前提条件(输入中精确指定的政策),并发布了所有评估工件。
英文摘要
Language-model agents increasingly answer questions over customer-relationship management (CRM) records, such as whether to qualify a sales lead. We identify a failure mode not addressed by a stronger model: when the context contains an assertion by a party with an incentive toward optimism - here the sales representative, a witness recorded in the CRM - the model treats the assertion as evidence and clears deals the company's own records deem unacceptable. Across 100 lead-qualification tasks from CRMArena-Pro, the representative asserts an acceptable timeline in every call and an acceptable budget in 76; on the 31 tasks where such an assertion contradicts the price list and installation policy, a model reading only the transcript clears the deal in 29 of 31 cases. The signature is consistent across seven models from four providers (misled on 87-97%); scale and explicit reasoning confer no resistance. Only 3 of 35 genuine failures involve no assertion: the failure is persuasion, not missing information. We contribute a diagnostic method rather than an architecture: (i) a bucket analysis that separates persuasion from information gaps, (ii) a same-information control showing that supplying the records to the model lowers strict accuracy from 41 to 18 while raising recall - precision collapses - and (iii) a compute-step control that holds extraction fixed and varies only who computes Budget and Timeline. The margin ranges from 42 points on an inexpensive model to 2-5 points on models that already compute correctly; on the strongest models the arms are within confidence intervals, so the pattern is a consistent direction and a soundness property, not a proved performance floor. We pre-specify a generalization test that returns a negative result, characterize the precondition (a policy exactly specified in the inputs), and release all evaluation artifacts.
Comments9 pages, 4 figures, IEEE conference format. Ancillary files contain the evaluation harness, pre-specifications, and per-run result files