arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CallScreenBench:将端侧模型作为电话秘书的基准测试

CallScreenBench: Benchmarking Small Language Models as Phone Secretaries

Jiaqi Gan, Haoyuan Tang, Jamey Z. Liang, Siying Chen, Ankit Raj, Kidus Zewde, Yuchen Zhou, Yuxin Zhang, Simiao Ren

arXiv 2608.01033首次发表:更新:

AI 中文总结

研究人员提出CallScreenBench基准,针对端侧模型作为电话秘书的场景,从五个维度评分,发现分类性能差异是测量假象,脚本退化代理提供基准线,未设通过/失败阈值。

AI 中文摘要

可在手机上运行、量化为几比特的小型语言模型,正日益具备替用户执行任务的能力,使端侧任务自动化成为新的可能,其中一项任务就是接听电话。电话秘书代表用户接听未知的来电,与大多数基准测试评估的智能体不同,它无需完成特定任务,也没有协作用户:来电者掌握目标,可能是对手,且必须在无任何先验信息的情况下从通话起始轮次做出判断。关键不在于任务是否成功,而在于用户是否会认可其代理处理通话的方式。我们提出CallScreenBench,该基准从五个质量维度对这一设置进行评分,每个维度都对应一个用于衡量的反指标,且绝不将其平均为单一数值。我们还报告了一个无工具、无凭证、不调用任何工具的代理的戒备状态概况。在六个端侧模型(参数规模0.6-4B,4比特量化)中,质量随能力提升而提升,但分类能力并非如此,其看似提升的表象是测量的假象。脚本化的退化代理提供了缺失的基准线:校正后,在预先注册的操作点上,分类性能存在差异的模型对数量从15对中的11对降至0对。一个仅挂断电话并重复来电者话语的代理也能获得完美的消息保真度。我们报告了这些基准线击败了我们自己的哪些指标,且未设定通过/失败阈值。

英文摘要

Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting on their user's behalf -- which makes on-device task automation newly plausible. One such task is answering the phone. A phone secretary takes an unknown inbound call on its owner's behalf, and unlike the agents most benchmarks evaluate, it has no cooperative caller-assigned task to complete: the caller holds the goal and may be an adversary, while the secretary must begin deciding how to respond without an oracle. What matters is not task success but whether the owner would endorse how their proxy handled the call. We evaluate only the text-domain conversational decision layer; speech recognition, audio interaction, end-to-end latency, and handset execution are outside scope. We present CallScreenBench, which reports five automated call-and-note measure groups motivated by owner endorsement. Each is paired, where available, with a counter-metric and an uncertainty estimate; no benchmark-wide Q1-Q5 composite or leaderboard score is defined. Three guardedness diagnostics identify candidate cases for a toolless proxy that holds no credentials and calls no tools. Across three model families represented by paired 4-bit checkpoints (0.6-4B), the primary scoring snapshot gives the larger checkpoint higher point estimates on several service, recall, and plausibility measures, while triage discrimination follows a different ordering. Bare scam-side TPR rewards universal suspicion, and pairwise separation changes when legitimate-side false positives are included and across judge snapshots. Scripted degenerate agents expose further floors, including a hangup-and-echo policy with entity recall 1.000. We report quality measures and guardedness channels separately so that a single pass/fail score does not hide their trade-offs.

Comments24 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑