arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CredLeakBench:评估LLM代理中的凭证泄露与恢复

CredLeakBench: Evaluating Credential Leakage and Recovery in LLM Agents

Rafid Ahmed, Joseph Fioresi, Mubarak Shah, Yuzhang Shang

arXiv 2610.08871首次发表:更新:

发表机构

University of Central Florida(中佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CredLeakBench通过全面基准测试评估LLM代理在钓鱼和身份验证场景中的凭证泄露风险,发现所有模型均易受攻击,且现有防御措施在减少泄露时损害真实任务性能,强调需兼顾安全与效用。

AI 中文摘要

语言模型代理越来越多地被部署来自动化日常数字事务,从管理电子邮件和社交媒体到处理银行和账单,使用户能够摆脱监督。然而,这种能力也使敏感信息暴露于网络钓鱼攻击。安全执行要求区分恶意请求与真实请求,而不是简单地拒绝行动。尽管具有实际重要性,但这一问题仍未得到充分探索,目前尚不清楚现有代理或现有防御措施能否实现这一目标。为了研究这一问题,我们首先提出CredLeak-Bench,一个全面的基准测试,旨在评估代理在面对网络钓鱼和身份验证时如何有效且安全地自动化人类工作流程。该基准涵盖用户指导的认证和自主收件箱监控,其中代理未被明确指示登录。它系统地变化欺骗性线索,并将钓鱼场景与合法对应场景配对,从而能够对真实任务上的信息泄露和效用进行联合评估。在沙盒环境中,泄露通过实际提交的信息来衡量,而非代理的自我报告行为。我们的评估显示,所有测试模型都容易受到泄露的影响。代理在自主收件箱监控期间也会披露敏感信息,表明钓鱼可以在没有用户请求认证的情况下诱导披露。此外,大多数评估的缓解措施在减少泄露的同时也损害了真实任务的性能,暴露了现有防御中的安全效用权衡。这些发现表明,仅减少泄露是不够的:有效的防御必须防止未经授权的披露,同时保持合法任务的完成。CredLeak-Bench为衡量这两个目标以及评估安全、有用代理的进展提供了一个受控框架。

英文摘要

Language model agents are increasingly deployed to automate everyday digital chores from managing emails and social media to handling banking and bills allowing users to step away from supervision. However, this capability also exposes sensitive information to phishing. Safe execution requires distinguishing malicious requests from genuine ones without simply refusing to act. Despite its practical importance, this problem remains underexplored and it is unclear whether current agents or existing defenses can achieve it. To study this problem, we first propose CredLeak-Bench, a comprehensive benchmark designed to evaluate how effectively and securely agents automate human workflows when confronted with phishing and identity verification. The benchmark covers both user-directed authentication and autonomous inbox monitoring, where agents are not explicitly instructed to log in. It systematically varies deceptive cues and pairs phishing scenarios with legitimate counterparts, enabling joint evaluation of information leakage and utility on genuine tasks. Within a sandboxed environment, leakage is measured through actual submissions of information rather than agents' self-reported behavior. Our evaluation reveals that all tested models are vulnerable to leakage. Agents also disclose sensitive information during autonomous inbox monitoring, demonstrating that phishing can induce disclosure without a user request to authenticate. Furthermore, most evaluated mitigations that reduce leakage also impair performance on genuine tasks, exposing a security utility trade off in existing defenses. These findings show why reducing leakage alone is insufficient: effective defenses must prevent unauthorized disclosure while preserving legitimate task completion. CredLeak-Bench provides a controlled framework for measuring both objectives and evaluating progress toward secure, useful agents.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑