HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human-LLM Collaborative Writing
HarDBench: 面向安全人机协作写作的基于草稿的合著越狱攻击基准
机构 * Korea University(韩国大学) ; Sogang University(ソガン大学)
AI总结 提出HarDBench基准,评估大语言模型在协作写作中面对恶意草稿填充的越狱攻击的鲁棒性,并通过偏好优化实现安全-效用平衡的对齐方法。
Comments ACL 2026 Main
Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 40785-40831 July 2-7, 2026