SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning
SafeMCP:基于环境接地前瞻推理的LLM智能体防御主动功率调节
机构 * Beijing Institute of Technology(北京理工大学) ; Beijing Academy of Artificial Intelligence(北京人工智能研究院) ; Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)
专题命中 复杂问题求解 :reasoning(title,abstract);分类 cs.CL、cs.AI
AI总结 针对LLM智能体因动作空间扩大而面临功率寻求风险,提出SafeMCP服务器端防御插件,通过内部世界模型进行前瞻推理,实现主动工具过滤和即时干预两级防御,在保持智能体效用的同时有效降低风险。
Comments Accepted to the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), Main Conference