arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-01-23 至 2026-01-23 共收录 4 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 4 篇

2406.12205 2026-01-23 cs.LG cs.AI cs.IT math.IT math.ST stat.ML stat.TH 82%

On the Exponential Convergence for Offline RLHF with Pairwise Comparisons

关于通过成对比较进行离线RLHF的指数收敛性

Zhirui Chen, Vincent Y. F. Tan

专题命中 偏好对齐 :RLHF(title,abstract);分类 cs.AI、cs.LG;alignment(comments)

AI总结 本文提出了一种在离线RLHF中通过成对比较实现指数收敛的算法RL-LOW,并推导了实例依赖的下界。

Comments Accepted as an oral presentation at AAAI 2026 (AI Alignment Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15120 2026-01-23 cs.AI 57%

Emerging from Ground: Addressing Intent Deviation in Tool-Using Agents via Deriving Real Calls into Virtual Trajectories

从地面崛起:通过推导真实调用为虚拟轨迹来解决工具使用代理中的意图偏差

Qian Xiong, Yuekai Huang, Bo Yang, Yujia Zheng, Tianhao Li, Ziyou Jiang, Zhiyuan Chang, Zhaoyang Li, Huanxiang Feng, Mingyang Li

机构 * Beijing Forestry University(北京林业大学) State Key Laboratory of Complex System Modeling and Simulation Technology(复杂系统建模与仿真技术国家重点实验室) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学) Duke University(杜克大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

AI总结 RISE通过推导真实调用为虚拟轨迹,解决工具使用代理中的意图偏差问题,提升任务完成和意图对齐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14727 2026-01-23 stat.ME 50%

Recent advances in the Bradley--Terry Model: theory, algorithms, and applications

布莱德利-蒂尔模型近期进展:理论、算法与应用

Shuxing Fang, Ruijian Han, Yuanhang Luo, Yiming Xu

专题命中 偏好对齐 :alignment(abstract)

AI总结 本文综述了布莱德利-蒂尔模型的最新理论、算法及应用,探讨了大规模场景下的统计推断和计算方法,并指出了未来研究的方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17915 2026-01-23 cs.MA 50%

DISPATCH -- Decentralized Informed Spatial Planning and Assignment of Tasks for Cooperative Heterogeneous Agents

DISPATCH -- 分布式知情空间规划与任务分配以实现协作异构代理

Yao Liu, Sampad Mohanty, Elizabeth Ondula, Bhaskar Krishnamachari

专题命中 偏好对齐 :alignment(abstract)

AI总结 DISPATCH提出两种算法,通过整合公平性和效率,在分布式环境下实现异构代理任务分配的公平与效率平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏