arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

International Conference on Learning Representations · 会议 · Machine Learning

2025-12-02 至 2025-12-02 共收录 2
2512.00656 2025-12-02 cs.CL cs.CY

Sycophancy Claims about Language Models: The Missing Human-in-the-Loop

语言模型中的趋炎附势主张:缺失的人工智能循环

Jan Batzner, Volker Stocker, Stefan Schmid, Gjergji Kasneci

机构 * Weizenbaum Institute(韦岑鲍姆研究所) Technical University Berlin(柏林技术大学) Technical University Munich(慕尼黑技术大学)

AI总结 本文探讨了大型语言模型中趋炎附势现象的测量挑战,提出五个核心操作化定义,并指出当前研究缺乏对人类感知的评估,为未来研究提供建议。

Comments NeurIPS 2025 Workshop on LLM Evaluation and ICLR 2025 Workshop on Bi-Directional Human-AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00351 2025-12-02 cs.LG stat.ML

Provable Memory Efficient Self-Play Algorithm for Model-free Reinforcement Learning

可证明的内存高效自博弈算法用于无模型强化学习

Na Li, Yuchen Jiao, Hangguan Shan, Shefeng Yan

机构 * College of Information Science and Electronic Engineering(信息科学与电子工程学院) Zhejiang University(浙江大学) School of Information and Electronics(信息与电子学院) Beijing Institute of Technology(北京理工大学) Institute of Acoustics(声学研究所) Chinese Academy of Sciences(中国科学院)

AI总结 本文提出了一种内存高效的自博弈算法,用于解决多智能体强化学习中的内存效率、样本复杂度和预热成本问题。

Comments ICLR 2024. arXiv admin note: substantial text overlap with arXiv:2110.04645 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏