arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Annual Meeting of the Association for Computational Linguistics · 会议 · Natural Language Processing

2026-04-27 至 2026-04-27 共收录 36
2510.21285 2026-04-27 cs.AI cs.CL

When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models

当模型超越安全限制:揭示并缓解大推理模型的自我突破

Yingzhi Mao, Chunkang Zhang, Junxiang Wang, Xinyan Guan, Boxi Cao, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun

机构 * Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所信息处理实验室) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 研究揭示大推理模型存在自我突破现象,提出CoG框架通过分步干预缓解安全问题,平衡安全与推理性能。

Comments ACL 2026. The first two authors contributed equally. The main text is 9 pages, with an appendix of 28 pages. The paper contains 20 figures and 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05038 2026-04-27 cs.CL

Toward Automated Robustness Evaluation of Mathematical Reasoning

迈向数学推理的自动化鲁棒性评估

Yutao Hou, Zeguan Xiao, Fei Yu, Yihan Jiang, Ma Shuguang, Zhaoqian Dai, Hailiang Huang, Yun Chen, Guanhua Chen

机构 * Shanghai University of Finance and Economics(上海金融学院) Ant Group(蚂蚁集团) Southern University of Science and Technology(南方科技大学) MoE Key Laboratory of Interdisciplinary Research of Computation and Economics(计算与经济学交叉学科研究教育部重点实验室)

AI总结 本文提出Math Stress Tester框架,通过多轮重写验证循环生成对抗样本,提升数学推理任务的鲁棒性评估能力,并验证其在非数学任务中的扩展性。

Comments Accepted by Findings of ACL2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20662 2026-04-27 cs.AI

AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage

AutoReproduce:基于论文溯源的自动AI实验复现

Xuanle Zhao, Zilin Sang, Yuxuan Li, Qi Shi, Weilun Zhao, Shuo Wang, Duzhen Zhang, Xu Han, Zhiyuan Liu, Maosong Sun

机构 * Tsinghua University(清华大学) Xidian University(西安电子科技大学) OpenBMB(开放大脑实验室) University of the Chinese Academy of Sciences(中国科学院大学)

AI总结 本文提出AutoReproduce,通过论文溯源系统自动复现实验代码,采用多智能体框架和采样单元测试策略,验证了其在复现精度和执行性能上的优势。

Comments Accepted by ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16994 2026-04-27 cs.LG cs.AI cs.CL

FADE: Why Bad Descriptions Happen to Good Features

FADE:为什么好的特征会遇到糟糕的描述

Bruno Puri, Aakriti Jain, Elena Golimblevskaia, Patrick Kahardipraja, Thomas Wiegand, Wojciech Samek, Sebastian Lapuschkin

机构 * Department of Artificial Intelligence, Fraunhofer Heinrich Hertz Institute(人工智能系,弗劳恩霍夫海因里希·赫兹研究所) Department of Electrical Engineering and Computer Science, Technische Universität Berlin(电气工程与计算机科学系,柏林技术大学) BIFOLD - Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究所) Centre of eXplainable Artificial Intelligence, Technological University Dublin(可解释人工智能中心,都柏林技术大学)

AI总结 FADE提出了一种模型无关的框架,用于自动评估特征与描述的一致性,通过四个指标量化特征与描述之间的不一致原因,揭示了生成特征描述的挑战。

Journal ref In Findings of the Association for Computational Linguistics: ACL 2025, pages 17138-17160, Vienna, Austria. Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00955 2026-04-27 cs.CL

Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search

高效多智能体系统训练中的数据影响导向树搜索

Wentao Shi, Zichun Yu, Fuli Feng, Xiangnan He, Chenyan Xiong

机构 * University of Science and Technology of China(中国科学技术大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出数据影响导向树搜索框架,通过影响评分指导树搜索和数据选择,提升多智能体系统训练效率与效果。

Comments Accepted by ACL 2026 Main;

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10138 2026-04-27 cs.CL

PL-MTEB: Polish Massive Text Embedding Benchmark

PL-MTEB:波兰大规模文本嵌入基准

Rafał Poświata, Sławomir Dadas, Michał Perełkiewicz

机构 * National Information Processing Institute(国家信息处理研究所)

AI总结 本文提出PL-MTEB,一个涵盖波兰语言文本嵌入的综合基准,包含30个不同NLP任务,新增12个波兰任务和两个新数据集,评估30种公开文本嵌入模型并公开结果。

Comments Accepted for ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏