arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Annual Meeting of the Association for Computational Linguistics · 会议 · Natural Language Processing

共收录 10304
2601.15625 2026-04-21 cs.LG cs.AI

Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors

通过Fission-GRPO实现鲁棒的工具使用:学习从执行错误中恢复

Zhiwei Zhang, Fei Zhao, Rui Wang, Zezhong Wang, Bin Liang, Jiakang Wang, Yao Hu, Shaosheng Cao, Kam-Fai Wong

机构 * The Chinese University of Hong Kong(香港中文大学) Xiaohongshu Inc.(小红书公司) MoE Key Laboratory of High Confidence Software Technologies(可信软件技术国家重点实验室)

AI总结 本文提出Fission-GRPO框架,通过将执行错误转化为在线纠正监督,提升模型在多轮执行中的错误恢复能力,实验显示其在BFCL v4 Multi-Turn上提升了Qwen3-8B的错误恢复率和整体准确率。

Comments 9 pages, 4 figures, 4 tables. Accepted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13099 2026-04-21 cs.CL

Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs

Alexandria:一个多领域方言阿拉伯语机器翻译数据集,用于文化包容和语言多样性的LLM

Abdellah El Mekki, Samar M. Magdy, Houdaifa Atou, Ruwa AbuHweidi, Baraah Qawasmeh, Omer Nacar, Thikra Al-hibiri, Razan Saadie, Hamzah Alsayadi, Nadia Ghezaiel Hammouda, Alshima Alkhazimi, Aya Hamod, Al-Yas Al-Ghafri, Wesam El-Sayed, Asila Al sharji, Mohamad Ballout, Anas Belfathi, Karim Ghaddar, Serry Sibaee, Alaa Aoun, Areej Asiri, Lina Abureesh, Ahlam Bashiti, Majdal Yousef, Abdulaziz Hafiz, Yehdih Mohamed, Emira Hamedtou, Brakehe Brahim, Rahaf Alhamouri, Youssef Nafea, Aya El Aatar, Walid Al-Dhabyani, Emhemed Hamed, Sara Shatnawi, Fakhraddin Alwajih, Khalid Elkhidir, Ashwag Alasmari, Abdurrahman Gerrio, Omar Alshahri, AbdelRahim A. Elmadany, Ismail Berrada, Amir Azad Adli Alkathiri, Fadi A Zaraket, Mustafa Jarrar, Yahya Mohamed El Hadj, Hassan Alhuzali, Muhammad Abdul-Mageed

机构 * The University of British Columbia(不列颠哥伦比亚大学) Canada Research Chair in NLP and ML(自然语言处理和机器学习研究主席) Mohammed VI Polytechnic University(穆莱·阿卜杜勒阿齐兹国王理工学院) Birzeit University(比尔泽特大学) Western Michigan University(西部密歇根大学) Tuwaiq Academy(图瓦伊克学院) King Khalid University(国王卡利德大学) American University of Beirut(贝鲁特美国大学) Ibb University(伊卜大学) University of Hail(海勒大学) University of Technology and Applied Sciences(技术与应用科学大学) Arab Open University(阿拉伯开放大学) Minia University(米尼亚大学) Nantes University(南特大学) Prince Sultan University(沙特王子大学) Umm Al-Qura University(乌姆·阿勒·卡拉大学) University of Nouakchott(努尔人大学) Fatabyyano(法塔比亚诺) Independent Researcher(独立研究者) Hadhramout University(哈德拉姆大学) Cairo University(开罗大学) Misurata University(米斯拉塔大学) Al-Balqa Applied University(巴勒斯坦应用大学) University of Khartoum(喀土穆大学) Sultan Qaboos Higher Centre for Culture and Science(穆罕默德·本·拉希德·阿勒马克图姆文化与科学高级中心) Arab Center for Research and Policy Studies(阿拉伯研究中心) Hamad Bin Khalifa University(哈马德·本·哈利法大学) Institut Supérieur du Numérique(数字高级学院)

AI总结 Alexandria数据集通过多领域方言阿拉伯语对话数据,提升LLM在不同阿拉伯方言中的翻译能力,揭示现有模型在方言翻译中的挑战。

Comments Accepted to ACL 2026 Main; Project resources will be available here: https://github.com/UBC-NLP/Alexandria

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11886 2026-04-21 cs.CL

Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence

忠实性与安全性:在反事实医学证据下评估LLM行为

Kaijie Mo, Siddhartha Venkatayogi, Chantal Shaib, Ramez Kouzy, Wei Xu, Byron C. Wallace, Junyi Jessy Li

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Northeastern University(东北大学) MD Anderson Cancer Center(MD安德森癌症中心) Georgia Institute of Technology(佐治亚理工学院)

AI总结 本文研究了在反事实医学证据下LLM的行为,构建了MedCounterFact数据集,发现模型在面对危险或不合理证据时仍提供自信回答,表明模型可能过度强调忠实性而忽视安全性。

Comments Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11038 2026-04-21 cs.CL

Budget-Aware Anytime Reasoning with LLM-Synthesized Preference Data

具有预算意识的任意时间推理与LLM合成的偏好数据

Xuanming Zhang, Shwan Ashrafi, Aziza Mirsaidova, Amir H. Rezaeian, Miguel Ballesteros, Lydia B. Chilton, Zhou Yu, Dan Roth

机构 * Columbia University(哥伦比亚大学) Oracle AI

AI总结 研究在有限计算预算下LLM的推理行为,提出任意时间推理框架和Anytime Index指标,通过LLM合成的偏好数据提升推理效率与质量。

Comments ACL 2026 Findings, 13 pages, 3 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08564 2026-04-21 cs.CR

MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization

MASH:通过风格人化规避黑盒AI生成文本检测器

Yongtong Gu, Songze Li, Xia Hu

AI总结 本文提出MASH框架,通过风格迁移技术规避黑盒检测器,实验显示其在6个数据集和5个检测器上优于11种基线方法,攻击成功率达92%。

Comments Accepted to Findings of the Association for Computational Linguistics (ACL 2026). 21 pages. Code is available at: https://github.com/githigher/MASH

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06767 2026-04-21 cs.CL cs.AI cs.LG

GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO

GanitLLM:通过课程-GRPO实现难度感知的孟加拉数学推理

Shubhashis Roy Dipta, Khairul Mahbub, Nadia Najjar

机构 * University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校) University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)

AI总结 本文提出GanitLLM模型及新的难度感知孟加拉数学语料库和基于课程的GRPO流程,提升低资源环境下孟加拉数学推理能力。

Comments Accepted at ACL 2026 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06328 2026-04-21 cs.AI

C-World: A Computer Use Agent Environment Creator

C-World:一个计算机使用代理环境创建者

Ziqiao Xi, Shuang Liang, Qi Liu, Jiaqing Zhang, Letian Peng, Fang Nan, Meshal Nayim, Tianhui Zhang, Rishika Mundada, Lianhui Qin, Biwei Huang, Kun Zhou

机构 * UC San Diego(UC圣迭戈大学)

AI总结 C-World通过四个组件构建代理环境,解决LLM代理在规划和推理中与人类的差距问题,展示其作为评估环境和数据引擎的双重价值。

Comments Submitted to ACL 2026 12 pages, 4 figures Ziqiao Xi and Shuang Liang contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05654 2026-04-21 cs.CL cs.AI

Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction

学习检索用户历史并生成用户画像以实现个性化说服预测

Sejun Park, Yoonah Park, Jongwon Lim, Yohan Jo

机构 * Graduate School of Data Science, Seoul National University(首尔国立大学数据科学研究生院)

AI总结 本文提出一种上下文感知的用户画像框架,通过生成查询和总结记录来提升说服预测模型性能,实验表明在Reddit数据集上显著提高预测精度。

Comments This paper has been accepted for publication at Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05062 2026-04-21 cs.CL cs.AI cs.LG

Compositional Steering of Large Language Models with Steering Tokens

通过引导标记对大型语言模型进行组合引导

Gorjan Radevski, Kiril Gashteovski, Giwon Hong, Carolin Lawrence, Goran Glavaš

机构 * NEC Laboratories Europe(NEC欧洲实验室) University of Edinburgh(爱丁堡大学) Center for Artificial Intelligence and Data Science(人工智能与数据科学中心) University of Würzburg(乌尔姆大学) CAIR, Ss. Cyril and Methodius University of Skopje(CAIR,斯·西里尔和方法ius大学)

AI总结 本文提出组合引导标记,用于多行为引导,通过自蒸馏将自然语言指令嵌入专用标记,实现更有效的零样本组合,实验表明其在可验证约束上的表现优于其他方法。

Comments Accepted at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04744 2026-04-21 cs.SD cs.AI

Semi-Supervised Diseased Detection from Speech Dialogues with Multi-Level Data Modeling

基于多级数据建模的语音对话中半监督疾病检测

Xingyuan Li, Mengyue Wu

机构 * X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China(X-LANCE实验室、计算机科学学院、上海交通大学、上海、中国)

AI总结 本文提出一种音频-only半监督学习框架,通过多粒度特征聚合提升医疗语音分析中弱监督学习的效果,实验表明其在少量标注数据下能取得接近全监督性能的成果。

Comments Accepted for publication as a Findings paper at the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04278 2026-04-21 cs.CL cs.AI cs.CR cs.LG

From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning

从领域到实例:双粒度数据合成用于大语言模型去学习

Xiaoyu Xu, Minxin Du, Zitong Li, Zi Liang, Zhibiao Guo, Shiyu Zhang, Peizhao Hu, Qingqing Ye, Haibo Hu

机构 * The Hong Kong Polytechnic University(香港理工大学) Huawei Technologies(华为技术)

AI总结 本文提出BiForget框架,通过双粒度数据合成提升大语言模型去学习的效率与效果,实验显示其在相关性、多样性和效率方面表现优异。

Comments ACL 2026 (Findings), accepted to appear

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03559 2026-04-21 cs.CL

DiffCoT: Diffusion-styled Chain-of-Thought Reasoning in LLMs

DiffCoT:在大语言模型中采用扩散风格的推理链推理

Shidong Cao, Hongzhan Lin, Yuxuan Gu, Ziyang Luo, Jing Ma

机构 * Hong Kong Baptist University(香港 Baptist 大学) Harbin Institute of Technology(哈尔滨工业大学)

AI总结 DiffCoT通过扩散风格的推理链框架,改进了大语言模型在多步数学问题解决中的鲁棒性和错误纠正能力,通过迭代去噪过程实现中间步骤的统一生成与回顾性修正。

Comments DiffCoT improves multi-step LLM reasoning by applying diffusion-based iterative denoising to correct intermediate Chain-of-Thought steps

Journal ref The 64th Annual Meeting of the Association for Computational Linguistics 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17408 2026-04-21 cs.AI cs.LG

The Impact of Off-Policy Training Data on Probe Generalisation

偏离策略训练数据对探测泛化的影响

Nathalie Kirch, Samuel Dower, Adrians Skapars, Helen Yannakoudakis, Ekdeep Singh Lubana, Dmitrii Krasheninnikov

机构 * King’s College London(伦敦国王学院) Imperial College London(伦敦帝国学院) LASR Labs(LASR实验室) University of Manchester(曼彻斯特大学) Goodfire AI University of Cambridge(剑桥大学)

AI总结 本文研究了偏离策略训练数据对八个不同LLM行为探测泛化的影响,发现训练策略显著影响探测性能,尤其在涉及响应意图的行为上表现更差,并提出预测泛化失败的测试方法。

Comments 10 pages, ACL 2026 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05993 2026-04-21 cs.CL cs.AI cs.LG

Revisiting Entropy in Reinforcement Learning for Large Reasoning Models

重新审视强化学习在大推理模型中的熵

Renren Jin, Pengzhi Gao, Yuqi Ren, Zhuowen Han, Tongxuan Zhang, Wuwei Huang, Wei Liu, Jian Luan, Deyi Xiong

机构 * TJUNLP Lab, School of Computer Science and Technology, Tianjin University, China(天津大学计算机科学与技术学院 TJUNLP 实验室,中国) Independent Researcher(独立研究者) College of Computer and Information Engineering, Tianjin Normal University, China(天津师范大学计算机与信息工程学院,中国)

AI总结 本文研究了强化学习可验证奖励中大语言模型熵崩溃问题,分析了熵与响应多样性、校准及性能的关系,提出正优势重加权方法以调节熵并保持性能。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19028 2026-04-21 cs.CL

Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues

他们是恋人还是朋友?评估LLMs在英语和韩语对话中的社交推理能力

Eunsu Kim, Junyeong Park, Juhyun Oh, Kiwoong Park, Seyoung Song, A. Seza Doğruöz, Alice Oh, Najoung Kim

机构 * KAIST(韩国科学技术院) LT3 IDLab(ID实验室) Universiteit Gent(根特大学) Boston University(波士顿大学)

AI总结 本文通过SCRIPTS数据集评估LLMs在英语和韩语对话中推断人物关系的能力,发现当前模型在英语和韩语中的准确率分别为75-80%和58-69%,且存在显著的社交推理局限性。

Comments Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17795 2026-04-21 cs.CL cs.AI cs.LG cs.MA cs.SE

What Makes AI Research Replicable? Executable Knowledge Graphs as Scientific Knowledge Representations

什么使AI研究可重复?可执行知识图谱作为科学知识表示

Yujie Luo, Zhuoyun Yu, Xuehai Wang, Yuqi Zhu, Ningyu Zhang, Lanning Wei, Lun Du, Da Zheng, Huajun Chen

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团) Zhejiang University - Ant Group Joint Laboratory of Knowledge Graph(知识图谱联合实验室)

AI总结 本文提出可执行知识图谱(xKG),通过整合文献中的代码片段和技术洞察,提升AI研究的可重复性,实验显示其在PaperBench上性能提升显著。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09354 2026-04-21 cs.CL

Logit Arithmetic Elicits Long Reasoning Capabilities Without Training

逻辑运算激发长推理能力而不需训练

Yunxiang Zhang, Muhammad Khalifa, Lechen Zhang, Xin Liu, Ayoung Lee, Xinliang Frederick Zhang, Farima Fatahi Bayat, Lu Wang

机构 * University of Michigan(密歇根大学)

AI总结 本文提出ThinkLogit方法,通过逻辑运算将小型推理引导器的能力转移至大模型,无需训练即可提升推理性能,实验显示在六个基准测试中实现21.5%-24.2%的提升。

Comments Accepted to ACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07745 2026-04-21 cs.CL cs.AI cs.LG

Parallel Test-Time Scaling for Latent Reasoning Models

潜在推理模型的并行测试时缩放

Runyang You, Yongqi Li, Meng Liu, Wenjie Wang, Liqiang Nie, Wenjie Li

机构 * Hong Kong Polytechnic University(香港理工大学) Shandong Jianzhu University(山东建筑大学) University of Science and Technology of China(中国科学技术大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

AI总结 本文提出通过改进采样和聚合方法,使潜在推理模型能有效利用并行测试时缩放,提升连续空间中的推理能力。

Comments Accepted at ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07248 2026-04-21 cs.CL

Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models

不要为小型语言模型适应工具;而是将工具模式适应模型

Jonggeun Lee, Woojung Song, Jongwook Han, Haesung Pyun, Yohan Jo

机构 * Graduate School of Data Science, Seoul National University(数据科学研究生院,首尔国立大学)

AI总结 本文提出PA-Tool方法,通过调整工具模式以匹配预训练知识,提升小型语言模型在工具使用任务中的准确性,实验显示性能提升达17%。

Comments Accepted at ACL 2026 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24328 2026-04-21 cs.CL

Speculative Verification: Exploiting Information Gain to Refine Speculative Decoding

推测验证:利用信息增益来细化推测解码

Sungkyun Kim, Jaemin Kim, Dogyung Yoon, Jiho Shin, Junyeol Lee, Jiwon Seo

机构 * Hanyang University(翰阳大学) Seoul National University(首尔国立大学) KAIST(韩国科学技术院) Machine Learning Systems Lab(机器学习系统实验室)

AI总结 本文提出推测验证(SV),通过动态预测推测准确性并调整验证长度来提升推测解码效率,实验证明SV在多个NLP任务中显著优于SD和标准解码。

Comments 16 pages, 8 figures, accepted to ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16538 2026-04-21 cs.CV cs.CL

VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis

VC-Inspector:通过事实分析推进无参考视频字幕评估

Shubhashis Roy Dipta, Tz-Ying Wu, Subarna Tripathi

机构 * University of Maryland, Baltimore County(马里兰大学巴尔的摩分校) Intel(英特尔)

AI总结 VC-Inspector提出了一种轻量级开源大多模态模型,用于无参考视频字幕评估,专注于事实准确性。通过生成可控制事实错误的字幕及评分注释,实现了可重复和事实感知的评估,实验显示其与人类判断的高相关性。

Comments Accepted at ACL 2026 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17458 2026-04-21 cs.CL

Evaluating the Impact of Verbal Multiword Expressions on Machine Translation

评估口语多词表达对机器翻译的影响

Linfeng Liu, Saptarshi Ghosh, Tianyu Jiang

机构 * University of Cincinnati(辛辛那提大学)

AI总结 研究分析了三种口语多词表达类型对英译多语言机器翻译质量的影响,发现其降低翻译质量主要归因于多词表达本身而非句子层面难度。

Comments ACL 2026, 29 pages, 10 figures, Code URL: https://github.com/cincynlp/vmwe-mt-eval

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01486 2026-04-21 cs.CL

Human-Centered Supervision for Sentiment Analysis in Telugu: A Systematic Inquiry Beyond Accuracy

面向泰卢格语情感分析的人本监督:超越准确性的系统探究

Vallabhaneni Raj Kumar, Ashwin S, Supriya Manna, Niladri Sett, Cheedella V S N M S Hema Harshitha, Kurakula Harshitha, Anand Kumar Sharma, Basina Deepakraj, Tanuj Sarkar, Bondada Navaneeth Krishna, Samanthapudi Shakeer

机构 * SRM University AP, India(印度AP SRM大学) University of Maryland, College Park, USA(美国马里兰大学 College Park分校) GITAM University, Bangalore, India(印度班加罗尔GITAM大学)

AI总结 本文提出TeSent数据集,通过人类选择的解释进行模型对齐,探讨在低资源语言中监督学习对预测性能和公平性的影响。

Comments Camera-ready version; ACL Findings, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21934 2026-04-21 cs.CL cs.AI cs.CY cs.IR cs.LG

Culinary Crossroads: A RAG Framework for Enhancing Diversity in Cross-Cultural Recipe Adaptation

烹饪交汇点:一种增强跨文化食谱适应多样性的RAG框架

Tianyi Hu, Andrea Morales-Garzón, Jingyi Zheng, Maria Maistro, Daniel Hershcovich

机构 * Aarhus University(奥胡斯大学) University of Copenhagen(哥本哈根大学) Dept. of Computer Science and Artificial Intelligence, University of Granada(格拉纳达大学计算机科学与人工智能系) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 本文提出CARRIAGE框架,通过增强检索与上下文组织来提升跨文化食谱适应的多样性,解决了RAG在生成多样化结果时的局限性。

Comments ACL 2026 (main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18141 2026-04-21 cs.CL cs.AI

Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language Models

稀疏特征共激活揭示大语言模型中的因果语义模块

Ruixuan Deng, Xiaoyang Hu, Miles Gilberti, Shane Storks, Aman Taxali, Mike Angstadt, Chandra Sripada, Joyce Chai

机构 * Georgia Institute of Technology(佐治亚理工学院) Brown University(布朗大学) University of Michigan(密歇根大学)

AI总结 通过稀疏自编码器特征的共激活,研究发现大语言模型中存在语义连贯且上下文一致的网络组件,揭示了模型的模块化结构及高效操控方法。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06485 2026-04-21 cs.CL cs.AI

Task Matters: Knowledge Requirements Shape LLM Responses to Context-Memory Conflict

任务相关性:知识需求影响LLM在上下文记忆冲突中的响应

Kaiser Sun, Fan Bai, Mark Dredze

机构 * Center for Language and Speech Processing, Data Science and AI Institute(语言与语音处理中心,数据科学与人工智能研究所)

AI总结 研究探讨了任务类型如何影响LLM在上下文与记忆冲突中的表现,发现任务对知识的依赖和冲突合理性共同影响性能,提出任务感知方法以平衡上下文与记忆。

Comments ACL2026 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03535 2026-04-21 cs.SE

Across Programming Language Silos: A Study on Cross-Lingual Retrieval-augmented Code Generation

跨越编程语言壁垒:关于跨语言检索增强型代码生成的研讨

Qiming Zhu, Jialun Cao, Xuanang Chen, Weili Zhang, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun, Shing-Chi Cheung

AI总结 本文研究了跨语言检索增强型代码生成中的知识转移,通过构建13种编程语言的14000个实例数据集,发现跨语言知识转移非 trivial,且依赖语言亲和力和预训练语料多样性。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02718 2026-04-21 cs.LG cs.AI

End-to-End Optimization of LLM-Driven Multi-Agent Search Systems via Heterogeneous-Group-Based Reinforcement Learning

通过异质组强化学习实现LLM驱动的多智能体搜索系统的端到端优化

Guanzhong Chen, Shaoxiong Yang, Chao Li, Wei Liu, Jian Luan, Zenglin Xu

机构 * MiLM Plus, Xiaomi Inc.(小米公司MiLM Plus实验室) Fudan University(复旦大学) Shanghai Academy of AI for Science(上海人工智能科学研究院)

AI总结 本文提出MHGPO方法,通过估计异质组间的相对优势,优化多智能体系统的整体成功而非单个智能体性能,提升了任务表现和计算效率。

Comments Accepted to ACL 2026 Main Conference. 20 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18128 2026-04-21 cs.CL

Frankentext: Stitching random text fragments into long-form narratives

Frankentext:将随机文本片段拼接成长篇叙述

Chau Minh Pham, Jenna Russell, Dzung Pham, Mohit Iyyer

机构 * University of Maryland, College Park(马里兰大学学院公园分校) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

AI总结 Frankentext通过将大量随机文本片段拼接成连贯故事,挑战LLM生成质量与原创性,同时引发作者权属问题。

Comments Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15404 2026-04-21 cs.CL

How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study

如何增强大推理模型的安全性:一项实证研究

Zhexin Zhang, Xian Qi Loye, Victor Shea-Jay Huang, Junxiao Yang, Qi Zhu, Shiyao Cui, Fei Mi, Lifeng Shang, Yingkang Wang, Hongning Wang, Minlie Huang

机构 * The Conversational AI (CoAI) group, DCST, Tsinghua University(清华大学对话人工智能(CoAI)小组,DCST,清华大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

AI总结 本文通过监督微调探讨如何提升大推理模型的安全性,发现直接蒸馏安全响应无效,但针对性处理危险模式可显著提升安全性,且短或模板化推理过程同样有效。

Comments ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏