arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Annual Meeting of the Association for Computational Linguistics · 会议 · Natural Language Processing

共收录 10304
2501.03191 2026-04-28 cs.CL

CLIX: Cross-Lingual Explanations of Idiomatic Expressions

CLIX: 习语的跨语言解释

Aaron Gluck, Katharina von der Wense, Maria Leonor Pacheco

机构 * University of Colorado Boulder(科罗拉多大学博尔德分校) Johannes Gutenberg University Mainz(美因茨约翰尼斯·古腾堡大学)

AI总结 研究提出CLIX任务,探讨NLP模型在生成习语跨语言解释中的能力,发现大语言模型有潜力,但需进一步解决错误分析中的关键挑战。

Comments Accepted to Findings of ACL 2025

Journal ref Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10141 2026-04-28 cs.CV cs.CL cs.MM

ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis

ANCHOR:基于大语言模型的文本到图像合成中的主体条件化

Aashish Anantha Ramakrishnan, Sharon X. Huang, Dongwon Lee

机构 * Optum AI The Pennsylvania State University(宾夕法尼亚州立大学)

AI总结 ANCHOR通过大规模抽象式标题数据集研究文本到图像合成中多主体理解与上下文推理的缺陷,提出基于大语言模型的主体感知微调方法,提升图像-标题一致性与人类偏好对齐。

Comments Accepted to The 64th Annual Meeting of the Association for Computational Linguistics (ACL) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22708 2026-04-27 cs.MA

Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems

看见整个象:面向基于大语言模型的多智能体系统故障归因的基准测试

Mengzhuo Chen, Junjie Wang, Fangwen Mu, Yawen Wang, Zhe Liu, Huanxiang Feng, Qing Wang

AI总结 本文提出TraceElephant基准测试,通过完整执行轨迹和可复现环境提升故障归因准确性,验证完整轨迹对故障原因识别的重要性。

Comments Accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18655 2026-04-27 cs.DC cs.AI cs.CL

Unlocking the Edge deployment and ondevice acceleration of multi-LoRA enabled one-for-all foundational LLM

解锁多LoRA支持的多用途基础大语言模型的边缘部署与设备端加速

Sravanth Kodavanti, Sowmya Vajrala, Srinivas Miriyala, Utsav Tiwari, Uttam Kumar, Utkarsh Kumar Mahawar, Achal Pratap Singh, Arya D, Narendra Mutyala, Vikram Nelvoy Rajendiran, Sharan Kumar Allur, Euntaik Lee, Dohyoung Kim, HyeonSu Lee, Gyusung Cho, JungBae Kim

机构 * Samsung Research Institute(三星研究院) Samsung Electronics(三星电子)

AI总结 本文提出一种硬件感知框架,实现基于LLaMA的多语言基础模型在三星S24/S25设备上的高效设备端推理,通过动态任务切换和多流解码机制,提升内存和延迟性能,验证了多用途大语言模型在边缘设备上的可行性。

Comments Accepted at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18572 2026-04-27 cs.CL

One Persona, Many Cues, Different Results: How Sociodemographic Cues Impact LLM Personalization

一个身份,多种线索,不同结果:社会人口学线索如何影响LLM个性化

Franziska Weeber, Vera Neplenbroek, Jan Batzner, Sebastian Padó

机构 * Institute for Natural Language Processing, University of Stuttgart(斯图加特大学自然语言处理研究所) Institute for Logic, Language and Computation, University of Amsterdam(阿姆斯特丹大学逻辑、语言和计算研究所) Weizenbaum Institute(Weizenbaum研究所;慕尼黑机器学习中心;慕尼黑技术大学) Munich Center for Machine Learning TU Munich

AI总结 研究探讨社会人口学线索对LLM个性化的影响,指出单一线索可能导致偏差,强调需考虑外部有效性。

Comments ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22631 2026-04-27 cs.CL

Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models

在自监督语音识别模型的音素级嵌入中识别和分类人口不公平现象

Felix Herron, Solange Rossato, Alexandre Allauzen, François Portet

机构 * MILES Team, LAMSADE, Université Paris Dauphine-PSL(巴黎-萨克雷大学巴黎第四大学LAMSADE团队) GETALP Team, LIG, Université Grenoble Alpes(格勒诺布尔阿尔卑斯大学LIG团队)

AI总结 本文研究了自监督语音识别模型中音素嵌入的不公平现象,提出两种误差类型:随机误差与系统误差,并发现音素嵌入中的偏差可能影响不同群体的性能。

Journal ref Findings of the Association for Computational Linguistics 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22558 2026-04-27 cs.LG cs.AI

SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning

SOLAR-RL:半在线长 horizon 分配强化学习

Jichao Wang, Liuyang Bian, Yufeng Zhou, Han Xiao, Yue Pan, Guozhi Wang, Hao Wang, Zhaoxiong Wang, Yafei Wen, Xiaoxin Chen, Shuai Ren, Lingfang Zeng

机构 * vivo AI Lab(vivo人工智能实验室) Zhejiang Lab(浙江实验室) CUHK MMLab(香港大学多模态实验室) Hubei University(湖北省大学) Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院)

AI总结 SOLAR-RL通过整合全局轨迹信息到离线学习中,提升GUI任务的长horizon完成率和鲁棒性,提供高效样本利用的自主导航方案。

Comments 14 pages, 11 figures. Accepted to Findings of the Association for Computational Linguistics: ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22520 2026-04-27 cs.CL

RouteLMT: Learned Sample Routing for Hybrid LLM Translation Deployment

RouteLMT: 为混合LLM翻译部署设计的学得样本路由

Yingfeng Luo, Hongyu Liu, Dingyang Lin, Kaiyan Chang, Chenglong Wang, Bei Li, Quan Du, Tong Xiao, Jingbo Zhu

机构 * School of Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程学院,中国沈阳) NiuTrans Research, Shenyang, China(牛译研究院,中国沈阳)

AI总结 本文提出RouteLMT,通过在模型内部预测大模型相对于小模型的边际增益,实现高效的样本路由,优于传统启发式和质量估计方法,达到更优的质量-预算帕累托前沿。

Comments Accepted to ACL 2026 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22517 2026-04-27 cs.CL

Aggregate vs. Personalized Judges in Business Idea Evaluation: Evidence from Expert Disagreement

聚合与个性化评判在商业创意评估中的对比:来自专家分歧的证据

Wataru Hirota, Tomoki Taniguchi, Tomoko Ohkuma, Kosuke Takahashi, Takahiro Omi, Kosuke Arima, Takuto Asakura, Chung-Chi Chen, Tatsuya Ishigaki

机构 * Stockmark Inc(Stockmark公司) Asahi Kasei Corporation(朝日化学株式会社) National Institute of Informatics(日本信息处理学会) National Institute of Advanced Industrial Science and Technology(国家先进工业科学与技术研究院)

AI总结 本文研究了在商业创意评估中,自动评判应接近聚合共识还是个体评判的问题,通过PBIG-DATA数据集发现个性化评判更符合评估者,而聚合评判存在较大分歧。

Comments ACL 2026 Industry Track (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22438 2026-04-27 cs.CR cs.AI cs.CL

SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking

SSG: 用于LLM水印的对数平衡词汇分区

Chenxi Gu, Xiaoning Du, John Grundy

机构 * AllenG-L

AI总结 本文提出SSG方法,通过对数平衡词汇分区提升水印检测效果,针对低熵场景改进水印强度下界。

Comments ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22374 2026-04-27 cs.CL

Selective Contrastive Learning For Gloss Free Sign Language Translation

选择性对比学习用于无 gloss 手语翻译

Changhao Lai, Rui Zhao, Xuewen Zhong, Jinsong Su, Yidong Chen

机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) Key Lab of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian-Taiwan (XMU), Ministry of Culture and Tourism, China(福建省 Fujian-Taiwan 非物质文化遗产数字保护与智能处理重点实验室,文化部,中国) National Language Resources Monitoring and Research Center for Education and Teaching Media, Xiamen University, China(教育与教学媒体语言资源监测与研究中心,厦门大学,中国)

AI总结 本文提出选择性对比学习方法,通过动态选择负样本提升手语翻译的跨模态对齐效果,减少噪声干扰。

Comments Accepted by ACL 2026 as the main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22367 2026-04-27 cs.CL cs.AI

CNSL-bench: Benchmarking the Sign Language Understanding Capabilities of MLLMs on Chinese National Sign Language

CNSL-bench:用于评估多模态大语言模型在中文国家手语理解能力的基准

Rui Zhao, Xuewen Zhong, Xiaoyun Zheng, Jinsong Su, Yidong Chen

机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) Key Lab of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian-Taiwan (XMU), Ministry of Culture and Tourism, China(福建省-台湾非物质文化遗产数字化保护与智能处理重点实验室(XMU),文化和旅游部,中国) National Language Resources Monitoring and Research Center for Education and Teaching Media, Xiamen University, China(教育与教学媒体语言资源监测与研究中心,厦门大学,中国)

AI总结 本文提出CNSL-bench,首个评估多模态大语言模型在中文国家手语理解能力的基准,通过权威 grounding、多模态覆盖和手部动作多样性,评估21个模型,发现当前模型在手语理解上仍显著劣于人类表现。

Comments Accepted as the Main Conference at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22345 2026-04-27 cs.CL

Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization

大型语言模型中的偏好头:可解释个性化机制框架

Weixu Zhang, Ye Yuan, Changjiang Han, Yuxing Tian, Zipeng Sun, Linfeng Du, Jikun Kang, Hong Kang, Xue Liu, Haolun Wu

机构 * McGill University(麦吉尔大学) Mila - Quebec AI Institute(魁北克人工智能研究所) MBZUAI(马克斯·普朗克人工智能研究所) University of Montreal(蒙特利尔大学) Salesforce(Salesforce公司)

AI总结 本文提出DPS框架,通过因果掩码分析识别偏好头并实现可控可解释的个性化,实验表明在保持内容连贯性的同时提升个性化精度。

Comments Accepted at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22335 2026-04-27 cs.CL

Context-Fidelity Boosting: Enhancing Faithful Generation through Watermark-Inspired Decoding

上下文保真度提升:通过水印启发式解码增强忠实生成

Weixu Zhang, Fanghua Ye, Qiang Gao, Jian Li, Haolun Wu, Yuxing Tian, Sijing Duan, Nan Du, Xiaolong Li, Xue Liu

机构 * Hunyuan AI Digital Human, Tencent(腾讯文深AI数字人) McGill University(麦吉尔大学) Mila - Quebec AI Institute(魁北克人工智能研究所) Wuhan University(武汉大学) University of Montreal(蒙特利尔大学) Tsinghua University(清华大学) MBZUAI

AI总结 本文提出CFB框架,通过增加源支持标记的生成概率减少大语言模型中的保真度幻觉,采用基于水印技术的logit调整策略,三种增强策略提升生成忠实度,无需重训练即可兼容多种LLM。

Comments Accepted at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22313 2026-04-27 cs.CL

CLARITY: A Framework and Benchmark for Conversational Language Ambiguity and Unanswerability in Interactive NL2SQL Systems

CLARITY:一种用于交互式NL2SQL系统中对话语言歧义和不可回答性的框架和基准

Tabinda Sarwar, Farhad Moghimifar, Cong Duy Vu Hoang, Xiaoxiao Ma, Shawn Chang Xu, Fahimeh Saleh, Poorya Zaremoodi, Avirup Sil, Katrin Kirchhoff

机构 * Oracle Corporation(Oracle公司)

AI总结 CLARITY框架通过多维度歧义和用户行为生成NL2SQL基准,揭示现有系统在复杂歧义下的性能下降,强调需增强歧义检测与解决能力。

Comments Accepted at ACL 2026 (Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22291 2026-04-27 cs.CR cs.SE

Train in Vain: Functionality-Preserving Poisoning to Prevent Unauthorized Use of Code Datasets

无效训练:功能保持污染以防止未经授权使用代码数据集

Yuan Xiao, Jiaming Wang, Yuchen Chen, Wei Song, Jun Sun, Shiqing Ma, Yanzhou Mu, Juan Zhai, Chunrong Fang, Jin Song Dong, Zhenyu Chen

AI总结 本文提出FunPoison,一种保持功能的污染方法,通过注入可编译的弱使用片段来有效防止未经授权的数据集使用,同时保持100%的编译正确性。

Comments Accepted to Findings of the Association for Computational Linguistics (ACL 2026). Code is available at: https://github.com/xiaoyuanpigo/FunPoison

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22239 2026-04-27 cs.CL cs.AI

Navigating Large-Scale Document Collections: MuDABench for Multi-Document Analytical QA

在大规模文档集合中导航:MuDABench用于多文档分析问答

Zhanli Li, Yixuan Cao, Lvzhou Luo, Ping Luo

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences (CAS)(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Wenlan School of Business, Zhongnan University of Economics and Law(中南财经政法大学文澜商学院)

AI总结 本文提出在大规模半结构化文档集合上进行分析问答的任务,介绍了MuDABench多文档分析问答基准,要求跨多个文档提取和综合信息以进行定量分析,实验发现标准RAG系统表现不佳,提出多代理工作流以提升性能。

Comments Findings of ACL 2026. The camera-ready version corrects some labeling errors. The accompanying repository is continuously updated based on community feedback; for the most up-to-date implementation and results, please refer to the repository

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22209 2026-04-27 eess.AS cs.AI cs.CL cs.SD

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

UniSonate:一种统一的语音、音乐和音效生成模型,通过文本指令控制

Chunyu Qiang, Xiaopeng Wang, Kang Yin, Yuzhe Liang, Yuxin Guo, Teng Ma, Ziyu Zhang, Tianrui Wang, Cheng Gong, Yushen Chen, Ruibo Fu, Chen Zhang, Longbiao Wang, Jianwu Dang

机构 * Tianjin University(天津大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

AI总结 UniSonate通过统一的文本指令接口生成语音、音乐和音效,采用动态令牌注入机制和多阶段课程学习策略,提升跨模态生成性能,实现指令驱动的TTS和TTM的SOTA表现。

Comments Accepted to ACL 2026 main conference (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22193 2026-04-27 cs.CL

How Large Language Models Balance Internal Knowledge with User and Document Assertions

大型语言模型如何在内部知识与用户和文档断言之间取得平衡

Shuowei Li, Haoxin Li, Wenda Chu, Yi Fang

机构 * Santa Clara University(圣克拉拉大学) Nanyang Technological University(南洋理工大学) California Institute of Technology(加州理工学院)

AI总结 本文研究了大型语言模型在处理内部知识与外部信息时的平衡问题,提出三源交互框架,评估27个模型,发现模型更依赖文档断言,且通过微调可提升其区分能力。

Comments Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22166 2026-04-27 cs.CL

Fine-Grained Analysis of Shared Syntactic Mechanisms in Language Models

语言模型中共享句法机制的细粒度分析

Ryoma Kumon, Hitomi Yanaka

机构 * The University of Tokyo(东京大学) RIKEN(日本理化学研究所) Tohoku University(东北大学)

AI总结 本研究通过细粒度分析探讨语言模型在不同句法结构中是否共享神经机制,发现填词-缺口依赖在早期至中期层有高度局部化的共享机制,而否定极性项目许可则无统一机制。

Comments Accepted to ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22098 2026-04-27 cs.CL

Knowledge-driven Augmentation and Retrieval for Integrative Temporal Adaptation

基于知识的增强与检索的整合时间适应

Weisi Liu, Guangzeng Han, Xiaolei Huang

机构 * University of Memphis(密苏里大学)

AI总结 本文提出KARITA方法,通过整合知识源和处理时间变化,提升多领域分类任务的性能,证明知识整合在时间增强学习中的关键作用。

Comments Accepted at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21806 2026-04-27 cs.CV

TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval

TEMA: 图像锚定,文本引导的多修改复合图像检索

Zixu Li, Yupeng Hu, Zhiheng Fu, Zhiwei Chen, Yongqi Li, Liqiang Nie

机构 * School of Software, Shandong University(山东大学软件学院) Department of Computing, Hong Kong Polytechnic University(香港理工大学计算机系) School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)计算机科学与技术学院)

AI总结 TEMA通过构建多修改数据集和提出文本导向实体映射架构,解决复合图像检索中的实体覆盖不足和句法-实体对齐问题,提升检索准确性和效率。

Comments Accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14651 2026-04-27 cs.CL

CURA: Clinical Uncertainty Risk Alignment for Language Model-Based Risk Prediction

CURA:语言模型基于风险预测的临床不确定性风险对齐

Sizhe Wang, Ziqi Xu, Claire Najjuuko, Charles Alba, Chenyang Lu

机构 * Washington University in St. Louis(华盛顿大学圣路易斯分校)

AI总结 CURA通过结合个体误差概率和群体模糊性,提升临床语言模型的风险预测不确定性校准,实验表明其能提高校准指标并减少过度自信的误判。

Comments Accepted at ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.10079 2026-04-27 cs.CL

Why Supervised Fine-Tuning Fails to Learn: A Systematic Study of Incomplete Learning in Large Language Models

为何监督微调失效:对大语言模型中不完全学习现象的系统研究

Chao Xue, Yao Wang, Mengqiao Liu, Di Liang, Xingsheng Han, Peiyang Liu, Xianjie Wu, Chenyao Lu, Lei Jiang, Yu Lu, Haibo Shi, Shuang Liang, Minlong Peng, Flora D. Salim

机构 * University of New South Wales(新南威尔士大学) Tencent Hunyuan(腾讯文言) Tencent Yuanbao(腾讯元宝) UESTC(电子科技大学) Peking University(北京大学)

AI总结 本文系统研究了大语言模型微调中不完全学习现象,揭示了五个导致学习不完整的原因,并提出诊断优先框架和缓解策略,证明监督微调的局限性。

Comments Accepted by ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04855 2026-04-27 cs.CL

HACHIMI: Scalable and Controllable Student Persona Generation via Orchestrated Agents

HACHIMI: 通过协调代理实现可扩展和可控的学生人设生成

Yilin Jiang, Fei Tan, Xuanyu Yin, Jing Leng, Aimin Zhou

机构 * East China Normal University(华东师范大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Shanghai Innovation Institute(上海创新研究院)

AI总结 HACHIMI通过协调代理框架生成理论对齐且分布可控的学生人设,生成100万个人设用于1-12年级,验证了其在教育理论和人口分布上的有效性。

Comments 46 pages, 7 figures, accepted by ACL 2026. The dataset is available at https://huggingface.co/datasets/sii-research/HACHIMI-1M

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.23022 2026-04-27 cs.CL

DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis

DimABSA:构建多语言多领域数据集用于维度方面基于情感分析

Lung-Hao Lee, Liang-Chih Yu, Natalia Loukashevich, Ilseyar Alimova, Alexander Panchenko, Tzu-Mi Lin, Zhe-Yu Xu, Jian-Yu Zhou, Guangmin Zheng, Jin Wang, Sharanya Awasthi, Jonas Becker, Jan Philip Wahle, Terry Ruas, Shamsuddeen Hassan Muhammad, Saif M. Mohammad

机构 * National Yang Ming Chiao Tung University Yuan Ze University Moscow State University Skoltech AIRI Yunnan University University of Cincinnati University of Göttingen Imperial College London National Research Council Canada

AI总结 本文提出DimABSA,首个多语言多领域数据集,通过引入连续valence-arousal评分提升细粒度情感分析能力,提供三个子任务及连续F1指标,推动多语言维度情感分析发展。

Comments accepted at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12979 2026-04-27 cs.CL

The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality Check

扩散语言模型在代理工作流中的苦涩教训:全面的现实检验

Qingyu Lu, Liang Ding, Kanjian Zhang, Jinxia Zhang, Dacheng Tao

机构 * Southeast University(东南大学) Alibaba(阿里巴巴) Southeast University Shenzhen Research Institute(东南大学深圳研究院) College of Computing and Data Science at Nanyang Technological University(南洋理工大学计算机与数据科学学院)

AI总结 本文评估了扩散语言模型在代理任务中的表现,发现其在长期规划和工具调用任务中存在系统性失败,提出需整合因果推理机制以提升代理能力。

Comments ACL 2026 - Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12430 2026-04-27 cs.CL

System-Mediated Attention Imbalances Make Vision-Language Models Say Yes

系统介导的注意力失衡使视觉-语言模型倾向于说‘是’

Tsan Tsai Chan, Varsha Suresh, Anisha Saha, Michael Hahn, Vera Demberg

机构 * Saarland Informatics Campus, Saarland University, Germany(萨尔兰州信息学校区,萨尔兰州大学,德国) Max Planck Institute for Informatics, Germany(马克斯·普朗克信息研究所,德国)

AI总结 本文研究了视觉-语言模型中系统介导的注意力失衡对‘是’偏见的影响,提出通过重新分配注意力以减少这种偏见,从而提升模型可靠性。

Comments Accepted to ACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11020 2026-04-27 cs.CL

From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models

从可解释性到性能:优化长上下文语言模型的检索头

Youmi Ma, Naoaki Okazaki

机构 * Department of Computer Science, Institute of Science Tokyo(东京科学研究所计算机科学系)

AI总结 本文研究检索头如何提升长上下文语言模型性能,提出RetMask方法通过对比正常输出与屏蔽检索头的输出生成训练信号,显著提升生成和重排序性能,验证了检索头的功能作用。

Comments Findings of ACL 2026; Source code available at https://github.com/YoumiMa/RetMask

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05414 2026-04-27 cs.CL cs.AI stat.ML

Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions

大语言模型是糟糕的骰子玩家:LLM在生成统计分布的随机数时表现不佳

Minda Zhao, Yilun Du, Mengyu Wang

机构 * Harvard University(哈佛大学)

AI总结 研究发现大语言模型在生成随机数时存在显著缺陷,其采样能力随分布复杂度和采样范围增加而下降,导致下游任务出现系统性偏差。

Comments Accepted to ACL 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏