arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

Annual Meeting of the Association for Computational Linguistics · 会议 · Natural Language Processing

共收录 10293
2607.07251 2026-07-09 cs.CL 新提交

Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models

视觉语言模型中使用空间指示表达式的多语言能力评估

Kaito Watanabe, Taisei Yamamoto, Tomoki Doi, Hitomi Yanaka

机构 * The University of Tokyo(东京大学) Riken(理化学研究所) Tohoku University(东北大学)

AI总结 研究聚焦视觉语言模型的空间推理能力,通过开发基准评估其使用四种语言空间指示表达式的多语言能力,实验发现测试模型使用指示词方式与人类不同,特别是在依距离选指示词方面。

Comments Accepted to ACL SRW 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.06818 2026-07-09 cs.CL cs.AI cs.LG 新提交

Ad Headline Generation using Self-Critical Masked Language Model

使用自批判掩码语言模型生成广告标题

Yashal Shakti Kanungo, Sumit Negi, Aruna Rajan

机构 * amazon(亚马逊)

AI总结 研究如何为电商网站生成吸引人的广告标题,核心方法是将强化学习策略梯度方法应用于基于Transformer的掩码语言模型,通过联合多种产品信息生成标题,该方法在指标和质量审核上优于现有方法,生成标题质量也优于人工提交的。

Comments Accepted at NAACL-HLT 2021 (Industry Track). 9 pages, 3 tables, 3 figures - ACL Anthology URL: https://aclanthology.org/2021.naacl-industry.33/ - Editors of the proceedings: Young-bum Kim, Yunyao Li, Owen Rambow - Bibkey: kanungo-etal-2021-ad

Journal ref Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Industry Papers, pages 263-271, June 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18360 2026-07-09 cs.SD cs.CL 版本更新

Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval

Omni-Embed-Audio: 利用多模态大语言模型实现鲁棒的音频-文本检索

HaeJun Yoo, Yongseop Shin, Insung Lee, Myoung-Wan Koo, Du-Seong Chang

机构 * Sogang University(首尔大学)

AI总结 提出Omni-Embed-Audio(OEA)检索编码器,利用多模态大语言模型原生理解音频,并通过用户意图查询(UIQ)和硬负样本挖掘,在文本到音频检索中达到与M2D-CLAP相当的性能,同时在文本到文本检索和硬负样本判别上显著优于现有方法。

Comments Accepted at ACL 2026 Main Conference. Camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00086 2026-07-09 cs.CL 版本更新

RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning

RIMRULE:通过MDL引导的规则学习改进使用工具的语言智能体

Xiang Gao, Yuguang Yao, Qi Zhang, Kaiwen Dong, Avinash Baidya, Ruocheng Guo, Hilaf Hasson, Kamalika Das

机构 * Intuit AI Research(Intuit AI研究)

AI总结 研究针对大型语言模型在特定领域使用工具的难题,提出RIMRULE神经符号方法,通过从失败轨迹提炼规则并用MDL目标巩固,以动态注入规则提升性能,实验证明该方法能提高工具使用准确性,且规则具有跨架构可移植性。

Comments Published as a long paper in the main conference of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05441 2026-07-08 cs.IR cs.AI 新提交

PORTS: Preference-Optimized Retrievers for Tool Selection with Large Language Models

PORTS:用于大语言模型工具选择的偏好优化检索器

Lorenzo Molfetta, Giacomo Frisoni, Nicolò Monaldini, Gianluca Moro

机构 * Department of Computer Science and Engineering, University of Bologna(计算机科学与工程系,博洛尼亚大学)

AI总结 研究针对大语言模型工具选择中现有检索器与LLMs不一致问题,提出PORTS方法,利用受困惑度启发的偏好信号,通过优化相关性及施加对比语义损失微调检索器,经多实验验证其通用性及提高工具选择准确性的能力,且计算需求低便于推广。

Comments Please cite the definitive, peer-reviewed version of this article published in the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, edited by Christos Christodoulopoulos et al., Association for Computational Linguistics, pp. 10007-10030, 2025. DOI: https://doi.org/10.18653/v1/2025.emnlp-main.507

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Association for Computational Linguistics, pp. 10007-10030, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31480 2026-07-08 cs.CL 版本更新

Language Models Can Resolve Reference Compositionally, But It's Not Their Native Strength: The Case of the Personal Relation Task

语言模型可以组合性地解析指代,但这并非其天然优势:以个人关系任务为例

Bart Evelo, Meaghan Fowlie, Denis Paperno

AI总结 通过个人关系任务,比较人类与大型语言模型在外延任务(确定指称对象)和内涵任务(结构化表示意义)上的表现,发现人类更擅长外延任务而LLM更擅长内涵任务,表明缺乏指称基础是LLM模拟人类语言理解的关键缺失。

Comments A pre-MIT Press publication version. Paper accepted to Transactions of the Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23597 2026-07-08 cs.CL cs.LG

Structure-Guided Entity Resolution: Fine-Tuning LLMs for Robust Name Matching in Complex Linguistic Contexts

结构引导的实体解析:微调大语言模型以实现复杂语言上下文中的鲁棒姓名匹配

Shivam Chourasia, Hitesh Kapoor, Nilesh Patil

机构 * Dream Sports

AI总结 提出结构引导实体解析(SGER)框架,通过两阶段课程微调大语言模型,先学习姓名语法结构再优化匹配任务,在印度身份数据上达到99.02%准确率和0.994 F1分数,已部署于Dream11平台服务2.5亿+用户。

Comments Accepted to ACL 2026. 8 pages, 1 figure, 2 tables

Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track), pages 1461-1468, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09159 2026-07-08 cs.IR 版本更新

LLMs Meet Isolation Kernel: Lightweight, Learning-free Binary Embeddings for Fast Retrieval

LLMs 与隔离内核:轻量、无学习的二进制嵌入用于快速检索

Zhibo Zhang, Yang Xu, Kai Ming Ting, Cam-Tu Nguyen

AI总结 本文提出无学习的隔离内核嵌入(IKE),将LLM嵌入转换为二进制嵌入,实现低内存和快速检索,实验显示其检索速度提升16.7倍,内存使用降低16倍,同时保持相似精度。

Comments Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10356 2026-07-08 cs.CL 版本更新

Decoding the Multimodal Mind: Generalizable Brain-to-Text Translation via Multimodal Alignment and Adaptive Routing

解码多模态思维:通过多模态对齐和自适应路由实现可泛化的脑到文本翻译

Chunyu Ye, Yunhao Zhang, Jingyuan Sun, Chong Li, Yang Zhao, Shaonan Wang

机构 * State Key Laboratory of Multimodal Artificial Intelligence System, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Department of Computer Science, The University of Manchester(曼彻斯特大学计算机科学系) Department of Language Science and Technology, Hong Kong Polytechnic University(香港理工大学语言科学与技术系)

AI总结 该研究针对脑机接口从人脑解码语言的挑战,提出利用多模态大语言模型和路由模块的统一框架,通过多模态对齐和自适应路由将脑信号与多模态语义空间对齐,在fMRI等数据集实验中性能领先,还扩展到EEG和MEG数据,为现实应用提供灵活方案。

Comments Accepted to ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.03196 2026-07-07 cs.CL cs.LG

Geometric Deviation as an Unsupervised Pre-Generation Reliability Signal: Probing LLM Representations for Answerability

几何偏差作为一种无监督预生成可靠性信号:探测大语言模型表示的可回答性

Yucheng Du

机构 * University of Southern California(南加州大学)

AI总结 研究能否通过测量隐藏状态与可回答参考集的偏差,利用表示几何提供预生成信号。在三个模型和三种提示形式上实验,发现几何主要编码任务形式,数学提示中有可区分性,代码提示有部分泛化,该信号在早期层出现。

Comments Accepted to TrustNLP 2026 (ACL Workshop). 11 pages, 3 figures, 3 tables

Journal ref In Proceedings of the 6th Workshop on Trustworthy NLP (TrustNLP 2026), pp. 353-363, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05174 2026-07-07 cs.AI 新提交

AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments

AgentGym2:在去理想化真实世界环境中评测大语言模型智能体

Zhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang, Zhonghang Lu, Chenyu Liu, Jiajun Sun, Jiazheng Zhang, Dingwei Zhu, Xin Guo, Junzhe Wang, Zhihao Zhang, Yuming Yang, Junjie Ye, Minghe Gao, Dongrui Liu, Jiaming Ji, Guohao Li, Tao Gui, Qi Zhang, Xuanjing Huang

机构 * Fudan University(复旦大学) Zhejiang University(浙江大学) Shanghaijiaotong University(上海交通大学) Peking University(北京大学)

AI总结 针对现有LLM智能体基准过于理想化的问题,提出去理想化评测框架AgentGym2,可全面测查智能体实操能力,实验显示当前SOTA模型仍存在明显性能缺口。

Comments Accepted as a main conference paper at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04854 2026-07-07 cs.AI 新提交

CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs

CARL:用于大语言模型规划的约束感知强化学习

Qiuyi Qi, Jinjian Zhang, Mutian Bao, Tian Liang, Guocong Li, Dongnan Liu, Wei Zhou, Jie Liu, Ming Kong, Linjian Mo, Feng Zhang, Qiang Zhu

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团) City University of Hong Kong(香港城市大学) College of Artificial Intelligence, Shanghai Institute for Advanced Study, Zhejiang University(浙江大学上海高等研究院人工智能学院) School of Earth Sciences, Zhejiang University(浙江大学地球科学学院)

AI总结 研究大语言模型生成违反任务约束计划的问题,提出约束感知强化学习框架CARL,通过比较不同输入下模型输出分布引入奖励,提升模型约束意识,实验证明其优于基线和现有模型。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04727 2026-07-07 cs.SE cs.AI cs.CV 新提交

Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive Dashboards

Dashboard2Code:在重建交互式仪表板上评估多模态模型

Tianhao Niu, Ziyu Han, Qiguang Chen, Shiqi Zhou, Baocai Shan, Hengjie Fang, Qingfu Zhu, Wanxiang Che

机构 * Research Center for Social Computing and Interactive Robotics(社会计算与交互机器人研究院)

AI总结 研究多模态模型在交互式仪表板重建任务,介绍Dashboard2Code任务,提出DashboardMimic基准及自动化评估框架,实验发现开源与闭源模型在此任务上有差距。

Comments Accepted to ACL2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03833 2026-07-07 cs.CL cs.AI 新提交

Beyond Static Rules: Automated Discovery of Latent Vulnerabilities in Text-to-SQL

超越静态规则:文本到SQL中潜在漏洞的自动发现

Hanqing Wang, Yongdong Chi, Jian Yang, Lei Yang, Jiehui Zhao, Yun Chen, Guanhua Chen

机构 * Shanghai University of Finance and Economics(上海财经大学) Beihang University(北京航空航天大学) Deepexi Technology Co. Ltd.(深圳市智元机器有限公司) Southern University of Science and Technology(南方科技大学)

AI总结 研究LLMs在文本到SQL任务中潜在可靠性问题,提出SAGE框架,通过生成漏洞假设、参考漏洞法典设计扰动来发现潜在故障模式,实验证明其能发现大量故障案例及法典的跨模型转移性。

Comments Accepted by Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03709 2026-07-07 cs.CL 新提交

GRASP: Graph-Reasoning Aided Survey Planning for High-Fidelity Related Work Generation

GRASP:用于高保真相关工作生成的图推理辅助综述规划

Haoming Li, Jessica Ouyang

机构 * Department of Computer Science(计算机科学系)

AI总结 研究如何写文献综述,核心方法是结合大语言模型规划与图算法,主要贡献是提出GRASP框架,其两层图结构及拓扑感知剪枝能生成与人工撰写匹配的相关工作部分。

Comments 23 pages, 3 figures. Published in Findings of the Association for Computational Linguistics: ACL 2026

Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pages 36427-36449, San Diego, California, United States. Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03466 2026-07-07 cs.CL cs.AI cs.LG 新提交

CaresAI at SMM4H-HeaRD 2026: Predicting TNM Staging

CaresAI在SMM4H-HeaRD 2026中的应用:预测TNM分期

Joseph Itopa Abubakar, Jorge Jarme, Favour Igwezeke, Mary Adewunmi

机构 * CaresAI(关爱人工智能) Ateneo De Naga University(那牙雅典耀大学) Faculty of Pharmaceutical Sciences, Nsukka, Enugu, Nigeria(尼日利亚埃努古州恩苏卡药学院) Menzies School of Health Research(孟席斯健康研究学院)

AI总结 该研究以癌症基因组图谱病理报告为基础,将预测TNM分期问题转化为多标签分类任务,探索经典和深度学习方法,结果显示组合嵌入能提升预测能力,虽有局限但提供了基线模型和可重复流程。

Journal ref Association for Computational Linguistics 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03213 2026-07-07 cs.CV cs.AI cs.CL cs.HC 新提交

OpenGlass: A Sensing-Computing Split Architecture for Local MLLM-Driven Real-Time Visual Assistance

OpenGlass:用于本地MLLM驱动的实时视觉辅助的传感-计算分离架构

Mengzhang Li, Yuan Yao

机构 * Shanghai Qizhi Institute(上海期智研究院) College of AI, Tsinghua University(清华大学人工智能学院)

AI总结 针对视障和低视力用户,OpenGlass以传感-计算分离解决云MLLM辅助需上传数据、有网络延迟,以及可穿戴眼镜计算和电量受限问题,在本地设备实现低延迟多模态视觉辅助,并给出评估结果。

Comments Accepted to ACL 2026 System Demonstrations. 11 pages, 5 figures, 8 tables

Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 829-839, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03003 2026-07-07 cs.CL 新提交

psytechlab at CLPsych 2026: Utilising Natural Language Processing methods and Large Language Models for Social Media Text Analysis

CLPsych 2026 中的心理技术实验室:利用自然语言处理方法和大语言模型进行社交媒体文本分析

Igor Buyanov, Nafisa Valieva, Ekaterina Mazurina

机构 * psytechlab(心理技术实验室)

AI总结 在 CLPsych 2026 共享任务中,利用自然语言处理方法和大语言模型对社交媒体文本进行自我状态和幸福感分析与总结,为改进心理健康支持系统做贡献。

Comments Accepted by CLPsych2026. CLPsych 2026 will be held at ACL in San Diego July 4th, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09118 2026-07-07 cs.AI 新提交

ComplexConstraints and Beyond: Expert Rubrics for RLVR

复杂约束与超越:RLVR的专家评分标准

Sushant Mehta, Liudas Panavas, Suhaas Garre, Edwin Chen

机构 * Surge AI

AI总结 提出专家设计的评分标准作为评估和训练信号,通过复杂指令遵循和企业智能体任务验证,在RL训练中显著提升模型性能。

Comments Accepted to the GEM workshop at ACL 2026: https://gem-workshop.com/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.02509 2026-07-07 cs.CL 版本更新

When Rating Scales Fall Short: LLM-Assisted Discovery of ADHD Signals in Turkish Teacher Narratives

当评分量表不足时:LLM辅助发现土耳其教师叙述中的ADHD信号

Baris Karacan, Irem Aktar Songur, Ahmet Ozaslan, Elvan Iseri

机构 * Department of Computer Science, University of Illinois Chicago(伊利诺伊大学芝加哥分校计算机科学系) Department of Child and Adolescent Psychiatry, Gazi University(加齐大学儿童与青少年精神病学系)

AI总结 本研究通过分析土耳其教师评估表中的结构化评分和开放式叙述,利用大语言模型辅助的主题发现方法,揭示了叙述文本中未被结构化量表捕捉的ADHD互补信号。

Comments 15 pages. Accepted to CLPsych 2026. Camera-ready author version. The final version will appear in the ACL Anthology

Journal ref In Proceedings of the 10th Workshop on Computational Linguistics and Clinical Psychology (CLPsych 2026), pages 368-382, San Diego, California, USA, 2026. Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20052 2026-07-07 cs.CL cs.AI

PromptRad: Knowledge-Enhanced Multi-Label Prompt-Tuning for Low-Resource Radiology Report Labeling

PromptRad: 基于知识的多标签提示微调用于低资源放射报告标注

Ying-Jia Lin, Tzu-Chin Lo, Ping-Chien Li, Chi-Tung Cheng, Chien-Hung Liao, Hung-Yu Kao

机构 * Department of Artificial Intelligence and AI Research Center, Chang Gung University(人工智能系及AI研究中心,长庚大学) Department of Radiology, Sijhih Cathay General Hospital(放射科,西吉医院) Department of Medical Imaging and Intervention, Chang Gung Memorial Hospital(医学影像与介入科,长庚纪念医院) Department of Trauma and Emergency Surgery, Chang Gung Memorial Hospital(创伤与急诊外科,长庚纪念医院) Department of Computer Science, National Tsing Hua University(计算机科学系,国立清华大学)

AI总结 本文提出PromptRad,一种基于知识的多标签提示微调方法,用于在低资源环境下进行放射报告标注,通过引入UMLS元词典中的同义词增强类别表示,以更少的标注数据实现优于传统方法的性能。

Comments BioNLP 2026 @ ACL (camera-ready version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01685 2026-07-07 cs.CL cs.AI cs.MA

Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage

用真理欺骗:通过生成蒙太奇进行开放式通道多智能体合谋以操纵信念

Jinwei Hu, Xinmiao Huang, Youcheng Sun, Yi Dong, Xiaowei Huang

机构 * University of Liverpool(利物浦大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

AI总结 本文研究了通过公开通道分发真实证据片段,利用多智能体合谋操纵信念的新威胁,提出了生成蒙太奇框架,展示了在14种LLM家族中74.4%的攻击成功率,并揭示了更强的推理能力反而增加了易受攻击的风险。

Comments Accepted to the ACL 2026 Main Conference (Oral Presentation)

Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 5979-5996, San Diego, California, United States, July 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19185 2026-07-07 cs.CL cs.AI

SCURank: Ranking Multiple Candidate Summaries with Summary Content Units for Enhanced Summarization

SCURank: 基于摘要内容单元的多候选摘要排序以提升摘要生成

Bo-Jyun Wang, Ying-Jia Lin, Hung-Yu Kao

机构 * Department of Computer Science and Information Engineering, National Cheng Kung University(国立成功大学计算机科学与资讯工程系) Artificial Intelligence Research Center, Chang Gung University(长庚大学人工智能研究中心) Department of Artificial Intelligence, Chang Gung University(长庚大学人工智能系) Department of Computer Science, National Tsing Hua University(国立清华大学计算机科学系)

AI总结 SCURank通过摘要内容单元提升摘要生成,优于传统指标和LLM排序方法,验证了信息导向排序在多LLM蒸馏中的优势。

Comments Accepted by ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03786 2026-07-07 cs.CL cs.LG

Compact Example-Based Explanations for Language Models

紧凑的基于示例的语言模型解释

Loris Schoenegger, Benjamin Roth

机构 * Faculty of Computer Science, University of Vienna(维也纳大学计算机科学学院) UniVie Doctoral School Computer Science, University of Vienna(维也纳大学计算机科学博士学院) Faculty of Philological and Cultural Studies, University of Vienna(维也纳大学语言文化研究学院)

AI总结 本文提出一种无需重新训练的选例相关性评分,用于评估示例集对模型输出解释的有效性,并展示了其在选择策略中的应用,提升解释质量。

Comments ACL 2026 Findings. 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16512 2026-07-07 cs.AI 版本更新

Framework of Thoughts: A Foundation Framework for Dynamic and Optimized Reasoning based on Chains, Trees, and Graphs

思维框架:基于链、树和图的动态优化推理基础框架

Felix Fricke, Simon Malberg, Georg Groh

机构 * School of Computation, Information and Technology(计算、信息与技术学院) Technical University of Munich(慕尼黑技术大学)

AI总结 研究针对现有思维提示方案的局限,提出通用基础框架FoT,内置超参调整等功能,通过在其中实现三种流行方案验证其能力,能加快执行、降低成本并获更好任务分数,还发布代码库促进相关推理方案发展。

Comments Published at SURGeLLM 2026, ACL 2026. Camera-ready version

Journal ref Proceedings of the First Workshop on Structured Understanding, Retrieval, and Generation in the LLM Era (SURGeLLM 2026), pp. 132-151, Association for Computational Linguistics, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11098 2026-07-07 cs.SD cs.CL 版本更新

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents

VCB Bench:用于音频基础大语言模型对话代理的评估基准

Jiliang Hu, Wenfu Wang, Zuchao Li, Chenxing Li, Yiyang Zhao, Hanzhao Li, Liqiang Zhang, Meng Yu, Dong Yu

机构 * School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院) Tencent AI Lab(腾讯AI实验室)

AI总结 针对现有音频语言模型基准局限,提出VCB Bench,基于真实人类语音构建,从指令跟随、知识理解、鲁棒性三个角度评估,揭示性能差距并指明改进方向,提供评估框架。

Comments 25 pages, 9 figures, accepted by ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04853 2026-07-07 cs.CL 版本更新

Decomposed Prompting Does Not Fix Knowledge Gaps, But Helps Models Say "I Don't Know"

分解式提示无法弥补知识差距,但有助于模型说“我不知道”

Dhruv Madhwal, Lyuxin David Zhang, Dan Roth, Tomer Wolfson, Vivek Gupta

机构 * Arizona State University(亚利桑那州立大学) University of Pennsylvania(宾夕法尼亚大学) Oracle AI

AI总结 研究大语言模型在闭卷问答中识别知识局限的问题,评估三种提示方式在不同模型规模和多跳问答基准下的影响,利用提示方式间的分歧信号实现无训练弃权策略,提升模型可靠性。

Comments Camera-ready version. Published in Findings of ACL 2026. Code and data: https://github.com/dhruvmadhwal/disagreement-based-abstention

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07475 2026-07-07 cs.LG cs.AI 版本更新

ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs

ARCQuant:通过增强残差通道提升大语言模型的NVFP4量化

Haoqian Meng, Yilun Luo, Yafei Zhao, Wenyuan Liu, Peng Zhang, Xindian Ma

机构 * School of Computer Science and Technology, Tianjin University(天津大学计算机科学与技术学院)

AI总结 提出ARCQuant框架,利用增强残差通道提升NVFP4性能,维持统一格式,将误差补偿融入矩阵降维以用标准GEMM内核,理论分析和实验表明达到先进精度且有实际效益。

Comments Accepted to ACL 2026 (Main Conference)

Journal ref In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8609-8623, San Diego, California, United States, July 2026. Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07233 2026-07-07 cs.AI

From "Thinking" to "Justifying": Aligning High-Stakes Explainability with Professional Communication Standards

从‘思考’到‘证明’:将高风险可解释性与专业沟通标准对齐

Chen Qian, Yimeng Wang, Yu Chen, Lingfei Wu, Andreas Stathopoulos

机构 * William & Mary(威廉玛丽学院) Anytime AI

AI总结 本文提出‘结果->证明’方法,通过结构化可解释性框架SEF提升系统输出的可验证性与可靠性,在三个领域四个任务中验证了该方法的有效性。

Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pp. 24628-24637

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19186 2026-07-07 cs.CL 版本更新

When Users Are Happy but Agents Are Wrong: Multi-Dimensional Evaluation of Tool-Augmented Dialogue

当用户满意但智能体出错:工具增强对话的多维度评估

Tanya Shourya, Yingfan Wang, Zhaoyi Joey Hou, Shamik Roy, Vinayshekhar Bannihatti Kumar, Rashmi Gangadharaiah

机构 * AWS AI Labs(AWS人工智能实验室) University of Pittsburgh(匹兹堡大学)

AI总结 针对工具增强对话系统中用户满意但智能体错误的问题,提出TRACE基准,通过系统合成多样错误案例,评估现有框架发现性能远未理想。

Comments The Fifth Generation, Evaluation & Metrics Workshop (GEM) at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏