arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Annual Meeting of the Association for Computational Linguistics · 会议 · Natural Language Processing

2026-01-23 至 2026-01-23 共收录 7
2601.15809 2026-01-23 cs.CL

SteerEval: Inference-time Interventions Strengthen Multilingual Generalization in Neural Summarization Metrics

SteerEval: 在推理时间干预增强神经摘要度量的多语言泛化

Silvia Casola, Ryan Soh-Eun Shim, Felicia Körner, Yuchen Mao, Barbara Plank

机构 * MaiNLP, Center for Information and Language Processing, LMU Munich(MaiNLP、信息与语言处理中心、慕尼黑大学) Language Science and Technology, Saarland University(语言科学与技术、萨尔兰大学)

AI总结 SteerEval通过在推理时干预激活向英语基准倾斜,提升多语言神经摘要度量的泛化能力。

Comments Submitted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15487 2026-01-23 cs.AI cs.CL cs.MA

MiRAGE: A Multiagent Framework for Generating Multimodal Multihop Question-Answer Dataset for RAG Evaluation

MiRAGE:一种多智能体框架,用于生成多模态多跳问题-答案数据集以评估RAG系统

Chandan Kumar Sahu, Premith Kumar Chilukuri, Matthew Hetrich

机构 * ABB Inc(ABB公司)

AI总结 MiRAGE通过多智能体框架生成多模态多跳问题-答案数据集,提升RAG系统评估的准确性和复杂性。

Comments 12 pages, 2 figures, Submitted to ACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15337 2026-01-23 cs.LG cs.CL

Language Models Entangle Language and Culture

语言模型融合语言与文化

Shourya Jain, Paras Chopra

机构 * Lossfunk

AI总结 研究发现语言选择影响LLM生成答案的文化上下文,导致低资源语言回答质量较低。

Comments Accepted at LM4UC Workshop at AAAI'26, Submitted to ACL 2026. 17 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12471 2026-01-23 cs.CL cs.AI

Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty

知何时退避:医疗大语言模型在临床不确定性中的表现

Sravanthi Machcha, Sushrita Yerra, Sahil Gupta, Aishwarya Sahoo, Sharmin Sultana, Hong Yu, Zonghai Yao

机构 * Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(马萨诸塞大学阿姆赫斯特曼宁信息与计算机科学学院) Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(医疗组织与实施研究中心) Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(米纳尔计算机与信息科学学院)

AI总结 本文提出MedAbstain基准,探讨医疗LLM在临床不确定性中的退避能力,发现显式退避选项能显著提升安全性,而模型规模和提示方法效果有限。

Comments Equal contribution for the first two authors; To appear in proceedings of the Main Conference of the European Chapter of the Association for Computational Linguistics (EACL) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12061 2026-01-23 cs.CL cs.AI

Codebook-Injected Dialogue Segmentation for Multi-Utterance Constructs Annotation: LLM-Assisted and Gold-Label-Free Evaluation

用于多轮对话结构标注的代码表注入对话分割:LLM辅助且无需黄金标签的评估

Jinsook Lee, Kirk Vanacore, Zhuqian Zhou, Bakhtawar Ahtisham, Jeanine Grutter, Rene F. Kizilcec

机构 * Cornell University(康奈尔大学) LMU Muinich(慕尼黑大学)

AI总结 本文提出一种基于LLM的对话分割方法,通过代码表注入提升分割一致性,并在无黄金标签情况下评估不同分割器的性能,发现需根据下游任务优化分割策略。

Comments Under Review for ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23761 2026-01-23 cs.SE cs.AI cs.MA

TDFlow: Agentic Workflows for Test Driven Development

TDFlow: 为测试驱动开发设计的代理工作流

Kevin Han, Siddharth Maddikayala, Tim Knappe, Om Patel, Austen Liao, Amir Barati Farimani

机构 * Carnegie Mellon University(卡内基梅隆大学) UC San Diego(南加州大学) Johns Hopkins University(约翰霍普金斯大学)

AI总结 TDFlow通过测试驱动的工作流实现人类水平的测试解析,提升软件修复性能。

Comments Published in the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026 Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10695 2026-01-23 cs.LG cs.AI cs.CL

Introducing Verification Task of Set Consistency with Set-Consistency Energy Networks

引入集合一致性验证任务与集合一致性能量网络

Mooho Song, Hyeryung Son, Jay-Yoon Lee

机构 * Seoul National University(首尔国立大学)

AI总结 本文提出集合一致性验证任务及SC-Energy模型,通过对比损失框架提升多陈述逻辑一致性验证性能,并发布新数据集

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025), Long Papers

详情

展开后加载摘要…

URL PDF HTML 收藏