arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Empirical Methods in Natural Language Processing · 会议 · Natural Language Processing

共收录 7861
2506.08123 2025-12-05 cs.CL

QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA

QA-LIGN:通过宪法分解的问答对对齐大语言模型

Jacob Dineen, Aswin RRV, Qin Liu, Zhikun Xu, Xiao Ye, Ming Shen, Zhaonan Li, Shijie Lu, Chitta Baral, Muhao Chen, Ben Zhou

AI总结 QA-LIGN通过分解奖励信号提升LLM对齐效果,降低攻击成功率并保持低拒绝率,实现安全与帮助性的帕累托最优。

Comments Findings of the Association for Computational Linguistics: EMNLP 2025, pages 20619-20642, Suzhou, China

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, pages 20619-20642, Suzhou, China, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14607 2025-12-04 cs.CL cs.CR

sudoLLM: On Multi-role Alignment of Language Models

sudoLLM: 关于语言模型的多角色对齐

Soumadeep Saha, Akshay Chaturvedi, Joy Mahapatra, Utpal Garain

机构 * ISI Kolkata(印度Kolkata ISI研究所) IRIT Toulouse(法国图卢兹 IRIT 研究所)

AI总结 sudoLLM通过注入用户偏见信号,实现多角色对齐的LLM,提升安全性和抗攻击能力。

Comments Accepted to EMNLP 2025 (findings)

Journal ref In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 366-384, Suzhou, China. Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15478 2025-12-04 cs.CL

A Group Fairness Lens for Large Language Models

为大语言模型引入群体公平性视角

Guanqun Bi, Yuqiang Xie, Lei Shen, Yanan Cao

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

AI总结 本文提出从群体公平性角度评估大语言模型的偏见,引入 GFAIR 数据集和 GF-THINK 方法,以缓解模型中的偏见问题。

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00195 2025-12-04 cs.CL cs.AI cs.HC

Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences

让它们轻松一些!LLM护栏对用户感知和偏好的情境影响

Mingqian Zheng, Wenjia Hu, Patrick Zhao, Motahhare Eslami, Jena D. Hwang, Faeze Brahman, Carolyn Rose, Maarten Sap

机构 * Carnegie Mellon University(卡内基梅隆大学) Simon Fraser University(西蒙弗雷泽大学) Allen Institute for AI(人工智能研究所)

AI总结 研究探讨了LLM护栏对用户感知和偏好的影响,发现部分合规策略能显著降低负面感知,强调应通过创造性的拒绝策略而非意图检测来提升安全性和用户体验。

Comments Accepted to Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22479 2025-12-03 cs.CL

NeLLCom-Lex: A Neural-agent Framework to Study the Interplay between Lexical Systems and Language Use

NeLLCom-Lex: 一种神经代理框架用于研究词汇系统与语言使用之间的相互作用

Yuqing Zhang, Ecesu Ürker, Tessa Verhoef, Gemma Boleda, Arianna Bisazza

机构 * Center for Language and Cognition, University of Groningen(语言与认知中心,格罗宁根大学) Department of Translation and Language Sciences, Universitat Pompeu Fabra(翻译与语言科学系,庞培法拉大学) Leiden Institute of Advanced Computer Science, Leiden University(莱顿高级计算机科学研究所,莱顿大学) Catalan Institution for Research and Advanced Studies (ICREA)(加泰罗尼亚研究与高级科学研究所(ICREA))

AI总结 NeLLCom-Lex通过神经代理框架模拟词汇系统演变,研究词汇与语言使用间的相互作用及语义变化机制。

Comments Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21265 2025-12-03 cs.CL cs.AI

Multilingual Pretraining for Pixel Language Models

多语言预训练用于像素语言模型

Ilker Kesen, Jonas F. Lotz, Ingo Ziegler, Phillip Rust, Desmond Elliott

机构 * Department of Computer Science, University of Copenhagen(计算机科学系,哥本哈根大学) ROCKWOOL Foundation Research Unit(ROCKWOOL基金会研究单位)

AI总结 PIXEL-M4通过多语言预训练提升了像素语言模型对多种语言的支持能力,尤其在非拉丁字母语言中表现更优。

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13487 2025-12-03 cs.CL

TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs

TurBLiMP:土耳其语的语言最小对基准

Ezgi Başar, Francesca Padovani, Jaap Jumelet, Arianna Bisazza

机构 * Center for Language and Cognition (CLCG), University of Groningen(语言与认知中心(CLCG)、格罗宁根大学)

AI总结 TurBLiMP是首个土耳其语语言最小对基准,用于评估语言模型的语法能力,揭示了先进模型在处理复杂语法现象上的不足。

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13127 2025-12-03 cs.CL

Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoning

通过监督的链式思维推理促进长上下文理解

Jingyang Lin, Andy Wong, Tian Xia, Shenghua He, Hui Wei, Mei Han, Jiebo Luo

机构 * University of Rochester(罗切斯特大学) PAII Inc.(PAII公司) University of California, Merced(加州梅尔德大学)

AI总结 本文提出基于属性的代理推理框架PAI,通过生成包含中间推理步骤的合成数据集LongFinanceQA,提升LLMs在金融领域的长上下文理解能力。

Comments Main Conference of EMNLP 2025, Project Page: https://long-pai.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23689 2025-12-02 cs.CL

Child-Directed Language Does Not Consistently Boost Syntax Learning in Language Models

儿童导向语言并不总是能提升语言模型的语法学习

Francesca Padovani, Jaap Jumelet, Yevgen Matusevych, Arianna Bisazza

机构 * Center for Language and Cognition (CLCG) University of Groningen(语言与认知中心(CLCG)格罗宁根大学)

AI总结 该研究发现儿童导向语言在提升语言模型语法学习方面效果不一致,且维基百科数据表现更优,提出FIT-CLAMS方法以更公平评估语法能力。

Comments 21 pages, 4 figures, 4 tables

Journal ref Proceedings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22273 2025-12-02 cs.CL

Comprehensive Evaluation on Lexical Normalization: Boundary-Aware Approaches for Unsegmented Languages

对词汇规范化进行综合评估:面向无分隔语言的边界感知方法

Shohei Higashiyama, Masao Utiyama

机构 * National Institute of Information and Communications Technology(信息与通信技术国家研究所)

AI总结 本文针对无分隔语言提出词汇规范化方法,构建大规模数据集并验证了编码器和解码器在多视角评估中的有效性。

Comments EMNLP 2025 (Findings), 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20073 2025-12-02 cs.CL cs.AI cs.MA

Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents

Collab-Overcooked: 大语言模型作为协作代理的基准测试与评估

Haochen Sun, Shuwen Zhang, Lujie Niu, Lei Ren, Hao Xu, Hao Fu, Fangkun Zhao, Caixia Yuan, Xiaojie Wang

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Li Auto Inc

AI总结 本文提出Collab-Overcooked基准测试,评估大语言模型在协作代理中的表现,揭示其在目标解释和协作能力上的优劣。

Comments Accepted to EMNLP 2025 Main Conference. Camera-Ready Version. 30 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.06434 2025-12-02 cs.CL cs.AI cs.MM cs.SD eess.AS

Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition

Whispering LLaMA: 一种用于语音识别的跨模态生成性错误校正框架

Srijith Radhakrishnan, Chao-Han Huck Yang, Sumeer Ahmad Khan, Rohit Kumar, Narsis A. Kiani, David Gomez-Cabrero, Jesper N. Tegner

AI总结 Whispering LLaMA通过跨模态融合技术提升语音识别的生成性错误校正性能,相较n-best假设提升37.66%的词错误率。

Comments Accepted to EMNLP 2023 as main paper. 10 pages. Revised math notations. GitHub: https://github.com/Srijith-rkr/Whispering-LLaMA

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23391 2025-12-01 cs.CL

Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization

模糊性意识优化:朝着直接偏好优化的语义消歧

Jian Li, Shenglin Yin, Yujia Zhang, Alan Zhao, Xi Chen, Xiaohui Zhou, Pengfei Xu

机构 * AI Technology Center of OVB, Tencent, China(腾讯OVB人工智能技术中心,中国) School of Computer Science, Peking University, China(北京大学计算机学院,中国)

AI总结 本文提出模糊性意识优化(AAO)方法,通过计算偏好对的语义相似性自动重新加权模糊内容,有效提升直接偏好优化的性能。

Comments Accepted at EMNLP 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23376 2025-12-01 cs.HC cs.CL

Is Passive Expertise-Based Personalization Enough? A Case Study in AI-Assisted Test-Taking

基于被动专家经验的个性化是否足够?一项在AI辅助考试中的案例研究

Li Siyan, Jason Zhang, Akash Maharaj, Yuanming Shi, Yunyao Li

机构 * Columbia University(哥伦比亚大学) Georgia Institute of Technology(佐治亚理工学院) Adobe(Adobe公司)

AI总结 本文研究了基于被动专家经验的个性化在AI辅助考试中的效果,发现其能降低任务负荷但存在局限,需结合主动个性化以提升用户体验。

Comments Accepted into Tailoring AI: Exploring Active and Passive LLM Personalization (PALS) workshop at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21725 2025-12-01 cs.CL cs.AI

PromptTailor: Multi-turn Intent-Aligned Prompt Synthesis for Lightweight LLMs

PromptTailor: 为轻量级大语言模型进行多轮意图对齐的提示合成

Yizhou Xu, Janet Davis

机构 * Whitman College(惠特曼学院)

AI总结 PromptTailor通过意图对齐的提示合成提升轻量级大语言模型的输出质量,以更少的模型调用次数实现更高的用户偏好率。

Comments EMNLP 2025 Workshop PALS. Additional note: There is a citation error on Evoke. The paper we are referring to is "Evoking critical thinking abilities in LLMs via reviewer-author prompt editing."

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21699 2025-12-01 cs.CL cs.AI

Cacheback: Speculative Decoding With Nothing But Cache

Cacheback: 基于缓存的推测解码

Zhiyao Ma, In Gim, Lin Zhong

机构 * Yale University(耶鲁大学)

AI总结 Cacheback是一种基于缓存的推测解码方法,通过利用语言局部性加速LLM推理,具有简洁设计和高效性能。

Journal ref In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 31067-31072. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18775 2025-12-01 cs.CL cs.AI

Financial Risk Relation Identification through Dual-view Adaptation

通过双视角适应识别金融风险关系

Wei-Ning Chiu, Yu-Hsiang Wang, Andy Hsiao, Yu-Shiang Huang, Chuan-Ju Wang

机构 * National Taiwan University(国立台湾大学) Academia Sinica(学术院)

AI总结 本文提出通过双视角适应方法,利用10-K文件提取企业间的风险关系,提升金融风险分析的自动化与准确性。

Comments 11 pages, 3 figures, EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.24199 2025-12-01 cs.CL

Linguistically-Controlled Paraphrase Generation

语言控制的改写生成

Mohamed Elgaar, Hadi Amiri

AI总结 LingConv通过细粒度控制40种语言属性生成高质量改写,显著降低属性误差并提升生成质量。

Comments This paper was published in Findings of ACL: EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17591 2025-11-27 cs.CL cs.AI cs.LG cs.SE

HGAdapter: Hypergraph-based Adapters in Language Models for Code Summarization and Clone Detection

HGAdapter: 基于超图的语言模型代码摘要与克隆检测适配器

Guang Yang, Yujie Zhu

机构 * Data Technology Group, Technology Research and Development Department, Guotai Haitong Securities, China(国泰君安证券数据技术部、技术研发部) School of Computer Science and Technology, East China Normal University, China(华东师范大学计算机科学与技术学院)

AI总结 HGAdapter通过引入高阶数据相关性,改进超图神经网络架构,提升预训练语言模型在代码摘要与克隆检测任务中的性能。

Comments Accepted by the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025) as a findings long paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01341 2025-11-26 cs.CL

TurnBench-MS: A Benchmark for Evaluating Multi-Turn, Multi-Step Reasoning in Large Language Models

TurnBench-MS: 一个评估大语言模型多轮多步推理能力的基准

Yiran Zhang, Mo Wang, Xiaoyang Li, Kaixuan Ren, Chencheng Zhu, Usman Naseem

AI总结 TurnBench-MS通过互动破码任务评估大语言模型的多轮多步推理能力,揭示当前模型在复杂推理任务中的显著不足。

Comments Accepted to Findings of the Association for Computational Linguistics: EMNLP 2025

Journal ref Findings of the ACL: EMNLP 2025, pp. 19892-19924, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18685 2025-11-26 cs.CL

From Generation to Detection: A Multimodal Multi-Task Dataset for Benchmarking Health Misinformation

从生成到检测:一个多模态多任务数据集用于基准测试健康误信息

Zhihao Zhang, Yiran Zhang, Xiyue Zhou, Liting Huang, Imran Razzak, Preslav Nakov, Usman Naseem

机构 * Macquarie University(麦考瑞大学) University of Sydney(悉尼大学) UTS(UTS大学) MBZUAI

AI总结 本文提出MM Health数据集,包含人类和AI生成的多模态健康误信息,用于基准测试可靠性检查、原创性检查和细粒度AI检测。

Comments Accepted to Findings of the Association for Computational Linguistics: EMNLP 2025

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 24245-24260, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19168 2025-11-25 cs.LG cs.CL

RAVEN++: Pinpointing Fine-Grained Violations in Advertisement Videos with Active Reinforcement Reasoning

RAVEN++: 通过主动强化推理精准定位广告视频中的细粒度违规

Deyi Ji, Yuekui Yang, Liqun Liu, Peng Shu, Haiyang Wu, Shaogang Tang, Xudong Chen, Shaoping Ma, Tianrun Chen, Lanyun Zhu

机构 * Tencent(腾讯) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学)

AI总结 RAVEN++通过主动强化推理提升广告视频中细粒度违规检测的精度与泛化能力

Comments EMNLP 2025 (Oral, Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18934 2025-11-25 cs.CL cs.AI cs.DB

Skeletons Matter: Dynamic Data Augmentation for Text-to-Query

骨架很重要:面向文本到查询的动态数据增强

Yuchen Ji, Bo Xu, Jie Shi, Jiaqing Liang, Deqing Yang, Yu Mao, Hai Chen, Yanghua Xiao

机构 * School of Data Science, Fudan University(复旦大学数据科学学院) School of Computer Science and Technology, Donghua University(东华大学计算机科学与技术学院) College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) Ant Group(蚂蚁集团) Shanghai Key Laboratory of Data Science(上海数据科学 key laboratory)

AI总结 本文提出一种动态数据增强框架,通过识别查询骨架作为共同优化目标,提升文本到查询任务的泛化能力和效率。

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12109 2025-11-25 cs.CL cs.AI

Personalized LLM Decoding via Contrasting Personal Preference

通过对比个人偏好实现个性化大语言模型解码

Hyungjune Bu, Chanjoo Jung, Minjae Kang, Jaehyung Kim

机构 * Yonsei University(延世大学) Opt-AI Inc.(Opt-AI公司)

AI总结 本文提出CoPe方法,通过奖励引导解码实现个性化,提升ROUGE-L指标10.57%。

Comments EMNLP 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23715 2025-11-25 cs.CL

Don't Take the Premise for Granted: Evaluating the Premise Critique Ability of Large Language Models

不要轻信前提:评估大语言模型的前提批判能力

Jinzhe Li, Gengxu Li, Yi Chang, Yuan Wu

机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE, China(知识驱动人机智能工程研究中心,教育部,中国) International Center of Future Science, Jilin University(未来科学国际中心,吉林大学)

AI总结 本研究提出前提批判基准,评估大语言模型在面对错误前提时的批判能力,揭示其在推理与前提批判上的差异及改进需求。

Comments EMNLP 2025 Findings camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22061 2025-11-25 cs.CL

Safeguarding Privacy of Retrieval Data against Membership Inference Attacks: Is This Query Too Close to Home?

在成员推断攻击中保护检索数据的隐私:这个查询太接近家了吗?

Yujin Choi, Youngjoo Park, Junyoung Byun, Jaewook Lee, Jinseong Park

AI总结 本文提出了一种基于相似性的成员推断攻击检测框架,用于保护RAG系统中检索数据的隐私,通过检测和隐藏策略有效防御攻击。

Comments Accepted for EMNLP findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22006 2025-11-25 cs.CL cs.LG

Enhancing Domain-Specific Encoder Models with LLM-Generated Data: How to Leverage Ontologies, and How to Do Without Them

通过LLM生成数据增强领域特定编码器模型:如何利用本体论,以及如何不依赖本体论

Marc Brinner, Tarek Al Mustafa, Sina Zarrieß

机构 * Computational Linguistics Department of Linguistics(计算语言学系) Institute of Computer Science(计算机科学研究所) Bielefeld University(比勒菲尔德大学)

AI总结 通过LLM生成数据增强领域特定编码器模型,利用本体论或自动提取概念,实现低资源环境下的高效预训练。

Comments Published in the Findings of the Association for Computational Linguistics: EMNLP 2025

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 22740-22754). Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17775 2025-11-25 cs.CL

FoREST: Frame of Reference Evaluation in Spatial Reasoning Tasks

FoREST: 空间推理任务中的参考框架评估

Tanawan Premsri, Parisa Kordjamshidi

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Michigan State University(密歇根州立大学)

AI总结 FoREST基准通过空间引导提示法提升LLMs在空间推理任务中的参考框架理解能力。

Comments 10 pages, 3 Figures, 4 Tables, EMNLP-2025 Main (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14734 2025-11-25 cs.CL

Sentence Smith: Controllable Edits for Evaluating Text Embeddings

Sentence Smith: 可控编辑用于评估文本嵌入

Hongji Li, Andrianos Michail, Reto Gubelmann, Simon Clematide, Juri Opitz

机构 * University of Zurich(苏黎世大学)

AI总结 Sentence Smith通过可控编辑生成文本,用于评估文本嵌入模型的细粒度性能。

Comments EMNLP 2025 (main), this version fixes a subscript typo in Eq 1

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18436 2025-11-25 cs.CL

Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switching

代码切换文本能否在LLMs中激活知识切换?一种英语-韩语代码切换案例研究

Seoyeon Kim, Huiseo Kim, Chanjun Park, Jinyoung Yeo, Dongha Lee

AI总结 本研究通过EnKoQA数据集探讨代码切换是否能激活LLMs中的知识,发现其在低资源语言任务中具有潜力。

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏