arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

Conference on Empirical Methods in Natural Language Processing · 会议 · Natural Language Processing

共收录 7861
2512.22712 2026-03-31 cs.CL

Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languages

从分歧开始:理解跨语言推理-答案不一致现象

Anaelia Ovalle, Candace Ross, Sebastian Ruder, Adina Williams, Karen Ullrich, Mark Ibrahim, Levent Sagun

机构 * FAIR, Meta(Meta 人工智能基础研究团队) Meta Superintelligence Labs(Meta 超级智能实验室)

AI总结 研究探讨了多语言模型推理与答案的一致性问题,发现非拉丁文字推理痕迹与结论的不一致程度是拉丁文字的两倍,提出基于人工标注的错误分类方法,揭示了证据错误和逻辑推理错误是主要问题。

Comments Accepted to 2025 EMNLP Multilingual Representation Learning Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18602 2026-03-30 cs.CL

Do LLMs suffer from Multi-Party Hangover? A Diagnostic Approach to Addressee Recognition and Response Selection in Conversations

大型语言模型是否受多方对话的困扰?一种诊断方法用于对话中的收件人识别和回应选择

Nicolò Penzo, Maryam Sajedinia, Bruno Lepri, Sara Tonelli, Marco Guerini

机构 * Fondazione Bruno Kessler, Italy(意大利布鲁诺·凯斯勒基金会) University of Trento, Italy(意大利特伦托大学) University of Turin, Italy(意大利都灵大学)

AI总结 本文提出了一种方法论流程,用于研究对话特定结构属性下的模型性能,通过诊断子数据集分析回应选择和收件人识别任务,揭示模型弱点。

Comments Accepted to EMNLP 2024 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.20093 2026-03-30 cs.CL cs.AI

Evaluating Neural Language Models as Cognitive Models of Language Acquisition

评估神经语言模型作为语言习得的认知模型

Héctor Javier Vázquez Martínez, Annika Lea Heuser, Charles Yang, Jordan Kodner

机构 * University of Pennsylvania(宾夕法尼亚大学) Stony Brook University(石溪大学)

AI总结 本文指出现有神经语言模型语法能力评估基准不够严谨,建议使用经过严格评估的语料库以更准确研究语言习得机制。

Comments To appear in the GenBench 2023 workshop proceedings, the first workshop on (benchmarking) generalisation in NLP. GenBench 2023 will be held at EMNLP 2023 on December 6, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16889 2026-03-27 cs.CL

Can GRPO Boost Complex Multimodal Table Understanding?

GRPO能否提升复杂多模态表格理解?

Xiaoqiang Kang, Shengen Wu, Zimu Wang, Yilin Liu, Xiaobo Jin, Kaizhu Huang, Wei Wang, Yutao Yue, Xiaowei Huang, Qiufeng Wang

机构 * School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学先进科技学院) Department of Computer Science, University of Liverpool(利物浦大学计算机科学系) Information Hub, Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)信息中心) University of Southern California(南加州大学) Duke Kunshan University(杜克大学昆山分校)

AI总结 本文提出Table-R1框架,通过预热、感知对齐GRPO和提示-完成GRPO三阶段提升多模态表格理解性能,优于SFT和GRPO。

Comments EMNLP 2025

Journal ref EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12164 2026-03-25 cs.CL cs.DB cs.LG

Table-LLM-Specialist: Language Model Specialists for Tables using Iterative Generator-Validator Fine-tuning

Table-LLM-Specialist: 用于表格任务的语言模型专家使用迭代生成-验证微调

Junjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong, Shi Han, Dongmei Zhang, Surajit Chaudhuri

机构 * University of Michigan(密歇根大学) Microsoft Corporation(微软公司)

AI总结 本文提出Table-LLM-Specialist,通过生成-验证框架实现表格任务的有效微调,无需人工标注数据,提升性能并降低部署成本。

Comments Full version of a paper in EMNLP 2025; code is available at: https://github.com/microsoft/Table-Specialist

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.05033 2026-03-24 cs.IR

Collaborative User Prompt for Personalized Generative Recommendation

协同用户提示用于个性化生成推荐

Jerome Ramos, Bin Wu, Aldo Lipani

AI总结 本文提出一种整合个体与集体偏好的人工智能框架,通过注意力机制融合相似兴趣用户的嵌入表示,提升推荐系统的个性化与协同性。

Comments Accepted at PALS Workshop@EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.09252 2026-03-20 cs.CL cs.AI cs.HC

DAVIS: Planning Agent with Knowledge Graph-Powered Inner Monologue

DAVIS:基于知识图谱的规划代理

Minh Pham Dinh, Munira Syed, Michael G Yankoski, Trenton W. Ford

机构 * Davis Institute for Artificial Intelligence(戴维斯人工智能研究所) Colby College(科伯学院) University of Notre Dame(圣约翰大学) William & Mary(威廉与玛丽学院)

AI总结 本文提出DAVIS,一种基于知识图谱的规划代理,通过结构化和时间记忆实现模型驱动规划,并采用多轮检索系统提升科学任务处理能力,在ScienceWorld基准测试中表现优异。

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.02887 2026-03-20 cs.CL cs.AI

Using Optimal Transport as Alignment Objective for fine-tuning Multilingual Contextualized Embeddings

利用最优传输作为对齐目标来微调多语言上下文嵌入

Sawsan Alqahtani, Garima Lalwani, Yi Zhang, Salvatore Romeo, Saab Mansour

机构 * AWS(亚马逊)

AI总结 本文提出在微调过程中使用最优传输作为对齐目标,以提升多语言上下文表示,无需预训练词对齐,通过上下文无监督学习词对齐,实现不同映射类型,实验在XNLI和XQuAD任务上取得优于基线的改进。

Journal ref EMNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20371 2026-03-17 cs.CL

DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain Translation

DMDTEval: LLMs在多领域翻译中歧义辨析的评估与分析

Zhibo Man, Yuanmeng Chen, Yujie Zhang, Jinan Xu

AI总结 本文提出DMDTEval框架,通过构建多领域歧义词标注测试集、设计精准的歧义度量标准,评估不同提示策略对LLMs在多领域翻译中歧义辨析能力的影响,揭示提升LLMs歧义辨析性能的关键发现。

Comments Accepted by EMNLP2025-main

Journal ref DMDTEval: An Evaluation and Analysis of LLMs on Disambiguation in Multi-domain Translation (Man et al., EMNLP 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16457 2026-03-17 cs.HC cs.AI cs.CL

Integrating Personality into Digital Humans: A Review of LLM-Driven Approaches for Virtual Reality

将人格融入数字人类:一种基于大语言模型的虚拟现实方法综述

Iago Alves Brito, Julia Soares Dollis, Fernanda Bufon Färber, Pedro Schindler Freire Brasil Ribeiro, Rafael Teixeira Sousa, Arlindo Rodrigues Galvão Filho

AI总结 本文综述了利用大语言模型驱动虚拟现实中的数字人类人格塑造方法,探讨了零样本、少样本和微调等技术,并指出计算需求、延迟及多模态交互评估框架缺乏等挑战。

Comments Revised and expanded version of the survey published in Findings of EMNLP 2025. Includes 14 pages and 2 figures

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, 9519--9532

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07597 2026-03-16 cs.CL

Instructing Large Language Models for Low-Resource Languages: A Systematic Study for Basque

为低资源语言指导大型语言模型:巴斯克语的系统研究

Oscar Sainz, Naiara Perez, Julen Etxaniz, Joseba Fernandez de Landa, Itziar Aldabe, Iker García-Ferrero, Aimar Zabala, Ekhi Azurmendi, German Rigau, Eneko Agirre, Mikel Artetxe, Aitor Soroa

AI总结 本文系统研究了巴斯克语在低资源场景下的指导模型方法,发现目标语言语料库至关重要,合成指令能产生稳健模型,且使用指导模型作为骨干优于非指导基础模型。

Comments Accepted at EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18395 2026-03-13 cs.CL

NormGenesis: Multicultural Dialogue Generation via Exemplar-Guided Social Norm Modeling and Violation Recovery

NormGenesis: 通过实例引导的社会规范建模与违反恢复实现多文化对话生成

Minki Hong, Jangho Choi, Jihie Kim

AI总结 NormGenesis通过实例引导的社会规范建模与违反恢复,实现多文化对话生成,提升对话自然度和文化适应性。

Comments 39 pages, 17 figures, EMNLP 2025 Main Conference, Senior Area Chair (SAC) Highlights Award

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16083 2026-03-13 cs.CL cs.AI cs.LG

Efficient Compositional Multi-tasking for On-device Large Language Models

高效设备端大语言模型的组合多任务处理

Ondrej Bohdal, Mete Ozay, Jijoong Moon, Kyeng-Hun Lee, Hyeonmok Ko, Umberto Michieli

AI总结 本文提出了一种高效的设备端多任务处理方法,通过任务融合提升大语言模型在复杂多任务场景中的性能。

Comments Accepted at EMNLP 2025 (main track, long paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22699 2026-03-12 cs.CL

Are you sure? Measuring models bias in content moderation through uncertainty

你确定吗?通过不确定性测量内容审核中的模型偏差

Alessandra Urbinati, Mirko Lai, Simona Frenda, Marco Antonio Stranisci

机构 * Laboratory for the Modeling of Biological and Socio-technical Systems, Northeastern University(生物与社会技术系统建模实验室,东北大学) Heriot-Watt University(赫瑞-瓦特大学) aequa-tech(aequa-tech公司) Università del Piemonte Orientale(皮埃蒙特东方大学) Università degli Studi di Torino(托里尼大学)

AI总结 本文提出通过模型预测不确定性来衡量内容审核中模型的偏差,揭示预训练模型对少数群体的预测准确性与置信度的差异,以改进模型公平性。

Comments accepted at Findings of ACL: EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18722 2026-03-11 cs.AI

VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft

VistaWise: 构建低成本代理的跨模态知识图谱用于Minecraft

Honghao Fu, Junlong Ren, Qi Chai, Deheng Ye, Yujun Cai, Hao Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Queensland(昆士兰大学) Tencent(腾讯)

AI总结 VistaWise通过整合跨模态知识图谱和专用模型,实现低成本、高效率的Minecraft代理构建。

Comments Accepted by EMNLP 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09260 2026-03-11 cs.IR cs.CL

ThinkQE: Query Expansion via an Evolving Thinking Process

ThinkQE: 通过演进的思考过程实现查询扩展

Yibin Lei, Tao Shen, Andrew Yates

AI总结 ThinkQE通过演进的思考过程和语料库交互策略,在测试时实现更全面的查询扩展,优于传统密集检索器和重排序器。

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08148 2026-03-10 cs.CL cs.AI

Gradually Excavating External Knowledge for Implicit Complex Question Answering

逐步挖掘外部知识以实现隐式复杂问题回答

Chang Liu, Xiaoguang Li, Lifeng Shang, Xin Jiang, Qun Liu, Edmund Y. Lam, Ngai Wong

AI总结 本文提出一种渐进式知识挖掘框架,通过迭代获取外部知识并动态调整策略,提升开放领域复杂问题回答的准确率和效率。

Comments 13 pages, 3 figures, EMNLP findings 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02593 2026-03-10 cs.CL cs.MA cs.NE cs.SD eess.AS

Spoken Conversational Agents with Large Language Models

基于大语言模型的语音对话代理

Chao-Han Huck Yang, Andreas Stolcke, Larry Heck

机构 * NVIDIA Research(NVIDIA研究)

AI总结 本文探讨了如何通过大语言模型构建语音对话代理,涵盖从级联到端到端系统的路径,以及跨模态对齐和联合训练等关键技术,同时讨论了隐私、安全和评估中的开放问题。

Comments Accepted to EMNLP 2025 Tutorial

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01646 2026-03-09 cs.CL cs.AI cs.LG

ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledge

ESGenius:对环境、社会和治理(ESG)及可持续性知识的LLM基准测试

Chaoyue He, Xin Zhou, Yi Wu, Xinjia Yu, Yan Zhang, Lei Zhang, Di Wang, Shengfei Lyu, Hong Xu, Xiaoqiao Wang, Wei Liu, Chunyan Miao

机构 * Alibaba-NTU Global e-Sustainability CorpLab (ANGEL)(阿里巴巴-NTU全球可持续性公司实验室) Alibaba Group(阿里巴巴集团)

AI总结 ESGenius是首个针对LLM在ESG及可持续性知识评估的综合问答基准,通过RAG方法显著提升模型性能。

Comments EMNLP'25 Main Oral (42 pages, 10 figures, 11 tables), Nominations for Resource Award & Theme Paper Award

Journal ref In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP 2025), pages 14612-14653

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04308 2026-03-05 cs.LG cs.AI

Activation Outliers in Transformer Quantization: Reproduction, Statistical Analysis, and Deployment Tradeoffs

Transformer量化中的激活异常:重现、统计分析与部署权衡

Pranav Kumar Kaliaperumal

机构 * University of Colorado Denver(科罗拉多大学丹佛分校)

AI总结 本文研究了Transformer量化中的结构化激活异常问题,通过实验发现量化导致精度大幅下降,并提出混合精度和按嵌入组量化等缓解方法,强调通道意识精度分配的重要性。

Comments 10 pages, 3 tables. Reproducible study of transformer PTQ activation outliers based on Bondarenko et al. (EMNLP 2021, Qualcomm AI Research). Code: https://github.com/pranavkkp4/TransQuant-Edge

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08714 2026-03-05 cs.CV cs.CL cs.LG

Generating Fine Details of Entity Interactions

生成实体交互的细节

Xinyi Gu, Jiayuan Mao

机构 * Massachusetts Institute of Technology(麻省理工学院)

AI总结 本文提出了一种基于多模态大语言模型的方法,通过分解和细化过程提升图像中实体交互细节的生成质量。

Comments EMNLP 2025. Project Page: https://detailscribe.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04812 2026-03-03 cs.CV cs.AI cs.CL cs.LG

LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning

LLaVE: 大规模语言和视觉嵌入模型与基于难度加权的对比学习

Zhibin Lan, Liqiang Niu, Fandong Meng, Jie Zhou, Jinsong Su

机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) Pattern Recognition Center, WeChat AI, Tencent Inc, China(腾讯人工智能研究院) Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室)

AI总结 LLaVE通过基于难度加权的对比学习提升多模态嵌入模型性能,实现SOTA表现和强泛化能力。

Comments Accepted by Findings of EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13061 2026-03-03 cs.CL cs.AI cs.CV cs.LG

Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection

鲁棒适应大规模多模态模型用于检索增强的仇恨表情包检测

Jingbiao Mei, Jinghong Chen, Guangyu Yang, Weizhe Lin, Bill Byrne

机构 * Department of Engineering University of Cambridge(工程系剑桥大学)

AI总结 本文提出了一种鲁棒适应框架,用于提升大规模多模态模型在检索增强的仇恨表情包检测中的性能与泛化能力,同时提高模型的可解释性。

Comments EMNLP 2025 Main (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23840 2026-03-02 cs.CL

Measuring Sycophancy of Language Models in Multi-turn Dialogues

在多轮对话中衡量语言模型的趋炎附势性

Jiseung Hong, Grace Byun, Seungone Kim, Kai Shu, Jinho D. Choi

机构 * Carnegie Mellon University(卡内基梅隆大学) Emory University(埃默里大学)

AI总结 本研究提出SYCON基准,评估多轮对话中语言模型的趋炎附势性,发现对齐调优会放大该行为,而模型规模和推理优化能增强抗压能力,采用第三人称视角可显著减少趋炎附势性。

Comments Accepted to Findings of EMNLP 2025

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, pages 2239-2259

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22696 2026-02-27 cs.CL

Enhancing Persuasive Dialogue Agents by Synthesizing Cross-Disciplinary Communication Strategies

通过合成跨学科沟通策略增强说服对话代理

Shinnosuke Nozue, Yuto Nakano, Yotaro Watanabe, Meguru Takasaki, Shoji Moriya, Reina Akama, Jun Suzuki

机构 * Tohoku University(东北大学) PKSHA Technology Inc.(PKSHA技术公司) NINJAL

AI总结 本文提出一种跨学科方法,通过合成策略提升说服对话代理的说服效果和泛化能力,尤其在说服低意图个体方面表现突出。

Comments Accepted to the EMNLP 2025 Industry Track; 26 pages

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 2287-2312

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04403 2026-02-27 cs.CV cs.CL cs.CR

Self-adaptive Dataset Construction for Real-World Multimodal Safety Scenarios

自适应数据集构建用于现实世界多模态安全场景

Jingen Qu, Lijun Li, Bo Zhang, Yichen Yan, Jing Shao

AI总结 本文提出了一种图像导向的自适应数据集构建方法,用于构建现实世界多模态安全场景的数据集,通过生成35,000个图像-文本对及其指导响应,提升了安全评估的有效性。

Comments Accepted at EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18582 2026-02-27 cs.CL

Parallel Continuous Chain-of-Thought with Jacobi Iteration

并行连续链式思维与雅可比迭代

Haoyi Wu, Zhihao Teng, Kewei Tu

AI总结 本文提出PCCoT,通过雅可比迭代提升连续链式思维的训练和推理效率,实现更快的训练速度和更稳定的性能。

Comments Accepted to EMNLP 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02452 2026-02-26 cs.CL cs.AI cs.LG

Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions

大型语言模型是否遵循标签定义?检验其对外部标签定义的接受性

Seyedali Mohammadi, Bhaskara Hanuma Vedula, Hemank Lamba, Edward Raff, Ponnurangam Kumaraguru, Francis Ferraro, Manas Gaur

机构 * UMBC(美国马里兰大学伯克利分校) IIIT Hyderabad(印度海得拉巴理工学院) Dataminr, Inc.(数据矿研公司) CrowdStrike(CrowdStrike公司)

AI总结 研究发现大型语言模型在处理外部标签定义时,通常依赖内部表示而非显式定义,尤其在通用任务中表现更弱,而在领域特定任务中显式定义更有效。

Comments EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03867 2026-02-24 cs.CL

EuroGEST: Investigating gender stereotypes in multilingual language models

EuroGEST:研究多语言语言模型中的性别刻板印象

Jacqueline Rowe, Mateusz Klimaszewski, Liane Guillou, Shannon Vallor, Alexandra Birch

机构 * University of Edinburgh(爱丁堡大学) Warsaw University of Technology(华沙技术大学) Aveni

AI总结 EuroGEST研究多语言语言模型中的性别刻板印象,通过跨29种欧洲语言的数据集揭示了性别刻板印象的普遍性及模型对刻板印象的编码强度。

Comments In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 32074-32096, Suzhou, China. Association for Computational Linguistics. 9 pages, 5 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00339 2026-02-24 cs.CL cs.LG

GRASP: Replace Redundant Layers with Adaptive Singular Parameters for Efficient Model Compression

GRASP: 用自适应奇异参数替代冗余层以实现高效模型压缩

Kainan Liu, Yong Zhang, Ning Cheng, Zhitao Li, Shaojun Wang, Jing Xiao

机构 * Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

AI总结 GRASP通过自适应奇异参数替代冗余层,实现高效模型压缩并保持高性能。

Comments EMNLP 2025(Main)

详情

展开后加载摘要…

URL PDF HTML 收藏