arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Empirical Methods in Natural Language Processing · 会议 · Natural Language Processing

共收录 7862
2509.16551 2025-10-07 cs.CL cs.AI

Rethinking the Role of Text Complexity in Language Model Pretraining

Dan John Velasco, Matthew Theodore Roque

机构 * Samsung R&D Institute Philippines(三星菲律宾研发中心)

Comments Camera-ready version for BabyLM Workshop at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11465 2025-10-07 cs.CL cs.LG

CEMTM: Contextual Embedding-based Multimodal Topic Modeling

Amirhossein Abaskohi, Raymond Li, Chuyuan Li, Shafiq Joty, Giuseppe Carenini

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02160 2025-10-07 cs.CL cs.AI

Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages

David Demitri Africa, Suchir Salhan, Yuval Weiss, Paula Buttery, Richard Diehl Martinez

机构 * University of Cambridge(剑桥大学)

Comments Accepted (poster) to 5th Workshop on Multilingual Representation Learning at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08163 2025-10-07 cs.CL cs.AI cs.LG

LPI-RIT at LeWiDi-2025: Improving Distributional Predictions via Metadata and Loss Reweighting with DisCo

Mandira Sawkar, Samay U. Shetty, Deepak Pandita, Tharindu Cyril Weerasooriya, Christopher M. Homan

机构 * Rochester Institute of Technology(罗切斯特理工学院)

Comments To appear in Proceedings of the EMNLP 2025 Workshop on Learning with Disagreements (LeWiDi)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00924 2025-10-07 cs.CL

XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML

Ernesto L. Estevanell-Valladares, Suilan Estevez-Velarde, Yoan Gutiérrez, Andrés Montoyo, Ruslan Mitkov

机构 * University of Alicante(阿利坎特大学) University of Havana(哈瓦那大学) University of Lancaster(兰卡斯特大学)

Comments 18 pages, 10 figures, 7 tables. Preprint. Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23158 2025-10-07 cs.CL

User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal

Yuhan Liu, Michael J. Q. Zhang, Eunsol Choi

机构 * New York University(纽约大学)

Comments EMNLP camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22968 2025-10-07 cs.CL cs.AI

C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations

Chengqian Ma, Wei Tao, Yiwen Guo

机构 * Peking University(北京大学) LIGHTSPEED

Comments EMNLP 2025 main; Project Page: https://step-out.github.io/C3-web/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13686 2025-10-07 cs.CR

TopicAttack: An Indirect Prompt Injection Attack via Topic Transition

Yulin Chen, Haoran Li, Yuexin Li, Yue Liu, Yangqiu Song, Bryan Hooi

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00534 2025-10-07 cs.CR cs.AI

The Security Threat of Compressed Projectors in Large Vision-Language Models

Yudong Zhang, Ruobing Xie, Xingwu Sun, Jiansheng Chen, Zhanhui Kang, Di Wang, Yu Wang

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) Large Language Model Department, Tencent(腾讯大语言模型部门) School of Computer and Communication Engineering, University of Science and Technology Beijing(北京科技大学计算机与通信工程学院) Faculty of Science and Technology, University of Macau(澳门大学科技学院)

Comments Accepted by EMNLP 2025 findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23832 2025-10-07 cs.CL cs.IR

LegalSearchLM: Rethinking Legal Case Retrieval as Legal Elements Generation

Chaeeun Kim, Jinu Lee, Wonseok Hwang

Comments EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21815 2025-10-07 cs.IR cs.AI cs.CL

Scientific Paper Retrieval with LLM-Guided Semantic-Based Ranking

Yunyi Zhang, Ruozhen Yang, Siqi Jiao, SeongKu Kang, Jiawei Han

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Korea University(韩国大学)

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18562 2025-10-07 cs.CL cs.AI

From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test

Xunlian Dai, Li Zhou, Benyou Wang, Haizhou Li

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Research Institute of Big Data(深圳大数据研究院)

Comments Cultural Analysis, Cultural Alignment, Word Association Test, Large Language Models. Accepted by EMNLP 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19064 2025-10-07 cs.CL

Can Large Language Models Outperform Non-Experts in Poetry Evaluation? A Comparative Study Using the Consensual Assessment Technique

Piotr Sawicki, Marek Grześ, Dan Brown, Fabrício Góes

Comments 18 pages, 3 figures. Accepted for publication at the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10246 2025-10-07 cs.LG

Detecting LLM Hallucination Through Layer-wise Information Deficiency: Analysis of Ambiguous Prompts and Unanswerable Questions

Hazel Kim, Tom A. Lamb, Adel Bibi, Philip Torr, Yarin Gal

机构 * University of Oxford(牛津大学)

Comments Accepted to EMNLP(main)2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03467 2025-10-07 cs.CL

Searching for the Most Human-like Emergent Language

Brendon Boldt, David Mortensen

机构 * Language Technologies Institute Carnegie Mellon University(语言技术研究所卡内基梅隆大学)

Comments Accepted for publication at the 2025 Conference on Empirical Methods in Natural Language Processing; 19 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03439 2025-10-07 cs.CL

Morpheme Induction for Emergent Language

Brendon Boldt, David Mortensen

机构 * Language Technologies Institute Carnegie Mellon University(语言技术研究所卡内基梅隆大学)

Comments Accepted for publication at the 2025 Conference on Empirical Methods in Natural Language Processing; 16 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02788 2025-10-06 cs.CL

XTRA: Cross-Lingual Topic Modeling with Topic and Representation Alignments

Tien Phat Nguyen, Vu Minh Ngo, Tung Nguyen, Linh Van Ngo, Duc Anh Nguyen, Sang Dinh, Trung Le

机构 * Hanoi University of Science and Technology(河内科学技术大学) University of Monash(墨尔本大学)

Comments 2025 EMNLP Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02665 2025-10-06 cs.CL

Self-Improvement in Multimodal Large Language Models: A Survey

Shijian Deng, Kai Wang, Tianyu Yang, Harsh Singh, Yapeng Tian

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Toronto(多伦多大学) University of Notre Dame(诺特丹大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02648 2025-10-06 cs.CL

SoT: Structured-of-Thought Prompting Guides Multilingual Reasoning in Large Language Models

Rui Qi, Zhibo Man, Yufeng Chen, Fengran Mo, Jinan Xu, Kaiyu Huang

机构 * Key Laboratory of Big Data & Artificial Intelligence in Transportation (Beijing Jiaotong University), Ministry of Education(大数据与人工智能交通 key laboratory(北京交通大学)) School of Computer Science and Technology, Beijing Jiaotong University(计算机科学与技术学院(北京交通大学)) University of Montreal(蒙特利尔大学)

Comments EMNLP 2025 (findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02328 2025-10-06 cs.CL cs.AI cs.MA

AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question Answering

Ziqing Wang, Chengsheng Mao, Xiaole Wen, Yuan Luo, Kaize Ding

机构 * Northwestern University(西北大学) Microsoft(微软公司)

Comments EMNLP Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16076 2025-10-06 cs.CL

The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models

Marlene Lutz, Indira Sen, Georg Ahnert, Elisa Rogers, Markus Strohmaier

机构 * University of Mannheim(曼海姆大学) GESIS - Leibniz Institute for the Social Sciences(莱布尼茨社会科学研究所) Complexity Science Hub Vienna(维也纳复杂科学中心)

Comments Accepted to EMNLP Findings 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04793 2025-10-06 cs.CL cs.AI

A Survey of Pun Generation: Datasets, Evaluations and Methodologies

Yuchen Su, Yonghua Zhu, Ruofan Wang, Zijian Huang, Diana Benavides-Prado, Michael Witbrock

机构 * School of Computer Science, University of Auckland(奥克兰大学计算机科学学院) Singapore University of Technology and Design(新加坡科技设计大学) School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦女王学院电子工程与计算机科学学院)

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02790 2025-10-06 cs.CV cs.CL

From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding

Xiangfeng Wang, Xiao Li, Yadong Wei, Xueyu Song, Yang Song, Xiaoqiang Xia, Fangrui Zeng, Zaiyi Chen, Liu Liu, Gu Xu, Tong Xu

机构 * University of Science and Technology of China(中国科学技术大学) ByteDance China(字节跳动中国)

Comments Accepted by EMNLP 2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01761 2025-10-06 cs.CL

Same evaluation, more tokens: On the effect of input length for machine translation evaluation using Large Language Models

Tobias Domhan, Dawei Zhu

Comments Accepted at EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07044 2025-10-06 cs.CL cs.AI

DatawiseAgent: A Notebook-Centric LLM Agent Framework for Adaptive and Robust Data Science Automation

Ziming You, Yumiao Zhang, Dexuan Xu, Yiwei Lou, Yandong Yan, Wei Wang, Huaming Zhang, Yu Huang

机构 * National Engineering Research Center for Software Engineering, Peking University(软件工程国家工程研究中心,北京大学) School of Software & Microelectronics, Peking University(软件与微电子学院,北京大学) School of Computer Science, Peking University(计算机科学学院,北京大学) Xi’an Jiaotong University(西安交通大学) Institute of Basic Theory of Chinese Medicine, China Academy of Chinese Medical Sciences(中医基础理论研究所,中国中医科学院)

Comments The camera-ready version for EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11191 2025-10-06 cs.CR cs.AI cs.CL

Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training

Yao-Ching Yu, Tsun-Han Chiang, Cheng-Wei Tsai, Chien-Ming Huang, Wen-Kwang Tsao

机构 * AI Lab, TrendMicro(TrendMicro人工智能实验室)

Comments Accepted to EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18406 2025-10-06 cs.CV cs.AI cs.CL

RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives

Jaehong Yoon, Shoubin Yu, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Nanyang Technological University(南洋理工大学)

Comments EMNLP 2025 main; The first two authors contribute equally. Project Page: https://raccoon-mllm-gen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02297 2025-10-03 cs.LG cs.AI cs.CL

Interactive Training: Feedback-Driven Neural Network Optimization

Wentao Zhang, Yang Young Lu, Yuntian Deng

机构 * University of Waterloo(滑铁卢大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

Comments EMNLP 2025 Demo

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02292 2025-10-03 cs.CL cs.CV

From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens

Hala Sheta, Eric Huang, Shuyu Wu, Ilia Alenabi, Jiajun Hong, Ryker Lin, Ruoxi Ning, Daniel Wei, Jialin Yang, Jiawei Zhou, Ziqiao Ma, Freda Shi

Comments EMNLP 2025 System Demonstration | Code: https://github.com/compling-wat/vlm-lens

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02133 2025-10-03 cs.AI cs.LG

FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models

Karan Dua, Hitesh Laxmichand Patel, Puneet Mittal, Ranjeet Gupta, Amit Agarwal, Praneet Pabolu, Srikant Panda, Hansa Meghwani, Graham Horwood, Fahad Shah

机构 * Oracle Corporation(Oracle公司)

Comments Accepted at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏