arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Conference on Empirical Methods in Natural Language Processing · 会议 · Natural Language Processing

2026-02-03 至 2026-02-03 共收录 7
2506.16123 2026-02-03 cs.CL

FinCoT: Grounding Chain-of-Thought in Expert Financial Reasoning

FinCoT:在专家金融推理中 grounding 思维链

Natapong Nitarach, Warit Sirichotedumrong, Panop Pitchayarthorn, Pittawat Taveekitworachai, Potsawee Manakul, Kunat Pipatanakul

机构 * SCB 10X, SCBX Group(SCB 10X集团)

AI总结 FinCoT通过整合专家金融推理蓝图,提升金融领域模型性能并减少推理成本,实现更可解释的推理过程。

Comments Accepted at FinNLP-2025, EMNLP (Oral Presentation)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00573 2026-02-03 cs.CL

Training a Utility-based Retriever Through Shared Context Attribution for Retrieval-Augmented Language Models

通过共享上下文归因训练基于效用的检索器以增强检索增强语言模型

Yilong Xu, Jinhua Gao, Xiaoming Yu, Yuanhai Xue, Baolong Bi, Huawei Shen, Xueqi Cheng

机构 * State Key Lab of AI Safety, Institute of Computing Technology, CAS(人工智能安全国家重点实验室,计算技术研究所,中国科学院) Key Lab of AI Safety, Chinese Academy of Sciences(人工智能安全重点实验室,中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 SCARLet通过共享上下文归因和多任务泛化提升检索增强语言模型的检索性能。

Comments EMNLP 2025 Main Conference (Long paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16718 2026-02-03 cs.SD cs.CL cs.LG

CAARMA: Class Augmentation with Adversarial Mixup Regularization

CAARMA: 基于对抗混合正则化的类别增强

Massa Baali, Xiang Li, Hao Chen, Syed Abdul Hannan, Rita Singh, Bhiksha Raj

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 CAARMA通过在嵌入空间中生成合成类别并采用对抗性细化机制,提升说话人验证和零样本语音分析任务的性能。

Comments Accepted to EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12769 2026-02-03 cs.CL cs.AI

How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination

LLMs在不同语言中产生幻觉的程度有多大?对多语言现实估计LLM幻觉的探讨

Saad Obaid ul Islam, Anne Lauscher, Goran Glavaš

机构 * WüNLP, CAIDAS, University of Würzburg(乌尔姆大学) Data Science Group, University of Hamburg(汉堡大学)

AI总结 研究评估了多语言LLM在长形式问答中的幻觉程度,发现高资源语言中幻觉率与语言规模无关,且支持更多语言的LLM幻觉率更高。

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23208 2026-02-03 cs.CL

A Structured Framework for Evaluating and Enhancing Interpretive Capabilities of Multimodal LLMs in Culturally Situated Tasks

一种评估和增强多模态大语言模型在文化情境任务中解释能力的结构框架

Haorui Yu, Ramon Ruiz-Dolz, Qiufeng Yi

机构 * DJCAD, University of Dundee, United Kingdom(邓迪大学DJCAD部门) ARG-tech, SSEN, University of Dundee, United Kingdom(邓迪大学) School of Computer Science, University of Birmingham, United Kingdom(伯明翰大学计算机科学学院)

AI总结 本研究提出了一种结构框架,用于评估和增强多模态大语言模型在文化情境任务中生成中国绘画批评的能力,通过量化评价特征和人设引导提示,揭示了VLMs在艺术批评领域的表现与局限。

Comments EMNLP 2025 submission, 10 pages, 6 figures, 5 tables

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, pages 1945-1971, Suzhou, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16584 2026-02-03 cs.CL cs.AI

From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations

从分数到步骤:诊断和改进证据医学计算中LLM的性能

Benlu Wang, Iris Xia, Yifan Zhang, Junda Wang, Feiyun Ouyang, Shuo Han, Arman Cohan, Hong Yu, Zonghai Yao

机构 * Department of Computer Science, Yale University, CT, USA(耶鲁大学计算机科学系) Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(VA贝福德医疗中心健康组织与实施研究中心) Miner School of Computer and Information Sciences, UMass Lowell, MA, USA(UMass洛厄尔矿尔计算机与信息科学学院) Manning College of Information and Computer Sciences, UMass Amherst, MA, USA(UMass阿默斯特马宁信息与计算机科学学院)

AI总结 本文提出MedRaC框架,通过分步评估和代码执行提升LLM在证据医学计算中的准确性,揭示现有评估方法的不足,并推动临床可信度的提升。

Comments Equal contribution for the first two authors. To appear as an Oral presentation in the proceedings of the Main Conference on Empirical Methods in Natural Language Processing (EMNLP) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06979 2026-02-03 cs.CL

MedTutor: A Retrieval-Augmented LLM System for Case-Based Medical Education

MedTutor: 一种基于检索的LLM系统用于基于病例的医学教育

Dongsuk Jang, Ziyao Shangguan, Kyle Tegtmeyer, Anurag Gupta, Jan Czerminski, Sophie Chheang, Arman Cohan

机构 * Department of Computer Science, Yale University(耶鲁大学计算机科学系) Department of Radiology and Biomedical Imaging, Yale School of Medicine(耶鲁医学院放射学与生物医学成像系) Interdisciplinary Program for Bioengineering, Seoul National University(首尔国立大学生物工程跨学科项目)

AI总结 MedTutor是一种基于检索的LLM系统,通过自动从临床病例报告生成教育内容和多项选择题,提升医学教育质量。

Comments Accepted to EMNLP 2025 (System Demonstrations)

Journal ref Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 319-353

详情

展开后加载摘要…

URL PDF HTML 收藏