arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Annual Meeting of the Association for Computational Linguistics · 会议 · Natural Language Processing

2026-08-06 至 2026-08-06 共收录 7
2608.04899 2026-08-06 cs.CL 新提交

Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification

基于大语言模型的分类置信度估计中的评估缺陷与稀疏性限制

Elena Merdjanovska, Omar Zaidan, Andreas Rücklé

机构 * Humboldt-Universität zu Berlin(柏林洪堡大学) Amazon(亚马逊公司)

AI总结 该研究指出 LLM 分类置信度估计中 verbalization 方法存在稀疏性缺陷,AUARC 评估的插值选择会影响排名,提出 verbalization logprobs 方法可解决稀疏性并提升 AUARC 且无额外推理成本。

Comments Published at Findings of ACL 2026

Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pages 33424-33435

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04576 2026-08-06 cs.CL 新提交

Causal Evidence Extraction and Triangulation in Crisis Reports using Large Language Models: A ReliefWeb-based Study

基于大型语言模型的危机报告中因果证据提取与三角验证:一项基于ReliefWeb的研究

Yuanjun Zhang, Mourad Oussalah

机构 * University of Oulu(奥卢大学) LUT University(拉普兰塔理工大学)

AI总结 该研究基于ReliefWeb数据,提出两阶段LLM流水线提取危机报告中结构化因果证据,结合上下文保留的三角验证方法,在100份专家标注报告中取得优异性能,为现金援助的食品相关结果提供了高收敛性证据。

Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pages 32478-32491, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04056 2026-08-06 cs.CL cs.CY cs.LG 新提交

Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization

基于多智能体视角偏好优化的性别歧视检测学习

Hadi Mohammadi, Tina Shahedi, Robert A. Bagheri, Mehdi Dastani, Masoume M. Raeissi

机构 * Utrecht University(乌得勒支大学) Wageningen University & Research(瓦赫宁根大学及研究中心)

AI总结 本研究针对性别歧视检测中标注分歧问题,提出MAP-PO框架,通过聚类标注者并训练对应智能体,结合个体与团队奖励协调智能体,实验验证了聚类训练及团队信号的必要性。

Comments 17 pages, 12 figures, 14 tables. Preprint; under review at EACL 2027 (ACL Rolling Review, August 2026 cycle). Code and data: https://github.com/mohammadi-hadi/MAP-PO

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02087 2026-08-06 cs.AI cs.CL cs.LG 版本更新

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy

基于非对称强化学习与自蒸馏的指令条件探索

Jim Dilkes, Vahid Yazdanpanah, Sebastian Stein

机构 * University of Southampton(南安普顿大学)

AI总结 该研究针对LLM强化学习中的探索挑战,提出指令条件探索(ICE)方法,结合非对称RL/SD训练目标,使Qwen3-1.7B数学推理性能提升5.0%且长上下文下仍有效。

Comments Submitted to ACL Rolling Review (ARR) May 2026 cycle. OpenReview submission record at https://openreview.net/forum?id=PV945lekMa

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18199 2026-08-06 cs.CL

Linear-Time and Constant-Memory Text Embeddings Based on Recurrent Language Models

基于循环语言模型的线性时间与常数内存文本嵌入

Tobias Grantner, Emanuel Sallinger, Martin Flechl

机构 * Dynatrace Research(DynaTrace研究)

AI总结 本文提出基于循环架构的高效文本嵌入方法,通过垂直分块推理策略实现线性时间复杂度和常数内存使用,验证了Mamba2等模型在文本嵌入任务中的有效性。

Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2026, pages 41459-41481

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21692 2026-08-06 cs.CL cs.AI 版本更新

Revisiting Generalization Across Difficulty Levels: It's Not So Easy

重新审视不同难度层级间的泛化:这并不容易

Yeganeh Kordi, Nihal V. Nayak, Max Zuo, Ilana Nguyen, Stephen H. Bach

机构 * Brown University(布朗大学) Harvard University(哈佛大学)

AI总结 本文研究了LLMs在不同任务难度间泛化的能力,发现训练数据的难度对泛化效果影响有限,强调在训练和评估中需涵盖多种难度以避免风险。

Comments Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19358 2026-08-06 cs.CL cs.AI

Benchmarking and Improving LLM Robustness for Personalized Generation

Chimaobi Okite, Naihao Deng, Kiran Bodipati, Huaidian Hou, Joyce Chai, Rada Mihalcea

机构 * University of Michigan(密歇根大学)

Comments First draft. First camera-ready version

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏