arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Annual Meeting of the Association for Computational Linguistics · 会议 · Natural Language Processing

共收录 10304
2512.01020 2026-05-04 cs.AI cs.CL

Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics

用法律问题树准则评估法律推理轨迹

Jinu Lee, Kyoung-Woon On, Simeng Han, Arman Cohan, Julia Hockenmaier

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) LBOX Stanford(斯坦福) Yale(耶鲁)

AI总结 本文提出 LEGIT 数据集,用于评估 LLM 在法律领域推理轨迹的质量,发现法律问题覆盖度和正确性影响 LLM 推理能力,RAG 和带准则的 RL 分别提升整体能力与正确性。

Comments ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10913 2026-05-04 cs.CL

ADVICE: Answer-Dependent Verbalized Confidence Estimation

ADVICE: 答案依赖性口头置信度估计

Ki Jung Seo, Sehun Lim, Taeuk Kim

机构 * Department of Computer Science, Hanyang University(韩国汉阳大学计算机科学系)

AI总结 本文提出ADVICE框架,通过增强答案依赖性改进大语言模型的置信度校准,减少过度自信现象,提升可信度。

Comments ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07630 2026-05-04 cs.CL cs.AI cs.CV

InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information

InterChart:跨分解与分布式图表信息的基准测试

Anirudh Iyengar Kaniyar Narayana Iyengar, Srija Mukhopadhyay, Adnan Qidwai, Shubhankar Singh, Dan Roth, Vivek Gupta

机构 * Arizona State University(亚利桑那州立大学) IIIT, Hyderabad(海得拉巴印度理工学院) Mercer Mettl University(梅森-梅特尔大学) University of Pennsylvania(宾夕法尼亚大学)

AI总结 InterChart通过评估视觉语言模型在多图表间推理的能力,揭示了模型在复杂图表集成中的局限性,推动多模态推理的发展。

Comments 22 pages, 8 figures, 14 tables. Accepted at IJCNLP-AACL 2025

Journal ref Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics 2025, 2046-2067

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06698 2026-05-04 cs.CL

SCAN: Structured Capability Assessment and Navigation for LLMs

SCAN:面向大语言模型的结构化能力评估与导航

Zongqi Wang, Tianle Gu, Chen Gong, Xin Tian, Siqi Bao, Yujiu Yang

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) School of Cyber Engineering, Xidian University(西安电子科技大学电子工程学院) Baidu, Inc(百度公司)

AI总结 本文提出SCAN框架,通过细粒度评估实现对LLM能力的深入分析,揭示了同一类别下不同子能力的显著性能差异,强调了细粒度评估的重要性。

Comments Accepted by ACL 2026 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28147 2026-05-01 cs.CL

On the Proper Treatment of Units in Surprisal Theory

关于在惊奇理论中正确处理单位的探讨

Samuel Kiegeland, Vésteinn Snæbjarnarson, Tim Vieira, Ryan Cotterell

机构 * ETH Zürich(苏黎世联邦理工学院) University of Copenhagen(哥本哈根大学)

AI总结 本文探讨了在惊奇理论中正确处理语言单位的重要性,提出应明确区分单位定义与预测区域选择,并统一框架处理任意单位库。

Comments ACL 2026 (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27934 2026-05-01 cs.AI cs.CL

MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection

MM-StanceDet:基于检索的多模态多智能体立场检测

Weihai Lu, Zhejun Zhao, Yanshu Li, Huan He

机构 * Peking University(北京大学) Baidu Inc(百度公司) Brown University(布朗大学) Amazon(亚马逊)

AI总结 本文提出MM-StanceDet框架,通过整合检索增强、多模态分析代理、辩论阶段和自我反思,解决多模态立场检测中的上下文 grounding、跨模态解释模糊和单次推理脆弱问题,实验表明其在五个数据集上优于现有方法。

Comments Accepted on ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27929 2026-05-01 cs.CL

DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models

DPN-LE: 大语言模型双人格神经元定位与编辑

Lifan Zheng, Xue Yang, Jiawei Chen, Chenyan Wu, Jingyuan Zhang, Fanheng Kong, Xinyi Zeng, Xiang Chen, Yu Tian

机构 * Southeast University(东南大学) Shanghai Jiao Tong University(上海交通大学) East China Normal University(华东师范大学) Zhongguancun Academy(中关村学院) Zhejiang University of Technology(浙江工业大学) Kuaishou Technology(快手科技) Northeastern University(东北大学) Tsinghua University(清华大学) Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

AI总结 本文提出DPN-LE方法,通过对比高特质与低特质样本的MLP激活,定位并编辑双人格神经元,实现精准的人格控制与能力保留。

Journal ref ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27766 2026-05-01 cs.CL cs.AI

Instruction-Guided Poetry Generation in Arabic and Its Dialects

基于指令的阿拉伯语诗歌生成及其方言

Abdelrahman Sadallah, Kareem Elozeiri, Mervat Abassy, Rania Elbadry, Mohamed Anwar, Abed Alhakim Freihat, Preslav Nakov, Fajri Koto

机构 * Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学)

AI总结 本文提出基于指令的阿拉伯语诗歌生成方法,通过构建大规模指令数据集,实现诗歌创作、修改和延续,并通过实验验证模型生成诗歌的能力。

Comments ACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10702 2026-05-01 cs.CL cs.AI cs.IR

Grounding Agent Memory in Contextual Intent

在上下文意图中 grounding 代理记忆

Ruozhen Yang, Yucheng Jiang, Yueqi Jiang, Priyanka Kargupta, Yunyi Zhang, Jiawei Han

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Stanford University(斯坦福大学)

AI总结 STITCH通过上下文意图索引提升长周期交互的记忆检索能力,有效减少歧义和干扰,实现在动态目标轨迹中的高效检索。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08611 2026-05-01 cs.IR cs.AI cs.CV cs.MM

VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking

VeriTaS:首个动态多模态自动事实核查基准

Mark Rothermel, Marcus Kornmann, Marcus Rohrbach, Anna Rohrbach

机构 * Multimodal AI Lab(多模态AI实验室) Technical University of Darmstadt(达姆施塔特技术大学)

AI总结 为应对在线虚假信息的扩大规模,本文提出VeriTaS动态基准,通过自动化流程持续更新,涵盖104个专业机构的25000个真实声明,支持多模态事实核查评估。

Comments ACL 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22099 2026-05-01 cs.LG cs.AI

Decomposed Trust: Privacy, Adversarial Robustness, Ethics, and Fairness in Low-Rank LLMs

分解信任:低秩大语言模型中的隐私、对抗鲁棒性、伦理与公平性

Daniel Agyei Asante, Md Mokarram Chowdhury, Yang Li

机构 * Department of Computer Science, Iowa State University, United States(爱荷华州立大学计算机科学系) Meta, United States(Meta)

AI总结 本文研究低秩因子化对大语言模型信任性的影响,发现其在隐私保护上有所保留,但会削弱对话中个人身份信息的保护,同时增强对抗鲁棒性,但伦理和公平性有所下降。

Comments Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07180 2026-05-01 cs.CL cs.AI cs.CV

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs

视频中的奉承:视频大语言模型中奉承行为的基准测试与分析

Wenrui Zhou, Mohamed Hendy, Shu Yang, Qingsong Yang, Zikun Guo, Yuyu Luo, Lijie Hu, Di Wang

机构 * Provable Responsible AI and Data Analytics (PRADA) Lab(可证负责任人工智能与数据 analytics 实验室) King Abdullah University of Science and Technology(国王阿卜杜勒阿齐兹科学与技术大学) HKUST(香港科技大学) MBZUAI(穆罕默德·本·拉希德人工智能研究所) University of Science and Technology of China(中国科学技术大学) Kyungpook National University(庆尚国立大学)

AI总结 本文提出VISE基准,用于评估视频大语言模型在面对误导性输入时的奉承行为,通过多类型分析和两种无训练缓解策略提升模型可靠性。

Comments 27 Pages, Accepted by ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27674 2026-05-01 cs.CL cs.AI cs.CR cs.IR

One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness

一个单一的枢纽文本破坏CLIP:通过枢纽性揭示跨模态编码器的漏洞

Hiroyuki Deguchi, Katsuki Chousa, Yusuke Sakai

机构 * NTT, Inc.(NTT公司) Nara Institute of Science and Technology(奈良科学技术大学)

AI总结 本文提出方法识别枢纽嵌入及其对应文本,揭示跨模态编码器的漏洞,实验显示单一枢纽文本在图像描述评估中表现优异,暴露了编码器的缺陷。

Comments Accepted at ACL2026 (main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27637 2026-05-01 cs.AI

Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading

优化优先于评估:未经优化的提示进行评估可能是误导的

Nicholas Sadjoli, Tim Siefken, Atin Ghosh, Yifan Mai, Daniel Dahlmeier

机构 * SAP(SAP公司) Stanford University(斯坦福大学)

AI总结 本文研究了提示优化对大语言模型评估的影响,发现优化显著影响模型排名,强调在评估中应针对每个模型进行优化以选择最佳模型。

Comments Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27616 2026-05-01 cs.CL cs.MA

RoadMapper: A Multi-Agent System for Roadmap Generation of Solving Complex Research Problems

RoadMapper: 一个用于解决复杂研究问题的路线图生成的多智能体系统

Jiacheng Liu, Zichen Tang, Zhongjun Yang, Xinyi Hu, Xueyuan Lin, Linwei Jia, Ruofei Bai, Rongjin Li, Shiyao Peng, Haocheng Gao, Haihong E

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) IDEA Research(IDEA研究院) Hithink RoyalFlush Information Network Co., Ltd.(Hithink RoyalFlush信息网络有限公司)

AI总结 本文提出RoadMapper,一个基于LLM的多智能体系统,通过分解任务生成高质量路线图,提升LLM在复杂问题解决中的性能,效率提升84%。

Comments Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27550 2026-05-01 cs.CL cs.AI

APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation

APPSI-139: 一个平行语料库,用于英语应用隐私政策的摘要与解释

Pengyun Zhu, Qiheng Sun, Long Wen, Yanbo Wang, Yang Cao, Junxu Liu, Deyi Xiong, Jinfei Liu, Zhibo Wang, Kui Ren

机构 * Tianjin University(天津大学) Zhejiang University(浙江大学) North University of China(北方工业大学) Institute of Science Tokyo(东京科学研究所) The Hong Kong Polytechnic University(香港理工大学) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术区(滨江)区块链与数据安全研究院)

AI总结 本文提出APPSI-139平行语料库,包含139个隐私政策及15,692个改写语料,结合TCSI-pp-V2框架提升摘要与解释的可读性与可靠性,优于GPT-4o等大模型。

Comments Accepted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27536 2026-05-01 cs.AI

Belief-Guided Inference Control for Large Language Model Services via Verifiable Observations

基于可验证观测的信念引导推理控制:通过可验证观测实现大语言模型服务

Wenhao Yuan, Chenchen Lin, Jian Chen, Jinfeng Xu, Shuo Yang, Edith Cheuk Han Ngai

机构 * The University of Hong Kong(香港大学) Sun Yat-sen University(中山大学)

AI总结 本文提出Veroic框架,通过构建轻量级可验证观测通道,利用聚合的异质质量信号形成信念状态,以实现更优的质量-成本权衡、更强的风险估计和更稳健的长期推理控制。

Comments Accepted by KnowFM@ACL2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27495 2026-05-01 cs.CL cs.AI

Debiasing Reward Models via Causally Motivated Inference-Time Intervention

通过因果驱动的推理时干预去偏奖励模型

Kazutoshi Shinoda, Kosuke Nishida, Kyosuke Nishida

机构 * Human Informatics Labs., NTT, Inc.(NTT人类信息学实验室)

AI总结 本文提出因果驱动的干预方法,用于在推理时减轻奖励模型中的多种偏差。通过抑制与预定义偏差属性强相关的神经元激活,减少对虚假特征的敏感性,同时保持性能。

Comments Accepted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27467 2026-05-01 cs.SE cs.CL

ScaleBox: Enabling High-Fidelity and Scalable Code Verification for Large Language Models

ScaleBox:实现大规模语言模型高保真和可扩展的代码验证

Jiasheng Zheng, Xin Zheng, Boxi Cao, Pengbo Wang, Zhengzhao Ma, Qiming Zhu, Jiazhen Jiang, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun

机构 * Chinese Information Processing Laboratory(中国科学院信息处理实验室) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 ScaleBox通过自动化特殊判题生成与管理、细粒度并行执行和配置驱动的评估套件,提升大规模代码训练的验证准确性和效率,显著优于基线方法。

Comments Accepted to ACL 2026 Demo. Our project is available at https://github.com/icip-cas/ScaleBox

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27398 2026-05-01 cs.CL

Why Mean Pooling Works: Quantifying Second-Order Collapse in Text Embeddings

为何均值池化有效:量化文本嵌入中的二次崩溃

Tomomasa Hara, Hiroto Kurita, Masaaki Imaizumi, Kentaro Inui, Sho Yokoi

机构 * Tohoku University(东洋大学) The University of Tokyo(东京大学) Kyoto University(京都大学) RIKEN(日本科学技术研究所) MBZUAI NINJAL

AI总结 本文研究均值池化在真实模型中的有效性,发现其可能造成信息丢失,但现代文本编码器对这种崩溃具有鲁棒性,且鲁棒性与下游任务性能相关。

Comments ACL 2026 Main Conference; GitHub: https://github.com/tohoku-nlp/socm-text-embedding

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27340 2026-05-01 cs.AI

Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation Perspective

探究更可解释且无分区的组合性估计方法:从规则生成的角度

Ziyao Xu, Cong Wang, Houfeng Wang

机构 * National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(国家多媒体信息处理重点实验室,计算机学院,北京大学) OPPO AI Center(OPPO人工智能中心)

AI总结 本文提出从规则生成角度评估LLM组合性的新方法,解决现有测试的局限性,通过字符串到网格任务实验揭示LLM的组合性特征与不足。

Comments Accepted at ACL 2026 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27296 2026-05-01 cs.SE cs.CL

To Diff or Not to Diff? Structure-Aware and Adaptive Output Formats for Efficient LLM-based Code Editing

要进行差异生成还是不进行差异生成?面向高效LLM代码编辑的结构感知和自适应输出格式

Wei Cheng, Yongchang Cao, Chen Shen, Binhua Li, Jue Chen, Yongbin Li, Wei Hu

机构 * State Key Laboratory for Novel Software Technology, Nanjing University, China(新型软件技术国家重点实验室,南京大学,中国) Tongyi Lab, Alibaba Group, China(通义实验室,阿里巴巴集团,中国)

AI总结 本文提出BlockDiff和FuncDiff两种结构感知差异格式及AdaEdit策略,通过减少生成复杂度提升长代码编辑效率,准确度与全代码生成相当,降低30%以上延迟和成本。

Comments Accepted in the Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17022 2026-05-01 cs.CL cs.AI

Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks

超越黑盒标签:可解释的诊断标准用于主观NLP任务

Nisrine Rair, Alban Goupil, Valeriu Vrabie, Emmanuel Chochoy

机构 * CReSTIC, Université de Reims Champagne-Ardenne(里摩日香槟-阿登大学CReSTIC研究中心) Chochoy Conseil(Chochoy咨询公司)

AI总结 本文提出一种在确定金标签前审计注释方案的可解释诊断方法,用于识别主观NLP任务中分歧的根源,通过分析多注释者标准判断,发现分歧集中在少数标准上,且近半数句子激活多个类别。

Comments Accepted to ACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21016 2026-05-01 cs.CL cs.AI cs.LG

Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO

通过排列感知GRPO缓解大语言模型中的选择偏差

Jinquan Zheng, Jia Yuan, Jiacheng Yao, Chenyang Gu, Pujun Zheng, Guoxiu He

机构 * School of Economics and Management, East China Normal University(东华大学经济管理学院)

AI总结 本文提出PA-GRPO方法,通过强制排列一致的语义推理来缓解大语言模型中的选择偏差,实验显示其在七个基准测试中表现优异,有效减少偏差同时保持高性能。

Comments Accepted to ACL 2026 Main Conference. 19 pages, 3 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19044 2026-05-01 cs.CL

MoRI: Learning Motivation-Grounded Reasoning for Scientific Ideation in Large Language Models

MoRI: 在大语言模型中学习基于动机的推理以进行科学构想

Chenyang Gu, Jiahao Cheng, Meicong Zhang, Pujun Zheng, Jinquan Zheng, Guoxiu He

机构 * School of Economics and Management, East China Normal University(东华大学经济管理学院)

AI总结 MoRI通过强化学习奖励机制提升大语言模型在科学构想中的推理能力,优于现有方法,在新颖性、技术严谨性和可行性方面表现突出。

Comments Accepted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00095 2026-05-01 cs.CV cs.AI cs.CY

EDU-CIRCUIT-HW: Evaluating Multimodal Large Language Models on Real-World University-Level STEM Student Handwritten Solutions

EDU-CIRCUIT-HW:评估多模态大语言模型在真实世界大学级STEM学生手写解答中的表现

Weiyu Sun, Liangliang Chen, Yongnuo Cai, Huiru Xie, Yi Zeng, Ying Zhang

机构 * Georgia Institute of Technology(佐治亚理工学院) Virginia Tech(弗吉尼亚理工学院)

AI总结 本文提出EDU-CIRCUIT-HW数据集,用于评估多模态大语言模型在处理包含数学公式、图表和文本推理的大学STEM学生手写解答中的性能,揭示模型在自动评分中的可靠性问题,并提出通过纠正识别错误提升AI评分系统鲁棒性的方法。

Comments Accepted to Findings of the Association for Computational Linguistics: ACL 2026. Project Website: https://gt-learning-innovation.github.io/CIRCUIT_EDU_HW_ACL GitHub and Dataset: https://gt-learning-innovation.github.io/CIRCUIT_EDU_HW_ACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14289 2026-05-01 cs.CL cs.AI

RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension

RPC-Bench: 一项针对研究论文理解的细粒度基准测试

Yelin Chen, Fanjin Zhang, Suping Sun, Yunhe Pang, Yuanchun Wang, Jian Song, Xiaoyan Li, Lei Hou, Shu Zhao, Jie Tang, Juanzi Li

机构 * Xinjiang University(新疆大学) Renmin University of China(中国人民大学) Anhui University(安徽大学) Sun Yat-sen University(中山大学) Tsinghua University(清华大学) University of Southampton(南安普顿大学)

AI总结 本文提出RPC-Bench,通过高质量计算机科学论文的评审反驳交流构建,包含15000对人工验证的问答对,设计细粒度分类评估模型在学术场景中理解并回答为何、什么和如何问题的能力,揭示即使最强模型在精确学术论文理解上仍有显著差距。

Comments ACL'26, 12 pages, 23 appendix pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11908 2026-05-01 cs.CL

PPA-Plan: Proactive Pitfall Avoidance for Reliable Planning in Long-Context LLM Reasoning

PPA-Plan: 长上下文LLM推理中的前瞻性坑洞避免规划

Byeongjin Kim, Gyuwan Kim, Seo Yeon Park

机构 * Hanyang University(翰阳大学) University of California, Santa Barbara(加州大学圣芭芭拉分校)

AI总结 针对长上下文推理中计划生成不可靠的问题,PPA-Plan通过前瞻性策略预防逻辑错误,提升计划执行效果。

Comments Accepted to the Main Conference of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026). 27 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02845 2026-05-01 cs.CL cs.AI

TiMem: Temporal-Hierarchical Memory Consolidation for Long-Horizon Conversational Agents

TiMem:用于长时域对话代理的时序-层次记忆巩固

Kai Li, Xuanqing Yu, Ziyi Ni, Yi Zeng, Yao Xu, Zheqing Zhang, Xin Li, Jitao Sang, Xiaogang Duan, Xuelei Wang, Chengbao Liu, Jie Tan

机构 * Institute of Automation, CAS(中国科学院自动化研究所) School of Artificial Intelligence, UCAS(中国科学技术大学人工智能学院) AI Lab, AIGility Cloud Innovation(AIGility云创新AI实验室) North China Electric Power University(华北电力大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Gaoling School of Artificial Intelligence, RUC(中国人民大学高陵人工智能学院) School of Biomedical Engineering, USTC(中国科学技术大学生物医学工程学院) Suzhou Institute for Advance Research, USTC(中国科学技术大学苏州市先进研究院) School of Computer Science and Technology, BJTU(北京理工大学计算机科学与技术学院) Hunan Central South Intelligent Equipment Co., Ltd.(湖南中南智能装备有限公司)

AI总结 TiMem通过时序记忆树实现对话记忆的系统性巩固,提升长时域个性化效果,实现75.30%和76.88%的基准测试准确率。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18731 2026-05-01 cs.CL cs.AI cs.LG

Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards

通过可验证准确性和回避奖励的课程强化学习缓解多轮对话中的迷失问题

Ming Li, Pei Chen, Zhenhao Zhang, Tao Yang, Xinyang Zhang, Han Li, Tianyu Cao, Ming Zeng, Zhuofeng Wu, Meng Jiang, Huasheng Li, Lihong Li, Bing Yin

机构 * University of Maryland(马里兰大学) Amazon(亚马逊)

AI总结 本文提出RLAAR框架,通过可验证准确性和回避奖励的课程强化学习,减少多轮对话中因提前回答导致的性能下降,提升模型可靠性。

Comments ACL2026, camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏