arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

Annual Meeting of the Association for Computational Linguistics · 会议 · Natural Language Processing

共收录 10418 篇
2609.33670 2026-09-29 cs.CL 新提交

Closing the Cross-Dialect Gap: Query Plans as a Portable Interface in Text-to-SQL

弥合跨方言差距:查询计划作为文本到SQL中的可移植接口

Corentin Royer, Robin Oester, Yotam Perlitz, Yannick Metz, Andrea Giovannini, Mennatallah El-Assady

机构 * IBM Research(IBM研究) ; ETH Zurich(苏黎世联邦理工学院)

AI总结 针对文本到SQL跨方言性能下降问题,提出以方言无关的关系代数查询计划为生成目标,经编译器转换,在十三个模型上恢复可移植性,并引入MetricName公平评估,证明计划监督优于SQL监督。

Comments Accepted at Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.32684 2026-09-29 cs.CL 新提交

Focusing Condition: Inference-Time Self-Contrastive Steering Elicits Better Conditional Text Embeddings in LLMs

聚焦条件:推理时自对比引导提升大语言模型中的条件文本嵌入

Zifeng Cheng, Lingyun Qian, Zhiwei Jiang, Cong Wang, Yafeng Yin, Fei Shen, Ao Zhou, Qing Gu

AI总结 本文提出推理时自对比引导(SCS)方法,通过构建无条件嵌入并干预注意力机制,提升大语言模型的条件文本嵌入质量,无需训练且即插即用。

Comments ACL 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.18383 2026-09-29 cs.LG cs.CL 版本更新

When Can We Trust the Sparse Lens? A Certification Framework for SAE Faithfulness

从稀疏特征到可信代理:认证基于SAE的可解释性

Dibyanayan Bandyopadhyay, Asif Ekbal

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Patna(印度理工学院巴特那分校计算机科学与工程系)

AI总结 提出一种后验泛化框架,通过稀疏代理(SAE重建)认证语言模型,推导期望风险上界,并在GPT-2 Small等模型上验证非平凡界,揭示深层更易认证且特征分解区分语义对齐与统计稀疏性。

Comments Accepted at Transactions of the Association for Computational Linguistics (TACL). Pre-MIT Press publication version

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07606 2026-09-29 cs.CL cs.AI

Nürnberg NLP at PsyDefDetect: Multi-Axis Voter Ensembles for Psychological Defence Mechanism Classification

Nürnberg NLP在PsyDefDetect中的多轴投票集成:心理防御机制分类

Philipp Steigerwald, Eric Rudolph, Jens Albrecht

机构 * Technische Hochschule Nürnberg(纽伦堡技术大学)

AI总结 本文提出多轴投票集成方法,通过类粒度、训练方法和基模型三个维度提升心理防御机制分类效果,系统在隐藏测试集上达到0.420的F1分数,排名第一。

Comments Accepted at the BioNLP 2026 PsyDefDetect Shared Task @ ACL 2026 (1st place, 21 registered teams)

Journal ref Proceedings of the BioNLP 2026 (Shared Tasks), pp. 59-65

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01137 2026-09-29 cs.CL

Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models

跟随流动:文本到图像模型中文本令牌间信息流动的研究

Guy Kaplan, Michael Toker, Yuval Reif, Yonatan Belinkov, Roy Schwartz

机构 * Hebrew University of Jerusalem(海法大学) ; Technion – Israel Institute of Technology(技术学院 – 以色列理工学院) ; Kempner Institute, Harvard University(哈佛大学凯普纳研究所)

AI总结 本文研究文本到图像模型中令牌表示间的信息分布,发现信息常集中于少数几个令牌,且令牌间存在孤立与交互影响,改进编码可提升生成质量。

Comments Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18951 2026-09-29 cs.CL

BnMMLU: Measuring Massive Multitask Language Understanding in Bengali

BnMMLU:测量孟加拉语大规模多任务语言理解

Saman Sarker Joy, Swakkhar Shatabda

机构 * University of Malaya(马来大学) ; BRAC University(BRAC大学)

AI总结 BnMMLU通过大规模多任务评估孟加拉语语言理解,揭示模型推理和应用能力的不足,并推动多语言NLP发展。

Comments 19 Pages, 10 Tables, 12 Figures

Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pp. 12211-12230

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.30768 2026-09-28 cs.AI 新提交

Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More

思考有助于公平吗?推理标记解决了一些偏见,但制造了更多

Deng Pan, Joe Germino, Yihong Ma, Elizabeth Daly, Nuno Moniz, Ting Hua, Nitesh Chawla

机构 * University of Notre Dame(圣母大学) ; IBM Research, Dublin(IBM研究院(都柏林))

AI总结 本研究通过消融实验发现,推理语言模型中的思考对反事实公平性具有不对称双重效应:虽解决部分偏见但创造更多(约5倍),并提出了CDPG和BTM两种工具来追踪和解释这一现象。

Comments Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08596 2026-09-28 cs.CL cs.AI 版本更新

MedHal: a Synthetic Dataset for Medical Hallucination Detection

MedHal:用于医学幻觉检测的评估数据集

Fabrice Lamarche, Gaya Mehenni, Neshat Elhami Fard, Odette Rios-Ibacache, Li Ming Wang, John Kildea, Amal Zouaq

机构 * LAMA-WeST Lab(LAMA-WeST实验室) ; Polytechnique Montréal(蒙特利尔理工学院) ; Mila - Quebec AI Institute(魁北克AI研究院) ; Medical Physics Unit - RI-MUHC(医学物理部门 - RI-MUHC) ; McGill University(麦吉尔大学)

AI总结 MedHal是一个大规模医学幻觉检测评估数据集,整合多样来源与任务,提供带解释的标注样本,训练基线模型优于通用方法,助力高效评估并推动医学AI发展。

Comments The 5th Asia-Pacific Chapter of the Association for Computational Linguistics and the 15th International Joint Conference on Natural Language Processing, November 6-10, 2026, Hengqin, China

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.26347 2026-09-24 cs.CL cs.AI cs.LG 版本更新

TransBERT: A Framework for Synthetic Translation in Domain-Specific Language Modeling

TransBERT:面向特定领域语言建模的合成翻译框架

Julien Knafou, Luc Mottin, Anaïs Mottaz, Alexandre Flament, Patrick Ruch

机构 * HES-SO(瑞士西部应用科学与艺术大学) ; SIB, Swiss Institute of Bioinformatics(瑞士生物信息学研究所)

AI总结 TransBERT提出仅用合成翻译数据预训练语言模型,在法语生命科学领域达到最先进性能,并发布工具包、语料库和模型。

Comments 17 pages

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, pages 19338-19354

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.26539 2026-09-23 cs.CL 新提交

A retrospective analysis on the use of LLMs to study infant syntax learning

关于使用LLM研究婴儿句法学习的回顾性分析

Hélie Bazin, Anouk Barberousse, François Yvon

AI总结 本文回顾性分析LLMs在婴儿句法学习研究中的应用,通过认识论评估指出BabyLM等方法论假设及发育现实语料库对基准性能影响有限,揭示LLMs与婴儿学习者的计算差异。

Journal ref EMNLP 2026 Main Conference, ACL SIGDAT, Oct 2026, Budapest (Hungary), Hungary

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.26346 2026-09-23 cs.CL 新提交

Blaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament

跨党派指责:丹麦议会中的政治对比与责任归因

Markus Lundsfryd Jensen, Rune Egeskov Trust, Kenneth Christian Enevoldsen, Sara Kolding

机构 * Aarhus University(奥胡斯大学) ; Center for Humanities Computing(人文计算中心)

AI总结 本研究利用BlameBERT分类器与多水平模型分析丹麦议会1997-2026年责任归因,发现其呈香蕉形轨迹,近年上升且右翼意识形态极端化加剧指责,揭示政治话语的意识形态不对称硬化。

Comments 8 Pages + appendix (25 total) Main paper 4 figures 2 tables: Appendix 9 figures 10 tables. Model found here: https://huggingface.co/Lundsfryd/BlameBERT , dataset here: https://huggingface.co/datasets/runetrust/blame-folketinget-dk. Markus Lundsfryd Jensen and Rune Egeskov Trust have contributed equally. Paper will be submitted through ACL rolling review (ARR), we are aiming for COLING 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26842 2026-09-23 cs.LG cs.CL 版本更新

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training

MONA: 基于Nesterov加速的Muon优化器用于可扩展语言模型训练

Jiacheng Li, Jianchao Tan, Hongtao Xu, Jiaqi Zhang, Yifan Lu, Yerui Sun, Yuchen Xie, Xunliang Cai

机构 * Meituan(美团)

AI总结 提出MONA优化器,通过将Nesterov加速项集成到Muon的梯度处理流程中,实现曲率感知加速,从而帮助逃离尖锐局部最小值,并在1B到68B参数的混合专家预训练中取得更优收敛和下游任务性能。

Comments Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20106 2026-09-23 cs.LG cs.AI 版本更新

Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors

基于偏好向量的自适应有益有害对齐

Ren-Wei Liang, Chin-Ting Hsu, Chan-Hung Yu, Saransh Agrawal, Shih-Cheng Huang, Chieh-Yen Lin, Shang-Tse Chen, Kuan-Hao Huang, Shao-Hua Sun

机构 * National Taiwan University(国立台湾大学) ; Texas A&M University(德克萨斯A&M大学) ; Appier AI Research(Appier人工智能研究院) ; Graduate Institute of Communication Engineering, National Taiwan University(国立台湾大学通信工程研究所)

AI总结 本文提出偏好向量框架,通过模块化方法实现细粒度的用户可控偏好调整,提升大语言模型的有益性和无害性平衡。

Comments Accepted at The 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026), Rabat, Morocco

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04332 2026-09-23 cs.CR cs.LG 版本更新

The Challenge of Identifying the Origin of Black-Box Large Language Models

识别黑盒大语言模型来源的挑战

Ziqing Yang, Yixin Wu, Yun Shen, Wei Dai, Michael Backes, Yang Zhang

机构 * CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心) ; NetApp(NetApp公司) ; TikTok Inc.(TikTok公司)

AI总结 针对黑盒大语言模型来源识别难题,提出PlugAE方法,利用对抗性嵌入与自定义版权令牌,在准确性和鲁棒性上超越现有水印与指纹方法,并验证了其实用性。

Comments To Appear in the Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.23716 2026-09-22 cs.CL cs.AI cs.LG 新提交

STEVE: Stabilizing Textual Gradient-Based Prompt Optimization via Error-Driven Refinement and Regularized Verification

STEVE:通过错误驱动精炼与正则化验证稳定文本梯度提示优化

Yifan Xu, Yixuan Li, Xinzhuo Li, Yixin Gu, Yifan Shen, Lijun Yu, Haohan Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Google DeepMind(谷歌DeepMind)

AI总结 STEVE通过错误驱动精炼和正则化验证两种机制,稳定文本梯度提示优化,减少性能退化,在多个基准上生成更稳健的提示。

Comments Accepted to Findings of the Association for Computational Linguistics: AACL-IJCNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.23257 2026-09-22 cs.LG cs.CL 新提交

CTRL: Control-Based Time Series Forecasting with LLM-Guided Residual Learning

CTRL:基于控制的时间序列预测与LLM引导的残差学习

Minkyoung Kim, Daeun Ji, Yohan Lee, Beomsoo Kim, Beakcheol Jang

机构 * Yonsei University(延世大学)

AI总结 CTRL框架通过冻结骨干网络生成基础预测,利用LLM智能体分析残差并输出控制信号进行修正,实现语义推理与定量预测解耦,提升非平稳时间序列预测的鲁棒性。

Comments Published in Findings of the Association for Computational Linguistics: ACL 2026. 18 pages, 9 figures, 22 tables

Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pages 21952-21968

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.22977 2026-09-22 cs.LG cs.CL 新提交

Beyond Similarity: Coverage-Aware Prompt Selection for Time Series Forecasting with LLMs

超越相似性:面向LLM时间序列预测的覆盖感知提示选择

Daeun Ji, Minkyoung Kim, Dongkuk Kim, Yohan Lee, Beomsoo Kim, Beakcheol Jang

机构 * Yonsei University(延世大学)

AI总结 针对基于相似性提示选择导致时间序列预测偏置的问题,提出覆盖感知框架CASP-LLM,通过使用跟踪与饱和门控正则化器(无额外参数)在多数基准上匹配或超越现有方法。

Comments Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026. 24 pages, 8 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.22884 2026-09-22 cs.CL cs.AI 新提交

Block-Sparse Attention with Semantic-Geometric Decoupled Routing

块稀疏注意力与语义-几何解耦路由

Xinwei Long, Weigao Sun, Weibo Gao, Pengkun Jiao, Biqing Qi, Feida Zhu, Yiran Zhong, Steven Hoi, Bowen Zhou

机构 * Tsinghua University(清华大学) ; Alibaba Group(阿里巴巴集团) ; University of Science and Technology of China(中国科学技术大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

AI总结 针对长上下文推理中块稀疏注意力路由困难的问题,提出语义-几何解耦路由,通过分离语义与几何信息实现无训练高效路由,在4K-128K上下文接近全注意力精度,并显著加速。

Comments Technical report; Submitted to ACL ARR 2026 May

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.22793 2026-09-22 cs.CL cs.AI 新提交

Diagnose, Then Repair: A Two-Stage MQM-Guided Post-Editing Framework for Domain-Specific Machine Translation

诊断,然后修复:面向领域特定机器翻译的两阶段MQM引导的后编辑框架

Ji Hun Wang, Siyu Wu

机构 * Amazon(亚马逊)

AI总结 提出两阶段评估器引导的后编辑框架,将MQM诊断转化为最小修复,提升可控性并减少释义漂移,在七种语言和七个LLM上优于一阶段方法。

Comments Accepted to the ACL 2026 Industry Track and presented at ACL 2026

Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track), pp. 1683-1698, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10925 2026-09-22 cs.LG cs.CL 版本更新

Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation

找到你的最优教师:通过路由器引导的多教师蒸馏实现个性化数据合成

Hengyuan Zhang, Shiping Yang, Xiao Liang, Chenming Shang, Yuxuan Jiang, Chaofan Tao, Jing Xiong, Hayden Kwok-Hay So, Ruobing Xie, Angel X. Chang, Ngai Wong

机构 * The University of Hong Kong(香港大学) ; Simon Fraser University(西蒙菲莎大学) ; University of California, Los Angeles(加州大学洛杉矶分校) ; Dartmouth College(达特茅斯学院) ; University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校) ; Tencent(腾讯)

AI总结 本文提出PerSyn策略,通过'路由然后生成'范式为每个学生模型定制数据,提升学习效率。实验表明其在指令微调和数学推理中表现优异。

Comments ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22030 2026-09-22 cs.CL 版本更新

From Outliers to Topics in Language Models: Anticipating Trends in News Corpora

从语言模型中的离群点到主题:预测新闻语料库中的趋势

Evangelia Zve, Benjamin Icard, Alice Breton, Lila Sainero, Gauvain Bourgne, Jean-Gabriel Ganascia

机构 * LIP6, Sorbonne University, CNRS, France(LIP6,索邦大学,国家科学研究中心,法国)

AI总结 本文提出利用语言模型嵌入和累积聚类方法,将主题建模中的离群点视为新兴主题的弱信号,并在法英新闻数据上验证了离群点随时间演变为连贯主题的规律。

Comments presented at ICNLSP 2025; to appear in the ACL Anthology; received the Best Full Paper Award

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11582 2026-09-22 cs.CL 版本更新

AskQE: Question Answering as Automatic Evaluation for Machine Translation

AskQE:将问答作为机器翻译的自动评估方法

Dayeon Ki, Kevin Duh, Marine Carpuat

机构 * University of Maryland(马里兰大学) ; Johns Hopkins University(约翰霍普金斯大学)

AI总结 针对单语使用者无法判断外语机器翻译质量的实际问题,提出依托LLaMA-3 70B和蕴含事实引导问题生成的AskQE问答框架,在BioMQM数据集上其与人类评分的相关性及决策准确率优于现有质量估计指标。

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18673 2026-09-22 cs.CL

Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models

跨语言陷阱:自动探测多语言大语言模型的跨语言弱点

Zixiang Xu, Yanbo Wang, Yue Huang, Xiuying Chen, Jieyu Zhao, Meng Jiang, Xiangliang Zhang

机构 * MBZUAI ; University of Notre Dame(诺丁汉大学) ; University of Southern California(南加州大学)

AI总结 本文提出一种利用束搜索和LLM模拟生成双语问题对的方法,构建16种语言超6000对数据集,精准识别多语言模型跨语言弱点,并发现语言相似性与性能模式相关。

Comments ACL 2025. Code available at https://github.com/xzx34/Cross-Lingual-Pitfalls

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8254-8284, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.00852 2026-09-22 cs.CL 版本更新

VOLTA: Improving Generative Diversity by Variational Mutual Information Maximizing Autoencoder

VOLTA:通过变分互信息最大化自编码器提升生成多样性

Yueen Ma, Dafeng Chi, Jingjing Li, Kai Song, Yuzheng Zhuang, Irwin King

机构 * The Chinese University of Hong Kong(香港中文大学) ; Huawei Noah’s Ark Lab(华为诺亚方舟实验室)

AI总结 本文提出VOLTA框架,通过交叉注意力连接Transformer与VAE,并引入InfoGAN风格潜码,在保持生成质量的同时显著提升自然语言生成的多样性。

Comments 15 pages. Published in Findings of the Association for Computational Linguistics: NAACL 2024

Journal ref Findings of the Association for Computational Linguistics: NAACL 2024, pages 364-378

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.21672 2026-09-21 cs.AI cs.CL 新提交

Accelerating Dense LLMs via L0-regularized Mixture-of-Experts

通过L0正则化混合专家加速稠密大语言模型

Zhenyu Zhang, Jiudong Yang, Zhaowen Tao, Meng Chen

机构 * YZW ; FuTu AI(富途人工智能) ; Wise AI

AI总结 本文提出L0-MoE,利用L0正则化的轻量级混合专家方法,在不明显损失性能的前提下加速稠密大语言模型,实现高达2.5倍推理加速,超越现有加速基线。

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18519 2026-09-21 cs.AI

LLM Safety From Within: Detecting Harmful Content with Internal Representations

从内部检测有害内容:利用内部表示进行LLM安全

Difan Jiao, Yilun Liu, Ye Yuan, Zhenwei Tang, Linfeng Du, Haolun Wu, Ashton Anderson

机构 * University of Toronto(多伦多大学) ; McGill University(麦吉尔大学) ; LMU Munich(慕尼黑路德维希-马克西米利安大学)

AI总结 SIREN通过利用内部层特征提升有害内容检测性能,相比现有模型更高效且泛化能力强。

Comments 17 pages,10 figures,6 tables

Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.19553 2026-09-18 cs.CL 新提交

From Parameters to Behaviors: A Survey of Model Fusion for Large Language Models

从参数到行为:大型语言模型融合综述

Shuo Cai, Yanggan Gu, Zihao Wang, Yuanyi Wang, Yibo Yan, Wenjun Wang, Yuhang Liu, Guanghao Zhu, Sirui Huang, Ming Li, Hongxia Yang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; The Hong Kong Polytechnic University(香港理工大学) ; PolyU-Daya Bay Technology and Innovation Research Institute(香港理工大学大亚湾技术创新研究院) ; The Chinese University of Hong Kong(香港中文大学)

AI总结 本文综述了大型语言模型融合领域,将其划分为参数级、表示级和行为级三个层次,并回顾了相关指标、基准与应用,旨在为该领域提供系统性图谱并指引未来研究方向。

Comments 25 pages, 4 figures. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.19238 2026-09-18 cs.CL cs.AI 新提交

YNU-HPCC at SemEval-2025 Task 11: Bridging the Gap in Text-Based Emotion Using Multiple Prediction Headers

YNU-HPCC 在 SemEval-2025 任务 11 中的参与:使用多预测头弥合基于文本的情感差距

Hao Yang, Jin Wang, Xuejie Zhang

AI总结 YNU-HPCC 团队在 SemEval-2025 任务 11 中采用 RoBERTa 模型,通过改进输出头并翻译数据集,实现了 0.44 的官方分数,并发现单预测头与统一英文训练更优。

Journal ref Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), pages 83-89, Vienna, Austria. Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.19154 2026-09-18 cs.CL 新提交

Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry

Neo-Classic:古典诗歌语言-审美推理评估基准

Han Zhang, Zihan Gu, Zhiyuan Wang, Tianyi Ma, Jiacheng Lu, Xinyan Zhang, Yuhao Wei, Cheng Hua

机构 * Shanghai Jiao Tong University(上海交通大学) ; Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)

AI总结 针对LLMs在古典诗歌基准上依赖预训练模式而非真正推理的问题,提出Neo-Classic基准,含当代专家创作的格律诗与探针测试,发现模型存在20-50%的性能差距及话语级排序准确率低(0-13%)的局限。

Comments Published in the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2026

Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 27442-27465, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.16814 2026-09-18 cs.AI cs.IR 版本更新

Can We Do Interpretable NLI with Graphs Based on Atomic Propositions?

我们能基于原子命题的图进行可解释的自然语言推理吗?

Younes Boufouss, Luc Pommeret, Thomas Gerald, Patrick Paroubek, Sophie Rosset

机构 * Université Paris-Saclay(巴黎-萨克雷大学) ; CNRS(法国国家科学研究中心) ; Laboratoire Interdisciplinaire des Sciences du Numérique(数字科学跨学科实验室)

AI总结 本文提出一种完全基于图的NLI流程,将句子分解为原子命题并构建ConceptNet图,输入0.8B参数模型,在SNLI上达89.7%准确率,虽略低于文本模型,但揭示了可解释性的代价,且图与文本模态互补。

Journal ref AKBC @ EMNLP, ACL, Oct 2026, Budapest, Hungary

详情

展开后加载摘要…

URL PDF HTML 收藏
↑