arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

Annual Meeting of the Association for Computational Linguistics · 会议 · Natural Language Processing

共收录 10367
2609.03620 2026-09-04 eess.AS cs.AI cs.SD 新提交

ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection

ToolDF:面向混合真实性音频深度伪造检测的工具集成推理框架

Taewoo Kim, Young Han Lee, Nam In Park, Chanwoo Kim

机构 * Multi-Modal Research Center, KETI(韩国电子通信研究院多模态研究中心) Korea University(高丽大学) Digital Analysis Section, National Forensic Service(国家法医服务局数字分析科)

AI总结 本研究针对混合真实性音频深度伪造检测问题,提出ToolDF工具集成推理框架,引入混合真实性ADD基准,在复合检测任务上性能优于基线,同时提供可解释证据。

Comments To appear in Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.03955 2026-09-04 cs.CL cs.LG 新提交

Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs

面向代码大语言模型的健全性与对抗性测试用例生成的两阶段强化学习

Jiacheng Xu, Wentao Zhang, Zhiyi Lyu, Fuxiang Zhang, Chaojie Wang, Yang Liu, Bo An

机构 * Nanyang Technological University(南洋理工大学) Skywork AI(天工智源)

AI总结 针对代码LLM测试用例稀缺问题,提出两阶段RL框架TCS,在TACO和LiveCodeBench上提升pass@1与答案选择能力,且可用于其他LLM输出选择。

Comments 21 pages, 7 figures. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.03953 2026-09-04 cs.CL 新提交

Beyond Majority Vote: Multi-Perspective Adjudication for Medical Hallucination Detection

超越多数投票:面向医学幻觉检测的多视角裁决

Joe Cecil, Marjorie Freedman

机构 * Information Sciences Institute(信息科学研究所) University of Southern California(南加州大学)

AI总结 本研究针对医学聊天机器人响应开展多视角标注研究,结合首轮标注、LLM-as-a-Judge候选发现及医学专家与证据裁决,发现单轮幻觉基准易低估事实错误,多轮裁决可提升覆盖范围。

Comments 34 pages, 6 figures, to be published in Findings of the ACL: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.03150 2026-09-04 cs.LG cs.CL 新提交

Routing Is Not Enough: Diagnosing Intra-Adapter Subspace Contention in MoE+LoRA Fine-Tuning

路由还不够:诊断MoE+LoRA微调中的适配器内子空间竞争

Mehreen Hossain Chowdhury, Nowshin Mahjabin, Ahmed Shafin Ruhan, Md Azam Hossain, Abu Raihan Mostofa Kamal, Md Tahmid Rahman Laskar

机构 * Islamic University of Technology(伊斯兰理工大学) York University(约克大学) Dialpad Inc.(Dialpad公司)

AI总结 该研究针对MoE+LoRA微调中路由分离无法防止负迁移的问题,提出SpawnLoRA方法,通过在MoE专家内动态添加门控子适配器减少负迁移,在Phi-tiny-MoE-instruct等模型上验证了其有效性。

Comments 13 pages, 1 figure, 16 tables. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20721 2026-09-04 cs.CL cs.AI cs.HC 版本更新

User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios

用户感知与代理LLM评判:隐私与帮助性在隐私敏感场景中的LLM响应

Xiaoyuan Wu, Roshni Kaushik, Wenkai Li, Lujo Bauer, Koichi Onoue

机构 * Fujitsu Research of America Inc.(富士通美国研究所) Carnegie Mellon University(卡内基梅隆大学)

AI总结 研究发现用户对LLM响应的隐私和帮助性感知与代理LLM评判存在显著差异,需加强用户为中心的评估以提升隐私保护与实用性平衡。

Comments Published as a main conference paper at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18407 2026-09-04 cs.CL cs.AI cs.LG 版本更新

AgentRM: Enhancing Agent Generalization with Reward Modeling

AgentRM:通过奖励建模增强智能体泛化能力

Yu Xia, Jingru Fan, Weize Chen, Siyu Yan, Xin Cong, Zhong Zhang, Yaxi Lu, Yankai Lin, Zhiyuan Liu, Maosong Sun

机构 * Tsinghua University(清华大学)

AI总结 本研究提出AgentRM奖励模型,通过显式、隐式建模及LLM评判三种方式构建,结合N选最优采样等方法,在9个任务上提升智能体泛化与专门化能力,效果优于现有模型。

Comments Published in ACL 2025 Main Conference (Long Papers)

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.01610 2026-09-03 cs.IR 新提交

Making Revisions Understandable: A Survey of Edit Intentions, Methods, and Applications

让修订可理解:编辑意图、方法与应用综述

Fangping Lan, Qi Zhang, Eduard Dragut

AI总结 该综述从编辑意图视角整合文本修订研究,梳理了相关数据集、方法与应用,明确了该领域的开放研究方向。

Comments 17 pages, 6 figures, 4 tables, accepted by Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.01647 2026-09-03 cs.LG astro-ph.IM 新提交

Efficient Context-Limited Telescope Bibliography Classification for the WASP-2025 Shared Task Using SciBERT

面向WASP-2025共享任务的基于SciBERT的高效上下文受限望远镜文献分类

Madhusudhana Naidu

AI总结 针对WASP-2025共享任务中望远镜文献分类的人工资源消耗问题,提出基于SciBERT的高效方法,在512token上下文限制下获0.89宏F1值,居任务排行榜榜首,为科学文本整理效率边界提供见解。

Comments 3 pages, 2 tables. 1st place system description for the TRACS shared task at WASP 2025 (Third Workshop for Artificial Intelligence for Scientific Publications), co-located with IJCNLP-AACL 2025. Published version: https://aclanthology.org/2025.wasp-main.21/ . Code: https://github.com/E0NIA/TRACS-WASP-2025-1st-Place

Journal ref Proceedings of the Third Workshop for Artificial Intelligence for Scientific Publications (WASP 2025), pages 192-194, Mumbai, India and virtual, December 2025. Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30930 2026-09-03 cs.HC cs.AI cs.CL cs.CY 版本更新

TUX: Measuring Human--AI Tacit Understanding

TUX:衡量人机默契理解

Yueshen Li, Hanyi Min, Vedant Das Swain, Koustuv Saha

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) New York University(纽约大学)

AI总结 通过光谱放置任务和TUX指数,量化人类与LLM之间的默契理解,发现人格特征影响对齐程度。

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16749 2026-09-03 cs.CL cs.LG 版本更新

Probing Cultural Signals in Large Language Models through Author Profiling

通过作者画像探测大语言模型中的文化信号

Valentin Lafargue, Ariel Guerra-Adames, Emmanuelle Claeys, Elouan Vuichard, Jean-Michel Loubes

机构 * IMT, Toulouse, France(法国图卢兹IMT学院) INRIA Bordeaux, France(法国波尔多INRIA) ANITI 2, Toulouse, France(法国图卢兹ANITI) IRIT, Toulouse, France(法国图卢兹IRIT) Université de Bordeaux, Bordeaux, France(法国波尔多大学) BPH, Inserm, France(法国Inserm BPH) CNRS IRL CROSSING, Adelaide, Australia(澳大利亚阿德莱德CNRS IRL CROSSING)

AI总结 研究通过零样本设置评估大语言模型能否从歌曲歌词中推断歌手性别和种族,发现模型存在系统性文化偏见,提出两种公平性指标量化差异。

Journal ref In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026), Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16848 2026-09-03 cs.CL 版本更新

Mediocrity is the key for LLM as a Judge Anchor Selection

中庸是LLM作为判断锚点选择的关键

Shachar Don-Yehiya, Asaf Yehudai, Leshem Choshen, Omri Abend

机构 * The Hebrew University of Jerusalem(海法大学) IBM Research(IBM研究院) MIT(麻省理工学院) MIT-IBM Watson AI Lab(麻省理工-IBM沃森人工智能实验室)

AI总结 本文研究了LLM作为判断锚点选择对结果可靠性的影响,发现选择不当的锚点会显著降低与人类排名的相关性,并提出可操作的推荐方案。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08590 2026-09-03 cs.CL cs.LG 版本更新

GMTRouter: Personalized LLM Router over Multi-turn User Interactions

GMTRouter:面向多轮用户交互的个性化大语言模型路由

Yihang Sun, Encheng Xie, Tao Feng, Jiaxuan You

机构 * Antiquus S. Hippocampus, Natalia Cerebro & Amelie P. Amygdale Department of Computer Science Cranberry-Lemon University Pittsburgh, PA 15213, USA(计算机科学系,Cranberry-Lemon大学) Ji Q. Ren & Yevgeny LeNet Department of Computational Neuroscience University of the Witwatersrand Joburg, South Africa(计算神经科学系,沃特瓦特斯兰大学)

AI总结 针对现有LLM路由个性化不足、难捕捉用户-LLM交互及用户偏好数据稀缺的问题,提出GMTRouter,通过异质图建模多轮交互并结合定制采样机制,在少样本下实现个性化,性能优于基线且可适配新用户。

Comments Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14845 2026-09-03 cs.LG cs.CL 版本更新

Prompting the Unknown: Understanding Response Uncertainty in Large Language Models

提示未知:理解大语言模型中的响应不确定性

Ze Yu Zhang, Arun Verma, Finale Doshi-Velez, Bryan Kian Hsiang Low

机构 * Alibaba Group(阿里巴巴集团) Harvard University(哈佛大学) National University of Singapore(新加坡国立大学)

AI总结 该研究针对大语言模型响应不确定性与提示的关系问题,提出含四类不确定性来源的提示-响应概念模型,经实验验证模型及相关理论结果。

Comments Accepted at ACL Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.01195 2026-09-02 cs.CL 新提交

CaRL-EM: Cost-Aware Reinforcement Learning for Entity Matching with LLMs

CaRL-EM:面向大语言模型的实体匹配的成本感知强化学习

Chaohui Guo, Michel Klein, Zhisheng Huang

机构 * Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)

AI总结 该研究针对实体匹配的LLM方法灵活性不足、忽略推理成本的问题,提出CaRL-EM强化学习控制器,自适应选择算子与模型规模,在7个基准上实现更优质量-成本权衡与零样本迁移。

Comments Accepted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2609.00845 2026-09-02 cs.AI 新提交

Towards Generalizable Visually Grounded Exploration of Household Devices

面向家庭设备的可泛化视觉接地探索

Linhao Zheng, Zeming Liu, Wangke Chen, Li Zeng, Wanxiang Che, Heyan Huang, Yuhang Guo

机构 * School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院) Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology(哈尔滨工业大学社会计算与交互机器人研究中心) School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)

AI总结 本文针对现有具身探索范式泛化能力不足的问题,构建了基准VGEBench及逻辑驱动状态机框架,实验发现现有VLMs在语义知识转物理执行和长程状态跟踪上存在挑战。

Comments Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00207 2026-09-02 cs.CL 版本更新

Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection

缩小英阿医学知识鸿沟:通过因果层选择实现针对性低秩适配

Chaimae Abouzahir, Musa Khan, Hala Ali-Hassan, Congbo Ma, Khaled Saleh, Yousra Sadqi, Jihad Mallat, Walid Al-Eisawi, Nizar Habash, Farah E. Shamout

机构 * New York University Abu Dhabi(纽约大学阿布扎比分校) Cleveland Clinic Abu Dhabi(克利夫兰诊所阿布扎比分校)

AI总结 该研究针对阿拉伯语医学LLMs性能弱于英文的问题,提出TLoRA方法并构建阿拉伯医学对话基准,通过机制诊断实现针对性适配,提升了阿拉伯语医学任务性能。

Journal ref Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09837 2026-09-02 cs.HC cs.AI 版本更新

Self-EmoQ: Plutchik-Guided Value-based Planning to Drive Streaming Emotional TTS

Self-EmoQ: 基于Plutchik引导的价值规划驱动流式情感TTS

Yue Zhao, Hongyan Li, Yong Chen, Luo Ji

机构 * Geely AI Lab(地平线人工智能实验室)

AI总结 提出一种情感规划框架,通过强化学习训练LLM模块,在文本生成前确定情感,以驱动流式TTS,结合Plutchik情感理论进行混合奖励,实验表明在情感确定和响应质量上优于基线。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01972 2026-09-02 cs.CL cs.AI cs.LG 版本更新

Hidden State Poisoning Attacks against Mamba-based Language Models

针对基于Mamba的语言模型的隐藏状态污染攻击

Alexandre Le Mercier, Chris Develder, Thomas Demeester

机构 * IDLab–T2K, Ghent University–imec(IDLab–T2K,根特大学–imec)

AI总结 本文研究了特定短输入短语对Mamba模型造成部分遗忘效应的攻击现象,揭示了SSM模型对隐藏状态污染攻击的脆弱性,并展示了攻击对Jamba和Nemotron-3-Nano模型的影响。

Comments 27 pages, 4 figures. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05364 2026-09-02 cs.CL

The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures

Alexander M. Fichtl, Jeremias Bohn, Josefin Kelber, Edoardo Mosca, Georg Groh

Comments 21 pages, 2 figures, 2 tables

Journal ref In Proceedings of The Big Picture v2 - Crafting a Research Narrative, pages 60-81, San Diego, CA, USA. Association for Computational Linguistics. 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11857 2026-09-02 cs.CL cs.AI cs.LG 版本更新

SupraTok: Cross-Boundary Tokenization for Enhanced Language Model Performance

SupraTok:用于提升语言模型性能的跨边界分词

Andrei-Valentin Tănase, Elena Pelican

机构 * Faculty of Mathematics and Computer Science, "Ovidius" University of Constanta(数学与计算机科学学院,"奥维迪乌"康斯坦察大学)

AI总结 针对语言建模中分词受空格边界限制的问题,提出SupraTok分词器,通过跨空格边界的模块化设计实现压缩率、训练速度及下游任务性能的提升。

Comments Published at TACL. Substantially revised from v1/v2 preprints

Journal ref Transactions of the Association for Computational Linguistics, Vol. 14, pp. 1787-1802 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.31014 2026-09-01 cs.CL 新提交

Evidence-Bounded Mental Health Reasoning from Heterogeneous Speech Protocols

基于异构语音协议的证据受限心理健康推理

Chengyuan Gao, Jiang Wu, Tao Lu, Jiayan Guo, Mingkun Xu, Tianyi Zang, Shangyang Li

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Harbin Institute of Technology(哈尔滨工业大学) Renyixun Health Technology Co., Ltd.(仁医讯健康科技有限公司) GDIIST(广东省智能制造研究所) Tencent(腾讯)

AI总结 本研究针对现有心理健康筛查模型的证据边界问题,提出协议感知的证据控制框架EviBound,在抑郁症筛查任务中达到0.8658的AUROC,且实现零宣称违规。

Comments Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026. Long paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30842 2026-09-01 cs.CL 新提交

Thesis Proposal: Toward a Human-Centered and Perspective-Aware Framework for Reproducible ML Evaluation and AI Alignment

论文提案:面向可复现的机器学习评估与AI对齐的以人为中心且视角感知的框架

Deepak Pandita, Christopher M. Homan

机构 * Rochester Institute of Technology(罗切斯特理工学院)

AI总结 针对当前LLM评估掩盖少数视角、加剧AI可复现性危机的问题,论文提出一种以人为中心且视角感知的框架,用于可复现的机器学习评估与AI对齐。

Comments Published at ACL SRW 2026: https://aclanthology.org/2026.acl-srw.74/

Journal ref Proc. ACL SRW (2026) vol. 4 pp. 827-843

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30649 2026-09-01 cs.CL cs.CV 新提交

Where Identity Lives: Localized, Retain-Free Identity Unlearning in Multimodal Large Language Models

身份信息的留存位置:多模态大语言模型中无保留集的本地化身份遗忘

Kangwook Ko, Jaehyuk Jang, Wonjun Lee, Hee-Seon Kim, Changick Kim

机构 * KAIST(韩国科学技术院)

AI总结 该研究针对多模态大语言模型身份遗忘需依赖难获取保留集的问题,提出PAVA方法,定位解码器早期到中期MLPs并结合遗忘损失与视觉属性锚定,在基准上取得良好遗忘-保留权衡。

Comments Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30270 2026-09-01 cs.CL 新提交

Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts

读懂房间,读懂图像:理解多模态视觉语境中的间接言语行为

Jaehee Kim, Ji Hoon Chung, Seoyoon Park, Unsol Kim, Kyungwon Park, Ji Hak Kim, Yi-Jun Chen, Hansaem Kim

机构 * Yonsei University(延世大学) LG AI Research(LG人工智能研究院)

AI总结 该研究推出多模态基准READI,以视觉语用问答任务评估间接言语行为理解,发现先进多模态模型在视觉接地间接言语行为上表现差,性能随间接性提升而下降,凸显相关基准的必要性。

Comments Accepted to Findings of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.30033 2026-09-01 cs.CL cs.AI 新提交

"Act Like a 5th Grader" is Not Enough: Bounding Knowledge in LLM-Based User Simulators

仅模仿五年级学生还不够:基于大语言模型的用户模拟器中的知识边界

Krisztian Balog, Arild Michel Bakken

机构 * University of Stavanger(斯塔万格大学)

AI总结 该研究针对基于大语言模型的用户模拟器存在的“超人偏差”问题,提出认知受限用户模拟器(CBUS)框架,通过明确建模认知边界提升模拟人类阅读行为的保真度。

Comments Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29995 2026-09-01 cs.CL cs.AI 新提交

Generating Clinical Vignettes that Preserve Cognitive Formulations

生成保留认知构想的临床案例 vignettes

Amit Oren, Nimrod Hertz-Palmor, Dean Ariel, Guy Laban

机构 * Ben-Gurion University of the Negev(内盖夫本-古里安大学) University of Cambridge(剑桥大学) Clalit Health Services(克拉利特医疗服务机构) Tel Aviv University(特拉维夫大学) School of Brain Sciences and Cognition, Ben-Gurion University of the Negev(内盖夫本-古里安大学脑科学与认知学院) Azrieli National Center for Autism and Neurodevelopment Research(阿兹里利国家自闭症与神经发育研究中心)

AI总结 本研究提出基于理论的 FORMA 框架,将认知模型编译为图来生成符合临床结构的创伤后应激障碍案例 vignettes,经评估其质量高且人口统计学差异小,可作为合成临床文本生成的可审计规范。

Comments Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026. Code and data: https://github.com/Amit-Oren/FORMA

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29865 2026-09-01 cs.CL 新提交

ManGo: Manga Active Narrative Grounding Optimization

ManGo:漫画主动叙事定位优化

Hao Qiu, Junyan Wang, Zheyuan Liu, Lei Fan, Hong Jia, Lianbo Guo, Zhulin Tao

机构 * Communication University of China(中国传媒大学) Australian Institute for Machine Learning, Adelaide University(阿德莱德大学澳大利亚机器学习研究所) University of New South Wales(新南威尔士大学) University of Auckland(奥克兰大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 本研究针对漫画视觉问答的分镜叙事结构问题,提出无监督框架ManGo,通过主动叙事草图结合双奖励的组相对策略训练,在标准基准上取得最优性能。

Comments 16 pages, 9 figures, 7 tables. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29623 2026-09-01 cs.CL cs.AI 新提交

MI-Distillation: Selecting from Model-Interpolated Instruct-Reasoning Data Spectrum for Chain-of-Thought Distillation

MI-Distillation:从模型插值的指令推理数据频谱中选择以进行思维链蒸馏

Yangsong Lan, Renkai Hu, HongKai Zheng, Bo Zhang, Renzhi Wang, Hongliang Dai, Piji Li

机构 * The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(教育部脑机智能技术重点实验室)

AI总结 本研究提出MI-Distillation框架,通过模型插值构建指令推理数据频谱,引入SeqLSS选择合适轨迹,在推理基准上显著提升了小模型的思维链蒸馏效果。

Comments Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29613 2026-09-01 cs.CL cs.LG 新提交

Cross-lingual Functional Vectors for Emotion Detection in Large Language Models

用于大语言模型情感检测的跨语言功能向量

Jieying Xue, Phuong Minh Nguyen, Minh Le Nguyen, Shogo Okada

机构 * Japan Advanced Institute of Science and Technology(日本先进科学技术学院)

AI总结 本研究以多语言多标签情感识别为基准,探究跨语言功能向量(FVs)在大语言模型中的应用,发现FVs可跨语言提升情感检测性能,是轻量可迁移的多语言任务适配机制。

Comments Findings of the Association for Computational Linguistics: EMNLP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.29345 2026-09-01 cs.AI cs.CL 新提交

BIRD-History: A Benchmark for History-Driven Text-to-SQL with Fine-Grained Knowledge Annotations

BIRD-History:带有细粒度知识标注的历史驱动型Text-to-SQL基准测试

Yunfan Zhou, Qiming Shi, Yizhou Yang, Di Weng, Yingcai Wu

机构 * Zhejiang University(浙江大学)

AI总结 本研究推出带细粒度知识标注的BIRD-History基准测试,针对现有Text-to-SQL系统无法利用历史查询日志处理隐含领域知识的问题,提出插件式检索器,可提升多个Text-to-SQL系统的未明确查询处理性能。

Comments Accepted at Findings of the Association for Computational Linguistics: EMNLP, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏