arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2026-02-17 至 2026-02-17 共收录 166 信号源:cs.CL, cs.AI, cs.LG

1. 推理评测 41 篇

2510.09510 2026-02-17 cs.IR 78%

MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval

MRMR:一个面向推理密集型多模态检索的现实且专家级多学科基准

Siyue Zhang, Yuan Gao, Xiao Zhou, Yilun Zhao, Tingyu Song, Arman Cohan, Anh Tuan Luu, Chen Zhao

专题命中 推理评测 :reasoning(title,abstract)

AI总结 MRMR是一个面向推理密集型多模态检索的现实且专家级多学科基准,通过多领域查询和混合模态数据推动检索模型的改进。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13891 2026-02-17 cs.SD cs.AI 77%

GSRM: Generative Speech Reward Model for Speech RLHF

GSRM:生成式语音奖励模型用于语音强化学习反馈机制

Maohao Shen, Tejas Jayashankar, Osama Hanna, Naoyuki Kanda, Yancheng Wang, Kateřina Žmolíková, Ruiming Xie, Niko Moritz, Anfeng Xu, Yashesh Gaur, Gregory Wornell, Qing He, Jilong Wu

机构 * Meta Superintelligence Labs(Meta超智能实验室) Massachusetts Institute of Technology(麻省理工学院) Arizona State University(亚利桑那州立大学) University of Southern California(南加州大学)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);verifier(abstract);分类 cs.AI

AI总结 GSRM通过生成式语音奖励模型提升语音生成的自然度,利用可解释的推理链实现更准确的自然度评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07672 2026-02-17 cs.SE cs.AI cs.LG cs.PL cs.SC 73%

Debugging code world models

调试代码世界模型

Babak Rahmani

机构 * Tübingen AI Center, University of Tübingen, Microsoft Research(图宾根人工智能中心、图宾根大学、微软研究院)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 代码世界模型通过模拟程序执行来预测运行时状态,研究发现其在长周期状态跟踪中存在token预算耗尽和字符串状态处理的局限性,提出改进监督和状态表示的方向。

Comments 8 pages, 4 figures, under review in conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03669 2026-02-17 cs.LG cs.CL 73%

Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning

令牌隐藏奖励:在组相对深度强化学习中引导探索-利用

Wenlong Deng, Yi Ren, Yushu Li, Boying Gong, Danica J. Sutherland, Xiaoxiao Li, Christos Thrampoulidis

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute(向量研究所) Amii(阿米人工智能研究所) UC Berkeley(加州大学伯克利分校)

专题命中 推理评测 :reasoning(abstract);math reasoning(abstract);分类 cs.CL、cs.LG

AI总结 令牌隐藏奖励通过调整组相对策略优化的学习信号,引导探索-利用平衡,提升大语言模型在推理任务中的表现。

Comments Full version of submission to 2nd AI for Math Workshop@ ICML 2025 (best paper)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14017 2026-02-17 cs.LG 70%

S2SServiceBench: A Multimodal Benchmark for Last-Mile S2S Climate Services

S2SServiceBench:一个多模态基准用于最后一公里S2S气候服务

Chenyue Li, Wen Deng, Zhuotao Sun, Mengxi Jin, Hanzhe Cui, Han Li, Shentong Li, Man Kit Yu, Ming Long Lai, Yuhao Yang, Mengqian Lu, Binhang Yuan

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Nanjing University of Information Science and Technology(南京信息工程大学) Beijing Normal University(北京师范大学)

专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.LG

AI总结 S2SServiceBench是一个多模态基准,用于评估S2S气候服务中最后一公里的可靠性,通过10种服务产品和1000多个评估项目,揭示了多模态大语言模型在不确定性下的决策推理挑战。

Comments 18 pages, 3 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19666 2026-02-17 cs.CL 70%

RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams

RoD-TAL:罗马尼亚驾照考试问答的基准测试

Andrei Vlad Man, Răzvan-Alexandru Smădu, Cristian-George Craciun, Dumitru-Clementin Cercel, Florin Pop, Mihaela-Claudia Cercel

机构 * National University of Science and Technology POLITEHNICA Bucharest, Faculty of Automatic Control and Computers(波兰技术大学布加勒斯特分校) Technical University of Munich(慕尼黑技术大学) National Institute for Research & Development in Informatics - ICI Bucharest(信息研究所-布加勒斯特) Paris 1 Panthéon-Sorbonne University(巴黎1大学) University of Bucharest(布加勒斯特大学)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL

AI总结 RoD-TAL是一个用于评估大型语言模型和视觉语言模型在罗马尼亚驾照法律问答中性能的多模态基准数据集,通过文本和图像问答任务验证了领域微调和推理优化对考试通过率的影响。

Comments 41 pages, 30 figures, Accepted by the Findings of EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13292 2026-02-17 cs.AI 70%

Mirror: A Multi-Agent System for AI-Assisted Ethics Review

镜:一种用于AI辅助伦理审查的多智能体系统

Yifan Ding, Yuhui Shi, Zhiyan Li, Zilong Wang, Yifeng Gao, Yajun Yang, Mengjie Yang, Yixiu Liang, Xipeng Qiu, Xuanjing Huang, Xingjun Ma, Yu-Gang Jiang, Guoyu Wang

机构 * Institute of Trustworthy Embodied AI(可信具身人工智能研究所) Institute of Technology Ethics for Human Future(人类未来技术伦理研究所) School of Philosophy(哲学学院) School of Life Sciences(生命科学学院) Ethics Committee of Zhongshan Hospital(中山医院伦理委员会) Department of Cardiology, Zhongshan Hospital of Fudan University, Institute of Cardiovascular Diseases, National Clinical Research Centre for Interventional Medicine(复旦大学中山医院心内科、心血管疾病研究所、介入医学临床研究中心) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI

AI总结 Mirror是一种多智能体系统,通过整合伦理推理和多智能体协商,提升AI辅助伦理审查的效率和专业性。

Comments 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05316 2026-02-17 cs.LG cs.AI cs.CL 67%

Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay

通过难度目标在线数据选择和回放提升大语言模型强化微调的数据效率

Yifan Sun, Jingyan Shen, Yibin Wang, Tianyu Chen, Zhendong Wang, Mingyuan Zhou, Huan Zhang

机构 * UIUC(伊利诺伊大学香槟分校) New York University(纽约大学) University of Texas at Austin(得克萨斯大学奥斯汀分校) Microsoft(微软)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出通过难度目标在线数据选择和回放机制提升大语言模型强化微调的数据效率,实验表明可减少62%的微调时间并保持同等性能。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02158 2026-02-17 cs.CL cs.AI cs.LG physics.geo-ph 67%

FormationEval, an open multiple-choice benchmark for petroleum geoscience

FormationEval,一个用于石油地球科学的开源多选题基准测试

Almaz Ermilov

机构 * UiT The Arctic University of Norway(UiT北极大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 FormationEval是一个开源多选题基准测试,用于评估语言模型在石油地球科学领域的表现,涵盖七个领域,包含505个问题,展示了不同模型的准确率及领域差异。

Comments v2: expanded related work, added validation details, difficulty-domain table, community feedback website (at https://www.formationeval.no). 28 pages, 8 figures, 11 tables. Benchmark and code at https://github.com/AlmazErmilov/FormationEval-an-Open-Benchmark-for-Oil-Gas-Geoscience-MCQ-Evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05430 2026-02-17 cs.CL cs.AI cs.IR cs.LG 67%

ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answering

ArtistMus: 一个全球多样、以艺术家为中心的基准,用于检索增强的音乐问答

Daeyong Kwon, SeungHeon Doh, Juhan Nam

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 ArtistMus提出一个全球多样、以艺术家为中心的基准,通过检索增强生成技术提升音乐问答的准确性和上下文推理能力。

Comments Accepted to LREC 2026. This work is an evolution of our earlier preprint arXiv:2507.23334

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18053 2026-02-17 cs.RO 67%

V2V-GoT: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-of-Thoughts

V2V-GoT: 基于多模态大语言模型和思维图的车对车协同自动驾驶

Hsu-kuang Chiu, Ryo Hachiuma, Chien-Yi Wang, Yu-Chiang Frank Wang, Min-Hung Chen, Stephen F. Smith

机构 * NVIDIA Carnegie Mellon University(卡内基梅隆大学)

专题命中 推理评测 :reasoning(abstract);planning(abstract)

AI总结 本文提出V2V-GoT框架,结合多模态大语言模型和图-思维方法,提升车对车协同自动驾驶的感知、预测和规划能力。

Comments Accepted by ICRA 2026 (IEEE International Conference on Robotics and Automation). Project: https://eddyhkchiu.github.io/v2vgot.github.io/ Code: https://github.com/eddyhkchiu/V2V-GoT Dataset: https://huggingface.co/datasets/eddyhkchiu/V2V-GoT-QA

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18776 2026-02-17 cs.CL cs.AI cs.LG 67%

AECBench: A Hierarchical Benchmark for Knowledge Evaluation of Large Language Models in the AEC Field

AECBench: 一个用于评估大型语言模型在AEC领域知识能力的分层基准

Chen Liang, Zhaoqi Huang, Haofen Wang, Fu Chai, Chunying Yu, Huanhuan Wei, Zhengjie Liu, Yanpeng Li, Hongjun Wang, Ruifeng Luo, Xianzhong Zhao

机构 * College of Civil Engineering, Tongji University(同济大学土木工程学院) Shanghai Qi Zhi Institute(上海启智研究院) Arcplus Group East China Architectural Design & Research Institute Co., Ltd.(Arcplus集团东部建筑设计与研究研究院有限公司) College of Design and Innovation, Tongji University(同济大学设计与创新学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 AECBench提出一个分层基准,评估大型语言模型在AEC领域知识能力,揭示模型在复杂推理和文档生成方面的不足。

Comments Accepted by Advanced Engineering Informatics. Code and data available at: https://github.com/ArchiAI-LAB/AECBench

Journal ref Advanced Engineering Informatics, Vol. 71, Article 104314 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14989 2026-02-17 cs.CV cs.AI cs.LG 62%

ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery

ThermEval: 一种用于评估视觉语言模型在热成像上的性能的结构化基准

Ayush Shrivastava, Kirtan Gangani, Laksh Jain, Mayank Goel, Nipun Batra

机构 * Indian Institute of Technology, Gandhinagar, India(印度理工学院加尔各答分校) Carnegie Mellon University, Pittsburgh, USA(卡内基梅隆大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 ThermEval通过结构化基准评估视觉语言模型在热成像上的性能,揭示其在温度推理和颜色映射转换上的不足,推动热视觉语言模型的发展。

Comments 8 Pages with 2 figures of main content. 2 pages of References. 10 pages of appendix with 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14158 2026-02-17 cs.CL cs.AI cs.MA 62%

A Multi-Agent Framework for Medical AI: Leveraging Fine-Tuned GPT, LLaMA, and DeepSeek R1 for Evidence-Based and Bias-Aware Clinical Query Processing

面向医疗AI的多智能体框架:利用微调的GPT、LLaMA和DeepSeek R1进行基于证据和偏见意识的临床查询处理

Naeimeh Nourmohammadi, Md Meem Hossain, The Anh Han, Safina Showkat Ara, Zia Ush Shamszaman

机构 * Department of Computing Games, Teesside University, Middlesbrough, United Kingdom Centre for Digital Innovation, Teesside University, Middlesbrough, United Kingdom Faculty of Business \& Technology, University of Sunderland, Sunderland, United Kingdom

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出一个多智能体框架,利用微调的GPT、LLaMA和DeepSeek R1,结合证据检索和偏见检查,提升医疗问答的可靠性与准确性。

Comments 27 pages, 14 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14002 2026-02-17 cs.CL cs.AI 62%

The Sufficiency-Conciseness Trade-off in LLM Self-Explanation from an Information Bottleneck Perspective

在信息瓶颈视角下,大型语言模型自我解释的充分性与简洁性权衡

Ali Zahedzadeh, Behnam Bahrak

机构 * Tehran Institute for Advanced Studies, Khatam University, Tehran, Iran(泰赫兰高级研究学院,卡坦大学,泰赫兰,伊朗)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文从信息瓶颈视角探讨了大型语言模型自我解释中充分性与简洁性之间的权衡,通过实验表明简洁解释在保持准确性的同时能显著减少长度,但过度压缩会损害性能。

Comments LREC 2026 submission; focuses on LLM self-explanation, interpretability, and information bottleneck analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13650 2026-02-17 cs.CV cs.AI cs.CL 62%

KorMedMCQA-V: A Multimodal Benchmark for Evaluating Vision-Language Models on the Korean Medical Licensing Examination

KorMedMCQA-V: 一种用于评估视觉语言模型在韩国医学资格考试中多模态多项选择问答能力的基准测试

Byungjin Choi, Seongsu Bae, Sunjun Kweon, Edward Choi

机构 * Ajou University School of Medicine(阿乔大学医学院) KAIST(韩国科学技术院)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 KorMedMCQA-V是一个用于评估视觉语言模型在韩国医学资格考试中多模态多项选择问答能力的基准测试,展示了不同模型在医疗领域中的表现差异。

Comments 17 pages, 2 figures, 6 tables. (Includes appendix.)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13272 2026-02-17 cs.AI cs.LG 62%

TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks

TemporalBench: 一个用于评估基于LLM的智能体在上下文和事件驱动的时间序列任务中的基准

Muyan Weng, Defu Cao, Wei Yang, Yashaswi Sharma, Yan Liu

机构 * University of Southern California(美国南加州大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 TemporalBench是一个多领域基准,用于评估基于LLM的智能体在上下文和事件驱动的时间序列任务中的时序推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12522 2026-02-17 cs.SE cs.AI cs.IR cs.LG cs.MA 62%

Improved Bug Localization with AI Agents Leveraging Hypothesis and Dynamic Cognition

改进的AI代理在利用假设和动态认知进行Bug定位中的应用

Asif Mohammed Samir, Mohammad Masudur Rahman

机构 * Dalhousie University(达尔豪西大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出CogniGent技术,利用AI代理进行因果推理和动态认知调试,提升bug定位的准确性和效率。

Comments 13 pages, 7 tables, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14857 2026-02-17 cs.AI 57%

World Models for Policy Refinement in StarCraft II

为《星际2》政策细化设计的世界模型

Yixin Zhang, Ziyi Wang, Yiming Rong, Haoxi Wang, Jinling Jiang, Shuang Xu, Haoran Wu, Shiyu Zhou, Bo Xu

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 StarWM为《星际2》政策细化设计的世界模型,通过预测未来观察和结构化文本表示提升策略决策性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01445 2026-02-17 cs.LG 57%

A Meta-Knowledge-Augmented LLM Framework for Hyperparameter Optimization in Time-Series Forecasting

一种增强元知识的LLM框架用于时间序列预测中的超参数优化

Ons Saadallah, Mátyás andó, Tamás Gábor Orosz

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

AI总结 LLM-AutoOpt通过结合贝叶斯优化与大型语言模型的上下文推理,提升时间序列预测中超参数优化的性能和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14166 2026-02-17 cs.CR cs.AI 57%

IntentMiner: Intent Inversion Attack via Tool Call Analysis in the Model Context Protocol

IntentMiner: 通过模型上下文协议中的工具调用分析进行意图倒置攻击

Yunhao Yao, Zhiqiang Wang, Haoran Cheng, Yihang Cheng, Haohua Du, Xiang-Yang Li

机构 * University of Science and Technology of China(中国科学技术大学) Beijing University of Aeronautics and Astronautics(北京航空航天大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 IntentMiner通过分析模型上下文协议中的工具调用,揭示了意图倒置攻击的机制,实现了超过85%的语义对齐,挑战了AI代理的隐私基础。

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14564 2026-02-17 cs.CL 57%

Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation

评估大型语言模型用于医疗问答:零样本和LLM作为评判者评估

Shefayat E Shams Adib, Ahmed Alfey Sani, Ekramul Alam Esham, Ajwad Abrar, Tareque Mohmud Chowdhury

机构 * Department of Computer Science and Engineering, Islamic University of Technology, Gazipur, Bangladesh(计算机科学与工程系,伊斯兰技术大学,加兹ipur,孟加拉国) Department of Computer Science(计算机科学系) Engineering, Islamic University of Technology, Gazipur, Bangladesh(工程,伊斯兰技术大学,加兹ipur,孟加拉国)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 本文评估了多种大型语言模型在医疗问答任务中的性能,发现大模型表现更优,同时指出在实际部署中需权衡效率与避让问题。

Comments Accepted in 28th ICCIT, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14456 2026-02-17 cs.LG 57%

Traceable Latent Variable Discovery Based on Multi-Agent Collaboration

基于多智能体协作的可追溯潜在变量发现

Huaming Du, Tao Hu, Yijie Huang, Yu Zhao, Guisong Liu, Tao Gu, Gang Kou, Carl Yang

机构 * Southwestern University of Finance and Economics(西南财经大学) Hunan University of Technology and Business(湖南工业大学) Xiangjiang Laboratory(湘江实验室) Emory University(埃默里大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

AI总结 TLVD通过结合LLMs的元数据推理与TCDA的数据驱动建模,实现潜在变量及其语义的可追溯推断。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14028 2026-02-17 cs.CL 57%

GRRM: Group Relative Reward Modeling for Machine Translation

GRRM:用于机器翻译的组相对奖励建模

Sen Yang, Shanbo Cheng, Lu Xu, Jianbing Zhang, Shujian Huang

机构 * National Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家实验室,南京大学) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 本文提出GRRM,通过组相对奖励模型提升机器翻译的排名准确性和推理能力。

Comments 19 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13832 2026-02-17 cs.CL 57%

Beyond Words: Evaluating and Bridging Epistemic Divergence in User-Agent Interaction via Theory of Mind

超越词语:通过心灵理论评估和弥合用户代理交互中的认知分歧

Minyuan Ruan, Ziyue Wang, Kaiming Liu, Yunghwei Lai, Peng Li, Yang Liu

机构 * Dept. of Comp. Sci. \& Tech., Institute for AI, Tsinghua University, Beijing, China Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 本文提出通过心灵理论评估和弥合用户代理交互中的认知分歧,通过强化学习提升模型对用户心理状态的推理能力,从而增强下游任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01144 2026-02-17 cs.CR cs.AI 57%

AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence

AthenaBench: 一个用于评估LLM在网络安全威胁情报中动态基准

Md Tanvirul Alam, Dipkamal Bhusal, Salman Ahmad, Nidhi Rastogi, Peter Worth

机构 * Rochester Institute of Technology(罗切斯特技术学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 AthenaBench通过改进的数据集和评估指标,评估了LLM在网络安全威胁情报中的表现,揭示了当前LLM在推理任务上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03490 2026-02-17 cs.CL 57%

Beyond Memorization: A Rigorous Evaluation Framework for Medical Knowledge Editing

超越记忆:面向医学知识编辑的严格评估框架

Shigeng Chen, Linhao Luo, Zhangchi Qiu, Yanan Cao, Carl Yang, Shirui Pan

机构 * School of Information and Communication Technology, Griffith University(格里菲斯大学信息与通信技术学院) Department of Data Science and AI, Monash University(莫纳什大学数据科学与人工智能系) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Department of Computer Science, Emory University(埃默里大学计算机科学系)

专题命中 推理评测 :reasoning(abstract);分类 cs.CL

AI总结 本文提出 MedEditBench 框架,评估医学知识编辑方法的有效性,发现现有方法仅能表面记忆信息,提出 SGR-Edit 方法提升泛化能力。

Comments Accepted to EACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13699 2026-02-17 cs.LG 57%

Attention Head Entropy of LLMs Predicts Answer Correctness

LLMs注意力头熵预测答案正确性

Sophie Ostmeier, Brian Axelrod, Maya Varma, Asad Aali, Yabin Zhang, Magdalini Paschali, Sanmi Koyejo, Curtis Langlotz, Akshay Chaudhari

机构 * Stanford University(斯坦福大学) University Hospital Zurich(苏黎世大学医院) Microsoft AI(微软人工智能)

专题命中 推理评测 :reasoning(abstract);分类 cs.LG

AI总结 LLMs注意力头熵预测答案正确性,通过注意力熵模式提升跨领域泛化能力,平均提升AUROC 8.5%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13243 2026-02-17 cs.CY cs.AI 57%

Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K--12 Science Instructional Materials

评判评判者:人类验证多LLM评估用于高质量K-12科学教学材料

Peng He, Zhaohui Li, Zeyuan Wang, Jinjun Xiong, Tingting Li

机构 * Washington State University, Pullman, WA, USA(华盛顿州立大学) University at Buffalo, State University of New York, Buffalo, NY, USA(布法罗大学)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 本研究通过人类专家验证多LLM评估,揭示LLM在K-12科学教学材料设计中的推理优劣,为生成式AI代理的发展提供指导。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13214 2026-02-17 cs.AI 57%

BotzoneBench: Scalable LLM Evaluation via Graded AI Anchors

BotzoneBench: 通过分级AI锚点实现可扩展的LLM评估

Lingfeng Li, Yunlong Lu, Yuefei Zhang, Jingyu Yao, Yixin Zhu, KeYuan Cheng, Yongyi Wang, Qirui Zheng, Xionghui Yang, Wenxin Li

机构 * School of Computer Science, Peking University(北京大学计算机科学学院) School of Software Engineering, South China University of Technology(华南理工大学软件学院) Department of Computer Science, Yale University(耶鲁大学计算机科学系) School of Psychological and Cognitive Sciences, Peking University(北京大学心理学与认知科学学院)

专题命中 推理评测 :reasoning(abstract);分类 cs.AI

AI总结 BotzoneBench通过分级AI锚点实现LLM战略能力的可扩展评估,评估八个多样化游戏中的177,047个状态-动作对,揭示性能差异并识别战略行为。

详情

展开后加载摘要…

URL PDF HTML 收藏