arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2025-12-08 至 2025-12-08 共收录 54 信号源:cs.CL, cs.AI, cs.LG

1. 数学推理 2 篇

2512.05580 2025-12-08 cs.CL 85%

Structured Reasoning with Tree-of-Thoughts for Bengali Math Word Problems

基于树状思维的结构化推理用于孟加拉语数学应用题

Aurprita Mahmood, Sabrin alam, Neloy kumer Sagor, Md. Abdul Hadi, Md. Sehab Al Islam, Minhajul Islam

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Ahsanullah University of Science and Technology(阿沙努尔大学科学与技术学院) Southeast University(东南大学)

专题命中 数学推理 :reasoning(title,abstract);chain-of-thought(abstract);CoT(abstract);分类 cs.CL

AI总结 本文提出基于树状思维的结构化推理方法,用于提升孟加拉语数学应用题的解决效果,通过实验验证ToT在中等至大规模模型中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05502 2025-12-08 cs.LG 83%

GRASP: Graph Reasoning Agents for Systems Pharmacology with Human-in-the-Loop

GRASP:具有人机交互循环的图推理代理用于系统药理学

Omid Bazgir, Vineeth Manthapuri, Ilia Rattsev, Mohammad Jafarnejad

机构 * Clinical Pharmacology, Genentech(基因泰克临床药理部) Preclinical & Translational PKPD, Genentech Inc.(基因泰克预临床与转化药代动力学)

专题命中 数学推理 :reasoning(title,abstract);CoT(abstract);分类 cs.LG

AI总结 GRASP通过图推理代理实现系统药理学模型开发的自动化与严谨性,提升生物医学建模的效率和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 代码与定理证明 3 篇

2512.05318 2025-12-08 cs.CL cs.AI cs.LG 87%

To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples

思考还是不思考:过度使用Co-T示例的元训练隐藏成本

Vignesh Kothapalli, Ata Fatahibaarzi, Hamed Firooz, Maziar Sanjabi

机构 * Stanford University(斯坦福大学) LinkedIn AI

专题命中 代码与定理证明 :CoT(title,abstract);reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出CoT-Recipe方法,通过调节元训练序列中CoT和非CoT示例的比例,提升大型语言模型在新任务上的推理准确性,实验显示在无CoT示例时准确率可提升300%。

Comments 26 pages, 45 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05760 2025-12-08 cs.AI 79%

Evolutionary System 2 Reasoning: An Empirical Proof

进化系统2推理:一个实证证明

Zeyuan Ma, Wenqi Huang, Guo-Huan Song, Hongshu Guo, Sijie Ma, Zhiguang Cao, Yue-Jiao Gong

专题命中 代码与定理证明 :reasoning(title,abstract);分类 cs.AI

AI总结 本文提出进化推理优化框架,通过进化策略提升LLM的推理能力,实验证实即使弱模型也能通过简单进化获得强大推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05836 2025-12-08 cs.AI 57%

Using Large Language Models to Create Personalized Networks From Therapy Sessions

利用大语言模型从治疗会谈中创建个性化网络

Clarissa W. Ong, Hiba Arnaout, Kate Sheehan, Estella Fox, Eugen Owtscharow, Iryna Gurevych

机构 * Department of Psychological and Brain Sciences, University of Louisville, U.S.A(心理与脑科学系,路易斯维尔大学,美国) Department of Computer Science, TU Darmstadt, Germany(计算机科学系,图恩-达姆斯塔特大学,德国) Department of Psychology, University of Toledo, U.S.A(心理学系,托莱多大学,美国)

专题命中 代码与定理证明 :planning(abstract);分类 cs.AI

AI总结 本研究利用大语言模型从治疗记录中自动生成个性化网络,通过上下文学习和聚类方法提升临床效用和可解释性,为治疗个性化提供新方法。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 逻辑推理 2 篇

2504.07640 2025-12-08 cs.AI 88%

Enhancing Large Language Models through Neuro-Symbolic Integration and Ontological Reasoning

通过神经符号整合和本体推理增强大型语言模型

Ruslan Idelfonso Magana Vsevolodovna, Marco Monti

专题命中 逻辑推理 :reasoning(title,abstract);logical reasoning(title,abstract);分类 cs.AI

AI总结 本文提出一种结合符号本体推理和机器学习的方法,通过迭代反馈提升LLM输出的一致性和事实准确性。

Comments Withdrawn because Version 1 contains inaccuracies in references and architecture description. A corrected and improved version will be submitted separately

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05365 2025-12-08 cs.AI q-bio.QM 79%

MCP-AI: Protocol-Driven Intelligence Framework for Autonomous Reasoning in Healthcare

MCP-AI:面向医疗领域自主推理的协议驱动智能框架

Zag ElSayed, Craig Erickson, Ernest Pedapati

机构 * School of Information Technology University of Cincinnati Ohio, USA(信息科技学院 俄亥俄州立大学 奥哈伊俄州) Adolescent Psychiatry Cincinnati Children’s Hospital Medical Center Ohio, USA(青少年精神病学 奥克兰儿童医院医疗中心 奥哈伊俄州)

专题命中 逻辑推理 :reasoning(title,abstract);分类 cs.AI

AI总结 MCP-AI是一种基于模型上下文协议的医疗自主推理框架,通过整合临床逻辑和安全协作,提升医疗决策的可解释性和适应性。

Comments 6 pages, 4 figures

Journal ref IEEE ICMLA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 规划推理 10 篇

2512.05256 2025-12-08 cs.CL q-bio.QM 83%

Enhancing Clinical Note Generation with ICD-10, Clinical Ontology Knowledge Graphs, and Chain-of-Thought Prompting Using GPT-4

利用ICD-10、临床本体知识图谱和链式推理提示法提升临床笔记生成

Ivan Makohon, Mohamad Najafi, Jian Wu, Mathias Brochhausen, Yaohang Li

机构 * Computer Science, Old Dominion University(旧 Dominion 大学计算机科学系) Biomedical Informatics, University of Arkansas for Medical Sciences(阿肯色医学科学大学生物医学信息学系)

专题命中 规划推理 :chain-of-thought(title,abstract);CoT(abstract);分类 cs.CL

AI总结 本研究利用ICD-10、临床本体知识图谱和链式推理提示法提升临床笔记生成质量,通过GPT-4在六个临床案例中验证了方法的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01592 2025-12-08 cs.CL cs.AI 81%

AURA: A Diagnostic Framework for Tracking User Satisfaction of Interactive Planning Agents

AURA:一个用于跟踪交互式规划代理用户满意度的诊断框架

Takyoung Kim, Janvijay Singh, Shuhaib Mehri, Emre Can Acikgoz, Sagnik Mukherjee, Nimet Beyza Bozdag, Sumuk Shashidhar, Gokhan Tur, Dilek Hakkani-Tür

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 规划推理 :planning(title,abstract);分类 cs.CL、cs.AI

AI总结 AURA提出一个诊断框架,用于跟踪交互式规划代理的用户满意度,通过评估代理行为阶段和中间行为来提升用户体验。

Comments NeurIPS 2025 MTI-LLM Workshop. Full version is under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15436 2025-12-08 cs.CV 78%

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs

基于动态视觉搜索和缩放的自适应聚焦推理方法用于高效VLMs

Xintong Zhang, Zhi Gao, Bofei Zhang, Pengxiang Li, Xiaowen Zhang, Yang Liu, Tao Yuan, Yuwei Wu, Yunde Jia, Song-Chun Zhu, Qing Li

机构 * organization= School of Computer Science \& Technology, Beijing Institute of Technology , city= Beijing , country= China organization= State Key Laboratory of General Artificial Intelligence, BIGAI , city= Beijing , country= China organization= School of Intelligence Science Technology, Peking University , city= Beijing , country= China organization= Guangdong Laboratory of Machine Perception Intelligent Computing, Shenzhen MSU--BIT University , city= Shenzhen , country= China organization= Department of Automation, Tsinghua University , city= Beijing , country= China

专题命中 规划推理 :reasoning(title,abstract)

AI总结 本文提出基于动态视觉搜索和缩放的自适应聚焦推理方法,提升VLMs的多模态推理效率和实际应用效果。

Comments https://github.com/xtong-zhang/Chain-of-Focus

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08621 2025-12-08 cs.CL 74%

From Simulation to Strategy: Automating Personalized Interaction Planning for Conversational Agents

从模拟到策略:为对话代理自动化个性化交互规划

Wen-Yu Chang, Tzu-Hung Huang, Chih-Ho Chen, Yun-Nung Chen

机构 * Department of Computer Science and Information Engineering(计算机科学与信息工程系) National Taiwan University(台湾大学)

专题命中 规划推理 :planning(title);分类 cs.CL

AI总结 本文提出了一种基于职业信息的轻量级策略,用于优化销售导向对话代理的个性化交互规划,通过调整对话意图提升对话效果。

Comments Accepted to IEEE ASRU 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08475 2025-12-08 cs.SE cs.AI 70%

Designing LLM-based Multi-Agent Systems for Software Engineering Tasks: Quality Attributes, Design Patterns and Rationale

设计基于大语言模型的多智能体系统用于软件工程任务:质量属性、设计模式与理由

Yangxiao Cai, Ruiyin Li, Peng Liang, Mojtaba Shahin, Zengyang Li

机构 * School of Computer Science, Wuhan University(武汉大学计算机学院) School of Computing Technologies, RMIT University(皇家墨尔本理工学院计算技术学院) School of Computer Science, Central China Normal University(中央师范大学计算机学院)

专题命中 规划推理 :reasoning(abstract);planning(abstract);分类 cs.AI

AI总结 本文研究了基于大语言模型的多智能体系统在软件工程任务中的设计,发现代码生成是最常见任务,功能性适应性是主要关注的质量属性,基于角色的合作是最常用的设计模式,提高生成代码质量是主要设计动机。

Comments 35 pages, 4 images, 7 tables, Manuscript submitted to a Journal (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05824 2025-12-08 cs.AI cs.CV 57%

Multimodal Oncology Agent for IDH1 Mutation Prediction in Low-Grade Glioma

多模态肿瘤代理用于低级别胶质瘤IDH1突变预测

Hafsa Akebli, Adam Shephard, Vincenzo Della Mea, Nasir Rajpoot

机构 * Department of Mathematics, Computer Science and Physics, University of Udine(乌迪大学数学、计算机科学与物理系) Tissue Image Analytics Centre, Department of Computer Science, University of Warwick(沃里克大学计算机科学系组织图像分析中心) Histofy Ltd(Histofy有限公司)

专题命中 规划推理 :reasoning(abstract);分类 cs.AI

AI总结 本研究提出多模态肿瘤代理,结合组织学工具和外部生物医学资源,实现低级别胶质瘤IDH1突变的高精度预测。

Comments 4 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05403 2025-12-08 cs.LG 57%

RevoNAD: Reflective Evolutionary Exploration for Neural Architecture Design

RevoNAD:基于反射的进化探索用于神经架构设计

Gyusam Chang, Jeongyoon Yoon, Shin han yi, JaeHyeok Lee, Sujin Jang, Sangpil Kim

机构 * Korea University(韩国大学) Samsung AI Center(三星人工智能中心)

专题命中 规划推理 :reasoning(abstract);分类 cs.LG

AI总结 RevoNAD通过多轮多专家共识、自适应反射探索和帕累托引导的进化选择,实现了基于LLM的神经架构设计的高效优化与可靠生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05336 2025-12-08 cs.LG 57%

PathFinder: MCTS and LLM Feedback-based Path Selection for Multi-Hop Question Answering

PathFinder: 基于MCTS和LLM反馈的多跳问答路径选择

Durga Prasad Maram, Kalpa Gunaratna, Vijay Srinivasan, Haris Jeelani, Srinivas Chappidi

专题命中 规划推理 :reasoning(abstract);分类 cs.LG

AI总结 PATHFINDER通过MCTS和LLM反馈优化多跳问答的路径选择,提升训练数据质量与推理准确性。

Comments 5 PAGES, 3 IMAGES

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05686 2025-12-08 eess.SY cs.SY 50%

LA-RL: Language Action-guided Reinforcement Learning with Safety Guarantees for Autonomous Highway Driving

LA-RL: 基于语言动作引导的安全强化学习用于自动驾驶高速公路驾驶

Yiming Shu, Jiahui Xu, Jiwei Tang, Ruiyang Gao, Chen Sun

专题命中 规划推理 :reasoning(abstract)

AI总结 LA-RL通过整合大语言模型的语义推理和改进的安全层,提升自动驾驶高速公路驾驶的效率与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05470 2025-12-08 cs.SE 50%

Everything is Context: Agentic File System Abstraction for Context Engineering

一切皆为上下文:面向上下文工程的代理文件系统抽象

Xiwei Xu, Robert Mao, Quan Bai, Xuewu Gu, Yechao Li, Liming Zhu

专题命中 规划推理 :reasoning(abstract)

AI总结 本文提出一种基于文件系统的上下文工程抽象,通过统一挂载、元数据和访问控制,实现可验证的上下文管理,支持可问责和以人类为中心的AI协作。

Comments Submitted

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 视觉空间推理 9 篇

2512.05809 2025-12-08 cs.CV cs.AI 83%

Probing the effectiveness of World Models for Spatial Reasoning through Test-time Scaling

通过测试时缩放探查世界模型在空间推理中的有效性

Saurav Jha, M. Jehanzeb Mirza, Wei Lin, Shiqi Yang, Sarath Chandar

机构 * MILA – Quebec AI Institute(魁北克人工智能研究所) Polytechnique Montréal(蒙特利尔理工学院) MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Institute for Machine Learning, Johannes Kepler University Linz(林茨约瑟夫·夫兰克大学机器学习研究所) Nankai University(南开大学)

专题命中 视觉空间推理 :reasoning(title,abstract);verifier(abstract);分类 cs.AI

AI总结 本文提出ViSA框架,通过可验证的微断言改进世界模型的空间推理能力,但发现当前模型在复杂任务中仍存在信息瓶颈。

Comments Extended abstract at World Modeling Workshop 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04563 2025-12-08 cs.CV 78%

COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence

COOPER:一种用于空间智能中协作感知与推理的统一模型

Zefeng Zhang, Xiangzhao Hao, Hengzhu Tang, Zhenyu Zhang, Jiawei Sheng, Xiaodong Li, Zhenyang Li, Li Gao, Daiting Shi, Dawei Yin, Tingwen Liu

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Baidu Inc.(百度公司)

专题命中 视觉空间推理 :reasoning(title,abstract)

AI总结 COOPER是一种统一的多模态大语言模型,通过整合深度和分割等辅助模态,提升空间感知与推理能力,实现空间智能的增强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18200 2025-12-08 cs.CV 78%

InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity

InfiniBench:用于可定制场景复杂度的无限基准测试

Haoming Wang, Qiyao Xue, Wei Gao

机构 * University of Pittsburgh(匹兹堡大学)

专题命中 视觉空间推理 :reasoning(title,abstract)

AI总结 InfiniBench通过可定制的3D场景生成方法,提升视觉语言模型在复杂空间推理任务中的评估能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19552 2025-12-08 cs.CV 78%

iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning

iFinder: 结构化零样本视觉基于LLM的地面定位用于行车记录仪视频推理

Manyi Yao, Bingbing Zhuang, Sparsh Garg, Amit Roy-Chowdhury, Christian Shelton, Manmohan Chandraker, Abhishek Aich

机构 * NEC Laboratories, America(NEC美国实验室) University of California, Riverside(加州大学河滨分校) University of California, San Diego(加州大学圣地亚哥分校)

专题命中 视觉空间推理 :reasoning(title,abstract)

AI总结 iFinder通过结构化语义接地框架,利用行车记录仪视频中的关键线索提升LLM在驾驶视频推理中的性能。

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18973 2025-12-08 cs.CL cs.LG 62%

Hierarchical Mamba Meets Hyperbolic Geometry: A New Paradigm for Structured Language Embeddings

层次Mamba与双曲几何:一种新的结构语言嵌入范式

Sarang Patil, Ashish Parmanand Pandey, Ioannis Koutis, Mengjia Xu

机构 * New Jersey Institute of Technology(新泽西理工学院)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL、cs.LG

AI总结 本文提出层次Mamba结合双曲几何,以学习具有层次意识的语言嵌入,提升复杂层次推理能力。

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07755 2025-12-08 cs.CV cs.AI cs.GR cs.RO 57%

SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models

SAT:多模态语言模型的动态空间能力训练

Arijit Ray, Jiafei Duan, Ellis Brown, Reuben Tan, Dina Bashkirova, Rose Hendrix, Kiana Ehsani, Aniruddha Kembhavi, Bryan A. Plummer, Ranjay Krishna, Kuo-Hao Zeng, Kate Saenko

机构 * Boston University(波士顿大学) University of Washington(华盛顿大学) Allen Institute for AI(人工智能研究院) Microsoft Research(微软研究院) New York University(纽约大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

AI总结 SAT通过模拟数据提升多模态语言模型在动态空间推理中的能力,实验表明其在多个基准测试中优于现有方法。

Comments Accepted to COLM 2025. Project webpage: https://arijitray.com/SAT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05815 2025-12-08 cs.RO 50%

Optimal Safety-Aware Scheduling for Multi-Agent Aerial 3D Printing with Utility Maximization under Dependency Constraints

多智能体空中3D打印的最优安全意识调度与效用最大化在依赖约束下

Marios-Nektarios Stamatopoulos, Shridhar Velhal, Avijit Banerjee, George Nikolakopoulos

机构 * Robotics and AI Group, Luleå University of Technology(机器人与人工智能组,卢勒奥技术大学)

专题命中 视觉空间推理 :planning(abstract)

AI总结 本文提出了一种基于效用最大化的优化框架,用于在依赖约束下实现多智能体空中3D打印的最优安全调度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05398 2025-12-08 cs.CV 50%

The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos

动态先验:为随意动态视频理解3D结构

Zhuoyuan Wu, Xurui Yang, Jiahui Huang, Yue Wang, Jun Gao

机构 * PKU(北京大学) NVIDIA(英伟达) USC(南加州大学) University of Michigan(密歇根大学)

专题命中 视觉空间推理 :reasoning(abstract)

AI总结 本文提出动态先验模型,利用Vision-Language Models和SAM2实现无需任务特定训练的动态物体识别,提升结构3D理解的准确性和鲁棒性。

Comments Code is available at https://github.com/wuzy2115/DYNAPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01046 2025-12-08 cs.RO 50%

STATE-NAV: Stability-Aware Traversability Estimation for Bipedal Navigation on Rough Terrain

STATE-NAV:面向粗糙地形双足导航的稳定性感知可通行性估计

Ziwon Yoon, Lawrence Y. Zhu, Jingxi Lu, Lu Gan, Ye Zhao

机构 * Institute for Robotics and Intelligent Machines, Georgia Institute of Technology(机器人与智能机械研究所,佐治亚理工学院)

专题命中 视觉空间推理 :planning(abstract)

AI总结 本文提出STATE-NAV框架,通过基于Transformer的神经网络预测双足机器人在粗糙地形中的稳定性,结合分层规划器实现风险感知的导航,提升导航性能和鲁棒性。

Comments Accepted to IEEE Robotics and Automation Letters (RA-L)

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 测试时计算 4 篇

2512.05325 2025-12-08 cs.CL cs.AI cs.LG 85%

LYNX: Learning Dynamic Exits for Confidence-Controlled Reasoning

LYNX: 为置信度控制推理学习动态退出

Ömer Faruk Akgül, Yusuf Hakan Kalaycı, Rajgopal Kannan, Willie Neiswanger, Viktor Prasanna

机构 * University of Southern California(南加州大学) DEVCOM ARL(美国陆军战争研究所)

专题命中 测试时计算 :reasoning(title,abstract);verifier(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 LYNX通过利用模型自身隐藏状态实现置信度控制的提前退出,提升推理效率和准确性,适用于多种任务和基准测试。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04996 2025-12-08 cs.LG cs.AI cs.CL stat.ML 67%

Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives

Reinforce-Ada: 非线性强化学习目标下的自适应采样框架

Wei Xiong, Chenlu Ye, Baohao Liao, Hanze Dong, Xinxing Xu, Christof Monz, Jiang Bian, Nan Jiang, Tong Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Microsoft Research(微软研究院) University of Amsterdam(阿姆斯特丹大学)

专题命中 测试时计算 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 Reinforce-Ada通过自适应采样框架提升大语言模型在非线性强化学习目标下的推理性能,有效恢复丢失信号并加速收敛。

Comments 27 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05542 2025-12-08 cs.LG cs.AI 66%

RoBoN: Routed Online Best-of-n for Test-Time Scaling with Multiple LLMs

RoBoN: 基于路由的在线最佳-n方法用于多LLM测试时间扩展

Jonathan Geuter, Gregor Kornhardt

机构 * Kempner Institute for the Study of Natural & Artificial Intelligence(自然与人工智能研究 institute)

专题命中 测试时计算 :reasoning(abstract,comments);分类 cs.AI、cs.LG

AI总结 RoBoN通过在线路由多LLM生成过程,提高测试时间扩展性能,无需额外训练,有效提升准确率。

Comments 20 pages, 3 figures. 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Foundations of Reasoning in Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05546 2025-12-08 cs.CV cs.AI 57%

Conscious Gaze: Adaptive Attention Mechanisms for Hallucination Mitigation in Vision-Language Models

有意识的注视:用于视觉-语言模型中幻觉抑制的自适应注意力机制

Weijue Bu, Guan Yuan, Guixian Zhang

机构 * School of Computer Science and Technology/School of Artificial Intelligence(计算机科学与技术学院/人工智能学院) China University of Mining and Technology(中国矿业大学)

专题命中 测试时计算 :reasoning(abstract);分类 cs.AI

AI总结 CG-VLM通过认知需求传感器和聚焦共识诱导模块,在推理时精准干预视觉-语言模型的注意力,有效抑制幻觉并提升性能。

Comments 6 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏