arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2025-12-30 至 2025-12-30 共收录 244 信号源:cs.CL, cs.AI, cs.LG

1. 长上下文与记忆 8 篇

2512.23343 2025-12-30 cs.CL cs.AI cs.CV 62%

AI Meets Brain: Memory Systems from Cognitive Neuroscience to Autonomous Agents

AI 与大脑:从认知神经科学到自主代理的记忆系统

Jiafeng Liang, Hao Li, Chang Li, Jiaqi Zhou, Shixin Jiang, Zekun Wang, Changkai Ji, Zhihao Zhu, Runxuan Liu, Tao Ren, Jinlan Fu, See-Kiong Ng, Xia Liang, Ming Liu, Bing Qin

机构 * Harbin Institute of Technology(哈尔滨工业大学) Fudan University(复旦大学) Peking University(北京大学) National University of Singapore(新加坡国立大学)

专题命中 长上下文与记忆 :LLM(abstract);分类 cs.CL、cs.AI

AI总结 本文从认知神经科学到自主代理,系统综合了记忆的跨学科知识,探讨了记忆的定义、功能、分类、存储机制及安全问题,并展望了多模态记忆系统和技能获取的未来研究方向。

Comments 57 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22435 2025-12-30 cs.AR cs.LG 57%

AnalogSAGE: Self-evolving Analog Design Multi-Agents with Stratified Memory and Grounded Experience

AnalogSAGE: 带分层记忆和基础经验的自进化模拟设计多智能体

Zining Wang, Jian Gao, Weimin Fu, Xiaolong Guo, Xuan Zhang

机构 * Northeastern University(东北大学) Kansas State University(堪萨斯州立大学)

专题命中 长上下文与记忆 :LLM(abstract);分类 cs.LG

AI总结 AnalogSAGE通过分层记忆和基础经验提升模拟设计自动化可靠性与自主性

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22624 2025-12-30 cs.CV 50%

Rethinking Memory Design in SAM-Based Visual Object Tracking

重新思考基于SAM的视觉目标跟踪中的记忆设计

Mohamad Alansari, Muzammal Naseer, Hasan Al Marzouqi, Naoufel Werghi, Sajid Javed

机构 * Department of Computer Science, Khalifa University(计算机科学系,卡利法大学)

专题命中 长上下文与记忆 :foundation model(abstract)

AI总结 本文提出了一种统一的混合记忆框架,通过分解记忆为短期外观记忆和长期干扰解决记忆,提升基于SAM的视觉目标跟踪在复杂场景下的鲁棒性。

Comments \textbf{This is a preprint. Some results are being finalized and may be updated in a future revision.}

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 推理与问题求解 41 篇

2512.07583 2025-12-30 cs.CL cs.AI 90%

Complementary Learning Approach for Text Classification using Large Language Models

基于大语言模型的文本分类互补学习方法

Navid Asgari, Benjamin M. Cole

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种基于大语言模型的文本分类互补学习方法,通过人机协作弥补各自弱点,以低成本技术处理评分差异问题。

Comments After further review, we identified substantive issues that materially affect the validity of the manuscript's core results and conclusions. Addressing these would require a fundamental reworking of the analysis and framing. To maintain the integrity of the public record, we request withdrawal of this version

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.03427 2025-12-30 cs.AI 90%

TPTU: Large Language Model-based AI Agents for Task Planning and Tool Usage

TPTU:基于大型语言模型的AI代理用于任务规划和工具使用

Jingqing Ruan, Yihong Chen, Bin Zhang, Zhiwei Xu, Tianpeng Bao, Guoqing Du, Shiwei Shi, Hangyu Mao, Ziyue Li, Xingyu Zeng, Rui Zhao

机构 * University of Chinese Academy of Sciences(中国科学院大学) SenseTime Research(商汤科技研究院) HKUST(香港科技大学)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出基于大型语言模型的AI代理框架,设计两种代理类型以提升任务规划和工具使用能力,并通过实验验证其在复杂任务中的有效性。

Comments Accepted in NeurIPS-2023 Workshop on Foundation Models for Decision Making

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18841 2025-12-30 cs.CL 89%

MDToC: Metacognitive Dynamic Tree of Concepts for Boosting Mathematical Problem-Solving of Large Language Models

MDToC:元认知动态概念树用于提升大语言模型的数学问题解决能力

Tung Duong Ta, Tim Oates, Thien Van Luong, Huan Vu, Tien Cuong Nguyen

机构 * University of Maryland, Baltimore County(马里兰大学巴尔的摩分校) National Economics University(国家经济大学) VNPT AI

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);prompting(abstract);分类 cs.CL

AI总结 MDToC通过构建概念树和多数投票机制,提升大语言模型在数学问题解决中的准确性与验证能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13082 2025-12-30 cs.CL 89%

Patience Is The Key to Large Language Model Reasoning

耐心是大型语言模型推理的关键

Yijiong Yu

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);preference optimization(abstract);分类 cs.CL

AI总结 本研究提出了一种无需额外知识或技能的简单方法,通过鼓励模型采用更耐心的推理风格,提升大型语言模型在复杂任务中的性能。

Comments The paper is not solid enough because the evaluation data is too less and the improvement is not significant

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01681 2025-12-30 physics.flu-dyn 89%

Large Language Model Driven Development of Turbulence Models

基于大语言模型的湍流模型开发

Zhongxin Yang, Yuanwei Bin, Yipeng Shi, Xiang I. A. Yang

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 本文提出利用大语言模型开发湍流模型,通过闭环迭代流程生成可解释且性能更优的近壁湍流模型,解决了不利压力梯度、系统旋转和表面粗糙度等问题。

Journal ref Flow (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23356 2025-12-30 cs.CL 88%

A Stepwise-Enhanced Reasoning Framework for Large Language Models Based on External Subgraph Generation

基于外部子图生成的大型语言模型逐步增强推理框架

Xin Zhang, Yang Cao, Baoxing Wu, Xinyi Chen, Kai Song, Siying Li

机构 * School of Information Science and Engineering, Chongqing Jiaotong University(重庆交通大学信息科学与工程学院) School of Computer Science and Technology, Chongqing University of Posts and Telecommunications(重庆邮电大学计算机科学与技术学院)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 SGR通过外部子图生成提升LLM的推理能力,通过多步骤推理和整合路径生成最终答案,实验表明其优于现有基线模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19466 2025-12-30 cs.CV cs.LG 88%

ForgerySleuth: Empowering Multimodal Large Language Models for Image Manipulation Detection

ForgerySleuth: 赋能多模态大语言模型进行图像篡改检测

Zhihao Sun, Haoran Jiang, Haoran Chen, Yixin Cao, Xipeng Qiu, Zuxuan Wu, Yu-Gang Jiang

机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理关键实验室,复旦大学计算机学院) Shanghai Collaborative Innovation Center of Intelligent Visual Computing(上海智能视觉计算协同创新中心)

专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.LG

AI总结 ForgerySleuth通过多模态大语言模型进行图像篡改检测,利用线索融合和数据集构建提升检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23511 2025-12-30 cs.SE cs.FL 87%

Beyond Correctness: Exposing LLM-generated Logical Flaws in Reasoning via Multi-step Automated Theorem Proving

超越正确性:通过多步骤自动定理证明暴露LLM生成的推理中的逻辑错误

Xinyi Zheng, Ningke Li, Xiaokun Luan, Kailong Wang, Ling Shi, Meng Sun, Haoyu Wang

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 MATP通过多步骤自动定理证明系统性验证LLM推理,揭示其逻辑错误并提升推理可信度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05822 2025-12-30 cs.CV 87%

Video Event Reasoning and Prediction by Fusing World Knowledge from LLMs with Vision Foundation Models

通过融合来自LLM的世界知识与视觉基础模型进行视频事件推理与预测

L'ea Dubois, Klaus Schmidt, Chengyu Wang, Ji-Hoon Park, Lin Wang, Santiago Munoz

机构 * INRIA(法国国家信息与自动化研究所) Max Planck Institute for Intelligent Systems(人工智能研究所) San Francisco State University(旧金山州立大学) Seoul AI Institute (SAII)(首尔人工智能研究所) Vision & Robotics Center, Tsinghua University(清华大学视觉与机器人中心) Polytechnic University of Madrid(马德里理工大学)

专题命中 推理与问题求解 :foundation model(title,abstract);LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文提出融合视觉基础模型与LLM的世界知识,以提升视频事件推理与预测能力,实现从简单识别到高级认知理解的突破。

Comments 22 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23167 2025-12-30 cs.AI cs.LG cs.MA 86%

SPIRAL: Symbolic LLM Planning via Grounded and Reflective Search

SPIRAL:通过 grounded 和 reflective 搜索实现符号 LLM 规划

Yifan Zhang, Giridhar Ganapavarapu, Srideepika Jayaraman, Bhavna Agrawal, Dhaval Patel, Achille Fokoue

机构 * IBM T.J. Watson Research Center(IBM TJ沃森研究中心)

专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 SPIRAL 通过 grounded 和 reflective 搜索实现更稳健高效的 LLM 规划

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22258 2025-12-30 cs.AI cs.LG cs.LO cs.SC 86%

Logic Sketch Prompting (LSP): A Deterministic and Interpretable Prompting Method

逻辑草图提示(LSP):一种确定性和可解释的提示方法

Satvik Tripathi

专题命中 推理与问题求解 :prompting(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 逻辑草图提示(LSP)通过引入类型变量、确定性条件评估器和规则验证器,提升了大语言模型在需要严格规则遵守、确定性和可审计性任务中的性能和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22275 2025-12-30 cs.CV cs.AI 85%

The Illusion of Clinical Reasoning: A Benchmark Reveals the Pervasive Gap in Vision-Language Models for Clinical Competency

临床推理的幻觉:一个基准揭示了视觉语言模型在临床能力方面的广泛差距

Dingyu Wang, Zimu Yuan, Jiajun Liu, Shanggui Liu, Nan Zhou, Tianxing Xu, Di Huang, Dong Jiang

专题命中 推理与问题求解 :language model(title,abstract);large language model(abstract);foundation model(abstract);分类 cs.AI

AI总结 本研究提出B&J基准测试,揭示了视觉语言模型在临床推理任务中存在显著的多模态整合能力不足问题,强调需在多模态整合和视觉理解上取得突破才能实现临床应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23515 2025-12-30 q-fin.TR cs.AI cs.CE cs.LG 84%

Alpha-R1: Alpha Screening with LLM Reasoning via Reinforcement Learning

Alpha-R1:通过强化学习实现基于LLM的Alpha筛选

Zuoyou Jiang, Li Zhao, Rui Sun, Ruohan Sun, Zhongjian Li, Jing Li, Daxin Jiang, Zuo Bai, Cheng Hua

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 推理与问题求解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 Alpha-R1通过强化学习实现基于LLM的上下文感知alpha筛选,提升非平稳市场中的投资策略性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17923 2025-12-30 q-fin.ST cs.AI cs.LG 84%

Inferring Latent Market Forces: Evaluating LLM Detection of Gamma Exposure Patterns via Obfuscation Testing

推断潜在市场力量:通过混淆测试评估大语言模型检测伽马暴露模式的能力

Christopher Regan, Ying Xie

机构 * Department of Computer Science Kennesaw State University Marietta, GA, Cobb(计算机科学系 凯尼恩州立大学 马里埃塔, 乔治亚州, 理查兹) Department of Information Technology Professor of Information Technology Kennesaw State University Marietta, GA(信息科技系 信息科技教授 凯尼恩州立大学 马里埃塔, 乔治亚州)

专题命中 推理与问题求解 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本研究通过混淆测试验证大语言模型能否通过因果推理识别结构性市场模式,发现其在无偏提示下检测伽马暴露模式的准确率为71.5%,展示了LLMs在金融机制识别方面的潜力。

Comments 10 pages, 8 figures. Accepted at IEEE Big Data 2025. Extended journal version in preparation. ISBN: 979-8-3315-9447-3/25. Page numbers: 7226-7235

Journal ref 2025 IEEE International Conference on Big Data (Big Data)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22931 2025-12-30 cs.AI cs.LG 81%

Geometric Structural Knowledge Graph Foundation Model

几何结构知识图谱基础模型

Ling Xin, Mojtaba Nayyeri, Zahra Makki Nayeri, Steffen Staab

机构 * University of Stuttgart(斯图加特大学) University of Southampton(南安普顿大学) Shahrood University of Technology(沙霍尔德大学)

专题命中 推理与问题求解 :foundation model(title,abstract);分类 cs.AI、cs.LG

AI总结 Gamma通过引入多头几何注意力机制,提升知识图谱推理的表达能力,优于现有方法Ultra,在零样本归纳链接预测中表现更优。

Comments Submitted to IEEE TPAMI, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18186 2025-12-30 cs.SD cs.CL eess.AS 79%

Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought

通过上下文感知和推理链引导语言模型实现稳定的语音情感识别

Zhixian Zhao, Xinfa Zhu, Xinsheng Wang, Shuiyuan Wang, Xuelong Geng, Wenjie Tian, Lei Xie

专题命中 推理与问题求解 :language model(title,abstract);分类 cs.CL

AI总结 C$^2$SER通过上下文感知和推理链提升语音情感识别的稳定性和准确性,优于现有模型。

Comments This work has been published in IEEE Transactions on Audio, Speech and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18262 2025-12-30 cs.RO cs.AI cs.CV cs.HC cs.LG 79%

ReSemAct: Advancing Fine-Grained Robotic Manipulation via Semantic Structuring and Affordance Refinement

ReSemAct:通过语义结构化和效用细化推进细粒度机器人操作

Chenyu Su, Weiwei Shang, Chen Qian, Fei Zhang, Shuang Cong

机构 * Department of Automation, University of Science and Technology of China(自动化系,中国科学技术大学)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);foundation model(abstract);分类 cs.AI、cs.LG

AI总结 ReSemAct 通过语义结构化和效用细化方法,在细粒度机器人操作中实现更精确的效用目标生成与动态环境适应。

Comments Code and videos: https://github.com/scy-v/ReSemAct and https://resemact.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23480 2025-12-30 cs.CR cs.AI 77%

Agentic AI for Autonomous Defense in Software Supply Chain Security: Beyond Provenance to Vulnerability Mitigation

面向软件供应链安全的代理AI:超越溯源到漏洞缓解

Toqeer Ali Syed, Mohammad Riyaz Belgaum, Salman Jan, Asadullah Abdullah Khan, Saad Said Alqahtani

机构 * Faculty of Computer and Information System(计算机与信息系统学院) Islamic University of Madinah(麦地那伊斯兰大学) Faculty of Computer Studies(计算机研究学院) Arab Open University-Bahrain(巴林阿拉伯开放大学)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出基于代理AI的软件供应链安全框架,结合LLM推理、强化学习和多代理协调,实现主动漏洞缓解,提升检测准确率和响应效率。

Comments Conference paper, accept in ACCA IEEE Bahrain

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23217 2025-12-30 cs.AI 77%

TCEval: Using Thermal Comfort to Assess Cognitive and Perceptual Abilities of AI

TCEval:利用热舒适性评估人工智能的认知与感知能力

Jingming Li

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 TCEval通过热舒适性场景评估AI的跨模态推理、因果关联和适应性决策能力,揭示当前LLM在热舒适性领域存在基础推理能力但缺乏精确因果理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25987 2025-12-30 cs.SE cs.AI 77%

R-Log: Incentivizing Log Analysis Capability in LLMs via Reasoning-based Reinforcement Learning

R-Log: 通过基于推理的强化学习激励大语言模型的日志分析能力

Yilun Liu, Ziang Chen, Song Xu, Minggui He, Shimin Tao, Weibin Meng, Yuming Xie, Tao Han, Chunguang Zhao, Jingzhou Du, Daimeng Wei, Shenglin Zhang, Yongqian Sun

机构 * Nankai University(南开大学) Huawei(华为)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);SFT(abstract);分类 cs.AI

AI总结 R-Log通过基于推理的强化学习提升大语言模型的日志分析能力,有效减少幻觉并提升泛化性能。

Comments Accepted by ICSE 2026 (SEIP Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01386 2025-12-30 cs.CL cs.CR cs.IR 77%

Topic-FlipRAG: Topic-Orientated Adversarial Opinion Manipulation Attacks to Retrieval-Augmented Generation Models

Topic-FlipRAG: 面向主题的对抗性观点操控攻击用于检索增强生成模型

Yuyang Gong, Zhuo Chen, Jiawei Liu, Miaokun Chen, Fengchang Yu, Wei Lu, Xiaofeng Wang, Xiaozhong Liu

机构 * Wuhan University(武汉大学) Nanyang Technological University(南洋理工大学) Worcester Polytechnic Institute(沃思堡理工学院)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出Topic-FlipRAG,一种针对检索增强生成模型的面向主题对抗性观点操控攻击方法,通过两阶段流程影响模型输出观点,揭示了RAG系统安全防护的迫切需求。

Comments Accepted by USENIX Security 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22627 2025-12-30 cs.CL 77%

Chain-of-thought Reviewing and Correction for Time Series Question Answering

时间序列问题回答的链式思考审查与修正

Chen Su, Yuanhe Tian, Yan Song

机构 * University of Science and Technology of China(中国科学技术大学) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 T3LLM通过引入显式修正机制,利用三个LLM协作实现时间序列问题回答的多步骤推理与自我修正,从而在多个基准测试中取得最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20368 2025-12-30 cs.AI 77%

AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning

AI-SearchPlanner: 通过帕累托最优多目标强化学习实现模块化代理搜索

Lang Mei, Zhihan Yang, Xiaohan Yu, Huanyao Zhang, Chong Chen

机构 * Huawei Cloud BU, China(华为云业务部,中国) School of Computer Science, Peking University(北京大学计算机学院)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 AI-SearchPlanner通过帕累托最优多目标强化学习提升冻结QA模型的搜索规划性能,实现模块化代理搜索。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23024 2025-12-30 cs.CV 75%

With Great Context Comes Great Prediction Power: Classifying Objects via Geo-Semantic Scene Graphs

大_context带来大预测能力:通过地-语义场景图进行物体分类

Ciprian Constantinescu, Marius Leordeanu

机构 * National University of Science and Technology(科学与技术国家大学) POLITEHNICA Bucharest(布加勒斯特POLITEHNICA)

专题命中 推理与问题求解 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文提出了一种基于地-语义场景图的上下文感知物体分类框架,通过整合深度估计与全景分割模型,显著提升分类准确率至73.4%,优于传统方法和多模态大语言模型。

Comments This paper is a development of the visual riddle game with Human-AI interaction, entitled "GuessWhat - Riddle Eye with AI", developed by Ciprian Constantinescu (POLItEHNICA Bucharest), Serena Stan (Instituto Cervantes Bucarest) and Marius Leordeanu (POLITEHNICA Bucharest), which was the winner (1st place) of the NeoArt Connect NAC 2025 Scholarship Program

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14404 2025-12-30 cs.CV 75%

ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations

ViC-Bench: 用自由式中间状态表示法对多模态大语言模型的视觉交错链式思维能力进行基准测试

Xuecheng Wu, Jiaxing Liu, Danlei Huang, Yifan Wang, Yunyun Shi, Kedi Chen, Junxiao Xue, Yang Liu, Chunlin Chen, Hairong Dong, Dingkang Yang

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) Meituan Inc.(美团公司) Institute of Advanced Technology, University of Science and Technology of China(中国科学技术大学先进技术研究院) School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) Research Center for Space Computing System, Zhejiang Lab(浙江实验室空间计算系统研究中心) College of Electronic and Information Engineering, Tongji University(同济大学电子与信息工程学院) School of Robotics and Automation, Nanjing University(南京大学机器人与自动化学院) College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);prompting(abstract)

AI总结 ViC-Bench 通过自由式中间状态表示法,系统评估多模态大语言模型的视觉交错链式思维能力,涵盖四个任务并提出新评估指标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22226 2025-12-30 cs.CV cs.AI cs.CL cs.LG 75%

VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs

VideoScaffold: 基于弹性尺度的视觉层次结构用于多模态大语言模型中的流媒体视频理解

Naishan Zheng, Jie Huang, Qingpei Guo, Feng Zhao

机构 * University of Science and Technology of China(中国科学技术大学) Ant Group(蚂蚁集团)

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 VideoScaffold通过弹性尺度事件分割和层次事件整合,实现了流媒体视频理解中的细粒度到抽象事件推理的动态转换,提升多模态大语言模型的视频处理性能。

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12620 2025-12-30 cs.CL cs.AI 73%

Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives

从形式和自然语言角度理解大语言模型中的三段论推理

Aheli Poddar, Saptarshi Sahoo, Sujata Ghosh

专题命中 推理与问题求解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文从形式和自然语言角度研究大语言模型的三段论推理能力,通过测试14种模型的符号推理和自然语言理解,探讨大语言模型是否正向形式化推理机制发展。

Comments 9 pages, 4 figures, 5 tables. Accepted at AAAI 2026 Bridge Program on Logic & AI. Code available at https://github.com/XAheli/Logic-in-LLMs

详情

展开后加载摘要…

URL PDF HTML 收藏