LCLA: Language-Conditioned Latent Alignment for Vision-Language Navigation
LCLA:语言引导的潜在对齐用于视觉-语言导航
专题命中 指令微调 :language model(abstract)
AI总结 LCLA通过将视觉-语言观测对齐到专家策略的潜在空间,实现轻量级的视觉-运动学习,提升在不同环境下的泛化能力。
AI 大模型
大语言模型、预训练、指令微调、后训练和语言模型应用。
LCLA:语言引导的潜在对齐用于视觉-语言导航
专题命中 指令微调 :language model(abstract)
AI总结 LCLA通过将视觉-语言观测对齐到专家策略的潜在空间,实现轻量级的视觉-运动学习,提升在不同环境下的泛化能力。
从人类反馈中鲁棒强化学习用于大语言模型微调
机构 * Department of Statistics, LSE(统计系,伦敦经济学院) ; Department of Mathematics, Tsinghua University(数学系,清华大学) ; School of Mathematics, University of Birmingham(数学学院,伯明翰大学) ; Department of Engineering Science, University of Oxford(工程科学系,牛津大学)
专题命中 后训练与偏好优化 :large language model(title,abstract);language model(title,abstract);LLM(abstract);RLHF(abstract)
AI总结 本文提出了一种鲁棒的强化学习算法,用于改进大语言模型微调中从人类反馈学习奖励函数的性能,通过减少方差和改进后悔界,实验证明其在基准数据集上表现优异。
超越成对:通过排名选择建模增强大语言模型对齐
机构 * Institute of Operations Research and Analytics(运营研究与分析研究所) ; National University of Singapore(新加坡国立大学) ; Department of Analytics and Operations(分析与运营系) ; NUS Business School(新加坡国立大学商学院)
专题命中 后训练与偏好优化 :LLM(title,abstract);large language model(abstract);language model(abstract);preference optimization(abstract)
AI总结 本文提出RCPO框架,通过最大似然估计结合排名选择建模,提升大语言模型对齐效果,实验显示其在多种模型和设置中均优于基线方法。
Comments Accepted by The Fourteenth International Conference on Learning Representations (ICLR 2026)
基于结构的分解直接偏好优化用于基于结构的药物设计
专题命中 后训练与偏好优化 :preference optimization(title,abstract);分类 cs.LG
AI总结 DecompDPO通过分解优化目标和引入物理信息能量项,提升基于结构的药物设计中分子生成与优化的性能。
Comments Accepted by TMLR
通过渐进自适应干预与再整合实现鲁棒编辑
机构 * Engineering Center for Space Utilization of Chinese Academy of Sciences(中国科学院空间利用工程技术中心) ; University of Chinese Academy of Sciences(中国科学院大学) ; School of Computer Science and Engineering(计算机科学与工程学院) ; Northeastern University(东北大学) ; New York University(纽约大学) ; School of Computing(计算机学院) ; National University of Singapore(新加坡国立大学) ; The Chinese University of Hong Kong(香港中文大学) ; Technical University of Munich(慕尼黑技术大学)
专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);post-training(abstract);分类 cs.CL、cs.AI
AI总结 REPAIR通过闭环反馈和动态内存管理实现鲁棒编辑,提升编辑准确性并减少知识遗忘。
ST4VLA:基于空间引导的视觉-语言-动作模型训练
机构 * Shanghai AI Laboratory(上海人工智能实验室) ; The Hong Kong University of Science and Technology(香港科技大学) ; Southern University of Science and Technology(南方科技大学) ; Fudan University(复旦大学)
专题命中 后训练与偏好优化 :language model(abstract);post-training(abstract);prompting(abstract)
AI总结 ST4VLA通过空间引导训练提升视觉-语言-动作模型的性能,实现对机器人任务的更稳健和可泛化的学习。
Comments Spatially Training for VLA, Accepted by ICLR 2026
VersaViT: 通过任务引导优化增强MLLM视觉骨干
机构 * School of Artificial Intelligence, Shanghai Jiao Tong University, China(上海交通大学人工智能学院) ; CMIC, Shanghai Jiao Tong University, China(上海交通大学计算机学院) ; WeChat AI, Tencent Inc., China(腾讯公司)
专题命中 后训练与偏好优化 :large language model(abstract);language model(abstract);post-training(abstract)
AI总结 VersaViT通过任务引导优化增强MLLM视觉骨干,解决其在密集预测任务中的性能问题,提升视觉任务的适应性与表现。
超越均匀信用:为策略优化的因果信用分配
机构 * Lexsi Labs(Lexsi实验室)
专题命中 后训练与偏好优化 :language model(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本文提出反事实重要性加权方法,通过因果信用分配提升语言模型推理性能,无需辅助模型或外部标注,实验验证其在GSM8K上的有效性。
Comments 12 pages, 1 figure
面向视觉生成的统一个性化奖励模型
机构 * Fudan University(复旦大学) ; Shanghai Innovation Institute(上海创新研究院) ; Shanghai Jiaotong University(上海交通大学) ; Shanghai AI Lab(上海人工智能实验室)
专题命中 后训练与偏好优化 :SFT(abstract);preference optimization(abstract)
AI总结 本文提出UnifiedReward-Flex,一种面向视觉生成的统一个性化奖励模型,通过结合奖励建模与灵活上下文适应的推理,提升视觉生成的准确性与对人类偏好的对齐性。
Comments Website: https://codegoat24.github.io/UnifiedReward/flex
UniARM: 向多目标测试时间对齐的统一自回归奖励模型迈进
机构 * School of Computer, Beihang University(北京航空航天大学计算机学院) ; Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究院)
专题命中 后训练与偏好优化 :LLM(abstract);分类 cs.CL
AI总结 UniARM通过统一自回归奖励模型实现多目标测试时间对齐,通过共享特征和偏好调节模块减少特征纠缠,提升对偏好权衡的控制能力。
Comments Under Review
从离线到在线:通过双层专家到策略融合提升GUI代理
机构 * Nanjing University(南京大学) ; Peking University(北京大学) ; Microsoft Research Asia(微软亚洲研究院)
专题命中 后训练与偏好优化 :language model(abstract);分类 cs.AI
AI总结 BEPA通过双层专家到策略融合方法,提升GUI代理在OSWorld-Verified等基准测试中的性能。
Comments Work In Progress
短上下文主导:自然语言实际上需要多少局部上下文?
机构 * University of British Columbia(不列颠哥伦比亚大学) ; Google DeepMind(谷歌DeepMind)
专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
AI总结 研究探讨了自然语言中短上下文预测的主导地位,并提出DaMCL方法用于检测长上下文序列,通过减少偏差提升模型性能。
Comments 38 pages, 7 figures, includes appendix and references
提示武器杀链:提示注入如何逐渐演变成多步骤恶意软件交付机制
机构 * Department of Software and Information Systems Engineering, Ben-Gurion University of the Negev(本·古里安大学软件与信息系统工程系) ; School of Electrical and Computer Engineering, Tel Aviv University(特拉维夫大学电气与计算机工程学院) ; Harvard Kennedy School, Harvard University, and Munk School, University of Toronto(哈佛大学哈佛肯尼迪学校及多伦多大学穆克学校)
专题命中 长上下文与记忆 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI
AI总结 本文提出提示武器杀链模型,揭示提示注入演变为多步骤恶意软件攻击机制的过程,并提出针对各阶段的防御策略。
注意力沉降与压缩山谷是大语言模型中的双面现象
机构 * University of Oxford(牛津大学) ; AITHYRA ; New York University(纽约大学)
专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
AI总结 本研究揭示了大语言模型中注意力沉降与压缩山谷的联系,提出信息流的Mix-Compress-Refine理论,解释LLM如何通过大规模激活控制注意力和压缩来组织深度计算。
TraceMem: 从用户对话轨迹编织叙事记忆图式
机构 * The University of Hong Kong(香港大学) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; Nankai University(南开大学)
专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL
AI总结 TraceMem通过三阶段流程从用户对话轨迹中编织叙事记忆图式,提升多跳和时间推理能力。
结构化事件记忆
机构 * Southeast University(东南大学) ; Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) ; Shenzhen Loop Area Institute(深圳河套学院) ; Alibaba Group(阿里巴巴集团)
专题命中 长上下文与记忆 :large language model(abstract);language model(abstract);分类 cs.CL
AI总结 SEEM通过结构化事件记忆框架提升自主代理的叙述连贯性和逻辑一致性。
OpenMonoGS-SLAM: 单目高斯点云SLAM与开放语义结合
机构 * Sungkyunkwan University(成均馆大学) ; Yonsei University(延世大学)
专题命中 长上下文与记忆 :foundation model(abstract)
AI总结 OpenMonoGS-SLAM通过结合3DGS与开放集语义,实现无需深度输入的单目SLAM,提升开放世界环境下的感知与建图性能。
Comments Work in progress. Project page: https://jisang1528.github.io/OpenMonoGS-SLAM/
无需训练的多模态仇恨定位与大型语言模型
机构 * Hybrid Intelligence Lab, University of Durham(杜伦大学混合智能实验室) ; Multimodal Intelligence Lab, University of Exeter(埃克塞特大学多模态智能实验室) ; The MIx Group University of Birmingham(伯明翰大学MIx集团)
专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);prompting(abstract)
AI总结 提出无需训练的多模态仇恨视频定位框架LELA,通过多阶段提示方案和跨模态推理机制实现高精度定位。
在大语言模型的搜索增强推理中知识整合衰减
机构 * Department of Electrical and Computer Engineering, Seoul National University, Seoul, Korea(电气与计算机工程系,首尔国立大学,韩国首尔) ; Interdisciplinary Program in Artificial Intelligence, Seoul National University, Seoul, Korea(人工智能交叉学科项目,首尔国立大学,韩国首尔)
专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL
AI总结 本研究提出SAKE方法,通过在推理过程前后锚定检索知识,缓解大语言模型在搜索增强推理中的知识整合衰减问题,提升多跳问答和复杂推理任务的性能。
通过逐步经验回忆实现大语言模型的自引导函数调用
机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) ; Nanjing University of Information Science & Technology(南京信息工程大学) ; Institute of Computing Technology, Chinese Academy of Sciences(计算技术研究所,中国科学院)
专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL
AI总结 SEER通过逐步经验回忆方法,提升大语言模型在多步骤工具使用中的准确性和效率。
Comments Accepted to EMNLP 2025
VideoAfford: 通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding
专题命中 推理与问题求解 :large language model(title,abstract);language model(title,abstract)
AI总结 VideoAfford通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding,结合动态交互先验和空间感知损失函数,提升机器人操作的可操作区域识别能力。
对多智能体大语言模型推理树进行审计优于多数投票和LLM作为裁判
机构 * University of Southern California, Los Angeles, CA, USA(美国南加州大学)
专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);preference optimization(abstract)
AI总结 通过构建推理树路径搜索机制,AgentAuditor在多智能体系统中实现了优于多数投票和LLM作为裁判的推理准确性提升
大语言模型推理可预测模型何时正确:来自编程课堂对话的证据
专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL
AI总结 本研究通过分析课堂对话中的教师发言,发现LLM生成的推理能有效预测模型预测的正确性,使用随机森林分类器达到F1得分0.83,表明基于推理的错误检测在教育对话分析中具有实用价值。
AnalyticsGPT: 一个基于大语言模型的科学度量问答工作流
机构 * Elsevier B.V.(埃塞尔弗斯有限公司)
专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL
AI总结 AnalyticsGPT通过LLM实现科学度量问答,结合检索增强生成和代理概念,提升科学数据分析效率。
基本推理范式诱导语言模型的领域外泛化
专题命中 推理与问题求解 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.CL
AI总结 本研究通过诱导基本推理范式提升语言模型的领域外泛化能力,实验表明方法在现实任务中性能提升显著。
ReAcTree:具有控制流的分层LLM代理树用于长时间任务规划
机构 * ETRI & UST(ETRI与UST) ; UST
专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
AI总结 ReAcTree通过分层代理树和控制流实现复杂任务规划,显著提升长周期任务的完成效率和成功率。
Comments Accepted as a Full Paper at AAMAS 2026. This is the extended version including full appendices. Code is available at https://github.com/Choi-JaeWoo/ReAcTree.git
PersonaX: 多模态数据集与LLM推断行为特征
机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) ; Carnegie Mellon University(卡内基梅隆大学) ; University of California San Diego(加州大学圣地亚哥分校) ; Australian National University(澳大利亚国立大学)
专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG
AI总结 PersonaX通过多模态数据集结合LLM推断行为特征,推动多模态特征分析与因果推理发展。
Comments ICLR 2026
HealthProcessAI: 一种基于大语言模型的医疗流程挖掘技术框架及证明概念
专题命中 推理与问题求解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
AI总结 HealthProcessAI通过整合大语言模型实现医疗流程挖掘的自动化解释和报告生成,提升其在医疗和流行病学中的可访问性和实用性。
Comments Figure 1 updated, typos corrected, references added, under review
Cochain: 在LLM代理工作流中平衡不足与过度的合作
机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) ; Department of Computer Science and Engineering, Southern University of Science and Technology(南方科技大学计算机科学与工程系) ; School of Urban Planning and Design, Peking University(北京大学城市规划与设计学院)
专题命中 推理与问题求解 :LLM(title);large language model(abstract);language model(abstract);prompting(abstract)
AI总结 Cochain通过整合知识图谱和提示树,有效解决业务工作流中LLM代理合作问题,优于基线模型。
Comments 35 pages, 23 figures
合同型深度伪造:大语言模型能生成合同吗?
专题命中 推理与问题求解 :large language model(title);language model(title);分类 cs.CL、cs.AI
AI总结 本文指出大语言模型生成的合同可能不具法律效力,质疑其在法律领域应用的可行性。
Comments Accepted for publication