arXivDaily每日学术速递，同步arXiv全量数据，AI总结、翻译，覆盖人工智能、机器人、计算机、金融、统计学、数学、物理学、生物学、经济学、电气&系统等方向。

2605.01391 2026-06-12 cs.CV 版本更新

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

VISTA：视频交互时空分析基准

Alejandro Aparcedo, Akash Kumar, Aaryan Garg, Dalton Pham, Wen-Kai Chen, Anirudh Bharadwaj, Aman Chadha, Yogesh Rawat

发表机构 * University of Central Florida（中央佛罗里达大学）； BITS Pilani（比特斯理工学院）； Ho Chi Minh City University of Science（胡志明市科学大学）； Amazon GenAI Project（亚马逊生成人工智能项目）

AI总结提出VISTA基准，通过分解视频为实体、动作和关系，实现开放集多实体多动作的时空理解评估，揭示传统指标掩盖的偏差。

Comments Accepted to CVPR 2026 Workshop on Pixel-level Video Understanding in the Wild (PVUW)

详情

AI中文摘要

现有的视觉-语言模型（VLM）基准主要评估简单单动作视频、封闭属性集和受限实体类型的时空理解，未能捕捉真实世界视频理解中多样实体之间的自由形式多动作交互。此外，缺乏一个系统性的框架来分析模型在互补时空轴上的失败，阻碍了全面评估。为解决这些问题，我们引入了VISTA，一个视频交互时空分析基准，专为VLM中的开放集、多实体和多动作时空理解设计。VISTA将视频分解为可解释的实体、其关联动作和关系动态，实现多轴诊断以及关系、空间和时间理解的统一评估。我们的基准将多个数据集整合到一个单一的交互感知分类法中，包含约12K个精心策划的视频-查询对，涵盖多样场景和复杂性。我们在VISTA上系统评估了11个最先进的VLM，并分解了跨分类法的聚合性能，揭示了传统指标掩盖的缺陷和显著的时空偏差。通过在具有挑战性的数据集上提供详细的、分类法驱动的诊断，VISTA提供了一个精细的框架来指导模型设计、预训练策略和评估协议的进步。总体而言，VISTA是第一个大规模、交互感知的VLM时空理解诊断基准。

英文摘要

Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute sets and restricted entity types, failing to capture the freeform, multi-action interactions between diverse entities which characterize real-world video understanding. Furthermore, the lack of a systematic framework for analyzing model failures across complementary spatio-temporal axes hinders comprehensive evaluation. To address these gaps, we introduce VISTA, a Video Interaction Spatio-Temporal Analysis benchmark designed for open-set, multi-entity and multi-action spatio-temporal understanding in VLMs. VISTA decomposes videos into interpretable entities, their associated actions, and relational dynamics, enabling multi-axis diagnostics and unified assessment of relational, spatial, and temporal understanding. Our benchmark integrates multiple datasets into a single interaction-aware taxonomy and comprises ~12K curated video-query pairs spanning diverse scenes and complexities. We systematically evaluate 11 state-of-the-art VLMs on VISTA, and break down aggregate performance across our taxonomy to reveal shortcomings and pronounced spatio-temporal biases obscured by traditional metrics. By providing detailed, taxonomy-driven diagnostics on a challenging dataset, VISTA offers a nuanced framework to guide advances in model design, pretraining strategies, and evaluation protocols. Overall, VISTA is the first, large-scale, interaction-aware diagnostic benchmark for spatio-temporal understanding in VLMs.

URL PDF HTML ☆

赞 0 踩 0

2601.19827 2026-06-12 cs.CL cs.AI cs.IR 版本更新

When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering

当迭代RAG优于理想证据：科学多跳问答中的诊断研究

Mahdi Astaraki, Mohammad Arshi Saloot, Ali Shiraee Kasmaee, Hamidreza Mahyar, Soheila Samiee

发表机构 * Faculty of Engineering, McMaster University, Canada（麦斯特大学工程学院，加拿大）； BASF Canada Inc., Canada（巴斯夫加拿大公司，加拿大）

AI总结通过化学多跳问答数据集，诊断发现迭代检索-推理循环在科学领域显著优于静态RAG上限，揭示了阶段式检索的优势与失败模式。

Comments 51 pages, 29 figures

详情

AI中文摘要

检索增强生成（RAG）将大型语言模型（LLMs）扩展到参数化知识之外，但目前尚不清楚迭代检索-推理循环何时能有效超越静态RAG，尤其是在涉及多跳推理、稀疏领域知识和异构证据的科学领域。我们首次进行了受控的、机制层面的诊断研究，以探究同步迭代检索和推理能否超越理想化的静态上限（Gold Context）RAG。我们在三种设置下对十一个最先进的LLM进行了基准测试：（i）无上下文，衡量对参数化记忆的依赖；（ii）Gold Context，一次性提供所有真实证据；（iii）迭代RAG，一种无需训练的控制器，交替进行检索、假设细化和证据感知停止。使用以化学为中心的ChemKGMultiHopQA数据集，我们分离出需要真正检索的问题，并通过诊断分析行为，涵盖检索覆盖缺口、锚点携带下降、查询质量、组合保真度和控制校准。在所有模型中，迭代RAG始终优于Gold Context，增益高达25.6个百分点，尤其对于非推理微调模型。阶段式检索减少了后期跳失败，缓解了上下文过载，并实现了对早期假设漂移的动态修正，但剩余的失败模式包括跳覆盖不完整、干扰物锁定轨迹、过早停止校准错误以及即使检索完美时的高组合失败率。总体而言，阶段式检索通常比理想证据的单纯存在更具影响力；我们为在专门科学环境中部署和诊断RAG系统提供了实用指导，并为更可靠、可控的迭代检索-推理框架奠定了基础。

英文摘要

Retrieval-Augmented Generation (RAG) extends large language models (LLMs) beyond parametric knowledge, yet it is unclear when iterative retrieval-reasoning loops meaningfully outperform static RAG, particularly in scientific domains with multi-hop reasoning, sparse domain knowledge, and heterogeneous evidence. We provide the first controlled, mechanism-level diagnostic study of whether synchronized iterative retrieval and reasoning can surpass an idealized static upper bound (Gold Context) RAG. We benchmark eleven state-of-the-art LLMs under three regimes: (i) No Context, measuring reliance on parametric memory; (ii) Gold Context, where all oracle evidence is supplied at once; and (iii) Iterative RAG, a training-free controller that alternates retrieval, hypothesis refinement, and evidence-aware stopping. Using the chemistry-focused ChemKGMultiHopQA dataset, we isolate questions requiring genuine retrieval and analyze behavior with diagnostics spanning retrieval coverage gaps, anchor-carry drop, query quality, composition fidelity, and control calibration. Across models, Iterative RAG consistently outperforms Gold Context, with gains up to 25.6 percentage points, especially for non-reasoning fine-tuned models. Staged retrieval reduces late-hop failures, mitigates context overload, and enables dynamic correction of early hypothesis drift, but remaining failure modes include incomplete hop coverage, distractor latch trajectories, early stopping miscalibration, and high composition failure rates even with perfect retrieval. Overall, staged retrieval is often more influential than the mere presence of ideal evidence; we provide practical guidance for deploying and diagnosing RAG systems in specialized scientific settings and a foundation for more reliable, controllable iterative retrieval-reasoning frameworks.

URL PDF HTML ☆

赞 0 踩 0

2605.00600 2026-06-12 cs.LG cs.AI cs.CV 版本更新

Possibilistic Predictive Uncertainty for Deep Learning

深度学习的可能性预测不确定性

Yao Ni, Jeremie Houssineau, Yew-Soon Ong, Piotr Koniusz

发表机构 * University of Cambridge（剑桥大学）； National University of Singapore（新加坡国立大学）； University of Warsaw（华沙大学）

AI总结提出基于可能性理论的Dirichlet近似可能性后验预测（DAPPr）框架，通过投影-近似策略实现高效且原则性的认知不确定性量化，在多个基准上达到竞争性能。

Comments Accepted by ICML 2026, 20 pages

详情

AI中文摘要

深度神经网络在多种应用中取得了令人印象深刻的结果，然而它们对未见输入的过度自信需要可靠的认知不确定性建模。现有的不确定性建模方法面临一个基本困境：贝叶斯方法提供原则性的估计，但计算成本高昂，而高效的二阶预测器在其特定目标与认知不确定性量化之间缺乏严格联系。为解决这一困境，我们引入了Dirichlet近似可能性后验预测（DAPPr），一个基于可能性理论的原则性框架。我们定义了参数上的可能性后验，通过上确界算子将其投影到预测空间，并使用可学习的Dirichlet可能性函数近似投影后的后验。这种投影-近似策略产生了一个具有闭式解的简单训练目标。尽管简单，跨多个不同基准的大量实验表明，DAPPr在保持原则性推导和计算效率的同时，实现了与最先进的二阶预测器相当或更优的不确定性量化性能。代码可在 https://github.com/MaxwellYaoNi/DAPPr 获取。

英文摘要

Deep neural networks achieve impressive results across diverse applications, yet their overconfidence on unseen inputs necessitates reliable epistemic uncertainty modeling. Existing methods for uncertainty modeling face a fundamental dilemma: Bayesian approaches provide principled estimates but remain computationally prohibitive, while efficient second-order predictors lack rigorous connections between their specific objectives and epistemic uncertainty quantification. To resolve this dilemma, we introduce Dirichlet-approximated possibilistic posterior predictions (DAPPr), a principled framework grounded in possibility theory. We define a possibilistic posterior over parameters, project it to the prediction space via supremum operators, and approximate the projected posterior using learnable Dirichlet possibility functions. This projection-and-approximation strategy yields a simple training objective with closed-form solutions. Despite its simplicity, extensive experiments across diverse benchmarks show that DAPPr achieves competitive or superior uncertainty quantification performance over state-of-the-art second-order predictors while maintaining both principled derivation and computational efficiency. Code is available at https://github.com/MaxwellYaoNi/DAPPr.

URL PDF HTML ☆

赞 0 踩 0

2604.27960 2026-06-12 cs.AI 版本更新

LLMs as ASP Programmers: Self-Correction Enables Task-Agnostic Nonmonotonic Reasoning

LLMs 作为 ASP 程序员：自我纠正实现任务无关的非单调推理

Adam Ishay, Joohyung Lee

发表机构 * Arizona State University（亚利桑那州立大学）； Samsung Research（三星研究院）

AI总结提出 LLM+ASP 框架，通过自我纠正循环将自然语言转化为回答集程序，实现无需任务特定工程的非单调推理，在多个基准上优于 SMT 方法。

Comments 30 pages

详情

AI中文摘要

近期的大语言模型（LLMs）在推理方面取得了令人瞩目的进展，但仍面临高计算成本、逻辑不一致性以及在高度复杂问题上性能急剧下降等问题。神经符号方法通过将 LLMs 与符号推理器结合来缓解这些问题，但现有方法通常依赖于单调逻辑（如 SMT），无法表示可废止推理——人类认知的重要组成部分。我们提出了“LLM+ASP”框架，该框架将自然语言转化为回答集编程（ASP），一种基于稳定模型语义的非单调形式化方法。与先前需要手动编写知识模块、领域特定提示或仅限于单一问题类别评估的“LLM+ASP”方法不同，我们的框架无需任何每任务工程，并统一适用于多种推理任务。我们的系统利用自动化的自我纠正循环，其中来自 ASP 求解器的结构化反馈能够实现迭代优化。在六个不同基准上的评估表明：（1）稳定模型语义使 LLMs 能够自然地表达默认规则和例外，在非单调任务上显著优于基于 SMT 的替代方法；（2）迭代自我纠正是性能的主要驱动力，有效替代了手工领域知识的需求；（3）紧凑的上下文参考指南显著优于冗长的文档，揭示了“上下文腐烂”现象，即过多上下文会阻碍约束遵循。

英文摘要

Recent large language models (LLMs) have achieved impressive reasoning milestones but continue to struggle with high computational costs, logical inconsistencies, and sharp performance degradation on high-complexity problems. While neuro-symbolic methods attempt to mitigate these issues by coupling LLMs with symbolic reasoners, existing approaches typically rely on monotonic logics (e.g., SMT) that cannot represent defeasible reasoning -- essential components of human cognition. We present "LLM+ASP," a framework that translates natural language into Answer Set Programming (ASP), a nonmonotonic formalism based on stable model semantics. Unlike prior "LLM+ASP" approaches that require manually authored knowledge modules, domain-specific prompts, or evaluation restricted to single problem classes, our framework operates without any per-task engineering and applies uniformly across diverse reasoning tasks. Our system utilizes an automated self-correction loop where structured feedback from the ASP solver enables iterative refinement. Evaluating across six diverse benchmarks, we demonstrate that: (1) stable model semantics allow LLMs to naturally express default rules and exceptions, outperforming SMT-based alternatives by significant margins on nonmonotonic tasks; (2) iterative self-correction is the primary driver of performance, effectively replacing the need for handcrafted domain knowledge; (3) compact in-context reference guides substantially outperform verbose documentation, revealing a "context rot" phenomenon where excessive context hinders constraint adherence.

URL PDF HTML ☆

赞 0 踩 0

2604.27277 2026-06-12 cs.LG cs.AI cs.CV 版本更新

BrainDINO: A Brain MRI Foundation Model for Generalizable Clinical Representation Learning

BrainDINO：一种用于通用临床表征学习的脑MRI基础模型

Yizhou Wu, Shansong Wang, Yuheng Li, Mojtaba Safari, Mingzhe Hu, Chih-Wei Chang, Harini Veeraraghavan, Xiaofeng Yang

发表机构 * Department of Radiation Oncology and Winship Cancer Institute, Emory University（放射肿瘤科和Winship癌症研究所，埃默里大学）； Department of Radiation and Cellular Oncology, The University of Chicago（放射肿瘤学与细胞肿瘤学部，芝加哥大学）； Department of Electrical and Computer Engineering, Georgia Institute of Technology（电气与计算机工程系，佐治亚理工学院）； Department of Biomedical Engineering, Georgia Institute of Technology（生物医学工程系，佐治亚理工学院）； Department of Biomedical Informatics, Emory University（生物医学信息学系，埃默里大学）； Department of Medical Physics, Memorial Sloan Kettering Cancer Center（医学物理系，纪念斯隆凯特琳癌症中心）

AI总结提出BrainDINO，一种基于自蒸馏的基础模型，在约660万张未标记轴向切片上训练，通过冻结编码器加轻量任务头，在多种脑MRI任务上达到或超越基线，尤其在小样本场景下优势显著。

Comments 25 pages, 5 figures

详情

AI中文摘要

脑MRI支撑着广泛的神经科学和临床应用，然而大多数基于学习的方法仍针对特定任务且需要大量标注数据。本文表明，单一的自监督表征可以泛化到异质的脑MRI终点。我们训练了BrainDINO，一个自蒸馏的基础模型，使用了来自20个数据集的约660万张未标记轴向切片，这些数据集涵盖了人群、疾病和采集设置的广泛变异。通过使用冻结编码器加轻量任务头，BrainDINO支持肿瘤分割、神经退行性和神经发育性疾病分类、脑年龄估计、卒中后时间预测、分子状态预测、MRI序列分类和生存建模等任务的迁移。在各种任务和监督机制下，BrainDINO始终等于或超过自然图像和MRI特定自监督基线，在标签稀缺时尤其具有优势。表征分析进一步显示，在缺乏任务特定监督的情况下，特征结构具有解剖学组织和病理敏感性。我们的发现表明，大规模切片级自监督学习可以产生统一的脑MRI表征，支持多样化的神经影像任务，无需体积预训练或全网络微调，为稳健且数据高效的脑影像分析建立了可扩展的基础。代码可在 https://github.com/mclwu22/BrainDINO 获取。

英文摘要

Brain MRI underpins a wide range of neuroscientific and clinical applications, yet most learning-based methods remain task-specific and require substantial labeled data. Here we show that a single self-supervised representation can generalize across heterogeneous brain MRI endpoints. We trained BrainDINO, a self-distilled foundation model, on approximately 6.6 million unlabeled axial slices from 20 datasets encompassing broad variation in population, disease, and acquisition setting. Using a frozen encoder with lightweight task heads, BrainDINO supported transfer across tumor segmentation, neurodegenerative and neurodevelopmental conditions classification, brain age estimation, post-stroke temporal prediction, molecular status prediction, MRI sequence classification, and survival modeling. Across tasks and supervision regimes, BrainDINO consistently equaled or exceeded natural-image and MRI-specific self-supervised baselines, with particularly strong advantages under label scarcity. Representation analyses further showed anatomically organized and pathology-sensitive feature structure in the absence of task-specific supervision. Our findings indicate that large-scale slice-wise self-supervised learning can yield a unified brain MRI representation that supports diverse neuroimaging tasks without volumetric pretraining or full-network fine-tuning, establishing a scalable foundation for robust and data-efficient brain imaging analysis. Code is available at https://github.com/mclwu22/BrainDINO

URL PDF HTML ☆

赞 0 踩 0

2604.26940 2026-06-12 cs.CL 版本更新

Select to Think: Unlocking SLM Potential with Local Sufficiency

Select to Think: 利用局部充分性解锁小语言模型潜力

Wenxuan Ye, Yangyang Zhang, Xueli An, Georg Carle, Yunpu Ma

发表机构 * University of Science and Technology of China（中国科学技术大学）

AI总结提出Select to Think (S2T)方法，通过将大语言模型角色从生成转为选择，并蒸馏选择逻辑到小语言模型，使其在推理时无需依赖大模型，显著提升性能。

Comments Accepted to ICML 2026. Code is available at https://github.com/YeRona/Select-to-Think

详情

AI中文摘要

小语言模型（SLM）部署高效，但在推理能力上常落后于大语言模型（LLM）。现有解决方案要么在推理分歧点调用LLM，导致大量延迟和成本，要么依赖标准蒸馏，受限于SLM准确模仿LLM复杂生成分布的能力。我们通过识别局部充分性来解决这一困境：在分歧点，LLM偏好的token通常位于SLM的top-K预测中，即使未能成为SLM的top-1选择。因此，我们提出Select to Think（S2T），将LLM的角色从开放式生成重新定义为在SLM的候选提案中进行选择，将监督信号简化为离散的候选排名。利用这一点，我们引入S2T-Local，将选择逻辑蒸馏到SLM中，使其能够在推理时自主重新排序，无需依赖LLM。实验表明，1.5B SLM的top-8候选包含32B LLM选择的命中率达95%，S2T-Local使1.5B SLM的数学平均相对贪心解码提升24.1%，以单轨迹效率达到8路径自一致性的效果。

英文摘要

Small language models (SLMs) offer efficient deployment, yet they often lag behind their larger counterparts (LLMs) in reasoning. Existing remedies either invoke an LLM at points of reasoning divergence, incurring substantial latency and cost, or rely on standard distillation, which is limited by the SLM's capacity to accurately mimic the LLM's complex generative distribution. We address this dilemma by identifying local sufficiency: at divergence points, the LLM's preferred token often resides within the SLM's top-K next-token predictions, even when failing to emerge as the SLM top-1 choice. We therefore propose Select to Think (S2T), which reframes the LLM's role from open-ended generation to selection among the SLM's proposals, simplifying the supervision signal to discrete candidate rankings. Leveraging this, we introduce S2T-Local, which distills the selection logic into the SLM, empowering it to perform autonomous re-ranking without inference-time LLM dependency. Empirically, a 1.5B SLM's top-8 candidates contain the 32B LLM's choice with a 95% hit rate, and S2T-Local improves the 1.5B SLM's Math Avg. over greedy decoding by 24.1% relative gain, matching the efficacy of 8-path self-consistency with single-trajectory efficiency.

URL PDF HTML ☆

赞 0 踩 0

2604.24079 2026-06-12 cs.CL cs.AI 版本更新

The Pragmatic Persona: Discovering LLM Persona through Bridging Inference

实用人格：通过桥接推理发现LLM人格

Jisoo Yang, Jongwon Ryu, Minuk Ma, Trung X. Pham, Junyeong Kim

发表机构 * Department of Artificial Intelligence, Chung-Ang University, Seoul, 06974, Republic of Korea（Chung-Ang大学人工智能系）； Department of Computer Science, University of British Columbia, Vancouver, BC V6T 1Z4, Canada（不列颠哥伦比亚大学计算机科学系）； Van Lang University, Ho Chi Minh City, Vietnam（文-lang大学）

AI总结提出基于桥接推理的框架，通过构建话语级知识图谱捕捉LLM对话中的隐含语义关联，实现从话语连贯性层面发现稳定人格特征，优于基于频率或风格的基线方法。

Comments 15 pages, 4 figures, accepted to ICPR 2026

详情

AI中文摘要

大型语言模型（LLM）通过对话展现出固有且独特的人格。然而，现有的大多数人格发现方法依赖于表面层面的词汇或风格线索，将对话视为平坦的token序列，未能捕捉维持人格一致性的更深层次话语结构。为解决这一局限，我们提出一种新颖的分析框架，通过桥接推理——即通过共享世界知识和话语连贯性连接话语的隐含概念关系——来解读LLM对话。通过将这些关系建模为结构化知识图谱，我们的方法捕捉了控制LLM在对话轮次间组织意义的潜在语义链接，从而在话语连贯性层面而非表面实现上实现人格发现。在多种推理骨干和从小型模型到80B参数系统的目标LLM上的实验结果表明，与基于频率或风格的基线相比，桥接推理图产生了显著更强的语义连贯性和更稳定的人格识别。这些结果表明，人格特质始终编码在话语的结构组织中，而非孤立的词汇模式中。本工作提出了一个系统框架，通过认知话语理论的视角来探测、提取和可视化潜在的LLM人格，桥接了计算语言学、认知语义学和大型语言模型中的人格推理。代码见：https://this URL

英文摘要

Large Language Models (LLMs) reveal inherent and distinctive personas through dialogue. However, most existing persona discovery approaches rely on surface-level lexical or stylistic cues, treating dialogue as a flat sequence of tokens and failing to capture the deeper discourse-level structures that sustain persona consistency. To address this limitation, we propose a novel analytical framework that interprets LLM dialogue through bridging inference -- implicit conceptual relations that connect utterances via shared world knowledge and discourse coherence. By modeling these relations as structured knowledge graphs, our approach captures latent semantic links that govern how LLMs organize meaning across turns, enabling persona discovery at the level of discourse coherence rather than surface realizations. Experimental results across multiple reasoning backbones and target LLMs, ranging from small-scale models to 80B-parameter systems, demonstrate that bridging-inference graphs yield significantly stronger semantic coherence and more stable persona identification than frequency or style-based baselines. These results show that persona traits are consistently encoded in the structural organization of discourse rather than isolated lexical patterns. This work presents a systematic framework for probing, extracting, and visualizing latent LLM personas through the lens of Cognitive Discourse Theory, bridging computational linguistics, cognitive semantics, and persona reasoning in large language models. Codes are available at https://github.com/JiSoo-Yang/Persona_Bridging.git

URL PDF HTML ☆

赞 0 踩 0

2508.04427 2026-06-12 cs.LG cs.AI 版本更新

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models

解码多模态迷宫：多模态注意力模型中可解释性采纳的系统综述

Md Raisul Kibria, Sébastien Lafond, Janan Arslan

发表机构 * University of Science and Technology of China（中国科学技术大学）

AI总结本文系统综述了2020年至2024年初多模态模型可解释性研究，发现多数工作集中于视觉-语言和纯语言模型，注意力机制是主要解释方法，但评估缺乏系统性和鲁棒性，并提出了改进建议。

详情

DOI: 10.1016/j.inffus.2026.104405

AI中文摘要

近年来，多模态学习取得了显著进展，特别是随着注意力模型的整合，在各种任务中带来了显著的性能提升。与此同时，对可解释人工智能（XAI）的需求推动了越来越多的研究，旨在解释这些模型的复杂决策过程。本系统文献综述分析了2020年1月至2024年初期间发表的、关注多模态模型可解释性的研究。在XAI更广泛目标的框架内，我们从多个维度审视文献，包括模型架构、涉及模态、解释算法和评估方法。我们的分析显示，大多数研究集中在视觉-语言和纯语言模型上，注意力机制是最常用的解释方法。然而，这些方法往往无法捕捉模态间交互的全谱系，这一问题因领域间的架构异质性而进一步加剧。重要的是，我们发现多模态环境中XAI的评估方法大多是非系统性的，缺乏一致性、鲁棒性，并且未考虑模态特定的认知和上下文因素。为解决这些不足，我们不仅综合了所调查研究的发现，还纳入了补充分析，整合了推动多模态可解释性的近期和新兴进展。基于这些见解，我们提出了一套全面的建议，旨在促进多模态XAI研究中严谨、透明和标准化的评估与报告实践。我们的目标是支持未来构建更可解释、可问责和负责任的多模态AI系统，并以可解释性为核心。

英文摘要

Multimodal learning has witnessed remarkable advancements in recent years, particularly with the integration of attention-based models, leading to significant performance gains across a variety of tasks. Parallel to this progress, the demand for explainable artificial intelligence (XAI) has spurred a growing body of research aimed at interpreting the complex decision-making processes of these models. This systematic literature review analyzes research published between January 2020 and early 2024 that focuses on the explainability of multimodal models. Framed within the broader goals of XAI, we examine the literature across multiple dimensions, including model architecture, modalities involved, explanation algorithms and evaluation methodologies. Our analysis reveals that most studies are concentrated on vision-language and language-only models, with attention-based techniques being the most commonly employed for explanation. However, these methods often fall short in capturing the full spectrum of interactions between modalities, a challenge further compounded by the architectural heterogeneity across domains. Importantly, we find that evaluation methods for XAI in multimodal settings are largely non-systematic, lacking consistency, robustness, and consideration for modality-specific cognitive and contextual factors. To address these gaps, we not only synthesize findings from the surveyed works but also incorporate a complementary analysis that integrates recent and emerging advances driving multimodal explainability. Based on these insights, we provide a comprehensive set of recommendations aimed at promoting rigorous, transparent, and standardized evaluation and reporting practices in multimodal XAI research. Our goal is to support future research in more interpretable, accountable, and responsible multimodal AI systems, with explainability at their core.

URL PDF HTML ☆

赞 0 踩 0

2604.23165 2026-06-12 cs.CV 版本更新

BSViT: A Burst Spiking Vision Transformer for Expressive and Efficient Visual Representation Learning

BSViT：用于高效表达视觉表征学习的脉冲视觉Transformer

Hongxiang Peng, Dewei Bai, Hong Qu

发表机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China（电子科技大学计算机科学与工程学院）

AI总结提出BSViT，通过双通道爆发脉冲自注意力机制和局部邻域掩码策略，解决脉冲视觉Transformer中二进制脉冲信息容量有限和全局自注意力密集交互的问题，在静态和事件视觉基准上取得更高精度和能效。

Comments Accepted by ECML PKDD 2026

详情

AI中文摘要

脉冲视觉Transformer（S-ViT）为节能视觉学习提供了有前景的框架。然而，现有设计仍受限于两个基本问题：二进制脉冲编码的信息容量有限以及全局自注意力引入的密集令牌交互。为应对这些挑战，本文提出BSViT，一种爆发脉冲驱动的视觉Transformer，具有双通道爆发脉冲自注意力（DBSSA）机制。DBSSA用二进制脉冲编码查询，用爆发脉冲编码键以增强表示能力。值通路采用双兴奋性和抑制性二进制通道，实现有符号调制和更丰富的脉冲交互。重要的是，整个注意力操作保持仅加法计算，确保与节能神经形态硬件的兼容性。为进一步降低脉冲活动并融入空间先验，引入补丁邻域掩码策略将注意力限制在局部邻域，实现结构感知稀疏性并减少计算开销。此外，爆发脉冲编码被系统地集成到网络中，以提升脉冲级表示能力，超越传统二进制脉冲。在静态和事件视觉基准上的大量实验表明，BSViT在精度上持续优于现有脉冲Transformer，同时保持有竞争力的能效。

英文摘要

Spiking Vision Transformers (S-ViTs) offer a promising framework for energy-efficient visual learning. However, existing designs remain limited by two fundamental issues: the restricted information capacity of binary spike coding and the dense token interactions introduced by global self-attention. To address these challenges, this work proposes BSViT, a burst spiking-driven Vision Transformer featuring a Dual-Channel Burst Spiking Self-Attention (DBSSA) mechanism. DBSSA encodes queries with binary spikes and keys with burst spikes to enhance representational capacity. The value pathway adopts dual excitatory and inhibitory binary channels, enabling signed modulation and richer spike interactions. Importantly, the entire attention operation preserves addition-only computation, ensuring compatibility with energy-efficient neuromorphic hardware. To further reduce spike activity and incorporate spatial priors, a patch adjacency masking strategy is introduced to restrict attention to local neighborhoods, resulting in structure-aware sparsity and reduced computational overhead. In addition, burst spike coding is systematically integrated across the network to increase spike-level representational capacity beyond conventional binary spiking. Extensive experiments on both static and event-based vision benchmarks demonstrate that BSViT consistently outperforms existing spiking Transformers in accuracy while maintaining competitive energy efficiency.

URL PDF HTML ☆

赞 0 踩 0

2506.18493 2026-06-12 cs.CV 版本更新

ShowFlow: From Robust Single Concept to Condition-Free Multi-Concept Generation

ShowFlow: 从鲁棒的单概念到无条件的多概念生成

Trong-Vu Hoang, Quang-Binh Nguyen, Thanh-Toan Do, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le

发表机构 * University of Science（科学大学）； Vietnam National University（越南国家大学）； Monash University（墨尔本大学）； University of Dayton（Dayton大学）

AI总结提出ShowFlow框架，通过KronA-WED适配器和语义感知注意力正则化增强单概念生成，并利用SAMA和布局一致性指导实现无额外条件的多概念生成。

详情

AI中文摘要

定制化图像生成仍然是可控图像合成中的核心挑战。对于单概念生成，保持身份保留和提示对齐是困难的。在多概念场景中，仅依赖提示而不使用布局框或语义掩码等额外条件，通常会导致身份丢失和概念遗漏。在本文中，我们介绍了ShowFlow，一个旨在应对这些挑战的全面框架。我们提出了用于单概念图像生成的ShowFlow-S，以及用于处理多个概念的ShowFlow-M。ShowFlow-S引入了一个KronA-WED适配器，它将Kronecker适配器与权重和嵌入分解相结合，并配合一种新颖的语义感知注意力正则化（SAR）训练目标，以增强单概念生成。在此基础上，ShowFlow-M直接重用由ShowFlow-S学习的鲁棒模型，以支持无需额外条件的多概念生成，并集成了主体自适应匹配注意力（SAMA）和布局一致性指导作为即插即用模块。大量实验和用户研究验证了ShowFlow的有效性，突显了其在广告和虚拟试穿等实际应用中的潜力。我们的源代码将在以下网址公开：this https URL。

英文摘要

Customizing image generation remains a core challenge in controllable image synthesis. For single-concept generation, maintaining both identity preservation and prompt alignment is challenging. In multi-concept scenarios, relying solely on a prompt without additional conditions like layout boxes or semantic masks, often leads to identity loss and concept omission. In this paper, we introduce ShowFlow, a comprehensive framework designed to tackle these challenges. We propose ShowFlow-S for single-concept image generation, and ShowFlow-M for handling multiple concepts. ShowFlow-S introduces a KronA-WED adapter, which integrates a Kronecker adapter with weight and embedding decomposition, and together with a novel Semantic-Aware Attention Regularization (SAR) training objective to enhance single-concept generation. Building on this foundation, ShowFlow-M directly reuses robust models learned by ShowFlow-S to support multi-concept generation without extra conditions, incorporating a Subject-Adaptive Matching Attention (SAMA) and a Layout Consistency guidance as the plug-and-play module. Extensive experiments and user studies validate ShowFlow's effectiveness, highlighting its potential in real-world applications like advertising and virtual dressing. Our source code will be publicly available at: https://htrvu.github.io/showflow.

URL PDF HTML ☆

赞 0 踩 0

2604.20236 2026-06-12 cs.LG 版本更新

Machine Learning-based Two-Stage Graph Sparsification for the Travelling Salesman Problem

基于机器学习的两阶段图稀疏化方法用于旅行商问题

Bo-Cheng Lin, Yi Mei, Mengjie Zhang

发表机构 * Centre for Data Science and Artificial Intelligence（数据科学与人工智能中心）； School of Engineering and Computer Science（工程与计算机科学学院）； Victoria University of Wellington（惠灵顿维多利亚大学）

AI总结提出两阶段方法，先结合α-Nearest和POPMUSIC得到近完美召回率的候选图，再用轻量级分类器修剪单源边，在保持≥99.69%最优边的同时降低37%-47%密度。

详情

AI中文摘要

高性能TSP求解器（如Lin-Kernighan-Helsgaun (LKH)）在\emph{候选图}（为求解器预先选定的边的小子集）中搜索，而不是在完整图上搜索。两种主要的稀疏化启发式方法，$\alpha$-Nearest和POPMUSIC，各自在密度-覆盖率平衡上存在不足：$\alpha$-Nearest密集且召回率稳定，而POPMUSIC更稀疏但其召回率随规模增大而下降。它们的并集在密度上远低于完整图的同时弥补了召回率差距，为进一步缩减留下了空间。现有的基于学习的稀疏化方法在完整图上对边评分，这种方法代价高昂且主要限于欧几里得实例。我们提出了一种两阶段方法，反转了这一逻辑。第一阶段取$\alpha$-Nearest和POPMUSIC的并集，在${\sim}6N$条边上实现近乎完美的召回率。关键在于，并集为每条边标注了其\emph{来源出处}——即它是由$\alpha$-Nearest、POPMUSIC还是两者共同支持的。第二阶段在这些标注边上训练一个轻量级分类器，并修剪得分最低的边。由于双源边几乎总是最优的，学习问题简化为过滤单源子集——这比从头开始对所有$O(N^2)$条边进行分类要容易得多。在四种距离类型、五种空间分布以及50到500的问题规模上，该流程将候选图密度降低了37%-47%，同时保留了${\geq}99.69\%$的最优旅行边，并且在TSP500上以更低的密度达到或超过了近期仅限欧几里得的神经稀疏化方法的覆盖率。

英文摘要

High-performance TSP solvers such as Lin-Kernighan-Helsgaun (LKH) search within a \emph{candidate graph} -- a small subset of edges pre-selected for the solver -- rather than over the complete graph. The two leading sparsification heuristics, $α$-Nearest and POPMUSIC, each fall short of the density-coverage balance: $α$-Nearest is dense with stable recall, while POPMUSIC is sparser but its recall degrades with scale. Their union closes the recall gap while remaining far below the complete graph in density, leaving room for further reduction. Existing learning-based sparsifiers score edges on the complete graph, an approach that is expensive and largely limited to Euclidean instances. We propose a two-stage method that inverts this logic. Stage~1 takes the union of $α$-Nearest and POPMUSIC, achieving near-perfect recall at ${\sim}6N$ edges. Crucially, the union annotates each edge with its \emph{source provenance} -- whether it was endorsed by $α$-Nearest, POPMUSIC, or both. Stage~2 trains a lightweight classifier on these annotated edges and prunes the lowest-scoring ones. Because dual-source edges are almost always optimal, the learning problem reduces to filtering the single-source subset -- a substantially easier task than classifying all $O(N^2)$ edges from scratch. Across four distance types, five spatial distributions, and problem sizes from 50 to 500, the pipeline reduces candidate-graph density by $37$-$47\%$ while retaining ${\geq}99.69\%$ of optimal-tour edges, and matches or exceeds the coverage of recent Euclidean-only neural sparsifiers at lower density at TSP500.

URL PDF HTML ☆

赞 0 踩 0

2604.18307 2026-06-12 cs.CL 版本更新

Reasoning Models Know What's Important, and Encode It in Their Activations

推理模型知道什么重要，并在其激活中编码

Yaniv Nikankin, Martin Tutek, Tomer Ashuach, Jonathan Rosenfeld, Yonatan Belinkov

发表机构 * Technion（技术离子大学）； University of Zagreb, FER（扎格雷布大学，FER）； MIT（麻省理工学院）； Kempner Institute, Harvard（哈佛大学凯普纳研究所）

AI总结通过分析模型激活而非仅依赖推理链文本，发现激活能更有效识别关键推理步骤，且模型在生成后续步骤前已内部编码步骤重要性。

详情

AI中文摘要

语言模型通常通过生成包含许多重要性不同的步骤的长推理链来解决复杂任务。虽然某些步骤对生成最终答案至关重要，但其他步骤是可移除的。确定哪些步骤最重要以及为什么，仍然是理解模型如何处理推理的核心开放问题。我们研究了这个问题是通过模型内部还是通过推理链本身的标记来最好地解决。我们发现，模型激活比标记包含更多信息，用于识别重要的推理步骤。关键的是，通过在模型激活上训练探针来预测重要性，我们表明模型在生成后续步骤之前就已经编码了步骤重要性的内部表示。不同模型中重要性的内部表示在哪些步骤重要上具有高度一致性。这种表示分布在各个层中，并且与表面特征（如步骤的相对位置或长度）不相关。我们的发现表明，分析激活可以揭示表面方法根本遗漏的推理方面，表明推理分析应该研究模型内部。

英文摘要

Language models often solve complex tasks by generating long reasoning chains, consisting of many steps with varying importance. While some steps are crucial for generating the final answer, others are removable. Determining which steps matter most, and why, remains an open question central to understanding how models process reasoning. We investigate if this question is best approached through model internals or through tokens of the reasoning chain itself. We find that model activations contain more information than tokens for identifying important reasoning steps. Crucially, by training probes on model activations to predict importance, we show that models encode an internal representation of step importance, even prior to the generation of subsequent steps. The internal representations of importance in different models yield high agreement on which steps are important. The representation is distributed across layers, and does not correlate with surface-level features, such as a step's relative position or its length. Our findings suggest that analyzing activations can reveal aspects of reasoning that surface-level approaches fundamentally miss, indicating that reasoning analyses should look into model internals.

URL PDF HTML ☆

赞 0 踩 0

2601.00921 2026-06-12 cs.LG cs.AI quant-ph 版本更新

Geometric and Quantum Kernel Methods for Predicting Skeletal Muscle Outcomes in chronic obstructive pulmonary disease

用于预测慢性阻塞性肺疾病骨骼肌结果的几何与量子核方法

Azadeh Alavi, Hamidreza Khalili, Stanley H. Chan, Fatemeh Kouchmeshki, Muhammad Usman, Ross Vlahos

发表机构 * School of Computing Technologies, RMIT University（计算技术学院，拉筹纳斯大学）； School of Health & Biomedical Sciences, STEM College, RMIT University（健康与生物医学科学学院，STEM学院，拉筹纳斯大学）； Pattern Recognition Pty Ltd, Melbourne（模式识别有限公司，墨尔本）； Data61, CSIRO（Data61，澳大利亚联邦科学与工业研究组织）

AI总结提出一种核几何量子混合方法，通过再生核希尔伯特空间映射合成SPD参考、随机投影压缩和低维量子回归电路，在COPD动物队列中预测肌肉重量、质量和力量，肌肉重量RMSE比最佳经典方法低约1.8%。

Comments 24 pages, 2 figures

详情

AI中文摘要

慢性阻塞性肺疾病（COPD）影响全球数亿人，骨骼肌功能障碍具有临床重要性。量子机器学习在生物医学预测中日益受到探索，但在小型生物标志物队列中的价值需要与强经典基线进行基准测试。我们分析了一个由213只动物组成的香烟烟雾COPD队列，利用血液和支气管肺泡灌洗生物标志物预测胫骨前肌重量、肌肉质量和力量。我们开发了一种核几何量子混合方法，其中合成对称正定（SPD）参考通过再生核希尔伯特空间映射，使用仅训练随机投影压缩，归一化，并输入低维量子回归电路。我们将该方法与经典岭/核模型、SPD关系表示和量子核回归（QKR）进行了基准测试。所有方法均使用条件分层重复交叉验证进行评估。最大的数值改进出现在肌肉重量上，所提出方法的平均均方根误差（RMSE）数值最低，比最佳经典比较器低约1.8%；配对折叠水平测试在Holm调整后未建立统计显著性优势，但该终点具有生物学意义。该方法在肌肉质量上也具有数值最低的平均RMSE。对于力量，仅使用生物标志物的岭回归表现最佳，表明更线性的终点结构。

英文摘要

Chronic obstructive pulmonary disease (COPD) affects hundreds of millions of people worldwide, and skeletal-muscle dysfunction is clinically important. Quantum machine learning is increasingly explored for biomedical prediction, but its value in small biomarker cohorts requires benchmarking against strong classical baselines. We analysed a cigarette-smoke COPD cohort of 213 animals with blood and bronchoalveolar-lavage biomarkers to predict tibialis anterior muscle weight, muscle quality, and force. We developed a kernel-geometric quantum hybrid method in which synthetic symmetric positive definite (SPD) references are mapped through a reproducing kernel Hilbert space, compressed using train-only random projection, normalised, and supplied to low-dimensional quantum regression circuits. We benchmarked this approach against classical ridge/kernel models, SPD relational representations, and quantum-kernel regression (QKR). All methods were evaluated using condition-stratified repeated cross-validation. The largest numerical improvement was observed for muscle weight, where the proposed method had the numerically lowest mean root mean squared error (RMSE), approximately 1.8% below the best classical comparator; paired fold-level testing did not establish statistically significant superiority after Holm adjustment, but the endpoint is biologically meaningful. The method also had the numerically lowest mean RMSE for muscle quality. For force, biomarker-only Ridge performed best, suggesting a more linear endpoint structure.

URL PDF HTML ☆

赞 0 踩 0

2604.16689 2026-06-12 cs.AI 版本更新

The Query Channel: Information-Theoretic Limits of Masking-Based Explanations

查询通道：基于掩码的解释的信息论极限

Erciyes Karakaya, Ozgur Ercetin

发表机构 * Department of Electrical and Computer Engineering, University of Maryland, College Park, USA（美国马里兰大学电气与计算机工程系）； Faculty of Engineering and Natural Sciences, Sabanci University, Turkiye（土耳其萨班奇大学工程与自然科学学院）

AI总结本文提出查询通道框架，将掩码后解释建模为通信过程，推导解释率与识别容量之间的信息论极限，并证明稀疏最大似然解码器可实现可靠恢复。

详情

AI中文摘要

基于掩码的事后解释方法，如KernelSHAP和LIME，通过随机扰动下的查询估计局部特征重要性。本文将这一过程建模为在查询通道上的通信，其中潜在解释作为消息，每次掩码评估作为一次信道使用。在此框架内，解释的复杂度由假设类的熵捕获，而查询接口以每次查询的识别容量确定的速率提供信息。我们推导了一个强逆定理，表明如果解释率超过该容量，则对于任何解释器和解码器序列，精确恢复的概率必然收敛到误差中的一。我们还证明了一个可达性结果，即当速率低于容量时，稀疏最大似然解码器可实现可靠恢复。互信息的蒙特卡洛估计器提供了一个非渐近查询基准，我们用它来比较最优解码与模拟LIME和KernelSHAP的基于Lasso和OLS的过程。实验揭示了在一定的查询预算范围内，信息论允许可靠解释，但标准凸替代方法仍然失败。最后，我们将神经语言模型的超像素分辨率和分词解释为一种源编码选择，它设定了解释的熵，并展示了高斯噪声和非线性曲率如何劣化查询通道，引发瀑布和错误平层行为，并使高分辨率解释无法实现。

英文摘要

Masking-based post-hoc explanation methods, such as KernelSHAP and LIME, estimate local feature importance by querying a black-box model under randomized perturbations. This paper formulates this procedure as communication over a query channel, where the latent explanation acts as a message and each masked evaluation is a channel use. Within this framework, the complexity of the explanation is captured by the entropy of the hypothesis class, while the query interface supplies information at a rate determined by an identification capacity per query. We derive a strong converse showing that, if the explanation rate exceeds this capacity, the probability of exact recovery necessarily converges to one in error for any sequence of explainers and decoders. We also prove an achievability result establishing that a sparse maximum-likelihood decoder attains reliable recovery when the rate lies below capacity. A Monte Carlo estimator of mutual information yields a non-asymptotic query benchmark that we use to compare optimal decoding with Lasso- and OLS-based procedures that mirror LIME and KernelSHAP. Experiments reveal a range of query budgets where information theory permits reliable explanations but standard convex surrogates still fail. Finally, we interpret super-pixel resolution and tokenization for neural language models as a source-coding choice that sets the entropy of the explanation and show how Gaussian noise and nonlinear curvature degrade the query channel, induce waterfall and error-floor behavior, and render high-resolution explanations unattainable.

URL PDF HTML ☆

赞 0 踩 0

2604.13924 2026-06-12 cs.LG cs.AI cs.CV 版本更新

ASTER: Latent Pseudo-Anomaly Generation for Unsupervised Time-Series Anomaly Detection

ASTER: 用于无监督时间序列异常检测的潜在伪异常生成

Romain Hermary, Samet Hicsonmez, Dan Pineau, Abd El Rahman Shabayek, Djamila Aouada

发表机构 * University of Montreal（蒙特利尔大学）； Université de Montréal（蒙特利尔大学）

AI总结提出ASTER框架，在潜在空间生成伪异常训练Transformer分类器，结合预训练LLM增强表示，在三个基准数据集上达到最优性能。

Comments Published in ICPR 2026

详情

AI中文摘要

时间序列异常检测（TSAD）在工业监控、医疗保健和网络安全等领域至关重要，但由于罕见且异质的异常以及标记数据的稀缺性，它仍然具有挑战性。这种稀缺性使得无监督方法占主导地位，但现有方法通常依赖于重建或预测（难以处理复杂数据），或依赖于需要领域特定异常合成和固定距离度量的基于嵌入的方法。我们提出ASTER，一个直接在潜在空间中生成伪异常的框架，避免了手工制作的异常注入和对领域专业知识的需求。潜在空间解码器生成定制的伪异常，用于训练基于Transformer的异常分类器，而预训练的LLM丰富了该空间的时间和上下文表示。在三个基准数据集上的实验表明，ASTER达到了最先进的性能，并为基于LLM的TSAD设立了新标准。

英文摘要

Time-series anomaly detection (TSAD) is critical in domains such as industrial monitoring, healthcare, and cybersecurity, but it remains challenging due to rare and heterogeneous anomalies and the scarcity of labelled data. This scarcity makes unsupervised approaches predominant, yet existing methods often rely on reconstruction or forecasting, which struggle with complex data, or on embedding-based approaches that require domain-specific anomaly synthesis and fixed distance metrics. We propose ASTER, a framework that generates pseudo-anomalies directly in the latent space, avoiding handcrafted anomaly injections and the need for domain expertise. A latent-space decoder produces tailored pseudo-anomalies to train a Transformer-based anomaly classifier, while a pre-trained LLM enriches the temporal and contextual representations of this space. Experiments on three benchmark datasets show that ASTER achieves state-of-the-art performance and sets a new standard for LLM-based TSAD.

URL PDF HTML ☆

赞 0 踩 0

2604.08958 2026-06-12 cs.LG cs.AI cs.RO 版本更新

WOMBET: World Model-Based Experience Transfer for Robust and Sample-efficient Reinforcement Learning

WOMBET：基于世界模型的经验迁移实现鲁棒且样本高效的强化学习

Mintae Kim, Koushil Sreenath

发表机构 * Hybrid Robotics, UC Berkeley（混合机器人技术，伯克利大学）

AI总结提出WOMBET框架，通过源任务中学习世界模型并生成不确定性惩罚的离线数据，再结合自适应采样进行在线微调，实现鲁棒且样本高效的强化学习迁移。

Comments 13 pages, 6 figures, 8th Annual Learning for Dynamics & Control Conference (L4DC)

详情

AI中文摘要

机器人领域的强化学习通常受限于数据收集的成本和风险，因此需要从源任务向目标任务进行经验迁移。离线到在线强化学习利用先验数据，但通常假设给定固定数据集，并未解决如何生成可靠数据进行迁移的问题。我们提出基于世界模型的经验迁移（WOMBET）框架，该框架联合生成和利用先验数据。WOMBET在源任务中学习世界模型，并通过不确定性惩罚规划生成离线数据，随后筛选出高回报和低认知不确定性的轨迹。然后，它通过在离线数据和在线数据之间进行自适应采样，在目标任务中进行在线微调，实现了从先验驱动的初始化到任务特定适应的稳定过渡。我们证明了不确定性惩罚目标提供了真实回报的下界，并推导了有限样本误差分解，捕捉了分布不匹配和近似误差。实验上，WOMBET在连续控制基准测试中相比强基线提高了样本效率和最终性能，展示了联合优化数据生成和迁移的益处。

英文摘要

Reinforcement learning (RL) in robotics is often limited by the cost and risk of data collection, motivating experience transfer from a source task to a target task. Offline-to-online RL leverages prior data but typically assumes a given fixed dataset and does not address how to generate reliable data for transfer. We propose World Model-Based Experience Transfer (WOMBET), a framework that jointly generates and utilizes prior data. WOMBET learns a world model in the source task and generates offline data via uncertainty-penalized planning, followed by filtering trajectories with high return and low epistemic uncertainty. It then performs online fine-tuning in the target task using adaptive sampling between offline and online data, enabling a stable transition from prior-driven initialization to task-specific adaptation. We show that the uncertainty-penalized objective provides a lower bound on the true return and derive a finite-sample error decomposition capturing distribution mismatch and approximation error. Empirically, WOMBET improves sample efficiency and final performance over strong baselines on continuous control benchmarks, demonstrating the benefit of jointly optimizing data generation and transfer.

URL PDF HTML ☆

赞 0 踩 0

2604.12497 2026-06-12 cs.LG stat.ML 版本更新

Allocating Human Oversight in AI-Enabled Analytics

AI赋能分析中的人类监督分配

Zikun Ye, Jiameng Lyu, Rui Tao

发表机构 * Michael G. Foster School of Business, University of Washington（华盛顿大学迈克尔·G·福斯特商学院）； Department of Management Science, School of Management, Fudan University（复旦大学管理学院管理科学系）； Guanghua School of Management, Peking University（北京大学光华管理学院）

AI总结针对AI预测可靠性异质且未知的问题，提出基于上置信界的在线学习策略，动态分配有限的人类验证预算，使终端效率损失随预算增长趋于零。

详情

AI中文摘要

组织越来越多地部署AI作为面向客户的决策过程中的低成本预测层，包括需求感知、服务质量监控、产品测试和市场研究，但AI生成的信号在不同任务、产品和客户细分中的可靠性并不均匀。因此，企业仍然需要稀缺的人类验证（标签、审计、调查回复或后续测量）来将AI输出锚定到真实情况。由于人类真实情况本身存在噪声，在不同标注者之间甚至重复判断中都有所变化，企业必须为每个任务收集并平均多个人类标签，这使得人类验证成本高昂。我们研究如何在可靠性异质且在部署前未知的情况下，将有限的人类验证预算分配到多个AI辅助任务中。我们将其置于调优的预测驱动推断框架内。每个人类标签既提高了AI辅助估计的精度，也揭示了任务的修正难度，即在使用AI预测作为控制变量后剩余的方差。如果难度已知，最优分配将遵循Neyman平方根规则；由于未知，我们提出一种基于上置信界的策略，该策略在线学习难度并将验证导向AI最不可靠的任务。我们证明，随着预算增长，该策略相对于最优分配的终端效率损失趋于零。在合成实验和一个包含68个任务和超过2000名受访者的真实数字孪生调查中，当可靠性异质时，该策略缩小了与最优分配的大部分差距，优于均匀分配和epsilon-贪婪分配；在调查数据上，它还优于先探索后提交的试点设计，并将均匀分配的10-12%差距缩小到2-6%。AI的价值不仅取决于模型准确性，还取决于将人类监督定向到AI错误影响最大的操作策略。

英文摘要

Organizations increasingly deploy AI as a low-cost prediction layer in customer-facing decision processes, including demand sensing, service-quality monitoring, product testing, and market research, but AI-generated signals are unevenly reliable across tasks, products, and customer segments. Firms therefore still need scarce human validation (labels, audits, survey responses, or follow-up measurements) to anchor AI outputs to ground truth. Because human ground truth is itself noisy, varying across labelers and even across repeated judgments, the firm must collect and average several human labels per task, which makes human validation costly. We study how to allocate a limited human-validation budget across many AI-assisted tasks when reliability is heterogeneous and unknown before deployment. We cast this within tuned prediction-powered inference. Each human label both sharpens the AI-assisted estimate and reveals the task's rectification difficulty, the variance that remains after the AI prediction is optimally used as a control variate. If difficulties were known, the optimal allocation would follow a Neyman square-root rule; because they are unknown, we propose a policy based on upper confidence bounds that learns them online and steers validation toward tasks where AI is least reliable. We prove that the policy's terminal efficiency loss relative to the oracle allocation vanishes as the budget grows. In synthetic experiments and a real digital-twin survey with 68 tasks and over 2000 respondents, it closes most of the gap to the oracle when reliability is heterogeneous, outperforming uniform and epsilon-greedy allocation; on the survey data it also outperforms explore-then-commit pilot designs and cuts uniform's 10--12% gap to 2--6%. The value of AI depends not only on model accuracy but also on the operational policy that targets human oversight where AI errors matter most.

URL PDF HTML ☆

赞 0 踩 0

2604.12002 2026-06-12 cs.CL 版本更新

Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

自蒸馏零：自我修订将二元奖励转化为密集监督

Yinghui He, Simran Kaur, Adithya Bhaskar, Yongjin Yang, Jiarui Liu, Narutatsu Ri, Liam Fowl, Abhishek Panigrahi, Danqi Chen, Sanjeev Arora

发表机构 * Princeton University（普林斯顿大学）； University of Toronto（多伦多大学）； Carnegie Mellon University（卡内基梅隆大学）

AI总结提出SD-Zero方法，通过让模型同时扮演生成器和修订者，利用二元奖励生成密集的token级自监督信号，显著提升训练样本效率，在数学和代码推理任务上超越RFT、GRPO等基线。

详情

AI中文摘要

当前在可验证设置下的后训练方法分为两类。强化学习（RLVR）依赖二元奖励，虽然广泛适用且强大，但在训练过程中仅提供稀疏监督。蒸馏提供密集的token级监督，通常从外部教师或使用高质量示范中获得。收集此类监督成本高昂或不可用。我们提出自蒸馏零（SD-Zero），一种比RL更高效利用训练样本的方法，且不需要外部教师或高质量示范。SD-Zero训练单个模型扮演两个角色：生成器，产生初始响应；修订者，基于该响应及其二元奖励生成改进的响应。然后我们进行在线自蒸馏，将修订者蒸馏到生成器中，使用修订者以生成器的响应及其奖励为条件的token分布作为监督。实际上，SD-Zero训练模型将二元奖励转化为密集的token级自监督。在数学和代码推理基准上，使用Qwen3-4B-Instruct和Olmo-3-7B-Instruct，SD-Zero相比基础模型性能提升至少10%，并在相同问题集和训练样本预算下优于强基线，包括拒绝微调（RFT）、GRPO和自蒸馏微调（SDFT）。大量消融实验显示了所提出算法的两个新特性：（a）token级自定位，其中修订者能够基于奖励识别生成器响应中需要修订的关键token；（b）迭代自进化，其中改进答案的修订能力可以通过定期教师同步蒸馏回生成性能。代码：此https URL。

英文摘要

Current post-training methods in verifiable settings fall into two categories. Reinforcement learning (RLVR) relies on binary rewards, which are broadly applicable and powerful, but provide only sparse supervision during training. Distillation provides dense token-level supervision, typically obtained from an external teacher or using high-quality demonstrations. Collecting such supervision can be costly or unavailable. We propose Self-Distillation Zero (SD-Zero), a method that is substantially more training sample-efficient than RL and does not require an external teacher or high-quality demonstrations. SD-Zero trains a single model to play two roles: a Generator, which produces an initial response, and a Reviser, which conditions on that response and its binary reward to produce an improved response. We then perform on-policy self-distillation to distill the reviser into the generator, using the reviser's token distributions conditioned on the generator's response and its reward as supervision. In effect, SD-Zero trains the model to transform binary rewards into dense token-level self-supervision. On math and code reasoning benchmarks with Qwen3-4B-Instruct and Olmo-3-7B-Instruct, SD-Zero improves performance by at least 10% over the base models and outperforms strong baselines, including Rejection Fine-Tuning (RFT), GRPO, and Self-Distillation Fine-Tuning (SDFT), under the same question set and training sample budget. Extensive ablation studies show two novel characteristics of our proposed algorithm: (a) token-level self-localization, where the reviser can identify the key tokens that need to be revised in the generator's response based on reward, and (b) iterative self-evolution, where the improving ability to revise answers can be distilled back into generation performance with regular teacher synchronization. Code: https://github.com/princeton-pli/Self-Distillation-Zero.

URL PDF HTML ☆

赞 0 踩 0

2604.10389 2026-06-12 cs.CL 版本更新

BLUEmed: Retrieval-Augmented Multi-Agent Debate for Clinical Error Detection

BLUEmed: 基于检索增强的多智能体辩论用于临床错误检测

Saukun Thika You, Nguyen Anh Khoa Tran, Wesley K. Marizane, Hanshu Rao, Qiunan Zhang, Xiaolei Huang

发表机构 * University of California, San Diego（加州大学圣地亚哥分校）

AI总结提出BLUEmed框架，结合混合检索增强生成与多智能体辩论，通过分解临床笔记、检索证据、专家辩论及安全层过滤，在术语替换错误检测中达到最优性能。

Comments Accepted to the IEEE International Conference on Healthcare Informatics (ICHI) 2026

详情

AI中文摘要

临床笔记中的术语替换错误（即一个医学术语被一个语言上有效但临床不同的术语替换）对医疗保健中的自动错误检测构成了持续挑战。我们引入了BLUEmed，一个多智能体辩论框架，增强有混合检索增强生成（RAG），该框架结合了基于证据的推理和多视角验证用于临床错误检测。BLUEmed将每个临床笔记分解为聚焦的子查询，通过密集、稀疏和在线检索检索来源分区的证据，并分配两个具有不同知识库的领域专家智能体以产生独立分析；当专家意见不一致时，一轮结构化的反论证和跨来源裁决解决冲突，随后是一个级联安全层，过滤常见的假阳性模式。我们在一个临床术语替换检测基准上评估BLUEmed，在零样本和少样本提示下，使用多个骨干模型（涵盖专有和开源系列）。实验结果表明，在少样本提示下，BLUEmed达到了最佳准确率（69.13%）、ROC-AUC（74.45%）和PR-AUC（72.44%），优于单智能体RAG和仅辩论基线。跨六个骨干模型和两种提示策略的进一步分析证实，检索增强和结构化辩论是互补的，并且该框架从具有足够指令遵循和临床语言理解的模型中受益最大。

英文摘要

Terminology substitution errors in clinical notes, where one medical term is replaced by a linguistically valid but clinically different term, pose a persistent challenge for automated error detection in healthcare. We introduce BLUEmed, a multi-agent debate framework augmented with hybrid Retrieval-Augmented Generation (RAG) that combines evidence-grounded reasoning with multi-perspective verification for clinical error detection. BLUEmed decomposes each clinical note into focused sub-queries, retrieves source-partitioned evidence through dense, sparse, and online retrieval, and assigns two domain expert agents distinct knowledge bases to produce independent analyses; when the experts disagree, a structured counter-argumentation round and cross-source adjudication resolve the conflict, followed by a cascading safety layer that filters common false-positive patterns. We evaluate BLUEmed on a clinical terminology substitution detection benchmark under both zero-shot and few-shot prompting with multiple backbone models spanning proprietary and open-source families. Experimental results show that BLUEmed achieves the best accuracy (69.13%), ROC-AUC (74.45%), and PR-AUC (72.44%) under few-shot prompting, outperforming both single-agent RAG and debate-only baselines. Further analyses across six backbone models and two prompting strategies confirm that retrieval augmentation and structured debate are complementary, and that the framework benefits most from models with sufficient instruction-following and clinical language understanding.

URL PDF HTML ☆

赞 0 踩 0

2511.18322 2026-06-12 cs.RO cs.CV cs.LG 版本更新

Learning Visually Interpretable Oscillator Networks for Soft Continuum Robots from Video

从视频中学习软体连续体机器人的视觉可解释振荡器网络

Henrik Krauss, Johann Licher, Naoya Takeishi, Annika Raatz, Takehisa Yairi

发表机构 * Department of Advanced Interdisciplinary Studies, The University of Tokyo（东京大学先进跨学科研究系）； Institute of Assembly Technology and Robotics, Leibniz University Hannover（莱比锡大学汉诺威装配技术与机器人研究所）； Research Center for Advanced Science and Technology, The University of Tokyo（东京大学先进科学研究中心）

AI总结提出注意力广播解码器（ABCD）和视觉振荡器网络（VONs），实现从视频中学习软体连续体机器人动力学的视觉和机械可解释性，多步预测误差降低5.8倍。

Comments Code available at: https://github.com/UThenrik/visual_oscillators_for_SCR Dataset available at: https://zenodo.org/records/17812071 Video available at: https://youtu.be/i80H8erVISM

详情

AI中文摘要

从视频中学习软体连续体机器人（SCR）动力学提供了灵活性，但现有方法缺乏可解释性或依赖先验假设。基于模型的方法需要先验知识和手动设计。我们通过引入以下内容来弥补这一差距：（1）注意力广播解码器（ABCD），一种用于基于自编码器的潜在动力学学习的即插即用模块，生成像素级注意力图，定位每个潜在维度的贡献，同时过滤静态背景，通过空间接地潜在变量和图像叠加实现视觉可解释性。（2）视觉振荡器网络（VONs），一种二维潜在振荡器网络，与ABCD注意力图耦合，用于学习到的质量、耦合刚度和力的图像可视化，从而实现机械可解释性。我们在单段和双段SCR上验证了我们的方法，表明基于ABCD的模型显著提高了多步预测精度，在双段机器人上，Koopman算子的误差降低了5.8倍，振荡器网络的误差降低了3.5倍。VONs自主发现了振荡器的链式结构。这种完全数据驱动的方法产生了紧凑、机械可解释的模型，对未来的控制应用具有潜在意义。

英文摘要

Learning soft continuum robot (SCR) dynamics from video offers flexibility but existing methods lack interpretability or rely on prior assumptions. Model-based approaches require prior knowledge and manual design. We bridge this gap by introducing: (1) The Attention Broadcast Decoder (ABCD), a plug-and-play module for autoencoder-based latent dynamics learning that generates pixel-accurate attention maps localizing each latent dimension's contribution while filtering static backgrounds, enabling visual interpretability via spatially grounded latents and on-image overlays. (2) Visual Oscillator Networks (VONs), a 2D latent oscillator network coupled to ABCD attention maps for on-image visualization of learned masses, coupling stiffness, and forces, thereby enabling mechanical interpretability. We validate our approach on single- and double-segment SCRs, demonstrating that ABCD-based models significantly improve multi-step prediction accuracy with 5.8x error reduction for Koopman operators and 3.5x for oscillator networks on a two-segment robot. VONs autonomously discover a chain structure of oscillators. This fully data-driven approach yields compact, mechanically interpretable models with potential relevance for future control applications.

URL PDF HTML ☆

赞 0 踩 0

2512.14937 2026-06-12 cs.CV cs.AI 版本更新

Improving Pre-trained Adult Glioma Segmentation Models Using only Post-processing Techniques

仅使用后处理技术改进预训练的成人胶质瘤分割模型

Abhijeet Parida, Daniel Capellán-Martín, Zhifan Jiang, Nishad Kulkarni, Krithika Iyer, Austin Tapp, Syed Muhammad Anwar, María J. Ledesma-Carbayo, Marius George Linguraru

发表机构 * Sheikh Zayed Institute for Pediatric Surgical Innovation（Sheikh Zayed儿童手术创新研究所）； Children’s National Hospital（儿童医院）； University of Madrid（马德里大学）； CIBER-BBN ； ISCIII ； School of Medicine and Health Sciences（医学与健康科学学院）； George Washington University（乔治·华盛顿大学）

AI总结针对预训练模型在胶质瘤分割中的系统误差，提出自适应后处理技术，在BraTS 2025挑战中使排名指标提升14.9%（撒哈拉以南非洲）和0.9%（成人胶质瘤），推动向高效、公平、可持续的后处理策略转变。

详情

DOI: 10.1007/978-3-032-16365-3_22

AI中文摘要

胶质瘤是成人中最常见的恶性脑肿瘤，也是最致命的肿瘤之一。尽管积极治疗，中位生存率仍低于15个月。准确的多参数MRI（mpMRI）肿瘤分割对于手术规划、放疗和疾病监测至关重要。虽然深度学习模型提高了自动分割的准确性，但大规模预训练模型泛化能力差且常表现不佳，产生系统性错误，如假阳性、标签交换和切片不连续。这些问题因GPU资源获取不平等和大规模模型训练日益增长的环境成本而进一步加剧。在这项工作中，我们提出自适应后处理技术，以改进为各种肿瘤类型开发的大规模预训练模型产生的胶质瘤分割质量。我们在多个BraTS 2025分割挑战任务中展示了这些技术，使撒哈拉以南非洲挑战的排名指标提升了14.9%，成人胶质瘤挑战提升了0.9%。该方法推动脑肿瘤分割研究从日益复杂的模型架构转向精确、计算公平且可持续的高效临床后处理策略。

英文摘要

Gliomas are the most common malignant brain tumors in adults and are among the most lethal. Despite aggressive treatment, the median survival rate is less than 15 months. Accurate multiparametric MRI (mpMRI) tumor segmentation is critical for surgical planning, radiotherapy, and disease monitoring. While deep learning models have improved the accuracy of automated segmentation, large-scale pre-trained models generalize poorly and often underperform, producing systematic errors such as false positives, label swaps, and slice discontinuities in slices. These limitations are further compounded by unequal access to GPU resources and the growing environmental cost of large-scale model training. In this work, we propose adaptive post-processing techniques to refine the quality of glioma segmentations produced by large-scale pretrained models developed for various types of tumors. We demonstrated the techniques in multiple BraTS 2025 segmentation challenge tasks, with the ranking metric improving by 14.9 % for the sub-Saharan Africa challenge and 0.9% for the adult glioma challenge. This approach promotes a shift in brain tumor segmentation research from increasingly complex model architectures to efficient, clinically aligned post-processing strategies that are precise, computationally fair, and sustainable.

URL PDF HTML ☆

赞 0 踩 0

2512.14648 2026-06-12 cs.CV eess.IV 版本更新

Adaptable Segmentation Pipeline for Diverse Brain Tumors with Radiomic-Guided Subtyping and Lesion-Wise Model Ensemble

适用于多样化脑肿瘤的自适应分割流程：放射组学引导的亚型分类与病灶级模型集成

Daniel Capellán-Martín, Abhijeet Parida, Zhifan Jiang, Nishad Kulkarni, Krithika Iyer, Austin Tapp, Syed Muhammad Anwar, María J. Ledesma-Carbayo, Marius George Linguraru

发表机构 * Sheikh Zayed Institute for Pediatric Surgical Innovation（Sheikh Zayed儿童外科创新研究所）； Children’s National Hospital（儿童医院）； University of Washington（华盛顿大学）； Universidad Politécnica de Madrid（马德里理工大学）； CIBER-BBN ； ISCIII ； School of Medicine and Health Sciences（医学与健康科学学院）

AI总结提出一种灵活模块化的自适应分割流程，通过放射组学特征检测肿瘤亚型并平衡训练，结合病灶级性能指标优化模型集成与后处理，在BraTS 2025挑战赛中达到顶尖性能，支持临床定量肿瘤测量。

Comments 12 pages, 5 figures, 3 tables. Algorithm presented at MICCAI BraTS 2025

详情

DOI: 10.1007/978-3-032-16365-3_41

AI中文摘要

在多参数磁共振成像（MRI）上对脑肿瘤进行鲁棒且可泛化的分割仍然困难，因为肿瘤类型差异很大。BraTS 2025 Lighthouse挑战赛在多种高质量成人及儿童肿瘤数据集上对分割方法进行基准测试：多联盟国际儿童脑肿瘤分割（PED）、术前脑膜瘤肿瘤分割（MEN）、脑膜瘤放射治疗分割（MEN-RT）以及治疗前后脑转移瘤分割（MET）。我们提出了一种灵活、模块化且自适应的流程，通过选择和组合最先进的模型，并在训练前后应用肿瘤和病灶特定的处理，来提高分割性能。从MRI中提取的放射组学特征有助于检测肿瘤亚型，确保更平衡的训练。自定义的病灶级性能指标决定了每个模型在集成中的影响力，并优化了进一步细化预测的后处理，使工作流能够针对每个病例定制每一步。在BraTS测试集上，我们的流程在多个挑战中取得了与顶尖算法相当的性能。这些发现证实，自定义的病灶感知处理与模型选择能够产生鲁棒的分割，而无需将方法锁定在特定的网络架构上。我们的方法在临床实践中具有定量肿瘤测量的潜力，支持诊断和预后。

英文摘要

Robust and generalizable segmentation of brain tumors on multi-parametric magnetic resonance imaging (MRI) remains difficult because tumor types differ widely. The BraTS 2025 Lighthouse Challenge benchmarks segmentation methods on diverse high-quality datasets of adult and pediatric tumors: multi-consortium international pediatric brain tumor segmentation (PED), preoperative meningioma tumor segmentation (MEN), meningioma radiotherapy segmentation (MEN-RT), and segmentation of pre- and post-treatment brain metastases (MET). We present a flexible, modular, and adaptable pipeline that improves segmentation performance by selecting and combining state-of-the-art models and applying tumor- and lesion-specific processing before and after training. Radiomic features extracted from MRI help detect tumor subtype, ensuring a more balanced training. Custom lesion-level performance metrics determine the influence of each model in the ensemble and optimize post-processing that further refines the predictions, enabling the workflow to tailor every step to each case. On the BraTS testing sets, our pipeline achieved performance comparable to top-ranked algorithms across multiple challenges. These findings confirm that custom lesion-aware processing and model selection yield robust segmentations yet without locking the method to a specific network architecture. Our method has the potential for quantitative tumor measurement in clinical practice, supporting diagnosis and prognosis.

URL PDF HTML ☆

赞 0 踩 0

2506.18438 2026-06-12 cs.CV 版本更新

CPAM: Context-Preserving Adaptive Manipulation for Zero-Shot Real Image Editing

CPAM: 保持上下文的自适应操作用于零样本真实图像编辑

Dinh-Khoi Vo, Thanh-Toan Do, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le

发表机构 * Faculty of Information Technology, University of Science, Ho Chi Minh City, Vietnam（越南科学大学信息科技学院）； Vietnam National University, Ho Chi Minh City, Vietnam（越南国家大学）； Faculty of Information Technology, Monash University, Melbourne, Victoria, Australia（莫纳什大学信息科技学院）； Department of Computer Science, University of Dayton, Dayton, Ohio, US（Dayton 大学计算机科学系）

AI总结提出CPAM零样本框架，通过保持上下文的自适应操作和掩码引导，实现复杂非刚性真实图像的编辑，保留纹理和身份，无需微调。

Comments Accepted to IEEE Transactions on Multimedia. Project page: https://vdkhoi20.github.io/CPAM

详情

AI中文摘要

使用文本描述在文本到图像扩散模型中编辑自然图像仍然是一个重大挑战，特别是在实现一致生成和处理复杂非刚性对象方面。现有方法通常难以保留纹理和身份，需要大量微调，并且在编辑特定空间区域或对象的同时保留背景细节方面存在局限性。本文提出了保持上下文的自适应操作（CPAM），一种用于复杂非刚性真实图像编辑的新型零样本框架。具体来说，我们提出了一个保留适应模块，该模块调整自注意力机制以有效保留并独立控制对象和背景。这确保了在编辑过程中使用掩码引导技术时，对象的形状、纹理和身份得以保持，同时背景不变形。此外，我们开发了一个局部提取模块，以减轻在交叉注意力机制的条件化过程中对非期望修改区域的干扰。我们还引入了各种掩码引导策略，以简单的方式促进多样化的图像操作任务。CPAM可以无缝集成到多个扩散骨干网络中，包括SD1.5、SD2.1和SDXL，展示了跨不同模型架构的强大泛化能力。在我们新构建的图像操作基准（IMBA）上进行的广泛实验表明，我们提出的方法是人类评估者的首选，优于现有的最先进编辑技术。源代码和数据将在项目页面公开发布：this https URL

英文摘要

Editing natural images using textual descriptions in text-to-image diffusion models remains a significant challenge, particularly in achieving consistent generation and handling complex, non-rigid objects. Existing methods often struggle to preserve textures and identity, require extensive fine-tuning, and exhibit limitations in editing specific spatial regions or objects while retaining background details. This paper proposes Context-Preserving Adaptive Manipulation (CPAM), a novel zero-shot framework for complicated, non-rigid real image editing. Specifically, we propose a preservation adaptation module that adjusts self-attention mechanisms to preserve and independently control the object and background effectively. This ensures that the objects' shapes, textures, and identities are maintained while keeping the background undistorted during the editing process using the mask guidance technique. Additionally, we develop a localized extraction module to mitigate the interference with the non-desired modified regions during conditioning in cross-attention mechanisms. We also introduce various mask-guidance strategies to facilitate diverse image manipulation tasks in a simple manner. CPAM can be seamlessly integrated with multiple diffusion backbones, including SD1.5, SD2.1, and SDXL, demonstrating strong generalization across different model architectures. Extensive experiments on our newly constructed Image Manipulation BenchmArk (IMBA), a robust benchmark dataset specifically designed for real image editing, demonstrate that our proposed method is the preferred choice among human raters, outperforming existing state-of-the-art editing techniques. The source code and data will be publicly released at the project page: https://vdkhoi20.github.io/CPAM

URL PDF HTML ☆

赞 0 踩 0

2604.08983 2026-06-12 cs.RO 版本更新

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly

AssemLM: 用于机器人装配的空间推理多模态大语言模型

Zhi Jing, Jinbin Qiao, Ouyang Lu, Jicong Ao, Shuang Qiu, Huazhe Xu, Yu-Gang Jiang, Chenjia Bai

发表机构 * Fudan University（复旦大学）； Institute of Artificial Intelligence (TeleAI), China Telecom（人工智能研究所（TeleAI），中国电信）； Tianjin University（天津大学）； Northwestern Polytechnical University（西北工业大学）； Tsinghua University（清华大学）； City University of Hong Kong（香港城市大学）

AI总结提出AssemLM，一种融合装配手册、点云和文本指令的多模态大语言模型，通过专用点云编码器提取几何与旋转特征，实现精确的6D装配位姿推理，并构建含90万样本的AssemBench基准，在真实机器人装配任务中取得最优性能。

Comments Project Page: https://assemlmhome.github.io/

详情

AI中文摘要

空间推理是具身智能的基本能力，尤其对于机器人装配等精细操作任务。当前基于视觉语言模型（VLM）的方法主要依赖粗粒度的2D感知，难以对复杂3D几何进行精确推理。为解决这一局限，我们提出AssemLM，一种用于机器人装配的空间多模态大语言模型，它整合装配手册、点云和文本指令，通过显式几何理解预测任务关键的6D装配位姿。为桥接原始3D感知与高层语言推理，AssemLM采用专用点云编码器提取细粒度几何与旋转特征，以实现装配任务中精确的3D空间推理。此外，我们引入AssemBench，一个面向装配空间推理的大规模基准，包含超过90万多模态样本和精确的6D位姿标注，将评估从2D定位扩展到完整的3D几何推理。大量实验和真实机器人评估表明，AssemLM在6D位姿推理性能上达到最优，并有效支持真实环境中的精细多步装配任务。代码、模型和AssemBench数据集将公开提供。

英文摘要

Spatial reasoning is a fundamental capability for embodied intelligence, especially for fine-grained manipulation tasks such as robotic assembly. Recent methods based on vision-language models (VLMs) largely rely on coarse 2D perception and struggle to perform accurate reasoning over complex 3D geometry. To address this limitation, we propose AssemLM, a spatial multimodal large language model for robotic assembly that integrates assembly manuals, point clouds, and textual instructions to predict task-critical 6D assembly poses with explicit geometric understanding. To bridge raw 3D perception and high-level linguistic reasoning, AssemLM employs a specialized point cloud encoder to extract fine-grained geometric and rotational features for accurate 3D spatial reasoning in assembly tasks. In addition, we introduce AssemBench, a large-scale benchmark for assembly-oriented spatial reasoning with over 900K multimodal samples and precise 6D pose annotations, extending evaluation from 2D grounding to full 3D geometric inference. Extensive experiments and real-robot evaluations demonstrate that AssemLM achieves state-of-the-art 6D pose reasoning performance and effectively supports fine-grained, multi-step assembly tasks in real-world settings. Code, models, and the AssemBench dataset will be made publicly available.

URL PDF HTML ☆

赞 0 踩 0

2603.29515 2026-06-12 cs.LG 版本更新

Variational Graph Neural Networks for Uncertainty Quantification in Inverse Problems

变分图神经网络用于反问题中的不确定性量化

David Gonzalez, Alba Muixi, Beatriz Moya, Elias Cueto

发表机构 * Keysight-UZ Chair of the Spanish National Strategy on AI（西班牙人工智能国家战略主席席位）； Aragon Institute of Engineering Research (I3A)（阿拉贡工程研究所（I3A））； Universidad de Zaragoza（萨拉戈塔大学）； Laboratori de Càlcul Numèric (LaCàN)（数值计算实验室（LaCàN））； Universitat Politècnica de Catalunya - BarcelonaTech (UPC)（加泰罗尼亚理工大学 - 巴塞罗那科技大学（UPC））； Centre Internacional de Mètodes Numèrics en Enginyeria (CIMNE)（国际数值工程方法中心（CIMNE））； PIMM Lab. Arts et Métiers Institute of Technology（巴黎艺术与技术理工学院PIMM实验室）

AI总结提出变分图神经网络（VGNN），通过在解码器引入变分层以较低成本量化认知和统计不确定性，在固体力学反问题中验证了高精度参数恢复与置信区间估计。

详情

AI中文摘要

深度学习技术在计算力学中的日益广泛应用显著加速了那些几年前还被认为是难以处理的问题的模拟。然而，在诸如工程或医学数字孪生等关键应用中，快速响应是不够的；还必须提供可靠的结果。在某些情况下，传统的确定性方法可能不是最优的，因为它们无法提供对其预测或结果的置信度度量，尤其是在反问题中，解可能不唯一或初始数据由于噪声等原因不完全可靠。经典的深度神经网络也缺乏明确的度量来量化其预测的不确定性。在这项工作中，我们提出了一种变分图神经网络（VGNN）架构，该架构将变分层集成到其架构中以建模权重的概率分布。与计算昂贵的全贝叶斯网络不同，我们的方法仅在解码器中策略性地引入变分层，从而能够以相对较低的成本估计认知不确定性和统计不确定性。在这项工作中，我们在两个固体力学案例中验证了所提出的方法：在二维弹性问题中识别具有非线性分布的弹性模量值，以及在三维超弹性梁中定位和量化施加的载荷，在这两种情况下仅使用每个测试的位移场作为输入数据。结果表明，该模型不仅以高精度恢复了物理参数，还提供了与问题物理特性一致的置信区间，并且能够定位施加载荷的位置并估计其值，为该实验提供了置信区间。

英文摘要

The increasingly wide use of deep machine learning techniques in computational mechanics has significantly accelerated simulations of problems that were considered unapproachable just a few years ago. However, in critical applications such as Digital Twins for engineering or medicine, fast responses are not enough; reliable results must also be provided. In certain cases, traditional deterministic methods may not be optimal as they do not provide a measure of confidence in their predictions or results, especially in inverse problems where the solution may not be unique or the initial data may not be entirely reliable due to the presence of noise, for instance. Classic deep neural networks also lack a clear measure to quantify the uncertainty of their predictions. In this work, we present a variational graph neural network (VGNN) architecture that integrates variational layers into its architecture to model the probability distribution of weights. Unlike computationally expensive full Bayesian networks, our approach strategically introduces variational layers exclusively in the decoder, allowing us to estimate cognitive uncertainty and statistical uncertainty at a relatively lower cost. In this work, we validate the proposed methodology in two cases of solid mechanics: the identification of the value of the elastic modulus with nonlinear distribution in a 2D elastic problem and the location and quantification of the loads applied to a 3D hyperelastic beam, in both cases using only the displacement field of each test as input data. The results show that the model not only recovers the physical parameters with high precision, but also provides confidence intervals consistent with the physics of the problem, as well as being able to locate the position of the applied load and estimate its value, giving a confidence interval for that experiment.

URL PDF HTML ☆

赞 0 踩 0

2601.06572 2026-06-12 cs.LG cs.AI 版本更新

Hellinger Multimodal Variational Autoencoders

Hellinger多模态变分自编码器

Huyen Vo, Isabel Valera

发表机构 * Department of Computer Science, Saarland University（萨尔兰大学计算机科学系）； MPI-SWS, Saarland Informatics Campus（萨尔兰信息学校区Max Planck研究所）

AI总结提出基于Hellinger距离的矩匹配近似方法HELVAE，避免子采样，在多模态变分自编码器中实现更优的生成一致性与质量权衡。

Comments Accepted at AISTATS 2026. Camera-ready version

详情

AI中文摘要

多模态变分自编码器（VAEs）广泛用于弱监督生成学习，涉及多种模态。主流方法通过专家乘积（PoE）、专家混合（MoE）或其组合来聚合单模态推理分布，以近似联合后验。本文从概率意见池化的优化视角重新审视多模态推理。我们从$\alpha=0.5$的Hölder池化出发，这是$\alpha\text{-散度}$族中唯一的对称成员，并推导出一种矩匹配近似，称为Hellinger。我们利用这种近似提出HELVAE，一种避免子采样的多模态VAE，从而得到一个高效且有效的模型，该模型：（i）随着观察到的模态增加，学习更具表达力的潜在表示；（ii）在生成一致性和质量之间实现更好的权衡，优于最先进的多模态VAE模型。

英文摘要

Multimodal variational autoencoders (VAEs) are widely used for weakly supervised generative learning with multiple modalities. Predominant methods aggregate unimodal inference distributions using either a product of experts (PoE), a mixture of experts (MoE), or their combinations to approximate the joint posterior. In this work, we revisit multimodal inference through the lens of probabilistic opinion pooling, an optimization-based approach. We start from Hölder pooling with $α=0.5$, which corresponds to the unique symmetric member of the $α\text{-divergence}$ family, and derive a moment-matching approximation, termed Hellinger. We then leverage such an approximation to propose HELVAE, a multimodal VAE that avoids sub-sampling, yielding an efficient yet effective model that: (i) learns more expressive latent representations as additional modalities are observed; and (ii) empirically achieves better trade-offs between generative coherence and quality, outperforming state-of-the-art multimodal VAE models.

URL PDF HTML ☆

赞 0 踩 0

2603.25450 2026-06-12 cs.AI 版本更新

Cross-Model Disagreement as a Label-Free Correctness Signal

跨模型分歧作为无标签正确性信号

Matt Gorbett, Suman Jana

发表机构 * Independent Researcher（独立研究者）； Department of Computer Science Columbia University（计算机科学系哥伦比亚大学）

AI总结提出跨模型分歧作为无标签正确性指标，通过验证模型对生成模型答案的困惑度或熵来检测错误，无需训练或标签，在多个基准上优于模型内不确定性方法。

详情

AI中文摘要

在没有真实标签的情况下检测语言模型何时出错是安全部署的一个基本挑战。现有方法依赖于模型自身的不确定性——例如令牌熵或置信度分数——但这些信号在最危险的失败模式：自信错误（模型错误但确定）上会严重失效。在这项工作中，我们引入跨模型分歧作为正确性指标——一种简单、无需训练的信号，可以无需修改地插入现有的生产系统、路由管道和部署监控基础设施。给定模型生成的答案，跨模型分歧通过单次前向传递计算第二个验证模型在读取该答案时的惊讶或不确定性程度。不需要验证模型生成任何内容，也不需要正确性标签。我们将这一原则实例化为跨模型困惑度（CMP），它衡量验证模型对生成模型答案令牌的惊讶程度，以及跨模型熵（CME），它衡量验证模型在这些位置的不确定性。CMP和CME在涵盖推理、检索和数学问题求解（MMLU、TriviaQA和GSM8K）的基准测试中均优于模型内不确定性基线。在MMLU上，CMP的平均AUROC为0.75，而模型内熵基线为0.59。这些结果确立了跨模型分歧作为一种实用的、无需训练的无标签正确性估计方法，可直接应用于部署监控、模型路由、选择性预测、数据过滤和生产语言模型系统的可扩展监督。

英文摘要

Detecting when a language model is wrong without ground truth labels is a fundamental challenge for safe deployment. Existing approaches rely on a model's own uncertainty -- such as token entropy or confidence scores -- but these signals fail critically on the most dangerous failure mode: confident errors, where a model is wrong but certain. In this work we introduce cross-model disagreement as a correctness indicator -- a simple, training-free signal that can be dropped into existing production systems, routing pipelines, and deployment monitoring infrastructure without modification. Given a model's generated answer, cross-model disagreement computes how surprised or uncertain a second verifier model is when reading that answer via a single forward pass. No generation from the verifying model is required, and no correctness labels are needed. We instantiate this principle as Cross-Model Perplexity (CMP), which measures the verifying model's surprise at the generating model's answer tokens, and Cross-Model Entropy (CME), which measures the verifying model's uncertainty at those positions. Both CMP and CME outperform within-model uncertainty baselines across benchmarks spanning reasoning, retrieval, and mathematical problem solving (MMLU, TriviaQA, and GSM8K). On MMLU, CMP achieves a mean AUROC of 0.75 against a within-model entropy baseline of 0.59. These results establish cross-model disagreement as a practical, training-free approach to label-free correctness estimation, with direct applications in deployment monitoring, model routing, selective prediction, data filtering, and scalable oversight of production language model systems.

URL PDF HTML ☆

赞 0 踩 0

2603.21563 2026-06-12 cs.AI 版本更新

Counterfactual Credit Policy Optimization for Multi-Agent Collaboration

多智能体协作的反事实信用策略优化

Zhongyi Li, Wan Tian, Jinju Chen, Huiming Zhang, Yang Liu, Yikun Ban, Fuzhen Zhuang

发表机构 * Beihang University（北航）； Peking University（北京大学）； Beijing University of Posts and Telecommunications（北京邮电大学）

AI总结针对多智能体大语言模型协作中信用分配难题，提出CCPO框架，通过反事实信用估计和验证器锚定的自评估两种分配器，将团队奖励转化为个体学习信号，提升数学推理任务表现。

详情

AI中文摘要

协作式多智能体大语言模型可以通过分解角色来解决复杂的推理任务，但此类系统的强化学习受到信用分配的限制：共享的终端奖励模糊了个体贡献，并可能鼓励搭便车行为。我们引入了协作信用策略优化（CCPO），这是一个与优化器无关的信用分配层，将团队层面的结果转化为智能体特定的学习信号。CCPO提供了两种互补的分配器。反事实信用通过比较实际团队结果与移除该智能体的反事实结果来估计智能体的边际贡献。验证器锚定的LLM自我评估是一种探索性分配器，它使用受限的自我评估和同伴评估来重新分配信用，同时保持外部验证器结果的主导地位。由此产生的角色特定奖励可以被GRPO风格的更新或其他策略梯度优化器（如GSPO和REINFORCE++）使用。我们在顺序的思考-求解设置中实例化CCPO，并在数学推理基准上评估它。结果表明，显式的信用分配通常能改善双智能体推理，尤其是在MATH500和几个分布外设置中，而增益因模型和数据集而异。

英文摘要

Collaborative multi-agent large language models (LLMs) can solve complex reasoning tasks by decomposing roles, but reinforcement learning for such systems is limited by credit assignment: shared terminal rewards obscure individual contributions and can encourage free-riding. We introduce two optimizer-agnostic credit assignment methods for converting joint outcomes into agent-specific learning signals. Counterfactual Credit for Policy Optimization (CCPO) estimates an agent's marginal contribution by comparing the realized joint outcome with a counterfactual outcome where that agent is removed. Self-Evaluated Credit for Policy Optimization (SEPO) uses constrained self- and peer-evaluations as a verifier-anchored credit signal while keeping the external task outcome dominant. Both operate at the reward-construction layer rather than as policy optimizers, producing role-specific rewards or advantages for GRPO, GSPO, or REINFORCE++. We instantiate these credit signals in a sequential Think--Solve setting and evaluate them on mathematical reasoning benchmarks. Results show that explicit credit assignment often improves dual-agent reasoning, especially on MATH500 and several out-of-distribution settings, while gains vary across models and datasets. Our code is available at: https://github.com/bhai114/ccpo.

URL PDF HTML ☆

赞 0 踩 0

2603.16013 2026-06-12 cs.RO cs.SE 版本更新

Safety Case Patterns for VLA-based driving systems: Insights from SimLingo

基于VLA的驾驶系统的安全案例模式：来自SimLingo的见解

Gerhard Yu, Fuyuki Ishikawa, Oluwafemi Odu, Alvine Boaye Belle

发表机构 * York University（约克大学）； National Institute of Informatics（国家信息研究所）

AI总结针对VLA驾驶系统提出RAISE安全案例设计方法，通过扩展HARA和定制模式，结合SimLingo案例验证其构建基于证据的安全声明的有效性。

详情

AI中文摘要

基于视觉-语言-动作（VLA）的驾驶系统代表了自动驾驶领域的重大范式转变，因为通过结合交通场景理解、语言解释和动作生成，这些系统能够实现更灵活、自适应和响应指令的驾驶行为。然而，尽管它们被越来越多地采用，并具有支持社会责任型自动驾驶以及理解高级人类指令的潜力，基于VLA的驾驶系统可能表现出新型的危险行为。例如，将开放式的自然语言输入（如用户或导航指令）集成到多模态控制回路中可能导致不可预测和不安全的行为，从而危及车辆乘员和行人。因此，确保这些系统的安全性对于建立对其运行的信任至关重要。为此，我们提出了一种名为RAISE的新型安全案例设计方法。我们的方法引入了针对基于指令的驾驶系统（如VLA驾驶系统）定制的新模式，扩展了危害分析和风险评估（HARA），详细说明了安全场景及其结果，并设计了一种创建VLA驾驶系统安全案例的技术。在SimLingo上的案例研究说明了如何使用我们的方法为这类新兴的自动驾驶系统构建严谨的、基于证据的安全声明。

英文摘要

Vision-Language-Action (VLA)-based driving systems represent a significant paradigm shift in autonomous driving since, by combining traffic scene understanding, linguistic interpretation, and action generation, these systems enable more flexible, adaptive, and instruction-responsive driving behaviors. However, despite their growing adoption and potential to support socially responsible autonomous driving as well as understanding high-level human instructions, VLA-based driving systems may exhibit new types of hazardous behaviors. For instance, the integration of open-ended natural language inputs (e.g., user or navigation instructions) into the multimodal control loop may lead to unpredictable and unsafe behaviors that could endanger vehicle occupants and pedestrians. Hence, assuring the safety of these systems is crucial to help build trust in their operations. To support this, we propose a novel safety case design approach called RAISE. Our approach introduces novel patterns tailored to instruction-based driving systems such as VLA-based driving systems, an extension of Hazard Analysis and Risk Assessment (HARA) detailing safe scenarios and their outcomes, and a design technique to create the safety cases of VLA-based driving systems. A case study on SimLingo illustrates how our approach can be used to construct rigorous, evidence-based safety claims for this emerging class of autonomous driving systems.

URL PDF HTML ☆

赞 0 踩 0

2603.14482 2026-06-12 cs.CV 版本更新

V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

V-JEPA 2.1: 解锁视频自监督学习中的密集特征

Lorenzo Mur-Labadia, Matthew Muckley, Amir Bar, Mido Assran, Koustuv Sinha, Mike Rabbat, Yann LeCun, Nicolas Ballas, Adrien Bardes

发表机构 * FAIR at Meta（Meta的FAIR）； Universidad de Zaragoza（萨拉戈萨大学）

AI总结提出V-JEPA 2.1系列自监督模型，通过密集预测损失、深度自监督、多模态分词器和有效缩放，学习图像和视频的密集高质量视觉表示，在多个基准上取得最优性能。

详情

AI中文摘要

我们提出V-JEPA 2.1，一系列自监督模型，能够学习图像和视频的密集、高质量视觉表示，同时保持强大的全局场景理解。该方法结合了四个关键组件。首先，密集预测损失使用基于掩码的目标，其中可见和掩码令牌都贡献于训练信号，鼓励显式的空间和时间接地。其次，深度自监督在多个中间编码器层上分层应用自监督目标，以提高表示质量。第三，多模态分词器实现了图像和视频的统一训练。最后，该模型受益于模型容量和训练数据的有效缩放。这些设计选择共同产生了空间结构、语义一致和时间连贯的表示。实验上，V-JEPA 2.1在几个具有挑战性的基准上取得了最先进的性能，包括在Ego4D上短期物体交互预测的7.71 mAP，在EPIC-KITCHENS上高级动作预测的40.8 Recall@5，以及在实际机器人抓取成功率上比V-JEPA-2 AC提高了20个百分点。该模型还在机器人导航（TartanDrive上5.687 ATE）、深度估计（NYUv2上线性探针0.307 RMSE）和全局识别（Something-Something-V2上77.7）方面表现出强大的性能。这些结果表明，V-JEPA 2.1显著推进了密集视觉理解和世界建模的最新技术。

英文摘要

We present V-JEPA 2.1, a family of self-supervised models that learn dense, high-quality visual representations for both images and videos while retaining strong global scene understanding. The approach combines four key components. First, a dense predictive loss uses a masking-based objective in which both visible and masked tokens contribute to the training signal, encouraging explicit spatial and temporal grounding. Second, deep self-supervision applies the self-supervised objective hierarchically across multiple intermediate encoder layers to improve representation quality. Third, multi-modal tokenizers enable unified training across images and videos. Finally, the model benefits from effective scaling in both model capacity and training data. Together, these design choices produce representations that are spatially structured, semantically coherent, and temporally consistent. Empirically, V-JEPA 2.1 achieves state-of-the-art performance on several challenging benchmarks, including 7.71 mAP on Ego4D for short-term object-interaction anticipation and 40.8 Recall@5 on EPIC-KITCHENS for high-level action anticipation, as well as a 20-point improvement in real-robot grasping success rate over V-JEPA-2 AC. The model also demonstrates strong performance in robotic navigation (5.687 ATE on TartanDrive), depth estimation (0.307 RMSE on NYUv2 with a linear probe), and global recognition (77.7 on Something-Something-V2). These results show that V-JEPA 2.1 significantly advances the state of the art in dense visual understanding and world modeling.

URL PDF HTML ☆

赞 0 踩 0

AI 大模型

视觉与机器人

科学与医疗

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering

Possibilistic Predictive Uncertainty for Deep Learning

LLMs as ASP Programmers: Self-Correction Enables Task-Agnostic Nonmonotonic Reasoning

BrainDINO: A Brain MRI Foundation Model for Generalizable Clinical Representation Learning

Select to Think: Unlocking SLM Potential with Local Sufficiency

The Pragmatic Persona: Discovering LLM Persona through Bridging Inference

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models

BSViT: A Burst Spiking Vision Transformer for Expressive and Efficient Visual Representation Learning

ShowFlow: From Robust Single Concept to Condition-Free Multi-Concept Generation

Machine Learning-based Two-Stage Graph Sparsification for the Travelling Salesman Problem

Reasoning Models Know What's Important, and Encode It in Their Activations

Geometric and Quantum Kernel Methods for Predicting Skeletal Muscle Outcomes in chronic obstructive pulmonary disease

The Query Channel: Information-Theoretic Limits of Masking-Based Explanations

ASTER: Latent Pseudo-Anomaly Generation for Unsupervised Time-Series Anomaly Detection

WOMBET: World Model-Based Experience Transfer for Robust and Sample-efficient Reinforcement Learning

Allocating Human Oversight in AI-Enabled Analytics

Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision

BLUEmed: Retrieval-Augmented Multi-Agent Debate for Clinical Error Detection

Learning Visually Interpretable Oscillator Networks for Soft Continuum Robots from Video

Improving Pre-trained Adult Glioma Segmentation Models Using only Post-processing Techniques

Adaptable Segmentation Pipeline for Diverse Brain Tumors with Radiomic-Guided Subtyping and Lesion-Wise Model Ensemble

CPAM: Context-Preserving Adaptive Manipulation for Zero-Shot Real Image Editing

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly

Variational Graph Neural Networks for Uncertainty Quantification in Inverse Problems

Hellinger Multimodal Variational Autoencoders

Cross-Model Disagreement as a Label-Free Correctness Signal

Counterfactual Credit Policy Optimization for Multi-Agent Collaboration

Safety Case Patterns for VLA-based driving systems: Insights from SimLingo

V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning