arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Harvard University(哈佛大学)

至 收录 1302
2607.17521 2026-07-21 cs.RO 新提交

GeoWorldAD: Geometry World Action Model for Autonomous Driving

GeoWorldAD:用于自动驾驶的几何世界行动模型

Songyan Zhang, Jinyuan Tian, Hanbing Li, Daqi Liu, Hao Chen, Wenhui Huang, Fang Li, Guang Chen, Hangjun Ye, Long Chen, Kuiyuan Yang, Chen Lv

机构 * Nanyang Technological University(南洋理工大学) Xiaomi EV(小米汽车) Zhejiang University(浙江大学) Harvard University(哈佛大学)

AI总结 研究自动驾驶中安全高效规划决策问题,提出GeoWorldAD模型,通过在自我对齐3D空间中规划轨迹、用潜在未来几何标记预测场景演变,并逐步聚合多尺度几何线索,实验证明该模型在自动驾驶方面性能先进。

详情
AI中文摘要

自动驾驶需要在动态3D环境中做出安全且高效的规划决策。尽管近期的视觉/视频行动模型能直接从视觉观察中学习策略并随视觉Transformer和大规模训练数据发展良好,但常缺乏明确的几何基础和对未来的空间引导。本文提出GeoWorldAD,一种在自我对齐的3D空间中进行轨迹规划并通过潜在未来几何标记预测短视距场景演变的几何世界行动模型。当前几何为安全规划提供基本空间约束,未来几何揭示周围物体和以自我为中心的自由空间如何演变,减少过度保守决策且不牺牲安全性。通过迭代轨迹细化逐步聚合多尺度当前几何和潜在未来几何以有效利用这些几何线索。在NAVSIM v1和v2上的实验证明了其先进性能,凸显了明确的3D几何基础和未来几何世界建模对安全高效自动驾驶的有效性。

英文摘要

Autonomous driving requires both safe and efficient planning decisions in dynamic 3D environments. Although recent Vision/Video-Action models learn policies directly from visual observations and scale well with advances in vision transformers and large-scale training data, they often lack explicit geometric grounding and future-aware spatial guidance, limiting their ability to balance collision avoidance and driving progress. In this work, we propose GeoWorldAD, a geometry world action model that grounds trajectory planning in ego-aligned 3D space and anticipates short-horizon scene evolution with latent future geometry tokens. Present geometry provides essential spatial constraints for safe planning, while future geometry reveals how surrounding agents and ego-centric free space may evolve, reducing overly conservative decisions without sacrificing safety. To efficiently exploit these geometric cues, GeoWorldAD progressively aggregates multi-scale present geometry and latent future geometry through iterative trajectory refinement. Experiments on NAVSIM v1 and v2 demonstrate state-of-the-art performance, highlighting the effectiveness of explicit 3D geometry grounding and future geometry world modeling for safe and efficient autonomous driving.

URL PDF HTML 收藏
2607.17412 2026-07-21 cs.LG 新提交

CORAL: Learning Amyloid Fibril Ligand Docking with Cooperative Binding Rewards

CORAL:通过协同结合奖励学习淀粉样纤维配体对接

Yasheng Sun, Bohan Li, Youqi Tao, Jürgen Schmidhuber

机构 * Center of Excellence for Generative AI, KAUST(沙特阿卜杜拉国王科技大学生成式人工智能卓越中心) MoE Key Lab of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能教育部重点实验室) Stern Laboratory, Brigham and Women’s Hospital, Harvard University(哈佛大学布莱根妇女医院斯特恩实验室)

AI总结 研究针对淀粉样纤维配体对接面临的挑战,提出CORAL强化学习框架,通过纳入协同配体 - 配体堆叠能量及蛋白质 - 配体对接亲和力训练模型,引入评估集,实验证明其在姿态质量和结合亲和力相关性上优于现有基线。

Comments 25 pages

详情
AI中文摘要

神经退行性疾病(如阿尔茨海默病和帕金森病)的一个标志是蛋白质异常聚集成淀粉样纤维,能选择性结合这些纤维的小分子有望用作诊断、成像探针和治疗剂。预测此类配体与纤维靶点的结合存在两个基本挑战:一是淀粉样配体复合物的共晶体结构极少,难以进行对接模型的监督训练;二是淀粉样纤维的结合模式与球状蛋白根本不同,现有对接模型无法捕捉。为应对这些挑战,我们提出CORAL,这是一个强化学习框架,训练生成对接模型以生成适合交叉β凹槽几何形状的配体姿态分布。我们的奖励明确纳入了配体 - 配体堆叠能量以及蛋白质 - 配体对接亲和力,直接捕捉淀粉样纤维独特的结合几何形状。我们还引入了由领域专家验证的模型生成姿态构建的淀粉样配体复合物精选评估集。实验表明,与现有对接基线相比,姿态质量和结合亲和力相关性得到了改善。

英文摘要

A hallmark of neurodegenerative diseases such as Alzheimer's and Parkinson's is the aberrant aggregation of proteins into amyloid fibrils, and small molecules that selectively bind to these fibrils hold promise as diagnostics, imaging probes, and therapeutics. Predicting how such ligands bind to fibril targets, however, presents two fundamental challenges. First, resolved co-crystal structures of amyloid-ligand complexes are exceptionally scarce; even with recent advances in cryo-EM only a handful have been structurally characterized, making supervised training of docking models impractical for this target class. Second, amyloid fibrils present a binding mode fundamentally different from globular proteins: ligands intercalate into longitudinal cross-$β$ grooves and stack cooperatively along the fibril axis, a geometry that existing docking models are not designed to capture. To address these challenges, we present CORAL (COopeRative Amyloid Ligand docking), a reinforcement learning framework that trains a generative docking model to produce ligand pose distributions tailored to the cross-$β$ groove geometry. Our reward explicitly incorporates cooperative ligand-ligand stacking energy alongside protein-ligand docking affinity, directly capturing the distinctive binding geometry of amyloid fibrils. We further introduce a curated evaluation set of amyloid-ligand complexes constructed from model-generated poses validated by domain experts. Experiments on both experimentally resolved structures and this evaluation set demonstrate improved pose quality and binding affinity correlation over existing docking baselines.

URL PDF HTML 收藏
2607.16224 2026-07-21 cs.CY cs.AI 新提交

International Agreements to Limit Frontier AI: Objectives and Exit

限制前沿人工智能的国际协议:目标与退出

Lennart Finke

机构 * ETH Zürich, Zurich, Switzerland(苏黎世联邦理工学院) Harvard University, Cambridge, United States(哈佛大学) MATS Research, Berkeley, United States(MATS研究)

AI总结 研究限制人工智能发展的国际协议,通过调查现有协议,明确适当条件属性并列出可能条件,建议设定固定时间段,由新组织规定发展条件,特殊情况可退出,以此说明相关协议的考虑因素。

详情
AI中文摘要

限制人工智能发展的国际协议对于减轻人工智能风险可能至关重要。然而,尚不清楚哪些条件应决定何时放宽限制措施。我们调查了现有国际协议,概述了适当条件应满足的属性,列出了可能的条件,并在一个示例场景中给出了建议。我们建议设定一个固定时间段,在此期间开始时成立的新组织规定人工智能何时可以安全发展的条件,在特殊情况下有可能退出。我们希望说明在限制人工智能的国际协议中可能会涉及的考虑因素。

英文摘要

An international agreement to limit AI development could be crucial to mitigate risks from AI. However, it remains unclear which conditions should determine when the limiting measures are relaxed. We survey existing international agreements, outline what properties appropriate conditions should satisfy, list possible conditions, and finally give a recommendation in an example scenario. We recommend a fixed time period after which a new organization established at the start of the period specifies conditions that address when AI development can be safely conducted, with a possibility of withdrawal in extraordinary circumstances. We hope to illustrate the considerations that would likely go into an international agreement to limit AI.

URL PDF HTML 收藏
2607.16554 2026-07-21 cs.LG cs.AI cs.IT math.IT 新提交

Capacity and Redundancy Trade-offs in Multi-Task Learning

多任务学习中的容量与冗余权衡

Asif Khan

机构 * Harvard Medical School(哈佛医学院)

AI总结 研究多任务学习中负迁移与容量、冗余的关系,通过容量 - 冗余恒等式及相关结果,如聚类差距分解和梯度 - TC 桥梁,证明聚类 LoRA 可降低残余耦合,优于随机划分,有显著收益。

Comments Accepted in 42nd Conference on Uncertainty in Artificial Intelligence (UAI) 2026

详情
AI中文摘要

在多任务学习(MTL)中,负迁移通常被视为一种优化假象,但它也可被视为共享容量有限和任务冗余薄弱的结果。我们通过容量 - 冗余(CR)恒等式来研究这种效应,该恒等式将每个任务的预测信息之和分解为联合预测信息(包括通过全相关(TC)定义的标签冗余)和一个残余耦合项(量化共享表示未解决的干扰)。此外,我们展示了两个关键结果:(i)聚类差距分解,给出聚类共享优于全局共享的充要条件;(ii)高斯多任务模型中的梯度 - TC 桥梁,从形式上证明梯度余弦相似度可作为冗余排序的代理。从经验上看,我们从验证残余相关性估计残余耦合$\Delta$,表明聚类 LoRA 显著降低$\widehat{\Delta}$,优于大小匹配的随机划分,并在多种子置信区间下带来统计上显著的收益。

英文摘要

In multi-task learning (MTL) negative transfer is often considered as an optimization artifact, but it can also be viewed as a consequence of limited shared capacity and weak task redundancy. We investigate this effect through a Capacity--Redundancy (CR) identity that decomposes the sum of per-task predictive informations into joint predictive information that includes label redundancy defined via total correlation (TC), and a residual coupling term that quantifies interference left unresolved by the shared representation. Additionally, we show two key results: (i) a clustering-gap decomposition that gives a necessary and sufficient condition for clustered sharing to outperform global sharing, and (ii) a gradient--TC bridge in a Gaussian multi-task model that formally justifies gradient cosine similarity as a proxy for redundancy ordering. Empirically, we estimate the residual coupling $Δ$ from validation residual correlations, showing that clustered LoRA substantially reduces $\widehatΔ$, outperforms size-matched random partitions, and results in statistically significant gains with multi-seed confidence intervals.

URL PDF HTML 收藏
2607.18209 2026-07-21 math.ST cs.LG stat.ME stat.ML stat.TH 新提交

Unveiling Invariant and Transferable Latent Factors Across Heterogeneous Environments via ATLAS

通过ATLAS揭示异构环境中的不变和可转移潜在因素

Yihong Gu, Katherine Liao, Tianxi Cai

机构 * Harvard University(哈佛大学)

AI总结 研究异构环境下多环境因素模型,提出ATLAS方法,利用不变性原理和辅助标签监督,分离不变与异构因素,用于潜在因素回归,在新环境中实现可转移预测,还建立了相关非渐近误差界。

Comments 47 pages, 3 figures

详情
AI中文摘要

本文考虑一个多环境因素模型,其中高维协变量从异构环境中收集,且在部分环境中有辅助标签。协变量的联合分布可能因环境而异,潜在结构分解为具有共享载荷的不变因素和具有特定环境载荷的异构因素。该模型受迁移学习和潜在因素回归的启发。利用不变性原理,我们表明在最小结构条件下可解开不变和异构因素。基于此,我们提出ATLAS,它通过跨异构环境的潜在对齐,利用辅助标签和不变性指导进行转移。ATLAS是一个统一过程,利用不变性原理分离对齐的不变和未对齐的异构因素,并利用辅助标签监督从那些未对齐的异构因素中提取预测不变和可转移因素。ATLAS在下游潜在因素回归中产生接近最优的性能,当有辅助标签时通过完整潜在信号在新环境中实现可转移预测,否则简化为仅健壮的不变因素预测。我们为恢复不变和异构因素、识别所有响应不变因素以及估计Y中的不变信号建立了精确的非渐近误差界。

英文摘要

This paper considers a multi-environment factor model in which high-dimensional covariates are collected from heterogeneous environments, with auxiliary labels available in a subset of these environments. The joint distribution of the covariates may vary across environments, whereas the latent structure is decomposed into invariant factors with shared loadings and heterogeneous factors with environment-specific loadings. Such a model is motivated by transfer learning and latent factor regression, where one seeks stable low-dimensional representations for both interpretation and robust out-of-sample prediction of the response $Y$. Leveraging the invariance principle, we show that the invariant and heterogeneous factors are disentangled under a minimal structural condition. Based on this, we propose ATLAS, an Auxiliary-label and invariance-guided Transfer via Latent Alignment across heterogeneous environmentS. ATLAS is a unified procedure that leverages the invariance principle to separate aligned invariant and unaligned heterogeneous factors, and further exploits supervision from auxiliary labels to extract prediction-invariant and transferable factors from those unaligned heterogeneous factors. ATLAS yields near-oracle performance for downstream latent factor regression, enables transferable prediction in new environments through the full latent signal when auxiliary labels are available, and reduces to robust invariant-factor-only prediction otherwise. We establish sharp non-asymptotic error bounds for recovering invariant and heterogeneous factors, identifying all the response-invariant factors, and estimating the invariant signal in $Y$.

URL PDF HTML 收藏
2607.08793 2026-07-21 stat.ML cs.AI cs.LG cs.SY eess.SY math.OC 版本更新

EHR-MPC: Inference-Time Control for Sepsis Treatment with Generative Patient Digital Twins

EHR-MPC:利用生成式患者数字孪生进行脓毒症治疗的推理时间控制

Joshua Pickard, Wei Qi, Na Li, Ann Woolley, Lisa Cosimi, Roy Kishony, Deborah Hung

机构 * Broad Institute of MIT and Harvard(MIT和哈佛大学Broad研究所) Harvard University(哈佛大学) Brigham and Women’s Hospital(布莱根妇女医院) Technion–Israel Institute of Technology(技术学院-以色列理工学院)

AI总结 针对脓毒症治疗策略有争议且现有强化学习方法适应性不足的问题,提出EHR-MPC框架,通过训练生成式电子健康记录模型形式的患者数字孪生解耦学习与治疗优化,经模拟评估性能优于强化学习基线,建立了决策通用框架。

详情
AI中文摘要

脓毒症是主要死因,但最佳治疗策略仍有争议。现有强化学习方法学习固定策略,限制了推理时对变化临床目标的适应性。我们提出EHR-MPC框架,通过训练生成式电子健康记录模型形式的患者数字孪生,将学习患者动态与优化治疗解耦。数字孪生预测干预下的临床轨迹,使模型预测控制通过推理时模拟规划优化治疗。我们用离策略重要性采样和基于策略的模拟评估在多中心ICU脓毒症队列上评估EHR-MPC。相对于强化学习基线,EHR-MPC实现了可比的离策略性能和更好的模拟性能。不同于强化学习,这项工作将脓毒症治疗优化框架化为对学习到的患者动态的推理时控制,建立了使用生成式临床模型进行决策的通用框架。

英文摘要

Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested. Existing reinforcement learning (RL) approaches learn fixed strategies for sepsis treatment, limiting adaptability to changing clinical objectives during inference. We propose EHRMPC, a framework that decouples learning patient dynamics from optimizing treatment by training a patient digital twin in the form of a generative electronic health record (EHR) model. The digital twin predicts clinical trajectories under interventions and enables model predictive control (MPC) to optimize treatments via inference-time planning over simulations. We evaluate EHR-MPC on a multicenter ICU sepsis cohort spanning 8 hospitals in the Mass General Brigham health system using both off-policy importance sampling and on-policy simulation-based evaluation. Relative to RL baselines, EHR-MPC achieves comparable off-policy performance and improved simulation performance. Unlike RL, this work frames sepsis treatment optimization as inference-time control over learned patient dynamics, establishing a general framework for decision making with generative clinical models.

URL PDF HTML 收藏
2606.32017 2026-07-21 cs.LG cs.AI 版本更新

TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

TRIAGE:面向智能体强化学习的角色类型化信用分配

Yuanda Xu, Zhengze Zhou, Hejian Sang, Xiaomin Li, Jiaxin Zhang, Xinchen Du, Sen Na, Zhipeng Wang, Alborz Geramifard

机构 * LinkedIn Corporation(领英公司) Harvard University(哈佛大学) Johns Hopkins University(约翰霍普金斯大学)

AI总结 提出TRIAGE框架,通过角色类型化信用分配修正GRPO仅依赖结果信用的盲点,在ALFWorld等任务中提升成功率并减少交互步数。

详情
AI中文摘要

智能体强化学习需要将信用分配给环境交互动作,如搜索、点击、编辑、导航命令和对象交互。标准GRPO使用最终验证器结果作为所有动作令牌的统一优势。该结果信号有用但结构不完整:它在失败轨迹中惩罚有用的探索,并在成功轨迹中强化冗余或倒退动作。我们提出TRIAGE,一种角色类型化信用分配框架,为结果信用添加语义角色轴。结构化评判器将每个片段分类为决定性进展、有用探索、无进展基础设施或倒退,固定角色条件规则将这些标签映射到有界片段级过程奖励。这保持了验证器结果作为优化方向的来源,同时纠正了仅结果信用的两个主要盲点。我们进一步证明,角色条件信用是从角色标签本身可表达的最优片段级修正——将每片段优势残差投影到角色变量上——因此当评判器可靠时,固定角色常数减少优势估计误差,并将其与低方差策略梯度联系起来。在ALFWorld、Search-QA和WebShop上,TRIAGE在两种策略模型上均优于GRPO的成功率,并优于标量评判器推导的过程奖励和结果监督共享骨干价值基线。消融实验表明,增益来自角色类型化而非仅仅添加密集奖励:成功轨迹内倒退的可靠检测是主要贡献者,而探索信用提供一致的次要增益;在完成的ALFWorld和WebShop轨迹上,TRIAGE相对于GRPO还额外减少了10.4%和14.8%的环境交互步数。

英文摘要

Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and object interactions. Standard GRPO uses the final verifier outcome as a uniform advantage over all action tokens. This outcome signal is useful but structurally incomplete: it punishes useful exploration in failed rollouts and reinforces redundant or regressive actions in successful rollouts. We propose TRIAGE, a role-typed credit assignment framework that adds a semantic role axis to outcome credit. A structured judge classifies each segment as decisive progress, useful exploration, no-progress infrastructure, or regression, and a fixed role-conditioned rule maps these labels to bounded segment-level process rewards. This keeps verifier outcomes as the source of optimization direction while correcting the two main blind spots of outcome-only credit. We further show that the Bayes-optimal role-measurable correction is the L2 projection of the per-segment advantage residual onto the role variable, and that TRIAGE's fixed role constants approximate this projection, reducing advantage estimation error whenever the judge is reliable; we connect this to lower-variance policy gradients. Across ALFWorld, Search-QA, and WebShop, TRIAGE improves success rates over GRPO for two policy models and outperforms both a scalar judge-derived process reward and an outcome-supervised shared-backbone value baseline. Ablations show that the gain comes from role typing rather than merely adding dense rewards: reliable detection of regression inside successful trajectories is the dominant contributor, while exploration credit provides a consistent secondary gain; on completed ALFWorld and WebShop rollouts, TRIAGE also reduces environment-facing turns by an additional $10.4\%$ and $14.8\%$ relative to GRPO.

URL PDF HTML 收藏
2605.12111 2026-07-21 cs.AI cs.DS 版本更新

Adaptive Multi-Round Allocation with Stochastic Arrivals

自适应多轮分配与随机到达

Yuqi Pan, Davin Choo, Haichuan Wang, Milind Tambe, Alastair van Heerden, Cheryl Johnson

机构 * Harvard University(哈佛大学) University of Witwatersrand(沃特沙兰大学)

AI总结 研究了受适应性网络招募启发的序列资源分配问题,提出基于截断概率生成函数的动态规划算法,实现多项式复杂度的多轮分配方案,并分析了模型误差下的鲁棒性。

Comments Accepted into ICML 2026

详情
AI中文摘要

我们研究了一个受适应性网络招募启发的序列资源分配问题,其中有限的相同资源必须在多个轮次中分配给具有随机推荐能力的个体。成功的推荐会内生地产生未来决策机会,而向个体分配额外资源则表现出边际递减回报。我们首先证明单轮分配问题可以通过边际生存概率获得精确贪心解。在多轮设置中,由于前沿的随机性和高维演化,贝尔曼递归难以处理。为此,我们引入了一个仅依赖剩余预算和前沿大小的群体级替代价值函数。该替代函数通过截断概率生成函数实现了精确的动态规划,得到一个具有多项式复杂度的规划算法。我们进一步分析了模型误差下的鲁棒性,证明了多轮误差界可分解为紧致的单轮前沿误差和群体级转换误差。最后,我们在真实世界启发的招募场景上评估了我们的方法。

英文摘要

We study a sequential resource allocation problem motivated by adaptive network recruitment, in which a limited budget of identical resources must be allocated over multiple rounds to individuals with stochastic referral capacity. Successful referrals endogenously generate future decision opportunities while allocating additional resources to an individual exhibits diminishing returns. We first show that the single-round allocation problem admits an exact greedy solution based on marginal survival probabilities. In the multi-round setting, the resulting Bellman recursion is intractable due to the stochastic, high-dimensional evolution of the frontier. To address this, we introduce a population-level surrogate value function that depends only on the remaining budget and frontier size. This surrogate enables an exact dynamic program via truncated probability generating functions, yielding a planning algorithm with polynomial complexity in the total budget. We further analyze robustness under model misspecification, proving a multi-round error bound that decomposes into a tight single-round frontier error and a population-level transition error. Finally, we evaluate our method on real-world inspired recruitment scenarios.

URL PDF HTML 收藏
2507.06445 2026-07-21 cs.LG cs.AI cs.CL 版本更新

Can Interpretation Predict Behavior on Unseen Data?

解释能否预测未见过的数据上的行为?

Victoria R. Li, Jenny Kaufmann, Tian Qin, Martin Wattenberg, David Alvarez-Melis, Naomi Saphra

机构 * Harvard University(哈佛大学) Boston University(波士顿大学)

AI总结 研究探讨能否用模型内部预测对未见数据的响应,通过在合成任务上训练Transformer,利用注意力模式预测OOD行为,实验解耦解释的忠实性与预测价值,证明可通过理解模型内部评估分布转移下的行为和可靠性。

详情
AI中文摘要

可解释性研究通常预测模型对目标机制干预的响应。但我们能否预测对未见输入数据的响应呢?我们通过使用模型内部结构来预测其分布外(OOD)行为,提出并证明了这个替代目标。我们在简单的合成任务上训练了数百个Transformer,在此完美的分布内准确率与多种OOD泛化规则兼容。我们成功地利用仅在分布内数据上观察到的注意力模式,来预测每个模型在OOD数据上遵循的规则。我们的实验将解释的机制忠实性与其预测价值解耦;消融实验表明,这样的内部模式可能抑制而非支持它们所预测的规则,这表明即使因果分析无法支持简单的因果联系,观察性分析也能预测行为。我们的发现是一个新的可解释性目标的概念验证:理解模型内部结构以预测行为并评估分布转移下的可靠性。

英文摘要

Interpretability research often predicts model responses to targeted mechanistic interventions. But can we predict responses to unseen input data? We propose and demonstrate this alternate objective by using model internals to predict their out-of-distribution (OOD) behavior. We train hundreds of Transformers on simple synthetic tasks, where perfect in-distribution accuracy is compatible with multiple OOD generalization rules. We successfully use attention patterns -- observed only on in-distribution data -- to predict which rule each model follows on OOD data. Our experiments decouple the mechanistic faithfulness of our interpretation from its predictive value; ablations reveal such internal patterns can suppress rather than support the rule they predict, showing observational analysis can forecast behavior even when causal analysis fails to support a simple cause-effect link. Our findings are a proof-of-concept for a new interpretability objective: understanding model internals to predict behavior and assess reliability under distribution shift.

URL PDF HTML 收藏
2412.18613 2026-07-21 q-bio.NC cs.CL cs.CV 版本更新

The Illusion-Illusion: Vision Language Models See Illusions Where There Are None

错觉-错觉:视觉语言模型在不存在错觉的地方看到错觉

Tomer Ullman

机构 * Harvard University(哈佛大学)

AI总结 研究通过给视觉语言模型呈现不应引发处理错误的‘错觉-错觉’,发现许多模型会误将其视为错觉,揭示了模型存在基本处理错误,此失败是文献中更广泛失败的一部分。

Comments 9 pages, 5 figures

详情
AI中文摘要

错觉既有趣,又是认知科学、哲学和神经科学中有用的诊断工具。典型错觉展示了事物‘实际情况’与‘看起来的样子’之间的差距,有助于理解导致事物呈现方式的心理过程。错觉对研究人工系统也有用,许多研究探讨了感知计算模型是否会像人类一样陷入同样的错觉。本文通过呈现‘错觉-错觉’(常见错觉的类似物,不应引发处理错误)来研究当前视觉语言模型的基本处理错误。结果表明许多当前视觉语言系统会将这些‘错觉-错觉’误视为错觉,这种失败是文献中已讨论的更广泛失败的一部分。

英文摘要

Illusions are entertaining, but they are also a useful diagnostic tool in cognitive science, philosophy, and neuroscience. A typical illusion shows a gap between how something `really is' and how something `appears to be', and this gap helps us understand the mental processing that led to how something appears to be. Illusions are also useful for investigating artificial systems, and much research has examined whether computational models of perception fall prey to the same illusions as people. Here, I invert the standard use of perceptual illusions to examine basic processing errors in current vision language models. I present these models with illusory-illusions, neighbors of common illusions that should not elicit processing errors. These include such things as perfectly reasonable ducks, crooked lines that truly are crooked, circles that seem to have different sizes because they are, in fact, of different sizes, and so on. I show that many current vision language systems mistakenly see these illusion-illusions as illusions. I suggest that such failures are part of broader failures already discussed in the literature.

URL PDF HTML 收藏
2404.01549 2026-07-21 cs.CL cs.SE 版本更新

Octopus: On-device language model for function calling of software APIs

章鱼:用于软件API函数调用的设备端语言模型

Wei Chen, Zhiyuan Li, Mingyuan Ma

机构 * Stanford University(斯坦福大学) Harvard University(哈佛大学)

AI总结 研究利用设备端大语言模型调用软件API,通过编译数据集微调不同参数模型,提升其API交互能力,提出条件掩码技术和新基准,经微调的Octopus模型在API调用上性能超GPT-4,推动自动化软件开发和API集成。

详情
AI中文摘要

在快速发展的人工智能领域,大语言模型(LLMs)因其先进的文本处理和生成能力发挥着关键作用。本研究引入一种在调用软件API时利用设备端LLMs的新策略。精心编译源自软件API文档的数据集,对2B、3B和7B参数的LLMs进行微调,以提高其在软件API交互方面的能力,专注提升模型对API结构和语法的理解,增强API函数调用准确性。还提出条件掩码技术确保输出格式正确并降低错误率,同时保持推理速度。提出新基准评估LLMs在API交互中的有效性。经微调的Octopus模型在软件API调用方面性能优于GPT-4,推动了自动化软件开发和API集成。

英文摘要

In the rapidly evolving domain of artificial intelligence, Large Language Models (LLMs) play a crucial role due to their advanced text processing and generation abilities. This study introduces a new strategy aimed at harnessing on-device LLMs in invoking software APIs. We meticulously compile a dataset derived from software API documentation and apply fine-tuning to LLMs with capacities of 2B, 3B and 7B parameters, specifically to enhance their proficiency in software API interactions. Our approach concentrates on refining the models' grasp of API structures and syntax, significantly enhancing the accuracy of API function calls. Additionally, we propose \textit{conditional masking} techniques to ensure outputs in the desired formats and reduce error rates while maintaining inference speeds. We also propose a novel benchmark designed to evaluate the effectiveness of LLMs in API interactions, establishing a foundation for subsequent research. Octopus, the fine-tuned model, is proved to have better performance than GPT-4 for the software APIs calling. This research aims to advance automated software development and API integration, representing substantial progress in aligning LLM capabilities with the demands of practical software engineering applications.

URL PDF HTML 收藏
2607.16057 2026-07-20 cs.CL cs.AI 新提交

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

跨商业学科的前沿人工智能性能:基于案例的知识工作和分析推理基准

Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, Karim Lakhani, Mitch Weiss

机构 * The Wharton School, University of Pennsylvania(宾夕法尼亚大学沃顿商学院) Carnegie Mellon University(卡内基梅隆大学) Harvard Business School, Harvard University(哈佛大学哈佛商学院)

AI总结 研究针对人工智能在白领分析性知识工作衡量上的差距,利用顶尖商学院案例教学法构建BusinessCaseBench基准,发现前沿AI模型在此基准上得分高且能力提升快,为商学院及相关职业角色带来启示。

详情
AI中文摘要

大语言模型在基准测试分数中迅速提升,但这些人工智能基准大多测试事实性回忆、狭义问答、数学问题解决、编码和代理工具使用等能力。对于白领专业人员日常进行的分析性知识工作,包括综合复杂信息、在不确定和不完整信息下做出判断、在多利益相关者环境中应用战略和对抗性思维、权衡取舍以及进行合理的结构化分析等方面的人工智能进展衡量不足。顶尖商学院采用的“案例教学法”为解决这一衡量差距提供了自然基础。我们构建了BusinessCaseBench基准,涵盖来自18个学科商业案例的数百个问题,并配有专家编写的教师案例解决方案得出的评分标准。前沿人工智能模型在BusinessCaseBench上已经能根据教师评分标准获得高分,且一个模型家族的能力在两年内有显著提升。这些结果表明人工智能在这类工作上的表现已经很高且正在迅速改善,对商学院和入门级专业角色有影响。

英文摘要

Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty and incomplete information, applying strategic and adversarial thinking in multi-stakeholder settings, weighing trade-offs, and producing defensible, structured analyses. This gap is even more pronounced for subjective components of such work, where success can be challenging to define. The "case method" form of education practiced by top business schools provides a natural foundation for addressing this measurement gap, and we construct BusinessCaseBench, a benchmark spanning hundreds of questions drawn from business cases across eighteen disciplines, each paired with a grading rubric derived from the expert-written instructor case solution. On BusinessCaseBench, frontier AI models already score highly against instructor rubrics, and capability within one model family improves substantially over two years. These results provide strong evidence that AI performance on this class of work is already high and rapidly improving, with implications for business schools, where case pedagogy trains undergraduates and MBAs in this kind of analytical reasoning, and for entry-level professional roles, where such skills have historically anchored early-career work.

URL PDF HTML 收藏
2604.23786 2026-07-20 cs.AI cs.LG 版本更新

FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment

FAIR_XAI: 通过可解释性提升多模态基础模型公平性以用于幸福感评估

Sophie Chiang, Tom Brennan, Fethiye Irmak Dogan, Jiaee Cheong, Hatice Gunes

机构 * Department of Computer Science & Technology, University of Cambridge(计算机科学与技术系,剑桥大学) Harvard University(哈佛大学)

AI总结 本文研究了多模态基础模型在幸福感评估中的公平性问题,通过可解释性干预框架改善诊断可靠性与公平性,发现不同模型在不同数据集上表现差异显著,且存在性别和种族偏见。

Comments 11 pages, 4 figures

详情
AI中文摘要

近年来,多模态机器学习在幸福感评估中的整合为心理健康监测提供了变革性潜力。然而,随着视觉-语言模型(VLMs)的快速发展,其在临床应用中的部署引发了透明度不足和潜在偏见的担忧。尽管先前研究探讨了公平性与可解释人工智能(XAI)的交集,但将其应用于VLMs进行幸福感评估和抑郁症预测仍显不足。本文研究了VLMs在实验室(AFAR-BSFT)和自然(E-DAIC)数据集上的表现,聚焦诊断可靠性与人口公平性。性能在不同环境和架构间差异显著;Phi3.5-Vision在E-DAIC上达到80.4%的准确率,而Qwen2-VL在同数据集上仅33.9%。此外,两种模型在AFAR-BSFT上均表现出对抑郁症的过度预测倾向。尽管两种架构均存在偏见,但Qwen2-VL显示出更高的性别差异,而Phi-3.5-Vision则表现出更多的种族偏见。我们的XAI干预框架产生了混合结果;公平性提示在Qwen2-VL上实现了完美的相等机会,但以严重的准确率代价。在AFAR-BSFT上,基于可解释性的干预提高了程序一致性,但未保证结果公平性,有时加剧了种族偏见。这些结果突显了程序透明度与公平结果之间的持续差距。我们分析了这些发现并提出了具体的解决建议,强调未来公平性干预必须共同优化预测准确性、人口平等性和跨领域泛化能力。

英文摘要

In recent years, the integration of multimodal machine learning in wellbeing assessment has offered transformative potential for monitoring mental health. However, with the rapid advancement of Vision-Language Models (VLMs), their deployment in clinical settings has raised concerns due to their lack of transparency and potential for bias. While previous research has explored the intersection of fairness and Explainable AI (XAI), its application to VLMs for wellbeing assessment and depression prediction remains under-explored. This work investigates VLM performance across laboratory (AFAR-BSFT) and naturalistic (E-DAIC) datasets, focusing on diagnostic reliability and demographic fairness. Performance varied substantially across environments and architectures; Phi3.5-Vision achieved 80.4% accuracy on E-DAIC, while Qwen2-VL struggled at 33.9%. Additionally, both models demonstrated a tendency to over-predict depression on AFAR-BSFT. Although bias existed across both architectures, Qwen2-VL showed higher gender disparities, while Phi-3.5-Vision exhibited more racial bias. Our XAI intervention framework yielded mixed results; fairness prompting achieved perfect equal opportunity for Qwen2-VL at a severe accuracy cost on E-DAIC. On AFAR-BSFT, explainability-based interventions improved procedural consistency but did not guarantee outcome fairness, sometimes amplifying racial bias. These results highlight a persistent gap between procedural transparency and equitable outcomes. We analyse these findings and consolidate concrete recommendations for addressing them, emphasising that future fairness interventions must jointly optimise predictive accuracy, demographic parity, and cross-domain generalisation.

URL PDF HTML 收藏
2512.22274 2026-07-20 cs.CV 版本更新

GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure

GeCo:通过运动和结构评估视频生成的几何一致性

Leslie Gu, Junhwa Hur, Charles Herrmann, Fangneng Zhan, Todd Zickler, Deqing Sun, Hanspeter Pfister

机构 * Harvard University(哈佛大学) Google DeepMind(谷歌DeepMind) MIT(麻省理工学院)

AI总结 GeCo通过融合残差运动和深度先验,检测静态场景中的几何变形和遮挡不一致问题,并用于评估视频生成模型的性能与缺陷。

详情
AI中文摘要

我们介绍了GeCo,一种基于几何的度量标准,用于联合检测静态场景中的几何变形和遮挡不一致伪影。通过融合残差运动和深度先验,GeCo生成可解释的密集一致性图,揭示这些伪影。我们使用GeCo系统地评估最近的视频生成模型,发现常见的失败模式,并进一步将其用作无训练指导损失,以减少视频生成中的变形伪影。

英文摘要

We introduce GeCo, a geometry-grounded metric for jointly detecting geometric deformation and occlusion-inconsistency artifacts in static scenes. By fusing residual motion and depth priors, GeCo produces interpretable, dense consistency maps that reveal these artifacts. We use GeCo to systematically benchmark recent video generation models, uncovering common failure modes, and further employ it as a training-free guidance loss to reduce deformation artifacts during video generation.

URL PDF HTML 收藏
2205.04599 2026-07-20 cs.LG cs.AI 版本更新

Perception-Aligned AI Outputs: End-to-End Visual Prediction for Uncertainty Communication in Clinical Decision-Making

感知对齐的人工智能输出:临床决策中用于不确定性通信的端到端视觉预测

Mohammad Eslami, Solale Tabarestani, Saber Kazeminasab, Ehsan Adeli, Glyn Elwyn, Tobias Elze, Mengyu Wang, Nazlee Zebardast, Lucia Sobrin, Nassir Navab, Daniel Shu Wei Ting, Malek Adjouadi

机构 * Harvard Ophthalmology AI Lab(哈佛眼科人工智能实验室) Schepens Eye Research Institute of Massachusetts Eye and Ear(马萨诸塞眼耳医院施佩恩眼科研究所) Harvard Medical School(哈佛医学院) Center for Advanced Technology and Education(先进教育技术中心) Florida International University(佛罗里达国际大学) Dartmouth Institute for Health Policy and Clinical Practice(达特茅斯健康政策与临床实践研究所) Dartmouth College(达特茅斯学院) Computer Aided Medical Procedures(医学辅助程序) Technical University of Munich(慕尼黑技术大学) Singapore Eye Research Institute(新加坡眼科研究所) Singapore National Eye Centre(新加坡国家眼科中心) Department of Ophthalmology, Byers Eye Institute, Stanford University(眼科部门,比尔斯眼科研究所,斯坦福大学)

AI总结 研究针对医疗保健中可解释人工智能的问题,提出以人为本的机器学习可视化学习框架VL4ML,通过直观视觉表示传达模型预测与不确定性,经多临床任务验证及评估,结果显示其能有效支持临床决策,具有广泛可及性。

详情
AI中文摘要

可解释人工智能(XAI)对医疗保健领域中可靠的人工智能至关重要,但现有许多方法依赖于临床医生和患者难以理解的技术解释。我们引入了机器学习可视化学习(VL4ML),这是一个以人为本的可解释性框架,通过直观的视觉表示而非数值或事后解释来传达模型预测和不确定性。通过在颜色、图案和空间结构中编码诊断信息,VL4ML使用户无需了解模型内部或统计专业知识就能解释预测。我们在包括分类、回归、纵向预测和多模态分析等多个临床任务中展示了该框架。通过一项涉及158名参与者(39.2%为临床专业人员)的以人为本的研究和专家可解释性评估对其有效性进行了评估。超过79%的参与者在评估维度上对视觉解释给予了积极评价,84.0%的人认为它们比数字输出更令人难忘,76.9%的人报告决策更快。超过82%的人在没有事先统计训练的情况下成功感知到视觉表示中嵌入的不确定性。临床医生和非临床医生之间以及男性和女性参与者之间没有观察到显著差异,表明其具有广泛的可及性。这些结果表明,VL4ML通过提供直观、普遍可解释的视觉解释来补充现有的XAI和不确定性量化方法,支持透明和可靠的临床决策。

英文摘要

Explainable Artificial Intelligence (XAI) is essential for trustworthy AI in healthcare, yet many existing methods rely on technical explanations that are difficult for clinicians and patients to interpret. We introduce Visualized Learning for Machine Learning (VL4ML), a human-centered explainability framework that communicates model predictions and uncertainty through intuitive visual representations rather than numerical or post-hoc explanations. By encoding diagnostic information in colors, patterns, and spatial structures, VL4ML enables users to interpret predictions without requiring knowledge of model internals or statistical expertise. We demonstrate the framework across multiple clinical tasks, including classification, regression, longitudinal prediction, and multimodal analysis. Its effectiveness was evaluated through a human-centered study involving 158 participants (39.2% clinical professionals) and an expert interpretability assessment. More than 79% of participants positively rated the visual explanations across evaluation dimensions, 84.0% found them more memorable than numeric outputs, and 76.9% reported faster decision-making. Over 82% successfully perceived uncertainty embedded in the visual representations without prior statistical training. No significant differences were observed between clinicians and non-clinicians or between male and female participants, indicating broad accessibility. These results suggest that VL4ML complements existing XAI and uncertainty quantification methods by providing intuitive, universally interpretable visual explanations that support transparent and trustworthy clinical decision-making.

URL PDF HTML 收藏
2607.15065 2026-07-17 cs.RO cs.CV cs.LG 新提交

DriftWorld: Fast World Modeling through Drifting

DriftWorld:通过漂移实现快速世界建模

Susie Lu, Haonan Chen, Weirui Ye, Yilun Du

机构 * Massachusetts Institute of Technology(麻省理工学院) Harvard University(哈佛大学)

AI总结 研究针对预测性世界模型生成展开慢的问题,提出基于漂移生成模型的DriftWorld,训练时学习动作条件漂移以快速生成未来帧,在机器人操作基准测试中实现快速准确决策,还能作离线模拟器,性能优于基于扩散的基线。

Comments Website at https://susie-lu.github.io/driftworld/

详情
AI中文摘要

预测性世界模型能让机器人通过想象行动结果来进行规划,但其对控制的价值取决于能否快速生成多个展开。这给基于扩散的世界模型带来瓶颈:多步采样使每个展开成本高昂,限制了推理时的大规模动作搜索。我们引入DriftWorld,一种基于漂移生成模型的动作条件世界模型。它在训练时学习动作条件漂移,而非在推理时迭代去噪,能在单次前向传播中以30+帧每秒的速度从当前观察和候选动作序列生成未来帧,比基于扩散的基线平均快17倍。我们在标准视觉机器人操作基准上评估DriftWorld,它生成的展开既准确又快速,在推理时间远少于基于扩散的世界模型基线的情况下实现了最优决策性能。此外,DriftWorld还可作为离线模拟器对现实世界机器人策略进行排序,基于展开的分数与地面真值的相关性高达0.99。这些结果表明漂移模型非常适合机器人世界建模,快速、高质量的想象能直接支持规划和策略评估。

英文摘要

Predictive world models enable robots to plan by imagining the outcomes of their actions, but their value for control hinges on generating many rollouts quickly. This creates a bottleneck for diffusion-based world models: multistep sampling makes each rollout expensive, limiting large-scale action search at inference time. We introduce DriftWorld, an action-conditioned world model based on drifting generative models. Rather than denoising iteratively at inference, DriftWorld learns an action-conditioned drift during training, allowing it to generate future frames from the current observation and a candidate action sequence in a single forward pass at 30+ fps, which is 17x faster on average than diffusion based baselines. We evaluate DriftWorld on standard vision-based robotic manipulation benchmarks, including Bridge-V2, RT-1, Language Table, Push-T, and Robomimic. By producing rollouts that are both accurate and fast, DriftWorld achieves state-of-the-art decision-making performance with far less inference time than diffusion-based world model baselines. Beyond online control, DriftWorld can also serve as an offline simulator for ranking real-world robot policies, with rollout-based scores correlating with ground truth at up to 0.99. These results show that drifting models are a strong fit for robot world modeling, where fast, high-quality imagination directly supports planning and policy evaluation.

URL PDF HTML 收藏
2607.14946 2026-07-17 cs.CV 新提交

DINE: Distance Is Not Enough -- Learning Global Deformation Priors for Robust Soft-Tissue Point Cloud Registration

DINE:距离并不够——学习用于鲁棒软组织点云配准的全局变形先验

Sara Monji-Azad, Rohit Beer, Marvin Kinz, Claudia Scherl, Jürgen Hesser

机构 * Mannheim Institute for Intelligent Systems in Medicine (MIISM), Medical Faculty Mannheim, Heidelberg University(曼海姆医学智能系统研究所(MIISM),海德堡大学曼海姆医学院) Department of Radiation Oncology, Brigham and Women’s Hospital, Dana-Farber Cancer Institute, Harvard Medical School(布莱根妇女医院放射肿瘤学系,达纳-法伯癌症研究所,哈佛医学院) Department of Otorhinolaryngology, Head and Neck Surgery, University Medical Center Mannheim, Medical Faculty Mannheim, Heidelberg University(海德堡大学曼海姆医学院曼海姆大学医学中心耳鼻咽喉头颈外科) Department of Otolaryngology, Head and Neck Surgery, Campus Klinikum Bielefeld Mitte, University Hospital OWL of Bielefeld University(比勒费尔德大学OWL大学医院比勒费尔德市中心校区耳鼻咽喉头颈外科) Interdisciplinary Center for Scientific Computing (IWR), Heidelberg University(海德堡大学跨学科科学计算中心(IWR)) Central Institute for Computer Engineering (ZITI), Heidelberg University(海德堡大学中央计算机工程研究所(ZITI)) CZS Heidelberg Center for Model-Based AI, Heidelberg University(海德堡大学CZS基于模型的人工智能中心)

AI总结 研究针对软组织点云配准中对应估计难题,提出DINE框架,通过学习统计先验增强基于距离的配准,应用于两个主干并采用两阶段策略,实验表明其能降低平均倒角距离,提高对变形和噪声的鲁棒性,凸显全局变形合理性的重要性。

详情
AI中文摘要

非刚性点云配准是软组织形状分析的核心,但大变形、噪声和离群值使对应估计具有挑战性。大多数基于学习的方法依赖局部目标,如倒角距离,虽鼓励逐点接近,但不约束预测变形场的全局合理性。我们用DINE解决此限制,它是一个最大后验框架,通过对位移向量场的学习统计先验增强基于距离的配准。DINE应用于两个配准主干,采用两阶段策略:第一阶段模型用倒角距离训练,其预测变形场用于估计先验,然后用组合距离和负对数先验目标细化模型。我们比较全场PCA高斯先验和逐向量归一化流先验。在DeformedTissue和SynBench上的实验表明,在变形和损坏情况下平均倒角距离更低。在DeformedTissue上,相对于相应的第一阶段主干,DINE - PCA在不同变形水平下将倒角距离降低约27 - 69%,对离群值的鲁棒性提高高达66%,对高斯噪声的鲁棒性提高83%。在SynBench上,在最小变形水平下改进不大,从中等变形到严重变形时提高约59 - 79%。这些结果表明全局变形合理性是可靠软组织点云配准的重要约束。

英文摘要

Non-rigid point cloud registration is central to soft-tissue shape analysis, but large deformations, noise, and outliers make correspondence estimation challenging. Most learning-based methods rely on local objectives such as Chamfer distance, which encourage point-wise proximity but do not constrain the global plausibility of the predicted deformation field. We address this limitation with DINE, a maximum a posteriori framework that augments distance-based registration with a learned statistical prior over displacement vector fields. DINE is applied to two registration backbones, Robust-DefReg and DefTransNet, using a two-stage strategy: a first-stage model is trained with Chamfer distance, its predicted deformation fields are used to estimate a prior, and the model is then refined with a combined distance and negative log-prior objective. We compare a full-field PCA Gaussian prior with a per-vector normalizing-flow prior. Experiments on DeformedTissue and SynBench show lower mean Chamfer distance under deformation and corruption. On DeformedTissue, DINE-PCA reduces Chamfer distance by approximately 27--69\% relative to the corresponding Stage-1 backbone across deformation levels, and improves robustness by up to 66\% for outliers and 83\% for Gaussian noise. On SynBench, improvements are modest at the smallest deformation levels and reach approximately 59--79\% from moderate to severe deformation. These results suggest that global deformation plausibility is an important constraint for reliable soft-tissue point cloud registration. (The code will be published soon.)

URL PDF HTML 收藏
2512.15948 2026-07-17 cs.AI q-bio.NC 版本更新

Subjective functions

主观函数

Samuel J. Gershman

机构 * Harvard University(哈佛大学) Department of Psychology(心理学系) Center for Brain Science(脑科学中心) Kempner Institute for the Study of Natural and Artificial Intelligence(自然与人工智能研究学院)

AI总结 本文探讨了主观函数的概念,通过预期预测误差作为例子,提出了一种赋予人工系统自主设定目标能力的方法,连接了心理学、神经科学和机器学习。

详情
AI中文摘要

客观函数从何而来?我们如何选择追求的目标?人类智能能够即时合成新的客观函数。本文提出了一种方法,从主观函数的概念出发,即一种内生于智能体的高阶目标函数。预期预测误差被作为主观函数的一个具体例子。该方法与心理学、神经科学和机器学习中的多个概念密切相关。

英文摘要

Where do objective functions come from? How do we select what goals to pursue? Human intelligence is adept at synthesizing new objective functions on the fly. How does this work, and can we endow artificial systems with the same ability? This paper proposes an approach to answering these questions, starting with the concept of a subjective function, a higher-order objective function that is endogenous to the agent (i.e., defined with respect to the agent's features, rather than an external task). Expected prediction error is studied as a concrete example of a subjective function. This proposal has many connections to ideas in psychology, neuroscience, and machine learning.

URL PDF HTML 收藏
2508.00923 2026-07-17 cs.LG 版本更新

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

超越基准:动态、自动和系统化的红队代理用于可信的医疗语言模型

Jiazhen Pan, Bailiang Jian, Paul Hager, Yundi Zhang, Che Liu, Friederike Jungmann, Hongwei Bran Li, Julian Canisius, Chenyu You, Junde Wu, Jiayuan Zhu, Fenglin Liu, Yuyuan Liu, Niklas Bubeck, Moritz Knolle, Chen, Chen, Christian Wachinger, Zhenyu Gong, Cheng Ouyang, Georgios Kaissis, Benedikt Wiestler, Daniel Rueckert

机构 * Technical University of Munich (TUM)(慕尼黑技术大学) University of Oxford(牛津大学) TUM University Hospital(慕尼黑技术大学医院) Imperial College London(伦敦帝国理工学院) Harvard Medical School(哈佛医学院) Stony Brook University(史泰兹布鲁克大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) University of Sheffield(谢菲尔德大学)

AI总结 本文提出DAS红队框架,通过动态压力测试揭示医疗语言模型在鲁棒性、隐私、偏见和幻觉方面的潜在风险,发现高静态基准性能与低动态可靠性之间的'基准差距'。

详情
AI中文摘要

确保大型语言模型(LLMs)在临床实践中的安全性和可靠性对于防止患者伤害至关重要。然而,LLMs的发展速度如此之快,静态基准很快变得过时或容易过拟合,导致对模型可信度的误导性图像。在这里,我们介绍了一个动态、自动和系统化的(DAS)红队框架,该框架在四个关键安全轴上持续对LLMs进行压力测试:鲁棒性、隐私、偏见/公平性和幻觉。经过获得认证的临床医生验证,一组对抗性代理会自动突变临床测试用例,以实时揭示漏洞。将DAS应用于15个专有和开源的LLMs,揭示了高静态基准性能与低动态可靠性之间的深刻差距——“基准差距”。尽管中位数MedQA准确率超过80%,但94%的先前正确答案在我们的动态鲁棒性测试中失败。至关重要的是,这种脆弱性扩展到了现实的、开放性的HealthBench数据集,在此数据集中,顶级模型的失败率超过70%,且在评估中模型排名出现了显著变化,表明在已建立的静态基准上的高分可能反映的是表面记忆。我们观察到在其他领域也出现了类似的高失败率:隐私泄露在86%的场景中被触发,认知偏见先验改变了81%的公平性测试中的临床建议,我们发现广泛使用的模型中的幻觉率超过74%。通过将医疗LLM安全分析从静态清单转换为动态压力测试,DAS提供了一个基础、可扩展和持续的平台,以揭示必须在下一代医疗AI安全部署之前解决的潜在风险。

英文摘要

Large language models (LLMs) are increasingly used to answer health-related questions and support healthcare workflows, yet evidence for their safety still relies heavily on static benchmarks that can rapidly become obsolete or be optimized against. Here we introduce a Dynamic, Automatic, and Systematic (DAS) red-teaming audit framework that continuously stress-tests LLMs for health across four safety-critical axes: robustness, privacy, bias/fairness, and hallucination/factual inaccuracies. Validated against board-certified clinicians with high concordance, a suite of adversarial agents autonomously mutates health-related test cases to uncover vulnerabilities in real time. Applying DAS to 15 proprietary and open-source LLMs revealed a profound gap between high static benchmark performance and low dynamic reliability--the "Benchmarking Gap". Despite median MedQA accuracy exceeding 80\%, 94\% of previously correct answers failed under dynamic robustness testing. This brittleness generalized to the realistic, open-ended HealthBench dataset, where top-tier models exhibited failure rates exceeding 70\% and sharp shifts in model rankings across evaluations, suggesting that high scores on established static benchmarks may reflect superficial memorization. We observed similarly high failure rates across other domains: privacy leaks were elicited in 86\% of scenarios, cognitive-bias priming altered recommendations in 81\% of fairness tests, and hallucination rates exceeded 74\% in widely used models. By converting LLM safety evaluation for health from a static checklist into a living adversarial audit, DAS provides a scalable framework for surfacing latent risks before such systems are deployed in consumer-facing health assistants, clinician-facing tools, and broader healthcare workflows. Code is available at https://github.com/JZPeterPan/DAS-Medical-Red-Teaming-Agents.

URL PDF HTML 收藏
2511.01680 2026-07-16 econ.EM cs.LG 版本更新

Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach

从非结构化数据中做出可解释的发现:一种高维多重假设检验方法

Jacob Carlson

机构 * Harvard University(哈佛大学)

AI总结 本文针对社会科学家利用非结构化数据获取实证见解的需求,提出通用灵活框架。借AI可解释性方法映射数据到概念嵌入,计算统计量检验假设,经选择性推断得出发现,并生成评估自然语言描述,具有低自由度、鲁棒性等优点。

详情
AI中文摘要

社会科学家越来越多地转向非结构化数据集以获取新的实证见解,如估计从文本、音频或视频数据得出的定量测量的描述性统计或因果效应。在许多情况下,无监督分析是主要关注点,因为研究者不想(或不能)手动预先指定非结构化数据的所有重要方面来测量,他们对“发现”感兴趣。本文提出了一个通用且灵活的框架,以统计原则的方式从非结构化数据中进行此类发现。该框架利用人工智能可解释性文献中的最新方法,将非结构化数据点映射到高维、稀疏且可解释的“概念嵌入”;从这些概念嵌入中计算统计量以检验可解释的、逐个概念的假设;使用高维中心极限理论的新结果验证的算法对这些假设进行选择性推断,产生一个选定的集合(“发现”);并生成和评估这些发现的人类可解释的自然语言描述。所提出的框架研究者自由度低,对数据窥探和其他选择后推断问题具有鲁棒性,并便于快速且廉价的敏感性分析和复制。还探索了其在实证经济学中非结构化数据的近期描述性和因果分析中的应用。

英文摘要

Social scientists are increasingly turning to unstructured datasets to unlock new empirical insights, e.g., estimating descriptive statistics of or causal effects on quantitative measures derived from text, audio, or video data. In many settings, unsupervised analysis is of primary interest, in that the researcher does not want to (or cannot) manually pre-specify all important aspects of the unstructured data to measure; they are interested in "discovery." This paper proposes a general and flexible framework for pursuing such discovery from unstructured data in a statistically principled way. The framework leverages recent methods from the literature on AI interpretability to map unstructured data points to high-dimensional, sparse, and interpretable "concept embeddings"; computes statistics from these concept embeddings for testing interpretable, concept-by-concept hypotheses; performs selective inference on these hypotheses using algorithms validated by new results in high-dimensional central limit theory, producing a selected set ("discoveries"); and both generates and evaluates human-interpretable natural language descriptions of these discoveries. The proposed framework has few researcher degrees of freedom, is robust to data snooping and other post-selection inference concerns, and facilitates fast and inexpensive sensitivity analysis and replication. Applications to recent descriptive and causal analyses of unstructured data in empirical economics are explored.

URL PDF HTML 收藏
2607.12631 2026-07-15 cs.CL cs.AI 新提交

Can Induced Emotion Bias LLM Behaviors in Sequential Decision Making?

诱导情绪会影响大语言模型在序列决策中的行为吗?

Minh Khoi Ho, Zihao Zhu, Runchuan Zhu, Levina Li, Zhiwen Fan, Zhangyang Wang, Junyuan Hong

机构 * MBZUAI(穆罕默德·本·扎耶德人工智能大学) Texas A&M University(德州农工大学) National University of Singapore(新加坡国立大学) UCLA(加州大学洛杉矶分校) University of Texas at Austin(德克萨斯大学奥斯汀分校) Mass General Hospital(麻省总医院) Harvard Medical School(哈佛医学院)

AI总结 研究探讨诱导情绪对大语言模型在序列决策中行为的影响,采用爱荷华赌博任务结合情绪诱导程序,发现诱导情绪平均不显著影响其决策动态,但愤怒有条件地影响决策,揭示了与人类行为的差异,为相关研究提供工具。

详情
AI中文摘要

随着大语言模型(LLMs)越来越多地在高风险领域中作为自主智能体部署,了解可能调节其决策的上下文因素变得至关重要。虽然LLMs经过训练以感知并与用户情绪产生共鸣,但诱导情绪是否会影响其序列决策仍不清楚。我们使用爱荷华赌博任务(IGT)(一种研究不确定性下决策的经典心理学范式)并结合基于想象的情绪诱导程序来研究这个问题。首先通过确认LLMs能从上下文中感知强烈且可区分的情绪,以及LLM智能体能以类似人类的速度从序列交互中学习,验证了该范式的可行性。在经过验证的设置下,我们发现,与人类不同,平均而言诱导情绪不会显著影响LLM智能体的决策动态。然而,愤怒的影响是有条件的:诱导愤怒会使LLM智能体对错误决策的惩罚不太敏感,并且在游戏早期,愤怒会降低探索,使决策过早锁定在少数选择上。这些发现揭示了诱导情绪对LLM决策与人类行为相比的细微但不同的影响,并为未来关于LLM智能体情感调节的研究提供了一种工具。

英文摘要

As Large Language Models (LLMs) are increasingly deployed as autonomous agents in high-stakes domains, understanding contextual factors that may modulate their decision-making becomes critical. While LLMs are trained to perceive and resonate with users' emotions, it remains unclear whether induced emotion can influence their sequential decision-making. We investigate this question using the Iowa Gambling Task (IGT), a classic psychological paradigm for studying decision-making under uncertainty, combined with an imagination-based emotion induction procedure. We first validate the feasibility of this paradigm by confirming that LLMs can sense strong, distinguishable emotions from context and that LLM agents can learn from sequential interactions in a human-like pace. With the validated setup, we find that, different from humans, induced emotion does not significantly bias the decision dynamics of LLM agents on average. However, the effects of anger are conditioned: inducing anger makes LLM agents less sensitive to penalties for bad decisions, and in early stages of the game, anger can lower exploration, locking decisions into a few choices early. These findings reveal the subtle yet distinct effects of induced emotion on LLM decision-making compared to human behavior, and provide a tool for future research on affective modulation of LLM agents.

URL PDF HTML 收藏
2606.12346 2026-07-15 cs.CV cs.AI cs.LG 版本更新

Atlas H&E-TME: Scalable AI-Based Tissue Profiling at Expert Pathologist-Level Accuracy

Atlas H&E-TME:基于AI的可扩展组织分析,达到专家病理学家级别的准确性

Kai Standvoss, Miriam Hägele, Rosemarie Krupar, Julika Ribbat-Idel, Jennifer Altschüler, Gerrit Erdmann, Hans Pinckaers, Evelyn Ramberger, Madleen Drinkwitz, Ádám Nárai, Alexander Möllers, Katja Lingelbach, Sebastian Kons, Lukas Hönig, Recepcan Adigüzel, Joana Baião, Alberto Megina Gonzalo, Marius Teodorescu, Marie-Lisa Eich, Paolo Chetta, Shakil Merchant, Verena Aumiller, Simon Schallenberg, Andrew Norgan, Klaus-Robert Müller, Lukas Ruff, Maximilian Alber, Frederick Klauschen

机构 * Aignostics, Germany(Aignostics,德国) Institute of Pathology, Charité – Universitätsmedizin Berlin, Germany(柏林夏里特医学院病理学研究所) Berlin Institute of Health, Charité – Universitätsmedizin Berlin, Germany(柏林夏里特医学院柏林健康研究所) Massachusetts General Hospital, Department of Pathology, Harvard Medical School, Boston, MA, US(哈佛医学院麻省总医院病理学系) Department of Laboratory Medicine and Pathology, Mayo Clinic, Rochester, MN, US(梅奥诊所检验医学与病理学系) Machine Learning Group, Technische Universität Berlin, Germany(柏林工业大学机器学习组) BIFOLD – Berlin Institute for the Foundations of Learning and Data, Germany(柏林学习与数据基础研究所) Department of Artificial Intelligence, Korea University, Republic of Korea(高丽大学人工智能系) Max-Planck Institute for Informatics, Germany(马克斯·普朗克信息学研究所) German Cancer Research Center (DKFZ) & German Cancer Consortium (DKTK), Berlin & Munich Partner Sites, Germany(德国癌症研究中心及德国癌症联盟柏林和慕尼黑合作站点) Institute of Pathology, Ludwig-Maximilians-Universität München, Germany(慕尼黑大学病理学研究所) Bavarian Cancer Research Center (BZKF), Germany(巴伐利亚癌症研究中心)

AI总结 提出Atlas H&E-TME系统,利用病理基础模型预测组织质量、区域和细胞类型,通过IHC共识验证和20万+注释基准,在多种癌症中达到或超越病理学家水平。

详情
AI中文摘要

苏木精和伊红(H&E)染色是组织病理学的基石,然而对H&E全切片图像(WSI)进行可扩展的定量分析仍然是计算病理学中的核心挑战。我们提出了Atlas H&E-TME,这是一个基于Atlas病理基础模型家族的AI系统,可预测多种癌症类型的组织质量、组织区域和细胞类型标签,在细胞级分辨率下每张切片产生超过4,500个定量读数。验证此类系统的关键挑战在于克服H&E-only金标准固有的形态模糊性,以及依赖免疫组织化学(IHC)等模态的更可靠参考的可扩展性有限。我们通过一个双重验证框架解决了这一问题,该框架将生物学深度的基础与技术及形态学的广度相结合。在深度方面,我们提出了一种IHC引导的多病理学家共识协议,该协议显著提高了相较于传统H&E-only注释的评分者间一致性。这产生了一个分子学基础的参考,我们据此比较Atlas H&E-TME和仅使用H&E的病理学家。在广度方面,我们在超过20万个高置信度H&E-only病理学家注释上对Atlas H&E-TME进行了基准测试,这些注释涵盖1,500多个病例,跨越八种癌症类型及其最常见的转移部位,亚型覆盖每种癌症类型>90%的临床病例,来自25个以上来源和8种以上扫描仪型号。与IHC引导的共识相比,Atlas H&E-TME达到或超过了病理学家仅使用H&E的性能,并在这一广泛的形态学和技术范围内一致且稳健地泛化。通过这种方式,Atlas H&E-TME将H&E切片——病理学中最普遍的数据——转化为一个可扩展的、定量的肿瘤及其微环境窗口,为转化和临床研究中下一代基于组织的生物标志物奠定了基础。

英文摘要

Hematoxylin and eosin (H&E) staining is the cornerstone of histopathology, yet scalable, quantitative analysis of H&E whole-slide images (WSIs) remains a central challenge in computational pathology. We present Atlas H&E-TME, an AI-based system built on the Atlas family of pathology foundation models that predicts tissue quality, tissue region, and cell type labels across multiple cancer types, yielding over 4,500 quantitative readouts per slide at cell-level resolution. A key challenge to validating such systems is overcoming morphological ambiguity inherent to H&E-only ground truth and the limited scalability of more informed references drawing on modalities such as immunohistochemistry (IHC). We address this with a dual validation framework combining biologically grounded depth with technical and morphological breadth. For depth, we propose an IHC-informed multi-pathologist consensus protocol that substantially improves inter-rater agreement over conventional H&E-only annotation. This yields a molecularly grounded reference against which we compare Atlas H&E-TME and pathologists working from H&E alone. For breadth, we benchmark Atlas H&E-TME on over 200,000 high-confidence H&E-only pathologist annotations across 1,500+ cases spanning eight cancer types and their most common metastatic sites, with subtypes covering >90% of clinical cases per cancer type, drawn from 25+ sources and 8+ scanner models. Benchmarked against the IHC-informed consensus, Atlas H&E-TME matches or exceeds pathologist H&E-only performance and generalizes consistently and robustly across this broad morphological and technical scope. In doing so, Atlas H&E-TME turns the H&E slide -- the most ubiquitous data in pathology -- into a scalable, quantitative window into the tumor and its microenvironment, laying a foundation for the next generation of tissue-based biomarkers in translational and clinical research.

URL PDF HTML 收藏
2605.12765 2026-07-15 cs.LG 版本更新

Inference-Time Machine Unlearning via Gated Activation Redirection

推理时的机器去学习 via 门控激活重定向

Vinícius Conte Turani, Otávio Parraga, João Vitor Boer Abitante, Kristen K. Arguello, Joana Pasquali, Ramiro N. Barros, Flavio du Pin Calmon, Christian Mattjie, Rodrigo C. Barros, Lucas S. Kupssinskü

机构 * MALTA, Machine Learning Theory and Applications Lab, PUCRS, Porto Alegre, Brazil(MALTA机器学习理论与应用实验室,PUCRS,波士顿-阿尔格雷,巴西) Harvard University(哈佛大学) Kunumi Institute, Brazil(库努米研究所,巴西)

AI总结 本文提出了一种无需训练和梯度的机器去学习方法GUARD-IT,通过在推理时依赖输入的激活引导来消除特定数据集的影响,同时保持模型性能,且在量化部署下仍有效。

详情
AI中文摘要

大型语言模型会记住大量训练数据,这引发了隐私、版权侵犯和安全方面的担忧。机器去学习旨在在不改变模型性能的情况下移除特定遗忘集的影响,理想上近似于从头重新训练模型而不包含遗忘集。现有方法通过梯度基方法更新模型参数来实现这一目标。然而,这些更新计算成本高,导致不可逆的权重变化,并在模型量化部署时性能下降。一种最近的替代方法是激活工程,在推理期间更改激活以引导模型行为。尽管绕过了权重编辑,但朴素的激活引导会引入自身的问题,因为单一的全局引导向量对每个输入应用相同的干预,导致模型行为的意外变化。我们引入了推理时的机器去学习 via 门控激活重定向(GUARD-IT),这是一种训练和梯度自由的方法,通过在推理时依赖输入的激活引导来实现去学习。所得到的干预作为残差流中的规范保持旋转应用,不改变模型权重。在TOFU和MUSE上的实验表明,GUARD-IT在三个模型规模上匹配或超过了12种基于梯度的基线方法,是唯一一个在所有设置中同时保持效用、抑制记忆和避免灾难性崩溃的方法。GUARD-IT进一步支持无需重新训练的连续去学习,并在参数编辑方法会退化的量化场景下仍有效。

英文摘要

Large Language Models memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety. Machine unlearning seeks to remove the influence of a targeted forget set while preserving model performance, ideally approximating a model retrained from scratch without the forget set. Existing approaches aim to achieve this by updating model parameters via gradient-based methods. However, these updates are computationally expensive, lead to irreversible weight changes, and degrade when the model is quantized for deployment. A recent alternative to changing model weights is activation engineering, where activations are changed during inference to steer model behavior. Despite circumventing weight editing, naive activation steering introduces its own failure modes, as a single global steering vector applies the same intervention to every input, leading to unintended changes in model behavior. We introduce Inference-Time Unlearning via Gated Activation Redirection (GUARD-IT), a training- and gradient-free method that unlearns via input-dependent activation steering at inference time. The resulting intervention is applied as a norm-preserving rotation in the residual stream, leaving model weights untouched. Experiments on TOFU and MUSE show that GUARD-IT matches or exceeds 12 gradient-based baselines across three model scales, while being the only method to simultaneously preserve utility, suppress memorization, and avoid catastrophic collapse across all settings. GUARD-IT further supports continual unlearning without retraining, and remains effective under quantization, a scenario in which parameter-editing methods degrade.

URL PDF HTML 收藏
2603.01568 2026-07-15 cs.LG cs.CV cs.IT math.IT q-bio.NC 版本更新

Same Compression Principle, Different Geometry: Rate-Distortion Signatures Dissociate Biological and Artificial Visual Systems

泛化与信息权衡的率-失真签名

Leyla Roksan Caglar, Pedro A. M. Mediano, Baihan Lin

机构 * Windreich Department of AI Human Health, Icahn School of Medicine at Mount Sinai, New York, NY, USA Department of Computing, Imperial College London, London, UK Department of Psychiatry, Icahn School of Medicine at Mount Sinai, New York, NY, USA Department of Neuroscience, Icahn School of Medicine at Mount Sinai, New York, NY, USA Berkman Klein Center for Internet \& Society, Harvard University, Cambridge, MA, USA

AI总结 本文提出率-失真理论框架,通过斜率和曲率签名分析系统泛化与鲁棒性权衡,揭示生物与人工系统在RD空间中的不同表现。

详情
AI中文摘要

向新的视觉条件泛化仍然是人类和机器视觉面临的重大挑战,然而标准的鲁棒性度量在揭示系统如何在准确性与鲁棒性之间进行权衡方面提供有限的见解。我们引入了一个率-失真理论框架,将刺激-响应行为视为一个有效的通信通道,从混淆矩阵中推导出率-失真(RD)前沿,并用两个可解释的几何签名——斜率(β)和曲率(κ)——来总结每个系统,这些签名捕捉了准确率-鲁棒性权衡的边际成本和突变性。将此框架应用于人类心理物理学和18个深度视觉模型在受控图像扰动下的表现,我们比较了不同模型架构和训练制度下的泛化几何结构。我们发现,生物和人工系统都遵循一种损失性压缩原则,但系统地占据RD空间的不同区域。特别是,人类表现出更平滑、更灵活的权衡,而现代深度网络即使在匹配的准确性下也处于更陡峭且更脆弱的区域。在不同的训练制度下,鲁棒性训练会引发系统但可区分的β/κ变化,揭示了某些情况下改进的鲁棒性或准确性并不转化为更像人类的泛化几何结构。这些结果表明,RD几何提供了紧凑且模型无关的视角,用于比较系统之间的泛化行为,超越了标准的准确性度量。

英文摘要

Efficient coding theory predicts that biological perceptual systems compress sensory input optimally under resource constraints, with the systematic structure of errors reflecting the geometry of that compression. Here we operationalize this principle using rate-distortion theory (RDT) to characterize how any system - biological or artificial - trades representational fidelity for informational efficiency. Treating stimulus-response behavior as an effective communication channel, we infer rate-distortion (RD) frontiers directly from confusion matrices and summarize each system with three geometric signatures: slope (beta), curvature (kappa), and area under the RD curve (AUC), capturing the marginal cost, abruptness, and overall efficiency of the accuracy-compression trade-off respectively. Applying this framework to human psychophysical data and 18 deep vision models across 12 families of controlled image perturbations at graded severities, we find that both biological and artificial systems follow a common lossy-compression principle but occupy systematically different regions of RD space. Humans exhibit smooth, flexible trade-offs characteristic of near-optimal efficient coding, while deep networks operate in steeper, more brittle regimes even at matched accuracy, with geometry dissociable from performance across training regimes. Critically, behavioral RD signatures track internal representational geometry, evidenced by the behaviorally inferred compression structure correlating with internal representational dissimilarity across all models. These results establish RD geometry as a compact diagnostic of perceptual compression strategy that recovers mechanistically interpretable structure in internal representations from behavioral input alone and extends naturally to the direct characterization of compression geometry in neural population activity.

URL PDF HTML 收藏
2512.01241 2026-07-15 cs.CY cs.AI 版本更新

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

首先,不伤害:迈向临床安全的大语言模型

David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj, Priyank Jain, Jessica Tran, Matthew Gwiazdon, Arjun Rustagi, Jenelle Jindal, Jacob M. Koshy, Vinay Kadiyala, Anup Agarwal, Bassman Tappuni, Brianna French, Sirus Jesudasen, Christopher V. Cosgriff, Rebanta Chakraborty, Jillian Caldwell, Susan Ziolkowski, David J. Iberri, Robert Diep, Rahul S. Dalal, Kira L. Newman, Kristin Galetta, J. Carl Pallais, Nancy Wei, Kathleen M. Buchheit, David I. Hong, Vartan Pahalyants, Ernest Y. Lee, Allen Shih, Tamara B. Kaplan, Vishnu Ravi, Sarita Khemani, Thomas A. Buckley, April S. Liang, Daniel Shirvani, Advait Patil, Nicholas Marshall, Kanav Chopra, Joel Koh, Adi Badhwar, Anastasia Perez, Austin J. Schoeffler, Mahbuba Tusty, Chase M. Walton, Liam G. McCoy, David J. H. Wu, Yingjie Weng, Sumant Ranji, Kevin Schulman, Nigam H. Shah, Jason Hom, Arnold Milstein, Arjun K. Manrai, Adam Rodman, Jonathan H. Chen, Ethan Goh

机构 * Harvard Combined Dermatology Program(哈佛联合皮肤科项目) Department of Dermatology, Mass General Brigham(麻省总医院皮肤科) Harvard Medical School(哈佛医学院) Stanford Center for Biomedical Informatics Research(斯坦福生物医学信息学研究中心) Stanford University(斯坦福大学) Division of Hospital Medicine, Department of Medicine, Stanford University School of Medicine(斯坦福大学医学院医院医学科) Department of Medicine, Cambridge Health Alliance(剑桥健康联盟医学科) Beth Israel Deaconess Hospital–Plymouth(贝塞斯达德acons医院-普利茅斯) Department of Medicine, University of California, San Francisco(加州大学旧金山分校医学科) Department of Neurology, Stanford University School of Medicine(斯坦福大学医学院神经科) Department of Medicine, Beth Israel Deaconess Medical Center(贝塞斯达德acons医学中心医学科) Division of Cardiology, Department of Medicine, Cambridge Health Alliance(剑桥健康联盟心脏病科) Department of Cardiovascular Medicine, Summa Health System(Summa健康系统心血管医学科) Division of Allergy, Pulmonary, and Critical Care Medicine, Department of Medicine, University of Wisconsin-Madison(威斯康星大学麦迪逊分校医学科过敏、呼吸科和危重医学科) Division of Pulmonary and Critical Care Medicine, Department of Medicine, Massachusetts General Hospital(麻省总医院呼吸科和危重医学科) Center for Immunology and Inflammatory Diseases, Department of Medicine, Massachusetts General Hospital(麻省总医院免疫和炎症疾病中心) Broad Institute of MIT and Harvard(MIT和哈佛Broad研究所) Division of Pulmonary, Critical Care, and Sleep Medicine, Cambridge Health Alliance(剑桥健康联盟呼吸科、危重医学科和睡眠医学科)

AI总结 提出NOHARM基准,包含1100个初级到专科咨询案例,评估28个LLM的医疗建议安全性,发现高达22.6%的案例存在严重危害风险,其中遗漏错误占80%以上。

详情
AI中文摘要

大语言模型(LLM)被医生和患者常规用于医疗建议,但其临床安全性特征仍不明确。我们提出NOHARM(医学风险评估的众多选项危害评估),一个包含1100个初级保健到专科咨询案例的基准,用于衡量LLM生成的医疗建议的危害频率和严重程度。NOHARM涵盖10个专科,包含4249个临床管理选项的12747个专家注释。在28个LLM中,建议在高达22.6%的案例中具有严重危害潜力,其中遗漏错误占严重错误的80%以上。在一项涉及101名全科医生的随机试验中,AI辅助显著提高了人类基准表现,但医生远未实现AI工具的潜力,经常忽略AI提出的重要建议。安全性表现与通用智能和医学知识基准在整个模型范围内相关,但在前沿模型上解耦。尽管在现有评估中表现强劲,广泛使用的AI模型可能以非平凡的比例产生具有严重危害潜力的医疗建议,凸显了明确测量临床安全性的重要性。

英文摘要

Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized. We present NOHARM (Numerous Options Harm Assessment for Risk in Medicine), a 1,100-task benchmark of primary care-to-specialist consultation cases to measure the frequency and severity of potentially harmful errors from LLM-generated medical consultation recommendations. NOHARM covers 10 specialties, with 12,747 expert annotations for 4,249 clinical management options. Across 20 notable LLMs and 4 widely used retrieval-augmented generation (RAG) clinical AI tools, direct application of recommendations carried potential for severe harm in up to 24.6% of cases, with errors of omission accounting for more than 80% of severe errors. Harm potential was not uniform across systems, with clinical AI tools outperforming generalist LLMs, and multi-agent AI teaming further improving performance in generalist models. In a randomized study of 101 U.S.-licensed generalist physicians, AI assistance improved physician performance compared to conventional resources. However, AI-assisted physicians frequently omitted valuable AI-generated recommendations and still scored lower than many AI systems alone. Had those recommendations been incorporated, combined human-AI responses would have outperformed both the human and AI system as used, suggesting complementary strengths and unrealized potential in human-AI teaming. Collectively, these results show that despite strong performance on medical knowledge benchmarks, widely used AI tools can produce medical consultation advice with the potential for severe harm, and highlight the need for explicit measurement of clinical safety. The benchmark and leaderboard are publicly available to support ongoing evaluation and improvement of AI systems used for clinical care.

URL PDF HTML 收藏
2512.04144 2026-07-15 cs.AI 版本更新

RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories

RippleBench: 利用现有知识库捕捉涟漪效应

Roy Rinberg, Usha Bhalla, Igor Shilov, Flavio P. Calmon, Rohit Gandikota

机构 * Harvard University(哈佛大学) Imperial College London(伦敦帝国学院) Northeastern University(东北大学)

AI总结 提出RippleBench-Maker自动管道,从知识库检索语义邻居生成选择题,评估八种遗忘方法在Llama3-8B-Instruct上的涟漪效应,发现准确率下降随语义距离衰减且跨模型一致。

详情
AI中文摘要

针对语言模型的目标干预,如遗忘或模型编辑,旨在修改特定信息,但其效果往往传播到相关的、非预期的领域(例如,删除病毒学内容可能降低对过敏任务的性能);这些副作用通常被称为涟漪效应。我们引入RippleBench-Maker,一个自动管道,从知识库中检索任何源概念的语义邻居,并生成不同语义距离的多选题。我们使用WikiRAG(一个基于英文维基百科的开源RAG系统)实例化该框架,构建RippleBench-WMDP-Bio(584个种子主题,352,961个问题),并在Llama3-8B-Instruct上评估八种遗忘方法。所有八种方法在遗忘目标附近准确率下降最大,并随语义距离衰减,每种方法具有不同的传播曲线。我们在Mistral-7B、Zephyr-7B和Yi-34B上复现了这些发现;跨模型的差值曲线几乎相同,表明涟漪效应是遗忘方法的属性而非基础模型。我们通过一项包含四个实验的Mechanical Turk研究(5,200+次响应,61名工作者)验证了所有主要管道阶段。我们发布所有代码、数据和基础设施。

英文摘要

Targeted interventions on language models, such as unlearning or model editing, aim to modify specific information, but their effects often propagate to related, unintended areas (e.g., removing virology content may degrade performance on allergies); these side-effects are commonly referred to as the ripple effect. We introduce RippleBench-Maker, an automatic pipeline that retrieves semantic neighbors of any source concept from a knowledge repository and generates multiple-choice questions at varying semantic distances. We instantiate this framework using WikiRAG, an open-source RAG system over English Wikipedia, to construct RippleBench-WMDP-Bio (584 seed topics, 352,961 questions), and evaluate eight unlearning methods on Llama3-8B-Instruct. All eight exhibit accuracy drops that are largest near the unlearned target and decay with semantic distance, each with a distinct propagation profile. We replicate these findings across Mistral-7B, Zephyr-7B, and Yi-34B; cross-model delta curves are nearly identical, suggesting ripple effects are a property of the unlearning method rather than the base model. We validate all major pipeline stages using a four-experiment Mechanical Turk study (5,200+ responses, 61 workers). We release all code, data, and infrastructure.

URL PDF HTML 收藏
2502.15131 2026-07-15 math.ST cs.LG stat.ME stat.ML stat.TH 版本更新

Optimal and Provable Calibration in High-Dimensional Binary Classification: Angular Calibration and Platt Scaling

高维二分类中的最优且可证明的校准:角度校准与Platt缩放

Yufan Li, Pragya Sur

机构 * Harvard University(哈佛大学)

AI总结 针对高维高斯特征下的线性二分类器,提出基于估计权重与真实权重夹角的角度校准方法,证明其可校准且唯一Bregman最优,并揭示Platt缩放在高维下收敛于该最优解。

详情
AI中文摘要

我们研究校准形如 $\sigma(\hat{w}^\top x)$ 的线性二分类器的基本问题,其中特征向量 $x$ 服从高斯分布,$\sigma$ 是链接函数,$\hat{w}$ 是真实线性权重 $w^\star$ 的估计量。通过与非信息性的 $\textit{机会分类器}$ 插值,我们构建了一个良好校准的预测器,其插值权重取决于估计量 $\hat{w}$ 与真实线性权重 $w_\star$ 之间的夹角 $\angle(\hat{w}, w_\star)$。我们证明,在样本量和特征量均以可比速率发散的高维机制下,这种角度校准方法可证明是良好校准的。夹角 $\angle(\hat{w}, w_\star)$ 可以一致地估计。此外,所得预测器是唯一 $\textit{Bregman最优}$ 的,即在合适的校准预测器类中最小化与真实标签分布的Bregman散度。我们的工作是首个在高维下同时满足校准和最优性可证明的校准策略。此外,我们识别了经典Platt缩放预测器收敛到我们的Bregman最优校准解的条件。因此,Platt缩放在高维下也继承了这些理想性质。

英文摘要

We study the fundamental problem of calibrating a linear binary classifier of the form $σ(\hat{w}^\top x)$, where the feature vector $x$ is Gaussian, $σ$ is a link function, and $\hat{w}$ is an estimator of the true linear weight $w^\star$. By interpolating with a noninformative $\textit{chance classifier}$, we construct a well-calibrated predictor whose interpolation weight depends on the angle $\angle(\hat{w}, w_\star)$ between the estimator $\hat{w}$ and the true linear weight $w_\star$. We establish that this angular calibration approach is provably well-calibrated in a high-dimensional regime where the number of samples and features both diverge, at a comparable rate. The angle $\angle(\hat{w}, w_\star)$ can be consistently estimated. Furthermore, the resulting predictor is uniquely $\textit{Bregman-optimal}$, minimizing the Bregman divergence to the true label distribution within a suitable class of calibrated predictors. Our work is the first to provide a calibration strategy that satisfies both calibration and optimality properties provably in high dimensions. Additionally, we identify conditions under which a classical Platt-scaling predictor converges to our Bregman-optimal calibrated solution. Thus, Platt-scaling also inherits these desirable properties provably in high dimensions.

URL PDF HTML 收藏
2312.17670 2026-07-15 cs.CV cs.LG q-bio.QM q-bio.TO 版本更新

The TopCoW Challenge -- Topology-Aware Circle of Willis Segmentation for CT and MR Angiography

TopCoW挑战——用于CT和MR血管造影的拓扑感知Willis环分割

Kaiyuan Yang, Fabio Musio, Yihui Ma, Norman Juchler, Johannes C. Paetzold, Rami Al-Maskari, Luciano Höher, Hongwei Bran Li, Ibrahim Ethem Hamamci, Anjany Sekuboyina, Suprosanna Shit, Houjing Huang, Chinmay Prabhakar, Ezequiel de la Rosa, Bastian Wittmann, Diana Waldmannstetter, Florian Kofler, Fernando Navarro, Martin J. Menten, Ivan Ezhov, Daniel Rueckert, Iris N. Vos, Ynte M. Ruigrok, Birgitta K. Velthuis, Hugo J. Kuijf, Pengcheng Shi, Wei Liu, Ting Ma, Maximilian R. Rokuss, Yannick Kirchhoff, Fabian Isensee, Klaus Maier-Hein, Chengcheng Zhu, Huilin Zhao, Philippe Bijlenga, Julien Hämmerli, Catherine Wurster, Laura Westphal, Jeroen Bisschop, Elisa Colombo, Hakim Baazaoui, Hannah-Lea Handelsmann, Andrew Makmur, James Hallinan, Amrish Soundararajan, Benedikt Wiestler, Jan S. Kirschke, Evamaria O. Riedel, Roland Wiest, Emmanuel Montagnon, Laurent Letourneau-Guillon, Kwanseok Oh, Dahye Lee, Orhun Utku Aydin, Adam Hilbert, Jana Rieger, Dimitrios Rallios, Satoru Tanioka, Alexander Koch, Dietmar Frey, Abdul Qayyum, Moona Mazher, Steven Niederer, Nico Disch, Julius C. Holzschuh, Dominic LaBella, Francesco Galati, Daniele Falcetta, Maria A. Zuluaga, Chaolong Lin, Haoran Zhao, Zehan Zhang, Minghui Zhang, Xin You, Hanxiao Zhang, Guang-Zhong Yang, Yun Gu, Sinyoung Ra, Jongyun Hwang, Hyunjin Park, Junqiang Chen, Marek Wodzinski, Henning Müller, Nesrin Mansouri, Florent Autrusseau, Cansu Yalcin, Rachika E. Hamadache, Clara Lisazo, Joaquim Salvi, Adrià Casamitjana, Xavier Lladó, Uma Maria Lal-Trehan Estrada, Valeriia Abramova, Luca Giancardo, Arnau Oliver, Paula Casademunt, Adrian Galdran, Matteo Delucchi, Oscar Camara, Jialu Liu, Haibin Huang, Yue Cui, Zehang Lin, Yusheng Liu, Shunzhi Zhu, Tatsat R. Patel, Adnan H. Siddiqui, Vincent M. Tutino, Maysam Orouskhani, Huayu Wang, Mahmud Mossa-Basha, Yuki Sato, Sven Hirsch, Susanne Wegener, Bjoern Menze

机构 * Department of Quantitative Biomedicine, University of Zurich, Zurich, Switzerland Institute of Computational Life Sciences, Zurich University of Applied Sciences (ZHAW), Waedenswil, Switzerland Department of Neuroradiology, University Hospital of Zurich, Zurich, Switzerland Department of Neurosurgery, Zhongnan Hospital of Wuhan University, Wuhan, China Department of Radiology at Weill Cornell Medicine, Cornell University, New York, USA Institute for Tissue Engineering School of Computation, Information Technology, Technical University of Munich, Germany Athinoula A. Martinos Center for Biomedical Imaging, Harvard Medical School, Boston, USA School of Medicine Health, TUM Klinikum, Technical University of Munich, Germany Munich Center for Machine Learning, Munich, Germany Department of Computing, Imperial College London, London, UK Image Sciences Institute, UMC Utrecht, Utrecht, The Netherlands Department of Neurology Neurosurgery, University Medical Center Utrecht, Utrecht, The Netherlands Department of Radiology, University Medical Center Utrecht, Utrecht, The Netherlands Electronic \& Information Engineering School, Harbin Institute of Technology (Shenzhen), China Peng Cheng Laboratory, Shenzhen, China Division of Medical Image Computing, German Cancer Research Center (DKFZ), Heidelberg, Germany Faculty of Mathematics Computer Science, Heidelberg University, Germany Helmholtz Imaging, German Cancer Research Center, Heidelberg, Germany Data Science School for Health, Karlsruhe/Heidelberg, Germany Learning Group, Department of Radiation Oncology, Heidelberg University Hospital Department of Radiology, University of Washington, Seattle, WA, USA Department of Radiology, Ren Ji Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China Department of Clinical Neurosciences, Division of Neurosurgery, Geneva University Hospitals, Geneva, Switzerland Department of Neurology, University Hospital of Zurich, Zurich, Switzerland Department of Physiology, University of Toronto, Canada Department of Neurosurgery, University Hospital of Zurich, Zurich, Switzerland Department of Diagnostic Imaging, National University Hospital, Singapore University of Chicago, USA Department of Diagnostic Interventional Neuroradiology, University Hospital Berne University of Berne, Berne, Switzerland Centre de Recherche du Centre Hospitalier de l’Université de Montréal (CRCHUM), Montréal, Québec, Canada DEEPNOID Inc., Seoul, South Korea Department of Artificial Intelligence, Korea University, Seoul, South Korea Charité Lab for AI in Medicine (CLAIM), Charité Universitätsmedizin Berlin, Berlin, Germany Lung Institute, Faculty of Medicine, Imperial College London, London, UK Centre for Medical Image Computing, Department of Computer Science, University College London, London, UK Department of Radiation Oncology, Duke University Medical Center, Durham, NC, USA Institute of Medical Technology, Peking University Health Science Center, Beijing, China Hangzhou Genlight MedTech Co., Ltd., China Institute of Medical Robotics, Shanghai Jiao Tong University, Shanghai, China Department of Automation, Shanghai Jiao Tong University, Shanghai, China Department of Artificial Intelligence, Sungkyunkwan University, Seoul, South Korea Department of Electrical Computer Engineering, Sungkyunkwan University, Seoul, South Korea Shanghai MediWorks Precision Instruments Co., Ltd., China Institute of Informatics, HES-SO Valais-Wallis, Switzerland Department of Measurement Electronics, AGH University of Krakow, Poland Laboratoire de Thermique et Energie de Nantes (LTeN), Université Nantes, Polytech’Nantes, Nantes, France Research Institute of Computer Vision Center for Precision Health, McWilliams School of Biomedical Informatics, University of Texas Health Science Center at Houston, USA Physense, BCN-Medtech, Department of Communication Information Technologies, Universitat Pompeu Fabra, Barcelona, Spain Department of Mathematical Modeling Machine Learning, University of Zurich, Zurich, Switzerland Laboratory of Brain Atlas Brain-inspired Intelligence, Institute of Automation, Chinese Academy of Sciences, Beijing, China School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China School of Computer Information Engineering, Xiamen University of Technology, Xiamen, China Vascular Research Center, University at Buffalo, NY, USA Department of Pathology Anatomical Sciences, University at Buffalo, NY, USA Department of Neurosurgery, University at Buffalo, NY, USA LPIXEL Inc., Tokyo, Japan

AI总结 组织TopCoW基准挑战,发布含125对MRA和CTA扫描的注释数据集,参与者提交CoW分割和变体分类算法,经评估,最佳算法在多任务中表现出色,证明CoW分割算法对下游临床应用有可解释性效用。

Comments Summary paper for the TopCoW Challenge: 4 figures, 1 table, and supplementary material in appendix. Accepted for publication in NEJM AI. Datasets and best-performing algorithm Dockers are available at https://zenodo.org/records/15692630 and https://zenodo.org/records/15665435

详情
AI中文摘要

Willis环(CoW)是连接大脑主要循环的重要动脉网络。其血管结构被认为会影响严重神经血管疾病的风险、严重程度和结果。然而,表征高度可变的CoW解剖结构仍然是一项人工且耗时的专家任务。CoW通常通过磁共振血管造影(MRA)和计算机断层血管造影(CTA)这两种非侵入性血管造影成像方式进行成像,但带注释的CoW解剖结构数据集很少,也没有用于比较CoW分割算法的既定基准。我们组织了TopCoW基准挑战,并发布了一个带注释的CoW数据集,其中包含来自同一患者的125对MRA和CTA扫描。使用虚拟现实技术创建了13个血管成分的体素级注释,并由临床专家进行了验证。参与者提交了CoW分割和变体分类算法,我们在包含来自五个以上中心的226次扫描的内部和外部测试集上进行了评估。该基准包括体素级分割、CoW成分检测、CoW变体分类和两个临床应用任务。我们收到了来自六大洲250多名参与者的提交。表现最佳的团队在几乎所有测试集中,CoW分割的Dice分数超过90%,关键血管成分检测的F1分数超过80%,CoW变体分类中的平衡准确率超过70%。最佳算法还通过准确分类胎儿型大脑后动脉并定位与CoW解剖结构相关的动脉瘤,支持了临床相关的下游任务。这个基准证明了CoW分割算法在一些具有可解释性的下游临床应用中的效用。

英文摘要

The Circle of Willis (CoW) is an important network of arteries connecting major circulations of the brain. Its vascular architecture is believed to influence the risk, severity, and outcome of serious neurovascular diseases. However, characterizing the highly variable CoW anatomy remains a manual and time-consuming expert task. The CoW is commonly imaged by two non-invasive angiographic imaging modalities, magnetic resonance angiography (MRA) and computed tomography angiography (CTA), yet few datasets with annotated CoW anatomy exist, and there have been no established benchmarks for comparing CoW segmentation algorithms. We organized the TopCoW benchmark challenge alongside the release of an annotated CoW dataset with 125 paired MRA and CTA scans from the same patients. Voxel-level annotations for 13 vessel components were created using virtual reality technology and verified by clinical experts. Participants submitted algorithms for CoW segmentation and variant classification, which we evaluated on internal and external test sets comprising 226 scans from over five centers. The benchmark includes voxel-level segmentation, CoW component detection, CoW variant classification, and two clinical application tasks. We received submissions from over 250 participants across six continents. Top-performing teams achieved over 90% Dice scores for CoW segmentation, over 80% F1 scores for detecting key vessel components, and over 70% balanced accuracy in CoW variant classification across nearly all test sets. The best algorithms also supported clinically relevant downstream tasks by accurately classifying fetal-type posterior cerebral arteries and localizing aneurysms in relation to CoW anatomy. This benchmark demonstrated the utility of CoW segmentation algorithms for some downstream clinical applications with explainability.

URL PDF HTML 收藏
2607.11052 2026-07-14 cs.LG cs.CL 新提交

Domain-Aware Scaling Laws Uncover Data Synergy

领域感知缩放定律揭示数据协同效应

Kimia Hamidieh, Lester Mackey, David Alvarez-Melis

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Microsoft Research(微软研究院) Harvard University(哈佛大学)

AI总结 研究语言模型预训练中的数据协同效应,利用开放权重语言模型的观测变化估计领域间协同效应,其框架提高预测精度,恢复稳定估计,经训练模型验证能正确预测性能排名。

详情
AI中文摘要

机器学习的进展通常归因于扩大模型规模和数据集大小,但数据的组成同样重要。实证研究反复表明,合并来自不同领域的数据集会产生重要的相互作用。例如,添加代码数据可提升数学推理能力,而某些混合数据会产生干扰,降低模型性能。我们将这些效应统称为数据协同效应,即多个领域的贡献超过或低于其单独贡献之和。在这项工作中,我们对语言模型预训练中的数据协同效应进行了形式化和量化。利用具有不同预训练混合的开放权重语言模型的观测变化,我们估计了直接的领域到基准协同效应(一个领域对另一个领域性能的贡献)和二阶领域间协同效应(需要多个领域共同出现的能力)。我们的框架提高了预测精度,恢复了稳定的协同效应估计。我们通过在预测的最优和预测的反最优混合数据上训练模型来验证这些估计,并确认我们的协同效应估计正确地预测了性能排名。

英文摘要

Machine learning progress is often attributed to scaling model size and dataset volume, yet the composition of data can be just as consequential. Empirical findings repeatedly show that combining datasets from different domains yields nontrivial interactions. For instance, adding code improves mathematical reasoning, while certain mixtures introduce interference that reduces model performance. We refer to these effects collectively as data synergy, where the contribution of multiple domains exceeds or falls short of the sum of their isolated contributions. In this work, we formalize and quantify data synergy in language model pretraining. Leveraging observational variation across open-weight LLMs with diverse pretraining mixtures, we estimate both direct domain-to-benchmark synergy (how one domain contributes to performance on another) and a second-order domain-domain synergy (capabilities that require co-occurrence of multiple domains). Our framework improves predictive accuracy over domain-agnostic scaling laws and recovers stable synergy estimates. We validate these estimates by training models on predicted optimal and predicted anti-optimal mixtures and confirm that our synergy estimates correctly predict performance rankings.

URL PDF HTML 收藏
2607.10892 2026-07-14 cs.RO 新提交

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer

用于多任务块推的单扩散策略控制器,具有零样本模拟到现实转移

Haitong Ma, Haldun Balim, Yang Hu, Bo Dai, Na Li

机构 * Harvard University(哈佛大学) Georgia Institute of Technology(佐治亚理工学院)

AI总结 研究旨在用强化学习从零训练单扩散策略用于多任务块推,提出含简单策略损失函数的框架,结合反向课程生成等应对探索挑战,评估其在不同条件下零样本从模拟到现实的转移能力,证明该流程有效。

Comments 8 pages, 7 figures

详情
AI中文摘要

扩散策略在通过行为克隆为机器人表示和学习复杂动作方面展现出了有前景的实证性能。本文中,我们探索使用强化学习从零开始训练扩散策略用于多任务机器人操纵。具体而言,我们旨在训练一个针对多种形状块推任务的单扩散策略。所提出的框架具有一个简单的策略损失函数,它是基于行为克隆的扩散策略训练中使用的重新加权证据下界,并且能无缝用作强化学习算法中的策略学习模块。为应对因缺乏示范而产生的探索挑战,我们纳入了反向课程生成和以目标为中心的表示。结合扩散策略的表现力,我们的设计支持在稀疏奖励模拟设置中学习多任务块推策略。我们进一步评估训练好的扩散策略在包括目标位置、块形状、块重量和表面摩擦等不同环境条件下能否零样本转移到现实世界任务,结果表明该流程在测试的变化情况下能转移到我们的现实世界块推设置中。

英文摘要

Diffusion policies have shown promising empirical performance in representing and learning complex maneuvers for robots using behavior cloning (BC). In this paper, we explore training diffusion policies from scratch using reinforcement learning (RL) for multi-task robotic manipulation. Specifically, we aim to train a single diffusion policy for block-pushing tasks with multiple shapes. The proposed framework features a simple policy loss function, which is a reweighted evidence lower bound used in BC-based diffusion policy training and can seamlessly serve as the policy learning module in RL algorithms. To address the exploration challenges arising from the absence of demonstrations, we incorporate reverse curriculum generation and objective-centric representations. Combined with the expressiveness of diffusion policies, our design supports learning of multi-task block-pushing policies in our sparse-reward simulation setting. We further evaluate whether the trained diffusion policy transfers in zero-shot to real-world tasks under varying environmental conditions including goal positions, block shapes, block weights and surface friction, providing evidence that this pipeline can transfer to our real-world block-pushing setup under the tested variations.

URL PDF HTML 收藏