arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Imperial College London(帝国理工学院)

至 收录 1191
2607.17201 2026-07-21 stat.ML cs.LG 新提交

Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning

在线强化学习中的非渐近最优策略识别保证

Joseph Lazzaro, Alessio Russo, Aldo Pacchiano

机构 * Imperial College London(伦敦帝国学院) Boston University(波士顿大学) Broad Institute of MIT and Harvard(MIT和哈佛大学Broad研究所)

AI总结 研究在线表格强化学习中的最优策略识别问题,通过为导航与停止(NaS)算法提供非渐近样本复杂度保证,揭示其样本复杂度不仅依赖特征时间,还与MDP连通性等有关,填补了相关空白。

Comments 64 pages, 2 figures

详情
AI中文摘要

在这项工作中,我们研究在线表格强化学习中的最优策略识别(BPI)问题。这是一个主动序贯假设检验问题,学习者的目标是以高置信度识别马尔可夫决策过程(MDP)中的最优策略,同时最小化预期样本复杂度。我们考虑具有确定性奖励的在线设置,智能体必须策略性地在MDP中导航以有效探索。先前文献为BPI提供了渐近最优方法,如导航与停止(NaS)算法及其变体,但现有分析仍是渐近的。我们通过为NaS提供首个非渐近样本复杂度保证来填补这一空白,表明其样本复杂度不仅取决于特征时间,还取决于基础MDP的连通性、最优特征时间的曲率以及其他依赖实例的量。我们识别出这些额外属性并明确它们对整体样本复杂度的贡献。

英文摘要

In this work we study the Best Policy Identification (BPI) problem in online, tabular Reinforcement Learning. This is an active sequential hypothesis testing problem in which the learner's objective is to identify an optimal policy in a Markov Decision Process (MDP) with high confidence, while minimizing the expected sample complexity to do so. We consider an online setting with deterministic rewards, where the agent must strategically navigate through the MDP in order to effectively explore. Previous works in the literature have provided asymptotically optimal methods for BPI, such as the Navigate and Stop (NaS) algorithm and its variants, however existing analysis remains asymptotic. In this work, we fill that gap by providing the first non-asymptotic sample complexity guarantees for NaS, showing that its sample complexity depends not only on the characteristic time, but also on the connectivity of the underlying MDP, the curvature of the optimal characteristic time, and other instance-dependent quantities. We identify these additional attributes and make explicit their contributions to the overall sample complexity.

URL PDF HTML 收藏
2607.16877 2026-07-21 eess.SP cs.LG 新提交

Hierarchical Wireless Foundation Model for Multi-Task Optimization

用于多任务优化的分层无线基础模型

Yangjing Wang, Ouya Wang, Shenglong Zhou, Geoffrey Ye Li

机构 * Department of Electrical and Electronic Engineering, Faculty of Engineering, Imperial College London(电气与电子工程系,工程学院,伦敦帝国理工学院) School of Mathematics and Statistics, Beijing Jiaotong University(数学与统计学学院,北京交通大学)

AI总结 针对下一代无线网络中人工智能技术泛化瓶颈问题,提出分层无线基础模型(WFM),通过几何感知交叉注意力耦合上下游模块,采用混合训练策略,实现多任务优化,学习高保真信道表示,降低延迟,有强大泛化能力。

Comments 15 pages, 10 figures, 3 tables

详情
AI中文摘要

下一代无线网络日益复杂,推动了人工智能融入无线通信。但现有多数研究专注单场景特定任务深度学习技术,限制了跨任务、信道条件和系统配置的泛化能力。为此提出分层无线基础模型(WFM)用于多任务优化,通过几何感知交叉注意力将上游基础信道编码器(FCE)与下游基础优化解码器(FOD)耦合。FCE经自监督掩码重建提取与任务无关的信道表示,FOD通过可微输出头生成多任务优化决策。采用混合监督到无监督训练策略克服纯监督学习性能上限,其模块化架构能高效适应未见通信任务且参数开销小。仿真结果表明,该模型能学习高保真信道表示,实现有竞争力的多任务优化性能,大幅降低优化推理延迟,对未见传播环境、变化约束参数和异构系统配置有强大泛化能力。

英文摘要

The increasing complexity of next-generation wireless networks has driven the integration of artificial intelligence (AI) into wireless communications. However, most existing studies focus on developing task-specific deep learning techniques for single scenarios, which limits their ability to generalize across diverse tasks, channel conditions, and system configurations. To address this generalization bottleneck, we propose a hierarchical wireless foundation model (WFM) for multi-task optimization. The proposed WFM couples an upstream foundation channel encoder (FCE) with a downstream foundation optimization decoder (FOD) via geometry-aware cross-attention. Specifically, the FCE extracts task-agnostic channel representations via self-supervised masked reconstruction while the FOD generates multi-task optimization decisions through differentiable output heads. Moreover, a hybrid supervised-to-unsupervised training strategy is employed to overcome the performance ceiling of purely supervised learning, and the modular architecture of the WFM enables efficient adaptation to unseen communication tasks with minimal parameter overhead. Simulation results show that the proposed WFM learns high-fidelity channel representations and achieves competitive multi-task optimization performance while substantially reducing optimization inference latency relative to numerical baselines. Furthermore, it exhibits robust generalization to unseen propagation environments, varying constraint parameters, and heterogeneous system configurations.

URL PDF HTML 收藏
2607.18195 2026-07-21 cs.CV cs.LG 新提交

Certified Training for Convolutional Perturbations

卷积扰动的认证训练

Benedikt Brückner, Alessio Lomuscio

机构 * Safe Intelligence(安全智能) Imperial College London(伦敦帝国理工学院)

AI总结 研究视觉模型受扰动问题,提出利用卷积扰动编码的认证训练方法,显著优于对抗训练,在CIFAR10上对运动模糊有超80%鲁棒准确率,且标准准确率相当。

详情
AI中文摘要

视觉模型易受运行时相机抖动引起的运动模糊等扰动影响,这阻碍了它们在关键应用中的部署。数据增强或对抗训练等方法缺乏形式安全保证。我们引入一种新颖的认证训练方法,利用卷积扰动的高效编码来训练可证明鲁棒的模型。该方法显著优于对抗训练,在CIFAR10上对合理强度运动模糊的鲁棒准确率超过80%,同时保持相当的标准准确率。

英文摘要

Vision models have been found to be susceptible to perturbations such as motion blur induced at runtime by a shaking camera. This impedes their deployment in critical applications since phenomena such as slightly blurred vision might lead to failures, for example an object detector missing objects. While methods such as data augmentation or Adversarial Training can improve empirical robustness, they lack formal safety guarantees, making it difficult to identify and mitigate hidden vulnerabilities. We introduce a novel Certified Training approach that leverages an efficient encoding of convolutional perturbations to train provably robust models. Our method significantly outperforms Adversarial Training, achieving, for example, over 80% robust accuracy against motion blur of reasonable intensity on CIFAR10 while maintaining comparable standard accuracy.

URL PDF HTML 收藏
2607.16372 2026-07-21 cs.SE cs.AI cs.LG cs.PL 新提交

AoA: Theorem Proving Agent over Abstract Syntax Tree of Redesigned Language

AoA:基于重新设计语言抽象语法树的定理证明智能体

Qiyuan Xu, Joshua Ong Jun Leang, Renxi Wang, Wenda Li, Haonan Li, Luke Ong, Conrad Watt

机构 * Nanyang Technological University Singapore(南洋理工大学新加坡分校) Imperial College London London, UK(伦敦帝国理工学院伦敦分校) University of Edinburgh Edinburgh, UK(爱丁堡大学爱丁堡分校)

AI总结 研究针对交互式定理证明中人工操作限制可扩展性及基于LLM的证明智能体成本高的问题,提出将智能体从源文本提升到抽象语法树的方法,实现了AoA,在多个方面有显著提升且解决更多难题。

Comments 13 pages

详情
AI中文摘要

交互式定理证明(ITP)是程序验证和形式化数学的基础,但人工操作限制了其可扩展性。基于大语言模型(LLM)的证明智能体有望减轻这一负担,但其高昂的令牌消耗和API成本仍是主要障碍。我们将此成本追溯到一个共同根源:当前智能体在序列化的具体语法上运行,将证明作为源文本发出,并通过单独的基于行号的查询恢复证明状态,因此每次编辑都会移动后续行,并迫使错误和状态重复重新定位。对具体语法的这种依赖也阻碍了Minilang的采用,Minilang是一种最近的证明语言,在基于LLM的证明方面达到了当前最优水平,但对于LLMs的训练语料库来说太新了。我们通过将智能体从源文本提升到抽象语法树(AST)来解决这两个问题:模型以Minilang的AST的JSON表示形式提供证明,这是工具调用LLMs原生的,并通过树编辑模型驱动证明器,该模型将证明操作和状态融合到一个证明树中,因此每个操作都携带其自己子目标的状态,可直接从树中读取。我们在“AST上的智能体”(AoA)中实现了这一设计。与亚马逊的Isabelle智能体在miniF2F和NTP4VC-Pearl常见成功集上相比,AoA将API成本降低了2.3-4.7倍(归一化输入缓存计算),使用的令牌减少了2.9-6.9倍,工具调用减少了3.9-8.9倍,并完成速度快1.4-2.0倍,同时在更难的验证基准上解决了更多问题。

英文摘要

Interactive theorem proving (ITP) underpins program verification and formalized mathematics, but its manual effort limits scalability. LLM-based proof agents promise to ease this effort, but their heavy token consumption and API cost remain a major obstacle. We trace this cost to a shared root: current agents operate on serialized concrete syntax, emitting proofs as source text and recovering proof states through separate, line-number-based queries, so every edit shifts later lines and forces repeated relocation of errors and states. This same dependence on concrete syntax also blocks adoption of Minilang, a recent proof language that reaches SOTA on LLM-based proving but is too new for LLMs' training corpora. We address both problems by lifting the agent off source text and onto the abstract syntax tree (AST): the model supplies proofs as JSON representations of Minilang's AST -- native to tool-calling LLMs -- and drives the prover through a tree-edit model that fuses proof operations and states into one proof tree, so each operation carries its own subgoal's state, readable directly off the tree. We realize this design in \emph{Agent over AST} (AoA). Against Amazon's Isabelle Agent on miniF2F and NTP4VC-Pearl common success sets, AoA cuts API cost by 2.3--4.7x (normalized input-cache accounting), uses 2.9--6.9x fewer tokens and 3.9--8.9x fewer tool calls, and finishes 1.4--2.0x faster -- while also solving far more problems on the harder verification benchmark.

URL PDF HTML 收藏
2607.17910 2026-07-21 cond-mat.mtrl-sci cs.AI cs.LG 新提交

Chemical filters for ultra-high-throughput materials screening and generation

用于超高通量材料筛选和生成的化学过滤器

Kinga O. Mastej, Panyalak Detrattanawichai, Hyunsoo Park, Anthony Onwuli, Masahiro Negishi, Aron Walsh

机构 * Department of Materials, Imperial College London(帝国理工学院材料系)

AI总结 研究针对AI生成材料成分不合理问题,引入化学有效性算子,基于SMACT包构建氧化态模型,通过可调阈值支持不同工作流程。经基准测试,该算子能过滤不合理成分,还可作强化学习奖励,为氧化态感知生成模型奠定基础。

Comments 19 pages, 7 figures, 3 tables, including Supplementary Information

详情
AI中文摘要

生成式人工智能正在迅速改变材料设计,通过对巨大化学空间进行从头探索。然而,很大一部分人工智能生成的成分仍然不合理,违反了既定的化学原理,这限制了生成式材料设计的可靠性和可解释性。在此,我们引入了一种化学有效性算子,将启发式化学规则重新塑造为一种可配置的算法先验,用于评估和指导生成式材料发现。基于开源SMACT包构建的一个数据驱动的氧化态模型揭示了可调阈值,允许用户在宽松和保守的化学约束之间连续插值,同时支持探索性和保守性材料设计工作流程。对六种无机晶体的最新生成模型进行基准测试表明,大多数模型能重现化学计量,但实际氧化态组合的代表性不足,过滤可去除依赖罕见氧化态的成分,同时保留凸包附近的低能化合物。除了筛选,同一个算子还可以作为强化学习奖励,引导潜在扩散模型生成基于化学的成分。通过编码化学启发式和观察结果,这项工作为氧化态感知生成模型奠定了基础。

英文摘要

Generative artificial intelligence is rapidly transforming materials design by enabling de novo exploration of immense chemical spaces. Yet a large proportion of AI-generated compositions remain implausible, violating established chemical principles, which limits the reliability and interpretability of generative materials design. Here, we introduce a chemical validity operator that recasts heuristic chemical rules as a configurable algorithmic prior for evaluating and guiding generative materials discovery. Built on the open-source SMACT package, a data-informed oxidation-state model exposes tunable thresholds, allowing users to interpolate continuously between permissive and conservative chemical constraints, while supporting both exploratory and conservative materials-design workflows. Benchmarking six state-of-the-art generative models for inorganic crystals shows that most reproduce stoichiometry but under-represent realistic oxidation-state combinations, and that filtering removes compositions reliant on rarely observed oxidation states while preserving low-energy compounds near the convex hull. Beyond screening, the same operator can also serve as a reinforcement-learning reward, steering a latent diffusion model towards chemically grounded compositions. By encoding chemical heuristics and observations, this work establishes a foundation for oxidation-state-aware generative models.

URL PDF HTML 收藏
2602.07008 2026-07-21 cs.CV cs.LG 版本更新

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making

不应学习的地方:基于子集归因约束的先验对齐训练以实现可靠的决策制定

Ruoyu Chen, Shangquan Sun, Xiaoqing Guo, Sanyi Zhang, Kangwei Liu, Shiming Liu, Zhangcheng Wang, Qunli Zhang, Wei Wang, Hua Zhang, Xiaochun Cao

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) University of Chinese Academy of Sciences(中国科学院大学) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) Department of Computer Science, Hong Kong Baptist University(香港 Baptist 大学计算机科学系) Communication University of China(中国传媒大学) Imperial College London(伦敦帝国学院) School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区网络科学与技术学院)

AI总结 本文提出了一种基于归因的先验对齐方法,通过子集选择归因技术约束模型依赖于人类先验区域,从而提升决策的可靠性。

详情
AI中文摘要

可靠的模型不仅要预测正确,还要能用可接受的证据来解释决策。然而,传统监督学习通常只提供类别级标签,使模型通过捷径相关性实现高精度,而非预期的证据。人类先验可以约束此类行为,但对齐模型到这些先验仍然具有挑战性,因为学习的表示往往偏离人类感知。为了解决这一挑战,我们提出了一种基于归因的人类先验对齐方法。我们将人类先验编码为模型应依赖的输入区域(例如边界框),并利用高度忠实的子集选择归因方法,在训练过程中暴露模型的决策证据。当归因区域显著偏离先验区域时,我们惩罚对非先验证据的依赖,促使模型将归因转向预期区域。这是通过一个训练目标实现的,该目标通过人类先验诱导归因约束。我们在基于MLLM的GUI代理模型上验证了我们的方法,涵盖图像分类和点击决策任务。在传统分类和自回归生成设置中,人类先验对齐一致提高了任务准确性,同时增强了模型的决策合理性。

英文摘要

Reliable models should not only predict correctly, but also justify decisions with acceptable evidence. Yet conventional supervised learning typically provides only class-level labels, allowing models to achieve high accuracy through shortcut correlations rather than the intended evidence. Human priors can help constrain such behavior, but aligning models to these priors remains challenging because learned representations often diverge from human perception. To address this challenge, we propose an attribution-based human prior alignment method. We encode human priors as input regions that the model is expected to rely on (e.g., bounding boxes), and leverage a highly faithful subset-selection-based attribution approach to expose the model's decision evidence during training. When the attribution region deviates substantially from the prior regions, we penalize reliance on off-prior evidence, encouraging the model to shift its attribution toward the intended regions. This is achieved through a training objective that imposes attribution constraints induced by the human prior. We validate our method on both image classification and click decision tasks in MLLM-based GUI agent models. Across conventional classification and autoregressive generation settings, human prior alignment consistently improves task accuracy while also enhancing the model's decision reasonability.

URL PDF HTML 收藏
2601.22259 2026-07-21 cs.LG 版本更新

Tabular Foundation Models Can Do Survival Analysis

表格基础模型也能进行生存分析

Da In Kim, Wei Siang Lai, Kelly W. Zhang

机构 * Department of Computing, Imperial College London(帝国理工学院计算机系) Department of Mathematics, Imperial College London(帝国理工学院数学系)

AI总结 本文提出了一种基于分类的框架,通过将生存分析转化为二分类问题,使表格基础模型能够无需显式训练即可进行生存分析,并在多个数据集上验证了其有效性。

详情
AI中文摘要

尽管表格基础模型在分类和回归任务中取得了显著成功,但将其应用于建模时间到事件结果以进行生存分析是具有挑战性的,因为存在右删失问题,即数据观测可能在事件发生前结束。我们开发了一个基于分类的框架,通过将事件时间离散化,将静态和动态生存分析重新表述为一系列二分类问题。删失观测自然地被处理为在特定时间点具有缺失标签的示例。这种分类表述使现有的表格基础模型能够通过上下文学习进行生存分析,而无需显式训练。我们证明,在标准删失假设下,最小化我们的二分类损失可以随着训练集大小的增加而恢复真实的生存概率。我们通过在53个现实世界数据集上的评估证明,使用这种分类表述的现成表格基础模型在多个生存度量指标上平均优于经典和深度学习基线。

英文摘要

While tabular foundation models have achieved remarkable success in classification and regression, adapting them to model time-to-event outcomes for survival analysis is non-trivial due to right-censoring, where data observations may end before the event of interest occurs. We utilize a classification-based framework that reformulates both static and dynamic survival analysis as a series of binary classification problems by discretizing event times. Censored observations are naturally handled as examples with missing labels at certain time points. This classification formulation enables existing tabular foundation models (TFMs) to perform survival analysis through in-context learning without explicit training. In contrast to classical approaches that use binary classifiers to model discrete-time hazards, our approach directly models cumulative failure probabilities, which we find empirically to be more robust to the number of discretization bins by avoiding multiplicative accumulation of per-bin errors. We prove that under standard censoring assumptions, minimizing our binary classification loss recovers the true survival probabilities as the training set size increases. We demonstrate through evaluation across 48 real-world datasets (43 static and 5 dynamic) that off-the-shelf TFMs with this classification formulation outperform classical and deep learning baselines on average over multiple survival metrics.

URL PDF HTML 收藏
2512.13247 2026-07-21 cs.CV 版本更新

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

STARCaster:用于身份和视图感知的会说话肖像的时空自回归视频扩散

Foivos Paraperas Papantoniou, Stathis Galanakis, Rolandos Alexandros Potamias, Bernhard Kainz, Stefanos Zafeiriou

机构 * Imperial College London, UK(伦敦帝国理工学院) FAU Erlangen–Nürnberg, Germany(埃朗根-纽伦堡大学)

AI总结 研究提出STARCaster模型,通过重新思考参考和几何范式,采用组合方法及解耦学习,解决语音驱动肖像动画和动态视点控制问题,能有效泛化,超越先前方法。

Comments ICML 2026

详情
AI中文摘要

本文提出了STARCaster,一种身份感知的时空视频扩散模型,在统一框架内,给定身份嵌入或参考图像,解决语音驱动的肖像动画和动态视点控制问题。现有2D语音到视频扩散模型严重依赖参考指导,导致运动多样性有限,3D感知动画通常依赖预训练三平面生成器的反演,常导致重建不完美和身份漂移。我们从两方面重新思考基于参考和几何的范式,一是通过引入更软的身份约束偏离预训练时的严格参考条件,二是通过利用视频数据固有的多视图性质在2D视频域内隐式解决3D感知问题。STARCaster采用组合方法,从身份感知运动建模,到基于唇读监督的视听同步,最后通过时间到空间的适配实现新颖视图动画。为克服4D视听数据稀缺问题,提出解耦学习方法,独立训练视图一致性和时间连贯性。综合评估表明,STARCaster在不同任务和身份上有效泛化,始终超越先前方法。

英文摘要

This paper presents STARCaster, an identity-aware spatio-temporal video diffusion model that addresses both speech-driven portrait animation and dynamic viewpoint control, given an identity embedding or reference image, within a unified framework. Existing 2D speech-to-video diffusion models depend heavily on reference guidance, leading to limited motion diversity. At the same time, 3D-aware animation typically relies on inversion through pretrained tri-plane generators, which often leads to imperfect reconstructions and identity drift. We rethink reference- and geometry-based paradigms in two ways. First, we deviate from strict reference conditioning at pretraining by introducing softer identity constraints. Second, we address 3D awareness implicitly within the 2D video domain by leveraging the inherent multi-view nature of video data. STARCaster adopts a compositional approach progressing from ID-aware motion modeling, to audio-visual synchronization via lip reading-based supervision, and finally to novel view animation through temporal-to-spatial adaptation. To overcome the scarcity of 4D audio-visual data, we propose a decoupled learning approach in which view consistency and temporal coherence are trained independently. Comprehensive evaluations demonstrate that STARCaster generalizes effectively across tasks and identities, consistently surpassing prior approaches in different benchmarks.

URL PDF HTML 收藏
2607.15951 2026-07-20 cs.GR cs.CV cs.DC 新提交

Rendering 3D Gaussians on a Graph Processor

在图形处理器上渲染3D高斯分布

Nicholas Fry, Ignacio Alzugaray, Mark Pupilli, Paul H. J. Kelly, Andrew J. Davison

机构 * Imperial College London Department of Computing(伦敦帝国学院计算机系)

AI总结 研究在IPU上实现3D高斯渲染器,输入为3D高斯地图,通过特定模式路由高斯基元,遵循BSP模型计算。评估了仅用SRAM实现的瓶颈及对性能和质量的影响,还探讨了对GPU和相关架构的意义。

Comments Project page: https://nmjfry.github.io/ipu-3dgs/

Journal ref Eurographics Symposium on Rendering (Symposium Track), 2026

详情
AI中文摘要

我们展示了在智能处理单元(IPU)上首次实现的3D高斯渲染器,该IPU由1472个仅带有片上静态随机存取存储器(SRAM)的独立瓦片组成,这些限制近似于高效传感器 - 处理器架构的属性。输入场景是来自真实世界序列的3D高斯地图。每个瓦片“拥有”帧缓冲区的一个屏幕空间区域;高斯基元通过在东北 - 西 - 南(NEWS)网格上的曼哈顿距离跳跃路由到目标瓦片,然后以扩展树模式分布到重叠邻居。计算遵循IPU的批量同步并行(BSP)模型,瓦片间通信在编译时定义。我们展示了这种硬件通过在核心之间实现本地数据传输来利用空间和时间局部性。我们评估了这种仅使用SRAM实现中的瓶颈:瓦片间带宽、每个瓦片的SRAM容量以及来自非均匀高斯密度的工作负载不平衡。我们分析了这些限制如何影响性能和渲染质量。这一探索为传统图形处理器(GPU)和3D表示提出了更广泛的问题,表明直接的流多处理器(SM)间通信可能为减少GPU内核中的动态随机存取存储器(DRAM)访问提供方法。我们讨论了这些对未来传感器上和无DRAM架构的影响。项目页面:此https URL

英文摘要

We present the first implementation of a 3D Gaussian renderer on an Intelligence Processing Unit (IPU), comprising 1,472 independent tiles with only on-chip SRAM; constraints that approximate properties of efficient sensor-processor architectures. Our input scenes are 3D Gaussian maps from real-world sequences. Each tile 'owns' a screen-space region of the framebuffer; Gaussian primitives are routed to destination tiles via Manhattan-distance hops on a north-east-west-south (NEWS) grid, then distributed to overlapping neighbours in an expanding tree pattern. Computation follows the IPU's Bulk Synchronous Parallel (BSP) model, with inter-tile communication defined at compile time. We show this hardware allows us to exploit spatial and temporal locality by enabling local data transfer between cores. We evaluate the bottlenecks in this SRAM-only implementation: inter-tile bandwidth, per-tile SRAM capacity, and workload imbalance from non-uniform Gaussian density. We analyse how these constraints affect performance and render quality. This exploration raises broader questions for conventional GPUs and 3D representations, suggesting that direct inter-SM (streaming multiprocessor) communication might offer ways to reduce DRAM access in GPU kernels. We discuss these implications for the future of on-sensor and DRAM-free architectures. Project page: https://nmjfry.github.io/ipu-3dgs/

URL PDF HTML 收藏
2607.15832 2026-07-20 math.NA cs.LG cs.NA math.DS 新提交

A zero-one law for one-shot system identification

一次性系统识别的零一律

Nicolas Boullé, Diana Halikias, Samuel E. Otto, Alex Townsend

机构 * Imperial College London(帝国理工学院) New York University(纽约大学) Cornell University(康奈尔大学)

AI总结 研究一次性系统识别问题,通过规定字典项线性参数化解析系统,证明零一律,将问题简化为关于退化输入的问题,能从单轨迹数据恢复多种系统并检测额外探测需求。

Comments 8 pages, 2 figures

详情
AI中文摘要

我们研究由规定字典项(如偏微分算子和动力系统)组合进行线性参数化的解析系统。对于单个输入 - 响应对,当评估的字典项线性独立时恢复是可能的。我们证明了一个精确的零一律:要么没有输入能唯一确定系数,要么几乎每个从非退化高斯测度中采样的随机输入都能。这种二分法将一次性系统识别简化为关于退化输入的问题,并为任何恢复的模型提供后验证书。数值例子从单轨迹数据中恢复动力系统、非线性偏微分方程和结构化矩阵族,同时还能检测何时需要额外探测。

英文摘要

Can a model be identified from one experiment? We study analytic systems that are linearly parameterized by a combination of prescribed dictionary terms, such as partial differential operators and dynamical systems. For a single input-response pair, recovery is possible exactly when the evaluated dictionary terms are linearly independent. We prove a sharp zero-one law: either no input uniquely determines the coefficients, or almost every random input sampled from a nondegenerate Gaussian measure does. This dichotomy reduces one-shot system identification to a question about degenerate inputs and provides an a posteriori certificate for any recovered model. Numerical examples recover dynamical systems, nonlinear partial differential equations, and structured matrix families from single trajectory data, while also detecting when an extra probe is necessary.

URL PDF HTML 收藏
2607.15291 2026-07-20 math.NA cs.LG cs.NA 新提交

A Physics-Informed Neural Network with a Modified Lorentzian Activation for Nonlocal Gradient-Flow Equations in Dynamic Density Functional Theory

一种具有修正洛伦兹激活函数的物理信息神经网络用于动态密度泛函理论中的非局部梯度流方程

Dimitrios Gourzoulidis, Soumaya Elkantassi, Serafim Kalliadasis

机构 * Department of Chemical Engineering, Imperial College London(帝国理工学院化学工程系) Department of Operations, University of Lausanne(洛桑大学运营系)

AI总结 该研究针对动态密度泛函理论中的非局部梯度流方程,开发物理信息神经网络框架。引入修正洛伦兹激活函数和预计算离散算子,经多维度测试,新激活函数加速收敛,框架与参考解吻合且捕捉到梯度流行为,展现求解此类方程的潜力。

详情
AI中文摘要

我们为动态密度泛函理论(DDFT)中出现的非局部偏微分方程开发了一个物理信息神经网络(PINN)框架。这类方程对标准PINN方法具有挑战性,因其包含非线性、非局部相互作用项及潜在的梯度流结构,常导致收敛慢和优化困难。我们将PINN方法应用于DDFT梯度流方程,引入两个计算组件:一个修正的洛伦兹激活函数,小输入时近似线性表现,输入量增大时趋于零;一个预计算离散算子,用于训练时有效评估非局部卷积项。该方法在一维和二维的四个例子上进行测试。第一个例子有精确稳态解,其余例子将神经网络近似与用连续和间断伽辽金有限元离散化计算的参考解进行验证。通过\(L^1\)、\(L^2\)和\(L^\infty\)误差以及质量守恒和自由能耗散评估准确性和物理一致性。结果表明,相对于标准双曲正切函数,所提出的激活函数加速了收敛,整体框架与参考解保持良好一致并捕捉到预期的梯度流行为。这些发现证明了所提出的PINN框架在求解DDFT中出现的非局部梯度流方程方面的潜力。

英文摘要

We develop a physics-informed neural network (PINN) framework for nonlocal partial differential equations arising in dynamic density functional theory (DDFT). Such equations are challenging for standard PINN methods because they involve nonlinearities, nonlocal interaction terms, and an underlying gradient-flow structure, often leading to slow convergence and difficult optimization. We adapt the PINN methodology to DDFT gradient-flow equations and introduce two computational components: a modified Lorentzian activation function that behaves approximately linearly for small inputs and decays toward zero as the input magnitude increases, and a precomputed discrete operator for evaluating the nonlocal convolution term efficiently during training. The method is tested on four examples in one and two space dimensions. In the first example, the exact stationary solution is known, while in the remaining cases the neural-network approximations are validated against reference solutions computed using continuous and discontinuous Galerkin finite element discretizations. Accuracy and physical consistency are assessed through $L^1$, $L^2$, and $L^\infty$ errors, together with mass conservation and free-energy dissipation. The results show that the proposed activation function accelerates convergence relative to the standard $\tanh$ function, while the overall framework maintains good agreement with the reference solutions and captures the expected gradient-flow behaviour. These findings demonstrate the potential of the proposed PINN framework for solving nonlocal gradient-flow equations arising in DDFT.

URL PDF HTML 收藏
2607.15472 2026-07-20 math.DS cs.LG physics.bio-ph 新提交

Ptolemy's Equant Equates to a Universal Dynamical Clock via Machine Learning

通过机器学习,托勒密等距点等同于通用动态时钟

Jingdong Zhang, Luan Yang, Murilo S. Baptista, Zefeng Zhang, Qunxi Zhu, Wei Lin, Celso Grebogi

机构 * School of Mathematical Sciences, Fudan University, Shanghai 200433, China(复旦大学数学科学学院) Research Institute of Intelligent Complex Systems, Fudan University, Shanghai 200433, China(复旦大学智能复杂系统研究所) Department of Mathematics, Imperial College London, London, SW7 2AZ, United Kingdom(伦敦帝国学院数学系) Institute for Complex Systems and Mathematical Biology, University of Aberdeen, Aberdeen AB24 3UE, United Kingdom(阿伯丁大学复杂系统与数学生物学研究所) Department of Psychiatry, University of Cambridge, Cambridge CB2 1TN, United Kingdom(剑桥大学精神病学系)

AI总结 研究非线性高维振荡中相位和相位动力学识别问题,基于托勒密等距点建立通用动态时钟原理,用机器学习框架证明其存在并构建相关动力学,通过四个发现展示价值,为振荡系统研究提供新途径。

Comments 56 pages, 12 figures, 3 tables

详情
AI中文摘要

振荡动力学在非线性系统中普遍存在,但在非线性高维振荡中识别具有物理可解释性的相位和相位动力学仍是一个未解决的核心问题。本文建立了通用动态时钟原理,受托勒密等距点启发,通过面积均匀性原理形式化,将任意维度和几何形状的振荡等效表示为通过等距点诱导的非线性观察坐标的匀速旋转。利用机器学习框架,证明了一类广泛振荡动力学中等距点的存在,并构建了相关动态时钟和相位动力学。通过四个发现展示了其在揭示新物理规则和现象方面的价值,包括大肠杆菌群体中的集体振荡遵循超线性缩放定律、工程遗传电路对基因表达和环境条件变化的响应机制、贝里几何相位的经典力学对应物自然出现以及最优等距点非均匀性为临界转变提供几何预警信号并预测临界参数。动态时钟提供了可直接从数据构建的操作和系统无关的相位动力学,实现了对振荡系统的分类、比较和控制,为理解特定动态机制如何支持网络系统中不同功能行为提供了新途径。

英文摘要

Oscillatory dynamics arise ubiquitously in nonlinear systems, yet identifying a physically interpretable phase and phase dynamics in nonlinear, high-dimensional oscillations remains a central unresolved problem. Here we establish the principle of a universal dynamical clock, a physical perspective in which oscillations of arbitrary dimensionality and geometry are equivalently represented as uniform rotation through an equant-induced nonlinear viewing coordinate, inspired by Ptolemy's equant and formalised through an areal-uniformity principle reminiscent of Kepler's second law. Using a machine-learning framework, we demonstrate the existence of such an equant for a broad class of oscillatory dynamics and construct the associated dynamical clock and phase dynamics under additive forces, including noise, periodic perturbations, and coupling. Its value in uncovering new physical rules and phenomena is demonstrated by four findings: (i) collective oscillations in Escherichia coli populations obey a previously unexplained superlinear scaling law, resolving a long-standing open problem posed in 2004; (ii) the response mechanisms of engineered genetic circuits to changes in gene expression and environmental conditions; (iii) a classical-mechanics counterpart of the Berry geometric phase emerges naturally from the phase of the dynamical clock; and (iv) optimal equant non-uniformity provides a geometric early-warning signal for critical transitions and enables prediction of critical parameters. By providing operational and system-agnostic phase dynamics that can be constructed directly from data, the dynamical clock enables principled classification, comparison, and control of oscillatory systems, and offers a new route to understanding how specific dynamical regimes support distinct functional behaviours in networked systems.

URL PDF HTML 收藏
2601.22818 2026-07-20 cs.CR cs.AI 版本更新

Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models

嵌入空间中的隐秘行动:基于几何的隐写术与大语言模型中的检测

Charles Westphal, Keivan Navaie, Fernando E. Rosas

机构 * UCL Centre for Artificial Intelligence, University College London, UK(伦敦大学学院人工智能中心,大学学院伦敦) School of Computing and Communications, Lancaster University, UK(兰卡斯特大学计算机与通讯学院) Department of Informatics, University of Sussex, UK(苏塞克斯大学信息学院) Centre for Psychedelic Research and Centre for Complexity Science, Imperial College London, UK(伦敦帝国学院迷幻研究与复杂科学中心) Centre for Eudaimonia and Human Flourishing, University of Oxford, UK(牛津大学幸福与人类繁荣中心)

AI总结 本研究提出了一种低恢复性隐写术,通过嵌入空间衍生映射提升秘密恢复率,同时减少负载恢复性,并提出基于机制可解释性的检测方法,提高微调模型中的秘密检测准确率。

详情
AI中文摘要

经过微调的LLMs可以通过隐写通道在输出中隐蔽地编码提示秘密。先前的工作展示了这种威胁,但依赖于可以轻易恢复的编码。我们通过分类器准确性正式化负载恢复性,并显示先前的方案实现了100%的恢复性。对此,我们引入了低恢复性隐写术,用嵌入空间衍生的映射取代任意映射。对于在TrojanStego提示上训练的Llama-8B(LoRA)和Ministral-8B(LoRA),精确秘密恢复率从17%上升到30%(+78%),从24%上升到43%(+80%),而在在Wiki提示上训练的Llama-70B(LoRA)中,从9%上升到19%(+123%),同时减少负载恢复性。我们随后讨论检测。我们认为,检测基于微调的隐写攻击需要超越传统隐写分析的方法。标准方法测量分布偏移,这是微调的预期副作用。相反,我们提出了一种机制可解释性方法:在后期层激活上训练的线性探针可以检测秘密,其在微调模型中的检测准确率比基模型高33%以上,即使对于低恢复性方案也是如此。这表明恶意微调留下了可解释性防御的可操作内部签名。

英文摘要

Fine-tuned LLMs can covertly encode prompt secrets into outputs via steganographic channels. Prior work demonstrated this threat but relied on trivially recoverable encodings. We formalize payload recoverability via classifier accuracy and show previous schemes achieve 100\% recoverability. In response, we introduce low-recoverability steganography, replacing arbitrary mappings with embedding-space-derived ones. For Llama-8B (LoRA) and Ministral-8B (LoRA) trained on TrojanStego prompts, exact secret recovery rises from 17$\rightarrow$30\% (+78\%) and 24$\rightarrow$43\% (+80\%) respectively, while on Llama-70B (LoRA) trained on Wiki prompts, it climbs from 9$\rightarrow$19\% (+123\%), all while reducing payload recoverability. We then discuss detection. We argue that detecting fine-tuning-based steganographic attacks requires approaches beyond traditional steganalysis. Standard approaches measure distributional shift, which is an expected side-effect of fine-tuning. Instead, we propose a mechanistic interpretability approach: linear probes trained on later-layer activations detect the secret with up to 33\% higher accuracy in fine-tuned models compared to base models, even for low-recoverability schemes. This suggests that malicious fine-tuning leaves actionable internal signatures amenable to interpretability-based defenses.

URL PDF HTML 收藏
2607.15077 2026-07-17 cs.LG 新提交

An Introduction to Sparse Identification of Nonlinear Dynamics for Engineering Applications

工程应用中的非线性动力学稀疏识别介绍

Yao Cheng Li, Ana Larrañaga, Steven L. Brunton, Urban Fasel

机构 * Department of Aeronautics, Imperial College London(伦敦帝国理工学院航空系) Department of Mechanical Engineering, University of Washington(华盛顿大学机械工程系) NSF AI Institute in Dynamic Systems, University of Washington(华盛顿大学动态系统领域美国国家科学基金会人工智能研究所)

AI总结 介绍工程应用中非线性动力学稀疏识别(SINDy)方法,通过对候选非线性项库稀疏回归解决代理建模局限性,教程介绍该方法及扩展,经案例研究表明其易实现且灵活,是工程应用有价值的识别工具。

Comments 15 pages, 4 figures

详情
AI中文摘要

许多工程问题涉及控制方程表征不佳或仅部分已知的现象。神经网络等代理建模技术虽能捕捉系统行为,但需大量难以获取的训练数据集,且模型物理可解释性有限。稀疏识别非线性动力学(SINDy)方法通过对候选非线性项库进行稀疏回归,从小规模数据集中恢复可解释的控制方程,解决了上述两个局限性。本教程介绍了SINDy方法,并逐步介绍其主要扩展,从抗噪声弱形式和基于集成的变体到约束和可参数化公式。本文及配套教程分为三个部分:第一部分介绍标准SINDy算法并逐步扩展,让无先验知识读者能跟随步骤并将方法应用于自身问题;其余两部分给出详细案例研究,一是无人机系统识别,二是混沌热虹吸换热器。通过这些例子,旨在证明SINDy易于实现且足够灵活,可作为先进工程应用的有价值识别工具。

英文摘要

Many engineering problems involve phenomena whose governing equations are poorly characterized or only partially known. Surrogate modeling techniques such as neural networks can capture the behavior of these systems, but they typically demand large training datasets that are difficult to obtain in engineering contexts and yield models with limited physical interpretability. The Sparse Identification of Nonlinear Dynamics (SINDy) method addresses both limitations by performing sparse regression over libraries of candidate nonlinear terms, recovering interpretable governing equations from comparatively small datasets. Although SINDy has been demonstrated extensively on canonical benchmark systems, its application to practical engineering problems is less widely documented. This tutorial introduces the SINDy method and progressively builds toward its main extensions, from noise-robust weak-form and ensembling-based variants to constrained and parametrizable formulations. The paper and the accompanying tutorial (available at https://github.com/paullililili/SINDy4Engineers) is organized in three parts: the first introduces the standard SINDy algorithm and progressively extends it, inviting readers without prior knowledge to follow each step and adapt the methods to their own problems; the remaining two parts present detailed case studies on (1) the system identification of an unmanned aerial vehicle and (2) a chaotic thermosyphon heat exchanger. Through these examples, we aim to demonstrate that SINDy is simple to implement yet flexible enough to serve as a valuable identification tool for advanced engineering applications.

URL PDF HTML 收藏
2607.14338 2026-07-17 cs.CV cs.AI cs.LG 新提交

Beyond scalar losses: calibrating segmentation models via gradient vector field surgery

超越标量损失:通过梯度向量场手术校准分割模型

Laurin Lux, Alexander H. Berger, Moritz Knolle, Daniel Rückert, Johannes C. Paetzold

机构 * School of Computation, Information and Technology, TUM(慕尼黑工业大学计算、信息与技术学院) Munich Center for Machine Learning(慕尼黑机器学习中心) Department of Radiology, Weill Cornell Medicine(威尔康乃尔医学院放射科) School of Medicine and Health, TUM University Hospital(慕尼黑工业大学医院医学与健康学院) Cornell Tech(康奈尔科技学院) Department of Computing, Imperial College London(伦敦帝国理工学院计算系)

AI总结 研究针对基于区域损失函数训练的分割模型校准不佳问题,提出对梯度向量场进行“手术”,即给损失偏导数添加因子,依预测误差线性缩放梯度大小,经2D和3D医学分割任务验证该方法有效且能保持高预测准确性。

Comments MIDL 2026. Published version: https://proceedings.mlr.press/v315/lux26a.html

Journal ref Proceedings of Machine Learning Research 315:3397-3423, 2026

详情
AI中文摘要

基于区域的损失函数,如骰子损失,已成为高度类别和区域不平衡分割任务的事实上的标准。然而,使用基于区域的损失函数训练的模型校准不佳,通常会产生过度自信的预测。在医学成像应用中,这种校准错误阻碍了临床应用。在这项工作中,我们概述了一种关于这种过度自信的新梯度观点,并展示了它如何影响基于区域的损失函数。我们提出对梯度向量场进行“手术”,作为一种简单而有效的干预措施来减轻校准问题。这种手术在损失的偏导数中添加一个因子,根据预测误差线性缩放梯度的大小。在2D和3D医学分割任务的实证评估中,我们证明了这种干预的有效性,同时当与任何基于区域的损失函数结合使用时保持高预测准确性。

英文摘要

Region-based loss functions, such as the Dice loss, have established themselves as the de facto standard for highly class- and region-imbalanced segmentation tasks. However, models trained using region-based loss functions are notoriously miscalibrated and typically yield over-confident predictions. In medical imaging applications, such as defining tumor resection margins, this miscalibration is hindering clinical adoption. In this work, we outline a novel gradient perspective on this overconfidence and show how it affects region-based loss functions. We propose a "surgery" on the gradient vector field as a simple, yet effective intervention to mitigate calibration issues. This surgery adds a factor to the loss's partial derivative, scaling the gradient's magnitude linearly with the prediction error. In empirical evaluations across 2D and 3D medical segmentation tasks, we demonstrate the effectiveness of this intervention while maintaining high prediction accuracy when used in conjunction with any region-based loss function.

URL PDF HTML 收藏
2607.14272 2026-07-17 cs.LG math.DS math.OC 新提交

Lyapunov Guidance: A Unified Framework for Stabilizing Generative Flows

李雅普诺夫引导:稳定生成流的统一框架

Jingdong Zhang, Xinze Li, Yize Jiang, Luan Yang, Minkai Xu, Junhong Liu

机构 * Imperial College London(伦敦帝国理工学院) Fudan University(复旦大学) Stanford University(斯坦福大学) MicroCyto(微赛生物)

AI总结 研究针对流匹配重新训练计算昂贵、现有训练后引导方法无稳定性保证的问题,提出LyaGuide框架,将流引导作为李雅普诺夫控制问题,统一多种引导策略,经实验验证其在多方面有改进且保持计算效率。

Comments 25 pages, 13 figures

详情
AI中文摘要

流匹配已成为学习复杂数据分布的有效框架,但将预训练流模型应用于新任务通常需要计算昂贵的重新训练。训练后引导提供了一种更有效的替代方法,但现有方法大多是启发式的,没有明确的稳定性保证。我们提出了LyaGuide,一个统一的李雅普诺夫引导框架来解决这一限制,将流引导表述为李雅普诺夫控制问题。主要理论结果建立了引导流匹配与李雅普诺夫控制之间的等价关系,统一了多种引导策略。引入伪投影算子以确保李雅普诺夫条件,支持模型驱动和数据驱动两种设置。实验表明在样本质量等方面有持续改进且保持计算效率。

英文摘要

Flow matching has emerged as an effective framework for learning complex data distributions, but adapting pretrained flow models to new tasks often requires computationally expensive retraining. Post-training guidance provides a more efficient alternative, but existing methods are largely heuristic and offer no explicit stability guarantees. We address this limitation by proposing LyaGuide, a unified Lyapunov-guided framework that formulates flow guidance as a Lyapunov control problem. Our main theoretical result establishes an equivalence between guided flow matching and Lyapunov control, thereby unifying common guidance strategies, such as classifier guidance, reward guidance, and energy-based guidance, within a single control-theoretic framework. To enforce the Lyapunov condition, we introduce a pseudo-projection operator with a closed-form expression that endows learned or heuristic guidance terms with explicit stability guarantees. LyaGuide supports two practical settings: a model-driven setting, where the target guidance distribution is specified through a known Lyapunov function, and a data-driven setting, where the guidance is adapted from task-specific downstream data. LyaGuide is compatible with existing guidance methods, introduces minimal additional computational overhead, and is straightforward to integrate in practice. Extensive experiments on synthetic benchmarks, image inverse problems, reinforcement learning planning, and energy-based modeling demonstrate consistent improvements in sample quality, guidance fidelity, and robustness, while maintaining computational efficiency.

URL PDF HTML 收藏
2508.00923 2026-07-17 cs.LG 版本更新

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

超越基准:动态、自动和系统化的红队代理用于可信的医疗语言模型

Jiazhen Pan, Bailiang Jian, Paul Hager, Yundi Zhang, Che Liu, Friederike Jungmann, Hongwei Bran Li, Julian Canisius, Chenyu You, Junde Wu, Jiayuan Zhu, Fenglin Liu, Yuyuan Liu, Niklas Bubeck, Moritz Knolle, Chen, Chen, Christian Wachinger, Zhenyu Gong, Cheng Ouyang, Georgios Kaissis, Benedikt Wiestler, Daniel Rueckert

机构 * Technical University of Munich (TUM)(慕尼黑技术大学) University of Oxford(牛津大学) TUM University Hospital(慕尼黑技术大学医院) Imperial College London(伦敦帝国理工学院) Harvard Medical School(哈佛医学院) Stony Brook University(史泰兹布鲁克大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) University of Sheffield(谢菲尔德大学)

AI总结 本文提出DAS红队框架,通过动态压力测试揭示医疗语言模型在鲁棒性、隐私、偏见和幻觉方面的潜在风险,发现高静态基准性能与低动态可靠性之间的'基准差距'。

详情
AI中文摘要

确保大型语言模型(LLMs)在临床实践中的安全性和可靠性对于防止患者伤害至关重要。然而,LLMs的发展速度如此之快,静态基准很快变得过时或容易过拟合,导致对模型可信度的误导性图像。在这里,我们介绍了一个动态、自动和系统化的(DAS)红队框架,该框架在四个关键安全轴上持续对LLMs进行压力测试:鲁棒性、隐私、偏见/公平性和幻觉。经过获得认证的临床医生验证,一组对抗性代理会自动突变临床测试用例,以实时揭示漏洞。将DAS应用于15个专有和开源的LLMs,揭示了高静态基准性能与低动态可靠性之间的深刻差距——“基准差距”。尽管中位数MedQA准确率超过80%,但94%的先前正确答案在我们的动态鲁棒性测试中失败。至关重要的是,这种脆弱性扩展到了现实的、开放性的HealthBench数据集,在此数据集中,顶级模型的失败率超过70%,且在评估中模型排名出现了显著变化,表明在已建立的静态基准上的高分可能反映的是表面记忆。我们观察到在其他领域也出现了类似的高失败率:隐私泄露在86%的场景中被触发,认知偏见先验改变了81%的公平性测试中的临床建议,我们发现广泛使用的模型中的幻觉率超过74%。通过将医疗LLM安全分析从静态清单转换为动态压力测试,DAS提供了一个基础、可扩展和持续的平台,以揭示必须在下一代医疗AI安全部署之前解决的潜在风险。

英文摘要

Large language models (LLMs) are increasingly used to answer health-related questions and support healthcare workflows, yet evidence for their safety still relies heavily on static benchmarks that can rapidly become obsolete or be optimized against. Here we introduce a Dynamic, Automatic, and Systematic (DAS) red-teaming audit framework that continuously stress-tests LLMs for health across four safety-critical axes: robustness, privacy, bias/fairness, and hallucination/factual inaccuracies. Validated against board-certified clinicians with high concordance, a suite of adversarial agents autonomously mutates health-related test cases to uncover vulnerabilities in real time. Applying DAS to 15 proprietary and open-source LLMs revealed a profound gap between high static benchmark performance and low dynamic reliability--the "Benchmarking Gap". Despite median MedQA accuracy exceeding 80\%, 94\% of previously correct answers failed under dynamic robustness testing. This brittleness generalized to the realistic, open-ended HealthBench dataset, where top-tier models exhibited failure rates exceeding 70\% and sharp shifts in model rankings across evaluations, suggesting that high scores on established static benchmarks may reflect superficial memorization. We observed similarly high failure rates across other domains: privacy leaks were elicited in 86\% of scenarios, cognitive-bias priming altered recommendations in 81\% of fairness tests, and hallucination rates exceeded 74\% in widely used models. By converting LLM safety evaluation for health from a static checklist into a living adversarial audit, DAS provides a scalable framework for surfacing latent risks before such systems are deployed in consumer-facing health assistants, clinician-facing tools, and broader healthcare workflows. Code is available at https://github.com/JZPeterPan/DAS-Medical-Red-Teaming-Agents.

URL PDF HTML 收藏
2607.14070 2026-07-16 q-bio.GN cs.LG 新提交

Screening of Biosecurity Features in Metagenomic Data with Evo 2 Probes

使用Evo 2探针筛选宏基因组数据中的生物安全特征

Jeremy Guntoro, Alexander Dack, Dylan Danno, Michaela Jančovičová, Križan Jurinović, Vanessa Smilansky

机构 * Department of Bioengineering, Imperial College London(帝国理工学院生物工程系) John Innes Centre(约翰·英尼斯研究中心) Independent Research Scientist(独立研究者)

AI总结 研究利用Evo 2探针在宏基因组数据中筛选生物安全特征,通过训练线性和注意力探针检测抗菌抗性及细菌毒力等,发现其在检测AMR等方面效果良好,可作为快速低成本首过检测层,明确了该方法的优势与局限。

详情
AI中文摘要

基因组基础模型如Evo 2能学习丰富的序列表示,但在生物安全筛选方面价值未充分探索。本文通过在冻结的Evo 2第26层激活上训练最小线性和注意力探针,研究这些表示中与生物安全相关信号的线性可及性。结果表明,探针能有效检测抗菌抗性(AMR),区分更细粒度的AMR药物类别子类别并与无关功能基因分离,细菌毒力也可解码但较弱。AMR探针在模拟短读上无需重新训练就有可比排名性能,还对比了与其他模型的情况。这些结果表明基于轻量级嵌入的探针可作为宏基因组生物监测的快速、低成本首过检测层,并明确了该方法的优缺点。

英文摘要

Genomic foundation models such as Evo 2 learn rich sequence representations, but their value for biosecurity screening is largely unexplored. We ask how much biosecurity-relevant signal is linearly accessible in these representations by training minimal linear and attention probes on frozen Evo 2 layer-26 activations, without fine-tuning the underlying model. Across held-out metagenomic test sets, the probes detect antimicrobial resistance (AMR) with strong discrimination: a linear probe reaches a region-level ROC-AUC of 0.888 (mean-pool), rising to 0.977 with a single-head attention probe. The probes resolve finer-grained AMR drug-class subcategories and separate them from unrelated functional genes, providing additional evidence that the learned signal is not explained solely by generic functional-gene status. Bacterial virulence is also decodable, though more weakly (region-level ROC-AUC 0.833). The AMR probe retains comparable ranking performance on simulated short reads without retraining, enabling evaluation before assembly in settings where assembly is computationally costly or unreliable. It achieves a read-level ROC-AUC of 0.898 (mean-pool), comparable to the mean-pooled full-region result. Within SynGenome, AMR-associated prompt labels are only weakly recoverable from Evo 1.5-generated sequences; these prompt-derived labels do not establish the function of the generated response sequences. A complementary sparse-autoencoder analysis recovers interpretable resistance-associated features but proves less consistent than the supervised probes. Together, these results position lightweight embedding-based probes as a fast, inexpensive first-pass detection layer for metagenomic biosurveillance and map both strengths and current limits of the approach. This work was conducted as part of the AIxBio Hackathon 2026 hosted by BlueDot Impact, Apart Research, and Cambridge Biosecurity Hub.

URL PDF HTML 收藏
2511.18685 2026-07-16 cs.CV cs.RO 版本更新

Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents

超越描述:为具身智能体进行细粒度动作的认知基准测试

Dayong Liu, Chao Xu, Weihong Chen, Suyu Zhang, Juncheng Wang, Jiankang Deng, Baigui Sun, Yang Liu

机构 * Zhejiang University(浙江大学) Wolf 1069 b(沃尔夫1069b) Sany Group(三一集团) The Hong Kong Polytechnic University(香港理工大学) Imperial College London(帝国理工学院)

AI总结 本文提出CFG-Bench基准测试,旨在评估具身智能体在物理交互中的细粒度动作能力,揭示现有MLLMs在高层次推理中的不足,并通过监督微调提升其性能。

Comments Accepted to ECCV2026

详情
AI中文摘要

多模态大语言模型(MLLMs)在复杂物理环境中作为具身智能体的决策引擎显示出有前景的结果。然而,现有基准测试往往优先考虑高层次规划或空间推理,忽视了具身物理交互所需的细粒度动作智能。为解决这一差距,我们引入了CFG-Bench,一个新基准测试,旨在系统评估这一关键能力。CFG-Bench包含1,368个精心挑选的视频和19,562个问题-答案对,涵盖三个评估范式,针对四种认知能力:1)物理交互,2)时间-因果关系,3)意图理解,4)评价判断。这些维度提供了系统评估模型将视觉观察转化为行动知识能力的框架,超越了表面层面的识别。我们在CFG-Bench上的全面评估表明,领先的MLLMs在生成物理交互的详细指令方面存在困难,并在意图和评价的高层次推理中表现出显著的局限性。此外,对我们的数据进行监督微调(SFT)表明,教导MLLMs直接表达细粒度动作会显著提升现有具身基准测试的性能。我们的分析指出了这些局限性,并为开发更具能力和基础的具身智能体提供了见解。项目页面:https://cfg-bench.github.io/

英文摘要

Multimodal Large Language Models (MLLMs) show promising results as decision-making engines for embodied agents operating in complex, physical environments. However, existing benchmarks often prioritize high-level planning or spatial reasoning, leaving the fine-grained action intelligence required for embodied physical interaction underexplored. To address this gap, we introduce CFG-Bench, a new benchmark designed to systematically evaluate this crucial capability. CFG-Bench consists of 1,368 curated videos paired with 19,562 question-answer pairs spanning three evaluation paradigms targeting four cognitive abilities: 1) Physical Interaction, 2) Temporal-Causal Relation, 3) Intentional Understanding, and 4) Evaluative Judgment. Together, these dimensions provide a systematic framework for assessing a model's ability to translate visual observations into actionable knowledge, moving beyond mere surface-level recognition. Our comprehensive evaluation on CFG-Bench reveals that leading MLLMs struggle to produce detailed instructions for physical interactions and exhibit profound limitations in the higher-order reasoning of intention and evaluation. Moreover, supervised fine-tuning (SFT) on our data demonstrates that teaching an MLLMs to articulate fine-grained actions directly translates to significant performance gains on established embodied benchmarks. Our analysis highlights these limitations and offers insights for developing more capable and grounded embodied agents. Project page: https://cfg-bench.github.io/

URL PDF HTML 收藏
2509.22415 2026-07-16 cs.CV cs.AI 版本更新

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models

多模态大语言模型中用于视觉归因的证据重组与预测上下文残差化

Jiawei Liang, Jianjie Huang, Ruoyu Chen, Xianghao Jiao, Siyuan Liang, Shiming Liu, Xiaochun Cao

机构 * Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区) Zhongguancun Academy(中关村学院) University of Chinese Academy of Sciences(中国科学院大学) Nanyang Technological University(南洋理工大学) Department of Mechanical Engineering, Imperial College London(伦敦帝国理工学院机械工程系)

AI总结 研究多模态大语言模型token级视觉证据难检查问题,提出基于证据重组和预测上下文残差化的ERCR框架,经实验验证该框架能改善目标token视觉证据、减轻上下文干扰,为视觉证据检查提供实用改进。

详情
AI中文摘要

多模态大语言模型(MLLMs)虽取得了强大的视觉语言性能,但其token级视觉证据难以检查。近期的logit-lens归因方法将每个视觉token隐藏状态投射到词汇空间来解释生成的单词,但这种token-wise读出存在不匹配。我们提出了ERCR归因框架,由证据重组(ER)和预测上下文残差化(PCR)构建。ER跨多个视图聚合目标证据,减少单一读出网格导致的归因碎片化。PCR估计前序token上下文图并从ER图中减去其拟合分量以抑制上下文token干扰。实验表明ERCR改善了目标token的视觉证据并减轻了前序token上下文干扰。

英文摘要

Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect. Recent logit-lens attribution methods project each visual-token hidden state into the vocabulary space to explain generated words, but this token-wise readout introduces a mismatch: visual tokens are context-mixed by the model, while the attribution score is decoded independently at each token location. This often produces fragmented attribution maps and can be further affected by autoregressive context signals from preceding text tokens. We propose ERCR, an attribution framework built from Evidence Recomposition (ER) and Predictive Context Residualization (PCR). ER aggregates target evidence across multiple views with different token-to-region assignments, reducing attribution fragmentation caused by a single readout grid. PCR estimates a preceding-token context map with RBO-based rank relevance and subtracts its fitted component from the ER map to suppress context-token interference. Experiments on LLaVA, Qwen2-VL, and InternVL families across COCO Caption, GranDf, and OpenPSG show that ERCR improves visual evidence for target tokens and mitigates preceding-token context interference under the existing evaluation protocol. On Qwen2-VL-2B, ERCR improves TAM F1-IoU from 39.10 to 44.45 on COCO Caption and from 30.83 to 37.20 on GranDf. Overall, ERCR provides a practical refinement for token-level visual evidence inspection.

URL PDF HTML 收藏
2506.20683 2026-07-16 eess.IV cs.AI cs.CV eess.SP

Global and Local Contrastive Learning for Joint Representations from Cardiac MRI and ECG

用于心脏MRI和ECG联合表示的全局和局部对比学习

Alexander Selivanov, Philip Müller, Özgün Turgut, Nil Stolt-Ansó, Daniel Rückert

机构 * Chair for AI in Healthcare and Medicine, Technical University of Munich (TUM) and TUM University Hospital, Munich, Germany(人工智能在医疗与医学中的研究中心,技术大学慕尼黑(TUM)及慕尼黑技术大学医院) School of Medicine, Klinikum rechts der Isar, TUM, Germany(医学院,右岸克里克医院,TUM,德国) Department of Computing, Imperial College London, UK(计算学院,伦敦帝国学院,英国) Munich Center for Machine Learning (MCML), Munich, Germany(慕尼黑机器学习中心(MCML),慕尼黑,德国)

AI总结 PTACL通过结合心脏MRI的时空信息提升ECG表示,实现患者级和时间级对比学习,提高心脏表型检索和功能参数预测的性能。

Comments accepted to MICCAI 2025 (Springer LNCS)

Journal ref Medical Image Computing and Computer Assisted Intervention - MICCAI 2025, Lecture Notes in Computer Science, vol. 15960, pp. 217-227, Springer, Cham (2026)

详情
AI中文摘要

心电图(ECG)是一种广泛使用的、成本效益高的工具,用于检测心脏的电异常。然而,它不能直接测量功能参数,如心室容积和射血分数,这些参数对评估心脏功能至关重要。心脏磁共振成像(CMR)是这些测量的金标准,提供详细的结构和功能信息,但成本高且可及性低。为弥合这一差距,我们提出了PTACL(患者和时间对齐对比学习),一种多模态对比学习框架,通过整合来自CMR的时空信息来增强ECG表示。PTACL使用全局患者级对比损失和局部时间级对比损失。全局损失通过将来自同一患者的ECG和CMR嵌入拉近,同时将不同患者的嵌入推远,对齐患者级表示。局部损失通过对比编码的ECG段与对应的编码CMR帧,强制在每个患者内部进行精细的时间对齐。这种方法通过超越仅靠全局对齐的对比学习,使ECG表示丰富化,获得诊断信息,同时在模态之间转移更多的见解。我们在英国生物银行中27,951名受试者的配对ECG-CMR数据上评估了PTACL。与基线方法相比,PTACL在两个临床相关任务中表现更好:(1)检索具有相似心脏表型的患者;(2)预测CMR衍生的心脏功能参数,如心室容积和射血分数。我们的结果强调了PTACL在利用ECG进行非侵入性心脏诊断中的潜力。代码可在:https://github.com/alsalivan/ecgcmr 上获得。

英文摘要

An electrocardiogram (ECG) is a widely used, cost-effective tool for detecting electrical abnormalities in the heart. However, it cannot directly measure functional parameters, such as ventricular volumes and ejection fraction, which are crucial for assessing cardiac function. Cardiac magnetic resonance (CMR) is the gold standard for these measurements, providing detailed structural and functional insights, but is expensive and less accessible. To bridge this gap, we propose PTACL (Patient and Temporal Alignment Contrastive Learning), a multimodal contrastive learning framework that enhances ECG representations by integrating spatio-temporal information from CMR. PTACL uses global patient-level contrastive loss and local temporal-level contrastive loss. The global loss aligns patient-level representations by pulling ECG and CMR embeddings from the same patient closer together, while pushing apart embeddings from different patients. Local loss enforces fine-grained temporal alignment within each patient by contrasting encoded ECG segments with corresponding encoded CMR frames. This approach enriches ECG representations with diagnostic information beyond electrical activity and transfers more insights between modalities than global alignment alone, all without introducing new learnable weights. We evaluate PTACL on paired ECG-CMR data from 27,951 subjects in the UK Biobank. Compared to baseline approaches, PTACL achieves better performance in two clinically relevant tasks: (1) retrieving patients with similar cardiac phenotypes and (2) predicting CMR-derived cardiac function parameters, such as ventricular volumes and ejection fraction. Our results highlight the potential of PTACL to enhance non-invasive cardiac diagnostics using ECG. The code is available at: https://github.com/alsalivan/ecgcmr

URL PDF HTML 收藏
2607.12112 2026-07-15 cs.LG cs.AI cs.CV cs.DC 新提交

Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning

用于联邦多模态大语言模型微调的弹性正则化和合成重放持续学习

Jing Liu, Chenxuanyin Zou, Jiayang Ren, Gaoyun Fang, Chengfang Li, Yan Wang, Zhenchao Ma, Bo Hu

机构 * The University of British Columbia(英属哥伦比亚大学) Fudan University(复旦大学) Royal College of Science, Imperial College London(伦敦帝国理工学院皇家科学学院) Dyson School of Design Engineering(戴森设计工程学院) Suzhou Institute of Biomedical Engineering and Technology (SIBET), Chinese Academy of Sciences(中国科学院苏州生物医学工程技术研究所) East China Normal University(华东师范大学)

AI总结 研究针对联邦多模态大语言模型微调中灾难性遗忘问题,提出FedCMM框架,在参数、数据、聚合三个层面嵌入持续学习保障,经实验验证该框架在准确性和反向迁移上优于基线,能实现跨异构网络AI部署的稳健进化适应。

Comments submitted to IEEE JSTSP

详情
AI中文摘要

跨分布式网络对多模态大语言模型(MLLM)进行联邦微调,可在保护隐私的情况下适应不断变化的数据流,但灾难性遗忘阻碍其在动态环境中的稳健部署。为应对这一挑战,我们提出了联邦持续多模态学习(FedCMM)框架,它在三个互补层面将持续学习保障嵌入联邦优化循环。在参数层面,模态感知弹性权重整合为视觉编码器、语言主干和跨模态投影仪计算单独的Fisher信息矩阵;在数据层面,每个客户端训练一个轻量级本地生成重放模块来合成无原始数据的嵌入级多模态重放元组;在聚合层面,任务相似性感知梯度聚合通过梯度余弦相似性自主过滤和重新加权客户端更新。实验表明,FedCMM在准确性和反向迁移方面优于近期基线,证实了整体、模态感知优化可实现跨异构网络AI部署的稳健进化适应。

英文摘要

Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust deployment in dynamic environments: catastrophic forgetting, wherein sequential task updates erase previously acquired knowledge across visual, linguistic, and cross-modal representations. Addressing this challenge is especially critical for autonomous networked AI operating in safety-sensitive domains, such as content moderation, where reliable retention of prior knowledge underpins system integrity. To overcome this, we propose Federated Continual Multimodal Learning (FedCMM), a framework that embeds continual-learning safeguards into the federated optimization loop at three complementary levels. At the parameter level, modality-aware elastic weight consolidation computes separate Fisher information matrices for the vision encoder, language backbone, and cross-modal projector, providing granular, asymmetry-aware protection against modality-specific forgetting. At the data level, each client trains a lightweight local generative replay module to synthesize raw-data-free embedding-level multimodal replay tuples without any raw data sharing. At the aggregation level, Task-similarity-aware gradient aggregation autonomously filters and reweights client updates by gradient cosine similarity, suppressing conflicting directions and stabilizing the global learning trajectory. Extensive experiments on two benchmarks demonstrate that FedCMM consistently outperforms recent baselines on accuracy and backward transfer, confirming that holistic, modality-aware optimization enables robust evolutive adaptation across heterogeneous networked AI deployments.

URL PDF HTML 收藏
2603.01568 2026-07-15 cs.LG cs.CV cs.IT math.IT q-bio.NC 版本更新

Same Compression Principle, Different Geometry: Rate-Distortion Signatures Dissociate Biological and Artificial Visual Systems

泛化与信息权衡的率-失真签名

Leyla Roksan Caglar, Pedro A. M. Mediano, Baihan Lin

机构 * Windreich Department of AI Human Health, Icahn School of Medicine at Mount Sinai, New York, NY, USA Department of Computing, Imperial College London, London, UK Department of Psychiatry, Icahn School of Medicine at Mount Sinai, New York, NY, USA Department of Neuroscience, Icahn School of Medicine at Mount Sinai, New York, NY, USA Berkman Klein Center for Internet \& Society, Harvard University, Cambridge, MA, USA

AI总结 本文提出率-失真理论框架,通过斜率和曲率签名分析系统泛化与鲁棒性权衡,揭示生物与人工系统在RD空间中的不同表现。

详情
AI中文摘要

向新的视觉条件泛化仍然是人类和机器视觉面临的重大挑战,然而标准的鲁棒性度量在揭示系统如何在准确性与鲁棒性之间进行权衡方面提供有限的见解。我们引入了一个率-失真理论框架,将刺激-响应行为视为一个有效的通信通道,从混淆矩阵中推导出率-失真(RD)前沿,并用两个可解释的几何签名——斜率(β)和曲率(κ)——来总结每个系统,这些签名捕捉了准确率-鲁棒性权衡的边际成本和突变性。将此框架应用于人类心理物理学和18个深度视觉模型在受控图像扰动下的表现,我们比较了不同模型架构和训练制度下的泛化几何结构。我们发现,生物和人工系统都遵循一种损失性压缩原则,但系统地占据RD空间的不同区域。特别是,人类表现出更平滑、更灵活的权衡,而现代深度网络即使在匹配的准确性下也处于更陡峭且更脆弱的区域。在不同的训练制度下,鲁棒性训练会引发系统但可区分的β/κ变化,揭示了某些情况下改进的鲁棒性或准确性并不转化为更像人类的泛化几何结构。这些结果表明,RD几何提供了紧凑且模型无关的视角,用于比较系统之间的泛化行为,超越了标准的准确性度量。

英文摘要

Efficient coding theory predicts that biological perceptual systems compress sensory input optimally under resource constraints, with the systematic structure of errors reflecting the geometry of that compression. Here we operationalize this principle using rate-distortion theory (RDT) to characterize how any system - biological or artificial - trades representational fidelity for informational efficiency. Treating stimulus-response behavior as an effective communication channel, we infer rate-distortion (RD) frontiers directly from confusion matrices and summarize each system with three geometric signatures: slope (beta), curvature (kappa), and area under the RD curve (AUC), capturing the marginal cost, abruptness, and overall efficiency of the accuracy-compression trade-off respectively. Applying this framework to human psychophysical data and 18 deep vision models across 12 families of controlled image perturbations at graded severities, we find that both biological and artificial systems follow a common lossy-compression principle but occupy systematically different regions of RD space. Humans exhibit smooth, flexible trade-offs characteristic of near-optimal efficient coding, while deep networks operate in steeper, more brittle regimes even at matched accuracy, with geometry dissociable from performance across training regimes. Critically, behavioral RD signatures track internal representational geometry, evidenced by the behaviorally inferred compression structure correlating with internal representational dissimilarity across all models. These results establish RD geometry as a compact diagnostic of perceptual compression strategy that recovers mechanistically interpretable structure in internal representations from behavioral input alone and extends naturally to the direct characterization of compression geometry in neural population activity.

URL PDF HTML 收藏
2512.04144 2026-07-15 cs.AI 版本更新

RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories

RippleBench: 利用现有知识库捕捉涟漪效应

Roy Rinberg, Usha Bhalla, Igor Shilov, Flavio P. Calmon, Rohit Gandikota

机构 * Harvard University(哈佛大学) Imperial College London(伦敦帝国学院) Northeastern University(东北大学)

AI总结 提出RippleBench-Maker自动管道,从知识库检索语义邻居生成选择题,评估八种遗忘方法在Llama3-8B-Instruct上的涟漪效应,发现准确率下降随语义距离衰减且跨模型一致。

详情
AI中文摘要

针对语言模型的目标干预,如遗忘或模型编辑,旨在修改特定信息,但其效果往往传播到相关的、非预期的领域(例如,删除病毒学内容可能降低对过敏任务的性能);这些副作用通常被称为涟漪效应。我们引入RippleBench-Maker,一个自动管道,从知识库中检索任何源概念的语义邻居,并生成不同语义距离的多选题。我们使用WikiRAG(一个基于英文维基百科的开源RAG系统)实例化该框架,构建RippleBench-WMDP-Bio(584个种子主题,352,961个问题),并在Llama3-8B-Instruct上评估八种遗忘方法。所有八种方法在遗忘目标附近准确率下降最大,并随语义距离衰减,每种方法具有不同的传播曲线。我们在Mistral-7B、Zephyr-7B和Yi-34B上复现了这些发现;跨模型的差值曲线几乎相同,表明涟漪效应是遗忘方法的属性而非基础模型。我们通过一项包含四个实验的Mechanical Turk研究(5,200+次响应,61名工作者)验证了所有主要管道阶段。我们发布所有代码、数据和基础设施。

英文摘要

Targeted interventions on language models, such as unlearning or model editing, aim to modify specific information, but their effects often propagate to related, unintended areas (e.g., removing virology content may degrade performance on allergies); these side-effects are commonly referred to as the ripple effect. We introduce RippleBench-Maker, an automatic pipeline that retrieves semantic neighbors of any source concept from a knowledge repository and generates multiple-choice questions at varying semantic distances. We instantiate this framework using WikiRAG, an open-source RAG system over English Wikipedia, to construct RippleBench-WMDP-Bio (584 seed topics, 352,961 questions), and evaluate eight unlearning methods on Llama3-8B-Instruct. All eight exhibit accuracy drops that are largest near the unlearned target and decay with semantic distance, each with a distinct propagation profile. We replicate these findings across Mistral-7B, Zephyr-7B, and Yi-34B; cross-model delta curves are nearly identical, suggesting ripple effects are a property of the unlearning method rather than the base model. We validate all major pipeline stages using a four-experiment Mechanical Turk study (5,200+ responses, 61 workers). We release all code, data, and infrastructure.

URL PDF HTML 收藏
2312.17670 2026-07-15 cs.CV cs.LG q-bio.QM q-bio.TO 版本更新

The TopCoW Challenge -- Topology-Aware Circle of Willis Segmentation for CT and MR Angiography

TopCoW挑战——用于CT和MR血管造影的拓扑感知Willis环分割

Kaiyuan Yang, Fabio Musio, Yihui Ma, Norman Juchler, Johannes C. Paetzold, Rami Al-Maskari, Luciano Höher, Hongwei Bran Li, Ibrahim Ethem Hamamci, Anjany Sekuboyina, Suprosanna Shit, Houjing Huang, Chinmay Prabhakar, Ezequiel de la Rosa, Bastian Wittmann, Diana Waldmannstetter, Florian Kofler, Fernando Navarro, Martin J. Menten, Ivan Ezhov, Daniel Rueckert, Iris N. Vos, Ynte M. Ruigrok, Birgitta K. Velthuis, Hugo J. Kuijf, Pengcheng Shi, Wei Liu, Ting Ma, Maximilian R. Rokuss, Yannick Kirchhoff, Fabian Isensee, Klaus Maier-Hein, Chengcheng Zhu, Huilin Zhao, Philippe Bijlenga, Julien Hämmerli, Catherine Wurster, Laura Westphal, Jeroen Bisschop, Elisa Colombo, Hakim Baazaoui, Hannah-Lea Handelsmann, Andrew Makmur, James Hallinan, Amrish Soundararajan, Benedikt Wiestler, Jan S. Kirschke, Evamaria O. Riedel, Roland Wiest, Emmanuel Montagnon, Laurent Letourneau-Guillon, Kwanseok Oh, Dahye Lee, Orhun Utku Aydin, Adam Hilbert, Jana Rieger, Dimitrios Rallios, Satoru Tanioka, Alexander Koch, Dietmar Frey, Abdul Qayyum, Moona Mazher, Steven Niederer, Nico Disch, Julius C. Holzschuh, Dominic LaBella, Francesco Galati, Daniele Falcetta, Maria A. Zuluaga, Chaolong Lin, Haoran Zhao, Zehan Zhang, Minghui Zhang, Xin You, Hanxiao Zhang, Guang-Zhong Yang, Yun Gu, Sinyoung Ra, Jongyun Hwang, Hyunjin Park, Junqiang Chen, Marek Wodzinski, Henning Müller, Nesrin Mansouri, Florent Autrusseau, Cansu Yalcin, Rachika E. Hamadache, Clara Lisazo, Joaquim Salvi, Adrià Casamitjana, Xavier Lladó, Uma Maria Lal-Trehan Estrada, Valeriia Abramova, Luca Giancardo, Arnau Oliver, Paula Casademunt, Adrian Galdran, Matteo Delucchi, Oscar Camara, Jialu Liu, Haibin Huang, Yue Cui, Zehang Lin, Yusheng Liu, Shunzhi Zhu, Tatsat R. Patel, Adnan H. Siddiqui, Vincent M. Tutino, Maysam Orouskhani, Huayu Wang, Mahmud Mossa-Basha, Yuki Sato, Sven Hirsch, Susanne Wegener, Bjoern Menze

机构 * Department of Quantitative Biomedicine, University of Zurich, Zurich, Switzerland Institute of Computational Life Sciences, Zurich University of Applied Sciences (ZHAW), Waedenswil, Switzerland Department of Neuroradiology, University Hospital of Zurich, Zurich, Switzerland Department of Neurosurgery, Zhongnan Hospital of Wuhan University, Wuhan, China Department of Radiology at Weill Cornell Medicine, Cornell University, New York, USA Institute for Tissue Engineering School of Computation, Information Technology, Technical University of Munich, Germany Athinoula A. Martinos Center for Biomedical Imaging, Harvard Medical School, Boston, USA School of Medicine Health, TUM Klinikum, Technical University of Munich, Germany Munich Center for Machine Learning, Munich, Germany Department of Computing, Imperial College London, London, UK Image Sciences Institute, UMC Utrecht, Utrecht, The Netherlands Department of Neurology Neurosurgery, University Medical Center Utrecht, Utrecht, The Netherlands Department of Radiology, University Medical Center Utrecht, Utrecht, The Netherlands Electronic \& Information Engineering School, Harbin Institute of Technology (Shenzhen), China Peng Cheng Laboratory, Shenzhen, China Division of Medical Image Computing, German Cancer Research Center (DKFZ), Heidelberg, Germany Faculty of Mathematics Computer Science, Heidelberg University, Germany Helmholtz Imaging, German Cancer Research Center, Heidelberg, Germany Data Science School for Health, Karlsruhe/Heidelberg, Germany Learning Group, Department of Radiation Oncology, Heidelberg University Hospital Department of Radiology, University of Washington, Seattle, WA, USA Department of Radiology, Ren Ji Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China Department of Clinical Neurosciences, Division of Neurosurgery, Geneva University Hospitals, Geneva, Switzerland Department of Neurology, University Hospital of Zurich, Zurich, Switzerland Department of Physiology, University of Toronto, Canada Department of Neurosurgery, University Hospital of Zurich, Zurich, Switzerland Department of Diagnostic Imaging, National University Hospital, Singapore University of Chicago, USA Department of Diagnostic Interventional Neuroradiology, University Hospital Berne University of Berne, Berne, Switzerland Centre de Recherche du Centre Hospitalier de l’Université de Montréal (CRCHUM), Montréal, Québec, Canada DEEPNOID Inc., Seoul, South Korea Department of Artificial Intelligence, Korea University, Seoul, South Korea Charité Lab for AI in Medicine (CLAIM), Charité Universitätsmedizin Berlin, Berlin, Germany Lung Institute, Faculty of Medicine, Imperial College London, London, UK Centre for Medical Image Computing, Department of Computer Science, University College London, London, UK Department of Radiation Oncology, Duke University Medical Center, Durham, NC, USA Institute of Medical Technology, Peking University Health Science Center, Beijing, China Hangzhou Genlight MedTech Co., Ltd., China Institute of Medical Robotics, Shanghai Jiao Tong University, Shanghai, China Department of Automation, Shanghai Jiao Tong University, Shanghai, China Department of Artificial Intelligence, Sungkyunkwan University, Seoul, South Korea Department of Electrical Computer Engineering, Sungkyunkwan University, Seoul, South Korea Shanghai MediWorks Precision Instruments Co., Ltd., China Institute of Informatics, HES-SO Valais-Wallis, Switzerland Department of Measurement Electronics, AGH University of Krakow, Poland Laboratoire de Thermique et Energie de Nantes (LTeN), Université Nantes, Polytech’Nantes, Nantes, France Research Institute of Computer Vision Center for Precision Health, McWilliams School of Biomedical Informatics, University of Texas Health Science Center at Houston, USA Physense, BCN-Medtech, Department of Communication Information Technologies, Universitat Pompeu Fabra, Barcelona, Spain Department of Mathematical Modeling Machine Learning, University of Zurich, Zurich, Switzerland Laboratory of Brain Atlas Brain-inspired Intelligence, Institute of Automation, Chinese Academy of Sciences, Beijing, China School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China School of Computer Information Engineering, Xiamen University of Technology, Xiamen, China Vascular Research Center, University at Buffalo, NY, USA Department of Pathology Anatomical Sciences, University at Buffalo, NY, USA Department of Neurosurgery, University at Buffalo, NY, USA LPIXEL Inc., Tokyo, Japan

AI总结 组织TopCoW基准挑战,发布含125对MRA和CTA扫描的注释数据集,参与者提交CoW分割和变体分类算法,经评估,最佳算法在多任务中表现出色,证明CoW分割算法对下游临床应用有可解释性效用。

Comments Summary paper for the TopCoW Challenge: 4 figures, 1 table, and supplementary material in appendix. Accepted for publication in NEJM AI. Datasets and best-performing algorithm Dockers are available at https://zenodo.org/records/15692630 and https://zenodo.org/records/15665435

详情
AI中文摘要

Willis环(CoW)是连接大脑主要循环的重要动脉网络。其血管结构被认为会影响严重神经血管疾病的风险、严重程度和结果。然而,表征高度可变的CoW解剖结构仍然是一项人工且耗时的专家任务。CoW通常通过磁共振血管造影(MRA)和计算机断层血管造影(CTA)这两种非侵入性血管造影成像方式进行成像,但带注释的CoW解剖结构数据集很少,也没有用于比较CoW分割算法的既定基准。我们组织了TopCoW基准挑战,并发布了一个带注释的CoW数据集,其中包含来自同一患者的125对MRA和CTA扫描。使用虚拟现实技术创建了13个血管成分的体素级注释,并由临床专家进行了验证。参与者提交了CoW分割和变体分类算法,我们在包含来自五个以上中心的226次扫描的内部和外部测试集上进行了评估。该基准包括体素级分割、CoW成分检测、CoW变体分类和两个临床应用任务。我们收到了来自六大洲250多名参与者的提交。表现最佳的团队在几乎所有测试集中,CoW分割的Dice分数超过90%,关键血管成分检测的F1分数超过80%,CoW变体分类中的平衡准确率超过70%。最佳算法还通过准确分类胎儿型大脑后动脉并定位与CoW解剖结构相关的动脉瘤,支持了临床相关的下游任务。这个基准证明了CoW分割算法在一些具有可解释性的下游临床应用中的效用。

英文摘要

The Circle of Willis (CoW) is an important network of arteries connecting major circulations of the brain. Its vascular architecture is believed to influence the risk, severity, and outcome of serious neurovascular diseases. However, characterizing the highly variable CoW anatomy remains a manual and time-consuming expert task. The CoW is commonly imaged by two non-invasive angiographic imaging modalities, magnetic resonance angiography (MRA) and computed tomography angiography (CTA), yet few datasets with annotated CoW anatomy exist, and there have been no established benchmarks for comparing CoW segmentation algorithms. We organized the TopCoW benchmark challenge alongside the release of an annotated CoW dataset with 125 paired MRA and CTA scans from the same patients. Voxel-level annotations for 13 vessel components were created using virtual reality technology and verified by clinical experts. Participants submitted algorithms for CoW segmentation and variant classification, which we evaluated on internal and external test sets comprising 226 scans from over five centers. The benchmark includes voxel-level segmentation, CoW component detection, CoW variant classification, and two clinical application tasks. We received submissions from over 250 participants across six continents. Top-performing teams achieved over 90% Dice scores for CoW segmentation, over 80% F1 scores for detecting key vessel components, and over 70% balanced accuracy in CoW variant classification across nearly all test sets. The best algorithms also supported clinically relevant downstream tasks by accurately classifying fetal-type posterior cerebral arteries and localizing aneurysms in relation to CoW anatomy. This benchmark demonstrated the utility of CoW segmentation algorithms for some downstream clinical applications with explainability.

URL PDF HTML 收藏
2405.19466 2026-07-15 cs.LG stat.ML 版本更新

Active Exploration via Autoregressive Generation of Missing Data

通过自回归生成缺失数据进行主动探索

Tiffany Tianhui Cai, Hongseok Namkoong, Daniel Russo, Kelly W Zhang

机构 * Columbia University(哥伦比亚大学) Imperial College London(帝国理工学院)

AI总结 将在线决策中的不确定性量化和探索问题转化为自回归序列模型的训练与生成,通过预测缺失结果而非潜在参数来建模不确定性,理论证明在线学习可归约为离线下一结果预测,并在新闻推荐中验证有效性。

详情
AI中文摘要

我们将在线决策中的不确定性量化和探索问题视为自回归序列模型的训练与生成问题,该领域正在经历快速创新。我们的方法将不确定性视为由于通过行动选择可能揭示的缺失未来结果而产生,而非来自环境不可观测的潜在参数。这种重新表述自然地与现代机器学习能力相一致:我们可以i)通过下一结果预测训练生成模型,而非拟合显式先验;ii)通过自回归生成评估不确定性,而非从后验中采样潜在参数;iii)通过扩展序列模型的上下文适应新信息,而非显式后验更新。我们的主要理论结果建立了从在线学习到离线下一结果预测的归约,表明贝叶斯遗憾由离线序列预测损失控制。半合成实验表明,我们的见解在具有挑战性的新闻推荐设置中成立,其中有效性能需要利用文章标题文本作为先验信息,将探索聚焦于解决剩余不确定性。

英文摘要

We pose uncertainty quantification and exploration in online decision-making as a problem of training and generation from an autoregressive sequence model, an area experiencing rapid innovation. Our approach rests on viewing uncertainty as arising from missing future outcomes that could be revealed through action choices, rather than from unobservable latent parameters of the environment. This reformulation aligns naturally with modern machine learning capabilities: we can i) train generative models through next-outcome prediction rather than fit explicit priors, ii) assess uncertainty through autoregressive generation rather than sampling latent parameters from posteriors, and iii) adapt to new information by extending the sequence model's context rather than explicit posterior updating. Our main theoretical result establishes a reduction from online decision-making to offline next-outcome prediction: Bayesian regret is controlled directly by the sequence model's offline prediction loss, without requiring an explicit latent-variable posterior. Experiments, including a semi-synthetic news recommendation task, show that autoregressive generation produces calibrated epistemic uncertainty and enables effective exploration by using article text as prior information to focus exploration on resolving remaining uncertainties.

URL PDF HTML 收藏
2607.11654 2026-07-14 cs.RO cs.SE 新提交

A Model for Mediating Multi-Modal Human Intent into Safe Maneuvers for UAVs

一种将多模态人类意图转化为无人机安全机动的中介模型

Sofia Nelson, Dalal Alrajeh, Pedro Antonio Alarcon Granadeno, Jane Cleland-Huang

机构 * University of Notre Dame(诺丁汉大学) Imperial College London(伦敦帝国理工学院)

AI总结 研究如何将多模态人类意图转化为无人机安全机动,提出需求导向的机动响应模型,通过结构化管道处理操作员输入,经多种约束验证后执行,还形式化为此类规范模型,经实验室验证可可靠解释并安全执行相关输入。

Comments 11 pages, 4 figures, preprint for MODRE 2026

详情
AI中文摘要

通过语音、手势和图形界面等方式可实现人类与自主无人机系统的直接交互。但将此类输入直接作为可执行命令会在动态环境中带来安全风险。本文提出一种需求导向的机动响应模型,将多模态人类意图转化为安全的无人机机动。把操作员输入视为有界机动请求,映射到受限运动原语,经结构化请求-评估-执行管道处理。每个请求都有置信度,针对多种约束进行验证,在持续运行监控下进行约束、拒绝或执行。还将该方法形式化为基于需求的规范模型,支持运行时验证等。通过基于实验室的初步验证,表明基于语音和GUI的输入能可靠解释并安全执行。

英文摘要

Direct human interaction with autonomous UAV systems can be enabled through modalities such as speech, gestures, and graphical interfaces. However, interpreting such inputs as directly executable commands introduces safety risks in dynamic environments. Operator requests may conflict with terrain constraints, inter-UAV separation requirements, or flight-envelope limitations. In this paper, we present a requirements-governed maneuver-response model that mediates multi-modal human intent into safe UAV maneuvers by treating operator inputs as bounded maneuver requests rather than direct commands. Requested maneuvers are mapped to constrained motion primitives and processed through a structured request-evaluate-execute pipeline. Each request is interpreted with associated confidence, validated against terrain, separation, workspace, and flight-envelope constraints, and either constrained, rejected, or executed under continuous runtime monitoring. We further formalize the approach as a requirements-based specification model in which maneuver primitives are associated with explicit preconditions, invariants, guard conditions, and postconditions governing admissibility, execution safety, and emergency handling. These requirements support runtime verification and future reactive synthesis approaches. We present an initial lab-based validation demonstrating that voice and GUI-based inputs can be reliably interpreted and safely executed as constrained maneuver requests.

URL PDF HTML 收藏
2607.11555 2026-07-14 cs.LG 新提交

Advancing Optimal Subset Oracle via Learning Relaxation of Neural Set Functions

通过学习神经集函数的松弛来推进最优子集预言机

Yongquan Shi, Zijing Ou, Shiping Wang, Yatao Bian

机构 * Fuzhou University(福州大学) Imperial College London(伦敦帝国学院) National University of Singapore(新加坡国立大学)

AI总结 研究神经集函数学习,针对现有最优子集预言机框架依赖蒙特卡罗采样估计梯度导致计算开销大且轨迹不稳定的问题,提出将证据下界重新解释为连续松弛并学习替代目标,实验证明该方法能减少开销、加速推理并优于现有基线。

详情
AI中文摘要

学习神经集函数对包括人工智能驱动的药物发现中的化合物选择和产品推荐在内的广泛重要应用至关重要。近期工作引入了最优子集预言机,在实际弱监督设置下隐式学习集函数,通过平均场变分推理优化模型参数。然而,这些框架在更新变分分布时依赖蒙特卡罗采样估计证据下界梯度。跨迭代重复采样会产生大量计算开销,且随机性会使优化轨迹不稳定。本文将证据下界重新解释为集函数的连续松弛,并学习一个替代目标,在变分优化期间取代基于采样的ELBO梯度估计。学习到的替代目标在整个连续域提供稳定高效的梯度,从而减少计算开销并加速推理。此外,我们为所提出框架在次模最大化下建立了近似保证,并刻画了其与变分自由能的联系。在各种实际任务上的实验表明,相对于现有基线有持续改进。

英文摘要

Learning neural set functions is pivotal to a wide range of important applications, including compound selection in AI-driven drug discovery and product recommendation. Recent work has introduced optimal subset oracles to implicitly learn set functions under practical weakly supervised settings, where model parameters are optimized through mean-field variational inference. However, these frameworks rely on Monte Carlo sampling to estimate gradients of the evidence lower bound when updating the variational distribution. Repeated sampling across iterations incurs substantial computational overhead, while the resulting stochasticity can destabilize the optimization trajectory. In this work, we reinterpret the evidence lower bound as a continuous relaxation of the set function and learn a surrogate objective that replaces sampling-based ELBO gradient estimation during variational optimization. The learned surrogate provides stable and efficient gradients throughout the continuous domain, thereby reducing computational overhead and accelerating inference. Furthermore, we establish an approximation guarantee for the proposed framework under submodular maximization and characterize its connection to variational free energy. Experiments on a variety of real-world tasks demonstrate consistent improvements over existing baselines.

URL PDF HTML 收藏
2607.11102 2026-07-14 cs.SD 新提交

CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection

CHARM:基于大语言模型的多模态讽刺检测中的电荷校准与声学救援

Qiyang Sun, Yi Chang, Yupei Li, Xi Shao, Zixing Zhang, Björn W. Schuller

机构 * GLAM – the Group on Language, Audio, & Music, Imperial College London(伦敦帝国理工学院语言、音频和音乐小组) College of Telecommunications and Information Engineering, Nanjing University of Posts and Telecommunications(南京邮电大学通信与信息工程学院) College of Computer Science and Electronic Engineering, Hunan University(湖南大学计算机科学与电子工程学院) Shenzhen Research Institute, Hunan University(湖南大学深圳研究院) CHI – Chair of Health Informatics, TUM University Hospital(慕尼黑工业大学医院健康信息学主席) relAI – the Konrad Zuse School of Excellence in Reliable AI(康拉德·楚泽可靠人工智能卓越学院) MDSI – Munich Data Science Institute(慕尼黑数据科学研究所) MCML – Munich Center for Machine Learning(慕尼黑机器学习中心)

AI总结 研究针对大语言模型在讽刺检测中过度预测积极类别及韵律线索利用不足的问题,提出CHARM框架,含双向电荷校准和声学后期融合救援两个模块,无需微调主干,提升检测性能,揭示跨文化韵律解耦,产生可解释的跨语言多模态检测器。

Comments under review

详情
AI中文摘要

讽刺检测是情感计算中的一项基本任务。然而,零样本指令微调的大语言模型(LLMs)在整个能力范围内系统性地过度预测积极(讽刺)类别,而人类依赖的韵律线索未得到充分利用且跨语言转移不均衡。我们引入了CHARM(用于多模态讽刺检测的电荷校准与声学救援),这是一个无需训练的框架,它结合了两个模块。双向电荷校准(BiCAL)沿着带电提示的对称轴引导大语言模型做出相反的讽刺和字面判断;诱导的方向偏差通过构造相互抵消,简单聚合可恢复无偏的语用信号。声学后期融合救援(ALFR)然后通过一个浅层分类器将校准后的投票与韵律描述符和大语言模型生成的听觉感知探针融合,积极降低饱和文本投票的权重以支持声学证据。在不微调任何主干的情况下,BiCAL在MUStARD上实现了报告的最高零样本纯文本Macro-F1为0.787,而ALFR在CMMA上使弱主干的Macro-F1提高了多达+0.382。斯托弗荟萃分析证实了在MUStARD和CMMA上的统计显著性。我们的分析还发现了跨文化韵律解耦:低级声学特征无法跨语言转移,而高级感知抽象则保持稳健。这些组件共同产生了一个可解释的跨语言多模态检测器。

英文摘要

Sarcasm detection, the identification of discrepancies between literal and intended meaning, is a fundamental task in affective computing. However, zero-shot instruction-tuned Large Language Models (LLMs) systematically over-predict the positive (sarcastic) class across the entire capability spectrum, while the prosodic cues humans rely on remain underexploited and transfer unevenly across languages. We introduce CHARM (Charge Calibration and Acoustic Rescue for Multimodal Sarcasm Detection), a training-free framework that couples two modules. Bidirectional Charge Calibration (BiCAL) steers the LLM toward opposing sarcastic and literal verdicts along a symmetric axis of charged prompts; the induced directional biases cancel by construction, and a simple aggregation recovers an unbiased pragmatic signal. Acoustic Late-Fusion Rescue (ALFR) then fuses the calibrated votes with prosodic descriptors and LLM-generated auditory-perception probes through a shallow classifier, actively down-weighting saturated text votes in favour of acoustic evidence. Without fine-tuning any backbone, BiCAL attains the highest reported zero-shot text-only Macro-F1 of 0.787 on MUStARD, while ALFR lifts weak backbones by up to +0.382 Macro-F1 on CMMA. A Stouffer meta-analysis confirms statistical significance on MUStARD and CMMA (Z = 13.89 and Z = 34.64, respectively; p < 10^-43). Our analysis further uncovers a cross-cultural prosodic decoupling: low-level acoustics fail to transfer across languages, whereas high-level perceptual abstractions remain robust. Together, these components yield an explainable, cross-lingual multimodal detector.

URL PDF HTML 收藏
2607.10891 2026-07-14 cs.AI 新提交

SETA: Scaling Environments for Terminal Agents

SETA:终端智能体的扩展环境

Qijia Shen, Zhiqi Huang, Vamsidhar Kamanuru, Aznaur Aliev, Jay Rainton, Ahmed Awelkair, Zhichen Zeng, Jiajun Li, Shi Dong, Yueming Yuan, Boyuan Ma, Qizheng Zhang, Jiwei Fu, Yuzhen Mao, Wendong Fan, Ping Nie, Philip Torr, Bernard Ghanem, Changran Hu, Jonathan Lingjie Li, Urmish Thakker, Guohao Li

机构 * Imperial College London(帝国理工学院) University College London(伦敦大学学院) SambaNova(桑巴诺瓦公司) KAUST(阿卜杜拉国王科技大学) Stanford University(斯坦福大学) University of Oxford(牛津大学) University of Waterloo(滑铁卢大学) RadixArk(基数方舟公司)

AI总结 研究针对终端智能体训练扩展难的问题,提出SETA框架,含SETA - Synth和SETA - Evol两个管道及统一验证机制,构建了SETA - Env数据集。实验显示该数据集能为终端智能体提供优质训练环境,推动相关研究发展。

详情
AI中文摘要

大语言模型正迅速向通过多种接口(包括网络和图形用户界面)解决任务的智能体转变。其中,终端命令行提供了基于文本的通用接口,涵盖从系统操作到数据科学和机器学习的任务。然而,扩展终端智能体训练仍具有挑战性,因为它需要多样且连贯的任务指令、可执行环境和可靠验证,同时缺乏自然基础的监督数据。在这项工作中,我们提出了SETA,这是一个用于为强化学习生成可验证终端环境的可扩展框架。该框架由两个共享统一验证机制的管道组成:SETA - Synth将各种来源转换为标准化的强化学习环境,SETA - Evol通过对难度和多样性的自适应控制从现有环境进一步扩展。我们共同构建并发布了SETA - Env,这是迄今为止最大的开源可验证终端强化学习数据集,包含超过4500个环境。我们通过在SETA - Env上使用GRPO训练Qwen3 - 8B来评估我们的数据集,在Terminal - Bench 2.0上实现了12%的通过率,这是8B规模的强化学习训练模型报告的最佳结果。我们还观察到在相同终端智能体框架下DeepSeek - V4 - Flash的收益,Terminal - Bench 2.0上的pass@1从40%提高到43%,pass@5从54%提高到58%。这些结果表明SETA - Env为终端智能体提供了高质量的训练环境,并作为推进基于终端的智能体学习研究的宝贵资源。

英文摘要

Large language models (LLMs) are rapidly shifting toward agents that solve tasks through diverse interfaces, including web and graphical user interfaces (GUIs). Among these, the terminal command line provides a text-based, general-purpose interface, covering tasks from system operations to data science and machine learning. However, scaling terminal-agent training remains challenging, as it requires diverse and coherent task instructions, executable environments, and reliable verification, while lacking naturally grounded supervision data. In this work, we propose SETA, a scalable framework for generating verifiable terminal environments for reinforcement learning (RL). The framework consists of two pipelines sharing a unified verification mechanism: SETA-Synth converts diverse sources into standardized RL environments, and SETA-Evol further expands from existing environments with adaptive control of difficulty and diversity. Together, we construct and release SETA-Env, the largest open-source verifiable terminal RL dataset to date, containing over 4,500 environments. We evaluate our dataset by training Qwen3-8B with GRPO on SETA-Env, achieving 12% pass rate on Terminal-Bench 2.0, the best reported result for an RL-trained model at the 8B scale. We further observe gains on DeepSeek-V4-Flash under the same terminal agent harness, with pass@1 on Terminal-Bench 2.0 improving from 40% to 43% and pass@5 improving from 54% to 58%. These results demonstrate that SETA- Env provides high-quality training environments for terminal agents and serves as a valuable resource for advancing research on terminal-based agent learning.

URL PDF HTML 收藏