arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Technion - Israel Institute of Technology(以色列理工学院)

至 收录 384
2607.17369 2026-07-21 cs.FL cs.DS cs.LG 新提交

Stringological sequence prediction II: Right-to-left automaticity and related complexity measures

字符串学序列预测 II:从右到左自动性及相关复杂度度量

Vanessa Kosoy

机构 * Faculty of Mathematics, Technion, Haifa, Israel(技术学院数学系, 哈ifa, 以色列) Computational Rational Agents Laboratory, Delaware, USA(计算理性代理实验室, 德拉瓦州, 美国)

AI总结 研究适用于字符串学复杂度度量的序列预测算法,展示了针对从右到左自动性的统计和计算高效算法,还演示了用于算术重复复杂度度量的预测算法,可用于预测混合自动序列。

Comments 73 pages

详情
AI中文摘要

在之前的论文中,我们开始研究适用于字符串学单词复杂度度量的序列预测算法。我们考虑的一种度量是从左到右(最高有效位优先)自动性。在此,我们展示了一种适用于“对偶”从右到左(最低有效位优先)自动性的统计和计算高效算法,事实证明该自动性在我们的目的下有很大不同。我们还演示了一种针对更具表现力的“算术重复复杂度”度量的预测算法。特别地,后者可用于预测所谓的混合自动序列。

英文摘要

In a previous paper, we began the study of sequence prediction algorithms adapted to stringological word complexity measures. One measure we considered was left-to-right (most-significant-digit-first) automaticity. Here, we show a statistically and computationally efficient algorithm adapted to the ``dual'' right-to-left (least-significant-digit-first) automaticity, which turns out to be substantially different for our purpose. We also demonstrate a prediction algorithm for a more expressive measure that we call ``arithmetic repetition complexity''. In particular, the latter can be used for predicting the so-called mix-automatic sequences.

URL PDF HTML 收藏
2607.17351 2026-07-21 cs.AI cs.RO 新提交

DeeperRadar: End-to-End MIMO Radar Design and Multi-Modal Fusion for Autonomous Vehicle Perception

深度雷达:用于自动驾驶车辆感知的端到端MIMO雷达设计与多模态融合

Eli Goldenshluger, Barak Pinkovich, Chaim Baskin

机构 * Technion–Israel Institute of Technology(以色列理工学院) Ben-Gurion University of the Negev(内盖夫本-古里安大学)

AI总结 深度雷达以雷达为中心,通过端到端学习稀疏采集模式,共同设计雷达传感与多模态3D检测。其可学习的MIMO设计模块在融合网络中训练,由其他传感器监督。在RADIal数据集上评估,能发现稀疏配置,降低成本与复杂性,表明最优设计取决于融合堆栈和感知任务。

Comments Accepted for publication at the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

详情
AI中文摘要

深度雷达是一个以雷达为中心、受传感器堆栈条件限制的框架,通过与融合模型端到端学习稀疏采集模式,共同设计用于自主移动的雷达传感和多模态3D检测。一个可学习的MIMO设计模块在融合网络中端到端训练,该网络直接对原始雷达ADC数据以及相机图像和激光雷达点云进行操作。训练期间,设计模块由其他传感器监督,使其能学习激活哪些接收天线及其有效数量。部署时,设计模块被学习到的稀疏子采样掩码取代,下游模型架构不变。在RADIal数据集上评估,深度雷达发现稀疏、任务感知的雷达配置,在使用更少接收器时匹配或超过全阵列基线,可能降低雷达成本和集成复杂性。这些结果表明,学习到的最优MIMO雷达设计取决于融合堆栈和下游感知任务。

英文摘要

DeeperRadar is a radar-centric, sensor-stack-conditioned framework that co-designs radar sensing and multi-modal 3D detection for autonomous mobility by learning a sparse acquisition pattern end-to-end with the fusion model. A learnable MIMO design module is trained end-to-end within a fusion network that operates directly on raw radar ADC data together with camera images and LiDAR point clouds. During training, the design module is supervised by the other sensors, enabling the system to learn both which receiver antennas to activate and the effective number of them. At deployment, the design module is removed and replaced by the learned sparse subsampling mask, leaving the downstream model architecture unchanged. Evaluated on the RADIal dataset, DeeperRadar discovers sparse, task-aware radar configurations that match or exceed full-array baselines while using fewer receivers, potentially reducing radar cost and integration complexity. These results show that learned optimal MIMO radar design depends on the fusion stack and the downstream perception task.

URL PDF HTML 收藏
2607.08793 2026-07-21 stat.ML cs.AI cs.LG cs.SY eess.SY math.OC 版本更新

EHR-MPC: Inference-Time Control for Sepsis Treatment with Generative Patient Digital Twins

EHR-MPC:利用生成式患者数字孪生进行脓毒症治疗的推理时间控制

Joshua Pickard, Wei Qi, Na Li, Ann Woolley, Lisa Cosimi, Roy Kishony, Deborah Hung

机构 * Broad Institute of MIT and Harvard(MIT和哈佛大学Broad研究所) Harvard University(哈佛大学) Brigham and Women’s Hospital(布莱根妇女医院) Technion–Israel Institute of Technology(技术学院-以色列理工学院)

AI总结 针对脓毒症治疗策略有争议且现有强化学习方法适应性不足的问题,提出EHR-MPC框架,通过训练生成式电子健康记录模型形式的患者数字孪生解耦学习与治疗优化,经模拟评估性能优于强化学习基线,建立了决策通用框架。

详情
AI中文摘要

脓毒症是主要死因,但最佳治疗策略仍有争议。现有强化学习方法学习固定策略,限制了推理时对变化临床目标的适应性。我们提出EHR-MPC框架,通过训练生成式电子健康记录模型形式的患者数字孪生,将学习患者动态与优化治疗解耦。数字孪生预测干预下的临床轨迹,使模型预测控制通过推理时模拟规划优化治疗。我们用离策略重要性采样和基于策略的模拟评估在多中心ICU脓毒症队列上评估EHR-MPC。相对于强化学习基线,EHR-MPC实现了可比的离策略性能和更好的模拟性能。不同于强化学习,这项工作将脓毒症治疗优化框架化为对学习到的患者动态的推理时控制,建立了使用生成式临床模型进行决策的通用框架。

英文摘要

Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested. Existing reinforcement learning (RL) approaches learn fixed strategies for sepsis treatment, limiting adaptability to changing clinical objectives during inference. We propose EHRMPC, a framework that decouples learning patient dynamics from optimizing treatment by training a patient digital twin in the form of a generative electronic health record (EHR) model. The digital twin predicts clinical trajectories under interventions and enables model predictive control (MPC) to optimize treatments via inference-time planning over simulations. We evaluate EHR-MPC on a multicenter ICU sepsis cohort spanning 8 hospitals in the Mass General Brigham health system using both off-policy importance sampling and on-policy simulation-based evaluation. Relative to RL baselines, EHR-MPC achieves comparable off-policy performance and improved simulation performance. Unlike RL, this work frames sepsis treatment optimization as inference-time control over learned patient dynamics, establishing a general framework for decision making with generative clinical models.

URL PDF HTML 收藏
2605.06605 2026-07-21 cs.LG 版本更新

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation

需要多少次迭代才能突破限制?多轮LLM评估中的动态预算分配

Shai Feldman, Yaniv Romano

机构 * Department of Computer Science(计算机科学系) Technion, Israel(技术ion, 以色列) Departments of Electrical and Computer Engineering and of Computer Science(电气与计算机工程系和计算机科学系)

AI总结 本文提出DAPRO框架,通过动态预算分配在多轮LLM交互中提供事件发生时间的界限,解决静态方法效率低的问题,实验表明其在覆盖性和方差方面优于传统方法。

详情
AI中文摘要

评估和预测大型语言模型(LLMs)在多轮对话设置中的性能至关重要但计算成本高;关键事件——例如突破限制或代理成功完成任务——往往在多次交互后才出现。这些事件可能罕见,在任何可行的计算预算下可能未被观察到。最近的符合生存框架构建了可靠的下界预测界限(LPBs)以确定触发事件所需的迭代次数,但依赖静态预算分配,在多轮设置中效率低下。为了解决这个问题,我们引入了动态分配通过投影优化(DAPRO),这是第一个理论上有效的动态预算分配框架,用于在多轮LLM交互中界定时间到事件。我们证明DAPRO满足预算约束,并提供分布无关、有限样本覆盖保证,而无需假设先前符合生存方法中条件独立性之间的截断和事件时间。关键理论贡献是新的覆盖界,其规模与均截断权重的平方根而不是最坏情况权重相关,从而比先前工作提供更紧的保证。此外,DAPRO可用于在有限计算资源下获得无偏、低方差的总体评估指标估计,如突破率。全面实验在代理任务成功、对抗性突破、有毒内容生成和RAG幻觉使用LLM如Llama 3.1和Qwen 2.5上显示,DAPRO在覆盖性和方差方面优于静态基线,同时满足预算约束。

英文摘要

Evaluating and predicting the performance of large language models (LLMs) in multi-turn conversational settings is critical yet computationally expensive; key events -- e.g., jailbreaks or successful task completion by an agent -- often emerge only after repeated interactions. These events might be rare, and under any feasible computational budget, remain unobserved. Recent conformal survival frameworks construct reliable lower predictive bounds (LPBs) on the number of iterations to trigger the event of interest, but rely on static budget allocation that is inefficient in multi-turn setups. To address this, we introduce \emph{Dynamic Allocation via PRojected Optimization} (DAPRO), the first theoretically valid dynamic budget allocation framework for bounding the time-to-event in multi-turn LLM interactions. We prove that DAPRO satisfies the budget constraint and provides distribution-free, finite-sample coverage guarantees without requiring the conditional independence between censoring and event times assumed by prior conformal survival approaches. A key theoretical contribution is a novel coverage bound that scales with the square root of the mean censoring weight rather than the worst-case weight, yielding provably tighter guarantees than prior work. Furthermore, DAPRO can be employed to obtain unbiased, low-variance estimates of population-level evaluation metrics, such as the jailbreak rate, under limited computing resources. Comprehensive experiments across agentic task success, adversarial jailbreaks, toxic content generation, and RAG hallucinations using LLMs such as Llama 3.1 and Qwen 2.5 demonstrate that DAPRO consistently achieves coverage closer to the nominal level with lower variance than static baselines, while satisfying the budget constraint.

URL PDF HTML 收藏
2603.17472 2026-07-17 cs.RO cs.MA cs.NI 版本更新

Bringing Network Coding into Multi-Robot Systems: Interplay Study for Autonomous Systems over Wireless Communications

将网络编码引入多机器人系统:自主系统在无线通信中的相互作用研究

Anil Zaher, Kiril Solovey, Alejandro Cohen

机构 * Viterbi Faculty of Electrical and Computer Engineering, Technion–Israel Institute of Technology(电气与计算机工程学院,技术学院–以色列理工学院)

AI总结 研究探讨了网络编码在多机器人系统中的应用,通过案例分析展示其在提升通信可靠性和任务执行效率方面的优势。

Comments To appear in IROS 2026

详情
AI中文摘要

通信是多机器人系统(MRS)的核心使能技术,为机器人交换状态信息、协调行动和满足安全约束提供机制。尽管许多MRS自主算法假设消息传输可靠且及时,但现实中的无线信道引入延迟、擦除和顺序停滞,可能影响性能并危及安全决策。本文研究了传输层可靠性机制如何影响自主-通信循环。传统非编码重传协议引入长延迟,与MRS应用的时间要求不匹配,可能导致接收数据无关。作为替代方案,本文倡导自适应和因果网络编码,通过主动注入编码冗余实现所需延迟和吞吐量,以支持机器人任务的数据交付。具体而言,该方法通过高效算法适应机器人间的信道条件,并通过因果调整通信速率。本文展示了两个案例研究:在延迟和丢失的机器人间通信下的协同定位,以及安全关键的超车 maneuver,其中及时的车辆间消息可用性决定自动驾驶车辆是否能紧急刹车以避免碰撞。结果表明,基于编码的通信显著减少了有序交付停滞,保持了在延迟下的估计一致性,并相对于基于重传的传输提高了截止时间的可靠性。总体而言,研究强调了自主算法和通信机制的联合设计需求,并将网络编码定位为无线网络上可靠多机器人操作的原理性工具。

英文摘要

Communication is a core enabler for multi-robot systems (MRS), providing the mechanism through which robots exchange state information, coordinate actions, and satisfy safety constraints. While many MRS autonomy algorithms assume reliable and timely message delivery, realistic wireless channels introduce delay, erasures, and ordering stalls that can degrade performance and compromise safety-critical decisions of the robot task. In this paper, we investigate how transport-layer reliability mechanisms that mitigate communication losses and delays shape the autonomy-communication loop. We show that conventional non-coded retransmission-based protocols introduce long delays that are misaligned with the timeliness requirements of MRS applications, and may render the received data irrelevant. As an alternative, we advocate for adaptive and causal network coding, which proactively injects coded redundancy to achieve the desired delay and throughput, enabling relevant data delivery for the robotic task. Specifically, this method adapts to channel conditions between robots and causally tunes the communication rates via efficient algorithms. We present two case studies: cooperative localization under delayed and lossy inter-robot communication, and a safety-critical overtaking maneuver where timely vehicle-to-vehicle message availability determines whether an ego vehicle can abort to avoid a crash. Our results demonstrate that coding-based communication significantly reduces in-order delivery stalls, keeps cooperative-localization accuracy close to the ideal baseline, and satisfies the overtaking abort deadline in 80% of the simulated runs, compared with 60% for a retransmission-based baseline. The study highlights the need to jointly design autonomy algorithms and communication mechanisms, and positions network coding as a principled tool for dependable MRS operation over wireless networks.

URL PDF HTML 收藏
2601.07599 2026-07-17 cs.CV 版本更新

Fundamental Recovery Bounds for SPAD Signals under Stationary Flux

SPAD信号中的扩散

Lior Dvir, Nadav Torem, Mohit Gupta, Yoav Y. Schechner

机构 * Viterbi Faculty of Electrical and Computer Engineering, Technion-Israel Institute of Technology(电气与计算机工程学院,技术学院-以色列理工学院)

AI总结 本文提出利用扩散模型表达图像先验,解决基于SPAD信号的逆问题,并探讨光子计数和检测事件时间对结果的影响。

详情
AI中文摘要

我们推导了在单光子雪崩二极管(SPAD)中给定固定光子通量下的原始信号似然性。原始信号包含检测事件的时间信息,这些时间信息与光子通量非线性相关。此外,这些事件本质上是随机的。我们随后推导了信号的分数函数。这是基于SPAD信号解决逆问题的关键。我们专注于推导涉及扩散模型的解决方案,以表达图像先验。我们展示了低或高光子计数的影响,以及利用检测事件时间信息的后果。

英文摘要

Single-photon avalanche diodes (SPADs) record light as a discrete stream of individual detections. The signal is stochastic. Its statistical structure depends on the sensor's operation mode: binary detection in fixed bins, timestamped detection in fixed bins, or free-running timestamped detection. We derive the likelihood score function for each of these three passive modes. From this single object, stem both fundamental limits of recovery (Cramer-Rao bounds) and practical recovery algorithms based on diffusion posterior sampling. The paper further generalizes fundamental limits to Bayesian Cramer-Rao lower bounds. This generalization makes use of a learned approximation of the score function of signal priors. In prior art, analyses and diffusion-based reconstruction for SPAD data have treated individual modes in isolation. Our unified treatment shows a qualitative high-flux gap between modes: binary counts saturate exponentially, while timestamped modes degrade only linearly. We further extend diffusion posterior sampling, previously restricted to binary SPAD data, to a full timestamped case using the suitable score function. We demonstrate experimentally that matching the score to the operation mode is beneficial for high-fidelity reconstruction. By tying the recovery bounds and diffusion to the score function, this work aims to establish a common foundation for both asking what is recoverable in single-photon sensing, and building methods that approach the bound.

URL PDF HTML 收藏
2607.13174 2026-07-16 physics.comp-ph cs.CE cs.RO 新提交

Towards end-to-end optimization in multimaterial 3D printing

迈向多材料3D打印中的端到端优化

Xue-Ling Luo, Steven Yang, Jingye Tan, Robert F. Shepherd, Noy Cohen, Nikolaos Bouklas

机构 * addressline= Sibley School of Mechanical Aerospace Engineering, Cornell University , city= Ithaca , state= NY , country= USA addressline= Department of Aerospace \& Mechanical Engineering, University of Southern California , city= Los Angeles , state= CA , country= USA addressline= Department of Materials Science Engineering, Technion - Israel Institute of Technology , city= Haifa , country= Israel addressline= Pasteur Labs , city= Brooklyn , state= NY , country= USA

AI总结 针对多材料3D打印中优化空间材料分布与结构拓扑的难题,提出将稀疏物理增强神经网络与有限元拓扑优化集成的端到端框架,通过提取本构定律实现精确微分,应用于软机器人抓手,可取代经验原型制作,建立实用的机器学习设计模型。

详情
AI中文摘要

多材料3D打印能够制造功能梯度部件,但由于高维设计空间和复杂的本构建模,在优化其空间材料分布和结构拓扑方面仍然是一个巨大挑战。本文提出了一个端到端计算框架,将稀疏物理增强神经网络与基于有限元的拓扑优化相结合。通过从实验数据中提取封闭形式、成分感知的超弹性本构定律,该方法借助FEniCSx实现的伴随状态法促进精确符号微分,有效规避应用神经网络本构模型的瓶颈。该流程应用于软机器人抓手应用,展示了针对高度各向异性接触响应的连续成分优化,以及在非失效拉伸约束下宏观拓扑和材料分布的协同优化。这种方法可以取代费力的经验原型制作,将可解释的机器学习模型确立为先进多材料增材制造实用、稳健的设计原语。

英文摘要

Multimaterial 3D printing enables the fabrication of functionally graded components, but optimizing their spatial material distribution alongside structural topology remains a formidable challenge due to high-dimensional design spaces and complex constitutive modeling. This paper presents an end-to-end computational framework integrating sparsified physics-augmented neural networks with finite-element-based topology optimization. By extracting closed-form, composition-aware hyperelastic constitutive laws from experimental data, this approach facilitates exact symbolic differentiation via the adjoint state method implemented with FEniCSx, efficiently circumventing the bottlenecks of applying neural network constitutive models. This pipeline is deployed on soft robotic gripper applications, demonstrating continuous composition optimization for highly anisotropic contact responses, and the concurrent optimization of macroscopic topology and material distribution under non-failure stretch constraints. This methodology could replace laborious empirical prototyping, establishing interpretable machine-learning models as practical, robust design primitives for advanced multimaterial additive manufacturing.

URL PDF HTML 收藏
2601.20496 2026-07-16 stat.ML cs.LG 版本更新

Leveraging Differentiable PDE Solvers for Semi-Neural Spatial Reconstruction From Sparse Measurements

利用可微偏微分方程求解器从稀疏测量中进行半神经空间重建

Ofek Aloni, Barak Fishbain

机构 * Department of Mathematics, Technion, Haifa, Israel(数学系,技术离子学院,海法,以色列) Department of Civil and Environmental Engineering, Technion, Haifa, Israel(土木与环境工程系,技术离子学院,海法,以色列)

AI总结 针对从稀疏测量生成密集物理场的问题,提出结合径向基函数重建、神经网络校正和可微偏微分方程求解器的混合建模方法,无需完全解析模拟状态示例训练神经网络,在流体力学基准测试中获优于其他方法的结果。

详情
AI中文摘要

从稀疏测量中生成密集物理场是采样、信号处理等众多应用中的基本问题。现有方法存在依赖忽略物理规律的空间统计、将物理规律集成到多目标优化过程或训练时需要完整模拟状态示例等问题,而这些示例在合成基准之外往往不可用。本文提出一种新方法,将径向基函数重建、神经网络校正和偏微分方程求解器相结合,构建混合建模管道,使数值模拟器直接嵌入学习组件的训练循环。通过实现可端到端微分的偏微分方程求解器,无需完全解析模拟状态示例即可训练神经网络。该灰箱方法在流体力学的三个标准基准上进行评估,取得优于基于统计和机器学习的重建方法的结果。

英文摘要

Generating dense physical fields from sparse measurements is a fundamental question in sampling, signal processing, and many other applications. State-of-the-art approaches to this problem either rely on spatial statistics that ignore the governing physics, integrate the physics into a multiple-objective optimization process, or require examples of the complete, fully-resolved simulation state during training, which are frequently unavailable outside of synthetic benchmarks. Here, we present a novel alternative that leverages recent advances in the integration of numerical simulators with data-driven models. Namely, we propose a hybrid modeling pipeline that couples Radial Basis Function (RBF) reconstruction with a Neural Network (NN) correction and a Partial Differential Equation (PDE) solver, so that the numerical simulator itself is embedded directly in the training loop of the learned component. Notably, the NN is trained without assuming availability of examples of the fully-resolved simulation state. This is made possible by implementing the PDE solver so that it is end-to-end differentiable, allowing gradients to be backpropagated through the simulation step during training. This grey-box methodology is evaluated on three standard benchmarks from fluid mechanics, where it achieves superior results over statistical and machine-learning-based reconstruction methods.

URL PDF HTML 收藏
2510.00180 2026-07-15 eess.AS cs.SD eess.SP 版本更新

DiffAU: Diffusion-Based Ambisonics Upscaling

DiffAU: 基于扩散的Ambisonics升阶

Amit Milstein, Nir Shlezinger, Boaz Rafaely

机构 * Technion - Israel Institute of Technology(技术学院 - 以色列理工学院)

AI总结 提出DiffAU方法,利用扩散模型和空间音频适配,从一阶Ambisonics生成三阶Ambisonics,实现快速可靠的升阶。

详情
AI中文摘要

空间音频通过再现3D声场增强沉浸感,Ambisonics为此提供了可扩展的格式。与高阶Ambisonics(HOA)相比,一阶Ambisonics(FOA)在硬件上高效地获取和存储声场,但其低空间分辨率限制了真实感,因此Ambisonics升阶(AU)作为增加Ambisonics信号阶数的方法显得尤为重要。本文提出DiffAU,一种级联的AU方法,利用扩散模型的最新进展并结合对空间音频的新颖适配,从FOA生成三阶Ambisonics。通过学习数据分布,DiffAU提供了一种原则性方法,能够在各种设置中快速可靠地再现HOA。在多个扬声器的消声条件下进行的实验,展示了强大的客观和感知性能。

英文摘要

Spatial audio enhances immersion by reproducing 3D sound fields, with Ambisonics offering a scalable format for this purpose. While first-order Ambisonics (FOA) notably facilitates hardware-efficient acquisition and storage of sound fields as compared to high-order Ambisonics (HOA), its low spatial resolution limits realism, highlighting the need for Ambisonics upscaling (AU) as an approach for increasing the order of Ambisonics signals. In this work we propose DiffAU, a cascaded AU method that leverages recent developments in diffusion models combined with novel adaptation to spatial audio to generate 3rd order Ambisonics from FOA. By learning data distributions, DiffAU provides a principled approach that rapidly and reliably reproduces HOA in various settings. Experiments in anechoic conditions with multiple speakers, show strong objective and perceptual performance.

URL PDF HTML 收藏
2605.14746 2026-07-14 cs.LG 版本更新

Selective Safety Steering via Value-Filtered Decoding

基于价值过滤解码的选择性安全引导

Bat-Sheva Einbinder, Hen Davidov, Yee Whye Teh, Yarin Gal, Yaniv Romano

机构 * Department of Electrical and Computer Engineering, Technion IIT(技术学院电气与计算机工程系) Department of Statistics, University of Oxford(牛津大学统计系) OATML, Department of Computer Science, University of Oxford(牛津大学计算机科学系) Department of Computer Science, Technion IIT(技术学院计算机科学系)

AI总结 本文提出一种测试时引导方法,通过价值基的安全标准过滤token,减少不必要的干预,提升不安全响应的安全性,同时在多个数据集上验证了其在安全与帮助性之间的更好平衡。

详情
AI中文摘要

尽管大型语言模型(LLMs)被训练以与人类价值观对齐,其生成仍可能违反安全约束。现有工作通过在解码时修改模型的采样策略来解决此问题,但这些方法往往干预过多,修改本应在基模型下安全的生成。本文提出一种新的测试时引导方法,旨在减少此类不必要的干预,同时提高不安全响应的安全性。我们的方法通过价值基的安全标准过滤token,并为误干预的概率提供明确界限。一个单一的阈值超参数控制此界限,使实践者能够权衡更高的不必要干预率以获得更好的输出安全性。在多个数据集和实验中,我们证明了我们的价值过滤解码方法优于现有基线,实现了在安全、帮助性和与基模型相似性之间的更好平衡。

英文摘要

While large language models (LLMs) are trained to align with human values, their generations may still violate safety constraints. A growing line of work addresses this problem by modifying the model's sampling policy at decoding time using a safety reward. However, existing decoding-time steering methods often intervene unnecessarily, modifying generations that would have been safe under the base model. Such unnecessary interventions are undesirable, as they can distort key properties of the base model such as helpfulness, fluency, style, and coherence. We propose a new test-time steering method designed to reduce such unnecessary interventions while improving the safety of unsafe responses. Our approach filters tokens using a value-based safety criterion and provides an explicit bound on the probability of false interventions. A single threshold hyperparameter controls this bound, allowing practitioners to trade off higher rates of unnecessary intervention for better output safety. Across multiple datasets and experiments, we show that our value-filtered decoding method outperforms existing baselines, achieving better trade-offs between safety, helpfulness, and similarity to the base model.

URL PDF HTML 收藏
2601.17090 2026-07-14 cs.LG cs.AI 版本更新

SFO: Learning PDE Operators via Spectral Filtering

SFO:通过谱滤波学习偏微分方程算子

Noam Koren, Rafael Moschopoulos, Kira Radinsky, Elad Hazan

机构 * Technion - Israel Institute of Technology(技术离子-以色列理工学院) Princeton University(普林斯顿大学)

AI总结 研究如何让神经算子有效捕捉偏微分方程解映射中的长程非局部相互作用,提出谱滤波算子SFO,利用通用谱基参数化积分核,通过学习快速衰减特征值的谱系数实现高效表示,在六个基准测试中达最优精度,大幅降误差并减少参数。

详情
AI中文摘要

偏微分方程(PDEs)控制着复杂系统,然而神经算子常常难以有效捕捉其解映射中固有的长程、非局部相互作用。我们引入了谱滤波算子(SFO),这是一种神经算子,它使用通用谱基(USB)对积分核进行参数化,USB是从谱滤波理论中的希尔伯特矩阵本征模式导出的固定全局正交基。基于理论发现,即平移不变PDE离散化的离散格林函数呈现空间线性动态系统(LDS)结构,我们证明这些核在USB中允许紧凑近似。通过仅学习快速衰减特征值的谱系数,SFO实现了高效表示。在包括反应扩散、流体动力学和3D电磁学在内的六个基准测试中,SFO达到了当前最优精度,相对于强大基线,误差降低了40%,同时使用的参数大幅减少。

英文摘要

Partial differential equations (PDEs) govern complex systems, yet neural operators often struggle to efficiently capture the long-range, nonlocal interactions inherent in their solution maps. We introduce Spectral Filtering Operator (SFO), a neural operator that parameterizes integral kernels using the Universal Spectral Basis (USB), a fixed, global orthonormal basis derived from the eigenmodes of the Hilbert matrix in spectral filtering theory. Motivated by our theoretical finding that the discrete Green's functions of shift-invariant PDE discretizations exhibit spatial Linear Dynamical System (LDS) structure, we prove that these kernels admit compact approximations in the USB. By learning only the spectral coefficients of rapidly decaying eigenvalues, SFO achieves a highly efficient representation. Across six benchmarks, including reaction-diffusion, fluid dynamics, and 3D electromagnetics, SFO achieves state-of-the-art accuracy, reducing error by up to 40% relative to strong baselines while using substantially fewer parameters.

URL PDF HTML 收藏
2607.07542 2026-07-09 cs.RO cs.CV 新提交

SonoRank: Towards Calibration-Free Real-Time Finger Flexion Detection from Forearm Ultrasound Sequences

SonoRank:迈向无需校准的实时前臂超声序列手指弯曲检测

Dean Zadok, Alon Wolf, Alex M. Bronstein, Oren Salzman

机构 * Technion(以色列理工学院)

AI总结 针对动力假肢手因依赖sEMG功能受限的问题,提出SonoRank方法,通过对超声序列对排序学习及微调,实现无需校准的手指弯曲检测,在交叉验证中F1分数比直接分类基线提高28%,推动超声假肢实用化。

详情
AI中文摘要

动力假肢手常被弃用,主要是因为依赖表面肌电图(sEMG)的现有设备功能有限。超声成像因其能实时观察肌肉活动并控制更多自由度,成为一种有前景的替代方案。然而,现有的基于超声的方法需要针对每个用户进行微调,限制了其商业化。我们提出了SonoRank,这是朝着无需校准的前臂超声视频手指弯曲检测迈出的重要一步。SonoRank首先通过五个手指的相对运动幅度对超声序列对进行排序学习。然后,利用操作开始时捕获的静止参考,对学习到的表示进行微调,以分类每个手指是否正在主动弯曲。在一个有十二个同步运动学受试者的数据集上进行12倍留一法交叉验证时,SonoRank在F1分数上比跳过排序阶段的直接分类基线提高了28%。这些结果表明成对排序是用于独立于受试者控制的有效预训练信号,使基于超声的假肢更接近实际的、无需校准的部署。

英文摘要

Powered prosthetic hands are frequently abandoned, largely due to the limited functionality of current devices that rely on surface electromyography (sEMG). Sonomyography (ultrasound) has emerged as a promising alternative, owing to its ability to observe muscle activity in real time and control a greater number of degrees of freedom. Yet, existing ultrasound-based methods require per-user fine-tuning, limiting their commercialization. We propose SonoRank, an important step towards calibration-free finger flexion detection from forearm ultrasound video. SonoRank first learns to rank pairs of ultrasound sequences by their relative motion magnitude for each of the five fingers. The learned representations are then fine-tuned to classify whether each finger is actively flexing, using a rest reference that is captured at the beginning of the operation. Under 12-fold leave-one-subject-out cross-validation on a dataset of twelve subjects with synchronized kinematics, SonoRank achieves a 28% improvement in F1 score over direct classification baselines that skip the ranking stage. These results establish pairwise ranking as an effective pretraining signal for subject-independent control, bringing ultrasound-based prosthetics closer to practical, calibration-free deployment.

URL PDF HTML 收藏
2606.15967 2026-07-09 cs.CV 新提交

CRIS: Cross-Plane Self-Supervised Isotropic Restoration for Anisotropic Volumetric Imaging Across Modalities

CRIS:跨模态各向异性体积成像的跨平面自监督各向同性恢复

Adi Ahituv, Anat Ilivitzki, Moti Freiman

机构 * Faculty of Data and Decision Sciences, Technion -- Israel Institute of Technology(数据与决策科学学院,技术离子技术学院) Faculty of Biomedical Engineering, Technion -- Israel Institute of Technology(生物医学工程学院,技术离子技术学院) The May-Blum-Dahl MRI Research Center, Technion -- Israel Institute of Technology(梅-布卢姆-达尔MRI研究中心,技术离子技术学院)

AI总结 提出CRIS,一种无需配对各向同性真值的跨平面自监督框架,通过正交重切2D条带补全实现3D各向同性恢复,在MRI和体积电镜上优于插值和多种方法。

Comments 24 pages, 8 figures, supplementary material included

详情
AI中文摘要

各向异性体积采集在临床MRI和体积电子显微镜(vEM)中很常见,其中稀疏的跨平面采样产生厚切片或截面,降低了正交重切和下游分析的质量。我们提出CRIS,一种跨平面自监督框架,无需配对各向同性真值即可实现各向同性恢复。CRIS将3D恢复视为各向同性网格正交重切上的2D条带补全:训练时,高分辨率面内切片被合成退化并周期性掩蔽;推理时,空白切片定义各向同性网格,恢复两个正交重切,并通过多视图平均融合预测。我们在两个MRI队列和两个显微镜基准上评估CRIS,各向异性高达8倍。在脑MRI上,CRIS达到32.921±0.436 dB PSNR和0.9631±0.0027 SSIM,优于插值、SMORE4、SIMPLE、SA-INR和ATME,并给出最佳分割一致性(Dice 0.940±0.004,ASSD 0.245±0.014 mm,HD99 1.275±0.061 mm)。在无参考腹部MRI上,CRIS将FID/KID降至48.714/0.023。在vEM上,CRIS优于插值、NIIV和vEMINR,在4倍时达到29.133 dB/0.834 3D PSNR/SSIM,在EPFL 8倍时达到27.123 dB/0.734,在噪声hemibrain数据上达到21.915 dB/0.699。在鲁棒性实验中,一个可变间隙CRIS模型在间隙因子3-7以及冠状、轴向和矢状退化下评估,保持比插值更高的PSNR/SSIM(36.36-31.14 dB和0.977-0.932对比33.07-27.85 dB和0.951-0.853)。这些结果支持CRIS作为一种模态灵活的途径,无需配对各向同性目标或特定配置的重新训练即可实现各向同性恢复。代码可在https://github.com/adi-hatav/CRIS获取。

英文摘要

Anisotropic volumetric acquisitions are common in clinical MRI and volume electron microscopy (vEM), where sparse through-plane sampling creates thick slices or sections that degrade orthogonal reformats and downstream analysis. We present CRIS, a cross-plane self-supervised framework for isotropic restoration without paired isotropic ground truth. CRIS casts 3D restoration as 2D stripe completion on orthogonal reformats of an isotropic grid: high-resolution in-plane slices are synthetically degraded and periodically masked for training, while at inference blank slices define the isotropic grid, two orthogonal reformats are restored, and predictions are fused by multi-view averaging. We evaluate CRIS on two MRI cohorts and two microscopy benchmarks up to 8x anisotropy. On brain MRI, CRIS achieves 32.921 +/- 0.436 dB PSNR and 0.963 +/- 0.003 SSIM, outperforming interpolation, ECLARE, SMORE4, SIMPLE, SA-INR, and ATME, and gives the best segmentation consistency (Dice 0.940 +/- 0.004, ASSD 0.245 +/- 0.014 mm, HD99 1.275 +/- 0.061 mm). On reference-free abdominal MRI, CRIS reduces FID/KID to 48.71/0.023, outperforming interpolation, ECLARE, SMORE4, and SIMPLE. On vEM, CRIS achieves 29.100 dB/0.830 3D PSNR/SSIM at 4x and 26.874 dB/0.722 at 8x on EPFL, and 21.935 +/- 0.437 dB/0.696 +/- 0.024 on noisy hemibrain data. In a dedicated robustness experiment, one variable-gap CRIS model evaluated across gap factors 3-7 and coronal, axial, and sagittal degradations maintained higher PSNR/SSIM than interpolation (36.36-31.14 dB and 0.977-0.932 vs. 33.07-27.85 dB and 0.951-0.853). These results support CRIS as a modality-flexible route to isotropic restoration without paired isotropic targets or configuration-specific retraining. Code is available at https://github.com/adi-hatav/CRIS.

URL PDF HTML 收藏
2602.18201 2026-07-09 cs.AI cs.LG 版本更新

SOMtime the World Ain$'$t Fair: Violating Fairness Using Self-Organizing Maps

有时世界并不公平:使用自组织映射违反公平性

Joseph Bingham, Netanel Arussy, Dvir Aran

机构 * Faculty of Biology, Technion University(生物学院,技术学院) Taub Faculty of Computer Science, Technion University(计算机科学学院,技术学院)

AI总结 研究发现基于自组织映射的SOMtime方法,能让敏感属性在无监督嵌入中成为主导潜在轴,恢复与敏感属性对齐的排序,其嵌入分割产生人口统计学倾斜聚类,揭示“通过无意识实现公平”在表示层面失败,公平审计需扩展到无监督组件。

Comments 12 pages, 2 figures, preprint

详情
AI中文摘要

人们普遍认为,当在训练中不考虑敏感属性时,无监督表示对这些属性是中立的。但本文表明这一假设是错误的。使用基于高容量自组织映射的拓扑保持表示方法SOMtime,研究发现年龄和收入等敏感属性会在纯无监督嵌入中成为主导潜在轴,即使明确从输入中排除。在两个大规模真实世界数据集上,SOMtime恢复了与被 withheld敏感属性对齐的单调排序,而PCA、UMAP、t-SNE和自动编码器表现较差。此外,SOMtime嵌入的无监督分割产生了人口统计学上倾斜的聚类。研究结果表明,“通过无意识实现公平”在序数敏感属性的表示层面上失败,公平性审计必须扩展到机器学习管道的无监督组件。

英文摘要

Unsupervised representations are widely assumed to be neutral with respect to sensitive attributes when those attributes are withheld from training. We show that this assumption is false. Using SOMtime, a topology-preserving representation method based on high-capacity Self-Organizing Maps, we demonstrate that sensitive attributes such as age and income emerge as dominant latent axes in purely unsupervised embeddings, even when explicitly excluded from the input. On two large-scale real-world datasets (the World Values Survey across five countries and the Census-Income dataset), SOMtime recovers monotonic orderings aligned with withheld sensitive attributes, achieving Spearman correlations of up to 0.85, whereas PCA and UMAP typically remain below 0.23 (with a single exception reaching 0.31), and against t-SNE and autoencoders which achieve at most 0.34. Furthermore, unsupervised segmentation of SOMtime embeddings produces demographically skewed clusters, demonstrating downstream fairness risks without any supervised task. These findings establish that \textit{fairness through unawareness} fails at the representation level for ordinal sensitive attributes and that fairness auditing must extend to unsupervised components of machine learning pipelines. We have made the code available at~ https://github.com/JosephBingham/SOMtime

URL PDF HTML 收藏
2604.21637 2026-07-07 cs.CL cs.CY 版本更新

Multilinguality at the Edge: Developing Language Models for the Global South

边缘侧多语言性:为全球南方开发语言模型

Lester James V. Miranda, Songbo Hu, Roi Reichart, Anna Korhonen

机构 * Language Technology Lab, University of Cambridge(剑桥大学语言技术实验室) Technion – Israel Institute of Technology(技术学院——以色列理工学院)

AI总结 针对全球南方非英语、硬件受限社区的语言模型落地难题,调研232篇相关文献,梳理多语言与边缘部署交叉领域进展,提出面向NLP生态各方的可行建议。

Comments Updated formatting and improved spacing. Project website is in https://ljvmiranda921.github.io/multilinguality-at-the-edge/

详情
AI中文摘要

语言模型(LMs)的部署位置与方式决定了哪些群体能够从中受益。然而,多项挑战阻碍了语言模型在全球南方的非英语且硬件受限社区的有效部署。我们将这一挑战称为最后一英里:即多语言性与边缘部署的交叉领域,二者的目标一致,但技术需求往往相互冲突。将这两个领域结合研究既是必要的——语言多样性丰富的社区往往面临最严峻的基础设施约束——也是机遇——当前边缘NLP与多语言NLP研究仍处于相互割裂的状态。为了解决这两个领域结合的最新进展与现存挑战,我们调研了覆盖语言建模全流程(从数据收集、开发到部署)的232篇相关论文。我们还讨论了待解决的开放问题,并为NLP生态中的不同利益相关方提供了可落地的建议。最终,我们希望本工作能够推动包容性、公平性语言技术的发展。

英文摘要

Where and how language models (LMs) are deployed determines who can benefit from them. However, there are several challenges that prevent effective deployment of LMs in non-English-speaking and hardware constrained communities in the Global South. We call this challenge the last mile: the intersection of multilinguality and edge deployment, where the goals are aligned but the technical requirements often compete. Studying these two fields together is both a need, as linguistically diverse communities often face the most severe infrastructure constraints, and an opportunity, as edge and multilingual NLP research remain largely siloed. To understand the state of the art and the challenges of combining the two areas, we survey 232 papers that tackle this problem across the language modelling pipeline, from data collection to development and deployment. We also discuss open questions and provide actionable recommendations for different stakeholders in the NLP ecosystem. Finally, we hope that this work contributes to the development of inclusive and equitable language technologies.

URL PDF HTML 收藏
2607.05359 2026-07-07 cs.AI 新提交

Graph Sparse Sampling: Breaking the Curse of the Horizon in Continuous MDP Planning

图稀疏采样:打破连续MDP规划中的视界诅咒

Idan Lev-Yehudi, Vadim Indelman

机构 * Technion Autonomous Systems Program (TASP), Technion – Israel Institute of Technology(以色列理工学院自主系统计划(TASP),以色列理工学院) Stephen B. Klein Faculty of Aerospace Engineering, Technion – Israel Institute of Technology(以色列理工学院斯蒂芬·B·克莱因航空航天工程学院) Faculty of Data and Decision Sciences, Technion – Israel Institute of Technology(以色列理工学院数据与决策科学学院)

AI总结 针对连续域不确定规划的视界指数增长难题,提出图稀疏采样算法跨决策共享采样未来,得到多项式视界依赖的性能保证,长视界下性能远超树型规划器。

详情
AI中文摘要

连续域下的不确定规划对自主系统至关重要,但计算量极大。蒙特卡洛树搜索等基于树的搜索方法应用广泛,但其分支结构在最坏情况下的采样预算会随前瞻深度指数增长。从树的视角看,连续状态或动作空间的规划难度极高,因为规划器必须在无限分支层级中确定搜索位置。本文提出图稀疏采样(GSS)这一在线规划算法,跨多个候选决策共享采样得到的未来轨迹,而非为每个候选动作采样独立的后继节点。该无分支图可生成适配GPU的大规模批处理,同时借助启发式方法聚焦计算资源。本文通过平滑回溯,证明了GSS针对满秩或低秩生成模拟器、离散或采样连续动作空间的有限样本性能保证。在满足重叠性、正则性和动作覆盖条件下,该界对规划视界呈多项式依赖关系,明确了共享未来轨迹可避免树型稀疏采样的指数级视界依赖的适用场景。连续控制仿真实验表明,GSS在长视界场景下性能大幅超越树型规划器,或可达到近优性能,证实无分支图规划可作为在线控制的补充设计原则。

英文摘要

Planning under uncertainty in continuous domains is essential for autonomous systems, yet computationally demanding. Tree-based search methods such as Monte Carlo Tree Search (MCTS) remain popular, but their branching structure can require sampling budgets that grow exponentially with lookahead depth in the worst case. From a tree perspective, continuous state or action spaces become especially challenging, since the planner must decide where to search in an infinite branching hierarchy. We propose Graph Sparse Sampling (GSS), an online planning algorithm that shares sampled futures across many candidate decisions, rather than sampling separate successors for each candidate action. This branch-free graph exposes large GPU-friendly batches, while using heuristics to focus computation. We prove finite-sample performance guarantees for GSS covering full-rank or low-rank generative simulators via smoothed backups, and discrete or sampled continuous action spaces. Under suitable overlap, regularity, and action-coverage conditions, these bounds have polynomial dependence on the planning horizon, formalizing when shared futures can avoid the exponential horizon dependence of tree-shaped sparse sampling. We demonstrate continuous-control simulations where GSS substantially outperforms tree-based planners on long horizons or achieves near-optimal performance, supporting no-branching graph planning as a complementary design principle for online control.

URL PDF HTML 收藏
2607.03875 2026-07-07 cs.CV 新提交

MACRO: Training-free Multi-plane Attention for Closeup Render Optimization

MACRO:用于特写渲染优化的免训练多平面注意力

Nitzan Hodos, Roy Amoyal, Lior Fritz, Ianir Ideses, Sagie Benaim, Netalee Efrat

机构 * Amazon Prime Video(亚马逊Prime视频) Technion, Israel Institute of Technology(以色列理工学院) Ben Gurion University(本古里安大学) Hebrew University of Jerusalem(耶路撒冷希伯来大学)

AI总结 针对特写渲染难题,分析现有方法失败原因,提出MACRO,利用场景3D结构解决比例差距,无需架构改变和额外训练,贡献新基准和评估协议,取得领先成果。

Comments Project page: https://nitzanhod.github.io/MACRO

详情
AI中文摘要

特写渲染对虚拟制作和交互式3D内容很重要但具挑战性。3D高斯溅射渲染质量近距会下降,基于扩散的方法有伪影。分析发现是特写与参考视图的比例差距问题。在此基础上介绍MACRO,一种免训练的高质量特写新视图合成方法,还贡献了新基准和评估协议,性能领先。

英文摘要

Close-up rendering, zooming into a scene well beyond any training camera, is important for virtual production and interactive 3D content, yet remains an open challenge. 3D Gaussian splatting (3DGS) enables high-fidelity, real-time novel view synthesis, but its rendering quality degrades at close range. Recent diffusion-based methods that enhance the rendering by conditioning on reference images from the training set produce significant artifacts in this setting. We analyze this failure and identify its root cause: the scale gap between the close-up and reference views. We show that the features in reference-conditioned enhancement models are not scale-invariant, causing cross-view attention to retrieve incorrect correspondences when the same content appears at different scales, and that this mismatch cannot be corrected in latent space because the VAE encoder is not scale-equivariant. Building on this analysis we introduce MACRO, Multi-plane Attention for Closeup Render Optimization, a training-free method for high-quality close-up novel view synthesis from 3DGS. MACRO resolves the scale gap by leveraging the scene's known 3D structure: it decomposes the close-up into depth planes, crops and resizes references in image space to match the scale of each plane before encoding, and applies a depth-aware attention mask so each token attends only to scale-matched references. The method requires no architectural changes or additional training. We further contribute two new close-up novel view synthesis benchmarks, the first standardized evaluation protocol for this setting, and demonstrate state-of-the-art results on both, outperforming existing 3DGS and diffusion-based methods on both reconstruction and perceptual metrics. Project page: https://nitzanhod.github.io/MACRO

URL PDF HTML 收藏
2607.03870 2026-07-07 cs.AI cs.CL cs.LG 新提交

Evaluating LLM Uncertainty in Long-Form Generation Using Deterministic Ground Truth

使用确定性真实数据评估长文本生成中语言模型的不确定性

Ido Amit, Ido Galil, Ran El-Yaniv

机构 * Technion(技术学院) Nvidia(英伟达)

AI总结 研究针对长文本生成中语言模型不确定性评估难题,引入含单一确定性长文本真值的SALT基准,通过对50多个模型分析揭示关键见解,如置信函数作用、错误驱动因素及推理影响等,为风险关键应用提供参考。

Comments Accepted to the 43rd International Conference on Machine Learning (ICML 2026). Code available at https://github.com/IdoAmit198/SALT

详情
AI中文摘要

随着语言模型生成的输出越来越长,有效的不确定性估计必须在细粒度级别识别错误,而不是丢弃整个响应。虽然存在这样的方法,但在任何分辨率(从令牌到整个生成)下评估不确定性都具有挑战性,并且对标签缺陷高度敏感,这使得零噪声基准至关重要;然而,长文本生成基准往往依赖于不可靠的标签,而不是确定性的真实数据。我们引入了单答案原子长文本目标(SALT),这是一个由六个程序生成的任务组成的基准,具有单一确定性长文本真值,无需外部评判即可进行正确性、校准和排名的单元级评估。借助SALT,我们对50多个语言模型的分析揭示了关键见解:我们确定了哪些置信函数主导每个不确定性方面,并表明即使在较粗的行级单元出现更清晰的可分离性时,置信排名在原子分辨率下也基本失效。SALT还能够在整个生成过程中进行可控的原子级干预,揭示了未来错误的两个可分离驱动因素:来自损坏前缀的传播,主要由全局上下文正确性主导,以及随着答案上下文长度增加而产生的有界退化。最后,我们证明,通过思维链提示或在训练中内化进行推理会带来权衡,提高准确性的同时降低置信排名。这些发现直接影响需要可靠错误识别和缓解的风险关键应用。

英文摘要

As LLMs generate increasingly long outputs, effective uncertainty estimation must identify errors at fine-grained levels rather than discard entire responses. While such methods exist, evaluating uncertainty at any resolution (token to an entire generation) is challenging and highly sensitive to label imperfections, making zero-noise benchmarks essential; yet, long-form generation benchmarks tend to rely on fallible labels rather than deterministic ground truth. We introduce Single-answer Atomic Long-form Target (SALT), a benchmark of six procedurally generated tasks with single deterministic long textual ground truths, enabling unit-level evaluation of correctness, calibration, and ranking without external judges. Equipped with SALT, our analysis of 50+ LLMs reveals key insights: We identify which confidence functions dominate each uncertainty aspect and show that confidence ranking largely breaks at atomic resolution, even when clearer separability emerges at coarser line-level units. SALT further enables controlled atom-level interventions throughout generation, revealing two separable drivers of future errors: propagation from corrupted prefixes, dominated by global context correctness, and bounded degradation from increasing answer-context length. Finally, we demonstrate that reasoning, via Chain-of-Thought prompting or internalized through training, introduces a trade-off, improving accuracy while degrading confidence ranking. These findings directly impact risk-critical applications requiring reliable error identification and mitigation.

URL PDF HTML 收藏
2607.03517 2026-07-07 cs.LG cs.CV 新提交

Mixture-of-Gaussians-Guided Schedule Design for Brownian Bridge Diffusion Models

用于布朗桥扩散模型的高斯混合引导调度设计

Ron Levi, Michael Elad

机构 * Technion, Israel(以色列理工学院)

AI总结 研究为布朗桥扩散模型开发调度设计框架,基于高斯混合先验分析其反向动力学,得出理想后验和去噪器,据此制定两个调度设计目标,揭示权衡并证明通用调度存在,实验验证其价值。

Comments 66 pages, 10 figures

详情
AI中文摘要

布朗桥扩散模型(BBDM)为图像恢复和反问题提供了一个有吸引力的框架,通过构建从干净信号直接到退化观测的随机桥,而不是到纯噪声。尽管有前景,但桥调度的选择通常是从启发式继承而来的,并且缺乏一个有原则的调度设计分析框架。在这项工作中,我们通过在高斯混合(MoG)先验下对BBDM反向动力学进行新颖的分析来开发这样一个框架。这种设置产生了一个封闭形式的理想后验和一个相应的最小均方误差去噪器,而BBDM诱导的重建定律通过一个易于处理的替代物进行分析捕获。基于这些表达式,我们制定了两个互补的调度设计目标:一个针对感知质量的瓦瑟斯坦准则和一个针对重建保真度的均方误差准则。我们的工作揭示了两者之间的内在权衡,并证明了存在与退化和先验无关的通用调度。在受控的MoG设置上进行的广泛实验证实了理论与实践之间的完全一致,并且在FFHQ数据集上针对修复、去模糊和超分辨率任务的实验验证了我们的调度设计标准的实用价值。

英文摘要

Brownian Bridge Diffusion Models (BBDM) offer an appealing framework for image restoration and inverse problems by constructing a stochastic bridge from the clean signal directly to the degraded observation, rather than to pure noise. Despite their promise, the choice of bridge schedule is typically inherited from heuristics, and a principled analytical framework for schedule design has been lacking. In this work, we develop such a framework by offering a novel analysis of BBDM reverse dynamics under a Mixture-of-Gaussians (MoG) prior. This setting yields a closed-form ideal posterior and a corresponding MMSE denoiser, while the BBDM-induced reconstruction law is captured analytically through a tractable surrogate. Building on these expressions, we formulate two complementary schedule-design objectives: a Wasserstein criterion targeting perceptual quality and an MSE criterion targeting reconstruction fidelity. Our work exposes an inherent tradeoff between the two and proves the existence of universal schedules for both that are independent of the degradation and prior. Extensive experiments on controlled MoG settings confirm full alignment between theory and practice, and experiments on the FFHQ dataset across inpainting, deblurring, and super-resolution tasks validate the practical value of our schedule-design criteria.

URL PDF HTML 收藏
2607.03815 2026-07-07 math.MG cs.LG 新提交

A simplex-based measure of symmetry

基于单纯形的对称度量

Egor Bakaev, Amir Yehudayoff

机构 * Department of Computer Science, University of Copenhagen(计算机科学系,哥本哈根大学) Department of Mathematics, Technion-IIT(数学系,技术学院-理工学院)

AI总结 研究基于\(n\)-单纯形\(\Delta\)的对称度量\(\rho_\Delta(L)\),推导其与经典闵可夫斯基对称度量关系,改进稳定性分析,给出单纯形新特征,还研究了多面体深度复杂度与该度量的关系及界。

详情
AI中文摘要

对于紧致凸集\(L,K\subset\mathbb{R}^n\),用\(\lambda_K(L)\)表示包含\(L\)的\(K\)的同胚的最小尺寸。定义基于\(n\)-单纯形\(\Delta=\Delta^n\subset\mathbb{R}^n\)的对称度量\(\rho_\Delta(L)=\frac{\lambda_{-\Delta}(L)}{\lambda_{\Delta}(L)}\),并推导了一系列结果。

英文摘要

For compact convex sets $L,K \subset \mathbb{R}^n$, denote by $λ_K(L)$ the smallest size of a homothet of $K$ that contains $L$. We define a measure of symmetry based on the $n$-simplex $Δ= Δ^n \subset \mathbb{R}^n$ as the ratio \[ ρ_Δ(L):=\frac{λ_{-Δ}(L)}{λ_Δ(L)}. \] We study this measure and deduce the following results: (1) The classical Minkowski measure of symmetry $m^*(L)$ can be defined as an affine-invariant version of $ρ_Δ(L)$. (2) We improve the stability analysis for the Minkowski measure of symmetry; if $m^*(L)\ge n-\varepsilon$ then $L$ is $\tfrac{1}{1-\varepsilon}$-close to $Δ$ in the Banach--Mazur distance. (3) We obtain a novel characterization of simplices as the only convex bodies $K$ for which the function $L \mapsto λ_K(L)$ is additive (a property we term ``outer additivity''). (4) Motivated by the expressivity of ReLU neural networks, we study the depth complexity of polytopes in $\mathbb{R}^n$ under the two operations: Minkowski sum and convex hull of a union. We prove the sharp bound $ρ_Δ(P) \leq 2^d -1$ for every polytope $P$ of depth complexity $d$. In other words, simplices cannot be approximated by low-depth polytopes.

URL PDF HTML 收藏
2602.06358 2026-07-07 cs.CL cs.AI 版本更新

SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass

SHINE:一种可扩展的上下文超网络,用于在单次传递中将上下文映射到LoRA

Yewei Liu, Xiyuan Wang, Yansheng Mao, Yoav Gelbery, Haggai Maron, Muhan Zhang

机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) Technion, NVIDIA(技术学院与NVIDIA) University of Oxford(牛津大学) School of Electronics Engineering(电子工程学院) Computer Science, Peking University(计算机科学,北京大学)

AI总结 本文提出SHINE,一种可扩展的上下文超网络,用于在单次传递中将多样且有意义的上下文映射到高质量的LoRA适配器,通过重用冻结LLM的自身参数和引入架构创新,克服了先前超网络的关键限制,以较少的参数实现了强大的表达能力。

详情
AI中文摘要

我们提出SHINE(可扩展的上下文超网络),一种可扩展的超网络,能够将多样且有意义的上下文映射到高质量的LoRA适配器,用于大型语言模型(LLMs)。通过在上下文超网络设计中重用冻结LLM的自身参数,并引入架构创新,SHINE克服了先前超网络的关键限制,以相对较少的参数实现了强大的表达能力。我们引入了预训练和指令微调流水线,并训练我们的超网络在单次前向传递中从多样且有意义的上下文中生成高质量的LoRA适配器。它在不进行微调的情况下更新LLM参数,并立即启用与上下文相关的复杂问答任务,而无需直接访问上下文,有效地将上下文知识转换为参数知识。我们的工作在各种任务上取得了出色的结果,相比基于SFT的LLM适应方法,大大节省了时间和计算和内存成本,并展示了良好的可扩展性潜力。我们的代码可在https://github.com/MuLabPKU/SHINE获取。

英文摘要

We propose SHINE (Scalable Hyper In-context NEtwork), a scalable hypernetwork that can map diverse meaningful contexts into high-quality LoRA adapters for large language models (LLMs). By reusing the frozen LLM's own parameters in an in-context hypernetwork design and introducing architectural innovations, SHINE overcomes key limitations of prior hypernetworks and achieves strong expressive power with a relatively small number of parameters. We introduce a pretraining and instruction fine-tuning pipeline, and train our hypernetwork to generate high quality LoRA adapters from diverse meaningful contexts in a single forward pass. It updates LLM parameters without any fine-tuning, and immediately enables complex question answering tasks related to the context without directly accessing the context, effectively transforming in-context knowledge to in-parameter knowledge in one pass. Our work achieves outstanding results on various tasks, greatly saves time, computation and memory costs compared to SFT-based LLM adaptation, and shows great potential for scaling. Our code is available at https://github.com/MuLabPKU/SHINE

URL PDF HTML 收藏
2604.09544 2026-07-07 cs.CL cs.AI cs.LG 版本更新

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types

大语言模型通过一种独特的统一机制生成有害内容

Hadas Orgad, Boyi Wei, Kaden Zheng, Martin Wattenberg, Peter Henderson, Seraphina Goldfarb-Tarrant, Yonatan Belinkov

机构 * Kempner Institute, Harvard University(哈佛大学肯普纳研究所) Princeton University(普林斯顿大学) Harvard University(哈佛大学) Cohere Technion—IIT(以色列理工学院)

AI总结 研究通过权重剪枝揭示大语言模型中有害生成的内部结构,发现有害内容生成依赖于一组通用且与良性能力不同的权重,表明对齐训练重塑了有害表示,解释了领域微调引发的广泛对齐偏差。

详情
AI中文摘要

大语言模型(LLMs)通过对齐训练来避免有害行为,但由此产生的安全措施仍然脆弱:越狱经常绕过它们,且在狭窄领域微调可诱导"涌现对齐偏差",这种偏差广泛泛化。是否这种脆弱性反映了有害性内部缺乏一致的组织结构仍不清楚。本文通过针对性权重剪枝作为因果干预,探测LLM中有害性的内部组织。研究发现,有害内容生成依赖于一组跨害类型通用且与良性能力不同的权重。对齐模型相比未对齐模型压缩了更多有害生成权重,表明对齐训练重塑了内部有害表示--尽管安全防护措施在表面层面仍脆弱。这种压缩解释了涌现对齐偏差:如果有害能力的权重被压缩,微调这些权重在一个领域可触发广泛对齐偏差。一致地,剪枝狭窄领域中的有害生成权重显著减少了涌现对齐偏差。值得注意的是,LLM的有害生成能力与其识别和解释此类内容的方式是分离的。这些结果揭示了LLM中有害性的一致内部结构,可能为更系统的方法提供基础。

英文摘要

Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks routinely bypass them, and fine-tuning on narrow domains can induce ``emergent misalignment'' that generalizes broadly. Whether this brittleness reflects a fundamental lack of coherent internal organization for harmfulness remains unclear. Here we use targeted weight pruning as a causal intervention to probe the internal organization of harmfulness in LLMs. We find that harmful content generation depends on a compact set of weights that are general across harm types and distinct from benign capabilities. Aligned models exhibit a greater compression of harm generation weights than unaligned counterparts, indicating that alignment reshapes harmful representations internally--despite the brittleness of safety guardrails at the surface level. This compression explains emergent misalignment: if weights of harmful capabilities are compressed, fine-tuning that engages these weights in one domain can trigger broad misalignment. Consistent with this, pruning harm generation weights in a narrow domain substantially reduces emergent misalignment. Notably, LLMs harmful generation capability is dissociated from how they recognize and explain such content. Together, these results reveal a coherent internal structure for harmfulness in LLMs that may serve as a foundation for more principled approaches to safety.

URL PDF HTML 收藏
2607.01823 2026-07-03 eess.AS cs.CL 新提交

Self-Supervised Test-Time Tuning for Packet Loss Concealment

自监督测试时调优用于丢包隐藏

Yehoshua Dissen, Joseph Keshet

机构 * Andrew and Erna Viterbi Faculty of Electrical and Computer Engineering, Technion–Israel Institute of Technology(安德鲁和伊尔纳·维特比电气与计算机工程学院,技术ion-以色列理工学院)

AI总结 提出TTT-PLC框架,在测试时利用接收到的音频包自监督地调整丢包隐藏模型,无需干净参考信号或外部数据,在非因果和因果场景下均提升性能。

Comments Under submission to IEEE TASLP

详情
AI中文摘要

丢包隐藏(PLC)重建接收端缺失的音频包,通常使用训练好的模型,其参数在部署时固定不变。这导致PLC模型被视为静态的,尽管每个通话或录音通过已到达的包暴露了信号特定信息。我们提出TTT-PLC,一个自监督的测试时调优框架,仅使用接收到的包来适应现有PLC模型。该方法通过合成地掩盖可用信号的部分来创建监督,训练模型用其原生PLC目标来隐藏这些部分,然后使用适应后的模型重建真实的丢包。不需要干净的参考信号、外部适应数据或架构修改。我们在两种部署设置中研究TTT-PLC。在非因果设置中,接收到的文件在重建前可用,允许重复的自监督适应过程,并提供每个文件的适应上限。在因果设置中,音频流式传输而不修改已发射的样本;适应仅在已完成的过去块上进行,更新后的参数仅影响未来音频。我们在两个公开的PLC骨干网络上实例化该框架:FRN(一个循环全带语音PLC模型)和PARCnet(一个用于网络音乐的混合自回归-神经模型)。在这些设置中,结果表明预训练的PLC系统在推理时不必被视为固定不变的,丢失信号中仍被观察到的部分可以为改善同一信号的隐藏提供有效的训练信号。

英文摘要

Packet loss concealment (PLC) reconstructs audio packets that are missing at the receiver, usually with a trained model whose parameters remain fixed at deployment time. This treats the PLC model as static, even though each call or recording exposes signal-specific information through the packets that did arrive. We present TTT-PLC, a self-supervised test-time tuning framework that adapts existing PLC models using only those received packets. The method creates supervision by synthetically masking portions of the available signal, training the model to conceal them with its native PLC objective, and then using the adapted model to reconstruct the true packet losses. No clean reference signal, external adaptation data, or architectural modification is required. We study TTT-PLC in two deployment settings. In the non-causal setting, the received file is available before reconstruction, allowing repeated self-supervised adaptation passes and providing a per-file adaptation ceiling. In the causal setting, audio is streamed without revising emitted samples; adaptation is performed only on completed past blocks, and updated parameters affect only future audio. We instantiate the framework on two public PLC backbones, FRN, a recurrent full-band speech PLC model, and PARCnet, a hybrid autoregressive-neural model for networked music. Across these settings, the results show that pretrained PLC systems do not need to be treated as fixed at inference time, the still-observed portions of a lossy signal can provide an effective training signal for improving concealment on that same signal.

URL PDF HTML 收藏
2603.17212 2026-07-03 cs.GT cs.AI cs.LG 版本更新

Adaptive Contracts for Cost-Effective AI Delegation

面向成本效益的AI委托自适应合同

Eden Saig, Tamar Garbuz, Ariel D. Procaccia, Inbal Talgam-Cohen, Jamie Tucker-Foltz

机构 * Tel Aviv University(特拉维夫大学) Harvard University(哈佛大学) Technion -- Israel Institute of Technology(技术学院——以色列理工学院) California Institute of Technology(加州理工学院) Yale School of Management(耶鲁管理学院)

AI总结 针对AI委托中评估噪声导致支付增加的问题,提出自适应合同机制,通过选择性详细评估降低成本,并给出最优合同计算算法与实证验证。

Comments ICML 2026

详情
AI中文摘要

当组织通过按绩效付费合同将文本生成任务委托给AI提供商时,若评估存在噪声,预期支付会上升。随着评估方法变得复杂,降低噪声的经济效益往往被增加的评估成本所掩盖。在这项工作中,我们引入了用于AI委托的自适应合同,该合同允许在观察初始粗略信号后选择性地执行详细评估以节约资源。我们做出了三组贡献:首先,我们在自然假设或核心问题维度较小的情况下,提供了计算最优自适应合同的高效算法,并证明了在一般无结构情况下近似计算的难度。然后,我们提出了随机自适应合同的替代模型,并讨论了其优缺点。最后,我们使用问答和代码生成数据集实证证明了自适应相对于非自适应基线的优势。

英文摘要

When organizations delegate text generation tasks to AI providers via pay-for-performance contracts, expected payments rise when evaluation is noisy. As evaluation methods become more elaborate, the economic benefits of decreased noise are often overshadowed by increased evaluation costs. In this work, we introduce adaptive contracts for AI delegation, which allow detailed evaluation to be performed selectively after observing an initial coarse signal in order to conserve resources. We make three sets of contributions: First, we provide efficient algorithms for computing optimal adaptive contracts under natural assumptions or when core problem dimensions are small, and prove hardness of approximation in the general unstructured case. We then formulate alternative models of randomized adaptive contracts and discuss their benefits and limitations. Finally, we empirically demonstrate the benefits of adaptivity over non-adaptive baselines using question-answering and code-generation datasets.

URL PDF HTML 收藏
2606.30342 2026-07-01 cs.CV 版本更新

A Classifier-Agnostic Zero-Shot Adversarial Attack Detection via CLIP

通过CLIP实现分类器无关的零样本对抗攻击检测

Hodaya Krakover, Meir Yossef Levi, Eyal Gofer, Guy Gilboa

机构 * Technion - Israel Institute of Technology(以色列理工学院)

AI总结 提出A4D框架,利用CLIP的提示相似度分数,在完全黑盒零样本设置下检测对抗攻击,无需攻击或分类器先验知识。

Comments Accepted to ECCV 2026

详情
AI中文摘要

对抗攻击对深度学习模型的可靠性构成挑战,促使了有效的检测方法。现有技术通常依赖于攻击特定假设、对抗样本的访问或对底层分类器的了解(白盒)。我们提出了$A^4D$(攻击和架构无关的对抗检测器),一个完全黑盒、零样本的对抗攻击检测框架,利用来自CLIP的基于提示的相似度分数。据我们所知,这是首次尝试将CLIP用于此类任务。该方法基于两个关键观察:(i)CLIP对即使微小的不可感知的非语义扰动也很敏感;(ii)CLIP嵌入空间中的偏移不是任意的,可以用作鲁棒的攻击指标。跨多个攻击、数据集和分类器的实验验证了$A^4D$在攻击无关和分类器无关的设置中达到了SOTA检测结果。

英文摘要

Adversarial attacks pose a challenge to the reliability of deep learning models, motivating effective detection methods. Existing techniques often rely on attack-specific assumptions, access to adversarial samples, or knowledge of the underlying classifier (white-box). We propose $A^4D$ Attack- and Architecture-Agnostic Adversarial Detector, a completely black-box, zero-shot adversarial attack detection framework that utilizes prompt-based similarity scores derived from CLIP. To the best of our knowledge this is the first attempt to utilize CLIP for such a task. The method is based on two key observations: (i) CLIP is sensitive even to small imperceptible non-semantic perturbations; (ii) The shift in CLIP embedding space is not arbitrary and can be used as a robust attack indicator. Experiments across multiple attacks, datasets and classifiers validate that $A^4D$ achieves SOTA detection results in the attack-agnostic and classifier-agnostic setting.

URL PDF HTML 收藏
2606.31258 2026-07-01 cs.CV 新提交

WarpHammer: Densifying Scene Warps with 3D Object Priors for Extreme View Synthesis

WarpHammer: 利用3D物体先验稠密化场景扭曲以实现极端视角合成

Michael Green, Gavriel Habib, Dvir Samuel, Tal Berkovitz Shalev, Issar Tzachor, Rami Ben-Ari, Or Litany

机构 * OriginAI, Israel(OriginAI,以色列) NVIDIA(英伟达) Technion(以色列理工学院)

AI总结 针对投影条件新视角合成在大轨道运动下扭曲稀疏导致生成失败的问题,提出无需训练的WarpHammer框架,通过引入3D生成先验的物体显式重建来补充缺失前景并遮挡不应可见的背景点,恢复外观和相机线索,并首次支持融合外部辅助物体视图。

详情
AI中文摘要

投影条件新视角合成(NVS)将输入视图的显式3D重建扭曲到目标相机,并对生成器进行条件化处理。这种方法在小视角变化下效果良好,但在大轨道运动下急剧退化:围绕轨道物体的扭曲变得稀疏,隐藏表面主导新视角,出现镜面伪影,导致生成器丢失像素内容和扭曲携带的隐式相机线索。我们提出WarpHammer,一个无需训练的框架,通过从原生3D生成先验(如SAM3D)获得的物体显式3D重建来增强扭曲场景,从而解决这种失败模式。重建的物体添加了缺失的前景表面,并遮挡了不应再可见的背景点,无需微调基础模型即可恢复外观和相机线索。相同的显式物体表示进一步解锁了当前NVS流水线不支持的能力:融合来自目标场景外部的物体辅助视图,例如,一辆车的随意快照与同一车型的制造商工作室照片。我们联合处理参考图像和辅助图像,使用预训练的多视图几何基础模型,预测统一点云,并将其融合到3D物体重建中。这比单图像重建产生更准确的几何形状,且无需用户提供辅助视图的相机位姿。在五个基准测试中,WarpHammer在强基线崩溃的视角偏差下生成稳定的新视图,并且是首个能够自然融合外部来源、位姿未知的辅助物体视图的场景级NVS方法。

英文摘要

Projection-conditioned novel view synthesis (NVS) warps an explicit 3D reconstruction of the input view into the target camera and conditions a generator on the warped rendering. This works well for small viewpoint changes but degrades sharply under large orbital motion: the warp becomes sparse around the orbited object, where hidden surfaces dominate the new view and mirror-like artifacts emerge, causing the generator to lose both pixel content and the implicit camera cue carried by the warp. We introduce WarpHammer, a training-free framework that resolves this failure mode by augmenting the warped scene with an explicit 3D reconstruction of the object obtained from a native 3D generative prior (e.g., SAM3D). The reconstructed object adds missing foreground surfaces and occludes background points that should no longer be visible, restoring both appearance and camera cues without fine-tuning the base model. The same explicit object representation further unlocks a capability current NVS pipelines do not support: incorporating auxiliary views of the object from sources outside the target scene, for example, a casual snapshot of a car paired with a manufacturer studio shot of the same model. We process the reference and auxiliary images jointly with a pretrained multi-view geometry foundation model, which predicts a unified point cloud that we fuse into the 3D object reconstruction. This yields substantially more faithful geometry than single-image reconstruction, without requiring user-provided camera poses for the auxiliary views. On five benchmarks, WarpHammer produces stable novel views at viewpoint deviations where strong baselines collapse, and is the first scene-level NVS method that can naturally fuse auxiliary, pose-unknown object views from an external source.

URL PDF HTML 收藏
2603.24036 2026-07-01 cs.CV 版本更新

SpectralSplats: Robust Differentiable Tracking via Spectral Moment Supervision

SpectralSplats: 通过谱矩监督实现鲁棒的可微跟踪

Avigail Cohen Rimon, Amir Mann, Mirela Ben Chen, Or Litany

机构 * Technion - Israel Institute of Technology(以色列理工学院) NVIDIA(英伟达)

AI总结 提出SpectralSplats框架,通过将优化目标从空间域转移到频率域,利用全局复正弦特征(谱矩)构建全局吸引域,解决3D高斯泼溅跟踪中因严重错位导致的梯度消失问题,并设计频率退火策略实现从全局凸性到精确空间对齐的平滑过渡。

Comments Accepted to ECCV 2026. Project page: https://avigailco.github.io/SpectralSplats/

详情
AI中文摘要

3D高斯泼溅(3DGS)实现了实时、照片级真实感的新视角合成,使其成为基于模型的视频跟踪中极具吸引力的表示。然而,在野外利用3DGS渲染器的可微性仍然非常脆弱。一个根本瓶颈在于高斯原语的紧凑局部支持。标准光度目标隐式依赖于空间重叠;如果严重的相机错位使得渲染对象位于目标局部足迹之外,梯度严格消失,导致优化器陷入困境。我们引入了SpectralSplats,一个鲁棒的跟踪框架,通过将优化目标从空间域转移到频率域来解决这个“梯度消失”问题。通过一组全局复正弦特征(谱矩)监督渲染图像,我们构建了一个全局吸引域,确保在整个图像域中存在指向目标的有效方向梯度,即使像素重叠完全不存在。为了利用这个全局吸引域而不引入与高频相关的周期性局部最小值,我们从第一性原理推导出一个原则性的频率退火调度,使优化器从全局凸性优雅地过渡到精确的空间对齐。我们证明了SpectralSplats可以无缝地作为空间损失的即插即用替代品,适用于各种变形参数化(从MLP到稀疏控制点),即使在标准外观跟踪灾难性失败的严重错位初始化下,也能成功恢复复杂变形。

英文摘要

3D Gaussian Splatting (3DGS) enables real-time, photorealistic novel view synthesis, making it a highly attractive representation for model-based video tracking. However, leveraging the differentiability of the 3DGS renderer "in the wild" remains notoriously fragile. A fundamental bottleneck lies in the compact, local support of the Gaussian primitives. Standard photometric objectives implicitly rely on spatial overlap; if severe camera misalignment places the rendered object outside the target's local footprint, gradients strictly vanish, leaving the optimizer stranded. We introduce SpectralSplats, a robust tracking framework that resolves this "vanishing gradient" problem by shifting the optimization objective from the spatial to the frequency domain. By supervising the rendered image via a set of global complex sinusoidal features (Spectral Moments), we construct a global basin of attraction, ensuring that a valid, directional gradient toward the target exists across the entire image domain, even when pixel overlap is completely nonexistent. To harness this global basin without introducing periodic local minima associated with high frequencies, we derive a principled Frequency Annealing schedule from first principles, gracefully transitioning the optimizer from global convexity to precise spatial alignment. We demonstrate that SpectralSplats acts as a seamless, drop-in replacement for spatial losses across diverse deformation parameterizations (from MLPs to sparse control points), successfully recovering complex deformations even from severely misaligned initializations where standard appearance-based tracking catastrophically fails.

URL PDF HTML 收藏
2606.30559 2026-06-30 cs.LG cs.NA math.NA math.OC stat.ML

Convergence of Continual Learning in Homogeneous Deep Networks

同质深度网络中持续学习的收敛性

Matan Schliserman, Gon Buzaglo, Itay Evron, Daniel Soudry

机构 * Blavatnik School of Computer Science and AI, Tel Aviv University(塔夫茨大学Blavatnik计算机科学与人工智能学院) Princeton University(普林斯顿大学) Department of Electrical and Computing Engineering, Technion(技术学院电子与计算工程系)

AI总结 将弱正则化持续分类建模为任务间隔集上的顺序投影,证明全局收敛一般失败,但利用非凸投影理论识别同质深度网络的局部线性收敛条件,并扩展至持续回归。

详情
AI中文摘要

我们将同质模型中的弱正则化持续分类刻画为任务间隔集上的顺序投影。这一结果推广了先前仅限于平稳(单任务)深度模型或持续线性模型的分析。我们证明,即使对于数据线性但参数非线性的简单模型,全局收敛通常也会失败。然而,通过利用非凸投影理论的结果,我们识别了同质深度网络的规则性性质,这些性质保证了在随机和循环任务序列下的局部线性收敛。最后,我们将分析扩展到持续回归,统一了同质模型的框架。

英文摘要

We characterize weakly regularized continual classification in homogeneous models as sequential projections onto task margin sets. This result generalizes prior analyses restricted to either stationary (single-task) deep models or continual linear models. We show that global convergence generally fails, even for simple models linear in data but nonlinear in parameters. Nevertheless, by leveraging results from nonconvex projection theory, we identify regularity properties of homogeneous deep networks that guarantee local linear convergence under random and cyclic task sequences. Finally, we extend our analysis to continual regression, unifying the framework for homogeneous models.

URL PDF HTML 收藏
2606.30410 2026-06-30 cs.LG cs.AI

Beyond IID: How General Are Tabular Foundation Models, Really?

超越IID:表格基础模型到底有多通用?

Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David Holzmüller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, Gaël Varoquaux, Frank Hutter

机构 * Prior Labs University of Freiburg(弗赖堡大学) University of Mannheim(曼海姆大学) INRIA Saclay(法国国家信息与自动化研究所萨克雷研究中心) Technion(技术学院) ELLIS Institute Tübingen(图宾根ELLIS研究所) Zuse School ELIZA(Zuse学校ELIZA)

AI总结 针对表格数据基础模型评估碎片化问题,提出统一基准BeyondArena,涵盖多种任务类型、数据规模和特征类型,发现现有模型在IID小数据上表现优异,但非IID、大数据和高维数据仍由传统模型主导。

详情
AI中文摘要

用于表格数据预测机器学习的基础模型最近在学术界和工业界引起了广泛关注。跨学科的研究社区越来越多地在各种数据集和任务上评估表格基础模型。然而,这些针对任务和学科的评估对模型研究者来说仍然难以获取,因为基准软件和评估协议是碎片化的。因此,模型研究者依赖标准基准,而这些基准主要定义在表格基础模型已经擅长的任务上。最具挑战性的场景被排除在外,限制了该领域的实质性进展,因为研究集中在IID数据上的边际改进,而非更广泛、更艰巨的挑战。为克服这一问题,我们引入了BeyondArena,这是第一个统一的表格数据整体基准,支持多种任务类型(IID、时间序列、分组),涵盖样本量和特征维度规模,以及来自广泛学科的各种特征类型(含文本、高基数)。为了实现超越标准基准的统一评估,我们引入了Data Foundry,一个用于策划预测机器学习表格数据集的Python框架和元数据模式。我们在11个模型和142个策划数据集上的结果表明,现有的表格基础模型在小型到中型IID数据上表现出色,而传统的基于树和深度学习模型在非IID、大型和高维数据集上仍然占主导地位。BeyondArena引导模型研究应对表格数据中最具挑战性的需求,推动实现真正基础性的表格模型。

英文摘要

Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry. Research communities across disciplines are increasingly evaluating tabular foundation models on diverse datasets and tasks. However, these task- and discipline-specific evaluations remain largely inaccessible to model researchers because benchmark software and evaluation protocols are fragmented. As a result, model researchers rely on standard benchmarks, which are mostly defined for tasks where tabular foundation models already excel. The most challenging scenarios are excluded, limiting meaningful progress in the field by focusing on marginal improvements on IID data rather than on broader, more demanding challenges. To overcome this, we introduce BeyondArena, the first unified holistic benchmark for tabular data that supports diverse task types (IID, temporal, grouped), across sample size and feature dimensionality scales, with diverse feature types (with text, with high cardinality) from a broad range of disciplines. To enable unified benchmarking beyond standard benchmarks, we introduce Data Foundry, a Python framework and metadata schema for curating tabular datasets for predictive machine learning. Our results across 11 models and 142 curated datasets show that existing tabular foundation models excel on tiny- to medium-sized IID data, while traditional tree-based and deep learning models still dominate on non-IID, large, and high-dimensional datasets. BeyondArena guides model research for the most demanding challenges in tabular data, enabling progress towards truly foundational tabular models.

URL PDF HTML 收藏
2606.29322 2026-06-30 cs.LG

SP-CACW: Convergence-Aware Client Weighting for Selfish Personalized Learning

SP-CACW:面向自私个性化学习的收敛感知客户端加权

Yaron Kiselman, Kfir Y. Levy

机构 * Technion Haifa, Israel(技术学院海法分校,以色列)

AI总结 提出SP-CACW框架,通过最小化目标客户端收敛误差上界选择聚合权重,平衡同伴偏差与随机方差,可分配零权重给有害同伴,在MNIST、CIFAR-100和LEAF Shakespeare上表现优于强基线。

Comments 31 pages, 6 figures

详情
AI中文摘要

协作学习只有在每个参与者都受益时才是可持续的。标准联邦学习优化全局平均目标,对于数据分布与总体差异较大的客户端,其表现可能不佳。我们研究自私个性化:指定目标客户端如何利用同伴梯度最小化自身风险,同时避免负迁移。我们提出SP-CACW,一个收敛感知的客户端加权框架,通过最小化目标客户端收敛误差的上界来选择聚合权重。得到的规则明确权衡同伴偏差与随机方差,并可以为有害同伴分配零权重。我们在平滑和有界方差假设下提供了收敛保证,并在MNIST、CIFAR-100和LEAF Shakespeare上评估了该方法,其性能与强个性化和聚类基线相当或更优。

英文摘要

Collaborative learning is sustainable only when it benefits each participant. Standard federated learning optimizes a global average objective, which can under perform for clients whose data distributions differ substantially from the population. We study selfish personalization: how a designated target client can use peer gradients to minimize its own risk while avoiding negative transfer. We propose SP-CACW, a convergence-aware client-weighting framework that selects aggregation weights by minimizing an upper bound on the target client's convergence error. The resulting rule explicitly trades off peer bias against stochastic variance and can assign zero weight to harmful peers. We provide convergence guarantees under smoothness and bounded-variance assumptions and evaluate the method on MNIST, CIFAR-100, and LEAF Shakespeare, where it is competitive with or improves over strong personalized and clustering baselines.

URL PDF HTML 收藏