arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Georgia Institute of Technology(佐治亚理工学院)

至 收录 1456
2607.17077 2026-07-21 cs.CV cs.AI 新提交

ALLUDE: A Unified Evaluation System for Configurable Attacks in Differentiable Environments

ALLUDE:可微环境中可配置攻击的统一评估系统

Mansi Phute, Alexander Greenhalgh, Matthew Hull, Haoran Wang, Alec Helbling, ShengYun Peng, Elliott Faa, Willian Lunardi, Martin Andreoni, Wenke Lee, Duen Horng Chau

机构 * Georgia Institute of Technology(佐治亚理工学院) Technological Innovation Institute(技术创新研究所)

AI总结 针对视觉模型对抗攻击评估条件有限的问题,ALLUDE提供统一评估系统,通过拉丁超立方抽样和压力测试现有攻击展示评估广度,利用端到端可微渲染针对实际部署优化攻击,且跨平台代码开源。

详情
AI中文摘要

对目标检测器等视觉模型的对抗攻击评估条件有限,性能未充分表征。整合模拟和可微渲染能实现更强大的端到端评估,但缺乏易用统一系统。ALLUDE填补了这些空白,通过双管齐下策略展示评估广度:一是拉丁超立方抽样,二是压力测试现有攻击。其端到端可微渲染能针对实际部署条件优化攻击,跨平台代码开源。

英文摘要

Adversarial attacks against vision models like object detectors are often evaluated under limited conditions, leaving their performance under-characterized. Bridging simulation and differentiable rendering enables more robust, end-to-end evaluation of these adversarial attacks, yet there is no easy-to-use, unified system that offers a rich set of customizable configurations for adversarial attacks across multiple scenes, objects, environmental and lighting conditions, and camera trajectories. We present ALLUDE, which addresses these gaps, offering first-of-its-kind evaluation capabilities across Linux and Windows. We comprehensively demonstrate ALLUDE's evaluation breadth through a two-pronged strategy: (1) using Latin Hypercube Sampling, we draw a representative subset from 5,400 configurations spanning 10 scene-object pairs, 9 weather conditions, 4 optimizers, 5 camera trajectories, and 3 detection models; (2) we stress-test existing attacks (CAMOU, RAUCA, FCA) under diverse weather conditions and continuous camera trajectories, revealing degradation of attack success across every attack, exposing evaluation gaps in prior work. Through ALLUDE's end-to-end differentiable rendering, adversarial attacks can be optimized against shifting real-world deployment conditions. Our cross-platform code is open source.

URL PDF HTML 收藏
2607.16712 2026-07-21 cs.AI cs.CL 新提交

DS@GT ARC at eRisk 2026: Hybrid Multi-Agent LLM System with Structured Algorithmic Guidance for Conversational Depression Screening

DS@GT在eRisk 2026的ARC:用于对话式抑郁症筛查的具有结构化算法指导的混合多智能体大语言模型系统

Victor Gong, David Guecha

机构 * Georgia Institute of Technology(佐治亚理工学院)

AI总结 介绍DS@GT参加eRisk 2026对话式抑郁症筛查挑战赛,其系统历经三个阶段,最终采用混合配置并添加算法组件。提交三次运行结果,混合运行3表现出色,以低成本超付费基线,验证较弱开源模型经算法监督可与专有模型竞争。

Comments Conference and Labs of the Evaluation Forum (CLEF), 19 pages, 5 figures

详情
AI中文摘要

我们描述了DS@GT提交给eRisk 2026对话式抑郁症筛查任务1挑战赛的内容。在该挑战赛中,系统要访谈模拟不同抑郁症特征个体的大语言模型角色,并给出贝克抑郁量表II(BDI-II)得分及每个角色的四个关键症状,且不直接询问敏感心理健康问题。我们的流程历经三个阶段:从整体单模型原型开始,到基线多智能体架构,再到最终混合配置,用开源的Gemma 27B取代付费的GPT-5-nano访谈器。为弥补模型推理和指令遵循能力较弱的问题,混合配置添加了三个算法组件。我们提交了针对所有20个角色的三次全自动运行结果。混合运行3的ADODL为0.9063,在所有完整提交运行中排名第三,在21个团队中DS@GT整体排名第二,且以约四分之一的人均API成本超过付费基线运行1。这些结果支持了我们的核心假设,即通过足够的算法监督,较弱的开源模型能在对话访谈角色中与较强的专有模型竞争。我们的源代码可在该https网址获取。

英文摘要

We describe DS@GT's submission to the eRisk 2026 Task 1 challenge on conversational depression screening, in which systems interview LLM personas that simulate individuals with varying depression profiles and produce a Beck Depression Inventory II (BDI-II) score plus four key symptoms per persona, without directly asking sensitive mental health questions. Our pipeline evolved through three stages: a monolithic single-model prototype to start off, a baseline multi-agent architecture that separates conversational interviewing from BDI-II scoring under a coordinating orchestration layer, and a final hybrid configuration that replaces the paid GPT-5-nano interviewer with the open-source Gemma 27B. To offset the model's weaker reasoning and instruction-following, the hybrid adds three algorithmic components: a precomputed dialogue tree that standardizes interview openers and follow-ups, a reliability-weighted consensus aggregation inspired by the Weaver framework, and a cluster-based imputation step for unprobed symptoms. We submitted three fully automated runs across all 20 personas, with Run 1 from the paid baseline and Runs 2 and 3 from the hybrid. Hybrid Run 3 achieved an ADODL of 0.9063, ranking 3rd among all complete-submission runs and placing DS@GT 2nd among the 21 teams overall, while outperforming our paid baseline Run 1 (0.8841) at roughly one-quarter of the per-persona API cost. These results support our central hypothesis that with sufficient algorithmic supervision, a weaker open-source model can compete with a stronger proprietary model in the conversational interviewer role. Our source code is available at https://github.com/dsgt-arc/erisk-task1-2026.

URL PDF HTML 收藏
2607.16657 2026-07-21 cs.SD 新提交

HARP: Harmonic-Aware Residual Partitioning for Neural Audio Codecs

HARP:用于神经音频编解码器的谐波感知残差划分

Qiaoyu Yang, Lixing He, Binyue Deng, Weifeng Zhao

机构 * Georgia Institute of Technology(佐治亚理工学院) The Chinese University of Hong Kong(香港中文大学) Tencent Music Entertainment(腾讯音乐娱乐集团)

AI总结 研究针对神经音频编解码器中码本频谱纠缠等问题,提出HARP训练策略,将RVQ阶段分组,在解码器能访问低频时各小组细化目标频带,重建泛音保留连贯性,该策略无需架构改变,性能优于标准RVQ和并行分解。

Comments Accepted to Interspeech 2026

详情
AI中文摘要

具有残差向量量化(RVQ)的神经音频编解码器通常对所有频率一视同仁,导致其码本频谱纠缠。截断阶段会去除不可预测的频率混合。并行频带分解通过将音频拆分为独立频带来解决此问题,但会使潜在空间碎片化并失去跨频率连贯性。我们引入了HARP(谐波感知残差划分),这是一种训练策略,将RVQ阶段划分为按频率排序的组,每个组在解码器仍可访问所有低频的同时细化其目标频带。泛音在基音的背景下重建,保留了并行方法所失去的连贯性。HARP无需架构更改,仅修改训练损失,推理与标准RVQ相同。在语音、音乐和一般音频上,HARP优于标准RVQ和并行分解。MUSHRA听力测试也显示出感知上的改进。

英文摘要

Neural audio codecs with residual vector quantization (RVQ) normally treat all frequencies uniformly, so their codebooks become spectrally entangled. Truncating stages then removes an unpredictable mix of frequencies. Parallel band decomposition addresses this by splitting audio into independent bands, but fragments the latent space and loses cross-frequency coherence. We introduce HARP (Harmonic-Aware Residual Partitioning), a training strategy that partitions RVQ stages into frequency-ordered groups where each group refines its target band while the decoder retains access to all lower frequencies. Overtones are reconstructed in the context of their fundamentals, preserving coherence that parallel methods lose. HARP requires no architectural changes; it only modifies the training loss, leaving inference identical to standard RVQ. On speech, music, and general audio, HARP outperforms both standard RVQ and parallel decomposition. MUSHRA listening tests also show perceptual improvements.

URL PDF HTML 收藏
2607.16453 2026-07-21 cs.CV 新提交

DS@GT ARC at AnimalCLEF 2026: Species-Aware Graph Construction for Multi-Species Animal Re-Identification

DS@GT ARC参加2026年动物CLEF:用于多物种动物重新识别的物种感知图构建

Evan Sinclair Smith, Anthony Miyaguchi, Snigdha Palamari, Danté Evangelista

机构 * Georgia Institute of Technology(佐治亚理工学院)

AI总结 该研究针对多物种动物重新识别问题,提出将其作为物种感知图构建,通过整合预处理、检索、验证、评分、边缘接纳和社区检测等流程,提升识别效果,所选提交在230个团队中排名第五,凸显视觉表示与多种约束校准整合的重要性。

Comments 18 pages, 8 figures. Accepted to the CLEF 2026 Working Notes. Code: https://github.com/dsgt-arc/animalclef-2026

详情
AI中文摘要

自动个体动物重新识别对于大规模生物多样性监测至关重要。然而,野外图像使得从姿态、光照、背景、分辨率和物种特定形态的干扰变化中分离身份线索变得复杂。DS@GT ARC提交给2026年动物CLEF的论文介绍了一种多物种图像聚类系统,用于重新识别欧亚猞猁、火蝾螈、蠵龟和德州角蜥。该方法将重新识别表述为对候选图像对进行物种感知图构建,而不是依赖单个描述符或最近邻检索。其流程整合了定制预处理、全局候选检索、基于LightGlue的多关键点家族局部验证、LightGBM对评分、保守边缘接纳和莱顿社区检测。跨物种的消融研究表明,局部特征支持、前景感知预处理和物种特定主干选择增强了对证据,而图操作点决定了碎片化和过度合并之间的权衡。所选提交在230个团队中排名第五,公开ARI为0.733,私有ARI为0.674。这些结果表明,强大的野生动物重新识别不仅需要强大的视觉表示,还需要对全局相似性、局部身份标记、邻域上下文和图级约束进行校准整合。代码可在该https URL找到。

英文摘要

Automated individual animal re-identification is essential for large-scale biodiversity monitoring; however, field imagery complicates separating identity cues from nuisance variation in pose, illumination, background, resolution, and species-specific morphology. The DS@GT ARC submission to AnimalCLEF 2026 introduces a multi-species image-clustering system for re-identifying Eurasian lynx, fire salamanders, loggerhead sea turtles, and Texas horned lizards. Instead of relying on a single descriptor or nearest-neighbor retrieval, this approach formulates re-identification as species-aware graph construction over candidate image pairs. The pipeline integrates tailored preprocessing, global candidate retrieval, LightGlue-based local verification with multiple keypoint families, LightGBM pair scoring, conservative edge admission, and Leiden community detection. This design directly addresses a primary failure mode of clustering-based re-identification: high-scoring false pairs that act as bridge edges and merge distinct individuals through transitive closure. Across species, ablation studies demonstrate that local feature support, foreground-aware preprocessing, and species-specific backbone selection enhance pair evidence, while graph operating points determine the trade-off between fragmentation and over-merging. The selected submission achieved a public ARI of 0.733 and a private ARI of 0.674, ranking fifth among 230 teams. These results indicate that robust wildlife re-identification requires not only strong visual representations but also calibrated integration of global similarity, local identity markings, neighborhood context, and graph-level constraints. The code can be found at https://github.com/dsgt-arc/animalclef-2026.

URL PDF HTML 收藏
2607.16384 2026-07-21 cs.LG math.OC math.PR stat.ML 新提交

Scaling Limits of Constant-Stepsize SGD at Flat Minima

常步长随机梯度下降在平坦极小值处的缩放极限

Jingyi Zhang, Cheng Mao, Debankur Mukherjee

机构 * Georgia Institute of Technology(佐治亚理工学院)

AI总结 研究常步长随机梯度下降在平坦极小值处的缩放极限,通过分析由压缩驱动链生成马尔可夫噪声的SGD,证明其在特定距离下的收敛性,给出不同平坦指数下的收缩因子及小步长缩放极限,揭示了与强凸情况不同的行为。

Comments 52 pages, 3 figures, 1 table

详情
AI中文摘要

对于具有常步长α的随机梯度下降(SGD),以极小值点为中心的迭代不变律描述了算法在长时间范围内的行为。在强凸情况下,该不变律具有熟悉的√α缩放,且当α↓0时极限为高斯分布。我们表明,对于具有平坦极小值和(次)二次尾部的凸目标H,这种行为会发生根本变化。具体而言,我们研究由压缩驱动链生成马尔可夫噪声的SGD。对于每个足够小的常步长α,我们证明了在由α相关度量诱导的Wasserstein距离中,存在、唯一且几何收敛到一个增强的不变律。当极小值点x*具有局部平坦指数m≥2时,我们得到收缩因子为1 - cα^(m - 1),在二次情况m = 2时恢复为1 - cα。然后我们分析小步长缩放极限。我们表明不变律集中在α^(1/m)尺度上,并且重新缩放后的迭代弱收敛到随机微分方程dY_t = -h_0(Y_t)dt + Σ^(1/2)dB_t的平稳分布,其中h_0是极小值点处的极限漂移,Σ表示渐近协方差。当m = 2时恢复高斯极限,在平坦情况m>2时通常给出非高斯平稳极限。最后,我们给出了具有不等平坦指数的坐标可分目标的相应结果。

英文摘要

For stochastic gradient descent (SGD) with a constant stepsize $α$, the invariant law of the iterates, centered at a minimizer, describes the behavior of the algorithm over long time horizons. In the strongly convex case, this invariant law has the familiar $\sqrtα$ scaling and a Gaussian limit as $α\downarrow 0$. We show that this behavior changes fundamentally for convex objectives $H$ with flat minima and (sub)quadratic tails. More specifically, we study SGD with Markovian noise generated by a contractive driving chain. For every sufficiently small constant stepsize $α$, we prove existence, uniqueness, and geometric convergence to an augmented invariant law in a Wasserstein distance induced by an $α$-dependent metric. When the minimizer $x_\star$ has local flatness exponent $m\ge2$, meaning that $\nabla^2 H(x)\asymp \lVert x-x_\star\rVert^{m-2} I_d$ as $x\to x_\star$, we obtain a contraction bound with factor $1-cα^{m-1}$, where $c>0$ is a constant. This recovers the factor $1-cα$ in the quadratic case $m=2$. We then analyze the small-stepsize scaling limit. We show that the invariant law concentrates on the scale $α^{1/m}$ and that the rescaled iterates converge weakly to the stationary distribution of the stochastic differential equation $$ dY_t=-h_0(Y_t)\,dt+Σ^{1/2}\,dB_t , $$ where $h_0$ is the limiting drift at the minimizer and $Σ$ denotes the asymptotic covariance. This recovers the Gaussian limit when $m=2$ and gives generally non-Gaussian stationary limits in the flat case $m>2$. Finally, we give corresponding results for coordinate-separable objectives with unequal flatness exponents.

URL PDF HTML 收藏
2607.17033 2026-07-21 cs.LG physics.chem-ph 新提交

ChemFusion: A Multimodal Cross-Attention Network for Reaction Yield Prediction

ChemFusion:用于反应产率预测的多模态交叉注意力网络

Qiwei Han, Chi Zhou

机构 * Duke University(杜克大学) Georgia Institute of Technology(佐治亚理工学院)

AI总结 研究过渡金属催化反应产率预测难题,提出ChemFusion多模态交叉注意力网络,融合电子特征与3D原子坐标,用交叉注意力机制,在交叉偶联库基准测试中性能出色,还能自主学习识别和惩罚空间位阻,提供物理可解释性。

Comments 10 pages, 4 figures, 2 tables

详情
AI中文摘要

预测过渡金属催化反应的结果因多种物理和化学变量的相互作用而极其复杂。一个长期存在的计算瓶颈是有效地将广泛的电子描述符与反应位点的局部三维几何结构合并。为弥合这种表示差距,我们提出了ChemFusion,一种将传统电子特征与明确的3D原子坐标融合的混合神经网络。该模型使用交叉注意力机制,使全局电子态能够动态关注未池化分子点云内的特定空间约束。在针对各种交叉偶联库进行基准测试时,此方法具有出色的预测性能,远超传统单模态框架。重要的是,提取注意力矩阵表明该架构能自主学习识别和惩罚限制性空间位阻。这提供了基于物理的可解释性,表明有空间意识的网络可以应对标准统计模型通常忽略的复杂反应空间位阻。

英文摘要

Forecasting the outcomes of transition-metal-catalyzed reactions is notoriously complex due to the interplay of diverse physical and chemical variables. A persistent computational bottleneck has been effectively merging broad electronic descriptors with the localized, three-dimensional geometry of the reactive site. To bridge this representation gap, we present ChemFusion, a hybrid neural network that fuses conventional electronic features with explicit 3D atomic coordinates. Using a cross-attention mechanism, the model enables global electronic states to dynamically attend to specific spatial constraints within un-pooled molecular point clouds. When benchmarked against a diverse library of cross-couplings, this approach delivers exceptional predictive performance, decisively surpassing traditional single-modality frameworks. Importantly, extracting the attention matrices reveals that the architecture autonomously learns to identify and penalize restrictive steric hindrances. This provides a physically grounded interpretability, demonstrating that spatially aware networks can navigate complex reaction sterics that standard statistical models typically miss.

URL PDF HTML 收藏
2409.12190 2026-07-21 cs.RO cs.CV

Bundle Adjustment in the Eager Mode

急切模式下的捆绑调整

Zitong Zhan, Huan Xu, Zihang Fang, Xinpeng Wei, Yaoyu Hu, Chen Wang

机构 * Spatial AI & Robotics (SAIR) Lab, University at Buffalo(空间人工智能与机器人实验室,布法罗大学) Georgia Institute of Technology(佐治亚理工学院) Purdue University(普渡大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出了一种与PyTorch无缝集成的高效急切模式捆绑调整库,通过稀疏感知的自动微分设计和GPU加速的稀疏运算,提升了在机器人应用中捆绑调整的运行效率和性能。

Journal ref IEEE Transactions on Robotics (T-RO), 2026

详情
AI中文摘要

捆绑调整(BA)是各种机器人应用中的关键技术,例如同步定位与建图(SLAM)、增强现实(AR)和摄影测量学。BA通过优化诸如相机姿态和3D地标等参数,使它们与观测结果对齐。随着深度学习在感知系统中的重要性日益增加,将BA与深度学习框架整合已成为提高可靠性和性能的迫切需求。然而,广泛使用的基于C++的BA库,如GTSAM、g²o和Ceres Solver,缺乏与现代深度学习库如PyTorch的原生整合。这种限制影响了它们的灵活性、调试简便性和整体实现效率。为了解决这一差距,我们引入了一种与PyTorch无缝集成的高效急切模式BA库。我们的方法包括稀疏感知的自动微分设计和针对二次优化设计的GPU加速稀疏运算。我们的GPU急切模式BA在所有基准测试中均实现了显著的运行时间效率,与GTSAM、g²o和Ceres相比,平均加速分别为18.5×、22×和23×。

英文摘要

Bundle adjustment (BA) is a critical technique in various robotic applications such as simultaneous localization and mapping (SLAM), augmented reality (AR), and photogrammetry. BA optimizes parameters such as camera poses and 3D landmarks to align them with observations. With the growing importance of deep learning in perception systems, there is an increasing need to integrate BA with deep learning frameworks for enhanced reliability and performance. However, widely-used C++-based BA libraries, such as GTSAM, g$^2$o, and Ceres Solver, lack native integration with modern deep learning libraries like PyTorch. This limitation affects their flexibility, ease of debugging, and overall implementation efficiency. To address this gap, we introduce an eager-mode BA library seamlessly integrated with PyTorch with high efficiency. Our approach includes a sparsity-aware auto-differentiation design and GPU-accelerated sparse operations designed for 2nd-order optimization. Our eager-mode BA on GPU demonstrates substantial runtime efficiency, achieving an average speedup of 18.5$\times$, 22$\times$, and 23$\times$ across all benchmarks compared to GTSAM, g$^2$o, and Ceres, respectively.

URL PDF HTML 收藏
2605.14473 2026-07-21 cs.CL cs.AI 版本更新

Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict

RAG 能知道检索错误吗?知识冲突下的上下文合规性诊断

Yihang Chen, Pin Qian, Su Wang, Sipeng Zhang, Huan Xu, Shuhuai Lin, Xinpeng Wei

机构 * Georgia Institute of Technology(佐治亚理工学院) Carnegie Mellon University(卡内基梅隆大学) University of California San Diego(加州大学圣地亚哥分校)

AI总结 提出上下文驱动分解(CDD)方法,在推理时探测并干预检索增强生成中的上下文与参数知识冲突,揭示上下文合规性模式并提升鲁棒性。

Comments Preprint. 3 figures, 3 tables. Diagnostic study of context compliance in RAG under knowledge conflict; closed-API evaluations (Gemini-2.5-Flash and Claude family)

详情
AI中文摘要

检索增强生成(RAG)中的上下文合规机制发生在检索到的上下文主导最终答案时,即使它与模型的参数化知识冲突。仅凭准确性并不能揭示在这种冲突下检索到的上下文如何因果性地塑造答案。我们引入了上下文驱动分解(CDD),这是一种在推理时运行的信念分解探针,并作为受控检索冲突的干预机制。通过跨Epi-Scale压力测试、TruthfulQA错误概念注入和跨模型重复实验,CDD揭示了三种模式。P1:上下文合规性在对抗性上界设置中是可测量的,标准RAG在TruthfulQA错误概念注入(N=500)上达到15.0%的准确率。P2:对抗性准确率提升跨模型家族迁移——CDD提高了Gemini-2.5-Flash以及Claude Haiku/Sonnet/Opus的准确率——但理由-答案因果耦合不迁移。CDD在Gemini-2.5-Flash上达到64.1%的错误注入因果敏感性,而所有三种Claude变体的敏感性落在[-3%, +7%]范围内,表明Claude侧的准确率提升通过一种与显式冲突解决轨迹不同的机制运作。P3:显式冲突分解提高了时间漂移和噪声干扰下的鲁棒性,CDD在完整Epi-Scale对抗性基准上对时间偏移达到71.3%,对干扰证据达到69.9%。这三种模式将上下文合规性识别为一个结构轴,沿此轴可以对标准RAG进行探测和干预,区别于检索质量或单一方法鲁棒性问题,并激励发布Epi-Scale以跨模型家族和检索管道进行系统研究。

英文摘要

Retrieval-Augmented Generation (RAG) is usually evaluated by whether the final answer is correct. Under knowledge conflict, this hides a key question: did the model follow retrieved evidence, rely on its parametric prior, or produce a post-hoc rationale? We study this as context compliance, the regime in which retrieved context controls the answer even when it conflicts with the model's prior knowledge. We introduce Context-Driven Decomposition (CDD), an inference-time diagnostic intervention that elicits contextual and prior answers, isolates the conflicting premise, and records a resolution trace that can be perturbed. Across Epi-Scale stress tests, TruthfulQA misconception injection, and cross-model reruns, CDD makes three behaviors visible. First, misleading retrieval can severely degrade accuracy: under a worst-case TruthfulQA misconception-injection probe, Standard RAG reaches only 15.0%. Second, better answers need not share the same mechanism: CDD improves adversarial accuracy on Gemini-2.5-Flash and shows directional gains across Claude variants, yet trace-perturbation sensitivity is high only on Gemini. Third, explicit decomposition improves controlled-conflict robustness over a conflict-aware instruction baseline on localized factual conflicts, with the clearest margins on Entity Swap (88.0% vs 79.3%) and Logical Contradiction (83.2% vs 75.4%). We frame RAG conflict handling as an observability problem.

URL PDF HTML 收藏
2601.00126 2026-07-21 cs.RO 版本更新

Compositional Diffusion with Guided Search for Long-Horizon Planning

用于长期规划的带引导搜索的组合扩散

Utkarsh A Mishra, David He, Yongxin Chen, Danfei Xu

机构 * Georgia Institute of Technology(佐治亚理工学院)

AI总结 研究针对组合生成模型在局部分布多模态时的模式平均问题,提出带引导搜索的组合扩散方法(CDGS),通过特定采样、过滤和重采样解决问题,在机器人操作任务上表现出色且跨领域通用。

Comments 38 pages, 18 figures, ICLR 2026 ORAL

详情
AI中文摘要

生成模型已成为规划的强大工具,组合方法通过组合局部、模块化生成模型,在对长期任务分布建模方面有独特前景,涵盖多领域。但组合生成模型面临关键挑战:局部分布多模态时,现有组合方法会平均不相容模式,产生局部不可行、全局不一致的计划。我们提出带引导搜索的组合扩散(CDGS),通过在扩散去噪过程中直接嵌入搜索解决模式平均问题。该方法通过基于种群的采样探索局部模式的不同组合,用基于似然的过滤去除不可行候选,通过重叠段间的迭代重采样强制全局一致性。CDGS在七个机器人操作任务上匹配神谕性能,优于缺乏组合性或需要长期训练数据的基线。该方法跨领域通用,通过有效的局部到全局消息传递实现连贯的文本引导全景图像和长视频。

英文摘要

Generative models have emerged as powerful tools for planning, with compositional approaches offering particular promise for modeling long-horizon task distributions by composing together local, modular generative models. This compositional paradigm spans diverse domains, from multi-step manipulation planning to panoramic image synthesis to long video generation. However, compositional generative models face a critical challenge: when local distributions are multimodal, existing composition methods average incompatible modes, producing plans that are neither locally feasible nor globally coherent. We propose Compositional Diffusion with Guided Search (CDGS), which addresses this mode averaging problem by embedding search directly within the diffusion denoising process. Our method explores diverse combinations of local modes through population-based sampling, prunes infeasible candidates using likelihood-based filtering, and enforces global consistency through iterative resampling between overlapping segments. CDGS matches oracle performance on seven robot manipulation tasks, outperforming baselines that lack compositionality or require long-horizon training data. The approach generalizes across domains, enabling coherent text-guided panoramic images and long videos through effective local-to-global message passing. More details: https://cdgsearch.github.io/

URL PDF HTML 收藏
2509.08765 2026-07-21 physics.comp-ph cs.LG cs.NA math.NA stat.ML 版本更新

One-shot acceleration of transient PDE solvers via online-learned preconditioners

通过在线学习预处理器实现瞬态偏微分方程求解器的一次性加速

Mikhail Khodak, Min Ki Jung, Brian Wynne, Edmond Chow, Egemen Kolemen

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Seoul National University(首尔国立大学) Princeton University(普林斯顿大学) Georgia Institute of Technology(佐治亚理工学院)

AI总结 研究针对瞬态偏微分方程数值模拟,提出用强盗算法利用线性求解器反馈在线学习预处理器配置的PCGBandit方法,直接在OpenFOAM上实现,可一次性加速模拟,在流体和MHD问题上验证了有效性。

Comments code available at https://github.com/mkhodak/PCGBandit

Journal ref Computer Physics Communications 327 (2026) 110304

详情
AI中文摘要

数据驱动的科学计算工作流程加速一直是机器学习用于科学领域的一个备受瞩目的目标,瞬态偏微分方程的数值模拟是主要应用之一。目前的重点是需要经典模拟来训练的方法,这与神经网络的数据饥渴和优化挑战相结合,使得难以证明相对于强大的经典基线有令人信服的优势。我们考虑一种替代范式,即学习者使用经典求解器自身的数据来加速它,实现模拟的一次性加速。具体而言,由于瞬态偏微分方程通常需要求解一系列相关的线性系统,强盗算法可以利用对线性求解器(如预条件共轭梯度法)的重复调用的反馈来在线学习求解器配置(如预处理器)的自适应序列。我们开发的方法PCGBandit直接在流行的开源软件OpenFOAM之上实现,我们用它来展示其在一组流体和磁流体动力学(MHD)问题上的有效性。

英文摘要

Data-driven acceleration of scientific computing workflows has been a high-profile aim of machine learning (ML) for science, with numerical simulation of transient partial differential equations (PDEs) being one of the main applications. The focus thus far has been on methods that require classical simulations to train, which when combined with the data-hungriness and optimization challenges of neural networks has caused difficulties in demonstrating a convincing advantage against strong classical baselines. We consider an alternative paradigm in which the learner uses a classical solver's own data to accelerate it, enabling a one-shot speedup of the simulation. Concretely, since transient PDEs often require solving a sequence of related linear systems, the feedback from repeated calls to a linear solver such as preconditioned conjugate gradient (PCG) can be used by a bandit algorithm to online-learn an adaptive sequence of solver configurations (e.g. preconditioners). The method we develop, PCGBandit, is implemented directly on top of the popular open-source software OpenFOAM, which we use to show its effectiveness on a set of fluid and magnetohydrodynamics (MHD) problems.

URL PDF HTML 收藏
2511.14592 2026-07-21 cs.RO cs.AI 版本更新

DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks

DSBench:用于评估外部和车内风险的综合基准测试

Xianhui Meng, Yuchen Zhang, Zhijian Huang, Zheng Lu, Ziling Ji, Yandan Lin, Yaoyao Yin, Hongyuan Zhang, Wei Zhou, Guangfeng Jiang, Li Zhang, Long Chen, Hangjun Ye, Jun Liu, Xiaoshuai Hao

机构 * University of Science and Technology of China(中国科学技术大学) Georgia Institute of Technology(佐治亚理工学院) Xiaomi EV(小米电动车) Fudan University(复旦大学) Xidian University(西安电子科技大学) South China University of Technology(华南理工大学) The University of Hong Kong(香港大学)

AI总结 针对视觉语言模型在安全关键场景适用性未充分探索的问题,引入DSBench综合基准测试,涵盖外部与车内风险,分多类别。评估发现模型在复杂情况下降,构建数据集微调可提升性能,推动自动驾驶技术发展,相关资源将公开。

详情
AI中文摘要

视觉语言模型(VLMs)在自动驾驶方面前景广阔,但在安全关键场景中的适用性尚未充分探索,引发安全担忧。由于缺乏同时评估外部环境风险和车内驾驶行为安全的综合基准测试,我们引入了DSBench,首个用于统一评估VLM对各种安全风险认知的综合驾驶安全基准测试。它涵盖外部环境风险和车内驾驶行为安全两大类别,分为10个关键类别和28个子类别。对多种主流开源和闭源VLM的广泛评估显示,在复杂安全关键情况下性能显著下降。为此,我们构建了一个包含98K实例的大型数据集,表明在此数据集上微调可显著提高现有VLM的安全性能,为推进自动驾驶技术铺平道路。基准测试工具包、代码和模型检查点将公开可用。

英文摘要

Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-critical scenarios is largely unexplored, raising safety concerns. This issue arises from the lack of comprehensive benchmarks that assess both external environmental risks and in-cabin driving behavior safety simultaneously. To bridge this critical gap, we introduce DSBench, the first comprehensive Driving Safety Benchmark designed to assess a VLM's awareness of various safety risks in a unified manner. DSBench encompasses two major categories: external environmental risks and in-cabin driving behavior safety, divided into 10 key categories and a total of 28 sub-categories. This comprehensive evaluation covers a wide range of scenarios, ensuring a thorough assessment of VLMs' performance in safety-critical contexts. Extensive evaluations across various mainstream open-source and closed-source VLMs reveal significant performance degradation under complex safety-critical situations, highlighting urgent safety concerns. To address this, we constructed a large dataset of 98K instances focused on in-cabin and external safety scenarios, showing that fine-tuning on this dataset significantly enhances the safety performance of existing VLMs and paves the way for advancing autonomous driving technology. The benchmark toolkit, code, and model checkpoints will be publicly accessible.

URL PDF HTML 收藏
2509.23185 2026-07-21 cs.RO 版本更新

Physically-Feasible Reactive Synthesis for Terrain-Adaptive Locomotion

用于地形自适应运动的物理可行反应式合成

Ziyi Zhou, Qian Meng, Hadas Kress-Gazit, Ye Zhao

机构 * Georgia Institute of Technology(佐治亚理工学院) Cornell University(康奈尔大学)

AI总结 针对四足动物在动态变化地形的运动规划问题,提出综合规划框架,结合反应式合成与混合整数凸规划,采用符号修复机制,经实验验证该框架能识别缺失技能并在关键环境有效响应。

详情
AI中文摘要

我们提出了一个用于四足动物在动态变化、不可预见地形上运动的综合规划框架。现有方法在实时立足点选择上常依赖启发式方法,限制了鲁棒性和适应性,或依赖复杂地形和长视野上计算密集的轨迹优化。相比之下,我们的方法将用于生成构造正确的符号级控制器的反应式合成与用于每个符号转换期间动态且物理可行的脚步规划的混合整数凸规划相结合。为减少对昂贵的混合整数凸规划求解的依赖并适应因物理不可行可能违反的规范,我们采用一种符号修复机制,仅选择性地生成所需的符号转换。在执行过程中,基于实际地形数据的实时混合整数凸规划重新规划,结合运行时符号修复和延迟感知协调,实现离线合成与在线操作之间的无缝衔接。通过广泛的模拟和硬件实验,我们验证了该框架识别缺失运动技能并在安全关键环境中有效响应的能力,包括散落的踏脚石和钢筋场景。

英文摘要

We present an integrated planning framework for quadrupedal locomotion over dynamically changing, unforeseen terrains. Existing methods often depend on heuristics for real-time foothold selection-limiting robustness and adaptability-or rely on computationally intensive trajectory optimization across complex terrains and long horizons. In contrast, our approach combines reactive synthesis for generating correct-by-construction symbolic-level controllers with mixed-integer convex programming (MICP) for dynamic and physically feasible footstep planning during each symbolic transition. To reduce the reliance on costly MICP solves and accommodate specifications that may be violated due to physical infeasibility, we adopt a symbolic repair mechanism that selectively generates only the required symbolic transitions. During execution, real-time MICP replanning based on actual terrain data, combined with runtime symbolic repair and delay-aware coordination, enables seamless bridging between offline synthesis and online operation. Through extensive simulation and hardware experiments, we validate the framework's ability to identify missing locomotion skills and respond effectively in safety-critical environments, including scattered stepping stones and rebar scenarios.

URL PDF HTML 收藏
2607.15079 2026-07-20 cs.AI 版本更新

BrainPilot: Automating Brain Discovery with Agentic Research

BrainPilot:通过智能研究实现大脑发现自动化

Haoxuan Li, Tianci Gao, Jianhe Li, Yang Fan, Runze Shi, Weiran Wang, Tianxiang Zhao, Zezhao Wu, Xiaoyang Jiang, Qihui Zhang, Jia Li, Xiao Xiao, Kai Du, Xiaoxuan Jia, Chao Xie, Lu Mi

机构 * College of AI, Tsinghua University(清华大学人工智能学院) Shanghai Qizhi Institute(上海期智研究院) Business School, Renmin University of China(中国人民大学商学院) School of Physics, Beihang University(北京航空航天大学物理学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院) Behavioral and Cognitive Neuroscience Center, Institute of Science and Technology for Brain-Inspired Intelligence, Fudan University(复旦大学脑科学与智能技术研究院行为与认知神经科学中心) College of Engineering, Georgia Institute of Technology(佐治亚理工学院工程学院) School of Life Sciences & IDG/McGovern Institute for Brain Research, Tsinghua University(清华大学生命科学学院&清华-IDG/麦戈文脑科学研究院) School of Computing and Artificial Intelligence, Southwest Jiaotong University(西南交通大学计算机与人工智能学院) Weixian College, Tsinghua University(清华大学未央学院) Department of Psychological and Cognitive Sciences, Tsinghua University(清华大学心理学与认知科学系)

AI总结 研究针对脑科学研究整合证据难、人工智能代理有缺陷的问题,提出完全开源的多智能体系统BrainPilot,它有可追溯日志和验证结果,含知识库与技能库,经实验评估,其开源模型以低成本达先进框架性能。

详情
AI中文摘要

理解大脑越来越依赖跨尺度、模态和学科整合证据。解决单个研究问题需要一系列协调操作。人工智能代理有望加速这一过程,但当前代理在脑科学领域缺乏专业知识,可能编造主张,在多步推理中偏离,且专家干预点少。我们提出了BrainPilot,一个完全开源的多智能体系统,它通过可追溯的日志和经智能体验证的结果加速脑科学研究。主要研究者(PI)智能体协调基于精心策划的领域知识的专家智能体,包括一个包含7233个索引条目的统一脑科学知识库和一个涵盖七个研究领域的72个可重复使用方法单元的技能库。每个主要步骤都记录在追踪图中,审核智能体将伪造检查集成到工作流程中。为了评估,我们运行了来自智能体期末考试的三个脑科学任务,引入了自己的基准BrainPilotBench-v0,并展示了其他端到端案例研究。在这些评估中,具有开源主干模型的BrainPilot以更低成本实现了与最先进智能体框架相当的性能。

英文摘要

Understanding the brain increasingly depends on integrating evidence across scales, modalities, and disciplines. Addressing a single research question therefore requires a coordinated sequence of operations, from surveying prior work to executing analyses and interpreting results in light of domain knowledge. AI agents promise to accelerate this process, but current agents lack domain expertise in brain science, may fabricate claims, drift during multi-step reasoning, and offer few defined points for expert intervention. These failures are especially costly in brain science, where conclusions feed into downstream scientific claims and depend on laboratory-specific expertise and careful human judgment. We present \textbf{BrainPilot} a \textbf{fully open-source} multi-agent system that accelerates brain science research with traceable logs and agent-verified results. A principal investigator (PI) agent coordinates specialist agents grounded in curated domain knowledge: a unified brain science knowledge base containing 7{,}233 indexed items and a skill library of 72 reusable methodology units across seven research domains. Every major step is recorded in the Graph of Trace, an auditable record that links subgoals, tool use, evidence, and claims and allows researchers to follow and inspect the workflow. An Auditor agent further integrates fabrication checking into the workflow. For evaluation, we run three brain science tasks from Agents' Last Exam, introduce our own benchmark, \textbf{BrainPilotBench-v0}, and present additional end-to-end case studies. Across these evaluations, BrainPilot with an open-source backbone model attains performance comparable to state-of-the-art agent framework with less costs.

URL PDF HTML 收藏
2607.03651 2026-07-20 cs.LG math.OC 版本更新

LLM-Guided Transportation Hub Capacity Planning with Textual Business Inputs

基于文本业务输入的大语言模型引导的交通枢纽容量规划

Xiaoyue Liu, Zheng Dong

机构 * Georgia Institute of Technology(佐治亚理工学院)

AI总结 研究提出用大语言模型代理依自然语言业务描述迭代提出枢纽容量决策的框架,核心机制是思维链推理协议,经反馈回路验证决策,在实际货运网络中表现优于传统模型。

详情
AI中文摘要

传统枢纽容量规划模型虽能有效优化定量输入,但常无法处理定性业务背景。我们提出一个新颖框架,其中大语言模型(LLM)代理根据自然语言业务背景描述迭代地提出枢纽容量决策。关键机制是思维链推理协议:LLM构建一个结构化决策表,根据变化的隐含方向和幅度将每个上下文项目映射到特定的容量调整。然后通过与优化模型的反馈回路验证新的容量决策,该优化模型提供基于路由的性能指标以指导代理的选择。在美国东南部一个真实的13枢纽货运网络上,相对于隐藏的真实情况,我们的框架实现了2.8%的最优差距,与没有文本业务输入的传统优化模型产生的11.0%的差距相比有显著改善。这表明大语言模型可以作为一个上下文桥梁,将定性业务见解整合到运筹学工作流程中。

英文摘要

While traditional hub capacity planning models optimize effectively for quantitative inputs, they often fail to digest qualitative business context. We propose a novel framework where a large language model (LLM) agent iteratively proposes hub capacity decisions guided by natural-language business context descriptions. The key mechanism is a chain-of-thought reasoning protocol: the LLM constructs a structured decision table that maps each contextual item to specific capacity adjustments based on the implied direction and magnitude of changes. The new capacity decision is then validated through a feedback loop with an optimization model, which provides routing-based performance metrics to guide the agent's selection. On a real-world 13-hub freight network in the southeastern US, our framework achieves a 2.8% optimality gap relative to the hidden ground-truth, a significant improvement over the 11.0% gap produced by the traditional optimization model without textual business inputs. This demonstrates that LLMs can serve as a contextual bridge, integrating qualitative business insights into Operations Research workflows.

URL PDF HTML 收藏
2607.15172 2026-07-17 cs.RO 新提交

AHEAD: Anticipatory Hand-Driven Teleoperation via Human Intent Prediction

AHEAD:通过人类意图预测实现预期的手动驱动遥操作

Seok Joon Kim, Junho Lee, Federica Spinola, Taein Kwon, Mohsen Moghaddam

机构 * Georgia Institute of Technology(佐治亚理工学院) Neuromeka Ltd.(Neuromeka有限公司) INRIA(法国国家信息与自动化研究所) University of Oxford(牛津大学)

AI总结 研究旨在减少机器人反应时间并降低操作员工作量,提出AHEAD实时VR遥操作系统,通过处理手和头部信号及场景上下文预测意图,转换为稳定目标,该系统意图预测准确率高,能有效减少延迟并降低操作员负荷。

Comments Accepted to IROS2026, 8 pages, 6 figures

详情
AI中文摘要

直接手动驱动遥操作能精确控制,但在接近、抓取和放置过程中需持续监控和校正,效率低且易疲劳。监督式遥操作虽简化流程,但有延迟。为解决如何减少机器人反应时间并降低操作员工作量的问题,提出AHEAD实时VR遥操作系统。在数字孪生中,操作员自然执行抓取和放置,AHEAD通过基于注意力的分类器处理手和头部信号及场景上下文来预测意图,状态机将意图预测转换为稳定目标。其意图预测模块在抓取对象和目标插槽上的Top1准确率达76%,用户研究表明AHEAD相对于基线分别减少了0.6秒(对象)和1.4秒(插槽)的机器人反应延迟,还降低了操作员负荷。

英文摘要

Direct hand-driven teleoperation maps an operator's hand motion to robot end-effector commands at every frame, enabling precise control, but it requires constant monitoring and correction during approach, grasp, and placement, which can be slow and fatiguing. For repetitive pick-and-place tasks, supervisory (goal-based) teleoperation simplifies this process: the operator specifies goals/waypoints, and the robot executes the motion using planning algorithms. Yet, this introduces latency, as the robot must wait for the next command before it can plan and act. "How can we reduce robot reaction time while lowering operator workload?" To tackle this question, we present AHEAD, a real-time VR teleoperation system that anticipates operator intent to enable proactive, hand-driven control. In a digital twin, the operator performs pick-and-place naturally, using hand motion to convey high-level commands rather than a continuous robot trajectory. AHEAD processes a short window of 3D hand and head signals together with scene context through an attention-based classifier to predict the intended grasp object and placement slot. A state machine converts intent predictions into stable robot goals, enabling early motion while remaining stable under noisy predictions and corrective hand movements. AHEAD's intent prediction module achieves Top1 accuracy: 76% for grasp objects and 76% for target slots. Moreover, our user study shows AHEAD reduces robot reaction latency by 0.6 s (object) and 1.4 s (slot) relative to baselines. Participants also reported lower operator load, indicating faster robot responses while maintaining low operator effort in practice.

URL PDF HTML 收藏
2607.14509 2026-07-17 cs.CV cs.AI cs.LG 新提交

Multi-Scale ViT Inference with Habitat-Fit Priors and kNN Retrieval for Multi-Species Plant Identification

基于栖息地适应先验和kNN检索的多尺度ViT推理用于多物种植物识别

Alper Erten, Murilo Gustineli, Adrian Cheung

机构 * Georgia Institute of Technology(佐治亚理工学院)

AI总结 针对植被样方图像多物种植物识别难题,基于微调的DINOv2 ViT-L/14分类器,采用多尺度切片分解、kNN检索、栖息地适应降级等方法,在PlantCLEF 2026挑战赛中获第三名,私有排行榜宏F1为0.43902。

详情
AI中文摘要

本文描述了DS@GT ARC在植被样方图像多物种植物识别的PlantCLEF 2026挑战赛中获得第三名的解决方案。系统要在仅使用单标签植物图像训练的情况下,预测高分辨率样方照片中的所有物种。该管道围绕微调后的DINOv2 ViT-L/14分类器构建,对每个样方进行多尺度切片分解,切片预测与FAISS kNN检索器融合,并通过源感知时间融合、栖息地适应降级和地理掩码后处理。消融实验表明栖息地适应降级和多尺度聚合贡献最大。一些训练方向未取得成果,推理时的实例感知分割裁剪也未提升性能。所选提交在私有排行榜上的宏F1为0.43902(第三名;公开为0.51096)。

英文摘要

This paper describes DS@GT ARC's third-place solution to the PlantCLEF 2026 challenge on multi-species plant identification in vegetation quadrat images, where systems must predict every species present in high-resolution (~3000 x 3000 pixel) plot photographs while training only on single-label images of individual plants. The pipeline is built around a fine-tuned DINOv2 ViT-L/14 classifier applied over a multi-scale tile decomposition of each quadrat, with per-tile predictions blended with a FAISS kNN retriever and post-processed by source-aware temporal fusion across repeated plot visits, a habitat-fit demotion that injects geographic and altitude priors from the training data, and a South-Western Europe geographic mask. Habitat-fit demotion and multi-scale aggregation are the largest individual contributors in the ablations. Two complementary training-centric directions, a cross-region transformer with noisy-student distillation on the LUCAS dataset and a label-as-query transformer decoder over synthetic CLS-domain pseudo-quadrats, yielded null results. An inference-time augmentation with instance-aware segmentation crops also did not improve performance. The selected submission reaches a private-leaderboard macro-F1 of 0.43902 (third place; public 0.51096); an unselected configuration of the same pipeline scored above 0.45 on the private set. Code: https://github.com/dsgt-arc/plantclef-2026.

URL PDF HTML 收藏
2607.14499 2026-07-17 cs.AI 新提交

Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions

通过动态多轮交互对视觉语言模型进行情境化评估

Yijiang Li, Huiqi Zou, Bingyang Wang, Ziang Xiao

机构 * UC San Diego(加州大学圣地亚哥分校) Northeastern University(东北大学) Georgia Institute of Technology(佐治亚理工学院) Johns Hopkins University(约翰·霍普金斯大学)

AI总结 研究多模态大语言模型现实有效性问题,提出CEDI框架,通过三方交互、多轮半结构化对话及多种策略评估,应用于视觉幻觉,发现能揭示更多接近实际情况的幻觉,凸显其对MLLMs能力评估的作用。

详情
AI中文摘要

多模态大语言模型(MLLMs)在基准测试中取得了显著进展,但其在现实世界中的有效性仍不确定。这种差距源于受控静态环境中的基准测试与现实世界应用的动态、交互和情境化性质之间的根本错位。为弥合这一差距,我们提出了CEDI(通过动态多轮交互对MLLMs进行情境化评估)框架,将评估重新构建为被评估模型、自动考官和评分者之间的三方交互。考官通过基于任务的图形表示进行多轮半结构化对话。通过导航状态空间转换,CEDI部署从澄清请求到对抗性探测等各种策略,以获取性能证据。我们将CEDI应用于视觉幻觉。多个模型、不同设置、数据集和领域的实证结果表明,情境化、交互式评估不仅比传统静态评估揭示出更多幻觉,而且揭示出的幻觉更接近实际用例中出现的幻觉。我们还表明,幻觉往往会通过自我强化的对话历史在长语境中累积,并且模型特别容易受到需要拒绝前提或拒绝的问题的影响。这些发现共同凸显了CEDI是朝着对MLLMs能力进行现实、系统和生态有效评估迈出的一步。

英文摘要

Multi-modal Large Language Models (MLLMs) have made substantial advances on benchmarks, yet their real-world effectiveness remains uncertain. This gap stems from the fundamental misalignment between benchmarks in controlled, static settings and the dynamic, interactive, and contextualized nature of real-world applications. To bridge this gap, we propose CEDI (Contextualized Evaluations of MLLMs through Dynamic, multi-round Interactions), a framework that recasts evaluation as a three-party interaction between an evaluatee model, an automated examiner, and a grader. The examiner conducts multi-turn, semi-structured conversation guided by a graph-based representation of the task. By navigating state-space transitions, CEDI deploys diverse strategies, from clarification requests to adversarial probes, to elicit performance evidence. We apply CEDI to visual hallucinations. Empirical results across multiple models, diverse settings, datasets, and domains show that contextualized, interactive evaluations reveal not only significantly more hallucinations than conventional static evaluation but also ones that more closely resemble those arising in practical use cases. We further show that hallucinations often accumulate over long contexts, through self-reinforcing dialogue history, and models are particularly vulnerable to questions requiring premise rejection or refusal. Together, these findings highlight CEDI as a step toward realistic, systematic, and ecologically valid assessments of MLLMs' capabilities. Code is available at github.com/williamium3000/cedi.

URL PDF HTML 收藏
2607.14474 2026-07-17 cs.SD cs.AI cs.LG 新提交

Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026

令牌能竞争吗?针对BirdCLEF+ 2026的与监督式CNN主干竞争的令牌表示

Anthony Miyaguchi, Murilo Gustineli, Adrian Cheung

机构 * Georgia Institute of Technology(佐治亚理工学院)

AI总结 针对BirdCLEF+ 2026中动物发声多标签检测任务,先构建监督基线,后对比神经音频编解码器的编解码表示与基础嵌入的语义表示,比较两个生物声学专家模型和四个基于令牌的编码器,探究基于令牌的表示能否竞争。

详情
AI中文摘要

本文详细介绍了DS@GT ARC团队针对BirdCLEF+ 2026的方法,即对潘塔纳尔湿地声景中的动物发声进行多标签检测。2026年版本增加了约一小时的标记声景,使任务转向适合标记集的监督管道。首先构建了一个有竞争力的监督基线,在90分钟CPU预算内,在排名1894时达到了0.936的私有排行榜分数。其次,对比神经音频编解码器的编解码表示与基础嵌入的语义表示,询问基于令牌的表示是否能竞争。比较了两个生物声学专家模型和四个在AudioSet上训练的基于令牌的编码器。

英文摘要

This paper details the DS@GT ARC team's approach to BirdCLEF+ 2026, multi-label detection of animal vocalizations in soundscapes from the Pantanal wetlands. The 2026 edition adds about an hour of labeled soundscapes, shifting the task toward supervised pipelines fit to the labeled set. First, we build a competitive supervised baseline that ensembles a frozen Perch v2 backbone, a trained HGNetV2-B0 sound-event-detection network, and a non-bird prototypical head, reaching a private leaderboard score of 0.936 at rank 1894 within a 90-minute CPU budget. Second, we ask whether token-based representations can compete, contrasting codec representations from neural audio codecs against semantic representations from foundational embeddings. We compare two bioacoustic specialist models against four token-based encoders trained on AudioSet. The repository for this work can be found at https://github.com/dsgt-arc/birdclef-2026.

URL PDF HTML 收藏
2607.14400 2026-07-17 cs.CL cs.IR 新提交

DS@GT ARC at LongEval: Citation Integrity and Factual Grounding in Scientific QA

DS@GT ARC参加LongEval:科学问答中的引用完整性和事实基础

Brandon Michaels, Brendon Johnson

机构 * Georgia Institute of Technology(佐治亚理工学院)

AI总结 研究在科学问答中传统评估指标与引用完整性的差异,通过Corrective RAG和CiteFix构建纠正管道,对比前沿模型,发现前沿模型答案生成不依赖文档上下文,而纠正管道提升了引用忠实度和答案基础,提出需奖励严格答案基础的评估指标。

Comments 12 pages, 4 figures. Accepted to the CLEF 2026 LongEval Lab Working Notes

详情
AI中文摘要

本文描述了DS@GT ARC提交给2026年CLEF LongEval任务4关于检索增强生成(RAG)的内容。我们研究了传统自然语言评估指标与应用于RAG问答系统的引用完整性之间的差异。使用Corrective RAG(CRAG)和CiteFix评估一个纠正管道,对比基线和前沿模型基准RAG问答分数。前沿模型最大化了答案相关性和流畅性分数,但我们的诊断表明前沿模型在生成答案时未使用文档上下文就能正确识别相关文档。而我们的纠正管道通过预生成过滤块和生成后对引用材料严格强制蕴含,略微提高了引用忠实度和答案基础。我们提出,对可信RAG问答的评估需要奖励严格答案基础的指标。

英文摘要

This paper describes DS@GT ARC's submission to the CLEF 2026 LongEval Task 4 on Retrieval-Augmented Generation (RAG). In this submission, we examine a divergence between traditional natural language evaluation metrics and citation integrity as applied to RAG QA systems. We evaluate a corrective pipeline using Corrective RAG (CRAG) and CiteFix against baseline and frontier model benchmark RAG QA scores. While frontier models maximized answer relevance and fluency scores, our RAGAs LLM-as-judge diagnostics indicate that frontier models would correctly identify relevant documents without using their context in answer generation. Conversely, by filtering chunks pre-generation and enforcing strict entailment of generated claims to the cited material post-generation, our corrective pipeline marginally improved citation faithfulness and answer grounding. We propose that evaluation of trustworthy RAG QA requires metrics that reward strict answer grounding.

URL PDF HTML 收藏
2607.14107 2026-07-17 cs.CL cs.AI 新提交

Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs

北极星:用于扩散语言模型高效推理的漂移感知缓存校准和令牌承诺

Mingyu Lee, Akshat Ramachandran, Souvik Kundu, Tushar Krishna

机构 * Georgia Institute of Technology(佐治亚理工学院) Intel AI Group(英特尔人工智能集团)

AI总结 研究针对扩散语言模型推理效率受双向注意力和静态阈值影响的问题,提出北极星框架,利用令牌表示漂移信号,由北极星缓存和北极星提交两组件构成,大幅提升了模型在准确性-吞吐量方面的表现及解码并行性。

详情
AI中文摘要

扩散大语言模型(dLLMs)的推理效率受到两个挑战的限制:双向注意力妨碍了高效的KV缓存重用,而使用静态置信阈值增加解码并行性可能会损害生成质量。我们发现这两个挑战都源于一个共同现象:随着令牌被解码,通过双向注意力的上下文整合会导致令牌表示在解码步骤中漂移(演变)。基于此,我们提出了北极星,一个无需训练的推理框架,它使用令牌表示漂移作为统一信号来共同应对这两个挑战。北极星由两个组件组成:北极星缓存,通过漂移识别过时的KV缓存位置并执行稀疏的KV缓存刷新以实现高效重用;北极星提交,检测急剧漂移事件以可靠地识别准备提交的令牌。在几个dLLM系列的数学和编码基准测试中,北极星在准确性-吞吐量帕累托前沿上创造了新的技术水平,与现有基线相比,准确性提高了10.73%,吞吐量提高了3.7倍,并且在前向传递中实现了3.67个令牌的高解码并行性。

英文摘要

The inference efficiency of diffusion large language models (dLLMs) is constrained by two challenges: bidirectional attention precludes efficient KV-cache reuse, while increasing decoding parallelism with static confidence thresholds can compromise generation quality. We observe that both challenges arise from a shared phenomenon: as tokens are decoded, their contextual integration through bidirectional attention causes token representations to drift (evolve) across decoding steps. This insight motivates Polestar, a training-free inference framework that uses token representation drift as a unified signal to jointly address both challenges. Polestar comprises two components: Polestar-Cache, which identifies stale KV-cache positions via drift and performs sparse KV-cache refreshes to enable efficient reuse, and Polestar-Commit, which detects sharp drift events to reliably identify commit-ready tokens. Across mathematics and coding benchmarks on several dLLM families, Polestar sets a new state of the art on the accuracy-throughput Pareto frontier, achieving up to 10.73% accuracy improvement, up to 3.7x higher throughput, and high decoding parallelism of 3.67 tokens per forward pass over existing baselines.

URL PDF HTML 收藏
2607.14943 2026-07-17 cs.RO cs.AI cs.LG cs.SY eess.SY math.OC 新提交

Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control

通过机制可解释性和最优控制将鲁棒性引入世界行动模型

Jihoon Hong, Julian Skifstad, Qiyue Dai, Alice Chan, Glen Chou

机构 * Georgia Institute of Technology(佐治亚理工学院)

AI总结 研究世界行动模型在分布转移下的脆弱性,利用机制可解释性通过对比激活方向和基于模型的最优控制实现WAM转向,产生WA-LQR,预测不同模型转向能力,在部分模型上提高了对多种扰动的鲁棒性。

详情
AI中文摘要

世界行动模型(WAMs)能实现语义和物理信息控制,但在分布转移下很脆弱。本文利用机制可解释性研究WAM激活空间中与鲁棒性相关的扰动如何表示。通过比较成功和失败展开的激活情况,发现一些WAM架构对关键鲁棒性特征具有低维线性可分性,从而推动使用对比激活方向进行无训练的WAM转向。还表明WAM激活动力学中的局部线性性可通过基于模型的最优控制实现高效反馈转向,产生世界行动线性二次调节器(WA-LQR)。通过机制评估,预测了不同模型的转向能力,在Cosmos-Policy和DiT4DiT上,WA-LQR将对比方向推广到新任务并提高了对多种扰动的鲁棒性。

英文摘要

World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we use mechanistic interpretability to study how robustness-relevant perturbations are represented in WAM activation space. Comparing activations across successful and unsuccessful rollouts, we find some WAM architectures exhibit low-dimensional linear separability for robustness-critical features, while others do not. This motivates the use of contrastive activation directions for training-free WAM steering. We also show that local linearity in WAM activation dynamics enables efficient feedback steering via model-based optimal control, yielding World-Action Linear Quadratic Regulator (WA-LQR), a minimally-invasive reduced-order LQR controller. Via mechanistic evaluations, we predict strong steerability in the Cosmos-Policy and DiT4DiT models but weak steerability in LingBot-VA, consistent with steering intervention results. On Cosmos-Policy and DiT4DiT, WA-LQR generalizes contrastive directions to new tasks and improves robustness to camera, gripper, and visual-noise perturbations over unsteered and prompt steering baselines.

URL PDF HTML 收藏
2412.20556 2026-07-17 stat.ML cs.LG math.OC 版本更新

Distributionally Robust Optimization via Iterative Algorithms in Continuous Probability Spaces

连续概率空间中基于迭代算法的分布鲁棒优化

Linglingzhi Zhu, Yunqin Zhu, Yao Xie

机构 * H. Milton Stewart School of Industrial and Systems Engineering(H. Milton Stewart工业与系统工程学院) Georgia Institute of Technology(佐治亚理工学院)

AI总结 研究连续概率空间中分布鲁棒优化的计算挑战,利用布雷尼尔定理提出迭代算法框架,建立收敛保证和复杂度界,基于神经网络传输映射的数值结果显示该方法可实现鲁棒分类器稳定训练及有效最坏情况推理。

详情
AI中文摘要

我们研究了在最坏情况分布为连续时用于鲁棒推理的分布鲁棒优化(DRO),由于优化问题的无限维性质,这带来了重大计算挑战。与传统离散DRO方法不同,我们的框架利用布雷尼尔定理将最不利分布表征为连续参考测度的传输映射的推送。我们提出了具有多种变体的迭代算法框架,并在温和假设下建立了全局收敛保证,得出了关于次梯度评估和不精确约旦 - 金德勒勒尔 - 奥托更新的复杂度界。基于神经网络传输映射的数值结果表明,该方法能实现鲁棒分类器的稳定训练和分类任务的有效最坏情况推理。

英文摘要

We study distributionally robust optimization (DRO) for robust inference when the worst-case distribution is continuous, leading to significant computational challenges due to the infinite-dimensional nature of the optimization problem. Unlike traditional discrete DRO approaches, which often suffer from scalability issues, limited generalization, and costly worst-case inference, our framework exploits Brenier's theorem to characterize the least favorable distribution as the pushforward of a transport map from a continuous reference measure. This characterization motivates our study of the minimax problem in Wasserstein space. We propose an iterative algorithmic framework with multiple variants and establish global convergence guarantees under mild assumptions, deriving complexity bounds in terms of subgradient evaluations and inexact Jordan-Kinderlehrer-Otto updates. Numerical results with neural network-based transport maps demonstrate that the proposed method enables both stable training of robust classifiers and effective worst-case inference for classification tasks.

URL PDF HTML 收藏
2607.12963 2026-07-16 cs.CL 版本更新

The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context

稳健性的错觉:聚合准确率掩盖了与任务无关的上下文下的预测翻转

Yanzhe Zhang, Sanmi Koyejo, Diyi Yang

机构 * Georgia Tech(佐治亚理工学院) Stanford University(斯坦福大学)

AI总结 研究大语言模型在含无关上下文环境中的表现,发现聚合准确率掩盖了单个示例预测的不稳定性,如随机伪词会改变部分预测,且此不稳定性受多种因素调节,揭示了尾部风险,推动对模型进行单个示例可靠性评估。

Comments Preprint

详情
AI中文摘要

随着大语言模型能力增强,它们越来越多地部署在上下文丰富的环境中,任务输入常伴有冗长且部分无关的上下文。在可控环境中,我们发现最先进的模型在聚合层面通常对与任务无关的上下文表现出稳健性:在基准问题前添加该上下文对整体准确率影响不大。然而,这种聚合稳定性掩盖了单个示例的显著不稳定性。即使是随机组合字符形成的语义无意义的伪词,也能在一小部分示例上显著改变模型预测,在一些示例上降低性能,在另一些上提高性能。这种双面效应在广泛的模型和数据集上持续存在,且受影响的示例很大程度上因模型而异。我们进一步表明,这种不稳定性受上下文类型、长度、测试时计算和模型开发阶段的调节。我们的发现揭示了聚合准确率掩盖下的上下文诱导的尾部风险,促使对语言模型进行单个示例的可靠性评估。

英文摘要

As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by long, partially irrelevant context. In a controlled setting, we find that state-of-the-art models often appear robust to task-irrelevant context at the aggregate level: prepending it to benchmark questions causes little change in overall accuracy. This aggregate stability, however, masks significant per-example instability. Even semantically meaningless pseudo-words, formed by randomly combining characters, can markedly shift model predictions on a small fraction of examples, degrading performance on some while improving it on others. This two-sided effect holds consistently across a wide range of models and datasets, yet the affected examples are largely model-specific. We further show that this instability is modulated by context type, context length, test-time compute, and model development stage. Together, our findings reveal context-induced tail risks concealed by aggregate accuracy, motivating per-example reliability evaluation of language models.

URL PDF HTML 收藏
2607.10936 2026-07-16 cs.LG stat.ML 版本更新

Bandit PCA with Minimax Optimal Regret

具有极小极大最优遗憾值的强盗主成分分析

Moïse Blanchard, Dmitrii Ostrovskii, Aadirupa Saha

机构 * Georgia Tech(佐治亚理工学院) UIC(伊利诺伊大学芝加哥分校)

AI总结 研究强盗反馈版本的在线主成分分析,改进了遗憾值的上下界,弥合差距至\(r\sqrt{dT}\)。上界通过新算法实现,结合在线镜像下降与多尺度探索;下界通过构造自适应对手,将遗憾值下界估计转化为子空间估计问题,并讨论其与量子断层扫描联系。

详情
AI中文摘要

我们研究在线主成分分析的强盗反馈版本(强盗主成分分析):在每一轮\(t = 1,\dots,T\)中,对手选择一个\(d \times d\)对称增益矩阵\(G_t\),其谱在\([0,1]\)内且秩至多为\(r\);学习者同时选择一个单位向量\(w_t \in S^{d - 1}\)并接收奖励\(w_t^\top G_t w_t\)。学习者没有其他反馈,旨在最小化与事后最佳单位向量相比的遗憾值。这个问题由Kotlowski和Neu(2019)提出,他们给出了一个遗憾值为\(O(d\sqrt{rT \log T})\)的算法,并证明了下界为\(\Omega(r\sqrt{T/\log T})\)。我们改进了这两个界并基本弥合了它们之间的差距,在\(d\)和\(T\)的多对数因子范围内建立了阶为\(r\sqrt{dT}\)的极小极大遗憾值。上界由一种新颖算法实现,该算法将(实)密度矩阵谱面体上的在线镜像下降与多尺度探索方案相结合,其中具有不同谱大小的特征子空间以不同速率更新。对于下界,我们构造了一个自适应对手,它根据学习者的行动细化一个隐藏的大奖励子空间,使得在不估计子空间的情况下不可能有低遗憾值;因此,对遗憾值进行下界估计归结为研究出现的子空间估计问题。最后,我们讨论了强盗主成分分析与自适应测量量子断层扫描的联系。

英文摘要

We study the bandit-feedback version of online principal component analysis (Bandit PCA): in each round $t = 1,\dots,T$, the adversary selects a $d \times d$ symmetric gain matrix $G_t$ with spectrum in $[0,1]$ and rank at most $r$; the learner simultaneously selects a unit vector $w_t \in S^{d-1}$ and receives the reward $w_t^\top G_t w_t$. The learner receives no other feedback, and aims to minimize the regret against the best unit vector in hindsight. This problem was introduced by Kotlowski and Neu (2019), who gave an algorithm with regret $O(d\sqrt{rT \log T})$ and showed the lower bound of $Ω(r\sqrt{T/\log T})$. We improve upon both of these bounds and essentially bridge the gap between them, establishing the minimax regret of order $r\sqrt{dT}$ up to polylogarithmic factors in $d$ and $T$. The upper bound is attained by a novel algorithm, which combines online mirror descent on the spectrahedron of (real) density matrices with a multiscale exploration scheme in which the eigenspaces with different spectral magnitudes are updated at different rates. For the lower bound, we construct an adaptive adversary that refines a hidden large-reward subspace based on the learner's actions, in such a way that low regret is impossible without estimating the subspace; as a result, lower-bounding the regret reduces to studying the arising subspace estimation problem. Finally, we discuss connections of Bandit PCA with adaptive-measurement quantum tomography.

URL PDF HTML 收藏
2607.12370 2026-07-15 cs.RO 新提交

StratMamba: Strategic and Reactive Stream Partitioning for Path-Efficient LiDAR-Based Obstacle Avoidance

StratMamba:用于基于路径高效的激光雷达避障的策略性和反应性流划分

Hung-Chieh Wu, Xiaopan Zhang, Kasra Sinaei, Ryan Abnavi, Kasun Weerakoon, Christopher Bradley, Seyed Fakoorian, Jiachen Li, Donald Ebeigbe

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) AlphaZ, Inc.(阿尔法兹公司) Georgia Institute of Technology(佐治亚理工学院) Massachusetts Institute of Technology(麻省理工学院)

AI总结 研究针对复杂环境中机器人导航问题,提出StratMamba双流时间建模架构,结合快慢衰减内存架构处理激光雷达数据。经多场景评估及与其他基线对比,其在时间推理效率、导航速度和路径最优性方面表现出色,在现实中性能更稳健。

Comments Accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026). 8 pages, 6 figures. Video: https://www.youtube.com/watch?v=Z0FfO_AVaSw

详情
AI中文摘要

本文提出了StratMamba,一种基于双流Mamba的时间建模架构,以更有效地捕捉复杂且障碍物多的环境中机器人导航所需的长期时间依赖性。StratMamba利用快速衰减和缓慢衰减内存架构的组合,快速衰减组件处理高频激光雷达数据以进行反应性避障,缓慢衰减组件维护长期目标信息用于策略规划。在IsaacLab和Gazebo中对不同避障场景进行了广泛评估,并在Unitree GO1四足机器人上验证了从模拟到现实的成功部署。与其他时间RL基线比较表明,StratMamba以更低的超时率实现了出色的时间推理效率,同时保持最快导航速度,还实现了最高路径最优性。现实世界评估显示,与普通Mamba和Transformer相比,StratMamba在扩展激光雷达范围内保持更稳健性能,证明双流划分在具有挑战性的传感条件下有效平衡了反应性安全与策略性导航。

英文摘要

This paper proposes StratMamba, a dual-stream Mamba-based temporal modeling architecture, to more efficiently capture long-horizon temporal dependencies required for robot navigation in complex and obstacle-rich environments. StratMamba leverages a combination of fast-decay and slow-decay memory architectures, where the fast-decay component processes high-frequency LiDAR data for reactive obstacle avoidance, while the slow-decay component maintains longer-horizon goal information for strategic planning. We perform extensive evaluations of different obstacle avoidance scenarios in IsaacLab and Gazebo, while also validating successful sim-to-real deployment on a Unitree GO1 quadruped robot navigating in the presence of static/dynamic obstacles. Comparisons with other temporal RL baselines, such as LSTM, Transformer, and Vanilla-Mamba, show that our StratMamba achieves exceptional temporal reasoning efficiency with a lower timeout rate, while maintaining the fastest navigation speed (576 median steps, 5.0% better than Vanilla-Mamba). It also achieves the highest path optimality (0.915 path efficiency) across all baselines. Real-world evaluation reveals that StratMamba maintains more robust performance across extended LiDAR ranges compared to vanilla Mamba and the Transformer, demonstrating that dual-stream partitioning effectively balances reactive safety with strategic navigation under challenging sensing conditions.

URL PDF HTML 收藏
2607.12334 2026-07-15 cs.CL cs.ET 新提交

QUBO-Optimized Evidence Selection for Retrieval-Augmented Question Answering with Unconventional Solvers

使用非常规求解器的QUBO优化证据选择用于检索增强问答

Rahul Singh, Madhav Vadlamani

机构 * University of California Santa Barbara(加利福尼亚大学圣巴巴拉分校) Georgia Institute of Technology(佐治亚理工学院)

AI总结 研究针对检索增强问答中证据选择问题,将其转化为QUBO问题构建能量函数,平衡多种因素选择证据段落,再由下游语言模型生成答案。在HotpotQA上评估,该方法性能与基于LLM的选择器相当,为RAG管道提供新思路。

详情
AI中文摘要

检索增强问答依赖于选择共同支持答案生成的证据段落。许多RAG管道依赖于top-k排名,尽管多跳问题通常需要满足多个信息需求的互补证据。基于LLM的选择器将检索视为集合选择来解决此问题,但在中间阶段使用LLM成本高且难以扩展。本文将证据选择公式化为二次无约束二元优化(QUBO)问题,构建能量函数平衡多种因素。所选段落再传递给下游语言模型生成答案。在HotpotQA上评估了QUBO选择器,结果表明多跳证据选择可转化为离散优化,为RAG管道开辟了一条路径。

英文摘要

Retrieval-augmented question answering depends on selecting evidence passages that jointly support answer generation. However, many RAG pipelines rely on top-\(k\) ranking, where passages are selected mainly by individual relevance scores, even though multi-hop questions often require complementary evidence satisfying multiple information requirements. Recent LLM-based selectors address this by treating retrieval as set selection, but using an LLM for this intermediate stage can be costly and difficult to scale. In this work, we formulate evidence selection as a Quadratic Unconstrained Binary Optimization (QUBO) problem. Given a question, candidate passages, and decomposed information requirements, our method constructs an energy function that balances relevance, requirement coverage, support strength, redundancy, complementarity, and compactness. Low-energy solutions correspond to compact evidence subsets that cover the needed requirements while avoiding unnecessary or repetitive context. The selected passages are then passed to a downstream language model for answer generation, separating combinatorial evidence selection from semantic answer generation. We evaluate the proposed QUBO selector on HotpotQA and compare it with LLM-based set selectors and non-LLM baselines including BM25, relevance top-\(k\), maximal marginal relevance, hybrid lexical--semantic ranking, greedy coverage, and random selection. The QUBO selector achieves competitive exact-match and token-F1 performance relative to LLM-based selectors while providing a solver-compatible formulation for structured evidence selection. These results suggest that multi-hop evidence selection can be cast as discrete optimization, opening a path toward RAG pipelines where LLMs are reserved for semantic processing and answer generation, while context selection is handled by Ising/QUBO-compatible solvers.

URL PDF HTML 收藏
2607.12304 2026-07-15 cs.CV cs.LG 新提交

What Does a Temporal Benchmark Score Measure? Decomposing Channel Use in Video VLM Evaluation

时间基准分数衡量的是什么?分解视频视觉语言模型评估中的通道使用情况

Farrukh Rahman

机构 * Microsoft(微软) Georgia Institute of Technology(佐治亚理工学院)

AI总结 研究视频视觉语言模型评估中时间基准分数衡量问题,提出无标签筛选方法“反转下降”区分模型通道使用情况,如Molmo2和Qwen3-VL,指出综合分数不能反映潜在失败模式,此区分在多基准和任务中成立。

Comments 9 pages, 11 pages supplemental

详情
AI中文摘要

时间视频问答基准上的分数旨在衡量模型是否具有时间理解能力,但它混淆了两个问题。一是任务问题,即问题是否具有时间性,是否需要多个帧及其顺序;二是通道问题,当需要时,模型是从像素中恢复顺序,还是从位置编码(RoPE)中读取顺序。大多数时间分数都无法回答这两个问题,单帧和答案先验往往就能得出分数。该领域的有效性检查、帧打乱敏感性以及从完整视频中获得的准确率,仅涉及任务问题。我们为通道问题贡献了一个无标签筛选方法——反转下降:当视觉序列反转而RoPE保持正向时损失的准确率。它可应用于兼容的时间基准,无需新的注释。配对的反向标签或标签在反转下确定性变换的任务,可区分遵循反向内容的模型和仅因冲突而被打乱的模型。Molmo2从位置读取正向事件的顺序,而Qwen3-VL读取其实际看到的反向事件的视觉顺序。我们称它们为位置主导型和视觉序列主导型。这种区分在两个基准和两个尺度的几个时间任务中都成立,激活修补表明这是一种真实的内部属性,而非冲突的假象。这种区分很重要,因为两个通道在相反的输入上会失败,所以两个分数相似的模型不可互换,即综合分数不能反映潜在的失败模式。

英文摘要

A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions. 1. The task question: is the question even temporal, does it need several frames and their order? and 2. The channel question, when it does, does the model recover the order from the pixels, or read it off the positional encoding (RoPE)? Most of a temporal score answers neither, a single frame and answer priors often carry it. The field's validity checks, frame-shuffle sensitivity and the accuracy gained from the full video, speak only to the task question. We contribute a label-free screen for the channel question, the reversal-drop: the accuracy lost when the visual sequence is reversed while RoPE remains forward. It can be applied to compatible temporal benchmarks without new annotations. Paired reverse labels, or tasks whose labels transform deterministically under reversal, distinguish models that follow reversed content from those merely disrupted by the conflict. Molmo2 answers the forward event reading order off positions, while Qwen3-VL answers the reversed event it actually sees, reading visual order (comparatively). We call them position-dominant and visual-sequence-dominant. The split holds across two benchmarks and several temporal tasks at two scales, and activation patching shows it is a real internal property, not an artifact of the conflict. The distinction matters, the two channels fail on opposite inputs so two models with similar score are not interchangable, i.e. an aggregate score does not reflect potential failure modes.

URL PDF HTML 收藏
2602.15892 2026-07-15 cs.CV cs.AI 版本更新

Egocentric Bias in Vision-Language Models

视觉语言模型中的自我中心偏差

Maijunxian Wang, Yijiang Li, Bingyang Wang, Tianwei Zhao, Ran Ji, Qingying Gao, Emmy Liu, Hokin Deng, Dezhi Luo

机构 * Cognitive Science Program, University of California, Berkeley(加州大学伯克利分校认知科学项目) Department of Electrical and Computer Engineering, University of California San Diego(加州大学圣地亚哥分校电气与计算机工程系) School of Computer Science, Georgia Institute of Technology & Emory University(佐治亚理工学院计算机科学学院及埃默里大学) Department of Computer Science, Johns Hopkins University(约翰霍普金斯大学计算机科学系) Department of Cognitive Science, University of California San Diego(加州大学圣地亚哥分校认知科学系) Equal Advising Department of Computer Science & Wilmer Eye Institute, Johns Hopkins University(约翰霍普金斯大学计算机科学系及威尔默眼科研究所) Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所) Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所) Weinberg Institute for Cognitive Science, University of Michigan(密歇根大学韦恩伯格认知科学研究所)

AI总结 本文提出FlipSet基准,揭示视觉语言模型在第二级视觉视角推理中存在系统性自我中心偏差,表明模型在整合社会认知与空间操作方面存在根本性不足。

Comments Accepted at CogSci 2026 (Best Undergraduate Student Paper)

详情
AI中文摘要

视觉视角的推理——推断从他人视角世界如何呈现——是社会认知的基础。我们引入FlipSet,一个用于评估视觉语言模型中第二级视觉视角推理(L2 VPT)的诊断基准。该任务要求模拟从另一个代理视角对二维字符字符串进行180度旋转,隔离空间变换与三维场景复杂性。评估103个VLMs发现系统性的自我中心偏差:绝大多数表现低于随机水平,其中约四分之三的错误重现了摄像机视角。控制实验揭示了组合缺陷——模型在孤立情况下实现高理论-心准确率和高于随机水平的内部旋转,但当需要整合时却彻底失败。这种分离表明当前VLMs缺乏将社会意识与空间操作结合的机制,暗示了基于模型的空间推理的根本限制。FlipSet为多模态系统诊断视角推理能力提供了认知基础的测试平台。

英文摘要

Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce FlipSet, a diagnostic benchmark for Level-2 visual perspective taking (L2 VPT) in vision-language models. The task requires simulating 180-degree rotations of 2D character strings from another agent's perspective, isolating spatial transformation from 3D scene complexity. Evaluating 103 VLMs reveals systematic egocentric bias: the vast majority perform below chance, with roughly three-quarters of errors reproducing the camera viewpoint. Control experiments expose a compositional deficit--models achieve high theory-of-mind accuracy and above-chance mental rotation in isolation, yet fail catastrophically when integration is required. This dissociation indicates that current VLMs lack the mechanisms needed to bind social awareness to spatial operations, suggesting fundamental limitations in model-based spatial reasoning. FlipSet provides a cognitively grounded testbed for diagnosing perspective-taking capabilities in multimodal systems.

URL PDF HTML 收藏
2607.11855 2026-07-14 cs.RO 新提交

Robust bipedal locomotion on flowable slopes via foot-driven terrain manipulation

通过足部驱动的地形操纵在可流动斜坡上实现稳健的双足运动

Deniz Kerimoglu, Junnosuke Kamohara, Jiyeon Maeng, Ziwon Yoon, Seth Hutchinson, Ye Zhao, Daniel I. Goldman

机构 * Georgia Institute of Technology(佐治亚理工学院) Northeastern University(东北大学)

AI总结 研究双足机器人在颗粒斜坡上的运动控制问题,通过研究带防滑钉足部的地面动力学,发现中等防滑钉间距利于行走,据此设计可调整防滑钉深度的足部,应用于大小不同的双足机器人,提出以肢体为中心调节地形相互作用的新控制方法。

Comments 38 pages, 12 figures

详情
AI中文摘要

双足机器人控制具有挑战性,因其接近不稳定状态,足部与地形接触的微小变化会迅速破坏运动稳定性。在刚性地形上,可通过成熟的接触力学和控制策略缓解这种脆弱性。而在可流动表面如颗粒斜坡上,足部接触会引发大的表面变形和类似固液转变,耦合地形效应与机器人动力学,导致性能不佳或失败,部分原因是缺乏可靠的可流动地形动力学表示方法。本文通过研究带防滑钉的足部(从鞋底伸出的薄板)的地面动力学,探讨控制地形响应如何改善颗粒斜坡上的双足运动。对小型(1.4千克)机器人物理双足的系统研究表明,防滑钉间距稀疏和密集分别会导致过度的地形屈服和阻力,降低性能并导致失败。中等防滑钉间距可分布相互作用力,使基底应力维持在(或低于)屈服阈值,从而能在高达30度的颗粒斜坡上行走。基于这些原理,设计了一种能主动调整防滑钉深度并适应刚性和颗粒地形的足部。还证明了有效的足部与地形相互作用原理可应用于更大(15千克)的自主双足机器人。本研究提出了一种替代传统以身体为中心的机器人控制方法的方案,即通过以肢体为中心的方法调节地形相互作用,而非通过身体运动调节地形诱导效应。

英文摘要

Bipedal robots are challenging to control because they operate close to instability, where small variations in foot-terrain contact can rapidly destabilize locomotion. On rigid terrain, bipedal robots mitigate this fragility by using well-established contact mechanics and control strategies. On flowable surfaces such as granular slopes, foot contact can induce large surface deformations and solid-fluid-like transitions, coupling terrain effects with robot dynamics, leading to underperformance or failure. This is partly due to the lack of reliable methods to represent the dynamics of flowable terrain, making it difficult to account for terrain effects in locomotion design. Here, we investigate how controlling terrain response can improve bipedal locomotion on granular slopes by studying the terradynamics of cleated feet, thin plates emanating from the foot soles. Systematic studies of a small-scale (1.4 kg) robophysical biped reveal that cleats with sparse and dense spacing lead to excessive terrain yielding and resistance, respectively, degrading performance and leading to failure. An intermediate cleat spacing distributes interaction forces to maintain substrate stresses near (or below) the yield threshold, enabling walking on granular slopes up to 30 degrees. Guided by these principles, we design a foot that actively adjusts cleat depth and accommodates both rigid and granular terrain. We also demonstrate that the principles of effective foot-terrain interaction translate to a larger (15 kg) autonomous biped. Our study presents an alternative to conventional body-centric robot control approaches, which regulate terrain-induced effects through body motion, by instead regulating terrain interactions through limb-centric approach.

URL PDF HTML 收藏
2607.10892 2026-07-14 cs.RO 新提交

A Single Diffusion-Policy Controller for Multi-Task Block Pushing with Zero-Shot Sim-to-Real Transfer

用于多任务块推的单扩散策略控制器,具有零样本模拟到现实转移

Haitong Ma, Haldun Balim, Yang Hu, Bo Dai, Na Li

机构 * Harvard University(哈佛大学) Georgia Institute of Technology(佐治亚理工学院)

AI总结 研究旨在用强化学习从零训练单扩散策略用于多任务块推,提出含简单策略损失函数的框架,结合反向课程生成等应对探索挑战,评估其在不同条件下零样本从模拟到现实的转移能力,证明该流程有效。

Comments 8 pages, 7 figures

详情
AI中文摘要

扩散策略在通过行为克隆为机器人表示和学习复杂动作方面展现出了有前景的实证性能。本文中,我们探索使用强化学习从零开始训练扩散策略用于多任务机器人操纵。具体而言,我们旨在训练一个针对多种形状块推任务的单扩散策略。所提出的框架具有一个简单的策略损失函数,它是基于行为克隆的扩散策略训练中使用的重新加权证据下界,并且能无缝用作强化学习算法中的策略学习模块。为应对因缺乏示范而产生的探索挑战,我们纳入了反向课程生成和以目标为中心的表示。结合扩散策略的表现力,我们的设计支持在稀疏奖励模拟设置中学习多任务块推策略。我们进一步评估训练好的扩散策略在包括目标位置、块形状、块重量和表面摩擦等不同环境条件下能否零样本转移到现实世界任务,结果表明该流程在测试的变化情况下能转移到我们的现实世界块推设置中。

英文摘要

Diffusion policies have shown promising empirical performance in representing and learning complex maneuvers for robots using behavior cloning (BC). In this paper, we explore training diffusion policies from scratch using reinforcement learning (RL) for multi-task robotic manipulation. Specifically, we aim to train a single diffusion policy for block-pushing tasks with multiple shapes. The proposed framework features a simple policy loss function, which is a reweighted evidence lower bound used in BC-based diffusion policy training and can seamlessly serve as the policy learning module in RL algorithms. To address the exploration challenges arising from the absence of demonstrations, we incorporate reverse curriculum generation and objective-centric representations. Combined with the expressiveness of diffusion policies, our design supports learning of multi-task block-pushing policies in our sparse-reward simulation setting. We further evaluate whether the trained diffusion policy transfers in zero-shot to real-world tasks under varying environmental conditions including goal positions, block shapes, block weights and surface friction, providing evidence that this pipeline can transfer to our real-world block-pushing setup under the tested variations.

URL PDF HTML 收藏