arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Tsinghua University(清华大学)

至 收录 4534
2607.18231 2026-07-21 cs.RO 新提交

FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation

FM-VLA:用于接触丰富操作中视觉-语言-动作模型的基于力的记忆

Ruicheng Li, Qixiu Li, Ruichun Ma, Yu Deng, Lin Luo, Zhiying Du, Jianfeng Xiang, Huizhi Liang, Ruicheng Wang, Jiaolong Yang, Baining Guo

机构 * Tsinghua University(清华大学) Microsoft Research(微软研究院) Fudan University(复旦大学) USTC(中国科学技术大学)

AI总结 研究针对视觉-语言-动作模型在接触丰富操作中的时间上下文推理问题,提出FM-VLA模型,通过基于力的记忆编码及投影,利用累积接触事件历史指导操作,在相关任务上评估,轻量级力记忆表现出色,显著优于基线方法。

详情
AI中文摘要

视觉-语言-动作(VLA)模型在机器人操作中实现了令人印象深刻的泛化,近期基于记忆增强的VLA通过以过去图像或语言摘要为条件放宽了马尔可夫假设。基于视觉的记忆方法通过对采样的过去图像帧进行条件设定来解决此问题,但在时间事件视觉模糊时计算成本高且有根本限制。我们提出FM-VLA,一种具有基于力的记忆的VLA模型,用于非马尔可夫、接触丰富操作的时间上下文推理。我们用变分自编码器(VAE)将力历史编码为紧凑的力记忆令牌,VAE通过力时间序列重建进行预训练。通过将力潜在表示和短状态历史投影为动作专家模块的额外条件令牌,使VLA能够利用累积的接触事件历史来指导操作。我们在三个依赖记忆的任务上评估FM-VLA,包括找到隐藏块、按按钮和擦拭盘子特定次数。我们的轻量级力记忆以最小推理开销实现了超过80%的成功率,显著优于基线方法。

英文摘要

Vision-language-action (VLA) models have achieved impressive generalization in robotic manipulation, and recent memory-augmented VLAs have relaxed the Markovian assumption by conditioning on past images or language summaries. Vision-based memory approaches address this by conditioning on sampled past image frames, but they are computationally expensive and fundamentally limited when temporal events are visually ambiguous, e.g., pushing a button multiple times with small movements. We propose FM-VLA, a VLA model with force-based memory, enabling temporal context reasoning for non-Markovian, contact-rich manipulation. We encode force histories into compact force memory tokens with a variational autoencoder (VAE) pretrained with force time series reconstruction. By projecting force latent representations and short state history as additional conditioning tokens to the action expert module, we enable VLAs to leverage accumulated contact event history to guide manipulation. We evaluate FM-VLA on three memory-dependent tasks, including finding a hidden block, pressing a button, and wiping a dish for a specific number of times. Our lightweight force memory achieves over 80% success rate with minimal inference overhead, significantly outperforming baseline approaches. Project page: https://qft-333.github.io/FM-VLA-Page/

URL PDF HTML 收藏
2607.18110 2026-07-21 cs.LG cs.CL 新提交

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

大语言模型作为教练:不可验证任务的体验式学习

Tianzhu Ye, Li Dong, Guanheng Chen, He Zhu, Xun Wu, Shaohan Huang, Furu Wei

机构 * Microsoft Research(微软研究院) Tsinghua University(清华大学) Peking University(北京大学)

AI总结 研究不可验证任务,提出体验式学习(EL),将LLM反馈模型从评判转为教练,通过提炼体验知识提供密集监督,在开放式任务上表现优于基于规则的RL,泛化性好且减轻奖励作弊。

详情
AI中文摘要

在开放式任务上的强化学习(RL)将基于大语言模型(LLM)的基于规则的评估压缩为标量奖励,丢弃了丰富的文本反馈,并将具有不同质量配置文件的响应混为一谈。我们提出了体验式学习(EL),它将反馈模型从作为评判的LLM重新用作作为教练的LLM。教练将其对每个策略响应的评估提炼为可转移的体验知识,该知识为教师模型提供条件,并通过策略上下文提炼被策略内化。与标量奖励相比,这个更高带宽的反馈通道提供了密集监督,并保留了高质量响应之间的细粒度偏好。在两个策略家族中,有来自策略本身或专有模型的反馈,EL在留出的和未见的开放式任务上始终优于基于规则的RL。值得注意的是,EL在训练分布之外具有更好的泛化能力,并减轻了奖励作弊。这些发现将体验知识确立为用于不可验证任务的训练后更丰富、更可泛化的学习信号。

英文摘要

Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning (EL), which repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach. The coach distills its assessment of each on-policy response into transferable experiential knowledge, which conditions a teacher model and is internalized by the policy through on-policy context distillation. Compared with scalar rewards, this higher-bandwidth feedback channel provides dense supervision and preserves fine-grained preferences among high-quality responses. Across two policy families, with feedback from the policy itself or a proprietary model, EL consistently outperforms rubric-based RL on held-out and unseen open-ended tasks. Notably, EL generalizes better beyond the training distribution, and mitigates reward hacking. These findings establish experiential knowledge as a richer and more generalizable learning signal for post-training on non-verifiable tasks.

URL PDF HTML 收藏
2607.17967 2026-07-21 cs.CV 新提交

Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement

基于自引导稀疏体细化的精细细节单目几何估计

Lingyu Kong, Ruicheng Li, Ruicheng Wang, Sicheng Xu, Chengtang Yao, Jianfeng Xiang, Jiaolong Yang

机构 * Tsinghua University(清华大学) USTC(中国科学技术大学) Microsoft Research(微软研究院)

AI总结 针对单目几何估计在局部3D结构精细细节上的失真问题,提出基于自引导稀疏体细化的方法,将建模从2D提升到3D空间,通过稀疏卷积避免特征混合,实验证明该方法在恢复精细3D几何方面显著优于现有方法。

详情
AI中文摘要

单目几何估计在不同场景中取得了显著性能,但当前最先进模型在局部3D结构尤其是精细细节上仍有明显失真。我们将此局限归因于架构不匹配,多数模型在2D参数化内解码3D几何,导致特征混合。本文提出自引导稀疏体细化(SSR)的精细细节单目几何估计,将单目几何建模从2D图像空间提升到3D空间。模型将基础模型的粗点图提升到稀疏体素壳上并通过SSR细化,SSR采用基于3D空间局部性聚合特征的稀疏卷积。实验表明该方法在恢复精细3D几何上显著优于现有方法。

英文摘要

Monocular geometry estimation has recently achieved impressive performance across diverse scenes. However, state-of-the-art models still face notable distortion in local 3D structure, especially in fine details, like thin structures and small objects. We attribute this limitation to an architectural mismatch: most current models decode 3D geometry within a 2D parameterization, where feature interactions are governed by image-plane proximity rather than true 3D spatial relationships. This inadvertently mixes features from geometrically distant surfaces, resulting in over-smoothed geometry particularly around thin or elongated structure. In this paper, we propose a fine-detail monocular geometry estimation with Self-Guided Sparse 3D Refinement (SSR) that lifts monocular geometry modeling from 2D image space to 3D space for high-fidelity metric-scale point maps. Our model lifts the coarse point map from a foundation base model onto a sparse voxel shell and refines it via SSR. The SSR employs sparse convolutions that aggregate features based on 3D spatial locality, avoiding feature mixing across depth discontinuities. Extensive experiments on diverse datasets demonstrate that our method significantly outperforms existing approaches in recovering fine detailed 3D geometry across both quantitative metrics and qualitative visualizations.

URL PDF HTML 收藏
2607.17897 2026-07-21 cs.LG 新提交

Distributional Soft Bellman Operator under the Cramér Geometry

在克拉默几何下的分布软贝尔曼算子

Keru Wang, Yixin Deng, Yao Lyu, Stephen Redmond, Shengbo Eben Li

机构 * School of Electrical and Electronic Engineering, University College Dublin(都柏林大学学院电气与电子工程学院) School of Vehicle and Mobility, Tsinghua University(清华大学车辆与运载学院) College of Artificial Intelligence, Tsinghua University(清华大学人工智能学院)

AI总结 研究分布软策略迭代中策略评估步骤里的固定策略分布软贝尔曼算子,基于克拉默几何,制定CDF级算子并证明其收缩性质,获得唯一不动点,还通过共轭得到谱域表示,为DSPI式算法研究提供参考。

详情
AI中文摘要

分布软策略迭代(DSPI)为结合分布强化学习(DRL)与最大熵控制提供了重要框架,其策略评估步骤由作用于熵正则化回报的分布软贝尔曼算子主导。本文聚焦于基于累积分布函数(CDF)且具有\(L^2\)结构的克拉默几何,研究固定策略分布软贝尔曼算子在此度量下是否具有收缩性质及唯一不动点。直接在允许的CDF场域上,制定CDF级分布软贝尔曼算子,证明其为\(\sqrt{\gamma}\)收缩,并获得相应唯一不动点及收敛的迭代策略评估。CDF公式表明此有限克拉默域性质源于联合一步奖励熵转移的均匀一阶矩条件。通过共轭将评估问题转移到谱域,得到相同决策过程的等效希尔伯特空间表示。这些结果确定了与DSPI策略评估步骤相关的克拉默几何贝尔曼不动点,为研究DSPI式算法中的近似评论家、评估误差和评论家损失设计提供了参考点。

英文摘要

Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the policy evaluation step is governed by a distributional soft Bellman operator acting on entropy-regularised returns. Theoretical analysis of such an evaluation step requires a probability metric under which Bellman updates can be controlled, typically by showing that the operator contracts the distance between any two candidate return-distribution estimates. In this paper, we focus on the Cramér geometry, a cumulative distribution function (CDF)-based metric with an $L^2$ structure, and study whether the fixed-policy distributional soft Bellman operator has this contraction property and hence a unique fixed point under this metric. Working directly on an admissible CDF field domain, we formulate the CDF-level distributional soft Bellman operator, prove that it is a $\sqrtγ$-contraction, and obtain the corresponding unique fixed point together with convergent iterative policy evaluation. The CDF formulation also shows that this finite-Cramér-domain property follows from a uniform first-moment condition on the combined one-step reward entropy shift, rather than from separate uniform boundedness assumptions on the reward and entropy terms. We then transport the same evaluation problem to the spectral domain by conjugation, obtaining an equivalent Hilbert-space representation of the same decision process. Taken together, these results identify the Cramér-geometric Bellman fixed point associated with the policy-evaluation step of DSPI, providing a reference point for studying approximate critics, evaluation error, and critic-loss design in DSPI-style algorithms.

URL PDF HTML 收藏
2607.17745 2026-07-21 cs.AI 新提交

WuYu-EnvLE-Bench: A Benchmark for Evaluating Large Language Models in Environmental Law Enforcement

吴语-环境执法基准:评估大语言模型在环境执法中的基准

Ziliang Yang, Yi Zhang, Kaijun Lin, Jiachao Ke, Haihong Xu, Zongguo Wen

机构 * School of Environment, Tsinghua University(清华大学环境学院) College of Economics and Management, Beijing University of Technology(北京工业大学经济与管理学院) State Key Laboratory of Iron and Steel Industry Environmental Protection, School of Environment, Tsinghua University(清华大学环境学院钢铁工业环境保护国家重点实验室) Appraisal Center for Environmental Engineering, Ministry of Ecology and Environment(生态环境部环境工程评估中心)

AI总结 该研究针对大语言模型在环境执法中生成可追溯决策能力不明的问题,构建吴语-环境执法基准,含多任务多子领域实例,用AES和IEI评估模型,发现其在部分任务表现不佳,强调证据与规则感知的执法推理需求。

Comments 98 pages, 45 figures,

详情
AI中文摘要

大语言模型(LLMs)在环境执法中的应用日益受到关注,但其生成可追溯执法决策的能力尚不明晰。我们引入了吴语-环境执法基准(WuYu-EnvLE-Bench),它基于实际执法案例、监管标准和专家评审构建。该基准包含2521个基准实例、14项任务和12个跨执法前、执法中和执法后工作流程的污染介质子领域。我们使用绝对环境执法得分(AES)和智能执法指数(IEI)评估了开源和闭源大语言模型的能力、响应质量和资源效率。结果表明,大语言模型在规则受限任务上表现良好,但在证据链构建、矛盾检测、多源整合和程序判断方面仍不可靠。模型扩展也显示出收益递减:中型模型在结构化任务中接近领先模型,而大型模型无法可靠地克服证据推理瓶颈。吴语-环境执法基准强调了基于证据、规则感知和任务自适应执法推理的必要性。

英文摘要

Large language models (LLMs) are increasingly considered for environmental enforcement, but their ability to produce traceable enforcement decisions remains unclear. We introduce WuYu-EnvLE-Bench, a benchmark built from real enforcement cases, regulatory standards, and expert review. It contains 2,521 benchmark instances, 14 tasks, and 12 pollution-medium subdomains across pre-enforcement, in-enforcement, and post-enforcement workflows. Using Absolute Environmental Enforcement Score (AES) and Intelligent Enforcement Index (IEI), we evaluate open-source and closed-source LLMs across capability, response quality, and resource efficiency. Results show that LLMs perform well on rule-bounded tasks but remain unreliable in evidence-chain construction, contradiction detection, multi-source integration, and procedural judgment. Model scaling also shows diminishing returns: medium-sized models approach leading models in structured tasks, while larger models do not reliably overcome evidence-reasoning bottlenecks. WuYu-EnvLE-Bench highlights the need for evidence-grounded, rule-aware, and task-adaptive enforcement reasoning.

URL PDF HTML 收藏
2607.17733 2026-07-21 cs.LG cs.AI 新提交

MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference

MXSens:用于高效大语言模型推理的灵敏度感知混合精度量化

Simla Burcu Harma, Danila Mishin, Zhengyuan Su, Ayan Chakraborty, Elizaveta Kostenok, Dongho Ha, Babak Falsafi, Martin Jaggi, Yunho Oh, Amir Yazdanbakhsh

机构 * EPFL(洛桑联邦理工学院) Tsinghua University(清华大学) MangoBoost Inc.(芒果助推公司) Korea University(韩国大学) Google DeepMind(谷歌深度思维)

AI总结 研究针对大语言模型推理中4位量化因异常值致精度降的问题,提出MXSens方法,基于列和层灵敏度分配混合尾数比特宽度,无需训练,利用MXINT块结构,在多模型任务中优于现有方法,平衡了量化的准确性与资源效率。

详情
AI中文摘要

4位量化可实现高效的大语言模型推理,但由于异常值会导致显著的精度下降。先前的工作通过数据旋转或混合精度整数量化来解决此问题,但通常依赖软件管理的缩放和频繁的反量化,带来大量开销。微缩放格式(如MXINT)通过在硬件中编码缩放来消除这些低效率,但仍与基于旋转的方法不兼容。我们的分析表明,异常值的严重程度各不相同,量化灵敏度在各层和各列中分布不均。这些见解促使我们采用一种细粒度、灵敏度引导的方法。我们引入了MXSens,这是一种无需训练的方法,它根据列和层的灵敏度分配混合尾数比特宽度(4/6/8),自然地利用了MXINT的块结构。MXSens在一系列模型和任务上优于现有量化方法。在W4A4KV4设置下,MXSens在LLaMA-2-70B和LLaMA-3-8B上分别实现了3.77和7.63的困惑度,在WikiText-2上比现有基线有显著改进。我们的工作在大语言模型量化的准确性和资源效率之间建立了新的平衡。

英文摘要

4-bit quantization enables efficient LLM inference, but suffers from significant accuracy degradation due to outliers. Prior work addresses this problem via data rotation or mixed-precision integer quantization, but often relies on software-managed scaling and frequent dequantization, incurring substantial overhead. Microscaling formats, such as MXINT, eliminate these inefficiencies by encoding scales in hardware, yet remain incompatible with rotation-based methods. Our analysis reveals that outliers vary in severity, from rare extremes to frequent mild deviations, and that quantization sensitivity is unevenly distributed across layers and columns. These insights motivate a fine-grained, sensitivity-guided approach. We introduce MXSens, a training-free method that assigns mixed mantissa bitwidths (4/6/8) based on column- and layer-wise sensitivity, naturally leveraging the block-wise structure of MXINT. MXSens outperforms state-of-the-art quantization methods across a range of models and tasks. Under the W4A4KV4 setting, MXSens achieves perplexities of 3.77 and 7.63 on LLaMA-2-70B and LLaMA-3-8B, respectively, substantially improving over existing baselines on WikiText-2. Our work establishes a new balance between accuracy and resource efficiency for LLM quantization.

URL PDF HTML 收藏
2607.17621 2026-07-21 cs.AI 新提交

Mechanistic Attention Guidance for Agent Memory Refinement

用于智能体内存优化的机制性注意力引导

Yechao Hong, Haiquan Qiu, Yaqing Wang, Quanming Yao

机构 * Tsinghua University(清华大学) Beijing Institute of Mathematical Sciences and Applications(北京应用数学科学研究院)

AI总结 研究如何优化智能体内存,提出注意力引导内存优化框架AGMR,利用检索头注意力揭示内存使用模式,指导段级内存更新,经实验验证其能提升任务性能与内存效率。

详情
AI中文摘要

现有的自我进化内存系统主要基于文本输出(如任务轨迹和反思)来改进智能体内存。然而,这种基于文本的范式很少纳入内部机制信号,导致在任务执行期间实际如何利用检索到的内存未得到充分探索。这一局限性可能导致不可靠的错误归因和幻觉性的内存修改。在这项工作中,我们表明检索头注意力提供了一个机制信号来揭示段级内存利用情况。通过在内存段和决策步骤上聚合注意力,我们构建了一个上下文利用矩阵,该矩阵揭示了反复出现的内存使用模式并指示相应的优化策略。在此基础上,我们提出了注意力引导内存优化(AGMR)框架,该框架利用注意力揭示的使用模式来指导有针对性的段级内存更新。AGMR对失败的执行进行内存纠正或增强,对成功的执行简化内存,并通过重新执行验证每次更新。在交互式决策基准上的实验表明,与仅基于文本的内存优化基线相比,AGMR提高了任务性能和内存效率。代码可在该https URL获取。

英文摘要

Existing self-evolving memory systems mainly improve agent memory based on textual outputs, such as task trajectories and reflections. However, this text-based paradigm rarely incorporates internal mechanistic signals, leaving how retrieved memory is actually utilized during task execution underexplored. This limitation can lead to unreliable error attribution and hallucinated memory modifications. In this work, we show that retrieval-head attention provides a mechanistic signal for revealing segment-level memory utilization. By aggregating attention over memory segments and decision steps, we construct a context utilization matrix that exposes recurring memory-use patterns and indicates corresponding refinement strategies. Building on this observation, we propose Attention-Guided Memory Refinement (AGMR), a framework that uses utilization patterns revealed by attention to guide targeted segment-level memory updates. AGMR corrects or enhances memory for failed executions, simplifies memory for successful executions, and verifies each update through re-execution. Experiments on interactive decision-making benchmarks show that AGMR improves both task performance and memory efficiency over text-only memory refinement baselines. Code is available at https://anonymous.4open.science/r/AGMR_code-3262/

URL PDF HTML 收藏
2607.17523 2026-07-21 cs.CV cs.AI cs.CL 新提交

Thinking in Video: Can Video Generators Really Reason About the Real World?

视频中的思考:视频生成器真的能对现实世界进行推理吗?

Yongheng Zhang, Guang Yang, Ruihan Hou, Qiguang Chen, Ziang Liu, Xiaolong Liu, Manman Zhang, Yanchao Hao, Zheng Wei, Hao Wu, Libo Qin, Peishan Dai, Yinghui Li, Di Yin, Xing Sun

机构 * Central South University(中南大学) Tencent(腾讯) Tsinghua University(清华大学)

AI总结 探讨视频生成器能否对现实世界推理,引入因果生成双判断(CGDJ)评估,发现开源模型无明确因果感知却有合理动态,先进闭源系统推理与生成一致性有限,还揭示了视听失调问题。

详情
AI中文摘要

世界模型和视频生成的最新进展引发了一种新的推理范式,利用视频生成模型来模拟、预测和推理现实世界动态,即“视频中的思考”。但这一设想未经证实,现有指标将感知保真度与语义逻辑分开。为评估视频生成器是否支持此类推理,引入因果生成双判断(CGDJ)从两个角度审核世界模型一致性。将CGDJ应用于代表性生成器发现感知与预测存在差距,开源模型虽无明确因果感知但有合理动态,而先进闭源系统推理与生成的一致性更强但仍有限。进一步分析揭示了视听失调问题。

英文摘要

Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative models to simulate, predict, and reason about real-world dynamics. We redefine this paradigm as Thinking in Video, where video is not merely an output artifact but a medium for constructing, extending, and verifying causal thought. However, this promise remains unverified: convincing rollouts may reflect memorized appearances rather than causal understanding, while existing metrics separate perceptual fidelity from semantic logic. To evaluate whether video generators support such reasoning, we introduce the Causal-Generative Dual-Judge (CGDJ), auditing World Model Consistency from two perspectives. Explicit Causal Perception tests whether a generator reads a video scenario as a reasoning problem through spatio-temporal flattened visual question answering, while Implicit Generative Perception-Prediction Gap evaluates whether it renders the causal consequence as a consistent future video. Applying CGDJ to representative open- and closed-source generators reveals a clear Perception-Prediction Gap: open-source models produce plausible dynamics despite near-zero explicit causal perception, whereas advanced closed-source systems show stronger but still limited alignment between reasoning and generation. Further analysis exposes audio-visual misalignment, where models verbalize correct causal logic more reliably than they render it, challenging the "world simulator" narrative.

URL PDF HTML 收藏
2607.17340 2026-07-21 cs.CV 新提交

Orthogonal Knowledge Refreshing for Domain-Incremental Object Detection

用于域增量目标检测的正交知识刷新

Aoting Zhang, Dongbao Yang, Chang Liu, Xiaopeng Hong, Can Ma, Yu Zhou

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Nankai University(南开大学) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院) Tsinghua University(清华大学) Harbin Institute of Technology(哈尔滨工业大学)

AI总结 研究域增量目标检测问题,提出正交知识刷新(OKR)框架,通过构建特定域子空间、采用正交刷新策略和拓扑感知一致性,减少知识干扰和语义碎片化,实验表明该方法在mAP上比最佳无范例方法有显著提升。

Comments Accepted by ECCV 2026

详情
AI中文摘要

域增量目标检测(DIOD)要求模型在保留先验知识的同时不断适应新域。参数高效微调虽有前景,但存在覆盖关键过去知识、引发域间干扰和性能下降的风险。我们提出正交知识刷新(OKR)框架,通过为每个域构建独立的特定域子空间并融合进行整体决策,还提出基于梯度的正交刷新策略及拓扑感知一致性来减少干扰和语义碎片化。实验验证了OKR的优越性,在Pascal VOC和BDD100K系列上分别比最佳无范例方法的mAP高出5.6%和6.5%。

英文摘要

Domain-incremental object detection (DIOD) requires models to continually adapt to new domains while preserving prior knowledge. Recently, parameter-efficient fine-tuning offers a promising avenue, wherein a pre-trained model is frozen and a small number of learnable parameters are injected for downstream tasks. However, these methods risk overwriting critical past knowledge, triggering inter-domain interference and performance degradation. To address this challenge, we propose Orthogonal Knowledge Refreshing (OKR), a simple yet effective framework for DIOD. OKR incrementally constructs independent domain-specific subspaces via dedicated low-rank branches for each domain, which are seamlessly fused for a holistic decision, enabling conflict-free capacity expansion without domain selection during inference. To minimize knowledge interference during fusion, we present a gradient-based orthogonal refreshing strategy that projects gradient updates of new domains onto the orthogonal complement of the fused historical subspace, supporting continual adaptation without forgetting. Moreover, to mitigate semantic fragmentation across domains, we enforce topology-aware consistency, aligning the semantic structures of old and new domains. Extensive experiments validate the superiority of OKR, outperforming the best exemplar-free method by significant margins of +5.6% and +6.5% mAP on the Pascal VOC and BDD100K series, respectively.

URL PDF HTML 收藏
2607.17244 2026-07-21 cs.LG 新提交

DynImmune-BERT: Dynamic Immune Repertoire Modeling with Neural ODE Driven Continuous Transformers

DynImmune-BERT:基于神经常微分方程驱动的连续变换器的动态免疫组库建模

Rong Fu, Yongtai Liu, Xiaowen Ma, Haoyu Zhao, Shuo Yin, Yiqing Lyu, Long Zhang, Wangyu Wu

机构 * University of Macau(澳门大学) Hanyang University(汉阳大学) Zhejiang University(浙江大学) Wuhan University(武汉大学) Tsinghua University(清华大学) South China University of Technology(华南理工大学) University of Liverpool(利物浦大学)

AI总结 研究针对纵向T细胞受体组库建模问题,提出DynImmune-BERT模型,结合多种方法,能在有纵向组库结构时补充静态编码器,通过特定评估方式得出结果,对小外部队列和协议差异需谨慎解读。

详情
AI中文摘要

纵向T细胞受体组库包含免疫扰动后克隆扩增、收缩、消失和重新出现的信号。静态组库语言模型通常将样本总结为一袋序列,因此采样间隔、测序深度和克隆存在模式仅得到微弱体现。本文提出了DynImmune-BERT,一种用于患者水平免疫状态预测的连续时间组库模型。该方法结合了深度自适应中心对数比初始化、克隆存在门控神经常微分方程动力学、有界邻域自注意力、基于事件的状态重启以及监督优势和稀有克隆质量的混合传输目标。一个低秩元适配器初始化重新出现的克隆型,同时保持参数数量与观察到的克隆数量无关。评估将文献报道的基线与内部控制的时间比较分开,报告小外部队列的不确定性,添加校准和阈值诊断,并可视化潜在克隆轨迹和注意力邻域。结果表明,当纵向组库结构可用时,事件感知时间建模可以补充强大的静态编码器,而小外部队列和协议差异需要谨慎解释。

英文摘要

Longitudinal T cell receptor repertoires contain signals of clonal expansion, contraction, disappearance, and reappearance after immune perturbation. Static repertoire language models usually summarize a sample as a bag of sequences, so the sampling interval, sequencing depth, and clone presence pattern are only weakly represented. This paper presents DynImmune-BERT, a continuous time repertoire model for patient level immune status prediction. The method combines depth adaptive centered log ratio initialization, clone presence gated Neural ordinary differential equation dynamics, bounded neighborhood self attention, event based state restart, and a hybrid transport objective that supervises dominant and rare clone mass. A low rank meta adapter initializes reappearing clonotypes while keeping the parameter count independent of the number of observed clones. The evaluation separates literature reported baselines from internally controlled temporal comparisons, reports uncertainty for small external cohorts, adds calibration and threshold diagnostics, and visualizes latent clone trajectories and attention neighborhoods. The results indicate that event aware temporal modeling can complement strong static encoders when longitudinal repertoire structure is available, while small external cohorts and protocol differences require cautious interpretation.

URL PDF HTML 收藏
2607.17099 2026-07-21 cs.CV cs.AI 新提交

DepthART: Scaling Foundation Monocular Depth to Tiny Models

DepthART:将基础单目深度模型扩展到小型模型

Feng Xue, Wu Chen, Mingshuai Zhao, Guofeng Zhong, Anlong Ming, Haozhe Wang, Dianqiao Lei, Zhaowen Lin, Haiyang Zhang, Nicu Sebe

机构 * University of Trento(特伦托大学) Beijing University of Posts and Telecommunications(北京邮电大学) The Hong Kong University of Science and Technology(香港科技大学) Tsinghua University(清华大学)

AI总结 研究旨在将基础单目深度模型扩展到小型模型,核心方法是结合抗偏差数据采样与相机条件微调策略的DepthART,主要贡献是在多数据集上取得更好的零样本泛化和度量精度,还提供了可扩展模型家族。

Comments Accepted in ACM Multimedia 2026 ; Code: https://github.com/xuefeng-cvr/DepthART ; Project page: https://xuefeng-cvr.github.io/DepthART ;

详情
AI中文摘要

近期的几何基础模型在跨场景泛化和度量尺度预测方面显著提升了单目深度估计(MDE),但这些进展未应用于小型模型。我们通过DepthART弥合这一差距,它是用于跨场景设备部署的紧凑型MDE模型。首先识别出小型模型中两个由容量驱动的瓶颈,然后结合两种有效策略:抗偏差数据采样方案和相机条件微调协议。在多个数据集上,DepthART在零样本泛化和度量精度上均超越先前小型基线,还提供了可扩展模型家族。

英文摘要

Recent geometric foundation models (e.g., Metric3D, Depth Anything and UniDepth) have substantially improved monocular depth estimation (MDE) in both cross-scene generalization and metric-scale prediction, yet these gains have not translated to tiny models. We bridge this gap with DepthART (Depth Anything Rethought for Tiny Models), which is a compact MDE model for on-device deployment across diverse scenes. We first identify two capacity-driven bottlenecks in tiny models: (i) overfitting to dataset-specific distribution bias and (ii) unstable metric adaptation under camera shift, where full fine-tuning easily damages transferable geometry. Accordingly, DepthART combines two simple but effective strategies: a bias-resistant data sampling scheme to reduce distribution bias under the same training budget, and a camera-conditioned fine-tuning protocol that freezes the distilled encoder and adjusts metric scale conditioned on intrinsics while better preserving cross-dataset generalization. Across datasets, DepthART consistently surpasses previous tiny baselines in both zero-shot generalization and metric accuracy (e.g., zero-shot $δ_1$=0.964 for DepthART-S on NYUD v2), and in some cases approaches heavy models. We further provide a scalable model family, with DepthART-S reaching 347/245 FPS (strict FP32) on an RTX A6000 at $224^2/448^2$, 102 FPS (TF32) on a Orin NX 8GB, and over 15 FPS (FP32) on a Jetson Nano 4GB.

URL PDF HTML 收藏
2607.17097 2026-07-21 cs.CV 新提交

HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis

HarmoHOI:用于多视图手-物体交互合成的外观与3D运动协调

Lingwei Dang, Juntong Li, Zonghan Li, Hongwen Zhang, Liang An, Wei Min, Yebin Liu, Qingyao Wu

机构 * South China University of Technology(华南理工大学) Beijing Normal University(北京师范大学) Tsinghua University(清华大学) Shadow AI(影谱科技)

AI总结 研究针对多视图手-物体交互合成难题,提出HarmoHOI统一扩散框架,用多视图扩散Transformer混合模型联合建模视频与点轨迹,结合全局运动对齐扩散确保一致性,采用混合数据学习策略,实现多视图HOI高质量合成。

详情
AI中文摘要

手-物体交互(HOI)合成是动画制作和具身人工智能的基石。尽管视频基础模型有强大先验,但由于手部动作复杂和遮挡,多视图一致的HOI合成仍具挑战性。我们提出HarmoHOI,一个统一的扩散框架,可联合并协调地生成同步多视图HOI视频和全局对齐的3D点轨迹。核心见解是强大的多视图一致性根本上需要全局对齐的3D几何和运动。为此,我们提出多视图扩散Transformer混合模型来共同建模RGB视频和3D点轨迹。通过将点轨迹表示为伪视频,使3D几何信号与基础模型的2D潜在空间对齐,最小化域差距并简化先验适应。为进一步确保几何一致性,引入全局运动对齐扩散,将粗糙点轨迹细化为度量尺度、全局对齐的3D轨迹。HarmoHOI能在去噪过程中实现2D外观和3D运动的实时协同进化。为克服多视图HOI数据稀缺问题,采用混合数据课程学习策略,成功将单视图数据的通用先验转移到同步多视图生成中。实验结果表明HarmoHOI在视觉质量、运动合理性和多视图几何一致性方面达到了当前最优性能。

英文摘要

Hand-Object Interaction (HOI) synthesis is a cornerstone for animation production and embodied AI. Despite the strong priors of video foundation models, multi-view consistent HOI synthesis remains challenging due to complex hand motions and occlusions. We present HarmoHOI, a unified diffusion framework that jointly and harmoniously generates synchronized multi-view HOI videos and globally aligned 3D point tracks. Our core insight is that robust multi-view consistency fundamentally requires globally aligned 3D geometry and motion. To this end, we propose a Mixture of Multi-view Diffusion Transformer that co-models RGB videos and 3D point tracks. By representing point tracks as pseudo-videos, we align 3D geometric signals with the 2D latent space of foundation models, thereby minimizing the domain gap and easing adaptation of priors. To further ensure geometry consistency, we introduce Global Motion Aligning Diffusion, which refines coarse point tracks into metric-scale, globally aligned 3D trajectories. HarmoHOI enables on-the-fly co-evolution of 2D appearance and 3D motion during denoising. To overcome the scarcity of multi-view HOI data, we employ a hybrid data curriculum learning strategy that successfully transfers generic priors from single-view data to synchronized multi-view generation. Experimental results show that HarmoHOI achieves state-of-the-art performance in visual quality, motion plausibility, and multi-view geometric consistency. Project page available at https://droliven.github.io/HarmoHOI_project.

URL PDF HTML 收藏
2607.17095 2026-07-21 cs.AI 新提交

Fourier Geometric Wind Power Forecasting with Numerical Weather Prediction

基于数值天气预报的傅里叶几何风力发电预测

Shiyuan Piao, Fan Zehui, Yang Liu, Hong Cheng, Juepeng Zheng, Jie Zhou, Fugee Tsung

机构 * The Hong Kong University of Science and Technology(香港科技大学) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) Goldwind Science and Technology Co.,Ltd(金风科技股份有限公司)

AI总结 研究针对风力发电预测难题,提出多模态框架,整合SCADA数据与NWP预测。先分解输入特征,再用几何编码器和傅里叶神经算子建模,实验表明该模型优于现有基线,凸显基于物理设计的有效性。

详情
AI中文摘要

准确的短期风力发电预测对电网稳定性和运营规划至关重要,但由于大气条件与涡轮机动力学之间的复杂相互作用,这一任务仍具有挑战性。现有方法未能有效整合天气预报与风力涡轮机数据(即SCADA),导致解决方案欠佳。为解决此问题,我们引入了一个多模态框架,将基于历史点的SCADA数据与基于网格的数值天气预报(NWP)预测相结合,这因异构输入和复杂的物理风力涡轮机相互作用而颇具挑战。我们的方法首先将输入明确分解为标量和矢量特征,以更好地捕捉特定地点和几何相关性,然后采用几何编码器从风矢量中提取旋转不变特征。我们还利用了傅里叶神经算子(FNO)架构,它在频域中执行全局卷积,以有效建模远程时空关系。在三个实际风电场进行的广泛实验表明,我们的模型始终优于现有基线,凸显了其基于物理的设计的有效性。我们方法的核心实现可在指定网址公开获取。

英文摘要

Accurate short-term wind power forecasting is essential for grid stability and operational planning, yet remains challenging due to the complex interactions between atmospheric conditions and turbine dynamics. However, existing methods fail to effectively incorporate weather forecasting with wind turbine data (i.e., SCADA), leading to suboptimal solutions. To address this, we introduce a multimodal framework that integrates historical point-based SCADA data with grid-based Numerical Weather Prediction (NWP) forecasts, which is challenging due to heterogeneous input and the complex physical wind-turbine interactions. Our approach first explicitly decomposes inputs into scalar and vector features to better capture both site-specific and geometric dependencies and then incorporates a geometric encoder to extract rotation-invariant features from wind vectors. We further leverages a Fourier Neural Operator (FNO) architecture, which performs global convolutions in the frequency domain to efficiently model long-range spatiotemporal relationships. Extensive experiments on three real-world wind farms, with weather forecasting data, demonstrate that our model consistently outperforms state-of-the-art baselines, highlighting the effectiveness of its physically-informed design. The core implementation of our method is publicly available at: https://github.com/shawn-sypiao/GWPF.

URL PDF HTML 收藏
2607.17070 2026-07-21 cs.AI 新提交

Bridging the Information Gap: Semantic Densification and Hindsight Distillation for Cold-Start Prediction

弥合信息差距:冷启动预测的语义致密化与事后蒸馏

Hao Duong Le, Yifei Gao, Huan Li, Lun Jiang, Chen Bai, Ke Xing, Chen Zhang

机构 * Tsinghua University(清华大学) Meituan(美团)

AI总结 针对电子商务平台新用户冷启动预测难题,SemRaD框架利用结构化语义推理管道和事后感知蒸馏网络,有效弥合信息差距,提升了LTV和CVR,减少训练数据用量,在工业数据集和在线测试中均取得良好效果。

详情
AI中文摘要

新用户冷启动是电子商务平台的关键瓶颈,即预测交互历史稀疏用户的终身价值(LTV)和转化率(CVR)。基于大语言模型的语义增强和使用特权信息学习这两个先前方向各有局限。本文提出SemRaD框架,通过结构化语义推理管道生成致密化语义配置文件和事后蒸馏目标,利用事后感知蒸馏网络传递特权知识。在大规模工业数据集上,SemRaD提升了LTV和CVR,在Keeta的四周在线A/B测试也有成效,还能用更少训练数据匹配生产系统的LTV并提升CVR。

英文摘要

New-user cold-start is a critical bottleneck for e-commerce platforms: predicting user lifetime value (LTV) and conversion rate (CVR) for users with sparse interaction history. Two prior directions -- LLM-based semantic augmentation and learning using privileged information (LUPI) -- each face a key limitation. First, LLM augmentation produces unstructured rationales that are noisy and hard to operationalize in production. Second, naive student-teacher distillation can be brittle due to an information gap between the privileged teacher and the sparse student; moreover, this gap is heterogeneous across users. We propose SemRaD, a Semantic Reasoning-aware Distillation framework addressing both limitations. First, a Structured Semantic Reasoning Pipeline replaces free-form rationales with a structured schema built via a discover-curate-audit workflow, producing per user a Densified Semantic Profile (consumed by the deployed student via a Semantic-Gated Encoder that focuses on the most informative dimensions) and a Hindsight Distillation Target reconciled from pre- and post-conversion reasoning (used only at training). Second, to bridge this gap and handle its heterogeneity, a Hindsight-Aware Distillation Network transfers privileged knowledge via the hindsight target, with Distillation Experts improving transfer under per-user variability. On a large-scale industrial dataset, SemRaD lifts +1.9% LTV (Gini) and +1.0% CVR (AUROC) over a production-grade base; a four-week online A/B at Keeta confirms +1.0% LTV / +0.43% CVR. SemRaD also matches the production system's LTV using only 9% of the training data while improving CVR by 0.8%.

URL PDF HTML 收藏
2607.17050 2026-07-21 cs.CV cs.AI 新提交

EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding

EvoGUI:用于GUI状态转换理解的进化感知基准测试

Yaohan Yang, Minglei Shi, Borui Zhang, Jie Zhou, Jiwen Lu

机构 * Tsinghua University(清华大学)

AI总结 研究GUI状态转换理解,提出EvoGUI诊断框架,将GUI轨迹转化为视觉问答探针,无需额外注释。通过Mind2Web和WebLINX实例化EvoGUI-Bench并零样本评估28种模型配置,结果显示其可补充端到端评估,揭示理解提升空间。

详情
AI中文摘要

GUI智能体必须推断动作如何改变界面状态,但端到端成功率将这种能力与感知、基础、规划和恢复能力纠缠在一起。我们引入了EvoGUI,这是一个诊断框架,它将标准化的GUI轨迹转换为三个互补的视觉问答探针:时间排序、反向动作/值预测和对比性单步后继判别。它们的标签来自轨迹顺序和记录的动作,在轨迹标准化后无需额外的任务标签注释。我们从Mind2Web和WebLINX实例化了EvoGUI-Bench,在120个域中产生了3000个实例,并对28种视觉语言模型配置进行了零样本评估。最强的模型仅达到60.4 EvoGain,而模型规模和GUI专业化并不能可靠地预测性能。这些结果将EvoGUI-Bench确立为端到端GUI智能体评估的可扩展诊断补充,同时揭示了状态转换理解方面的巨大提升空间。源代码可在这个https URL上公开获取。

英文摘要

GUI agents must reason about how actions transform interface states, but end-to-end success rates entangle this ability with perception, grounding, planning, and recovery. We introduce EvoGUI, a diagnostic framework that converts normalized GUI trajectories into three complementary visual question answering probes: temporal ordering, inverse action/value prediction, and contrastive one-step successor discrimination. Their labels are derived from trajectory order and logged actions, requiring no additional task-label annotation after trajectory normalization. We instantiate EvoGUI-Bench from Mind2Web and WebLINX, yielding 3,000 instances across 120 domains, and evaluate 28 vision-language model configurations zero-shot. The strongest model reaches only 60.4 EvoGain, while model scale and GUI specialization do not reliably predict performance. These results establish EvoGUI-Bench as a scalable diagnostic complement to end-to-end GUI-agent evaluation while exposing substantial headroom in state-transition understanding. The source code is publicly available at https://github.com/Yyhhh6/EvoGUI.

URL PDF HTML 收藏
2607.16943 2026-07-21 cs.RO 新提交

SinD 2.0: A Multi-City UAV Dataset with Semantic Risk Annotations for SOTIF-Oriented Safety Validation at Signalized Intersections

SinD 2.0:一个用于信号交叉口面向SOTIF安全验证的具有语义风险注释的多城市无人机数据集

Yunwei Li, Shengjie Fu, Chunrong Chen, Chengxiang Zhao, Yuchen Fan, Mingyu Zhu, Yanchao Xu, Yuxin Zhang, Lan Yang, Chuzhao Li, Jie Ji, Yi He, Abhijit Sarkar, Akash Sonth, Hong Wang, Jun Li

机构 * Tsinghua University(清华大学) Beijing Institute of Technology(北京理工大学) Guangzhou Automobile Group Co., Ltd.(广州汽车集团股份有限公司) Jilin University(吉林大学) Chang’an University(长安大学) Chongqing University(重庆大学) Southwest University(西南大学) Wuhan University of Technology(武汉理工大学) Virginia Tech Transportation Institute(弗吉尼亚理工大学交通研究所) Virginia Tech(弗吉尼亚理工大学)

AI总结 针对自动驾驶系统在信号交叉口安全验证的瓶颈,介绍SinD 2.0多城市无人机数据集。通过跨域多样、高密度风险交互、分层语义注释及全栈测试工具链,能有效暴露算法性能局限,为ADS安全分析提供有力支持。

详情
AI中文摘要

信号交叉口的安全验证仍然是自动驾驶系统(ADS)部署的关键瓶颈,因为这些场景涉及密集的异质交通、有争议的路权和长尾安全关键交互,对预期功能安全(SOTIF)构成重大挑战。现有自然驾驶数据集存在地理同质性、安全关键事件稀疏和缺乏语义风险注释等问题,限制了算法泛化性评估和针对性SOTIF验证。本文介绍了SinD 2.0,一个用于跨域ADS安全分析的大规模基于无人机的交叉口数据集。其主要贡献包括跨域多样性,涵盖中国四个城市的六个信号交叉口;高密度风险交互,通过替代安全措施提取32682个安全关键事件;分层语义注释,提供包括交通违规等多维语义标签;全栈测试工具链,支持多种测试方式。基准实验表明SinD 2.0在不同城市间存在显著域转移,语义风险子集能有效暴露ADS算法性能局限。数据集、注释和测试工具链可通过链接获取。

英文摘要

Safety validation at signalized intersections remains a critical bottleneck for the deployment of autonomous driving systems (ADS), as these scenarios involve dense heterogeneous traffic, contested right of way, and long-tail safety-critical interactions, posing significant challenges to the Safety of the Intended Functionality (SOTIF). Existing naturalistic driving datasets often suffer from geographical homogeneity, sparsity of safety-critical events, and lack of semantic risk annotations, which limit the evaluation of algorithmic generalizability and targeted SOTIF verification. To address these gaps, this paper introduces SinD 2.0, a large-scale drone-based intersection dataset dedicated to cross-domain ADS safety analysis. The main contributions of SinD 2.0 are: (1) Cross-domain diversity: It covers six signalized intersections across four Chinese cities, capturing distinct intersection topologies and regional driving behavior characteristics; (2) High-density risk interactions: A total of 32,682 safety-critical events are extracted via surrogate safety measures, significantly enriching the density of boundary test scenarios; (3) Hierarchical semantic annotations: Besides integration with high-definition (HD) maps and Signal Phase and Timing (SPaT) data, it provides multi-dimensional semantic labels including traffic violations, high-risk interactions, visual shielding, and narrow feasible areas; (4) Full-stack testing toolchain: It supports automated scenario extraction, prediction-only evaluation, open-loop replay, reactive closed-loop testing, and photorealistic rendering. Benchmark experiments demonstrate that SinD 2.0 exhibits significant domain shifts across cities, and the semantic risk subsets can effectively expose the performance limitations of ADS algorithms. The dataset, annotations, and testing toolchain are available at https://github.com/SOTIF-AVLab/SinD/tree/main.

URL PDF HTML 收藏
2607.16841 2026-07-21 cs.CV cs.MM 新提交

Look Clearly Before Answering: Mitigating Hallucinations in LVLMs via Saliency-Driven Perceptual Realignment

回答前看清楚:通过显著性驱动的感知重新对齐减轻LVLMs中的幻觉

Pengxu Chen, Yao Zhu, Guangming Zhu, Jun Sheng, Jincai Huang, Xiangyang Ji, Liang Zhang

机构 * Xidian University(西安电子科技大学) Tsinghua University(清华大学) Shanghai Road Transport Development Center(上海市道路运输发展中心) Hunan Institute of Advanced Technology(湖南先进技术研究院)

AI总结 研究针对LVLMs易产生幻觉问题,提出无需训练的SDPR框架,通过显著性驱动注意力重新分配、缓存对齐及先验约束对比解码,整体对齐视觉意识,在多基准测试中优于现有方法,无需额外训练且开销小。

Comments Accepted by ACM Multimedia 2026

详情
AI中文摘要

大型视觉语言模型(LVLMs)在多模态理解方面展现出卓越能力,但仍易产生与视觉证据不一致的幻觉。现有缓解方法多关注语言先验偏差或跨模态不平衡,而感知和记忆中的渐进视觉退化未被充分探索。本文提出显著性驱动的感知重新对齐(SDPR),这是一个无需训练的框架,可减轻推理过程中视觉意识的退化。具体包括:通过显著性驱动的注意力重新分配释放被非语义下沉令牌劫持的注意力,恢复关键视觉证据;识别KV缓存中的空间失真,提出显著性驱动的缓存对齐以在生成过程中保留与查询相关的视觉特征;引入先验约束对比解码以惩罚由主导语言先验引起的不忠实预测。大量实验表明,SDPR在幻觉和通用基准测试中均优于现有方法,无需额外训练且运行时开销最小。

英文摘要

Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding. However, they remain prone to hallucinations, generating responses that are inconsistent with the visual evidence. Existing mitigation methods largely address language-prior bias or cross-modal imbalance, while progressive visual degradation across perception and memory remains underexplored. In this work, we propose Saliency-Driven Perceptual Realignment (SDPR), a training-free framework that mitigates the degradation of visual awareness throughout inference. Specifically, we first introduce saliency-driven attention redistribution to release attention hijacked by non-semantic sink tokens, thereby recovering critical visual evidence. Second, we identify spatial distortion in the KV cache and propose saliency-driven cache alignment to preserve query-relevant visual features during generation. Finally, we introduce prior-constrained contrastive decoding to penalize unfaithful predictions induced by dominant language priors. Our proposed SDPR is robust against hallucinations due to its holistic alignment of visual awareness across the entire generative trajectory. Extensive experiments across diverse LVLM architectures show that SDPR outperforms state-of-the-art methods on both hallucination and general-purpose benchmarks, requiring no additional training and incurring minimal runtime overhead. The code is available \href{https://github.com/PengSyuChen/SDPR}{\color{blue}{here}}.

URL PDF HTML 收藏
2607.16828 2026-07-21 cs.CV 新提交

UniNDM: A Unified Noise-driven Detection and Mitigation Framework Against Sexual Content in Text-to-Image Generation

UniNDM:文本到图像生成中针对性内容的统一噪声驱动检测与缓解框架

Yao Huang, Yitong Sun, Huanran Chen, Ruochen Zhang, Shouwei Ruan, Ranjie Duan, Maoxun Yuan, Yinpeng Dong, Hui Xue, Xiaochun Cao, Xingxing Wei

机构 * Institute of Artificial Intelligence, State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室人工智能研究院) College of Artificial Intelligence, Tsinghua University(清华大学人工智能学院) Security Department, Alibaba Group(阿里巴巴集团安全部) School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-Sen University(中山大学深圳校区网络科学与技术学院)

AI总结 针对文本到图像生成易受隐式性提示影响的问题,提出UniNDM统一噪声驱动框架。利用早期预测噪声的可分离性开发轻量级检测器,引入噪声增强自适应负引导缓解问题,扩展到扩散变压器架构,实验显示比现有方法有显著改进。

Comments 18 pages, 10 figures, accepted by TPAMI

详情
AI中文摘要

尽管文本到图像扩散模型具有强大的生成能力,但它们容易受到隐式性提示的影响,由于模型偏差或训练数据中的潜在相关性,微妙线索会意外生成不当内容。现有安全机制存在根本局限性。为此,我们提出UniNDM,一个统一的噪声驱动框架,通过扩散过程中的噪声动态来重新思考安全机制。我们发现早期预测噪声在正常和性明确内容之间具有内在可分离性,并理论证明其语义浓度随时间步长二次增加。利用此特性,我们开发了轻量级基于噪声的检测器,准确率高且几乎无计算开销。对于缓解,我们引入噪声增强自适应负引导,通过大语言模型动态生成特定上下文负提示,同时通过抑制对明确令牌的注意力集中来优化初始噪声。我们还将框架扩展到新兴的扩散变压器架构。综合实验表明,我们的方法比现有方法有显著改进。

英文摘要

Despite the impressive generative capabilities of text-to-image diffusion models, they remain vulnerable to implicit sexual prompts, where subtle cues disguised as benign terms or adversarial tokens unexpectedly generate the inappropriate content due to model biases or latent correlations in training data. Existing safety mechanisms face fundamental limitations: detection methods primarily identify explicit content and fail to capture implicit malicious intent, while mitigation approaches rely on static negative prompts inadequate for diverse implicit scenarios. To address these challenges, we propose UniNDM, a unified noise-driven framework that rethinks safety mechanisms through the lens of noise dynamics in diffusion processes. Our key insight is that early-stage predicted noise exhibits inherent separability between normal and sexually explicit content, which we theoretically demonstrates quadratically increasing semantic concentration with timestep. Leveraging this property, we develop a lightweight noise-based detector achieving superior accuracy with virtually no computational overhead. For mitigation, we introduce noise-enhanced adaptive negative guidance: dynamically generating context-specific negative prompts via large language models to handle diverse implicit content, while optimizing initial noise by suppressing attention concentration on explicit tokens to provide comprehensive protection. Besides the U-Net-based diffusion models, we further extend our framework to emerging Diffusion Transformer architectures through region-constrained semantic guidance tailored for their unified multimodal attention. Comprehensive experiments across U-Net models and DiT models on both natural and adversarial datasets demonstrate substantial improvements over state-of-the-art methods, including SLD, UCE, Safree, etc. Our code is publicly available at https://github.com/Aries-iai/UniNDM.

URL PDF HTML 收藏
2607.16692 2026-07-21 cs.SE cs.CL 新提交

Dependency-Guided Code Generation: Structured Matrix Decomposition and Consistency-Guided Refinement

依赖引导的代码生成:结构化矩阵分解与一致性引导的细化

Mingqiao Mo, Yangchen Zeng, Zikai Xiao, Xin Xiao, Wenhua Nie, Zhaolu Kang, Guangyuan Dong, Kai Shu, Hao Zhang, Xiaodong Fan

机构 * University of the Chinese Academy of Sciences(中国科学院大学) ByteDance Inc.(字节跳动公司) Zhejiang University(浙江大学) National Taiwan University(台湾国立大学) Peking University(北京大学) Alibaba Group(阿里巴巴集团) Tsinghua University(清华大学) Liaoning Technical University(辽宁技术大学)

AI总结 针对现有代码生成方法无法充分捕捉代码实体依赖关系的问题,提出依赖感知代码生成框架,通过结构化矩阵分解和一致性引导细化生成代码,经实验验证该方法能生成语义对齐和结构保真度更高的代码。

Comments 12 pages

详情
AI中文摘要

现代软件系统日益复杂,使自动代码生成成为软件工程中的一项基本任务。然而,现有方法往往无法充分捕捉代码实体间复杂的多层次依赖关系,导致生成的代码逻辑不完整或难以集成到实际系统中。为解决此局限,我们提出一个依赖感知代码生成框架,通过基于图的表示明确建模代码实体间的交互。我们将依赖分解为两个互补组件:一个捕获强显式关系的量化矩阵和一个对弱隐式交互建模的稀疏低秩分解。通过交替优化过程有效学习分解。在代码生成期间,将学习到的依赖结构作为约束纳入,确保生成代码的语义连贯和结构一致。此外,我们为强依赖引入稀疏三元组表示,显著提高存储效率和计算可扩展性。大量实验表明,与现有方法相比,我们的方法始终能生成具有更高语义对齐和结构保真度的代码。

英文摘要

The increasing complexity of modern software systems has made automated code generation a fundamental task in software engineering. However, existing approaches often fail to adequately capture the intricate, multi-level dependencies among code entities, leading to generated code that is logically incomplete or difficult to integrate into real-world systems. To address this limitation, we propose a dependency-aware code generation framework that explicitly models interactions among code entities through a graph-based representation. We decompose dependencies into two complementary components: a quantized matrix that captures strong, explicit relations, and a sparse low-rank factorization that models weaker, implicit interactions. The decomposition is efficiently learned via an alternating optimization procedure. During code generation, the learned dependency structure is incorporated as a constraint, ensuring both semantic coherence and structural consistency of the generated code. Furthermore, we introduce a sparse triplet representation for strong dependencies, significantly improving storage efficiency and computational scalability. Extensive experiments demonstrate that our approach consistently produces code with superior semantic alignment and structural fidelity compared to existing methods.

URL PDF HTML 收藏
2607.16427 2026-07-21 cs.CL 新提交

Multi-level context Modeling for consistent expert selection in Mixture-of-Experts

用于混合专家模型中一致专家选择的多层次上下文建模

Shuhan Huang, Naifan Zhang, Yuanbo Tang, Yang Li, Wai Kin Victor Chan

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) School of AI, The Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳)人工智能学院)

AI总结 研究混合专家模型中专家选择问题,提出多层次上下文融合MoE框架,通过整合跨层语义聚合和局部令牌级交互信号构建上下文感知表示,提升路由一致性和下游性能。

详情
AI中文摘要

混合专家模型(MoE)通过将令牌路由到一小部分专家来实现Transformer模型的高效扩展。然而,现有路由器通常基于浅层或孤立的令牌表示来进行专家选择,这往往会在各层产生不稳定且语义不一致的路由决策。在这项工作中,我们从表示角度重新审视专家选择,并将上下文不完整性识别为限制有效专家专业化的关键瓶颈。为解决此问题,我们提出了多层次上下文融合MoE(MCF-MOE)框架,该框架通过整合来自跨层语义聚合和局部令牌级交互的互补信号来构建上下文感知表示,从而实现更具信息性和一致性的专家选择。在语言建模和理解基准测试上的实验表明,MCF-MOE相对于强大的MoE基线持续提高了路由一致性和下游性能,突出了上下文完整性在专家路由中的重要性。代码可在该https URL获取。

英文摘要

Mixture-of-Experts (MoE) enables efficient scaling of Transformer models by routing tokens to a small subset of experts. However, existing routers typically condition expert selection on shallow or isolated token representations, which often produce unstable and semantically inconsistent routing decisions across layers. In this work, we revisit expert selection from a representation perspective and identify context incompleteness as a key bottleneck limiting effective expert specialization. To address this issue, we propose Multi-level Context Fusion MOE (MCF-MOE), a framework that constructs context-aware representations by integrating complementary signals from cross-layer semantic aggregation and local token-level interactions, enabling more informative and consistent expert selection. Experiments on language modeling and understanding benchmarks demonstrate that MCF-MOE consistently improves routing consistency and downstream performance over strong MoE baselines, highlighting the importance of contextual completeness in expert routing. The code is available at https://anonymous.4open.science/r/MCFMOE.

URL PDF HTML 收藏
2607.16247 2026-07-21 cs.LG cs.CV 新提交

Self-Evolving Just-In-Time Memory for Proactive Embodied Safety

用于主动具身安全的自进化即时记忆

Bingrui Sima, Lizhong Wang, Xiaoya Lu, Kun He, Xiao Yang

机构 * Huazhong University of Science and Technology(华中科技大学) Tsinghua University(清华大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 研究视觉语言模型在闭环交互中应对动态危险的问题,提出自进化即时记忆框架,含RSG、事实记忆和经验记忆,并通过自动测试-验证-写入循环完善元技能,实验证明该框架大幅提升安全成功率且不阻碍任务进展。

详情
AI中文摘要

虽然视觉语言模型(VLMs)使具身智能体能够执行复杂的家庭任务,但它们在闭环交互中难以主动应对动态出现的危险。现有安全方法常依赖运行时护栏来阻止不安全行为或导致过度谨慎,严重阻碍任务进展。为打破安全与进展的权衡,我们引入自进化即时记忆框架,将具身安全从阻碍进展的护栏转变为主动减轻危险。该框架由用于部分可观测下持续安全相关状态跟踪的风险充足拓扑信念图(RSG)、用于精确危险预测的基于智能体的事实记忆以及注入程序元技能以指导可执行、保持进展的减轻危险的经验记忆组成。此外,我们提出自动测试-验证-写入循环,让智能体在测试时能从执行轨迹中不断完善减轻危险的元技能。在IS-Bench上的实验表明,我们的框架大幅提高了多个VLM主干的安全成功率(如在Qwen3-VL-8B上提高30.3%),使智能体能够主动减轻危险而不阻碍任务进展。

英文摘要

While Vision-Language Models (VLMs) have empowered embodied agents to execute complex household tasks, they struggle to proactively handle dynamically emerging hazards during closed-loop interactions. Existing safety approaches often rely on runtime guardrails to block unsafe actions or induce excessive caution, which severely stalls task progress instead of actively resolving the underlying risks. To break this safety-progress trade-off, we introduce the Self-Evolving Just-In-Time Memory framework, which reframes embodied safety from progress-stalling guardrails to proactive hazard mitigation. The framework consists of a Risk-Sufficient Topological Belief Graph (RSG) for persistent safety-relevant state tracking under partial observability, an Agency-Grounded Factual Memory for precise hazard anticipation, and an Experience Memory that injects procedural Meta-Skills to guide executable, progress-preserving mitigation. Furthermore, we propose an automated Test-Verify-Write loop, allowing agents to continually refine their mitigation Meta-Skills from execution traces at test time. Experiments on IS-Bench demonstrate that our framework substantially boosts the Safe-Success rate across multiple VLM backbones (e.g., +30.3% on Qwen3-VL-8B), enabling agents to proactively mitigate hazards without stalling task progress. Code is available at https://github.com/DyMessi/JIT-Memory.

URL PDF HTML 收藏
2607.04438 2026-07-21 cs.CV cs.AI cs.HC cs.MA cs.MM 版本更新

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

ResearchStudio-Reel:实现从论文到海报、视频和博客的研究最后一公里自动化

Lingao Xiao, Yalun Dai, Yangyu Huang, Qihao Zhao, Wenshan Wu, Hugo He, Ruishuo Chen, Jin Jiang, Qianli Ma, Jiahuan Zhang, Xin Zhang, Ying Xin, Yang Ou, Yan Xia, Scarlett Li, Longbo Huang, Zhipeng Zhang, Yang He, Yap Kim Hui, Yan Lu

机构 * Microsoft Research(微软研究院) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) Tsinghua University(清华大学) Peking University(北京大学) Shanghai Jiao Tong University(上海交通大学) Westlake University(西湖大学) CFAR, A*STAR(计算科学与工程研究所,新加坡科技研究局)

AI总结 研究传播自动化困难,以往方法有局限。该研究提出将最后一公里构建为技能组合,实例化ResearchStudio-Reel,包括共享提取器、可编辑生成器和交互式收敛层,能产出多种可编辑工件,效果优于现有系统。

详情
AI中文摘要

研究传播,即将论文转化为海报、演讲视频和博客文章,仍然是手动的最后一公里。以前的自动化方法孤立地处理每个工件,每个都从头重新提取论文,通常提供单向渲染,作者无法在PowerPoint或Word中重新打开,并且根据软VLM偏好分数来评估质量,而在承载部分仍为空时分数会趋于平稳。我们认为这最后一公里最好构建为技能组合:瘦代理可读契约,共享一个上游提取器,并在测量填充循环中包装确定性原语,其出口是硬通过/失败渲染门。我们将其实例化为ResearchStudio-Reel,五个Claude代码和Codex技能组织成一个共享提取器(Paper2Assets)、三个可编辑生成器(Paper2Poster、Paper2Video、Paper2Blog)和一个交互式收敛层(Paper2Reel)。Paper2Assets将每篇论文提取一次到一个共享包中,供每个下游技能重用;三个生成器生成一个可打印的海报、一个同步的演讲视频和一个双语博客,它们在事实层面上保持一致,并能通过PowerPoint或Word进行往返;Paper2Reel然后将这三个绑定到一个独立的HTML查看器中,其部分级点击会使视频、幻灯片、字幕和博客跳转到匹配的内容。在Paper2Poster基准测试中,我们的海报在美学和信息子标准方面领先于先前的自动化系统和单镜头前沿语言模型,在两名外部VLM评委的评估下,在美学方面超过了作者自己的海报,并在84%至93%的论文中总体获胜;能力审计进一步表明,通过将与叙述对齐的幻灯片亮点与由布局感知DOCX修复控制的双语博客独特配对,ResearchStudio-Reel是唯一能够提供所有三个可编辑工件的管道。项目可在此https URL上获取

英文摘要

Despite growing automation, turning a paper into a coherent poster, talk video, and blog piece often remains a labor-intensive last mile. Recent systems increasingly generate multiple dissemination formats, but a practical workflow must also keep the outputs editable in native tools and bound into one navigable deliverable for revision and reuse. We present ResearchStudio-Reel, a native-editable dissemination workspace that binds its three artifacts into one interactive deliverable at the experience level, implemented as five skills executable in Claude Code and Codex: one shared extractor, three editable artifact generators, and one interactive convergence layer. A shared asset bundle feeds a PowerPoint poster and video deck, plus a bilingual Word blog; rather than re-rendering the paper into a fourth format, Paper2Reel converges these already-produced artifacts at the experience level, binding poster regions, video segments, and blog passages into one interactive viewer. Artifact-specific release checks make this delivery contract testable, and Paper2Poster additionally uses a measured-fill loop. On the Paper2Poster benchmark, our Claude Code configuration achieves the best scores among automated systems on all three aesthetic sub-criteria and the best or tied-best scores on two of three information sub-criteria. Under two VLMjudges, it exceeds the authors' posters in average aesthetics (3.56 vs. 3.03) and wins on overall quality on 74 and 95 of the 100 papers under the two judges. The full pipeline additionally packages the native-editable source artifacts and their aligned viewer. Project is available at https://aka.ms/ResearchStudio

URL PDF HTML 收藏
2606.22394 2026-07-21 cs.CV 版本更新

Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning

曲率自适应一致性流匹配:基于强化学习的自主轨迹优化

Songtao Tian, Guhan Chen, Bohan Li, Jingyi Ma, Zixiong Yu

机构 * Tsinghua University(清华大学)

AI总结 提出曲率自适应一致性流匹配(CACFM),利用强化学习代理动态构建效率导向课程,优先处理关键区域,在FLUX和SDXL等大规模模型上实现最先进结果,极少数步下保持高视觉保真度。

Comments Accepted by ECCV 2026

详情
AI中文摘要

一致性蒸馏显著加速了扩散模型的推理。本文揭示了一个有趣的不对称性:虽然Logit-Normal采样先验对于标准迭代生成非常有效,但一致性蒸馏表现出截然不同的难度分布(例如U形)。我们发现主要优化瓶颈位于边界阶段(初始化或最终细化)而非中间步骤。为解决静态采样无法适应不断变化的学习需求的问题,我们提出了曲率自适应一致性流匹配(CACFM)。通过将蒸馏建模为动态决策过程,CACFM采用轻量级强化学习代理主动探测概率流ODE轨迹,自动构建效率导向的课程,优先处理关键区域而无需手动调度。结合新颖的流分布匹配蒸馏(DMD)目标,我们的方法在FLUX和SDXL等大规模模型上取得了新的最先进结果。它有效缓解了结构畸形,并在极端少步数情况下保留了高频细节,实现了前所未有的视觉保真度。

英文摘要

Consistency distillation has significantly accelerated diffusion-model inference, but its sampling dynamics remain underexplored. We reveal an asymmetry: although Logit-Normal sampling priors work well for standard iterative generation, consistency distillation exhibits a different difficulty profile (e.g., U-shaped), with optimization bottlenecks concentrated at the boundary stages rather than intermediate steps. To address the limitations of static sampling under evolving learning demands, we propose Curvature-Adaptive Consistency Flow Matching (CACFM). By formulating distillation as a dynamic decision process, CACFM uses a lightweight reinforcement learning agent to probe Probability Flow ODE trajectories and construct an efficiency-oriented curriculum that prioritizes critical regions without manual scheduling. Combined with Flow-adapted DMD and adversarial consistency objectives, our RL-based scheduler achieves state-of-the-art results on large-scale models such as FLUX and SDXL, mitigating structural deformities and preserving high-frequency details in extreme few-step regimes.

URL PDF HTML 收藏
2606.14409 2026-07-21 cs.RO cs.AI 版本更新

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack

Hy-Embodied-0.5-VLA:从视觉-语言-动作模型到真实世界机器人学习栈

He Zhang, Lingzhu Xiang, Haitao Lin, Zeyu Huang, Minghui Wang, Dingyan Zhong, Yubo Dong, Yihao Wu, Yongming Rao, Dongsheng Zhang, Wanjia He, Ling Chen, Kai Huang, Jiahao Chen, Sichang Su, Xumin Yu, Ziyi Wang, Chengwei Zhu, Xiao Teng, Yuchun Guo, Yufeng Zhang, Yuandong Liu, Rui Wang, Zisheng Lu, Han Hu, Zhengyou Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学)

AI总结 提出端到端机器人学习栈HyVLA-0.5,涵盖数据收集、模型设计、预训练与微调、RL后训练及真实部署,各组件协同工作。

详情
AI中文摘要

在本报告中,我们提出Hy-Embodied-0.5-VLA,简称HyVLA-0.5,一个覆盖完整机器人学习栈的端到端系统:数据收集、模型设计、持续预训练和监督微调、RL后训练以及真实世界部署。每个组件在该栈中扮演着独特的角色。

英文摘要

In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full robot learning stack: data collection, model design, continued pre-training and supervised fine-tuning, RL post-training, and real-world deployment. Each component serves a distinct role in this stack.

URL PDF HTML 收藏
2606.12995 2026-07-21 cs.RO 版本更新

GenHOI: Contact-Aware Humanoid-Object Interaction by Imitating Generated Videos without Task-Specific Training

GenHOI: 通过模仿生成视频实现接触感知的人形机器人-物体交互,无需任务特定训练

Zhihai Bi, Qiang Zhang, Guoyang Zhao, Jiahang Cao, Xueyin Luo, Yushan Zhang, Jinglan Xu, Ruoyu Geng, Yulin Li, Andrew F. Luo, Jun Ma

机构 * The University of Tokyo(东京大学) National University of Singapore(新加坡国立大学) University of California, Los Angeles(加州大学洛杉矶分校) Tsinghua University(清华大学)

AI总结 提出GenHOI框架,通过模仿单个生成视频实现人形机器人零样本执行多种物体交互任务,无需任务特定训练或物理演示数据,利用接触事件和手-物接触区域编码为几何约束优化轨迹。

详情
AI中文摘要

人形机器人-物体交互(HOI)是人形机器人的基本能力,但由于动态平衡与与多样物体稳定交互之间的紧密耦合,它仍然具有挑战性。现有方法通常需要耗时的任务特定策略训练或依赖于刚性轨迹回放,这限制了它们适应新颖交互场景的能力。在这项工作中,我们提出了\textit{GenHOI},一个简单而有效的框架,通过直接模仿单个生成视频,使人类形机器人能够以零样本方式执行多样化的物体交互任务,无需任务特定训练或物理演示数据。GenHOI首先在仿真中重建机器人-物体场景并渲染第一帧图像,该图像与语言命令一起条件化任务导向交互视频的合成。然后分析生成的视频以识别交互相关的接触事件并估计手-物体接触区域,这些被编码为以物体为中心的几何约束,将视觉交互线索转化为物理基础的优化先验。在这些先验的指导下,从视频中恢复的参考运动被细化和平滑,以解决2D视频生成中固有的尺度模糊性,同时将单个参考轨迹适应于未见过的机器人-物体相对姿态。优化后的轨迹最终由闭环跟踪控制器执行。我们在包括箱子抓取、非对称双臂椅子搬运、从下方抬桌子和圆柱物体包裹在内的多样化物体交互任务中,通过大量仿真和真实世界实验验证了所提出的框架。

英文摘要

Humanoid-Object Interaction (HOI) is a fundamental capability for humanoid robots, yet it remains challenging due to the tight coupling between dynamic balance and stable interaction with diverse objects. Existing methods often require time-consuming task-specific policy training or rely on rigid trajectory replay, which limits their ability to accommodate novel interaction scenarios. In this work, we present \textit{GenHOI}, a simple yet effective framework that enables humanoid robots to perform diverse object-interaction tasks in a zero-shot manner by directly imitating a single generated video, without task-specific training or physical demonstration data. GenHOI first reconstructs the robot-object scene in simulation and renders a first-frame image, which, together with the language command, conditions the synthesis of a task-oriented interaction video. The generated video is then analyzed to identify interaction-relevant contact events and estimate hand-object contact regions, which are encoded as object-centric geometric constraints that convert visual interaction cues into physically grounded optimization priors. Guided by these priors, the reference motion recovered from the video is refined and smoothed to resolve the scale ambiguity inherent in 2D video generation, while adapting a single reference trajectory to unseen robot-object relative poses. The optimized trajectory is finally executed by a closed-loop tracking controller. We validate the proposed framework in extensive simulation and real-world experiments across diverse object-interaction tasks, including box grasping, asymmetric bimanual chair carrying, table lifting from below, and cylindrical-object enveloping.

URL PDF HTML 收藏
2606.09079 2026-07-21 cs.LG cs.AI 版本更新

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

FlashMemory-DeepSeek-V4: 通过前瞻稀疏注意力实现闪电索引超长上下文

Yan Wang, Qifan Zhang, Jiachen Yu, Tian Liang, Dongyang Ma, Xiang Hu, Zibo Lin, Chunyang Li, Zhichao Wang, Miao Peng, Nuo Chen, Jia Li, Yujiu Yang, Haitao Mi, Dong Yu

机构 * Independent Researchers(独立研究者) Tencent(腾讯) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tsinghua University(清华大学)

AI总结 提出前瞻稀疏注意力(LSA),基于DeepSeek-V4架构的神经记忆索引器,通过预测未来上下文需求仅保留关键KV块,在超长上下文场景下将物理KV缓存压缩至全上下文的13.5%,同时保持或略微提升下游准确率。

Comments Technical report. 11 pages. Code and model available at https://github.com/libertywing/FlashMemory-Deepseek-V4 and https://huggingface.co/libertywing/FlashMemory-Deepseek-V4

详情
AI中文摘要

传统大语言模型在解码过程中保持完整的KV缓存,导致超长上下文服务出现严重的GPU内存瓶颈。在本报告中,我们提出前瞻稀疏注意力(LSA),一种基于DeepSeek-V4架构构建的神经记忆索引器驱动的新型推理范式。LSA并非被动地关注所有历史令牌,而是主动预测未来的上下文需求,并仅在GPU内存中保留查询关键的KV块。关键的是,我们通过无骨干的解耦训练策略实例化该架构。通过将索引器制定为标准双编码器架构,我们使用标准检索训练框架独立训练它,而无需将庞大的骨干模型加载到GPU内存中。我们证明这种“少即是多”的范式显著最大化服务效率,同时在依赖长期全局记忆的任务中充当有效的注意力去噪器。在主要的长上下文评估套件(例如LongBench-v2、LongMemEval和RULER)中,FM-DS-V4将平均物理KV缓存占用压缩至全上下文基线的仅13.5%,同时一致地保持或略微提升下游准确率(平均绝对边际+0.6%)。关键的是,在极端500K规模下,FlashMemory将物理KV缓存开销抑制超过90%,而不会破坏骨干的核心推理能力。

英文摘要

Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In this report, we propose \textbf{Lookahead Sparse Attention (LSA)}, a novel inference paradigm powered by a Neural Memory Indexer built upon the DeepSeek-V4 architecture. Rather than passively attending to all historical tokens, LSA proactively predicts future context demands and preserves only the query-critical KV chunks in the GPU memory. Crucially, we instantiate this architecture via a \textbf{backbone-free decoupled training} strategy. By formulating the indexer as a standard dual-encoder architecture, we train it independently using standard retrieval training frameworks without ever loading the massive backbone model into GPU memory. We demonstrate that this ``less is more'' paradigm significantly maximizes serving efficiency while acting as an effective attention denoiser in tasks that rely on long-term global memory. Across primary long-context evaluation suites (e.g., LongBench-v2, LongMemEval, and RULER), \texttt{FM-DS-V4} compresses the average physical KV cache footprint down to merely 13.5\% of the full-context baseline, while consistently preserving or slightly elevating downstream accuracy (+0.6\% absolute margin on average). At 1M context, per-decode-token compute drops to 0.30$\times$ of the baseline and GPU KV cache shrinks by 90\% (3.73$\to$0.37 GB), translating into \textbf{2.8$\times$ aggregate throughput and 2.7$\times$ concurrency gains} in PD-disaggregated serving on 8$\times$H20 GPUs.

URL PDF HTML 收藏
2511.08423 2026-07-21 cs.CV 版本更新

OmniAID: Decoupling Semantics and Artifacts for Universal AI-Generated Image Detection in the Wild

OmniAID: 解耦语义与伪影以实现通用AI生成图像野外检测

Yuncheng Guo, Junyan Ye, Chenjue Zhang, Hengrui Kang, Haohuan Fu, Conghui He, Weijia Li

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Sun Yat-Sen University(中山大学) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 提出OmniAID框架,通过解耦混合专家架构分离语义缺陷和通用伪影,结合两阶段训练策略和Mirage数据集,实现跨生成模型和语义内容的鲁棒AI生成图像检测。

Comments Accepted by ICML 2026

详情
AI中文摘要

一个真正通用的AI生成图像(AIGI)检测器必须同时泛化到多种生成模型和不同的语义内容。当前方法学习单一的、纠缠的伪造表示,混淆了内容相关的缺陷与内容无关的伪影,并进一步受到过时基准的限制。我们提出OmniAID,一种以解耦混合专家(MoE)架构为核心的新框架,该架构分离了:(1)通过可路由的专门语义专家在不同内容领域中的语义缺陷,以及(2)通过固定的通用伪影专家从内容相关缺陷中分离出内容无关的通用伪影。两阶段训练策略首先通过领域特定的困难采样独立专门化专家,然后训练一个轻量级门控网络以实现有效的输入路由。通过明确解耦“生成了什么”(内容特定缺陷)与“如何生成”(通用伪影),OmniAID实现了鲁棒的泛化。我们还引入了Mirage,一个大规模、当代的数据集,包含现代训练集和具有挑战性的测试集。大量实验表明,OmniAID超越了现有检测器,为针对现代野外威胁的AIGI检测建立了新标准。代码可在https://github.com/yunncheng/OmniAID获取。

英文摘要

A truly universal AI-Generated Image (AIGI) detector must simultaneously generalize across diverse generative models and varied semantic content. Current methods learn a single, entangled forgery representation, conflating content-dependent flaws with content-agnostic artifacts, and are further constrained by outdated benchmarks. We propose OmniAID, a novel framework centered on a decoupled Mixture-of-Experts (MoE) architecture that separates: (1) semantic flaws across distinct content domains via Routable Specialized Semantic Experts, and (2) content-agnostic universal artifacts from content-dependent flaws via a Fixed Universal Artifact Expert. A two-stage training strategy first specializes experts independently with domain-specific hard-sampling, then trains a lightweight gating network for effective input routing. By explicitly decoupling "what is generated" (content-specific flaws) from "how it is generated" (universal artifacts), OmniAID achieves robust generalization. We also introduce Mirage, a large-scale, contemporary dataset comprising a modern training set and a challenging test set. Extensive experiments demonstrate that OmniAID surpasses existing detectors, establishing a new standard for AIGI detection against modern, in-the-wild threats. Code is available at https://github.com/yunncheng/OmniAID.

URL PDF HTML 收藏
2604.08523 2026-07-21 cs.CL cs.AI 版本更新

ClawBench: Can AI Agents Complete Everyday Online Tasks?

ClawBench:AI代理能否完成日常在线任务?

Yuxuan Zhang, Yubo Wang, Yipeng Zhu, Penghui Du, Junwen Miao, Xuan Lu, Zhuofeng Li, Xingwei Qu, Zhengkang Guo, Yuanzhe Shen, Dingjie Song, Han Zhou, Tuney Zheng, Xian Wu, Hao Yu, Songcheng Cai, Yi Lu, Yunzhuo Hao, Minyi Lei, Liang Chen, Kai Zou, Huifeng Yin, Wendong Xu, Dongfu Jiang, Ping Nie, Jiaheng Liu, Wenhu Chen, Kelsey R. Allen

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute(向量研究所) Etude AI Carnegie Mellon University(卡内基梅隆大学) University of Waterloo(滑铁卢大学) Shanghai Jiao Tong University(上海交通大学) UniPat AI Zhejiang University(浙江大学) HKUST(香港科技大学) Tsinghua University(清华大学)

AI总结 ClawBench通过153个日常任务测试AI代理能力,涵盖15类144个平台,挑战多步骤流程和复杂操作,揭示现有模型在真实环境中的局限性。

Comments Project page: https://claw-bench.com

详情
AI中文摘要

AI代理可能能自动化邮箱,但能否自动化其他日常任务?日常在线任务为评估下一代AI代理提供了现实且未解的测试环境。我们引入ClawBench,一个包含153个简单任务的评估框架,涵盖144个活跃平台,从完成购买和预约到提交工作申请。这些任务需要超越现有基准的能力,如从用户提供的文档中获取信息、跨不同平台的多步骤流程导航以及填写大量详细表单。不同于现有在离线沙盒中评估的基准,ClawBench在生产网站上运行,保留了真实世界网页交互的全部复杂性、动态性和挑战。轻量级拦截层仅捕获并阻止最终提交请求,确保安全评估无实际影响。对7个前沿模型的评估显示,无论是专有还是开源模型,只能完成少量任务。例如,Claude Sonnet 4.6仅完成33.3%。ClawBench的进步使我们更接近于能够作为可靠通用助手的AI代理。

英文摘要

AI agents may be able to assist with emails and documents, but can they reliably complete everyday online workflows on real websites? Everyday online tasks offer a realistic yet unsolved testbed for evaluating the next generation of AI agents. To this end, we introduce ClawBench, an evaluation framework comprising 153 everyday online tasks that people need to accomplish regularly in their lives and work, spanning 144 platforms across 15 categories, from completing purchases and booking appointments to submitting job applications. These tasks require capabilities beyond existing benchmarks, such as obtaining relevant information from user-provided documents, navigating multi-step workflows across diverse platforms, and write-heavy operations like filling in many detailed forms correctly. Unlike existing benchmarks that evaluate agents in offline sandboxes with static pages, ClawBench operates on production websites, preserving the full complexity, dynamic nature, and interaction challenges of real-world web environments. An interception layer captures and blocks the final submission request, ensuring safe evaluation without real-world side effects. Our evaluations of 8 frontier models show that both proprietary and open-source models complete only a small portion of these tasks. For example, Claude Sonnet 4.6 achieves only 33.3%, which exposes gaps in current AI agents. Progress on ClawBench brings us closer to AI agents that can function as general-purpose assistants.

URL PDF HTML 收藏
2603.26648 2026-07-21 cs.SE cs.AI 版本更新

Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification

Vision2Web: 一个用于视觉网站开发的分层基准,包含代理验证

Zehai He, Wenyi Hong, Zhen Yang, Ziyang Pan, Mingdao Liu, Xiaotao Gu, Jie Tang

机构 * Tsinghua University(清华大学)

AI总结 本文提出Vision2Web基准,用于评估视觉网站开发能力,包含193个任务和16类,通过GUI代理验证器和基于VLM的裁判器评估多个视觉语言模型,揭示了在全栈开发中现有模型的性能差距。

详情
AI中文摘要

近年来,大型语言模型的进步提升了编码代理的能力,但对复杂、端到端网站开发的系统评估仍有限。为解决这一差距,我们引入了Vision2Web,一个用于视觉网站开发的分层基准,涵盖从静态UI到代码生成、交互多页面前端复现到长周期全栈网站开发。该基准由真实世界网站构建,包含193个任务,共16类,918个原型图像和1,255个测试用例。为支持灵活、全面和可靠的评估,我们提出基于工作流的代理验证范式,基于两个互补组件:GUI代理验证器和基于VLM的裁判器。我们评估了多个在不同编码代理框架下实例化的视觉语言模型,揭示了在所有任务级别上存在显著的性能差距,最先进的模型在全栈开发上仍面临挑战。

英文摘要

Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end website development remains limited. To address this gap, we introduce Vision2Web, a hierarchical benchmark for visual website development, spanning from static UI-to-code generation, interactive multi-page frontend reproduction, to long-horizon full-stack website development. The benchmark is constructed from real-world websites and comprises a total of 193 tasks across 16 categories, with 918 prototype images and 1,255 test cases. To support flexible, thorough and reliable evaluation, we propose workflow-based agent verification paradigm based on two complementary components: a GUI agent verifier and a VLM-based judge. We evaluate multiple visual language models instantiated under different coding-agent frameworks, revealing substantial performance gaps at all task levels, with state-of-the-art models still struggling on full-stack development.

URL PDF HTML 收藏
2602.21492 2026-07-21 cs.LG cs.AI cs.CL 版本更新

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning

GradAlign: 用于大语言模型强化学习的梯度对齐数据选择

Ningyuan Yang, Weihua Du, Weiwei Sun, Sean Welleck, Yiming Yang

机构 * Institute for Interdisciplinary Information Sciences (IIIS), Tsinghua University(交叉信息学院(IIIS)、清华大学) Language Technologies Institute (LTI), Carnegie Mellon University(语言技术研究所(LTI)、卡内基梅隆大学)

AI总结 GradAlign通过梯度对齐数据选择方法提升大语言模型强化学习的训练稳定性与性能。

Comments 20 pages. Accepted by COLM 2026

详情
AI中文摘要

强化学习(RL)已成为大语言模型(LLMs)后训练阶段的核心范式,但其性能对训练问题的质量非常敏感。这种敏感性源于RL的非平稳性:回放是由演进的策略生成的,学习受到探索和奖励反馈的影响,不同于具有固定轨迹的监督微调(SFT)。因此,先前的工作常常依赖于人工编目或简单的启发式过滤器(例如准确性),这可能会允许错误或低效用的问题。我们提出了GradAlign,这是一种用于LLM强化学习的梯度对齐数据选择方法,它使用一个小而可信的验证集来优先选择那些策略梯度与验证梯度对齐的训练问题,从而产生一个适应性的课程。我们评估了GradAlign在三个具有挑战性的数据领域:不可靠的奖励信号、分布不平衡和低效用的训练语料库,显示GradAlign在所有情况下都优于现有基线,突显了方向梯度信号在导航非平稳策略优化中的重要性,并产生了更稳定的训练和改进的最终性能。我们在此处发布我们的实现:https://github.com/StigLidu/GradAlign

英文摘要

Reinforcement learning (RL) has become a central post-training paradigm for large language models (LLMs), but its performance is highly sensitive to the quality of training problems. This sensitivity stems from the non-stationarity of RL: rollouts are generated by an evolving policy, and learning is shaped by exploration and reward feedback, unlike supervised fine-tuning (SFT) with fixed trajectories. As a result, prior work often relies on manual curation or simple heuristic filters (e.g., accuracy), which can admit incorrect or low-utility problems. We propose GradAlign, a gradient-aligned data selection method for LLM reinforcement learning that uses a small, trusted validation set to prioritize training problems whose policy gradients align with validation gradients, yielding an adaptive curriculum. We evaluate GradAlign across three challenging data regimes: unreliable reward signals, distribution imbalance, and low-utility training corpus, showing that GradAlign consistently outperforms existing baselines, underscoring the importance of directional gradient signals in navigating non-stationary policy optimization and yielding more stable training and improved final performance. We release our implementation at https://github.com/StigLidu/GradAlign

URL PDF HTML 收藏