arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

The Hong Kong University of Science and Technology(香港科技大学)

至 收录 2743
2607.18217 2026-07-21 cs.CV 新提交

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement

HOMIE:通过多模态智能增强实现以人为对象的视频个性化

Yiyang Cai, Nan Chen, Rongchang Xie, Junwen Pan, Chunyang Jiang, Cheng Chen, Wen Zhou, Zhenbang Sun, Wei Xue, Wenhan Luo, Yike Guo

机构 * Hong Kong University of Science and Technology(香港科技大学)

AI总结 研究以人为对象的视频个性化问题,提出HOMIE框架,统一处理主体间和主体内输入设置。通过更好的MLLM集成策略、自注意力中的全局多模态引导及模态参考嵌入,在多任务中达最优性能。

Comments 28 pages, 14 figures

详情
AI中文摘要

以人为对象的视频个性化(HOCVP)是主题驱动视频生成中的核心任务。现有方法存在两个关键局限。多数关注主体间个性化的方法难以在高主体保真度与人和多样物体间准确交互模式间平衡,尤其物体为抽象概念如图标时。虽主体内参考有望增强保真度,但多数现有工作缺乏理解潜在对应关系的机制。为应对这些挑战,我们提出HOMIE框架,统一处理主体间和主体内输入设置。与先前方法相比,HOMIE提出更好的MLLM集成策略,在不影响文本编码器可控性或产生高成本重新对齐的情况下提取参考级关系知识。具体而言,我们在自注意力中引入全局多模态引导,更好地对齐MLLM语义特征与VAE令牌;还提出模态参考嵌入来区分MLLM特征和VAE令牌并关联主体内参考图像令牌。大量实验验证了该方法在各种HOCVP任务中达到了当前最优性能。

英文摘要

Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on inter-subject personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and diverse objects, especially when objects represent abstract concepts such as logos. Second, while intra-subject references (e.g., OCR maps, multi-view inputs) are expected to enhance subject fidelity, most existing works lack mechanisms to understand such latent correspondence. To address both challenges, we propose HOMIE, an HOCVP framework that tackles both inter- and intra-subject input settings in a unified manner. Compared to previous approaches, HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment. Specifically, we introduce global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens. Furthermore, we propose modality-reference embedding to differentiate tokens from MLLM features and VAE tokens and associate intra-subject reference image tokens. Extensive experiments validate that our method achieves state-of-the-art performance across various HOCVP tasks. Project Page: https://yiyangcai.github.io/homie-page.github.io/

URL PDF HTML 收藏
2607.18181 2026-07-21 cs.CL cs.SE 新提交

VEHBench: A Stage-Local Diagnostic Benchmark for LLM-Assisted Vibration Energy Harvester Design

VEHBench:用于大语言模型辅助振动能量采集器设计的阶段局部诊断基准测试

Depeng Su, Yuyu Luo, Guobiao Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 研究针对无电池物联网中振动能量采集器设计,引入VEHBench基准测试,评估大语言模型在耦合物理设计不同阶段的表现,通过763个任务及四个设计角色测试,发现模型能力依赖阶段,为评估、选择等工程大语言模型提供阶段感知基础。

详情
AI中文摘要

无电池物联网需要在耦合物理约束下对振动能量采集器进行迭代设计,而大语言模型正成为工程工作流程的接口层。然而,现有的工程基准测试主要评估最终工件的有效性,对于大语言模型在耦合物理设计的不同阶段的表现提供的见解有限。我们引入了VEHBench,这是一个用于大语言模型辅助振动能量采集器设计的工程原生诊断基准测试,具有763个基于文献的任务,并由一个分析物理预言机评分。VEHBench评估四个设计角色:规范分类、验证器引导搜索、损坏状态恢复和策略条件选择。实验结果表明,大语言模型的能力强烈依赖于阶段:没有一个单一模型能在整个工作流程中始终占据主导地位,并且响应控制配置文件揭示了不同设计角色之间不同的行为模式。因此,VEHBench为评估、选择、路由和改进基于验证器的工程大语言模型提供了一个阶段感知基础。基准工件可在这个https URL上获取。

英文摘要

Battery-free Internet of Things (IoT) requires iterative design of vibration energy harvesters (VEHs) under coupled physical constraints, while LLMs are emerging as interface layers for engineering workflows. However, existing engineering benchmarks primarily assess final artifact validity, offering limited insights into how LLMs behave across different stages of coupled physical design. We introduce VEHBench, an engineering-native diagnostic benchmark for LLM-assisted VEH design, featuring 763 literature-grounded tasks scored by an analytical physical oracle. VEHBench evaluates four design roles: specification triage, verifier-guided search, corrupted-state recovery, and policy-conditioned selection. Experimental results reveal that LLM capability is strongly stage-dependent: no single model consistently dominates the entire workflow, and response-control profiles expose distinct behavioral patterns across design roles. VEHBench thus provides a stage-aware foundation for evaluating, selecting, routing, and improving verifier-grounded engineering LLMs. The benchmark artifact is available at https://huggingface.co/datasets/AnonymousVehbench/vehbench

URL PDF HTML 收藏
2607.17924 2026-07-21 cs.MA cs.LG 新提交

Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy Optimization

优势聚合而非比率聚合:合作多智能体策略优化的规范形式分析

Zijian Zhao, Sen Li

机构 * The Hong Kong University of Science and Technology(香港科技大学)

AI总结 研究合作多智能体策略优化中聚合相邻智能体的问题,将设计选择形式化为支持矩阵,证明规范结构,得出聚合应在优势中按耦合邻域大小进行且保持每个智能体比率的设计原则。

详情
AI中文摘要

以基于近端策略优化(PPO)的方法为例的多智能体策略优化,是合作多智能体强化学习(MARL)的关键分支。一个核心设计问题是聚合多少相邻智能体以有效利用全局信息进行合作。此决策须从优势(哪些智能体的奖励对信用信号有贡献)和比率(哪些智能体的似然比形成裁剪后的重要性权重)两个维度做出。现有方法在这两个轴上占据分散且未充分探索的点。我们将这两个设计选择形式化为支持矩阵$\SA$和$\SR$,并证明了一个规范结构:预期的多智能体策略优化目标仅通过它们的矩阵乘积$\tS=\SR\SA$依赖于对$(\SA,\SR)$。这产生了两个关键结果:(i)冗余性:两个支持矩阵在信号方面是可互换的,意味着没有一种聚合模式本质上更优越。(ii)方差排序:优势以和的形式聚合奖励(在耦合邻域处具有内部偏差 - 方差最优的加性方差),而比率以乘积的形式聚合似然比(随着支持大小呈指数增长的乘性方差,且没有伴随的偏差减少)。由此产生的设计原则很明确:在优势中聚合邻居,其大小与耦合邻域匹配,并保持每个智能体的比率。

英文摘要

Multi-agent policy optimization, exemplified by PPO-based methods, is a key branch of cooperative Multi-Agent Reinforcement Learning (MARL). A central design question is how many neighboring agents\footnote{In this paper, "neighbors" refer not only to physical proximity but also to agents whose actions influence one another.} to aggregate in order to effectively utilize global information for cooperation. This decision must be made along two dimensions: in the advantage (which agents' rewards contribute to the credit signal) and in the ratio (which agents' likelihood ratios form the clipped importance weight). Existing methods occupy scattered, underexplored points on these two axes: IPPO treats both separately; MAPPO pairs a team-level advantage with per-agent ratios; HAPPO employs sequential ratios with per-agent advantages; and single-agent reductions operating on factorized joint policies aggregate both into fully joint products. We formalize these two design choices as support matrices $\SA$ and $\SR$, and prove a canonical structure: the expected multi-agent policy optimization objective depends on the pair $(\SA,\SR)$ only through their matrix product $\tS=\SR\SA$. This yields two key consequences: (i) Redundancy: the two support matrices are interchangeable with respect to the signal, meaning neither aggregation pattern is inherently superior.(ii) Variance Ordering: the advantage aggregates rewards as a sum (additive variance with an interior bias-variance optimum at the coupling neighborhood), whereas the ratio aggregates likelihood ratios as a product (multiplicative variance that grows exponentially with support size, with no accompanying bias reduction). The resulting design principle is unambiguous: aggregate neighbors in the advantage, sized to the coupling neighborhood, and keep the ratio per-agent.

URL PDF HTML 收藏
2607.17780 2026-07-21 cs.PL cs.AI cs.LG cs.MA 新提交

ETAS: An Effect-Typed Language for Agent Systems

ETAS:一种用于智能体系统的效果类型化语言

Huiri Tan, Yikun Wang, Puyang Zhang, Shangyu Li, Jiasi Shen

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

AI总结 研究针对智能体系统设计ETAS语言,核心方法是通过规范一致性分配类型并用行为索引跟踪计算,主要贡献是为智能体执行相关推理提供编程语言基础,涵盖授权、非确定性等方面,还实现了该语言并形式化相关性质。

详情
AI中文摘要

ETAS是一种用于智能体系统的编程语言,它将模型支持的智能体、工具调用、提示、类型化内存、人工审批、策略和执行跟踪视为语义程序元素,而非库约定。它将确定性计算与智能体的非确定性和外部可见动作分离,同时保持直接的编程风格。本文介绍了ETAS的核心设计。其静态语义通过规范一致性分配普通类型,并用两个行为索引跟踪每个计算:一个逃逸效果行和它可能请求的类型化动作跟踪的持久抽象。规范形成一个终止的编译时约束演算。动态语义区分请求、处理、拒绝和提交的事件。还形式化了核心演算和状态保存等性质,并在Rust中实现了ETAS。ETAS为智能体执行前和执行期间的授权、非确定性、恢复和审计证据推理提供了编程语言基础。

英文摘要

ETAS is a programming language for agent systems that treats model-backed agents, tool calls, prompts, typed memory, human approvals, policies, and execution traces as semantic program elements rather than library conventions. It separates deterministic computation from agentic nondeterminism and externally visible actions while preserving a direct programming style. We present the core design of ETAS. Its static semantics assigns ordinary types through spec conformance and tracks each computation with two behavioral indices: an escaping effect row and a persistent abstraction of the typed action trace it may request. Specs form a terminating compile-time constraint calculus: type specs provide evidence for polymorphism and resource facts, callable specs constrain function and stage shapes, and trace specs express allow, deny, and temporal constraints. Typing checks requested traces against compiled monitors and emits residual obligations when dynamic resources preclude a complete static proof. The dynamic semantics distinguish requested, handled, denied, and committed events; handlers interpret typed actions without making their requests invisible to authorization or audit. We formalize a core calculus and state preservation, progress, type/effect soundness, handler trace-transparency, and policy safety. We also implement ETAS in Rust with a command-line interface, typed HIR checks, effect and policy diagnostics, handler checks, and trace-aware execution hooks. ETAS provides a programming-language foundation for reasoning about authorization, nondeterminism, recovery, and audit evidence before and during agent execution.

URL PDF HTML 收藏
2607.17291 2026-07-21 cs.LG cs.CL cs.IR 新提交

DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments

DRNOISE:在误导性证据环境中对深度研究代理进行基准测试

Jun Nie, Zhiqin Yang, Zhenheng Tang, Yonggang Zhang, Xiaowen Chu, Xinmei Tian, Bo Han

机构 * Hong Kong Baptist University(香港浸会大学) University of Science and Technology of China(中国科学技术大学) The Hong Kong University of Science and Technology(香港科技大学)

AI总结 研究在误导性证据环境下深度研究代理的表现,引入DRNOISE基准,每个任务含正确答案及冲突文档,涵盖多类证据操作。测试发现代理存在验证惰性,通用提示可缩小差距,强调可靠深度研究需积极协调主张与证据。

Comments 16 pages, 2 figures, 11 tables

详情
AI中文摘要

深度研究代理越来越多地在开放网络上运行,相关记录与冗余摘要、过时报告和误导性文档共存。现有评估对于当一份看似普通的虚假文档被故意植入可搜索环境并提供与冲突答案的直接捷径时,代理是否能保持合理的证据标准了解有限。我们引入了DRNOISE,一个用于在误导性证据下答案恢复的100任务基准。每个任务都有一个由两个相互佐证的间接记录链支持的唯一正确答案;配对的噪声条件添加了一个直接给出冲突答案的似是而非的文档。该基准涵盖十个证据操作类别。在具有强大的干净任务性能的代理中,这种单一干预导致准确率下降66 - 88个百分点。跟踪分析确定验证惰性是主要的失败模式:代理经常检索到真实记录,但在完成和协调证据链之前就停止了,而是听从类似答案的文档。通用验证提示缩小了但并未消除这一差距。该设置与开放网络部署特别相关,在开放网络中似是而非的虚假信息通过看似普通的页面而非明确攻击出现。因此,可靠的深度研究不仅需要检索和引用;还需要将直接主张与记录级证据进行积极协调。

英文摘要

Deep research agents increasingly operate over the open web, where relevant records coexist with redundant summaries, outdated reports, and misleading documents. Existing evaluations offer limited insight into whether agents preserve sound evidential standards when an ordinary-looking false document is deliberately seeded into a searchable environment and offers a direct shortcut to a conflicting answer. We introduce DRNOISE, a 100-task benchmark for answer recovery under misleading evidence. Each task has a unique gold answer supported by two corroborating indirect record chains; the paired noisy condition adds one plausible document that states a conflicting answer directly. The benchmark spans ten families of evidence operations. Across agents with strong clean-task performance, this single intervention causes 66-88 percentage-point accuracy drops. Trace analyses identify verification inertia as the dominant failure mode: agents often retrieve truthful records but stop before completing and reconciling the evidence chain, instead deferring to the answer-like document. Generic verification prompts reduce but do not close this gap. The setting is especially relevant to open-web deployment, where plausible falsehoods arrive through ordinary-looking pages rather than explicit attacks. Reliable deep research therefore requires more than retrieval and citation; it requires active reconciliation of direct claims with record-level evidence.

URL PDF HTML 收藏
2607.17262 2026-07-21 cs.CL 新提交

Should Missing Modalities Always Be Necessary to Repair for Multi-modal Sentiment Analysis?

多模态情感分析中缺失模态总是需要修复吗?

Yubo Gao, Haotian Wu, Xiaoyu Xu, Yibo Yan, Hong Chen, Ruoshui Peng, Fei Pan, Puay Siew Tan, Zhuoran Gao, Yonghua Hei, Jie Zhang, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学) Nanyang Technological University(南洋理工大学) Singapore Institute of Manufacturing Technology, A*STAR(新加坡制造技术研究所,新加坡科技研究局) Lingnan University(岭南大学)

AI总结 研究多模态情感分析中缺失模态是否总需修复,提出SIEVE方法,通过比较直接预测与修复分支,从样本损失差距得经验充分性信号,经证据门路由输入,与修复无关,实验证明其能改进修复主干并接近最优值。

详情
AI中文摘要

现有的多模态情感分析(MSA)中缺失模态的方法通常遵循先修复的范式。我们重新审视这一假设并提出疑问:每个缺失模态都应该被修复吗?样本级的神谕分析表明并非总是如此:全模态输入仅对一小部分样本最优,某些样本偏好每个模态子集。这表明添加或修复模态不一定总能改善预测,且每个模态的效用取决于样本。基于此,我们提出SIEVE,将‘是否修复’转化为样本级可学习的决策。它比较直接预测分支和修复分支,从样本损失差距中得出经验充分性信号,通过证据门路由输入。SIEVE与修复无关,可在任何修复模块之上即插即用。在CMU - MOSI和IEMOCAP上的实验表明,SIEVE在不同缺失率下持续改进代表性修复主干,并接近样本级双分支可实现的最优值。

英文摘要

Existing methods for multimodal sentiment analysis (MSA) under missing modalities usually follow a repair-first paradigm. We revisit this assumption and ask: \emph{should every missing modality be repaired?} A per-sample oracle analysis shows the answer is not always: full-modality input is optimal for only a small fraction of samples, and every modality subset is preferred by some samples. These results suggest that adding or repairing modalities may not always improve prediction, and that the utility of each modality is sample-dependent. Building on this finding, we propose \textbf{S}ufficiency-\textbf{I}nformed \textbf{E}vidential \textbf{V}al\textbf{vE} (\textbf{SIEVE}) that turns ``whether to repair'' into an explicit, learnable decision at the sample level. SIEVE compares a direct prediction branch with a repair branch, derives an empirical sufficiency signal from their per-sample loss gap, and routes each input through an evidential gate that jointly models sufficiency and its epistemic uncertainty. SIEVE is repair-agnostic: it operates as a plug-and-play decision on top of any explicit or implicit repair module, without modifying its internal design. Experiments on CMU-MOSI and IEMOCAP show that SIEVE consistently improves representative repair backbones across evaluated missing rates, and approaches the per-sample dual-branch achievable optimum.

URL PDF HTML 收藏
2607.17250 2026-07-21 cs.CL 新提交

EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World

EvolvingWorld:用于交互式文学世界中角色与世界模型协同进化的开放架构框架

Qing Zong, Yue Guo, Mengxin Yang, Yiwen Guo, Yangqiu Song

机构 * Hong Kong University of Science and Technology(香港科技大学) LIGHTSPEED(光速) Huazhong University of Science and Technology(华中科技大学)

AI总结 研究交互式文学世界中角色与世界协同进化问题,提出EvolvingWorld框架,含角色智能体和世界模型两个耦合模块,制定7个可训练任务,构建数据集并引入评估协议,实验证明其能有效改进长期模拟。

详情
AI中文摘要

本文介绍了EvolvingWorld,一个用于交互式文学世界中角色与世界协同进化的框架和基准。现有系统要么将交互式文学模拟视为静态角色模仿,要么视为孤立场景生成,无法捕捉角色和世界如何随时间共同进化。为解决此问题,EvolvingWorld将文学模拟建模为一个长期过程,其中角色互动、场景推进,且角色和世界状态持续更新。与依赖固定架构的先前系统不同,EvolvingWorld采用开放架构框架来支持跨不同文学世界的模拟。该框架由两个耦合模块组成:一个用于多角色角色扮演和持续角色档案进化的角色智能体,以及一个基于大语言模型的用于全局和位置/实体级状态维护及场景推进的世界模型。基于此架构,制定了7个可训练任务用于场景初始化、互动生成和状态更新。从57本书构建数据集,产生138596个监督训练样本和222个测试快照。还引入了涵盖10个维度和20个指标的轨迹级大语言模型评判评估协议。实验表明,EvolvingWorld能通过有效维持持续、连贯的角色和世界发展来改进长期模拟。

英文摘要

This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds. Existing systems either treat interactive literary simulation as static persona imitation or isolated scene generation, failing to capture how characters and worlds evolve together over time. To address this, EvolvingWorld models literary simulation as a long-horizon process where characters interact, scenes progress, and character and world states are persistently updated. Unlike prior systems relying on fixed schemas, EvolvingWorld adopts an open-schema framework to support simulation across diverse literary worlds. The framework consists of two coupled modules: a Character Agent for multi-character role-play and persistent profile evolution, and an LLM-based World Model for global and location/entity-level state maintenance and scene progression. Based on this architecture, we formulate 7 trainable tasks for scene initialization, interaction generation, and state update. We construct a dataset from 57 books, producing 138,596 supervised training samples and 222 snapshots for testing. Furthermore, we introduce a trajectory-level LLM-as-Judge evaluation protocol spanning 10 dimensions and 20 metrics. Experiments show that EvolvingWorld can improve long-horizon simulation by effectively maintaining persistent, coherent character and world development.

URL PDF HTML 收藏
2607.17166 2026-07-21 cs.LG 新提交

Explaining and Tuning Transformer-based LLMs in Arithmetic Tasks with Human Strategies

用人类策略解释和调整基于Transformer的语言模型在算术任务中的表现

Luyu Qiu, Jianing Li, Hwanhee Kim, Xiaoyong Wei, Yueyuan Zheng, Janet Hsiao, Lei Chen

机构 * The Hong Kong University of Science and Technology(香港科技大学) The Hong Kong Polytechnic University(香港理工大学) University of California, Berkeley(加州大学伯克利分校)

AI总结 研究基于Transformer的语言模型在算术任务中的表现,通过分解任务、分析损失收敛等,应用人类策略和方法提升其性能,并经多种验证展示有效性,探索其与人类学习者相似性以增强关键应用中的信任。

详情
AI中文摘要

基于Transformer的大型语言模型(LLMs)在各种自然语言处理任务中持续取得领先性能。然而,它们在基本算术等看似简单的问题上表现不佳,引发了对模型可靠性、安全性和道德部署的担忧。本研究表明,使用对人类学习者有效的方法可以提高在整数算术任务上训练的普通Transformer模型的性能。首先将算术任务分解为明确的子任务,并对每个子任务进行损失收敛阶分析和消融研究。发现LLMs呈现出与人类学习者相似的学习模式,简单子任务学习速度更快。此外,应用解决问题策略和认知强化方法提高了LLMs的准确性,这表明基于Transformer的LLMs在算术中可能与人类学习者共享认知过程。最后通过显著的准确性改进实验、可视化验证和基于解释的分析全面展示了方法的有效性,探索了基于Transformer的LLMs与人类学习者之间潜在的相似性,增强了对LLMs在关键和高风险应用中的信任。

英文摘要

Transformer-based large language models (LLMs) continue to achieve state-of-the-art performance across various natural language processing tasks. However, their subpar performance on seemingly elementary problems, such as basic arithmetic, raises concerns about model reliability, safety, and ethical deployment. In this study, we demonstrate that the performance of a vanilla Transformer model trained on integer arithmetic tasks can be improved using methods effective for human learners. We begin by decomposing the arithmetic task into well-defined subtasks and conducting loss convergence order analysis together with ablation studies for each subtask. Our findings reveal that LLMs exhibit learning patterns similar to those of human learners, with a faster learning speed for simpler subtasks compared to more complex ones. In addition, we successfully improved the accuracy of LLMs by applying problem-solving strategies and cognitive empowerment methods shown to enhance the performance of human learners. This suggests that transformer-based LLMs may share cognitive processes with human learners in arithmetic. Lastly, we provide a comprehensive demonstration of our method's effectiveness, including significant accuracy improvement experiments, visualization verification, and explanation-based analysis to illuminate the intricacies of LLMs in arithmetic learning. In general, this work explores the potential similarities between transformer-based LLMs and human learners, supported by explainable AI (XAI) verifications, ultimately fostering trust in LLMs for critical and high-stakes applications.

URL PDF HTML 收藏
2607.17099 2026-07-21 cs.CV cs.AI 新提交

DepthART: Scaling Foundation Monocular Depth to Tiny Models

DepthART:将基础单目深度模型扩展到小型模型

Feng Xue, Wu Chen, Mingshuai Zhao, Guofeng Zhong, Anlong Ming, Haozhe Wang, Dianqiao Lei, Zhaowen Lin, Haiyang Zhang, Nicu Sebe

机构 * University of Trento(特伦托大学) Beijing University of Posts and Telecommunications(北京邮电大学) The Hong Kong University of Science and Technology(香港科技大学) Tsinghua University(清华大学)

AI总结 研究旨在将基础单目深度模型扩展到小型模型,核心方法是结合抗偏差数据采样与相机条件微调策略的DepthART,主要贡献是在多数据集上取得更好的零样本泛化和度量精度,还提供了可扩展模型家族。

Comments Accepted in ACM Multimedia 2026 ; Code: https://github.com/xuefeng-cvr/DepthART ; Project page: https://xuefeng-cvr.github.io/DepthART ;

详情
AI中文摘要

近期的几何基础模型在跨场景泛化和度量尺度预测方面显著提升了单目深度估计(MDE),但这些进展未应用于小型模型。我们通过DepthART弥合这一差距,它是用于跨场景设备部署的紧凑型MDE模型。首先识别出小型模型中两个由容量驱动的瓶颈,然后结合两种有效策略:抗偏差数据采样方案和相机条件微调协议。在多个数据集上,DepthART在零样本泛化和度量精度上均超越先前小型基线,还提供了可扩展模型家族。

英文摘要

Recent geometric foundation models (e.g., Metric3D, Depth Anything and UniDepth) have substantially improved monocular depth estimation (MDE) in both cross-scene generalization and metric-scale prediction, yet these gains have not translated to tiny models. We bridge this gap with DepthART (Depth Anything Rethought for Tiny Models), which is a compact MDE model for on-device deployment across diverse scenes. We first identify two capacity-driven bottlenecks in tiny models: (i) overfitting to dataset-specific distribution bias and (ii) unstable metric adaptation under camera shift, where full fine-tuning easily damages transferable geometry. Accordingly, DepthART combines two simple but effective strategies: a bias-resistant data sampling scheme to reduce distribution bias under the same training budget, and a camera-conditioned fine-tuning protocol that freezes the distilled encoder and adjusts metric scale conditioned on intrinsics while better preserving cross-dataset generalization. Across datasets, DepthART consistently surpasses previous tiny baselines in both zero-shot generalization and metric accuracy (e.g., zero-shot $δ_1$=0.964 for DepthART-S on NYUD v2), and in some cases approaches heavy models. We further provide a scalable model family, with DepthART-S reaching 347/245 FPS (strict FP32) on an RTX A6000 at $224^2/448^2$, 102 FPS (TF32) on a Orin NX 8GB, and over 15 FPS (FP32) on a Jetson Nano 4GB.

URL PDF HTML 收藏
2607.17095 2026-07-21 cs.AI 新提交

Fourier Geometric Wind Power Forecasting with Numerical Weather Prediction

基于数值天气预报的傅里叶几何风力发电预测

Shiyuan Piao, Fan Zehui, Yang Liu, Hong Cheng, Juepeng Zheng, Jie Zhou, Fugee Tsung

机构 * The Hong Kong University of Science and Technology(香港科技大学) The Chinese University of Hong Kong(香港中文大学) Tsinghua University(清华大学) Goldwind Science and Technology Co.,Ltd(金风科技股份有限公司)

AI总结 研究针对风力发电预测难题,提出多模态框架,整合SCADA数据与NWP预测。先分解输入特征,再用几何编码器和傅里叶神经算子建模,实验表明该模型优于现有基线,凸显基于物理设计的有效性。

详情
AI中文摘要

准确的短期风力发电预测对电网稳定性和运营规划至关重要,但由于大气条件与涡轮机动力学之间的复杂相互作用,这一任务仍具有挑战性。现有方法未能有效整合天气预报与风力涡轮机数据(即SCADA),导致解决方案欠佳。为解决此问题,我们引入了一个多模态框架,将基于历史点的SCADA数据与基于网格的数值天气预报(NWP)预测相结合,这因异构输入和复杂的物理风力涡轮机相互作用而颇具挑战。我们的方法首先将输入明确分解为标量和矢量特征,以更好地捕捉特定地点和几何相关性,然后采用几何编码器从风矢量中提取旋转不变特征。我们还利用了傅里叶神经算子(FNO)架构,它在频域中执行全局卷积,以有效建模远程时空关系。在三个实际风电场进行的广泛实验表明,我们的模型始终优于现有基线,凸显了其基于物理的设计的有效性。我们方法的核心实现可在指定网址公开获取。

英文摘要

Accurate short-term wind power forecasting is essential for grid stability and operational planning, yet remains challenging due to the complex interactions between atmospheric conditions and turbine dynamics. However, existing methods fail to effectively incorporate weather forecasting with wind turbine data (i.e., SCADA), leading to suboptimal solutions. To address this, we introduce a multimodal framework that integrates historical point-based SCADA data with grid-based Numerical Weather Prediction (NWP) forecasts, which is challenging due to heterogeneous input and the complex physical wind-turbine interactions. Our approach first explicitly decomposes inputs into scalar and vector features to better capture both site-specific and geometric dependencies and then incorporates a geometric encoder to extract rotation-invariant features from wind vectors. We further leverages a Fourier Neural Operator (FNO) architecture, which performs global convolutions in the frequency domain to efficiently model long-range spatiotemporal relationships. Extensive experiments on three real-world wind farms, with weather forecasting data, demonstrate that our model consistently outperforms state-of-the-art baselines, highlighting the effectiveness of its physically-informed design. The core implementation of our method is publicly available at: https://github.com/shawn-sypiao/GWPF.

URL PDF HTML 收藏
2607.16806 2026-07-21 cs.RO 新提交

Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation

用于动态视觉语言导航的从慢速推理器到快速规划器的逐令牌潜在流

Tianshuai Hu, Yangyi Zhong, Zeying Gong, Lingdong Kong, Xiaodong Mei, Guoyang Zhao, Xiaolu Liu, Song Wang, Rong Li, Junwei Liang

机构 * The Hong Kong University of Science and Technology(香港科技大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) National University of Singapore(新加坡国立大学) Zhejiang University(浙江大学)

AI总结 针对动态视觉语言导航中语言推理慢与规划需即时的矛盾,提出SPARK-VLN双系统框架,通过三个模块将慢速VLM推理器知识流到快速规划器,引入新基准套件,提高了导航成功率、社会合规性及推理效率。

详情
AI中文摘要

在动态、以人类为中心的环境中的视觉语言导航存在一个基本矛盾:语言推理缓慢且深思熟虑,而安全、符合社会规范的规划应该即时且具有反应性。由此产生的观测陈旧性对安全至关重要:推理过程中选择的动作在执行时可能已经不安全。我们观察到,在VLM完成推理之前很久,其中间隐藏状态就已经编码了与动作相关的意图。我们提出了SPARK-VLN,这是一个用于动态社会VLN的双系统框架,在整个令牌生成过程中将慢速VLM推理器的知识流到快速流匹配专家规划器,在推理过程中提供新的和不断演变的指导。该设计由三个模块实现:一个逐令牌隐藏流提取器,一个序列到插槽潜在桥接器,一个不断演变的潜在调节器。我们还引入了一个用于动态社会视觉语言导航的以人类为中心的基准套件,该套件在整个推理过程中使行人和机器人保持活跃,并报告导航成功、社会合规性、人类碰撞和明确的陈旧性统计数据。在这些设置中,SPARK-VLN提高了导航成功率和社会合规性,并保持了推理效率。

英文摘要

Vision-Language Navigation in dynamic, human-centric environments exposes a fundamental tension: linguistic reasoning is slow and deliberative, whereas safe, socially compliant planning should be instant and reactive. The resulting observation staleness is safety-critical: a maneuver chosen during inference can already be unsafe by the time it executes. We observe that, long before a VLM finishes its inference, its intermediate hidden states already encode action-relevant intent. We propose SPARK-VLN, a dual-system framework for dynamic social VLN that streams the slow VLM reasoner's knowledge to a fast flow-matching expert planner throughout token generation, providing fresh and evolving guidance during inference. This design is realized by three modules: a Token-Wise Hidden Streamer that extracts intermediate hidden states along the token generation process, a Sequence-to-Slot Latent Bridge that projects them into fixed-size latent slots, and an Evolving Latent Conditioner that infuses them into the expert planner. We also introduce a human-centric benchmark suite for dynamic social vision-language navigation that keeps pedestrians and the robot active throughout inference and reports navigation success, social compliance, human collisions, and explicit staleness statistics. Across these settings, SPARK-VLN mproves navigation success and social compliance while sustaining inference efficiency. Webpage: https://hutslib.github.io/SPARK-VLN/.

URL PDF HTML 收藏
2607.16577 2026-07-21 cs.CV cs.GR 新提交

CNS-Edit++: Category-Agnostic 3D Editing with Coupled Neural Shape Representation

CNS-Edit++:基于耦合神经形状表示的类别无关3D编辑

Jingyu Hu, Weilong Yan, Zhengzhe Liu, Haipeng Li, Ka-Hei Hui, Hao, Zhang, Chi-Wing Fu

机构 * The Chinese University of Hong Kong(香港中文大学) Lingnan University(岭南大学) National University of Singapore(新加坡国立大学) The Hong Kong University of Science and Technology(香港科技大学) Autodesk AI Lab(欧特克人工智能实验室) Simon Fraser University(西蒙弗雷泽大学)

AI总结 研究提出基于耦合神经形状表示和神经特征体积优化的潜在空间3D形状编辑框架CNS-Edit++,能在特定类别和类别无关模型上实例化,有多种编辑操作符及区域控制机制,经评估其性能优于现有方法。

详情
AI中文摘要

本文提出了一个基于耦合神经形状(CNS)表示和神经特征体积优化的潜在空间3D形状编辑框架。该工作将基于耦合神经形状优化的CNS-Edit扩展到CNS-Edit++,通过将特定类别的耦合表示推广到使用基础模型的类别无关3D形状编辑。耦合神经形状(CNS)表示将捕获高级形状语义的全局潜在代码与为局部形状操作提供空间上下文的3D神经特征体积耦合。然后制定了一个耦合神经形状优化过程,以根据给定的编辑操作共同优化这两个组件。该框架可以在特定类别的3D反演模型和类别无关的3D基础模型上实例化。提供了各种形状编辑操作符,并引入两种互补的区域控制机制以保留编辑区域外的区域。不同3D生成模型的广泛定量和定性评估证明了该方法优于现有解决方案的强大能力。

英文摘要

This paper presents a latent-space 3D shape editing framework built upon a coupled neural shape (CNS) representation and a neural feature volume optimization. This work extends CNS-Edit, built on Coupled Neural Shape optimization, to CNS-Edit++, by generalizing the category-specific coupled representation to category-agnostic 3D shape editing with foundation models. The Coupled Neural Shape (CNS) representation couples a global latent code that captures high-level shape semantics with a 3D neural feature volume that provides spatial context for local shape manipulation. Then we formulate a coupled neural shape optimization procedure that co-optimizes these two components subject to a given editing operation. Our framework can be instantiated on both the category-specific 3D inversion model and category-agnostic 3D foundation models. We provide various shape editing operators, including copy, resize, delete, mix, point-wise drag, and region-wise drag, each of which is formulated as an objective to guide the CNS optimization. To preserve regions outside the editing area, we further introduce two complementary region-wise control mechanisms, i.e., KV-cache replacement and latent feature regularization. Extensive quantitative and qualitative evaluations across different 3D generative models demonstrate the strong capabilities of our approach over state-of-the-art solutions.

URL PDF HTML 收藏
2607.16354 2026-07-21 cs.LG cs.AI 新提交

A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting

基于少样本连续上下文博弈的预测-校正循环用于需求预测

Zhiwei Lei, Benedict Jun Ma, Ilya Jackson

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Massachusetts Institute of Technology(麻省理工学院)

AI总结 研究针对零售需求预测难题,提出预测-校正框架,运用少样本连续上下文博弈校正策略等,经实验在多需求模式下显著降低误差、提高RMSE并降低库存成本,证明在线预测校正能连接离线需求学习与实时零售决策。

详情
AI中文摘要

当需求变化速度超过静态预测模型的重新训练速度时,零售需求预测仍然困难,尤其是在新观察标签稀疏的早期需求周期。为解决此问题,本研究提出预测-校正(PtC)框架,保留第一阶段机器学习预测,并应用少样本连续上下文博弈校正策略及相似SKU增强和top-p掩码更新。通过沃尔玛零售数据和独家饮料数据集实验,PtC在多种需求模式下显著降低MAPE、MAE和RMSE,消融研究中平均RMSE比仅用机器学习基线提高9.52%,且库存成本更低。结果表明在线预测校正可通过适应稀疏反馈连接离线需求学习和实时零售决策,而无需完全重新训练基础预测模型。

英文摘要

Retail demand forecasting remains difficult when demand shifts faster than static forecasting models can be retrained, especially in early demand cycles where newly observed labels are sparse. To address this, this study aims to improve adaptive retail forecasting by proposing a predict-then-correct (PtC) framework that retains a first-stage machine learning (ML) forecast and applies a few-shot continuous contextual bandit correction policy with similar-SKUs augmentation and top-p masked updating. Across Walmart retail data and an exclusive beverage dataset, PtC delivers statistically significant reductions in MAPE, MAE, and RMSE across stable & high volume, stable & low volume, and erratic & intermittent demand patterns, improves average RMSE by 9.52% over the ML-only baseline in the ablation study, and yields lower inventory costs than base-stock, proximal policy optimization, and soft actor-critic policies under the tested lead-time settings. These findings show that online forecast correction can bridge offline demand learning and real-time retail decision-making by adapting to sparse feedback without fully retraining the base forecasting model.

URL PDF HTML 收藏
2607.16320 2026-07-21 cs.CV 新提交

The Devil is in the Dark Pixels: Toward Brightness Bias-Robust Denoising

魔鬼藏于暗像素中:迈向亮度偏差鲁棒去噪

Sungjun Cho, Zhuangzhuang Chen, Xiaomeng Li

机构 * The Hong Kong University of Science and Technology(香港科技大学)

AI总结 研究图像去噪中暗区域信噪比低、MSE训练去噪器加剧偏差的问题,提出亮度偏差鲁棒去噪(BBRD)方法,将像素分亮度带,归一化误差并应用Group-DRO,实验证明该方法能同时改善各亮度带,在暗区域增益最大。

Comments 6 figures, 4 tables. Code: https://github.com/xmed-lab/BBRD

详情
AI中文摘要

本文揭示了图像去噪中一个重要却被忽视的问题:在信号依赖的相机噪声模型下,暗区域的信噪比(SNR)天生较低,信号强度衰减比噪声方差减小快得多,使得暗区域细节恢复极具挑战。然而,基于均方误差(MSE)训练的去噪器非但没有弥补这一困难,反而加剧了它——重构暗像素比其每波段噪声本底差6倍。这种偏差源于两个因素:信号依赖噪声使亮像素残差膨胀,且网络的雅可比范数随亮度单调增加。为此,我们提出亮度偏差鲁棒去噪(BBRD),它是MSE损失的替代方法,将像素划分为亮度带,通过经验噪声方差对每带误差进行归一化,并应用组分布鲁棒优化(Group-DRO)动态加重当前最差的带,无需额外参数或推理成本。实验表明,BBRD是13种测试方法中唯一能同时改善各亮度带的方法,在暗带可达+0.45 dB,亮带可达+0.32 dB,在SIDD数据集上总峰值信噪比(PSNR)可达+0.65 dB,在最暗区域增益最大。代码可从该https链接获取。

英文摘要

In this paper, we reveal an important yet overlooked problem in image denoising: under signal-dependent camera noise models, dark regions suffer from inherently low Signal-to-Noise Ratio (SNR), as signal intensity decays far faster than noise variance diminishes, making detail recovery in dark areas fundamentally challenging. Yet rather than compensating for this difficulty, MSE-trained denoisers exacerbate it -- reconstructing dark pixels up to 6x worse relative to their per-band noise floor. This bias stems from two compounding factors: signal-dependent noise inflates bright-pixel residuals, and the network's Jacobian norm increases monotonically with brightness. Together, these cause bright regions to chronically dominate gradient updates at the expense of dark ones. To this end, we propose Brightness Bias-Robust Denoising (BBRD), a drop-in replacement for MSE loss that partitions pixels into brightness bands, normalizes per-band error by empirical noise variance, and applies Group Distributionally Robust Optimization (Group-DRO) to dynamically upweight whichever band is currently worst, with zero additional parameters or inference cost. Across 8 architectures and 2 datasets in our experiments, BBRD is the only method among 13 tested alternatives that improves each brightness band simultaneously, achieving up to +0.45 dB on dark bands, +0.32 dB on bright bands, and +0.65 dB aggregate Peak Signal-to-Noise Ratio (PSNR) on SIDD, with the largest per-band gains in the darkest regions where detail recovery matters most. Code is available at https://github.com/xmed-lab/BBRD

URL PDF HTML 收藏
2607.16251 2026-07-21 cs.LG 新提交

Learning Spatio-Temporal Foundation Models from Pure Synthetic Data

从纯合成数据中学习时空基础模型

Yutong Feng, Shiyuan Piao, Yutong Xia, Xu Liu, Wenqi Fan, Fugee Tsung, See-Kiong Ng, Yuxuan Liang

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) National University of Singapore(新加坡国立大学) Hong Kong University of Science and Technology(香港科技大学) Hong Kong Polytechnic University(香港理工大学)

AI总结 研究旨在学习时空基础模型,提出NeoST,通过在程序生成的合成系统上预训练,引入可扩展语料库、潜在空间推理架构和目标,实验证明其在多样真实世界时空系统中性能优越,有长期稳定性和推理效率。

详情
AI中文摘要

时空基础模型(STFMs)旨在学习复杂动力系统在时空上的可泛化表示。现有方法存在诸多问题,如真实世界预训练数据的分布偏差、自回归或基于扩散范式的结构瓶颈以及过度强调噪声观测中点状重建的目标。本文提出了NeoST,首个仅在程序生成的合成系统上预训练的时空基础模型。它引入可扩展合成预训练语料库减轻真实世界偏差,有潜在空间推理架构及潜在空间目标。实验表明NeoST在多样真实世界时空系统中优于现有模型,具有卓越的长期稳定性和推理效率。

英文摘要

Spatio-Temporal Foundation Models (STFMs) aim to learn generalizable representations of complex dynamical systems across space and time. However, existing approaches suffer from distributional bias in real-world pre-training data, structural bottlenecks of autoregressive or diffusion-based paradigms, and objectives that overemphasize point-wise reconstruction in noisy observation space.We propose \textbf{NeoST}, the first spatio-temporal foundation model pre-trained solely on procedurally generated synthetic systems. NeoST introduces a scalable synthetic pre-training corpus to mitigate real-world bias, a latent-space reasoning architecture that generates and iteratively refines multiple future trajectories without sequential error accumulation, and latent-space objectives that emphasize structural dynamics and enable inference-time correction under distribution shifts.Extensive experiments across diverse real-world benchmarks show that NeoST consistently outperforms existing STFMs in diverse real-world spatio-temporal systems, achieves superior long-horizon stability and inference efficiency.

URL PDF HTML 收藏
2607.04334 2026-07-21 cs.AI 版本更新

Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure

图形用户界面代理相信它们的眼睛吗?诊断状态信念对像素与结构的依赖

Guijia Zhang, Yuxun Chen, Yuheng Qi, Harry Yang

机构 * Shenzhen University(深圳大学) The Hong Kong University of Science and Technology(香港科技大学)

AI总结 研究多模态GUI代理状态信念来源,通过对310个真实探针进行单通道干预形式化视觉状态依赖并测量,核心指标是感知融合差距,发现文本状态信念依赖结构,图像精度高,错误会导致行动失败。

Comments 17 pages, 3 figures

详情
AI中文摘要

多模态GUI代理通过屏幕截图的渲染像素和序列化结构(如DOM或可访问性树)读取界面。现有基准测试不关注代理状态信念是否来自像素。我们形式化视觉状态依赖,通过对310个真实网络、移动和桌面探针的单通道干预进行测量。核心指标是感知融合差距,结果显示文本状态信念依赖结构,错误会导致行动失败。

英文摘要

Multimodal GUI agents read an interface through two redundant channels: the rendered pixels of a screenshot and a serialized structure such as a document object model or accessibility tree. Before acting, an agent forms a belief about the current interface state, but existing benchmarks score task success, element grounding, or attack resistance and do not ask whether that belief is drawn from the pixels. We formalize visual state reliance, the attribution of a state belief to pixels, structure, or priors, and measure it with paired single-channel interventions over 735 probes spanning real web, mobile, and desktop interfaces, of which 225 are zero-edit divergences mined from live production websites, all scored by deterministic forced choice with no model judge. Our central metric is the Perception-Fusion Gap (PFG), the fraction of probes a model perceives correctly yet resolves toward structure under conflict; a stricter variant that re-verifies perception on a tight crop of the target region leaves the gap intact. Across models from four vendors, textual state beliefs defer to structure while image-only accuracy stays near ceiling, and on unedited stale snapshots from live pages the same models follow the outdated structure on up to 0.88 of probes. A white-box ablation traces the textual effect to a single copied structural value, and gradient attribution shows the visual evidence is processed yet overridden. In live multi-step environments, one mis-sourced belief at the first step compounds into task failure with a self-recovery rate of at most 0.03. Comparing four mitigations on identical probes, prompt-level cues fail at the action level, certificate checks buy safety with refusals, and a training-free consistency gate is alone in reducing both hijack and task error. Visual state reliance thus gives a measurable diagnostic of whether agent state beliefs are visually grounded.

URL PDF HTML 收藏
2607.03920 2026-07-21 cs.RO 版本更新

LH-AVLN: A Benchmark for Long-Horizon Audio-Visual-Language Navigation

LH-AVLN:长距离视听语言导航基准测试

Rufeng Chen, Yue Chang, Zili Shao, Zhaofan Zhang, Li Chen, Hechang Chen, Hui Xiong, Sihong Xie

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Jilin University(吉林大学)

AI总结 介绍长距离视听语言导航基准LH-AVLN,结合多目标任务执行等。提出免训练参考智能体PAG-Nav,可维护语义地图并规划。实验表明现有智能体完成任务有困难,PAG-Nav提供更强诊断基线。

Comments 12 pages, 3 figures

详情
AI中文摘要

具身导航正朝着长距离任务发展,但现有长距离基准测试大多无声,视听导航任务通常聚焦单一目标。我们引入LH-AVLN,一个结合多目标任务执行、异构目标规范和持续空间声学线索的长距离视听语言导航基准测试。

英文摘要

Embodied navigation is moving toward long-horizon missions, yet existing long-horizon benchmarks are largely acoustically silent, and audio-visual navigation tasks typically focus on a single goal. We introduce LH-AVLN, a benchmark for Long-Horizon Audio-Visual-Language Navigation that combines multi-goal mission execution, heterogeneous goal specifications, and persistent spatialized acoustic cues. In LH-AVLN, an agent receives a global mission of two to four goals specified by category, language description, or reference image, and navigates with RGB-D observations, pose, and binaural audio in indoor 3D environments. The benchmark supports both ordered and unordered missions, where alternating goal-associated sounds can guide non-line-of-sight search but may also become distractors as mission progress changes. We further develop PAG-Nav, a training-free reference agent that maintains a temporal uniform semantic map and performs progressive goal-state planning, using sound for search while reserving completion for visual-semantic verification. Experiments show that existing vision-language, memory-based, and audio-visual agents struggle to complete full LH-AVLN missions, and that PAG-Nav provides a stronger diagnostic baseline while leaving substantial room for future progress.

URL PDF HTML 收藏
2606.09079 2026-07-21 cs.LG cs.AI 版本更新

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

FlashMemory-DeepSeek-V4: 通过前瞻稀疏注意力实现闪电索引超长上下文

Yan Wang, Qifan Zhang, Jiachen Yu, Tian Liang, Dongyang Ma, Xiang Hu, Zibo Lin, Chunyang Li, Zhichao Wang, Miao Peng, Nuo Chen, Jia Li, Yujiu Yang, Haitao Mi, Dong Yu

机构 * Independent Researchers(独立研究者) Tencent(腾讯) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tsinghua University(清华大学)

AI总结 提出前瞻稀疏注意力(LSA),基于DeepSeek-V4架构的神经记忆索引器,通过预测未来上下文需求仅保留关键KV块,在超长上下文场景下将物理KV缓存压缩至全上下文的13.5%,同时保持或略微提升下游准确率。

Comments Technical report. 11 pages. Code and model available at https://github.com/libertywing/FlashMemory-Deepseek-V4 and https://huggingface.co/libertywing/FlashMemory-Deepseek-V4

详情
AI中文摘要

传统大语言模型在解码过程中保持完整的KV缓存,导致超长上下文服务出现严重的GPU内存瓶颈。在本报告中,我们提出前瞻稀疏注意力(LSA),一种基于DeepSeek-V4架构构建的神经记忆索引器驱动的新型推理范式。LSA并非被动地关注所有历史令牌,而是主动预测未来的上下文需求,并仅在GPU内存中保留查询关键的KV块。关键的是,我们通过无骨干的解耦训练策略实例化该架构。通过将索引器制定为标准双编码器架构,我们使用标准检索训练框架独立训练它,而无需将庞大的骨干模型加载到GPU内存中。我们证明这种“少即是多”的范式显著最大化服务效率,同时在依赖长期全局记忆的任务中充当有效的注意力去噪器。在主要的长上下文评估套件(例如LongBench-v2、LongMemEval和RULER)中,FM-DS-V4将平均物理KV缓存占用压缩至全上下文基线的仅13.5%,同时一致地保持或略微提升下游准确率(平均绝对边际+0.6%)。关键的是,在极端500K规模下,FlashMemory将物理KV缓存开销抑制超过90%,而不会破坏骨干的核心推理能力。

英文摘要

Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In this report, we propose \textbf{Lookahead Sparse Attention (LSA)}, a novel inference paradigm powered by a Neural Memory Indexer built upon the DeepSeek-V4 architecture. Rather than passively attending to all historical tokens, LSA proactively predicts future context demands and preserves only the query-critical KV chunks in the GPU memory. Crucially, we instantiate this architecture via a \textbf{backbone-free decoupled training} strategy. By formulating the indexer as a standard dual-encoder architecture, we train it independently using standard retrieval training frameworks without ever loading the massive backbone model into GPU memory. We demonstrate that this ``less is more'' paradigm significantly maximizes serving efficiency while acting as an effective attention denoiser in tasks that rely on long-term global memory. Across primary long-context evaluation suites (e.g., LongBench-v2, LongMemEval, and RULER), \texttt{FM-DS-V4} compresses the average physical KV cache footprint down to merely 13.5\% of the full-context baseline, while consistently preserving or slightly elevating downstream accuracy (+0.6\% absolute margin on average). At 1M context, per-decode-token compute drops to 0.30$\times$ of the baseline and GPU KV cache shrinks by 90\% (3.73$\to$0.37 GB), translating into \textbf{2.8$\times$ aggregate throughput and 2.7$\times$ concurrency gains} in PD-disaggregated serving on 8$\times$H20 GPUs.

URL PDF HTML 收藏
2602.01167 2026-07-21 cs.AI

Do All Individual Layers Help? An Empirical Study of Task-Interfering Layers in Vision-Language Models

所有个体层都有帮助吗?视觉-语言模型中任务干扰层的实证研究

Zhiming Liu, Yujie Wei, Lei Feng, Xiu Su, Xiaobo Xia, Weili Guan, Zeke Xie, Shuo Yang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Harbin Institute of Technology(哈尔滨工业大学) Southeast University(东南大学) Central South University(中南大学) National University of Singapore(新加坡国立大学) The Hong Kong University of Science and Technology, Guangzhou(香港科学与技术大学(广州))

AI总结 研究通过层干预发现部分层阻碍下游任务,提出任务自适应层剔除方法提升性能,揭示预训练VLM的意外模块化特性。

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 2026, pp. 9597-9607

详情
AI中文摘要

当前VLM在多种多模态任务中表现出色,但默认启用所有层可能阻碍任务表现。通过干预单层参数发现,某些层反而抑制任务性能。系统研究各层对不同任务的影响,提出任务-层交互向量量化方法,并引入无需训练的测试时适应方法TaLo,动态剔除最干扰的层,提升模型在多个任务和数据集上的性能,包括提升Qwen-VL在ScienceQA地图任务上的准确率。

英文摘要

Current VLMs have demonstrated capabilities across a wide range of multimodal tasks. Typically, in a pretrained VLM, all layers are engaged by default to make predictions on downstream tasks. We find that intervening on a single layer, such as by zeroing its parameters, can improve the performance on certain tasks, indicating that some layers hinder rather than help downstream tasks. We systematically investigate how individual layers influence different tasks via layer intervention. Specifically, we measure the change in performance relative to the base model after intervening on each layer and observe improvements when bypassing specific layers. This improvement can be generalizable across models and datasets, indicating the presence of Task-Interfering Layers that harm downstream tasks' performance. We introduce Task-Layer Interaction Vector, which quantifies the effect of intervening on each layer of a VLM given a task. These task-interfering layers exhibit task-specific sensitivity patterns: tasks requiring similar capabilities show consistent response trends under layer interventions, as evidenced by the high similarity in their task-layer interaction vectors. Inspired by these findings, we propose TaLo (Task-Adaptive Layer Knockout), a training-free, test-time adaptation method that dynamically identifies and bypasses the most interfering layer for a given task. Without parameter updates, TaLo improves performance across various models and datasets, including boosting Qwen-VL's accuracy on the Maps task in ScienceQA by up to 16.6%. Our work reveals an unexpected form of modularity in pretrained VLMs and provides a plug-and-play, training-free mechanism to unlock hidden capabilities at inference time. The source code will be publicly available.

URL PDF HTML 收藏
2604.08523 2026-07-21 cs.CL cs.AI 版本更新

ClawBench: Can AI Agents Complete Everyday Online Tasks?

ClawBench:AI代理能否完成日常在线任务?

Yuxuan Zhang, Yubo Wang, Yipeng Zhu, Penghui Du, Junwen Miao, Xuan Lu, Zhuofeng Li, Xingwei Qu, Zhengkang Guo, Yuanzhe Shen, Dingjie Song, Han Zhou, Tuney Zheng, Xian Wu, Hao Yu, Songcheng Cai, Yi Lu, Yunzhuo Hao, Minyi Lei, Liang Chen, Kai Zou, Huifeng Yin, Wendong Xu, Dongfu Jiang, Ping Nie, Jiaheng Liu, Wenhu Chen, Kelsey R. Allen

机构 * University of British Columbia(不列颠哥伦比亚大学) Vector Institute(向量研究所) Etude AI Carnegie Mellon University(卡内基梅隆大学) University of Waterloo(滑铁卢大学) Shanghai Jiao Tong University(上海交通大学) UniPat AI Zhejiang University(浙江大学) HKUST(香港科技大学) Tsinghua University(清华大学)

AI总结 ClawBench通过153个日常任务测试AI代理能力,涵盖15类144个平台,挑战多步骤流程和复杂操作,揭示现有模型在真实环境中的局限性。

Comments Project page: https://claw-bench.com

详情
AI中文摘要

AI代理可能能自动化邮箱,但能否自动化其他日常任务?日常在线任务为评估下一代AI代理提供了现实且未解的测试环境。我们引入ClawBench,一个包含153个简单任务的评估框架,涵盖144个活跃平台,从完成购买和预约到提交工作申请。这些任务需要超越现有基准的能力,如从用户提供的文档中获取信息、跨不同平台的多步骤流程导航以及填写大量详细表单。不同于现有在离线沙盒中评估的基准,ClawBench在生产网站上运行,保留了真实世界网页交互的全部复杂性、动态性和挑战。轻量级拦截层仅捕获并阻止最终提交请求,确保安全评估无实际影响。对7个前沿模型的评估显示,无论是专有还是开源模型,只能完成少量任务。例如,Claude Sonnet 4.6仅完成33.3%。ClawBench的进步使我们更接近于能够作为可靠通用助手的AI代理。

英文摘要

AI agents may be able to assist with emails and documents, but can they reliably complete everyday online workflows on real websites? Everyday online tasks offer a realistic yet unsolved testbed for evaluating the next generation of AI agents. To this end, we introduce ClawBench, an evaluation framework comprising 153 everyday online tasks that people need to accomplish regularly in their lives and work, spanning 144 platforms across 15 categories, from completing purchases and booking appointments to submitting job applications. These tasks require capabilities beyond existing benchmarks, such as obtaining relevant information from user-provided documents, navigating multi-step workflows across diverse platforms, and write-heavy operations like filling in many detailed forms correctly. Unlike existing benchmarks that evaluate agents in offline sandboxes with static pages, ClawBench operates on production websites, preserving the full complexity, dynamic nature, and interaction challenges of real-world web environments. An interception layer captures and blocks the final submission request, ensuring safe evaluation without real-world side effects. Our evaluations of 8 frontier models show that both proprietary and open-source models complete only a small portion of these tasks. For example, Claude Sonnet 4.6 achieves only 33.3%, which exposes gaps in current AI agents. Progress on ClawBench brings us closer to AI agents that can function as general-purpose assistants.

URL PDF HTML 收藏
2602.02138 2026-07-21 cs.SE cs.AI 版本更新

CAM: A Causality-based Analysis Framework for Multi-Agent Code Generation Systems

CAM:基于因果分析的多智能体代码生成系统分析框架

Zongyi Lyu, Zhenlan Ji, Songqiang Chen, Liwen Wang, Yuheng Huang, Shuai Wang, Shing-Chi Cheung

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) The University of Tokyo(东京大学)

AI总结 CAM提出一种基于因果分析的框架,用于分析多智能体代码生成系统中中间特征对系统正确性的贡献,通过实证分析揭示了依赖上下文的特征和混合架构的优势。

Comments 20 pages, 12 tables, 5 figures

详情
AI中文摘要

尽管多智能体代码生成系统(MACGS)取得了显著成功,但多智能体架构的内在复杂性产生了大量中间输出。截至目前,这些中间输出对系统正确性的个体重要性仍然不透明,这阻碍了针对MACGS设计的有针对性优化。为了解决这一挑战,我们提出了CAM,第一个基于因果分析的MACGS分析框架,系统地量化了不同中间特征对系统正确性的贡献。通过全面分类中间输出并系统地模拟中间特征上的现实错误,我们识别出对系统正确性重要的特征并汇总其重要性排名。我们对识别的重要性排名进行了广泛的实证分析。我们的分析揭示了引人注目的发现:首先,我们发现了依赖上下文的特征——那些重要性主要通过与其他特征的相互作用出现,揭示了MACGS的质量保证应纳入跨特征一致性检查;其次,我们发现混合后端MACGS,通过根据相对优势分配不同后端LLM,实现了高达7.2%的Pass@1提升,凸显了混合架构作为未来MACGS设计的有前途方向。我们进一步通过两个应用证明了CAM的实用价值:(1)故障修复通过优化前3名重要性排名的特征,实现了73.3%的成功率;(2)特征修剪将中间token消耗减少了高达66.8%,同时保持生成性能。我们的工作为MACGS设计和部署提供了可操作的见解,确立了因果分析作为理解和改进MACGS的强大方法。

英文摘要

Despite the remarkable success that Multi-Agent Code Generation Systems (MACGS) have achieved, the inherent complexity of multi-agent architectures produces substantial volumes of intermediate outputs. To date, the individual importance of these intermediate outputs to the system correctness remains opaque, which impedes targeted optimization of MACGS designs. To address this challenge, we propose CAM, the first \textbf{C}ausality-based \textbf{A}nalysis framework for \textbf{M}ACGS that systematically quantifies the contribution of different intermediate features for system correctness. By comprehensively categorizing intermediate outputs and systematically simulating realistic errors on intermediate features, we identify the important features for system correctness and aggregate their importance rankings. We conduct extensive empirical analysis on the identified importance rankings. Our analysis reveals intriguing findings: first, we uncover context-dependent features\textemdash features whose importance emerges mainly through interactions with other features, revealing that quality assurance for MACGS should incorporate cross-feature consistency checks; second, we reveal that hybrid backend MACGS with different backend LLMs assigned according to their relative strength achieves up to 7.3\% Pass@1 improvement, underscoring hybrid architectures as a promising direction for future MACGS design. We further demonstrate CAM's practical utility through two applications: (1) failure repair which achieves a 73.6\% success rate by optimizing top-3 importance-ranked features and (2) feature pruning that reduces up to 33.6\% intermediate token consumption while maintaining generation performance. Our work provides actionable insights for MACGS design and deployment, establishing causality analysis as a powerful approach for understanding and improving MACGS.

URL PDF HTML 收藏
2509.09371 2026-07-21 stat.ME cs.LG 版本更新

Representation-Aware Distributionally Robust Optimization: A Knowledge Transfer Framework

表示感知分布鲁棒优化:一种知识转移框架

Zitao Wang, Nian Si, Molei Liu

机构 * Department of Statistics, Columbia University(哥伦比亚大学统计系) Department of Industrial Engineering and Decision Analytics, Hong Kong University of Science and Technology(香港科技大学工业工程与决策分析系) Department of Biostatistics, Peking University Health Science Center(北京大学北京医科大学生物统计学系) Beijing International Center for Mathematical Research, Peking University(北京大学北京国际数学研究中心)

AI总结 研究提出表示感知分布鲁棒估计(READ)框架,利用外部表示指导鲁棒性几何,增加改变表示坐标扰动的运输成本。在当前目标推断和未来总体部署中研究READ,模拟和应用证明其在多源多任务转移学习中有优势。

详情
AI中文摘要

分布鲁棒优化(DRO)通过在一组扰动分布上优化最坏情况性能来保护统计学习免受分布变化影响。然而,标准DRO公式通常平等对待所有特征扰动。当外部知识表明预测信号嵌入协变量的低维表示中时,这可能过于保守。我们提出了表示感知分布鲁棒估计(READ),这是一个Wasserstein DRO框架,使用外部表示来指导鲁棒性几何。READ增加改变表示坐标的扰动的运输成本,同时保持对与表示正交的变化的保护。我们在两种情况下研究READ。首先,对于当前目标的推断,我们渐近地刻画估计器,并开发一种Wasserstein轮廓推断方法来构建表示对齐的置信区域,同时实现自动超参数调整。其次,对于部署到与当前目标不同但由相同表示不变随机系数模型生成的未来总体,我们表明所得区域比标准方法实现更高的未来模型参数覆盖率。模拟和单细胞多组学应用证明了READ在多源和多任务转移学习设置中的优势。

英文摘要

Distributionally robust optimization (DRO) protects statistical learning against distributional shifts by optimizing the worst-case performance over a set of perturbed distributions. However, standard DRO formulations often treat all feature perturbations equally. This can be unnecessarily conservative when external knowledge suggests that the predictive signal is embedded in a low-dimensional representation of covariates. We propose REpresentation-Aware Distributionally robust estimation (READ), a Wasserstein DRO framework that uses external representations to guide the geometry of robustness. Rather than uniformly perturbing all covariate directions, READ increases the transport cost of perturbations that change representation coordinates, thereby reshaping the dual regularization toward the representation subspace. Meanwhile, it preserves protection against variations orthogonal to the representation. We study READ in two regimes. First, for inference on the current target, we characterize our estimator asymptotically and develop a Wasserstein profile inference approach to construct representation-aligned confidence regions while enabling automatic hyperparameter tuning. Second, for deployment to future populations that differ from the current target but are generated from the same representation-invariant random-coefficient model, we show that the resulting regions achieve higher coverage of future model parameters than standard methods. Simulations and a single-cell multi-omics application demonstrate the advantages of READ in multi-source and multitask transfer learning settings.

URL PDF HTML 收藏
2403.12537 2026-07-21 cs.CV 版本更新

Prompt-Guided Foundation Model Tuning for Pathology Image Classification

用于病理图像分类的提示引导基础模型调优

Yi Lin, Zhengjie Zhu, Kwang-Ting Cheng, Hao Chen

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系) Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(香港科技大学电子与计算机工程系) Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology(香港科技大学化学与生物工程系) HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute(香港科技大学深圳-香港协同创新研究院) State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology(香港科技大学神经系统疾病国家重点实验室)

AI总结 研究针对病理图像分类中基础模型预训练与下游任务的差异问题,提出PAMT框架,通过引入RPS、PVP及AMT,在14个公开数据集上严格评估,显著提升分类准确率,确立了PAMT作为病理图像分类新基准的地位。

详情
AI中文摘要

基础模型在推进计算病理学方面至关重要,特别是对于全切片图像(WSI)分类。然而,现有方法常依赖冻结的预训练模型进行特征提取,忽视了预训练与下游任务间显著的领域转移和任务差异。为此提出PAMT,一种新颖的提示引导自适应模型转换框架,能使通用基础模型精确适应组织病理学的不同领域。引入代表性补丁采样(RPS)和原型视觉提示(PVP)来封装组织病理学数据的复杂分布特征,通过特征提取管道中的适配器模块纳入自适应模型转换(AMT)以有效弥合领域差距。在14个公开数据集上进行严格评估,分类准确率有持续显著提高。这些结果确立了PAMT作为病理图像分类的新基准,并强调了计算病理学中目标模型适应的关键价值。

英文摘要

Foundation models have become pivotal in advancing computational pathology, particularly for whole slide image (WSI) classification. However, prevailing methodologies often rely on frozen, pre-trained models for feature extraction, overlooking the pronounced domain shift and task discrepancy between the pre-training and downstream tasks. To address this challenge, we propose PAMT, a novel Prompt-guided Adaptive Model Transformation framework that enables precise adaptation of general foundation models to the distinct domain of histopathology. To encapsulate the intricate distributions characteristic of histopathological data, we introduce Representative Patch Sampling (RPS) and Prototypical Visual Prompt (PVP), which reconstruct the input into compact yet highly informative representations. Further, to effectively bridge the domain gap, we incorporate Adaptive Model Transformation (AMT) via adapter modules within the feature extraction pipeline, facilitating the acquisition of domain-specific features by the foundation model. We conduct rigorous evaluation across 14 publicly available datasets and demonstrate consistent, substantial improvements in classification accuracy. These results establish PAMT as a compelling new benchmark for pathology image classification and underscore the critical value of targeted model adaptation within computational pathology.

URL PDF HTML 收藏
2607.15933 2026-07-20 cs.CV 新提交

Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework

用于矢量量化的分布匹配:一个统一的理论和实证框架

Xianghong Fang, Litao Guo, Hengchao Chen, Yuxuan Zhang, XiaofanXia, Dingjie Song, Yexin Liu, Hao Wang, Harry Yang, Qiang Sun, Yuan Yuan

机构 * University of Toronto(多伦多大学) The Hong Kong University of Science and Technology(香港科技大学) Boston College(波士顿学院) Lehigh University(里海大学) Southern University of Science and Technology(南方科技大学)

AI总结 针对现有矢量量化方法训练不稳定和码本崩溃问题,提出分布匹配框架,通过对齐特征和码向量分布缓解上述问题,经理论分析和实验验证,基于瓦瑟斯坦距离目标实例化该框架,在视觉tokenization基准上表现有效且鲁棒。

Comments 33 pages, 17 figures, and 16 tables. arXiv admin note: substantial text overlap with arXiv:2506.15078

详情
AI中文摘要

现代视觉表征学习和自回归模型的有效性严重依赖矢量量化(VQ),它使用可学习码本离散化连续特征表示。尽管广泛使用,但现有VQ方法常因直通估计器导致的梯度失配和码向量利用不足而存在训练不稳定和码本崩溃问题。本文表明这两个问题可追溯到特征向量和码向量分布的根本不匹配,导致表示效率低下和信息损失。基于此,提出分布匹配框架,引入理想VQ行为的原则标准,经理论分析和实证评估表明对齐特征和码向量分布可缓解训练不稳定和码本崩溃。使用基于瓦瑟斯坦距离的目标在温和高斯近似下实例化该框架,还表明基于最大均值差异的非参数替代方法性能相当。在视觉tokenization基准上的大量实验支持了该方法的有效性和鲁棒性。

英文摘要

The effectiveness of modern visual representation learning and autoregressive models critically depends on vector quantization (VQ), which discretizes continuous feature representations using a learnable codebook. Despite its widespread use, existing VQ methods often suffer from training instability and codebook collapse, arising from gradient mismatch induced by the straight-through estimator and the under-utilization of code vectors. In this work, we show that both issues can be traced to a fundamental mismatch between the distributions of feature vectors and code vectors, leading to inefficient representation and information loss. Building on this observation, we propose a distributional matching framework for vector quantization. We introduce principled criteria for desirable VQ behavior and demonstrate through theoretical analysis and empirical evaluation that aligning feature and code vector distributions provides a unifying mechanism for mitigating training instability and codebook collapse. We instantiate this framework using a Wasserstein-based objective with an efficient closed-form under a mild Gaussian approximation, and further show that a nonparametric alternative based on maximum mean discrepancy yields comparable performance. Extensive experiments on visual tokenization benchmarks support the effectiveness and robustness of the proposed approach.

URL PDF HTML 收藏
2607.15901 2026-07-20 cs.AI 新提交

DSWorld: A Data Science World Model for Efficient Autonomous Agents

DSWorld:用于高效自主智能体的数据科学世界模型

Zherui Yang, Fan Liu, Hao Liu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 研究自主数据科学智能体依赖试错流程效率低的问题,提出数据科学世界模型概念及DSWorld框架,含多种技术。构建数据集并引入优化策略,实验显示其加速智能体训练和推理,在转换预测任务上超基线。

详情
AI中文摘要

尽管自主数据科学智能体在数据理解和决策方面能力较强,但仍严重依赖涉及昂贵计算的试错工作流程。这一瓶颈促使人们开发能够在实际执行前预测数据科学操作效果的模型。本文引入数据科学世界模型的概念,通过根据当前工作流程状态和候选操作预测环境状态转换来对数据科学执行环境进行建模。我们进一步提出了DSWorld,这是一个实用框架,结合了结构化状态构建、成本感知路由、轻量级实际执行以及用于昂贵操作的基于大语言模型的模拟器。为支持训练,我们构建了一个8K规模的转换轨迹数据集,并引入了反思世界模型优化,这是一种用于改进转换预测的误差感知强化学习策略。实验表明,DSWorld在保持竞争力的同时,将基于强化学习的智能体训练速度提高了约14倍,将基于搜索的推理速度提高了约3至6倍,并且在转换预测任务上比最强的大语言模型基线高出35.6%。代码可在该https网址获取。

英文摘要

Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily rely on trial-and-error workflows that involve expensive computation. This bottleneck motivates models that can anticipate the effects of data science operations before real execution. In this paper, we introduce the concept of Data Science World Model, which model the data science execution environment by predicting environment state transitions conditioned on current workflow states and candidate operations. We further propose DSWorld, a practical framework that combines structured state construction, cost-aware routing, lightweight real execution, and an LLM-based simulator for expensive operations. To support training, we construct an 8K-scale transition trajectory dataset and introduce Reflective World Model Optimization, an error-aware reinforcement learning strategy for improving transition prediction. Experiments show that DSWorld accelerates RL-based agent training by approximately $14\times$ and search-based inference by approximately $3$-$6\times$ while maintaining competitive performance, and outperforms the strongest LLM baseline by 35.6% on transition prediction tasks. The code is available at https://anonymous.4open.science/r/DSWorld.

URL PDF HTML 收藏
2607.15830 2026-07-20 cs.AR cs.AI cs.LG 新提交

RTL-Sequencer: Towards Scalable RTL Timing Prediction with the Sequence-based Paradigm

RTL-Sequencer:基于序列范式实现可扩展的RTL时序预测

Ziyan Guo, Wenji Fang, Wenkai Li, Yuchao Wu, Shang Liu, Zhiyao Xie

机构 * Hong Kong University of Science and Technology (HKUST)(香港理工大学)

AI总结 针对RTL精确时序预测难题,现有基于图的方法存在局限。RTL-Sequencer提出基于序列的范式,通过线性化逻辑锥、应用序列模型及定制协同技术实现可扩展预测,实验证明其相比基线有显著改进,推动早期时序优化。

Comments Accepted by Design Automation Conference (DAC) 2026

详情
AI中文摘要

寄存器传输级(RTL)的精确时序预测是设计自动化中一项长期存在的挑战。现有的基于图的方法存在感受野有限、复杂度高和缺乏信号方向性等问题。我们提出了RTL-Sequencer,这是一种新颖的基于序列的范式,通过广度优先遍历线性化逻辑锥并应用现代线性序列模型来实现可扩展的RTL时序预测。此外,序列模型通过序列洗牌、双向建模、可微建模和混合图序列架构这四种协同技术进行定制。大量实验表明,RTL-Sequencer比现有基线有显著改进,推动了早期时序优化。

英文摘要

Accurate timing prediction at the register-transfer level (RTL) is a longstanding challenge in design automation. Existing graph-based methods struggle with limited receptive fields, high complexity, and a lack of signal directionality. We present RTL-Sequencer, a novel sequence-based paradigm that enables scalable RTL timing prediction via linearizing logic cones by breadth-first traversal and applying modern linear sequence models. Furthermore, sequence models are customized by four synergistic techniques, including sequence shuffling, bidirectional modeling, differentiable modeling, and a hybrid graph-sequence architecture. Extensive experiments demonstrate significant improvements of RTL-Sequencer over state-of-the-art baselines, advancing early-stage timing optimization.

URL PDF HTML 收藏
2607.15806 2026-07-20 cs.CV 新提交

HybridSim: A Physics-Learning Hybrid Digital Twin for mmWave Human Sensing

HybridSim:用于毫米波人体感知的物理学习混合数字孪生

Weitao Xiong, Tianyu Liu, Peng Li, Kok Chung Chua, Toa Chean Khim, Pu Wang, Hongfei Xue

机构 * Xiamen University Malaysia(厦门大学马来西亚分校) The Hong Kong University of Science and Technology(香港科技大学) University of North Carolina at Charlotte(北卡罗来纳大学夏洛特分校)

AI总结 研究针对毫米波雷达信号模拟成本高问题,提出HybridSim混合模拟器,用三平面和图卷积网络等提取人体特征、稳定优化,通过逆渲染等建模信号路径,实验显示其能有效增强特定地点数据,提升人体感知任务表现。

Comments Accepted to ECCV 2026. Project Page: https://weitao-xiong.github.io/HybridSim/

详情
AI中文摘要

毫米波雷达信号对动态人体运动的高保真模拟,对于开发基于雷达的人体感知模型很有价值;但为特定部署地点收集精确标记的测量数据仍然成本高昂。我们提出了HybridSim,这是一种物理学习混合模拟器,它能在固定室内房间配置下,从动态人体网格合成毫米波雷达信号,将传播明确解耦为两个分量。为参数化人体对象,我们使用三平面表示提取人体特征,并使用图卷积网络稳定优化并减轻梯度不稳定性。直接信号路径通过具有微面元双向反射分布函数的逆渲染公式建模,以捕获主要表面反射。同时,间接路径通过将3D高斯点云与虚拟接收器几何相结合来近似,以拟合和再现特定地点的多径干扰模式,计算成本远低于显式全光线追踪。在固定房间设置中的实验表明,当使用HybridSim进行特定地点数据增强时,与基于物理的参考有更好的一致性,并且在下游基于雷达的人体感知任务上有持续的增益。

英文摘要

High-fidelity simulation of mmWave radar signals for dynamic human motion is valuable for developing radar-based human sensing models; yet collecting accurately labeled measurements for a specific deployment site remains expensive. We present HybridSim, a physics-learning hybrid simulator that synthesizes mmWave radar signals from dynamic human meshes under a fixed indoor room configuration, explicitly decoupling propagation into two components. To parameterize the human subject, we use a tri-plane representation to extract human features and a Graph Convolutional Network to stabilize optimization and mitigate gradient instability. The direct signal path is modeled via an inverse-rendering formulation with a microfacet BRDF to capture primary surface reflections. In parallel, the indirect path is approximated by combining 3D Gaussian Splatting with a virtual-receiver geometry to fit and reproduce site-specific multipath interference patterns, achieving substantially lower computational cost than explicit full ray tracing. Experiments in a fixed-room setting show improved agreement with a physically based reference and consistent gains on downstream radar-based human sensing tasks when using HybridSim for site-specific data augmentation.

URL PDF HTML 收藏
2607.04020 2026-07-20 cs.CV 版本更新

Paired Uterine Whole-Slide Images and Pathology Reports for Multimodal Computational Pathology

用于多模态计算病理学的配对子宫全切片图像和病理报告

Han Li, Jingsong Liu, Ayako Ura, Junlin Hou, Zhengyang Xu, Azar Kazemi, Oskar Thaeter, Christian Grashei, Fabian Gülhan, Reza Nasirigerdeh, Xun Ma, Rui Yan, Hao Chen, S. Kevin Zhou, Nassir Navab, Carolin Mogler, Peter Schüffler

机构 * Institute of Pathology, Technical University of Munich(慕尼黑工业大学病理研究所) Computer Aided Medical Procedures (CAMP), Technical University of Munich(慕尼黑工业大学计算机辅助医疗程序(CAMP)) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Department of Human Pathology, Juntendo University Graduate School of Medicine(顺天堂大学医学研究生院人体病理学部) The Hong Kong University of Science and Technology(香港科技大学) Munich Data Science Institute (MDSI)(慕尼黑数据科学研究所)

AI总结 研究子宫疾病病理诊断,针对全切片图像与病理报告配对数据集稀缺问题,引入TUM-Uteria数据集,含多对病例及切片级配对,经验证,为计算病理学研究提供支持。

详情
AI中文摘要

子宫疾病是妇科病理学的重要类别,全切片图像推动了病理学工作流程的数字化转型。联合分析组织病理学图像和病理报告的多模态模型在自动生成病理报告和人工智能辅助诊断方面显示出潜力。然而,此类系统的发展受到全切片图像与有临床意义的病理报告配对数据集稀缺的限制。我们引入了TUM-Uteria,这是一个子宫病理学数据集,包含从三级医疗中心收集的病例和切片级的全切片图像与诊断病理报告配对。该数据集包含216个临床病例,包括455个切片级全切片图像-报告对。该数据集经过了一个结构化的多阶段验证程序,涉及经董事会认证的病理学家,以确保可靠的注释。TUM-Uteria支持计算病理学的研究,包括全切片图像分析、多模态学习和自动病理报告生成。

英文摘要

Uterine diseases represent an important category of gynecologic pathology and require accurate histopathological assessment for diagnosis and treatment planning. Whole-slide images (WSI) have enabled the digital transformation of pathology workflows and provided new opportunities for artificial intelligence (AI) in computational pathology. In particular, multimodal models that jointly analyze histopathology images and pathology reports have shown promising potential for automated pathology report generation and AI-assisted diagnosis. However, the development of such systems remains limited by the scarcity of datasets that pair whole-slide images with clinically meaningful pathology reports. Instead, existing pathology datasets focus on patch- or slide-level annotations of a single endpoint (e.g., disease class), which do not fully capture the rich information in full clinical diagnostic workflow reports. Here, we introduce TUM-Uteria, a uterine pathology dataset comprising WSIs paired with diagnostic pathology reports at both the case and slide levels, collected from a tertiary medical center. The dataset contains 216 clinical cases, comprising 455 slide-level WSI-report pairs. The dataset underwent a structured multi-stage validation procedure involving board-certified pathologists to ensure reliable annotations. TUM-Uteria supports research in computational pathology, including whole-slide image analysis, multimodal learning, and automated pathology report generation.

URL PDF HTML 收藏
2605.25878 2026-07-20 eess.IV cs.CV 版本更新

A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation

临床验证的基础模型用于全面肺部病理解读

Zhengrui Guo, Zhengyu Zhang, Jiabo Ma, Yihui Wang, Fengtao Zhou, Yingxue Xu, Ling Liang, Chenglong Zhao, Qi Xie, Jinbang Li, Shujing Guo, Fangyi Han, Zhijian Cen, Ziyi Liu, Cheng Jin, Junlin Hou, Zhixuan Chen, Yu Cai, Lijuan Qu, Shifu Chen, Yueping Liu, Zhe Wang, Xiuming Zhang, Muyan Cai, Li Liang, Hao Chen

机构 * Department of Pathology, Nanfang Hospital, Southern Medical University, Guangzhou, China(南方医科大学南芳医院病理科,广州,中国) Department of Pathology, School of Basic Medical Sciences, Southern Medical University, Guangzhou, China(南方医科大学基础医学学院病理科,广州,中国) Department of Computer Science and Engineering, Hong Kong University of Science and Technology, Hong Kong, China(香港科技大学计算机科学与工程系,香港,中国) Guangdong Provincial Key Laboratory of Molecular Tumor Pathology, Guangzhou, China(广东省分子肿瘤病理重点实验室,广州,中国) Department of Pathology, Shandong Provincial Qianfoshan Hospital, Jinan, Shandong, China(山东省青岛坊山医院病理科,济南,山东,中国)

AI总结 提出PulmoFoundation,一种基于Virchow2和约4万张H&E染色全切片图像进行亚专科预训练的肺部病理基础模型,通过32项临床任务和前瞻性随机对照试验验证,在诊断准确性、效率和一致性上显著提升。

详情
AI中文摘要

病理评估指导肺癌诊断、治疗选择和预后评估,但当前的CPath方法依赖于针对孤立目标的任务特定模型。尽管泛癌基础模型提供了多功能性,但它们缺乏亚专科深度,且未在临床工作流程中评估或在真实世界环境中进行前瞻性验证。我们介绍了PulmoFoundation,这是一个多中心、前瞻性验证、随机对照试验(RCT)评估的基础模型,用于术前、术中和术后护理的全面肺部病理评估。PulmoFoundation基于Virchow2,通过使用约40,000张诊断性H&E染色全切片图像(WSI)进行亚专科特定预训练构建,并在约26,000张WSI上系统评估了32项临床相关任务。除了准确预测分子标记和患者生存率外,我们的模型在活检、冰冻切片和手术切除切片的核芯诊断任务中达到了临床级性能。在一项针对1,357名患者、涵盖11项诊断任务的注册前瞻性研究中,我们的模型实现了平均AUC 92.3%。使用预设的分诊阈值,PulmoFoundation可以减少68.8%的活检和83.0%的冰冻切片的额外二次复核负担,并推迟44.5%的IHC染色订单,阳性预测值分别为1.0、0.991和0.966。除了前瞻性验证,我们还进行了一项交叉RCT,涉及八名病理学家,AI辅助在4,928个病例-阅片者对中提高了诊断准确性(有AI为91.7%,无AI为83.8%)。AI辅助还使中位诊断时间减少了19.6%,诊断信心提高了8.7%,并将阅片者间一致性从中等(kappa=0.56)提高到显著(kappa=0.76)。这些评估共同支持PulmoFoundation作为临床验证的肺部病理决策支持系统。

英文摘要

Pathological assessment guides lung cancer diagnosis, treatment selection, and prognostic evaluation, yet current CPath approaches rely on task-specific models for isolated objectives. Although pan-cancer foundation models offer versatility, they lack subspecialty-level depth and have not been evaluated across clinical workflows or prospectively validated in real-world settings. We introduce PulmoFoundation, a multi-center, prospectively validated, randomized controlled trial (RCT)-evaluated foundation model for comprehensive lung pathology assessment across pre-operative, intra-operative, and post-operative care. Built upon Virchow2 via subspecialty-specific pretraining using ~40,000 diagnostic H&E-stained whole-slide images (WSIs), PulmoFoundation was systematically evaluated on ~26,000 WSIs across 32 clinically relevant tasks. In addition to accurately predicting molecular markers and patient survival, our model achieves clinical-grade performance in core diagnostic tasks across biopsy, frozen section, and surgical resection slides. In a registered prospective study of 1,357 patients across 11 diagnostic tasks, our model achieved an average AUC of 92.3%. Using pre-specified triage thresholds, PulmoFoundation could reduce additional second-review burden for 68.8% of biopsies and 83.0% of frozen sections, and defer 44.5% of IHC stain orders, with PPVs of 1.000, 0.991, and 0.966. Beyond prospective validation, we conducted a crossover RCT with eight pathologists, in which AI assistance improved diagnostic accuracy across 5,264 case-reader pairs (91.7% w/ AI vs. 83.2% w/o AI). AI assistance also reduced median diagnostic time by 18.3%, increased diagnostic confidence by 9.0%, and improved inter-rater agreement from moderate (kappa = 0.55) to substantial (kappa = 0.76). Together, these evaluations support PulmoFoundation as a clinically validated decision-support system for lung pathology.

URL PDF HTML 收藏
2605.15677 2026-07-20 cs.CL cs.CV 版本更新

VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing

VCG-Bench:迈向统一的视觉导向基准,用于结构化生成与编辑

Xiaoyan Su, Peijie Dong, Zhenheng Tang, Song Tang, Yuyao Zhai, Kaitao Lin, Liang Chen, Gai Yuhang, Yuyu Luo, Qiang Wang, Xiaowen Chu

机构 * The Hong Kong University of Science and Technology (GuangZhou)(香港科学与技术大学(广州)) Huawei Technologies Co., Ltd(华为技术有限公司) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) South China University of Technology(华南理工大学)

AI总结 本文提出VCG-Bench,一个统一的视觉导向mxGraph任务基准,通过符号逻辑和XML实现精确的图表生成与编辑,解决现有方法在结构化任务中的局限性。

Comments Accepted by ICML2026, 37 pages, 10 figures

详情
AI中文摘要

尽管视觉语言模型(VLMs)迅速发展,但在处理专业工作流程中至关重要的结构化、可控图表任务方面仍存在关键差距。现有方法主要依赖像素级合成,其在可编辑性和保真度上存在固有限制。本文提出一种新的图表即代码范式,利用mxGraph可扩展标记语言(XML)进行精确的图表生成与编辑。我们提出了VCG-Bench,一个统一的视觉导向mxGraph任务基准。VCG-Bench包括:(1)一个包含1,449种不同图表的分类数据集,涵盖6个领域和15个子领域;(2)一种整合生成(视觉到代码)和可编辑性(代码到代码)的范式定义;(3)一种定制的评估协议,采用多维指标,如mxGraph执行成功率、风格一致性分数(SCS)等。实验结果突显了当前最先进(SOTA)VLMs在结构保真度和指令合规性方面的挑战,反映了其视觉和推理能力。

英文摘要

Despite the rapid advancements in Vision-Language Models (VLMs), a critical gap remains in their ability to handle structured, controllable diagrammatic tasks essential for professional workflows. Existing methods predominantly rely on pixel-based synthesis, which operates in probabilistic pixel spaces and is inherently limited in editability and fidelity. Instead, we propose a new Diagram-as-Code paradigm with symbolic logic that leverages mxGraph Extensible Markup Language (XML) for precise diagram generation and editing. We present VCG-Bench, a unified benchmark for visual-centric \texttt{mxGraph} tasks. VCG-Bench comprises: (1) a taxonomized dataset of 1,449 diverse diagrams spanning 6 domains and 15 sub-domains, (2) a paradigm definition that integrates Generation (Vision-to-Code) and Editability (Code-to-Code), (3) a Tailored Evaluation Protocol employing multi-dimensional metrics such as \texttt{mxGraph} Execution Success Rate, Style Consistency Score (SCS), etc. Experimental results highlight the challenges faced by current State-of-the-Art (SOTA) VLMs in structured fidelity and instruction compliance, reflecting their vision and reasoning capabilities.

URL PDF HTML 收藏