arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

Huawei(华为)

至 收录 1281
2607.18060 2026-07-21 cs.RO 新提交

RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning

RoboHarness:用于长期规划的异构机器人策略的内存驱动编排

Jinbang Huang, Yuanzhao Hu, Zhiyuan Li, Ran Qi, Yixin Xiao, Zhanguang Zhang, Mark Coates, Tongtong Cao, Yingxue Zhang

机构 * Huawei Noah’s Ark Lab(华为诺亚方舟实验室) University of British Columbia(英属哥伦比亚大学) University of Toronto(多伦多大学) McGill University(麦吉尔大学) Labs(2012实验室)

AI总结 针对长期机器人任务需多种能力、异构策略编排难的问题,提出RoboHarness框架,通过多模态执行内存等表征策略能力边界,经内存桥接稳定策略交接,实验验证其在长期规划和分布外鲁棒性上有显著提升。

Comments 21 pages, 8 figures

详情
AI中文摘要

长期的机器人任务需要多种能力,单一策略无法可靠提供。异构策略具有互补优势,但编排它们需要处理不确定的能力边界和跨策略分布不匹配问题,现有基于同质、预定义技能且适用性固定的规划方法大多忽略了这些问题。我们提出了RoboHarness,一个统一框架,将独立开发的机器人控制系统封装为可重用的智能技能。它使用多模态执行内存和在线证据来表征策略能力边界以进行能力感知分解和路由。其内存桥接可稳定策略交接,通过检索与下一策略相关的执行轨迹等方式引导机器人。在多个基准测试、定制任务及真实机器人实验中验证了其有效性,在零样本长期规划和分布外鲁棒性方面有显著提升。

英文摘要

Long-horizon robotic tasks require diverse capabilities that no single policy can reliably provide. Heterogeneous policies offer complementary strengths, but orchestrating them requires reasoning over uncertain capability boundaries and cross-policy distribution mismatch, which are largely overlooked by existing planning methods built on homogeneous, predefined skills with fixed applicability. We propose RoboHarness, a unified framework that encapsulates independently developed robot control systems as reusable agentic skills. Although instantiated in this work with VLAs, RL policies, and task-and-motion planning (TAMP) systems, RoboHarness is designed as a general framework compatible with a broader range of robot policies, such as navigation policies, model predictive controllers, and world-action models. RoboHarness uses multi-modal execution memory and online evidence to characterize policy capability boundaries for capability-aware decomposition and routing. To stabilize policy handoffs, its Memory Bridge retrieves execution trajectories associated with the next policy, estimates its in-distribution state region, and guides the robot toward that region without joint policy retraining. Extensive experiments on three public benchmarks, 500 customized tasks, and 135 real-robot experiments demonstrate effective capability-aware routing and stable policy orchestration, yielding substantial improvements in zero-shot long-horizon planning and out-of-distribution robustness.

URL PDF HTML 收藏
2607.18042 2026-07-21 cs.CV cs.AI 新提交

Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation

行动前预测:基于未来状态条件的视觉语言导航

Lingfeng Zhang, Zhanguang Zhang, Liheng Ma, Tongtong Cao, Yingxue Zhang

机构 * Noah’s Ark Lab, 2012 Labs, Huawei(诺亚方舟实验室,2012实验室,华为)

AI总结 研究视觉语言导航中标准行为克隆问题,提出FSC-VLN方法,通过添加未来查询令牌并经训练后移除的目标分支将其隐藏状态与未来视觉嵌入对齐,实验表明该方法在R2R val-unseen上提升了相关指标。

Comments 10 pages, 1 figure, 5 tables

详情
AI中文摘要

端到端视觉语言导航(VLN)利用因果视觉语言模型可将指令和自我中心观察直接映射到行动,但标准行为克隆仅监督下一个行动,未明确训练策略状态以预测未来视觉结果。首先提出诊断性问题:若在训练和测试时给策略提供专家轨迹未来图像作为特权输入,该额外视觉证据对选择当前行动是否有用?答案是肯定的。接着提出可部署问题:推理时不访问未来图像,仅用压缩未来视觉潜在特征作为训练监督能否从未来信息中受益?提出了未来状态条件的VLN(FSC-VLN),在R2R val-unseen上,FSC-VLN在两种训练数据机制下比StreamVLN风格基线提高了SR/OSR/SPL,在长视野情节上增益更大;消融实验进一步支持了双查询设计。

英文摘要

End-to-end vision-language navigation (VLN) with causal vision-language models can map instructions and egocentric observations directly to actions, but standard behavior cloning supervises only the next action and does not explicitly train the policy state to be predictive of future visual outcomes. We first ask a diagnostic question: if the policy is given an expert-trajectory future image as privileged input at training and testing time, is that additional visual evidence useful for choosing the current action? (These expert-trajectory future images are unavailable at test time in real deployment, so we use this setting only as a privileged-input diagnostic.) The answer is yes; this sanity check shows that future observations can provide rich, actionable cues. We then ask a deployable question: without accessing future images at inference, can we still benefit from future information by using a compressed future visual latent only as training supervision? We propose Future-State-Conditioned VLN (FSC-VLN), which adds a future-query token and aligns its hidden state to a frozen visual embedding $Δ$ steps ahead via a training-only target branch that is removed after training. On R2R val-unseen, FSC-VLN improves SR/OSR/SPL over a StreamVLN-style baseline under two training-data regimes, with larger gains on long-horizon episodes; ablations further support the dual-query design (separating future and action queries).

URL PDF HTML 收藏
2606.12562 2026-07-21 cs.CV cs.GR 版本更新

HairPort: In-context 3D-aware Hair Import and Transfer for Images

HairPort: 上下文感知的3D发型导入与迁移

Alireza Heidari, Amirhossein Alimohammadi, Ali Mahdavi-Amiri

机构 * Simon Fraser University(西蒙菲莎大学) Huawei Canada(华为加拿大)

AI总结 提出HairPort框架,通过显式分离发型移除与迁移,并利用3D感知管道实现大姿态差异下的发型迁移,结合LoRA适配的秃头转换器和条件流匹配生成器,实现高质量、身份保持的发型迁移。

Comments Accepted to SIGGRAPH 2026 (Conference Papers Track). 23 pages, 15 figures, 10 tables, including supplementary material as appendices. Project page: https://deepmancer.github.io/HairPort/

详情
AI中文摘要

在图像之间迁移发型是计算机图形学、计算机视觉和视觉效果中一个重要但具有挑战性的任务。它使用户能够在无需实际改变发型的情况下探索新造型,应用于虚拟试穿系统、增强现实和娱乐等领域。大多数先前的方法在姿态差异较小时表现最佳,但在视角和尺度差异较大时效果不佳,此时缺失的发型内容必须合成而非迁移。我们提出HairPort,一个3D感知的发型迁移框架,通过显式分离发型移除与迁移,并在合成前强制几何一致性来解决这些问题。我们引入了一个秃头转换器,通过基于LoRA的上下文适配FLUX.1 Kontext生成逼真的秃头人脸版本。为了训练我们的秃头转换器,我们引入了一个新数据集Baldy,包含6000对在不同身份和条件下的秃头和原始图像。我们还使用了一个3D感知迁移管道,在将参考发型合成到源图像之前,从目标视角重建并重新渲染该发型。由于具有3D感知能力,我们的方法支持源和目标之间的大姿态和尺度差异。最后,一个条件流匹配生成器从秃头源和几何对齐的参考引导中合成迁移结果。综合来看,我们的方法实现了准确、姿态一致且身份保持的发型迁移,在定性和定量上均优于现有方法。

英文摘要

Transferring hairstyles between images is an important but challenging task in computer graphics, computer vision, and visual effects. It enables users to explore new looks without physically altering their hair, with applications in virtual try-on systems, augmented reality, and entertainment. Most prior works operate best under small pose gaps, and they fall short under large viewpoint and scale differences, where missing hair content must be synthesized rather than transferred. We propose HairPort, a 3D-aware hairstyle transfer framework that attempts to solve these issues by explicitly separating hair removal from transfer and enforcing geometric consistency before synthesis. We introduce a Bald Converter, which produces realistic bald versions of faces through LoRA-based in-context adaptation of FLUX.1 Kontext. To train our Bald Converter, we introduce a new dataset, Baldy, containing 6,000 paired bald and original images across diverse identities and conditions. We also use a 3D-Aware Transfer Pipeline that reconstructs and re-renders the reference hairstyle from the target viewpoint before compositing it onto the source image. Being 3D aware, our method supports large pose and scale discrepancies between the source and target. Finally, a conditional flow-matching generator synthesizes the transferred result from the bald source and geometry-aligned reference guidance. Together, our method enables accurate, pose-consistent, and identity-preserving hairstyle transfer, outperforming existing methods both qualitatively and quantitatively.

URL PDF HTML 收藏
2603.16461 2026-07-21 cs.CV 版本更新

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models

GAP-MLLM:几何对齐预训练以激活多模态大语言模型中的3D空间感知

Jiaxin Zhang, Junjun Jiang, Haijie Li, Youyu Chen, Kui Jiang, Dave Zhenyu Chen

机构 * Harbin Institute of Technology(哈尔滨工业大学) School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院) Huawei(华为)

AI总结 本文提出GAP-MLLM,通过几何对齐预训练激活多模态大语言模型中的3D空间感知,改进了传统方法在3D空间感知上的不足。

Comments Accepted by ECCV 2026. Project page: https://gapmllm.github.io/

详情
AI中文摘要

多模态大语言模型(MLLMs)在语义推理方面表现出色,但在纯RGB输入限制下难以实现3D空间感知。尽管利用了3D重建模型的隐式几何先验,基于图像的方法在性能上仍逊于使用显式3D数据的方法。我们认为,这种差距并非源于几何先验不足,而是训练范式不匹配:以文本为主的微调未能激活MLLMs中的几何表示。现有方法通常采用简单的特征拼接并直接优化下游任务,缺乏几何特定的监督。为此,我们提出GAP-MLLM,一种几何对齐的预训练范式,旨在在下游适应前显式激活结构感知。具体而言,我们引入了一个视觉提示的联合任务,迫使MLLMs预测稀疏点云图并同时预测语义标签,从而强制几何意识。此外,我们设计了多级渐进融合模块,具有令牌级门控机制,使几何先验能够自适应地整合而不抑制语义推理。大量实验表明,GAP-MLLM显著增强了几何特征融合,并在3D视觉定位、3D密集标注和3D视频目标检测任务中持续提升了性能。

英文摘要

Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging implicit geometric priors from 3D reconstruction models, image-based methods still exhibit a notable performance gap compared to methods using explicit 3D data. We argue that this gap does not arise from insufficient geometric priors, but from a misalignment in the training paradigm: text-dominated fine-tuning fails to activate geometric representations within MLLMs. Existing approaches typically resort to naive feature concatenation and optimize directly for downstream tasks without geometry-specific supervision, leading to suboptimal structural utilization. To address this limitation, we propose GAP-MLLM, a Geometry-Aligned Pre-training paradigm that explicitly activates structural perception before downstream adaptation. Specifically, we introduce a visual-prompted joint task that compels the MLLMs to predict sparse pointmaps alongside semantic labels, thereby enforcing geometric awareness. Furthermore, we design a multi-level progressive fusion module with a token-level gating mechanism, enabling adaptive integration of geometric priors without suppressing semantic reasoning. Extensive experiments demonstrate that GAP-MLLM significantly enhances geometric feature fusion and consistently enhances performance across 3D visual grounding, 3D dense captioning, and 3D video object detection tasks.

URL PDF HTML 收藏
2510.23472 2026-07-21 cs.LG cs.AI cs.AR cs.NE 版本更新

BBOPlace-Bench: Benchmarking Black-Box Optimization for Chip Placement

BBOPlace-Bench:用于芯片布局的黑盒优化基准测试

Ke Xue, Ruo-Tong Chen, Rong-Xi Tan, Xi Lin, Yunqi Shi, Siyuan Xu, Mingxuan Yuan, Chao Qian

机构 * National Key Laboratory for Novel Software Technology(国家新型软件技术重点实验室) School of Artificial Intelligence, Nanjing University(人工智能学院,南京大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

AI总结 针对芯片布局中黑盒优化缺乏统一基准的问题,提出BBOPlace-Bench基准,整合问题表述、芯片案例和算法家族,提供模块化框架,通过共享协议系统评估算法性能,助力开发高效方案并拓宽应用场景。

Comments IEEE TEvC

详情
AI中文摘要

芯片布局是现代芯片设计的关键阶段,黑盒优化(BBO)已应用数十年。早期受限于问题表述和算法设计,效率等不如主流分析方法。虽近期BBO有进展,但缺乏统一基准。为此提出BBOPlace-Bench,首个针对芯片布局BBO算法评估与开发的基准。它整合三种问题表述,提供模块化框架,汇总芯片案例并标准化格式,集成算法家族并系统评估性能。在共享协议下,部分BBO配置具竞争力,既助力开发高效BBO驱动的芯片布局方案,又拓宽BBO社区急需的实际应用场景。

英文摘要

Chip placement is a vital stage in modern chip design, and black-box optimization (BBO) has been applied to it for decades. Early BBO efforts, however, were limited by immature problem formulations and inefficient algorithm designs, leading to worse efficiency, quality, and scalability than mainstream analytical methods. Recent advances in BBO have shown strong potential, but a unified, BBO-specific benchmark for thoroughly assessing various problem formulations and BBO algorithms is lacking. To fill this gap, we propose BBOPlace-Bench, the first benchmark tailored for evaluating and developing BBO algorithms for chip placement. It integrates three BBO problem formulations and offers a modular, flexible framework that enables users to seamlessly implement, test, and compare their own algorithms. It aggregates representative modern chip cases and standardizes their formats, providing uniform and comprehensive information to support BBO optimization. Moreover, it integrates representative BBO algorithm families, including simulated annealing, population-based search (including GA, CMA-ES, and PSO), and Bayesian optimization, and systematically evaluates their performance across different problem formulations using key chip-placement metrics. We position these experiments primarily as illustrative case studies under a shared evaluation protocol, including common benchmark instances, metric definitions, evaluation pipeline, and search budgets. Under this protocol, some BBO configurations (e.g., GA under the mask-guided optimization formulation) are competitive with representative analytical and reinforcement learning baselines. BBOPlace-Bench not only facilitates the development of efficient BBO-driven solutions for chip placement but also broadens the practical application scenarios urgently needed by the BBO community.

URL PDF HTML 收藏
2509.23071 2026-07-21 cs.CL cs.AI 版本更新

From Evidence to Trajectory: Abductive Reasoning Path Synthesis for Retrieval-Augmented Generation Agents Development

从证据到轨迹:用于检索增强生成智能体开发的溯因推理路径合成

Muzhi Li, Jinhu Qi, Yihong Wu, Minghao Zhao, Liheng Ma, Yifan Li, Xinyu Wang, Zhenghan Tai, Zixing Song, Yingxue Zhang, Ho-fung Leung, Irwin King

机构 * The Chinese University of Hong Kong, Sha Tin, NT, Hong Kong(香港中文大学) Université de Montréal, Montréal, Quebéc, Canada(蒙特利尔大学) McGill University, Montréal, Quebéc, Canada(麦吉尔大学) Mila - Quebéc AI Institute, Montréal, Quebéc, Canada(魁北克AI研究院) Huawei Noah’s Ark Lab, Montréal, Quebéc, Canada(华为诺亚实验室)

AI总结 针对检索增强生成智能体开发缺乏可执行轨迹问题,提出EviPath范式,通过溯因子任务规划、忠实子问题回答、对话微调三个阶段合成推理路径,实验表明基于此训练的模型在开放域问答中显著优于基线。

Comments KDD 2026 Research Track

详情
AI中文摘要

检索增强生成(RAG)智能体开发因缺乏可执行的真实智能体与环境交互轨迹而受阻。现有数据集提供问题、答案和证据,但缺乏对检索器调用、动态规划和逐步决策的细粒度监督。强化学习有潜在解决方案,但在基础大语言模型缺乏足够推理能力时会面临稀疏奖励和冷启动失败。同时,现有数据合成方法主要生成事后理由而非可执行的环境交互轨迹。本文提出EviPath,一种用于RAG智能体开发的证据锚定推理路径合成范式。EviPath通过三个阶段从问答对和支持证据中反向工程可执行轨迹:(i)溯因子任务规划,分解问题并规划依赖感知的解决方案路径;(ii)忠实子问题回答,使用支持证据作为代理环境生成有根据的中间思想和答案;(iii)对话微调,将完整轨迹转换为对话格式进行监督微调。在广泛使用的问答基准上的实验表明,在我们的合成语料库上训练的8B模型显著且持续优于现有最先进基线,在开放域问答中实现了14.7%的绝对精确匹配增益。

英文摘要

Retrieval-augmented generation (RAG) agent development is hindered by the lack of executable ground-truth agent-environment interaction trajectories. Existing datasets provide questions, answers, and evidence, but lack fine-grained supervision for retriever invocation, dynamic planning, and stepwise decision-making. Reinforcement learning offers a potential solution, but often suffers from sparse rewards and cold-start failures when base large language models (LLMs) lack sufficient reasoning capability. Meanwhile, existing data synthesis methods mainly generate post-hoc rationales rather than executable environment-interaction trajectories. In this paper, we propose EviPath, an evidence-anchored reasoning path synthesis paradigm for RAG agent development. EviPath reverse-engineers executable trajectories from question-answer pairs and supporting evidence through three stages: (i) Abductive Subtask Planning, which decomposes questions and plans dependency-aware solution paths; (ii) Faithful Sub-question Answering, which uses supporting evidence as a proxy environment to generate grounded intermediate thoughts and answers; and (iii) Conversational Fine-Tuning, which converts complete trajectories into a dialogue format for supervised fine-tuning. Experiments on widely used question-answering benchmarks show that an 8B model trained on our synthetic corpus significantly and consistently outperforms state-of-the-art baselines, achieving a 14.7% absolute Exact Match gain in open-domain question answering.

URL PDF HTML 收藏
2607.15621 2026-07-20 cs.RO cs.AI 新提交

Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving

以5Hz思考,20Hz行动:用于闭环驾驶的异步快慢视觉-语言-动作推理

Yun Li, Jiachen Gong, Simon Thompson, Ehsan Javanmardi, Qunli Zhang, Zifan Zeng, Shiming Liu, Peng Wang, Zixuan Guo, Manabu Tsukada

机构 * The University of Tokyo(东京大学) TIER IV, Inc.(TIER IV公司) Huawei(华为)

AI总结 研究针对大语言模型用于闭环驾驶时推理延迟与车辆控制速率冲突的问题,提出快慢架构,慢系统用冻结主干处理历史,快专家基于缓存和当前帧回归航路点,经训练在CARLA上提升路线完成率,降低误差且可零样本转移。

Comments 13 pages, 5 figures, 4 tables

详情
AI中文摘要

大语言模型为端到端驾驶带来了指令跟随和场景推理能力,但其推理延迟与车辆所需的控制速率相冲突。现有闭环智能体通过在交替的模拟时间步调用模型并在其间重放先前命令来掩盖这一差距,导致一半的控制输出忽略了最新观测。我们提出了一种快慢架构来消除这种折衷。一个冻结的7B视觉-语言主干作为慢系统,低频消化导航指令和视觉历史,同时将其每层的键值缓存作为场景的固定表示。一个轻量级动作专家作为快系统,在每个模拟时间步关注此缓存和当前相机帧,通过单次前向传递回归航路点。由于缓存与实际场景存在延迟,我们在随机陈旧性下训练专家,使训练与异步执行对齐。在CARLA的LangAuto-Short路线上,我们的系统每50毫秒模拟时间步产生新的控制,将路线完成率从37.0提升到94.0,超过了跳帧基线。具有相同专家的跳帧消融实验分离了两个起作用的因素:专家自身提高了驾驶分数,而每时间步的新鲜度将完成率从82.提升到94.0,并将闯红灯违规减少了三分之一。在单个城镇训练的专家可以零样本转移到两个未见城镇,保持84-94%的路线完成率,而基线仅为31-41%。与主干自身的动作头相比,它将开环航路点误差降低了近四倍,在单个消费级GPU上每时间步模型成本为32毫秒,且与历史长度无关。

英文摘要

Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires. Existing closed-loop agents hide this gap by invoking the model on alternate simulation ticks and replaying the previous command in between, so half of all control outputs ignore the newest observations. We present a fast-slow architecture that removes this compromise. A frozen 7B vision-language backbone acts as the slow system, digesting navigation instructions and visual history at low frequency while exposing its per-layer key-value cache as a standing representation of the scene. A lightweight action expert acts as the fast system, attending to this cache and to the current camera frame at every simulation tick to regress waypoints in a single forward pass. Since the cache lags behind the world at deployment, we train the expert under randomized staleness, aligning training with asynchronous execution. On LangAuto-Short routes in CARLA, our system produces fresh control at every 50 ms simulation tick and lifts route completion from 37.0 to 94.0 over the frame-skipping baseline. A frame-skip ablation with the same expert separates the two factors at work: the expert raises the driving score on its own, while per-tick freshness raises completion from 82.1 to 94.0 and cuts red-light violations by a third. Trained on a single town, the expert transfers zero-shot to two unseen towns, holding 84-94% route completion where the baseline reaches 31-41%. It reduces open-loop waypoint error by nearly a factor of four compared to the backbone's own action head, at a per-tick model cost of 32 ms that is independent of history length on a single consumer GPU.

URL PDF HTML 收藏
2606.09421 2026-07-20 cs.CL 版本更新

What Should a Skill Remember? Quality--Cost Trade-offs in Cost-Aware Skill Rewriting for Language Model Agents

技能应记住什么?语言模型代理中成本感知技能重写的质量-成本权衡

Qinghua Xing, Yinda Chen, Yaping Jin, Zhenhe Wu, Bohan Lin, Hang Zhou, Xinghao Chen, Hanting Chen, Zhiwei Xiong

机构 * University of Science and Technology of China(中国科学技术大学) Huawei Technologies(华为技术有限公司) Tianjin University(天津大学)

AI总结 研究语言模型代理中技能重写的质量-成本权衡,提出信息保留策略,在SkillsBench上实现成本降低7%-14.7%且保持验证质量。

详情
AI中文摘要

大型语言模型代理越来越依赖技能:可重用的程序文档,编码工作流程、工具使用、实现模式、验证检查和领域规则。技能重写通常被视为提示压缩,但较短的技能可能通过移除防止探索、调试和恢复的稀疏操作锚点而使代理更昂贵。我们通过这种经济视角研究技能重写。我们的受控框架剖析技能结构,使用信息保留策略重写技能,并在固定任务指令、环境和验证器下评估重写。在SkillsBench上的实验揭示了不同策略间明显的质量-成本权衡:API/代码锚定、工作流保护和规则/公式锚定有利于不同的任务族,没有普遍主导的模板。在主要的留出评估中,学习到的策略将总成本降低7.0%,下游代理令牌成本降低6.0%;在冻结的跨模型迁移中,相应的降低平均为14.7%和13.7%,同时验证器质量保持不变。这些结果将技能设计定位为成本感知的操作知识工程,而非提示压缩。资源:\href{https://github.com/1Reminding/Skill_EE}{SkillEE}。

英文摘要

Large language model agents increasingly rely on skills: reusable procedural documents encoding workflows, tool use, implementation patterns, validation checks, and domain rules. Skill rewriting is often treated as prompt compression, but shorter skills can make agents more expensive by removing sparse operational anchors that prevent exploration, debugging, and recovery. We study skill rewriting through this economic lens. Our controlled framework profiles skill structure, rewrites skills using information-preservation strategies, and evaluates the rewrites under fixed task instructions, environments, and verifiers. Experiments on SkillsBench reveal distinct quality--cost trade-offs across strategies: API/code anchoring, workflow guarding, and rule/formula anchoring benefit different task families, with no universally dominant template. In the main held-out evaluation, the learned policy reduces total cost by 7.0% and downstream agent-token cost by 6.0%; in frozen cross-model transfer, the corresponding reductions average 14.7% and 13.7%, while verifier quality is preserved. These results position skill design as cost-aware operational knowledge engineering rather than prompt compression. Resources: https://github.com/1Reminding/Skill_EE.

URL PDF HTML 收藏
2605.15677 2026-07-20 cs.CL cs.CV 版本更新

VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing

VCG-Bench:迈向统一的视觉导向基准,用于结构化生成与编辑

Xiaoyan Su, Peijie Dong, Zhenheng Tang, Song Tang, Yuyao Zhai, Kaitao Lin, Liang Chen, Gai Yuhang, Yuyu Luo, Qiang Wang, Xiaowen Chu

机构 * The Hong Kong University of Science and Technology (GuangZhou)(香港科学与技术大学(广州)) Huawei Technologies Co., Ltd(华为技术有限公司) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) South China University of Technology(华南理工大学)

AI总结 本文提出VCG-Bench,一个统一的视觉导向mxGraph任务基准,通过符号逻辑和XML实现精确的图表生成与编辑,解决现有方法在结构化任务中的局限性。

Comments Accepted by ICML2026, 37 pages, 10 figures

详情
AI中文摘要

尽管视觉语言模型(VLMs)迅速发展,但在处理专业工作流程中至关重要的结构化、可控图表任务方面仍存在关键差距。现有方法主要依赖像素级合成,其在可编辑性和保真度上存在固有限制。本文提出一种新的图表即代码范式,利用mxGraph可扩展标记语言(XML)进行精确的图表生成与编辑。我们提出了VCG-Bench,一个统一的视觉导向mxGraph任务基准。VCG-Bench包括:(1)一个包含1,449种不同图表的分类数据集,涵盖6个领域和15个子领域;(2)一种整合生成(视觉到代码)和可编辑性(代码到代码)的范式定义;(3)一种定制的评估协议,采用多维指标,如mxGraph执行成功率、风格一致性分数(SCS)等。实验结果突显了当前最先进(SOTA)VLMs在结构保真度和指令合规性方面的挑战,反映了其视觉和推理能力。

英文摘要

Despite the rapid advancements in Vision-Language Models (VLMs), a critical gap remains in their ability to handle structured, controllable diagrammatic tasks essential for professional workflows. Existing methods predominantly rely on pixel-based synthesis, which operates in probabilistic pixel spaces and is inherently limited in editability and fidelity. Instead, we propose a new Diagram-as-Code paradigm with symbolic logic that leverages mxGraph Extensible Markup Language (XML) for precise diagram generation and editing. We present VCG-Bench, a unified benchmark for visual-centric \texttt{mxGraph} tasks. VCG-Bench comprises: (1) a taxonomized dataset of 1,449 diverse diagrams spanning 6 domains and 15 sub-domains, (2) a paradigm definition that integrates Generation (Vision-to-Code) and Editability (Code-to-Code), (3) a Tailored Evaluation Protocol employing multi-dimensional metrics such as \texttt{mxGraph} Execution Success Rate, Style Consistency Score (SCS), etc. Experimental results highlight the challenges faced by current State-of-the-Art (SOTA) VLMs in structured fidelity and instruction compliance, reflecting their vision and reasoning capabilities.

URL PDF HTML 收藏
2603.26556 2026-07-20 cs.CL cs.AI 版本更新

When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models

当困惑度谎言:面向生成的混合序列模型蒸馏

Juan Gabriel Kostelec, Qinghai Guo

机构 * Huawei Zurich Research Center(华为苏黎世研究中心) ACS Lab, Huawei Technologies(华为技术有限公司ACS实验室)

AI总结 本文提出Hybrid-KDA架构与GenDistill蒸馏流程,通过生成导向评估揭示传统困惑度评估的局限性,并展示改进后的模型在知识基准上的高准确率与内存效率提升。

Comments 13 pages, 4 figures, 4 tables

详情
AI中文摘要

将预训练的Transformer转换为更高效的混合模型通过蒸馏,提供了一种减少推理成本的有希望的方法。然而,实现蒸馏模型的高质量生成需要仔细设计学生架构和蒸馏过程。许多先前的蒸馏工作通过排名候选答案的对数似然来评估下游多项选择基准,而不是要求自回归生成,这可能掩盖了模型质量的重要差异。例如,我们展示了一个7B参数蒸馏模型在对数似然评分下几乎与教师模型相差0.2pp,但在必须自回归生成时却落后20.8pp。我们提出Hybrid-KDA架构与GenDistill多阶段蒸馏流程,并在整个过程中使用基于生成的评估来指导设计决策。将此方法应用于Qwen3-0.6B,我们系统地消除了六个设计轴:训练目标、损失掩码、训练持续时间、数据集选择、参数冻结和架构选择。我们发现基于对数似然的评估一致低估了教师和学生之间的差距,并且在某些情况下会反转设计选择的排名,意味着仅基于困惑度的评估结论可能是误导的。在我们研究的因素中,数据集选择、仅完成掩码和训练后冻结注意力层对生成质量影响最大。我们的最佳Hybrid-KDA模型在知识基准上保留了86-90%的教师准确性,同时将KV缓存内存减少高达75%,并在128K-token上下文中将时间到第一个token提高2-4倍。

英文摘要

Converting a pretrained Transformer into a more efficient hybrid model through distillation offers a promising approach to reducing inference costs. However, achieving high-quality generation in distilled models requires careful joint design of both the student architecture and the distillation process. Many prior distillation works evaluate downstream multiple-choice benchmarks by ranking candidate answers with log-likelihood rather than requiring autoregressive generation, which can obscure important differences in model quality. For example, on overlapping benchmarks, we show that a 7B distilled model that nearly matches its teacher to within 0.2 pp under log-likelihood scoring falls behind by 20.8 pp when it must generate answers autoregressively. We investigate this phenomenon with GenDistill, a multi-stage pipeline we designed for distilling a pretrained Transformer into an efficient Hybrid Kimi Delta Attention (Hybrid-KDA) student. Using it as a controlled testbed on Qwen3-0.6B, we systematically ablate six design axes (training objective, loss masking, training duration, dataset selection, parameter freezing, and architecture choice) and evaluate every choice under both log-likelihood and generation-based protocols. We find that log-likelihood-based evaluation consistently underestimates the gap between teacher and student, and can in some cases reverse the ranking of design choices, so conclusions drawn from perplexity-only evaluation may be misleading. Among the factors we study, dataset selection, completion-only masking, and freezing attention layers during post-training have the largest impact on generation quality. Our best distillation recipe, using a Hybrid-KDA model as the student, retains 86-90% of teacher accuracy on knowledge benchmarks while reducing KV cache memory by up to 75% and improving time-to-first-token by 2-4x at 128K-token contexts.

URL PDF HTML 收藏
2601.12222 2026-07-20 cs.SD cs.MM eess.AS 版本更新

Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling

基于多茎注意力和分层不确定性建模的歌曲美学评估

Yishan Lv, Jing Luo, Boyuan Ju, Yang Zhang, Xinda Wu, Bo Yuan, Xinyu Yang

机构 * School of Computer Science and Technology, Xi’an Jiaotong University, Xi’an, China(计算机科学与技术学院,西安交通大学,西安,中国) Central Media Technology Institute, Huawei(中央媒体技术研究院,华为)

AI总结 针对音乐生成人工智能带来的歌曲美学评估需求,提出含多茎注意力融合与分层粒度感知区间聚合模块的评估框架,在两个数据集上评估并与两个SOTA模型比较,该方法在多维歌曲美学评估中性能更强。

Comments Accepted to the 27th International Society for Music Information Retrieval Conference (ISMIR 2026)

详情
AI中文摘要

音乐生成人工智能正在迅速扩展音乐内容,因此需要自动化的歌曲美学评估。然而,现有研究大多集中在语音、音频或演唱质量上,歌曲美学研究不足。此外,传统方法通常直接预测精确的平均意见得分(MOS)值,难以捕捉歌曲美学评估中人类感知的细微差别。本文提出了一个面向歌曲的美学评估框架,具有两个新颖的模块:多茎注意力融合(MSAF)在混合人声和混合伴奏对之间建立双向交叉注意力,融合它们以捕捉复杂的音乐特征;分层粒度感知区间聚合(HiGIA)学习多粒度得分概率分布,将它们聚合到一个得分区间,并在区间内应用回归以产生最终得分。我们在两个全长歌曲数据集上进行了评估:SongEval数据集(人工智能生成)和一个内部美学数据集(人类创作),并与两个最先进的(SOTA)模型进行了比较。结果表明,所提出的方法在多维歌曲美学评估中取得了更强的性能。推理代码和检查点可在这个https URL上公开获得。

英文摘要

Music generative artificial intelligence (AI) is rapidly expanding music content, necessitating automated song aesthetics evaluation. However, existing studies largely focus on speech, audio or singing quality, leaving song aesthetics underexplored. Moreover, conventional approaches often predict a precise Mean Opinion Score (MOS) value directly, which struggles to capture the nuances of human perception in song aesthetics evaluation. This paper proposes a song-oriented aesthetics evaluation framework, featuring two novel modules: 1) Multi-Stem Attention Fusion (MSAF) builds bidirectional cross-attention between mixture-vocal and mixture-accompaniment pairs, fusing them to capture complex musical features; 2) Hierarchical Granularity-Aware Interval Aggregation (HiGIA) learns multi-granularity score probability distributions, aggregates them into a score interval, and applies a regression within the interval to produce the final score. We evaluated on two datasets of full-length songs: SongEval dataset (AI-generated) and an internal aesthetics dataset (human-created), and compared with two state-of-the-art (SOTA) models. Results show that the proposed method achieves stronger performance for multi-dimensional song aesthetics evaluation. The inference code and checkpoint are publicly available at https://github.com/yisan33/song-aesthetics-evaluation.

URL PDF HTML 收藏
2607.14277 2026-07-17 cs.CL 新提交

Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making

多头潜在控制:大语言模型智能体决策的统一接口

Amirhosein Ghasemabadi, Ruichen Chen, Bahador Rashidi, Di Niu

机构 * University of Alberta(阿尔伯塔大学) Huawei Technologies Canada Co., Ltd.(华为加拿大技术有限公司)

AI总结 研究大语言模型作智能体时的决策控制问题,提出多头潜在控制方法,通过读取冻结模型的隐藏状态轨迹生成控制信号,能在不修改模型的情况下进行事后适配,改善多模型系统质量成本权衡,减少大模型使用并提高工具使用决策质量。

详情
AI中文摘要

大语言模型越来越多地被用作智能体,但可靠的智能体行为需要的不仅仅是下一个token预测。在推理时,智能体最好能决定是继续当前推理、听从更强的模型、请求更多信息、调用外部工具还是弃权。现有方法通过提示级路由、外部编排或特定任务微调来处理这些决策,主要依赖输入端信号,且随着模型主干的发展成本高且难以维护。本文提出多头潜在控制,从冻结的大语言模型或视觉语言模型中读取隐藏状态轨迹以生成部署时控制信号,包括能力头和分辨率头,仅在相同冻结模型主干的潜在轨迹上训练,能在不修改模型的情况下进行事后适配。跨语言和视觉语言设置,该方法持续改善多模型系统的质量成本权衡,减少大模型使用,提高工具使用决策质量。

英文摘要

Large language models are increasingly deployed as agents, but reliable agentic behavior requires more than next-token prediction. At inference time, it is preferred that an agent can decide whether to proceed with its current reasoning, defer to a stronger model, request additional information, invoke external tools, or abstain under the given setup. Existing approaches address these decisions through prompt-level routing, external orchestration, or task-specific fine-tuning, which primarily rely on input-side signals, and are often costly and difficult to maintain as model backbones evolve. We ask whether such control decisions can be inferred directly from a model's latent generation process. We introduce Multi-Head Latent Control, a lightweight layer that reads hidden-state trajectories from a frozen LLM or VLM to produce deployment-time control signals. A Capability Head predicts whether the current model can solve the instance or should defer to a stronger collaborator, while a Resolution Head predicts appropriate resolution decision Clarification, Tool Use, Abstention, or Direct Answering. Both heads are trained only on latent traces from the same frozen LLM backbone, enabling post hoc adaptation without modifying the model. Across language and vision-language settings, Multi-Head Latent Control consistently improves the quality-cost tradeoff of multi-model systems, enabling early handoff from partial generations and more accurate intervention decisions. In routed execution (small + large model), it reduces large-model usage by up to 90.7 percent on AndroidWorld and 27-53 percent on average across benchmarks, while retaining most of large-model performance. Additionally, the learned control signals improve tool-use decision quality, yielding up to +158 percent relative score gain and 65.5 percent fewer missed-required tool calls.

URL PDF HTML 收藏
2607.09815 2026-07-17 cs.RO cs.CV 版本更新

RASR: Range-Aware Scale Recovery for Metric UAV Navigation

RASR:用于度量无人机导航的距离感知尺度恢复

Hongtao Liang, Xinyu Shao, Chenxu Wang, Yiyao Wan, Jiahuan Ji, Fangwei Ye, Fuhui Zhou, Qihui Wu

机构 * College of Electronic and Information Engineering, Nanjing University of Aeronautics and Astronautics(南京航空航天大学电子信息工程学院) Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Noah Ark Lab, Huawei(华为诺亚方舟实验室) College of Automation Engineering, Nanjing University of Aeronautics and Astronautics(南京航空航天大学自动化工程学院) College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院) College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院)

AI总结 研究在GNSS信号阻断时无人机精确最后一米导航问题,提出RASR方法,将尺度恢复核心与校准模块分离,核心压缩几何成描述符并全局校准,校准模块进行残差校正和对齐,在PairUAV评估中取得较好结果。

Comments 5 pages, 4 figures. Technical report for the UAVM 2026 PairUAV Challenge

详情
AI中文摘要

在全球导航卫星系统(GNSS)信号被阻断的情况下,无人机控制器仍需要可执行的距离和航向指令,这使得精确的最后一米度量导航至关重要。密集对几何基础模型能很好地传递相对结构,但其原始度量输出的距离尺度校准不佳。在PairUAV的相对误差度量下,仅校正平均尺度仍会在目标附近留下与距离相关的高成本残差。为解决这种尺度不匹配问题,距离感知尺度恢复(RASR)在推理时固定的每对系统中,将可转移的尺度恢复核心与特定协议校准模块分离。核心将冻结的匹配和立体3D重建(MASt3R)风格的几何压缩成紧凑描述符,并使用全局校准恢复主导度量信号。距离桶残差校正和命令网格对齐保留在校准模块内,以匹配PairUAV的命令格式和评估协议。在2026年多媒体PairUAV在线评估中的无人机上,RASR的总分为0.0031⑧9。在PairUAV协议下,冻结的对几何可产生稳定的每对距离和航向估计,而每个特定协议的调整都局限于推理前固定的校准模块。代码和材料可在该https网址获取。

英文摘要

A central challenge in image-goal UAV navigation under Global Navigation Satellite System (GNSS) denial is estimating metric distance and heading between current and goal views. Dense pairwise geometry models capture relative scene structure, but without a calibrated metric scale, they cannot directly provide reliable distance estimates for navigation. Although global scale calibration corrects the dominant scale bias, the remaining errors vary systematically with distance. In this paper, Range-Aware Scale Recovery (RASR) is proposed, which complements global scale calibration with range-aware residual correction. RASR encodes pairwise geometry extracted by a frozen Matching And Stereo 3D Reconstruction (MASt3R) backbone as a compact descriptor and separates the scale-recovery core from task-specific command calibration. On the official online evaluation of the UAVs in Multimedia 2026 PairUAV challenge, RASR achieved a total error of 0.003189, achieving a lower total error than global scale calibration alone. The results demonstrate that range-aware residual correction improves metric distance estimation beyond global scale calibration. Code and materials are available at https://github.com/lht-research/rasr-pairuav.

URL PDF HTML 收藏
2607.13927 2026-07-16 cs.CV 新提交

Cyclone: Diffusion Model for Cycle-Consistent Weather Editing from Unpaired Driving Data

Cyclone:基于未配对驱动数据的循环一致天气编辑扩散模型

Thang-Anh-Quan Nguyen, Moussab Bennehar, Luis Guillermo Roldao Jimenez, Nathan Piasco, Dzmitry Tsishkou, Laurent Caraffa, Jean-Philippe Tarel, Roland Brémond

机构 * Huawei Paris Research Center(华为巴黎研究中心) Gustave Eiffel University(古斯塔夫·埃菲尔大学) IGN-ENSG(法国国家地理信息与森林和环境信息研究所)

AI总结 针对自动驾驶系统在不同天气条件下可靠感知的挑战,提出Cyclone框架,基于潜在扩散,利用循环一致约束和图像-文本模型知识,无需配对数据生成多种天气条件,实验表明其输出更优,还可提炼为视频扩散模型。

Comments Project page: https://ntaquan0125.github.io/weather-cyclone/

详情
AI中文摘要

在不同天气条件下的可靠感知仍然是自动驾驶系统的一个主要挑战。一种提高鲁棒性的常见策略是为训练感知模型合成不利天气条件,或应用天气去除技术来恢复干净的输入。然而,现有方法通常依赖于合成数据增强或基于物理的特定任务模型,这些模型需要配对的训练数据,并且往往难以生成逼真的天气效果或稳健地推广到域外场景。针对这个问题,我们提出了Cyclone,一个基于潜在扩散的天气编辑统一框架,配备了循环一致约束和来自图像-文本模型的知识。Cyclone能够在不同场景中生成多种天气条件,同时无需配对数据。实验结果表明,我们的方法比现有基线产生更逼真、保留结构的输出,并在几个下游驾驶感知任务中带来一致的改进。此外,我们证明Cyclone可以提炼为一个用于时间一致天气编辑的视频扩散模型。

英文摘要

Reliable perception under diverse weather conditions remains a major challenge for autonomous driving systems. A common strategy to improve robustness is either to synthesize adverse weather conditions for training perception models or to apply weather-removal techniques to recover clean inputs. However, existing approaches typically rely on synthetic data augmentation or physics-based, task-specific models that require paired training data and often struggle to generate realistic weather effects or generalize robustly to out-of-domain scenarios. Toward this problem, we present Cyclone, a unified framework for weather editing based on latent diffusion, equipped with cycle-consistent constraints and knowledge from image-text models. Cyclone enables the generation of multiple weather conditions across diverse scenes while eliminating the need for paired data. Experimental results show that our approach produces more realistic, structure-preserving outputs than existing baselines and leads to consistent improvements across several downstream driving perception tasks. Furthermore, we demonstrate that Cyclone can be distilled to a video diffusion model for temporally consistent weather editing.

URL PDF HTML 收藏
2607.13770 2026-07-16 cs.AR cs.AI 新提交

Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations

Kaleido:通过利用潜在空间相关性对视频扩散变压器进行算法-硬件协同设计

Wenxuan Miao, Haosong Liu, Weiming Hu, Zihan Liu, Aiyue Chen, Jianlin Yu, Yiwu Yao, Yiming Gan, Jieru Zhao, Jingwen Leng, Minyi Guo, Yu Feng

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Jiao Tong University, Shanghai Qi Zhi Institute(上海交通大学、上海颀智研究所) Huawei Technologies(华为技术有限公司) ICT, Chinese Academy of Sciences(信息科技研究所、中国科学院)

AI总结 针对视频扩散变压器计算成本高的问题,提出Kaleido算法-硬件协同设计,利用潜在空间通道级时空相关性加速操作,有轻量级重用算法,设计了加速器,实验表明其相比现有加速器有显著加速和节能效果。

详情
AI中文摘要

视频扩散变压器(vDiTs)能生成高质量视频,但由于扩散时间步长和自注意力计算,计算成本极高。随着扩散时间步长减少,自注意力计算成本成为主要瓶颈。现有加速方法大多继承大语言模型的稀疏注意力技术,未考虑视频数据独特的时空相关性。本文提出Kaleido,一种算法-硬件协同设计,通过利用潜在空间中的通道级时空相关性加速vDiTs中的所有操作。基于此,提出轻量级通道级重用算法,在保留比现有方法更高生成质量(>17dB)的同时跳过冗余计算。还设计了具有可重构处理元件的脉动阵列加速器和轻量级数据调度器。对三个主流vDiT模型的评估表明,Kaleido比现有加速器加速高达5.9倍,节能16.0倍。

英文摘要

Video diffusion transformers (vDiTs) generate high quality video but introduce extremely high compute cost due to the long diffusion timesteps and self attention computation. As diffusion timesteps are reduced, the computation cost of self attention becomes the dominant bottleneck. Existing acceleration approaches largely inherit sparse attention techniques from large language models, which fail to consider the unique spatiotemporal correlation of video data. This paper presents Kaleido, an algorithm hardware codesign that accelerates all operations in vDiTs by exploiting channel-wise spatiotemporal correlations in latent space. Based on this insight, we propose a lightweight channelwise reuse algorithm that skips redundant computations by reusing partial results while preserving higher generative quality than prior methods (>17 dB). To efficiently support this algorithm, we design a systolic array like accelerator with reconfigurable processing elements and a lightweight data dispatcher to mitigate irregular sparsity and data access patterns introduced by our reuse algorithm. Evaluations across three mainstream vDiT models show that Kaleido achieves up to 5.9x speedup and 16.0x energy savings over state of the art accelerators.

URL PDF HTML 收藏
2607.13428 2026-07-16 cs.LG 新提交

PUe: Biased Positive-Unlabeled Learning Enhancement by Causal Inference

PUe:基于因果推断的有偏正无标记学习增强

Xutao Wang, Hanting Chen, Tianyu Guo, Yunhe Wang

机构 * Huawei Noah’s Ark Lab(华为诺亚方舟实验室)

AI总结 研究正无标记学习问题,基于SAR-PU倾向加权框架提出PUe框架,运用归一化倾向得分和NIPW,有归一化逆概率加权风险公式等贡献,在多个数据集实验中,在非均匀标签分布下优于多个PU基线。

Comments Extended arXiv version of the NeurIPS 2023 paper; includes additional discussion of related SAR-PU work

详情
AI中文摘要

正无标记(PU)学习旨在利用有限的标记正例和大量未标记例实现高精度二分类。现有基于成本敏感的方法常依赖强假设,即观察到正标记的示例是完全随机选择的。但实际中标签分布不均,存在选择偏差。基于Bekker等人的SAR-PU倾向加权框架,研究使用归一化倾向得分和归一化逆概率加权(NIPW)的PU学习增强(PUe)框架。其主要贡献包括归一化逆概率加权的PU风险公式、偏差标记下归一化样本权重误差和常见PU估计器的理论分析、正则化深度倾向得分估计、与现代成本敏感PU方法集成以及对选择性标记负类的支持。在MNIST、CIFAR-10和ADNI上的实验表明,在非均匀标签分布下优于多个PU基线。

英文摘要

Positive-Unlabeled (PU) learning aims to achieve high-accuracy binary classification with limited labeled positive examples and numerous unlabeled ones. Existing cost-sensitive-based methods often rely on strong assumptions that examples with an observed positive label were selected entirely at random. In fact, the uneven distribution of labels is prevalent in real-world PU problems, indicating that most actual positive and unlabeled data are subject to selection bias. Building on the SAR-PU propensity-weighted framework of Bekker et al., we study a PU learning enhancement (PUe) framework using normalized propensity scores and normalized inverse probability weighting (NIPW). PUe's main contributions are a normalized inverse-probability-weighted PU risk formulation; additional theoretical analyses of normalized sample-weight error and common PU estimators under biased labeling; regularized deep propensity-score estimation; integration with modern cost-sensitive PU methods; and support for selectively labeled negative classes. Experiments on MNIST, CIFAR-10, and ADNI demonstrate improvements over several PU baselines under non-uniform label distributions.

URL PDF HTML 收藏
2606.30248 2026-07-16 cs.CV cs.LG 版本更新

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation

你的数据流形实际上是一个奖励模型:Shell-LCC 用于文本到视频生成

Shihao Zhang, Yunzhi Li, Yuguang Yan, Junzhe Zhang, Wei Zhao, Bohan Wang, Hanwang Zhang

机构 * Huawei Central Research Institute(华为中央研究院) Guangdong University of Technology(广东技术大学)

AI总结 提出 Shell-LCC 方法,通过建模高质量 SFT 数据的流形结构,提供密集、可微且几乎零成本的奖励信号,以提升文本到视频生成质量,减少低层失真。

Comments ECCV 2026

详情
AI中文摘要

最近的文本到视频(T2V)扩散模型严重依赖辅助奖励信号(例如通过奖励模型或DPO)来使生成内容与人类美学对齐并提高真实感。然而,这些信号会带来大量的计算开销,需要昂贵的人工标注,并且通常在细粒度局部细节上改进有限。在本文中,我们认为你的数据流形实际上是一个奖励模型。通过显式建模高质量监督微调(SFT)数据的流形结构,并鼓励视频潜在变量位于该流形上,我们推导出密集、可微且几乎零成本的奖励信号,显著提高了视频质量,特别是在减轻低层失真方面。我们的建模基于局部坐标编码(LCC),它捕捉流形的“骨架”。然而,直接应用LCC会遭受均值回归,将潜在变量拉向几何均值并丢失高频细节。因此,我们将其扩展为壳局部坐标编码(Shell-LCC),它将流形“表面”建模为各向同性壳,以与真正的高密度区域对齐。实验表明,我们的方法提高了真实感,增强了高频细节,减少了过度平滑伪影,并减轻了运动模糊。

英文摘要

Recent text-to-video (T2V) diffusion models rely heavily on auxiliary reward signals (e.g., via reward models or DPO) to align generated content with human aesthetics and improve realism. These signals, however, incur substantial computational overhead, require costly human annotations, and often yield limited improvement in fine-grained local details. In this paper, we argue that your data manifold is secretly a reward model. By explicitly modeling the manifold structure of high-quality Supervised Fine-Tuning (SFT) data and encouraging video latents to lie on this manifold, we derive dense, differentiable, and nearly cost-free reward signals that significantly improve video quality, particularly in mitigating low-level distortions. Our modeling builds upon Local Coordinate Coding (LCC), which captures the `skeleton' of the manifold. However, directly applying LCC suffers from mean regression, pulling latents toward the geometric mean and losing high-frequency details. We therefore extend it to Shell Local Coordinate Coding (Shell-LCC), which models the manifold `surface' as an isotropic shell to align with the true high-density region. Experiments demonstrate that our approach improves realism, enhances high-frequency details, reduces over-smoothing artifacts, and alleviates motion blur.

URL PDF HTML 收藏
2510.24803 2026-07-16 cs.MA cs.AI 版本更新

MASPRM: Multi-Agent System Process Reward Model

MASPRM:多智能体系统过程奖励模型

Milad Yazdani, Mahdi Mostajabdaveh, Zirui Zhou, Ying Xiong

机构 * Department of Electrical and Computer Engineering, University of British Columbia(英属哥伦比亚大学电气与计算机工程系) Huawei Technologies Canada(华为技术加拿大公司)

AI总结 MASPRM通过过程奖励模型在多智能体系统推理中提升搜索效率和质量,提高Hit@1和排序质量。

详情
AI中文摘要

多智能体系统(MAS)的实用部署需要在测试时具有强大的性能,这推动了在推理过程中引导搜索并选择性地花费计算资源以提高质量的方法。我们提出了多智能体系统过程奖励模型(MASPRM)。该模型为每个动作和每个智能体对部分智能体转录文本分配值,并在推理过程中充当控制器。MASPRM通过将回报传播到局部目标,从多智能体蒙特卡洛树搜索(MCTS)的回放中进行训练,仅使用终端结果奖励进行标记,而不需要人类的步骤级注释。在推理过程中,MASPRM指导步骤级束搜索(SBS)和MCTS,将计算重点放在有希望的分支上,并修剪不值得的分支。我们跨不同的任务和领域训练和测试MASPRM,使用GSM8K、MATH、MMLU和LogiQA作为基准。在这些基准的平均表现中,MASPRM将Hit@1超过策略似然度提高了多达+13.4个点,并提高了排序质量,将Hit@1→Hit@5的差距减少了多达10.3个点。MASPRM通过评分中间路由的转录文本来补充推理时间的搜索,以引导具有固定时间表的MAS的回放。代码:https://github.com/milad1378yz/MASPRM

英文摘要

Inference-time search over multi-agent systems (MAS) wastes compute when it cannot identify which agent's intermediate message advanced progress. We present the Multi-Agent System Process Reward Model (MASPRM), which scores routed transcripts (ordered sequences of messages between agents) and acts as an inference controller for step-level beam search (SBS) and Monte Carlo Tree Search (MCTS). MASPRM is trained from multi-agent MCTS rollouts labeled only with terminal outcome rewards, without human step-level annotations. We evaluate on GSM8K, MATH, MMLU, and LogiQA. Under matched scorer size and comparable MCTS budget, MASPRM exceeds a size-matched ORM by $+2.0$ to $+3.0$ points at 1.5B and $+4.1$ to $+14.5$ at 7B across all four benchmarks, with additional scorer-scaling gains over policy likelihood at 7B (avg $+13.4$ under MCTS). MASPRM also improves ranking quality, reducing Hit@1 to Hit@5 gaps by up to $10.3$ points, with the largest gains under stepwise search that uses intermediate decisions. Code: https://github.com/milad1378yz/MASPRM

URL PDF HTML 收藏
2607.12746 2026-07-15 cs.CV 新提交

Color Pass-Through via Camera-Display Coupling

通过相机-显示器耦合实现颜色直通

Ruikang Li, Molin Li, Jiarui Wu, Zhe Wei, Pengpeng Liu, Tianfan Xue

机构 * CUHK MMLab(香港中文大学多媒体实验室) Zhejiang University(浙江大学) Central Media Technology Institute, Huawei(华为中央媒体技术研究所)

AI总结 研究智能手机相机捕捉场景与屏幕显示颜色差异问题,提出颜色直通的端到端学习框架,将相机和显示器视为耦合系统,经实验验证该方法能提升原始场景感知颜色的再现效果。

Comments 35 pages, 20 figures, including supplementary material. Project page: https://lyricccco.github.io/color-pass-through/

详情
AI中文摘要

当智能手机相机捕捉现实场景并在屏幕上显示时,显示图像在颜色、亮度和对比度上往往与原始场景有显著差异。尽管现代相机和显示器有很大进步,但这种差距依然存在。主要原因是大多数流程将高维的捕捉到显示过程分为两个单独校准的相机和显示阶段,通过低维颜色变换连接,导致信息瓶颈和误差积累。为解决这一系统性挑战,我们提出颜色直通,这是一个直接对捕捉图像操作的端到端学习框架。我们将相机和显示器视为耦合系统而非单独校准。耦合带来两个实际优势:通过端到端优化将整个现实场景带到显示器,为每个不同观察者通过完整的捕捉到显示路径进行高效一步校准。我们用数字和人类观察者验证了颜色直通。与代表性基线相比,我们的方法在5分制用户研究中平均增益2.0分,在定量指标上提高了2倍多,证明了对原始场景感知颜色的更好再现。

英文摘要

When a real-world scene is captured by a smartphone camera and viewed on its screen, the displayed image often differs noticeably from the original scene in color, brightness, and contrast. This gap persists despite substantial advances in both modern cameras and displays. A key reason is that most pipelines factor the high-dimensional capture-to-display process into two separately calibrated camera and display stages, and then connect them through low-dimensional color transforms, leading to information bottlenecks and inevitable error accumulation. To address this systemic challenge, we propose Color Pass-Through, an end-to-end learned framework that operates directly on captured images. Our key insight is to treat the camera and display as a coupled system rather than calibrating them in isolation. Coupling the camera and display yields two practical advantages: (1) it brings the entire real-world scenes to the display via end-to-end optimization, and (2) it allows efficient one-step calibration for each distinct observer via complete capture-to-display path. We validate Color Pass-Through using both digital and human observers. Compared with representative baselines, our method achieves an average gain of +2.0 points on a 5-point user study and more than 2x improvement on quantitative metrics, demonstrating improved reproduction of the perceived color of the original scene.

URL PDF HTML 收藏
2607.06233 2026-07-15 cs.AI 版本更新

Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale

展示TOFFEE:一个大规模合成数据代理轨迹的学习系统

Ziting Wang, Yin Li, Zuhao Yang, Xiuchang Li, Jiale Bai, Gao Cong

机构 * Nanyang Technological University(南洋理工大学) Huawei(华为) Industrial and Commercial Bank of China Limited(中国工商银行)

AI总结 针对现有数据代理难以适应新环境的问题,提出TOFFEE系统,通过蒙特卡洛树搜索等方法,能从给定数据环境合成高质量数据代理轨迹,可用于监督微调与上下文学习,还展示了系统框架、界面及工作流程与应用场景。

Comments Accepted to VLDB 2026

详情
AI中文摘要

由大语言模型驱动的数据代理在数据驱动的决策中发挥着越来越重要的作用。然而,现有数据代理难以推广到未见过的数据环境和分析工作流程,特别是在异构企业环境中。这使得合成高质量数据代理轨迹的需求日益增长,这些轨迹可用于监督微调数据和上下文学习演示。因此,我们引入了TOFFEE系统,它通过蒙特卡洛树搜索、自适应模型选择和跨任务前缀重用,从给定数据环境中合成高质量数据代理轨迹。我们展示了TOFFEE能够为异构环境中的复杂分析任务有效生成可扩展的轨迹数据。在本演示中,我们介绍了TOFFEE的系统框架,包括任务池构建、轨迹探索器和学习成本模型。我们还介绍了TOFFEE的网络界面及其工作流程,并展示了两个端到端场景:数据代理微调的轨迹合成和演示增强的数据代理推理。

英文摘要

LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle to generalize to unseen data environments and analytical workflows, especially in heterogeneous enterprise settings. This creates a growing need for synthesizing high-quality data agent trajectories that capture complex analytical workflows for given data environments. Such trajectories support two key downstream uses: they can serve as supervised finetuning (SFT) data that adapts data agent models to the target domain, and as in-context learning (ICL) demonstrations to guide general-purpose LLMs in unfamiliar data environments. Thus, we introduce TOFFEE, a system for synthesizing high-quality data agent trajectories from given data environments via Monte Carlo Tree Search (MCTS) with adaptive model selection and cross-task prefix reuse. We show that TOFFEE can effectively generate scalable trajectory data for complex analytical tasks across heterogeneous environments. In this demonstration, we present the system framework of TOFFEE, including its task pool construction, trajectory explorer, and learned cost model. We also introduce the web interface of TOFFEE and its workflow, and demonstrate two end-to-end scenarios: trajectory synthesis for data agent finetuning, and demonstration-augmented data agent reasoning.

URL PDF HTML 收藏
2607.11429 2026-07-14 cs.LG 新提交

Physics-Aware Conditional SetGAN for Spatially Consistent Multi-User TR 38.901 Channel Generation

用于空间一致多用户TR 38.901信道生成的物理感知条件集生成对抗网络

Mauro Gonzalo Tarazona-Levano, David Lopez-Perez, Nicola Piovesan, David Gomez-Barquero

机构 * Institute of Telecommunications and Multimedia Applications (iTEAM), Universitat Polit\`ecnica de Val\`encia (UPV), Spain(电信与多媒体应用研究所) Beihang Valencia Polytechnic Institute (BVPI), China(北京航空航天大学瓦大理工大学) Huawei Technologies, France(华为技术)

AI总结 研究能否用训练好的生成模型更快生成多用户TR 38.901信道且保持空间相关性,提出物理感知条件SetGAN,经训练可分离大尺度与小尺度衰落相关信息,在UMa/NLoS基准测试中大幅加速信道生成且保持空间一致性。

Comments Submitted to IEEE GLOBECOM 2026

详情
AI中文摘要

基于TR 38.901的信道模型如Sionna虽可靠,但生成多用户信道实现成本高。本文提出问题:训练好的生成模型能否比Sionna更快生成多用户TR 38.901信道且不丢失用户几何结构带来的空间相关性?为此提出物理感知、几何条件的SetGAN并在Sionna参考数据上训练。该方法分离大尺度接收功率与归一化小尺度衰落,用主成分分析压缩后者,在潜在空间学习条件信道分布并保留几何相关相关性。在UMa/NLoS基准测试中,模型保持接收功率分布与参考接近,再现空间一致性曲线。此外,相比Sionna,生成时间缩短3.45倍,CPU总成本降低6.15倍。结果表明训练好的生成模型可大幅加速TR 38.901信道生成且不破坏评估多用户系统所需的空间一致性。

英文摘要

TR 38.901-based channel models such as Sionna are reliable, but generating many multi-user channel realizations remains expensive. This paper asks a practical question: can a trained generative model produce multi-user TR 38.901 channels faster than Sionna without losing the spatial correlations imposed by user geometry? To answer this question, we propose a physics-aware, geometry-conditioned SetGAN trained on Sionna reference data. The method separates large-scale received power from normalized small-scale fading, compresses the latter with principal component analysis, and learns the conditional channel distribution in a latent space while preserving geometry-dependent correlations. On the UMa/NLoS benchmark, the model keeps the received-power distributions close to the reference, with about 0.41 dB Wasserstein distance, and reproduces spatial-consistency profiles with mean deviations below 0.03 on median curves versus distance. In addition, it reduces elapsed generation time by a factor of 3.45 and CPU-total cost by a factor of 6.15 relative to Sionna under matched user positions in the fixed-position CPU-vs-CPU benchmark. These results show that a trained generative model can substantially accelerate TR 38.901 channel generation without breaking the spatial consistency needed to evaluate multi-user systems.

URL PDF HTML 收藏
2607.11357 2026-07-14 cs.AI cs.SE 新提交

OpsMem: Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis

OpsMem:用于故障诊断的具有跨内存共振的双内存推理

Yongqian Sun, Rongchen Gao, Yu Luo, Wenwei Gu, Shenglin Zhang, Qingyi Guo, Qiuai Fu, Yaoliang Wu, Dan Pei

机构 * Huawei(华为)

AI总结 针对现代软件系统故障诊断中缺乏协调诊断状态与操作经验机制的问题,提出双内存框架OpsMem,通过跨内存共振等进行多智能体诊断,实验表明其在华为微服务故障诊断数据集上优于基线,提升了匹配度和相关性。

Comments 6 pages, 5 figures

详情
AI中文摘要

现代软件系统中的故障诊断需要通过操作经验进行迭代证据获取和假设推理。现有的基于大语言模型的方法通过智能体推理或知识增强来改进诊断,但在迭代诊断过程中往往缺乏将不断演变的诊断状态与操作经验相协调的机制。我们提出了OpsMem,一个双内存框架,它为当前诊断状态维护短期内存,为可重复使用的操作经验维护长期内存。OpsMem使用跨内存共振来激活与状态相关的长期内存,基于短期和激活的长期内存进行多智能体诊断,并将已解决事件中的可重复使用经验整合回长期内存。在真实世界的华为微服务故障诊断数据集上的实验表明,OpsMem优于代表性的智能体推理和知识增强基线,分别比最强基线在匹配度和相关性上提高了46.88%和18.39%。

英文摘要

Failure diagnosis in modern software systems requires iterative evidence acquisition and hypothesis reasoning guided by operational experience. Existing LLM-based methods improve diagnosis through agentic reasoning or knowledge augmentation, but they often lack a mechanism to coordinate the evolving diagnostic state with operational experience during iterative diagnosis. We propose OpsMem, a dual-memory framework that maintains a short-term memory for the current diagnostic state and a long-term memory for reusable operational experience. OpsMem uses cross-memory resonance to activate state-relevant long-term memory, conditions multi-agent diagnosis on the short-term and activated long-term memories, and consolidates reusable experience from solved incidents back into long-term memory. Experiments on a real-world Huawei microservice failure diagnosis dataset show that OpsMem outperforms representative agentic-reasoning and knowledge-augmented baselines, improving Match and Relevant by up to 46.88% and 18.39% over the strongest baseline, respectively.

URL PDF HTML 收藏
2607.10995 2026-07-14 cs.CV 新提交

AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling

AsySplat:用于长序列场景建模的高效非对称3D高斯点云渲染

Yingji Zhong, Dave Zhenyu Chen, Fuzhao Ou, Youyu Chen, Zhihao Li, Lanqing Hong, Dan Xu

机构 * HKUST(香港科技大学) Huawei Noah’s Ark Lab(华为诺亚方舟实验室) CityU(城市大学)

AI总结 研究针对长序列场景建模中3D高斯点云渲染的冗余计算问题,提出非对称架构解耦几何与外观建模,通过双边连接交互,减少计算冗余,提高参数效率,在32视图960P输入上大幅提升效率并超越零样本性能。

Comments The project page is at https://zhongyingji.github.io/asysplat/

详情
AI中文摘要

近期可推广的3D高斯点云渲染模型推动了长序列新视图合成(NVS)发展,但存在大量冗余计算。基于高精度几何对高质量NVS非严格必需及外观学习通常比几何恢复容易的观察,提出非对称架构解耦几何与外观建模。几何分支处理粗粒度令牌用于多视图重建,外观分支处理细粒度令牌捕捉细节,二者通过双边连接交互。该任务感知非对称减少计算冗余,更合理分配计算,提高参数效率,使小模型性能强劲。在32视图960P输入上,模型匹配基于优化的方法且加速近800倍,超越零样本性能,减少训练/推理开销,实现整体效率提升。

英文摘要

Recent generalizable 3D Gaussian Splatting models have advanced long-sequence novel view synthesis (NVS), but at the cost of substantial redundant computation. We identify that the redundancy can be mitigated based on two observations: (i) high-precision geometry is not strictly required for high-quality NVS; (ii) appearance learning is generally easier than geometry recovery. Motivated by these insights, we propose an asymmetric architecture that decouples geometry and appearance modeling. The geometry branch processes coarse-grained tokens with most of the parameters for multi-view reconstruction, while the appearance branch operates on fine-grained tokens to capture details using significantly fewer parameters. The two branches interact through bilateral connections, enabling mutual guidance for their respective tasks. This task-aware asymmetry reduces the computational redundancy and allocates the computation more judiciously, thereby increasing parameter efficiency and enabling smaller models to achieve strong performance. On 32-view 960P inputs, our model matches optimization-based methods while delivering nearly 800x speedup, and surpasses the zero-shot performance of state-of-the-art generalizable models with markedly fewer parameters and reduced training/inference overhead, achieving an overall efficiency improvement.

URL PDF HTML 收藏
2607.10911 2026-07-14 cs.CL cs.AI cs.DB 新提交

The Nuts and Bolts of Natural Language to SQL Translation: A Systematic Analysis of Model Pipeline Optimisation Approaches and their Interactions

自然语言到SQL翻译的关键要素:模型管道优化方法及其交互的系统分析

Filip Klubicka, Vasudevan Nedumpozhimana, Sneha Rautmare, Bora Caglayan, Mingxue Wang, John D. Kelleher

机构 * Huawei Ireland Research Centre(华为爱尔兰研究中心)

AI总结 研究NL2SQL翻译问题,通过集成NatSQL中间表示、增加预处理和微调步骤、开发重排器模型等方法,并结合SmBoP和RASAT架构进行消融及Shapley分析,揭示组件交互对结果的影响,为轻量级模型开发提供参考。

详情
AI中文摘要

在大语言模型时代,自然语言到SQL(NL2SQL)翻译仍是一个存在诸多实用应用的开放问题。我们探索了几种NL2SQL管道扩展之间的交互,以促进更轻量级模型的开发。具体而言,我们集成了NatSQL中间表示,纳入基于合成数据的预处理步骤和微调步骤,并开发了一种新颖的重排器模型以改进最终束搜索中的SQL选择。我们结合SmBoP和RASAT这两种骨干架构,对这些不同组件进行了消融研究并辅以Shapley分析。我们发现简单组合所有组件并不会带来最佳结果,其影响取决于它们与基线系统以及彼此之间的交互。

英文摘要

In the age of large language models, Natural Language to SQL (NL2SQL) translation remains an open problem with many useful applications. We explore interactions between several NL2SQL pipeline extensions to inspire development of more lightweight models. Specifically, we integrate the NatSQL intermediate representation, include a preprocessing step and a fine-tuning step based on synthetic data, and develop a novel reranker model to improve SQL selection in the final beam. We perform an ablation study supplemented by a Shapley analysis of these different components integrated with two backbone architectures, SmBoP and RASAT. We find that simply combining all of them does not lead to best results, but that their impact depends on their interactions with the baseline system, as well as each other.

URL PDF HTML 收藏
2607.10661 2026-07-14 cs.CL cs.AI 新提交

Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting

通过渐进式树状草稿的推测性解码解锁自回归语言模型中的并行性

Zipeng Gao, Zhi Zheng, Qingrong Xia, Junda Lin, Ziwei Zhao, Tong Xu, Zhefeng Wang, Enhong Chen

机构 * University of Science and Technology of China(中国科学技术大学) Huawei Technologies Co., Ltd.(华为技术有限公司)

AI总结 研究提出渐进式树状草稿(PTD)方法,通过结构化、引导式并行草稿策略利用模型并行潜力,结合渐进树结构与逐步修剪机制,在单次前向传播中引导LLM探索多条语义路径,实现高达2倍解码加速且无需训练、与模型无关。

详情
AI中文摘要

推测性解码通过缓解内存受限瓶颈,显著加速了大语言模型(LLM)的推理。然而,传统的推测性解码通常依赖于辅助草稿模块,会产生大量的训练和通信开销。尽管最近的方法试图在目标模型本身内部生成草稿,但由于缺乏结构协调,它们往往无法充分利用其潜在的并行能力。在本文中,我们提出了渐进式树状草稿(PTD),它采用结构化、引导式并行草稿策略来利用模型的并行潜力。通过将渐进树结构与逐步修剪机制相结合,PTD在单次前向传播中积极引导LLM探索多条语义路径,确保草稿的多样性和连贯性。实验表明,PTD在各种基准测试中实现了高达2倍的解码加速,同时无需训练且与模型无关。我们的代码可在以下网址获取:此https网址。

英文摘要

Speculative decoding has significantly accelerated Large Language Model (LLM) inference by alleviating memory-bound bottlenecks. However, traditional speculative decoding typically relies on auxiliary draft modules, incurring significant training and communication overhead. Although recent methods attempt to generate drafts within the target model itself, they often fail to fully exploit its latent parallel capacity due to a lack of structural coordination. In this paper, we propose \textbf{Progressive Tree Drafting (PTD)}, which employs a structured, guided parallel drafting strategy to harness the model's parallel potential. By coupling a progressive tree structure with a stepwise pruning mechanism, PTD actively guides the LLM to explore multiple semantic paths in a single forward pass, ensuring both draft diversity and coherence. Experiments demonstrate that PTD achieves up to $2\times$ decoding speedup across various benchmarks while remaining training-free and model-agnostic. Our code is available at: https://github.com/MINE-USTC/PTD.

URL PDF HTML 收藏
2607.09696 2026-07-14 cs.LG cs.AI 新提交

Mitigating Early Training Collapse in CTR Models

缓解CTR模型早期训练崩溃

Ergun Biçici, Erkan Çetinyamaç

机构 * Huawei Türkiye R&D Center(华为土耳其研发中心)

AI总结 研究针对点击率预测模型早期训练崩溃问题,通过大规模工业数据集分析,发现降低学习率效果不佳,而控制特征稀疏性如去除高度稀疏特征、聚合罕见特征值可稳定训练,提升离线和在线性能。

Comments 4 pages, 1 figure

详情
AI中文摘要

用于点击率预测的深度神经模型通常在第一个训练周期后验证性能急剧下降,尽管训练损失持续改善。这种不稳定性限制了有效学习和模型性能。本研究使用大规模工业数据集分析此行为并评估实际缓解策略。降低学习率效果有限,控制特征稀疏性有显著改善。去除高度稀疏特征和聚合罕见特征值可稳定训练,延长有效学习,提升离线评估指标和在线系统性能。

英文摘要

Deep neural models for click-through rate prediction often exhibit a sharp decline in validation performance immediately after the first training epoch despite continued improvement in training loss. This instability restricts effective learning and limits model performance. In this study, we analyze this behavior using large-scale industrial datasets and evaluate practical mitigation strategies. While reducing the learning rate provides only incremental gains, controlling feature sparsity yields substantial improvements. Removing highly sparse features and aggregating infrequent feature values stabilizes training, extends useful learning beyond a single epoch, and improves both offline evaluation metrics and online system performance.

URL PDF HTML 收藏
2607.10295 2026-07-14 physics.optics cs.AI 新提交

Program-Synthesis-Driven Autodesign of Universal Unitary Operators

程序合成驱动的通用酉算子自动设计

Yifei Zhang, Dong Chen, Fan Wang, Wenrui Zhang, Yan Chen, Dingding Han, Jianmin Yuan, Xiangjin Kong, Yu-Gang Ma

机构 * Key Laboratory of Nuclear Physics and Ion-beam Application (MOE), Institute of Modern Physics, Fudan University, Shanghai 200433, China(核物理与离子束应用重点实验室(教育部),现代物理研究所,复旦大学,上海200433,中国) Research Center for Theoretical Nuclear Physics, NSFC and Fudan University, Shanghai 200438, China(理论核物理研究中心,国家自然科学基金委员会和复旦大学,上海200438,中国) Huawei Technologies Co., Ltd, Beijing 100095, China(华为技术有限公司,北京100095,中国) Department of Industrial Engineering and Decision Analytics, Hong Kong University of Science and Technology, HongKong, China(工业工程与决策分析系,香港科技大学,香港,中国) Hunan Key Laboratory of Mechanism and Technology of Quantum Information, Changsha 410073, China(湖南量子信息机制与技术重点实验室,长沙410073,中国) School of Information Science and Technology, Fudan University, Shanghai 200433, China(信息科学与技术学院,复旦大学,上海200433,中国) Research Institute of Intelligent Complex Systems, Fudan University, Shanghai 200433, China(智能复杂系统研究所,复旦大学,上海200433,中国) Institute of Atomic and Molecular Physics, Jilin University, Changchun 130012, China(原子与分子物理研究所,吉林大学,长春130012,中国) School of Physics, East China Normal University, Shanghai 200062, China(物理学院,华东师范大学,上海200062,中国)

AI总结 研究利用人工智能驱动的程序合成,通过扩展DreamCoder到复值线性代数,发现光子网络中酉矩阵分解策略,其生成的程序编码与维度无关的不变量及构造规则,还能利用矩阵结构减少干涉仪数量,为酉算子自动设计提供可扩展范式。

详情
AI中文摘要

我们证明了人工智能驱动的程序合成可以自主发现光子网络中酉矩阵分解的基本策略。通过将DreamCoder扩展到复值线性代数,该系统生成的分解程序使用最少的$N(N - 1)/2$个马赫曾德尔干涉仪,不同于雷克和克莱门茨架构。学习到的程序编码与维度无关的不变量,为$5×5$矩阵发现的策略可推广到更高维度如$64×64$。发现的程序编码可解释的、与维度无关的构造规则,无需重新训练即可跨矩阵大小推广。该系统还能利用矩阵结构将干涉仪数量减少到通用理论界限以下,如对Householder矩阵发现仅需$2N - 3$个MZIs的规则,对稀疏矩阵奇异值分解得到的矩阵,在95%稀疏度时比通用理论界限少38%的MZIs。这些减少直接转化为可扩展光子实现的实际硬件优势。总之,该系统作为一个统一引擎,能发现通用分解规则和特定矩阵优化,无需输入矩阵的结构或分析属性。

英文摘要

We demonstrate that AI-driven program synthesis can autonomously discover fundamental strategies for decomposing unitary matrices in photonic networks. By extending DreamCoder to complex-valued linear algebra, the system generates decomposition programs achieving the minimal $N(N-1)/2$ Mach-Zehnder interferometers, distinct from both Reck and Clements architectures. Learned programs encode dimension-agnostic invariants: strategies discovered for $5 \times 5$ matrices generalize to higher dimensions such as $64 \times 64$. The discovered programs encode interpretable, dimension-agnostic construction rules. These rules generalize across matrix sizes without retraining, demonstrating that autonomous program synthesis can serve as a scalable paradigm for algorithm discovery and the automated design of universal unitary operators. Beyond universal decompositions, the system automatically exploits matrix structure to reduce the interferometer count below the universal theoretical bound. For instance, for Householder matrices, it discovers a dimension-independent rule that requires only $2N-3$ MZIs. This achieves linear, rather than quadratic, scaling and generalizes to arbitrary $N$ without retraining. For matrices obtained from the singular value decomposition of sparse matrices, reductions generally increase with sparsity, reaching up to 38% fewer MZIs than the universal theoretical bound $N(N-1)/2$ at 95% sparsity. These MZI reductions translate directly into practical hardware benefits for scalable photonic implementations. Taken together, the system functions as a single unified engine that discovers both universal decomposition rules and matrix-specific optimizations, without being provided with the structural or analytical properties of the input matrices.

URL PDF HTML 收藏
2602.00620 2026-07-14 cs.LG cs.AI 版本更新

Rethinking Zero-Shot Time Series Classification: From Task-specific Classifiers to In-Context Inference

重新思考零样本时间序列分类:从特定任务分类器到上下文推理

Juntao Fang, Shifeng Xie, Shengbin Nie, Yuhui Ling, Yuming Liu, Zijian Li, Keli Zhang, Lujia Pan, Themis Palpanas, Ruichu Cai

机构 * Guangdong University of Technology, Guangzhou, China(广东工业大学) Paris Descartes University, Paris, France(巴黎笛卡尔大学) Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, United Arab Emirates(马尔代夫 bin Zayed 人工智能大学) Huawei Noah’s Ark Lab, Paris, France(华为诺亚实验室) Huawei Noah’s Ark Lab, Shenzhen, China(华为诺亚实验室)

AI总结 该研究针对时间序列基础模型零样本分类中存在的问题,提出TIC-FM上下文学习框架,将训练集作上下文,单次前向传播预测标签,无需参数更新,经实验验证其在多数据集上表现良好,实现无训练迁移。

详情
AI中文摘要

时间序列基础模型(TSFMs)用于分类的零样本评估通常使用冻结编码器和特定任务分类器。然而,这种做法违反了零样本部署无需训练的前提,并因依赖分类器的训练选择而引入评估偏差。为解决此问题,我们提出TIC-FM,一种上下文学习框架,将标记训练集视为上下文,在单次前向传播中预测所有测试实例的标签,无需参数更新。TIC-FM将时间序列编码器和轻量级投影适配器与分割掩码潜在记忆Transformer配对。我们进一步提供理论依据,证明上下文推理可包含训练好的分类器,并能在单次前向传播中模拟基于梯度的分类器训练。在128个UCR数据集上的实验显示出高准确率,在极低标签情况下有持续提升,突出了时间序列的无训练迁移。

英文摘要

The zero-shot evaluation of time series foundation models (TSFMs) for classification typically uses a frozen encoder followed by a task-specific classifier. However, this practice violates the training-free premise of zero-shot deployment and introduces evaluation bias due to classifier-dependent training choices. To address this issue, we propose TIC-FM, an in-context learning framework that treats the labeled training set as context and predicts labels for all test instances in a single forward pass, without parameter updates. TIC-FM pairs a time series encoder and a lightweight projection adapter with a split-masked latent memory Transformer. We further provide theoretical justification that in-context inference can subsume trained classifiers and can emulate gradient-based classifier training within a single forward pass. Experiments on 128 UCR datasets show strong accuracy, with consistent gains in the extreme low-label situation, highlighting training-free transfer for time series classification.The source code is publicly available at https://github.com/fangjuntao/TIC-FM.

URL PDF HTML 收藏
2511.23231 2026-07-14 cs.CV 版本更新

Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering

通过表示工程解锁大语言模型和大视觉语言模型的多语言推理能力

Qiming Li, Xiaocheng Feng, Yixuan Ma, Zekai Ye, Ruihan Chen, Xiachong Feng, Bing Qin

机构 * Harbin Institute of Technology(哈尔滨工业大学) Peng Cheng Laboratory(鹏城实验室) The University of Hong Kong(香港大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳分校) Huawei Technologies Co., Ltd(华为技术有限公司)

AI总结 研究针对LLMs和LVLMs在多语言推理中英语表现优于低资源语言的问题,提出无训练的推理时方法MRRE,通过在特定层注入两个预计算向量增强多语言推理能力,实验证明该方法有效提升非英语推理及输入输出语言一致性。

Comments ACL2026 main

详情
AI中文摘要

大语言模型(LLMs)和大视觉语言模型(LVLMs)虽展现出强大推理能力,但在英语上的表现远超低资源语言,引发多语言应用中的公平性问题。现有方法或依赖昂贵的多语言训练,或借助外部翻译工具提示,资源密集且受翻译质量影响。为此,我们提出一种无训练的推理时方法MRRE,通过表示工程增强多语言推理能力,无需额外训练数据或工具。MRRE在推理过程中特定层顺序注入两个预计算向量:跨语言推理增强向量引导非英语推理表示进入英语空间以解锁多语言推理,目标语言输出锚定向量恢复目标语言分布以保持输入输出语言一致性。在四个推理基准上对六个先进的LLMs和LVLMs进行的综合实验表明,MRRE持续增强非英语推理,低资源语言(泰语和斯瓦希里语)平均提升5.48%,最高达7.54%,同时将输入输出语言一致性提高3.78%。

英文摘要

Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) demonstrate strong reasoning capabilities, yet their performance in English significantly outperforms that in low-resource languages, raising fairness concerns in multilingual applications. Existing approaches either rely on costly multilingual training or employ prompting with external translation tools, both of which are resource-intensive and sensitive to translation quality. To address these limitations, we propose a training-free inference-time method to enhance Multilingual Reasoning capabilities via Representation Engineering (MRRE) without using any additional training data or tools. MRRE sequentially injects two precomputed vectors at specific layers during inference processing: cross-lingual reasoning enhancement vectors, which steer non-English reasoning representations toward English space to unlock multilingual reasoning, and target-language output anchoring vectors, which restore the distribution of the target language to preserve input-output language consistency. Comprehensive experiments across six advanced LLMs and LVLMs on four reasoning benchmarks demonstrate that MRRE consistently enhances non-English reasoning by an average gain of 5.48% and up to 7.54% in low-resource languages (Thai and Swahili), while improving input-output language consistency by 3.78%.

URL PDF HTML 收藏
2508.10123 2026-07-14 cs.LG cs.AI cs.CL 版本更新

Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts

Nested-ReFT:通过离策略展开实现大语言模型微调的高效强化学习

Maxime Heuillet, Yufei Cui, Boxing Chen, Audrey Durand, Prasanna Parthasarathi

机构 * Mila - Québec AI Institute, Canada(魁北克人工智能研究所) Huawei Noah's Ark Lab (Montreal Research Center), Canada(华为诺亚实验室(蒙特利尔研究中心)) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

AI总结 研究针对大语言模型在数学推理等领域的高级推理难题,提出Nested-ReFT框架,利用离策略展开及动态层跳过降低训练推理成本,经理论与实证分析验证其有效性,还探索偏差缓解变体以维持性能。

详情
AI中文摘要

在诸如数学推理等具有挑战性的领域中,大语言模型的高级推理可通过基于可验证奖励的强化微调(ReFT)来解决。在标准ReFT框架中,行为模型针对每个问题生成多个带有答案的完成结果,然后由奖励函数对答案进行评分。虽然这种强化学习后训练方法在具有挑战性的推理领域中显示出显著性能提升,但训练期间通过多个推理步骤生成完成结果的计算成本使得训练成本颇高。为解决此问题,我们从离策略强化学习和推测性解码中汲取灵感,引入了一种新颖的ReFT框架,称为Nested-ReFT,其中目标模型的部分层在训练期间充当行为模型以生成离策略完成结果。与标准ReFT框架相比,训练期间按批次配置动态层跳过的行为模型降低了推理成本。我们的理论分析表明,Nested-ReFT产生具有可控方差的无偏梯度估计。我们的实证分析表明,在多个数学推理基准和模型规模上,以每秒令牌数衡量的计算效率有所提高。此外,我们探索了三种偏差缓解变体,以最小化梯度更新中的离策略性,从而保持与基线ReFT性能相匹配的性能。

英文摘要

Advanced reasoning in LLMs on challenging domains like mathematical reasoning can be tackled using verifiable rewards based reinforced fine-tuning (ReFT). In standard ReFT frameworks, a behavior model generates multiple completions with answers per problem, for the answer to be then scored by a reward function. While such RL post-training methods demonstrate significant performance improvements across challenging reasoning domains, the computational cost of generating completions during training with multiple inference steps makes the training cost non-trivial. To address this, we draw inspiration from off-policy RL, and speculative decoding to introduce a novel ReFT framework, dubbed Nested-ReFT, where a subset of layers of the target model acts as the behavior model to generate off-policy completions during training. The behavior model configured with dynamic layer skipping per batch during training decreases the inference cost compared to the standard ReFT frameworks. Our theoretical analysis shows that Nested-ReFT yields unbiased gradient estimates with controlled variance. Our empirical analysis demonstrates improved computational efficiency measured as tokens/sec across multiple math reasoning benchmarks and model sizes. Additionally, we explore three variants of bias mitigation to minimize the off-policyness in the gradient updates that allows for maintaining performance that matches the baseline ReFT performance.

URL PDF HTML 收藏