arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

National University of Singapore(新加坡国立大学)

至 收录 2372
2607.18228 2026-07-21 cs.AI cs.CL 新提交

Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes

压力下的逻辑判断:用学习到的软前缀诊断三段论稳定性

Brian K Chen

机构 * National University of Singapore(新加坡国立大学)

AI总结 研究在三段论推理基准前加软前缀,探究学习到的上下文压力对模型逻辑判断的影响,发现软前缀能改变正确答案,在多模型测试中效果优于随机对照,主要影响是答案偏好,不同模型有显著差异。

Comments 41 pages, 6 figures

详情
AI中文摘要

为测试正确的逻辑判断如何响应学习到的上下文,我们在一个精确标注的三段论推理基准前添加软前缀,同时保持模型不变。软前缀是不透明的连续向量,通过它们在逻辑形式和接口的受控变化中引发的行为来表征。通过研究哪些前缀成功及其效果如何泛化,我们刻画了学习到的上下文压力如何推翻正确判断并揭示模型逻辑稳定性的局限性。在多个模型上,学习到的前缀改变了许多正确答案,在未见过的形式和接口变化中依然有效,且在多次测试中优于随机对照。诊断测试表明,主要影响是对一种答案含义的广泛偏好,不同模型的这种偏差形式不同。这些结果表明,成功的软前缀的主要行为影响是广泛的答案偏好,同时其余响应揭示了逻辑稳定性方面模型特定的显著差异。

英文摘要

To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syllogistic reasoning benchmark while keeping the model fixed. Soft prefixes are opaque continuous vectors, so we characterize them through the behavior they induce across controlled variations in logical form and interface. By studying which prefixes succeed and how their effects generalize, we characterize how learned contextual pressure can override correct judgments and expose limits in a model's logical stability. Across Qwen3.6-35B-A3B MoE, Qwen3-8B, and Gemma 4 31B, learned prefixes redirect many correct answers and remain effective across unseen forms and interface changes. In repeated tests with Qwen3.6 MoE and Gemma, they outperform paired random controls in all 16 model--direction--split comparisons by 37 to 99 percentage points. Qwen3.6 MoE flip rates remain between 72% and 90% across wording and prompt changes, while Gemma validity prefixes retain 54% to 56% flip compared with less than 1% for matched random prefixes. Diagnostic tests show that the dominant effect is a broad preference for one answer meaning rather than fixed-symbol forcing or a logical operation that transfers reliably between tasks. The form of this bias differs across models. In both Qwen models, simple score models often predict which judgments will flip but not how far their margins will move, whereas Gemma's overall response is more closely approximated by the same models. These results show that the dominant behavioral effect of successful soft prefixes is a broad answer preference, while the remaining response reveals substantial model-specific differences in logical stability.

URL PDF HTML 收藏
2607.18072 2026-07-21 cs.LG cs.AI 新提交

SGN: A Similarity-based Generative Network for Data Generation under Distribution Shift

SGN:一种用于分布偏移下数据生成的基于相似度的生成网络

Jiaqi Zhu, Xincheng Chen, Yuncheng Wu, Zhaojing Luo, Beng Chin Ooi

机构 * National University of Singapore(新加坡国立大学) Renmin University of China(中国人民大学) Beijing Institute of Technology(北京理工大学) Zhejiang University(浙江大学)

AI总结 研究分布偏移下数据生成问题,提出基于相似度的生成网络SGN,在源数据训练后无需参数更新用于新目标域,通过学习潜在空间和利用目标域小代表性集,实现目标引导数据增强,实验验证其有效性。

详情
AI中文摘要

在源域上训练的生成模型生成的样本与偏移后的目标域往往匹配不佳,限制了其在目标域数据增强方面的有效性。虽然特定于目标的适应可以减少这种不匹配,但通常需要额外的优化和特定于域的参数。我们提出了一种基于相似度的生成网络(SGN),这是一个可重复使用的框架,在有标签的源数据上训练一次,无需参数更新即可应用于新的目标域。SGN学习由标签诱导的成对相似度构建的潜在空间,同时通过编码器-解码器架构保留重构信息。在生成时,对来自目标域的一个小的有标签代表性集进行编码,并在学习到的潜在空间中进行组合,使生成的样本在保持类一致性的同时继承目标特定特征。我们进一步分析了所提出的相似度结构的可实现性和维度要求。在图像和表格数据集上的实验证明了SGN在源到目标分布偏移下进行目标引导数据增强的有效性。

英文摘要

Generative models trained on a source domain often produce samples that are poorly aligned with shifted target domains, limiting their effectiveness for target-domain data augmentation. Although target-specific adaptation can reduce this mismatch, it typically requires additional optimization and domain-specific parameters. We propose a Similarity-based Generative Network (SGN), a reusable framework that is trained once on labeled source data and applied to new target domains without parameter updates. SGN learns a latent space structured by label-induced pairwise similarities while preserving reconstructive information through an encoder-decoder architecture. At generation time, a small labeled representative set from the target domain is encoded and combined in the learned latent space, allowing the generated samples to inherit target-specific characteristics while maintaining class consistency. We further analyze the realizability and dimensionality requirements of the proposed similarity structure. Experiments on image and tabular datasets demonstrate the effectiveness of SGN for target-guided data augmentation under source-to-target distribution shifts.

URL PDF HTML 收藏
2607.16873 2026-07-21 cs.CV 新提交

InfoDense: Density-Aware Regional Decisive Replay for Memory-Efficient Incremental Face Forgery Detection

InfoDense:用于内存高效增量式人脸伪造检测的密度感知区域决定性重放

Jikang Cheng, Hao Shen, Xueyi Zhang, Guangcheng Wang, Zhongyuan Wang, Renye Yan, Baojin Huang

机构 * Huazhong Agricultural University(华中农业大学) Peking University(北京大学) National University of Singapore(新加坡国立大学) Nantong University(南通大学) Wuhan University(武汉大学)

AI总结 针对人脸伪造技术发展带来的检测挑战及传统方法的问题,提出密度感知区域决定性重放策略InfoDense,通过定位决定性补丁、排序候选片段、自适应合并样本等步骤,有效减轻灾难性遗忘,提升跨域泛化能力。

详情
AI中文摘要

人脸伪造技术的快速发展带来了越来越多的操纵手段。增量式人脸伪造检测(IFFD)作为一种应对不断演变的伪造威胁的有前途的方法应运而生,它通过增量添加新的伪造数据来微调先前训练的模型。然而,传统的基于重放的IFFD方法存在灾难性遗忘问题。在有限内存下存储完整历史图像往往无法保留细微的伪造线索或引入域偏差,降低了模型学习内在和可转移操纵特征的能力。本文提出了一种密度感知区域决定性重放策略InfoDense来应对这些挑战。InfoDense优先考虑伪影密集和伪造关键区域,在保持高保真伪造证据的同时显著降低存储需求。首先引入InfoDense Cut使用基于CLIP的嵌入定位决定性补丁,然后InfoDense Select通过结合潜在空间代表性和决定性补丁数量对候选片段进行排序,确保重放缓冲区的多样性和信息密度,最后InfoDense Fuse通过将存储的片段与当前任务样本自适应合并来重建无偏训练输入,增强知识保留和泛化能力。在具有挑战性的增量深度伪造基准上的大量实验表明,InfoDense有效地减轻了灾难性遗忘,同时提高了跨域泛化能力。

英文摘要

The rapid evolution of face forgery techniques has introduced an increasing variety of manipulations. Incremental Face Forgery Detection (IFFD), which incrementally adds new forgery data to fine-tune previously trained models, has emerged as a promising approach to handle evolving forgery threats. However, conventional replay-based IFFD methods suffer from catastrophic forgetting. Storing full historical images under limited memory often either fails to preserve subtle forgery cues or introduces domain bias, reducing the model's ability to learn intrinsic and transferable manipulation characteristics. In this paper, we propose a Density-Aware Regional Decisive replay strategy, termed InfoDense, to address these challenges. InfoDense prioritizes artifact-dense and forgery-critical regions, significantly reducing storage requirements while maintaining high-fidelity forgery evidence. We first introduce InfoDense Cut to localize decisive patches using CLIP-based embeddings. Then, InfoDense Select ranks candidate segments by combining latent-space representativeness and decisive patch counts, ensuring both diversity and information density in the replay buffer. Finally, InfoDense Fuse reconstructs unbiased training inputs by adaptively merging stored segments with current-task samples, enhancing knowledge retention and generalization. Extensive experiments on challenging incremental deepfake benchmarks demonstrate that InfoDense effectively mitigates catastrophic forgetting while improving cross-domain generalization.

URL PDF HTML 收藏
2607.16806 2026-07-21 cs.RO 新提交

Token-Wise Latent Streaming from Slow Reasoners to Fast Planners for Dynamic Vision Language Navigation

用于动态视觉语言导航的从慢速推理器到快速规划器的逐令牌潜在流

Tianshuai Hu, Yangyi Zhong, Zeying Gong, Lingdong Kong, Xiaodong Mei, Guoyang Zhao, Xiaolu Liu, Song Wang, Rong Li, Junwei Liang

机构 * The Hong Kong University of Science and Technology(香港科技大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) National University of Singapore(新加坡国立大学) Zhejiang University(浙江大学)

AI总结 针对动态视觉语言导航中语言推理慢与规划需即时的矛盾,提出SPARK-VLN双系统框架,通过三个模块将慢速VLM推理器知识流到快速规划器,引入新基准套件,提高了导航成功率、社会合规性及推理效率。

详情
AI中文摘要

在动态、以人类为中心的环境中的视觉语言导航存在一个基本矛盾:语言推理缓慢且深思熟虑,而安全、符合社会规范的规划应该即时且具有反应性。由此产生的观测陈旧性对安全至关重要:推理过程中选择的动作在执行时可能已经不安全。我们观察到,在VLM完成推理之前很久,其中间隐藏状态就已经编码了与动作相关的意图。我们提出了SPARK-VLN,这是一个用于动态社会VLN的双系统框架,在整个令牌生成过程中将慢速VLM推理器的知识流到快速流匹配专家规划器,在推理过程中提供新的和不断演变的指导。该设计由三个模块实现:一个逐令牌隐藏流提取器,一个序列到插槽潜在桥接器,一个不断演变的潜在调节器。我们还引入了一个用于动态社会视觉语言导航的以人类为中心的基准套件,该套件在整个推理过程中使行人和机器人保持活跃,并报告导航成功、社会合规性、人类碰撞和明确的陈旧性统计数据。在这些设置中,SPARK-VLN提高了导航成功率和社会合规性,并保持了推理效率。

英文摘要

Vision-Language Navigation in dynamic, human-centric environments exposes a fundamental tension: linguistic reasoning is slow and deliberative, whereas safe, socially compliant planning should be instant and reactive. The resulting observation staleness is safety-critical: a maneuver chosen during inference can already be unsafe by the time it executes. We observe that, long before a VLM finishes its inference, its intermediate hidden states already encode action-relevant intent. We propose SPARK-VLN, a dual-system framework for dynamic social VLN that streams the slow VLM reasoner's knowledge to a fast flow-matching expert planner throughout token generation, providing fresh and evolving guidance during inference. This design is realized by three modules: a Token-Wise Hidden Streamer that extracts intermediate hidden states along the token generation process, a Sequence-to-Slot Latent Bridge that projects them into fixed-size latent slots, and an Evolving Latent Conditioner that infuses them into the expert planner. We also introduce a human-centric benchmark suite for dynamic social vision-language navigation that keeps pedestrians and the robot active throughout inference and reports navigation success, social compliance, human collisions, and explicit staleness statistics. Across these settings, SPARK-VLN mproves navigation success and social compliance while sustaining inference efficiency. Webpage: https://hutslib.github.io/SPARK-VLN/.

URL PDF HTML 收藏
2607.16577 2026-07-21 cs.CV cs.GR 新提交

CNS-Edit++: Category-Agnostic 3D Editing with Coupled Neural Shape Representation

CNS-Edit++:基于耦合神经形状表示的类别无关3D编辑

Jingyu Hu, Weilong Yan, Zhengzhe Liu, Haipeng Li, Ka-Hei Hui, Hao, Zhang, Chi-Wing Fu

机构 * The Chinese University of Hong Kong(香港中文大学) Lingnan University(岭南大学) National University of Singapore(新加坡国立大学) The Hong Kong University of Science and Technology(香港科技大学) Autodesk AI Lab(欧特克人工智能实验室) Simon Fraser University(西蒙弗雷泽大学)

AI总结 研究提出基于耦合神经形状表示和神经特征体积优化的潜在空间3D形状编辑框架CNS-Edit++,能在特定类别和类别无关模型上实例化,有多种编辑操作符及区域控制机制,经评估其性能优于现有方法。

详情
AI中文摘要

本文提出了一个基于耦合神经形状(CNS)表示和神经特征体积优化的潜在空间3D形状编辑框架。该工作将基于耦合神经形状优化的CNS-Edit扩展到CNS-Edit++,通过将特定类别的耦合表示推广到使用基础模型的类别无关3D形状编辑。耦合神经形状(CNS)表示将捕获高级形状语义的全局潜在代码与为局部形状操作提供空间上下文的3D神经特征体积耦合。然后制定了一个耦合神经形状优化过程,以根据给定的编辑操作共同优化这两个组件。该框架可以在特定类别的3D反演模型和类别无关的3D基础模型上实例化。提供了各种形状编辑操作符,并引入两种互补的区域控制机制以保留编辑区域外的区域。不同3D生成模型的广泛定量和定性评估证明了该方法优于现有解决方案的强大能力。

英文摘要

This paper presents a latent-space 3D shape editing framework built upon a coupled neural shape (CNS) representation and a neural feature volume optimization. This work extends CNS-Edit, built on Coupled Neural Shape optimization, to CNS-Edit++, by generalizing the category-specific coupled representation to category-agnostic 3D shape editing with foundation models. The Coupled Neural Shape (CNS) representation couples a global latent code that captures high-level shape semantics with a 3D neural feature volume that provides spatial context for local shape manipulation. Then we formulate a coupled neural shape optimization procedure that co-optimizes these two components subject to a given editing operation. Our framework can be instantiated on both the category-specific 3D inversion model and category-agnostic 3D foundation models. We provide various shape editing operators, including copy, resize, delete, mix, point-wise drag, and region-wise drag, each of which is formulated as an objective to guide the CNS optimization. To preserve regions outside the editing area, we further introduce two complementary region-wise control mechanisms, i.e., KV-cache replacement and latent feature regularization. Extensive quantitative and qualitative evaluations across different 3D generative models demonstrate the strong capabilities of our approach over state-of-the-art solutions.

URL PDF HTML 收藏
2607.16303 2026-07-21 cs.CV cs.AI 新提交

Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation

Med-OPD:通过证据感知策略蒸馏改进医学视觉语言模型

Yunhang Qian, Jiaquan Yu, Jiawei Liu, Meng Wang, Hongwei Bran Li, Xiaobin Hu

机构 * National University of Singapore(新加坡国立大学) University of Science and Technology of China(中国科学技术大学)

AI总结 研究针对医学视觉语言模型依赖语言先验而非视觉证据推理的问题,提出Med-OPD框架,引入医学证据优势信号,在令牌和轨迹级别重新分配蒸馏信号,实验证明该方法能加强模型对关键视觉证据的依赖,提升多模态医学推理能力。

详情
AI中文摘要

医学视觉语言模型(Med-VLMs)需要从细粒度视觉证据进行可靠推理,但现有模型常依赖语言先验或医学模板给出看似合理的临床答案,而非真正关注关键诊断区域。策略蒸馏(OPD)能对学生生成轨迹进行密集令牌级监督且隐私兼容。然而,标准OPD均匀蒸馏所有令牌,使依赖证据的稀疏令牌被大量临床叙述令牌稀释。受OPD在大语言模型社区成功的启发,我们提出Med-OPD,这是首个将策略蒸馏与医学证据感知监督集成的统一训练后框架。我们引入医学证据优势(MEA),一种基于教师的反事实信号,通过答案感知提示聚焦教师对支持目标诊断证据的评分,并通过比较原始和证据退化成像模式下教师的可能性来衡量每个令牌对医学视觉证据的依赖。基于MEA,Med-OPD在令牌和轨迹级别重新分配蒸馏信号,强调关键诊断令牌和依赖证据的展开。在OmniMedVQA子集上的实验表明,Med-OPD在CT、MRI、疾病诊断和病变分级方面始终优于SFT和标准OPD。这些结果表明,证据感知蒸馏可以更好地加强医学VLMs对关键视觉证据的依赖,提高可靠的多模态医学推理能力。源代码和数据可在指定链接公开获取。

英文摘要

Medical Vision-Language Models (Med-VLMs) require reliable reasoning from fine-grained visual evidence, yet existing models can produce plausible clinical answers by relying on language priors or medical templates rather than truly attending to diagnosis-critical regions. On-Policy Distillation (OPD) offers dense token-level supervision on student-generated trajectories and provides a privacy-compatible means of capability transfer without requiring the redistribution of raw patient data. However, standard OPD uniformly distills all tokens, causing sparse evidence-dependent tokens to be diluted by abundant clinical narrative tokens. Inspired by the success of OPD in the large language model community, we propose \textbf{Med-OPD}, to our knowledge the first unified post-training framework that integrates on-policy distillation with medical evidence-aware supervision for Med-VLMs. We introduce \textbf{Medical Evidence Advantage} (MEA), a teacher-grounded counterfactual signal that uses an answer-aware hint to focus teacher scoring on evidence supporting the target diagnosis, and measures each token's dependence on medical visual evidence by comparing teacher likelihoods under the original and evidence-degraded imaging modalities. Based on MEA, Med-OPD redistributes the distillation signal at both the token and trajectory levels, emphasizing diagnosis-critical tokens and evidence-reliant rollouts. Experiments on OmniMedVQA subsets show that Med-OPD consistently outperforms SFT and standard OPD across CT, MRI, Disease Diagnosis, and Lesion Grading. These results demonstrate that evidence-aware distillation can better strengthen medical VLMs' reliance on key visual evidence and improve reliable multimodal medical reasoning. The source code and data is publicly available at: https://github.com/yunhang8658/MedOPD.git

URL PDF HTML 收藏
2607.16295 2026-07-21 cs.CV cs.AI 新提交

Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm

基于群对比前向算法的涌现分层单语义神经元

Yiming Tang, Qinglin Qi, Zhaoqian Yao, Harshvardhan Saini, Dianbo Liu

机构 * National University of Singapore(新加坡国立大学) Lund University(隆德大学) Chinese University of Hong Kong(香港中文大学) Indian Institute of Technology, Dhanbad(印度理工学院(丹巴德分校))

AI总结 研究探讨神经网络表示的可解释性,针对稀疏字典学习范式的局限,提出群对比前向算法GCFF,通过架构约束实现单语义性,能捕捉非线性概念,在CLIP表示上表现良好,还能从头训练网络并在图像分类基准中达最优性能。

详情
AI中文摘要

机械可解释性在理解神经网络表示方面取得了显著进展,稀疏字典学习(SDL)方法是核心范式,但存在局限性。我们假设存在不同的单语义性途径,生物视觉系统有高度选择性的神经元分层组织,源于局部、逐层学习规则。为此提出群对比前向算法(GCFF),通过架构约束而非稀疏性实现单语义性,能捕捉非线性概念。在CLIP表示上,单个训练的GCFF模块可恢复抽象度随深度递增的单语义神经元,且无需稀疏约束或抽象级别监督。此外,GCFF能从头训练网络,在各种图像分类基准上达到前向算法的最优性能。

英文摘要

Mechanistic interpretability has made significant strides in understanding neural network representations, with sparse dictionary learning (SDL) methods, most prominently sparse autoencoders, as a central paradigm. However, recent work has reported several limitations of this paradigm: SDL objectives are non-identifiable; SDL methods rely heavily on the Linear Representation Hypothesis; and a growing body of evidence points to concepts that are encoded non-linearly and are therefore not expressible as any single direction. We hypothesise that a different route to monosemanticity is available. Biological visual systems exhibit highly selective neurons organised into hierarchies of increasing abstraction, and this organisation emerges from local, layer-wise learning rules rather than from a global error signal; we therefore ask whether a biologically plausible learning algorithm will likewise yield monosemantic neurons. To test this, we propose Group-Contrastive Forward-Forward (GCFF), a forward-forward training algorithm that combines class-specific routing with within-class contrastive objectives, reaching monosemanticity through architectural constraints rather than sparsity. Because GCFF attaches multiple non-linear layers to the representation under study, its neurons can therefore capture the non-linear concepts. On CLIP representations, a single trained GCFF module recovers monosemantic neurons whose abstraction increases progressively with depth, reaching environmental properties that hold independently of an image's foreground, without any sparsity constraint or supervision of abstraction level. We further demonstrate that GCFF can train networks from scratch, achieving state-of-the-art performance among forward-forward algorithms on various image classification benchmarks.

URL PDF HTML 收藏
2607.16251 2026-07-21 cs.LG 新提交

Learning Spatio-Temporal Foundation Models from Pure Synthetic Data

从纯合成数据中学习时空基础模型

Yutong Feng, Shiyuan Piao, Yutong Xia, Xu Liu, Wenqi Fan, Fugee Tsung, See-Kiong Ng, Yuxuan Liang

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) National University of Singapore(新加坡国立大学) Hong Kong University of Science and Technology(香港科技大学) Hong Kong Polytechnic University(香港理工大学)

AI总结 研究旨在学习时空基础模型,提出NeoST,通过在程序生成的合成系统上预训练,引入可扩展语料库、潜在空间推理架构和目标,实验证明其在多样真实世界时空系统中性能优越,有长期稳定性和推理效率。

详情
AI中文摘要

时空基础模型(STFMs)旨在学习复杂动力系统在时空上的可泛化表示。现有方法存在诸多问题,如真实世界预训练数据的分布偏差、自回归或基于扩散范式的结构瓶颈以及过度强调噪声观测中点状重建的目标。本文提出了NeoST,首个仅在程序生成的合成系统上预训练的时空基础模型。它引入可扩展合成预训练语料库减轻真实世界偏差,有潜在空间推理架构及潜在空间目标。实验表明NeoST在多样真实世界时空系统中优于现有模型,具有卓越的长期稳定性和推理效率。

英文摘要

Spatio-Temporal Foundation Models (STFMs) aim to learn generalizable representations of complex dynamical systems across space and time. However, existing approaches suffer from distributional bias in real-world pre-training data, structural bottlenecks of autoregressive or diffusion-based paradigms, and objectives that overemphasize point-wise reconstruction in noisy observation space.We propose \textbf{NeoST}, the first spatio-temporal foundation model pre-trained solely on procedurally generated synthetic systems. NeoST introduces a scalable synthetic pre-training corpus to mitigate real-world bias, a latent-space reasoning architecture that generates and iteratively refines multiple future trajectories without sequential error accumulation, and latent-space objectives that emphasize structural dynamics and enable inference-time correction under distribution shifts.Extensive experiments across diverse real-world benchmarks show that NeoST consistently outperforms existing STFMs in diverse real-world spatio-temporal systems, achieves superior long-horizon stability and inference efficiency.

URL PDF HTML 收藏
2607.17872 2026-07-21 quant-ph cs.DC cs.ET cs.LG 新提交

Entanglement geometry separates circuit cutting, classical hardness, and trainability

纠缠几何分离电路切割、经典硬度和可训练性

Maria Gragera Garces, Sabina Drăgoi, Lirandë Pira

机构 * Quantum Software Lab(量子软件实验室) University of Edinburgh, UK(爱丁堡大学) IBM Research(IBM研究院) Centre for Quantum Technologies(量子技术中心) National University of Singapore(新加坡国立大学)

AI总结 研究表明纠缠几何约束电路切割等属性,具恒定接缝键维度的MPS和TTN电路可经典模拟,构建的双块电路家族可廉价切割,MPS硬度和可训练性深度范围不兼容,用魔法作硬度资源可避免冲突,浅Clifford+\(T\)电路有相应特性。

Comments 4 pages, 2 figures

详情
AI中文摘要

电路切割有望扩展量子计算规模,但变分量子优势还需低切割开销、经典硬度和可训练性。我们表明这些属性受纠缠几何强烈约束。具有恒定接缝键维度的矩阵乘积态(MPS)和树张量网络(TTN)电路可在\(O(1/\varepsilon^2)\)采样开销下切割,但仍可高效经典模拟,排除了这些家族内的渐近量子优势。通过独立控制接缝和块内纠缠,我们构建了一个双块电路家族,它可廉价切割且需要超多项式全局MPS键维度,数值上支持到\(n = 100\)。然而,MPS硬度和可训练性需要不兼容的深度范围,分别为\(d=\omega(\log n)\)和\(d=O(\log n)\)。使用魔法而非纠缠作为硬度资源可避免此冲突:浅的Clifford+\(T\)电路可切割且可训练,同时其稳定器模拟成本随\(T\)计数呈指数增长。

英文摘要

Circuit cutting promises to scale quantum computations beyond current hardware, but variational quantum advantage also requires low cutting overhead, classical hardness, and trainability. We show that these properties are strongly constrained by entanglement geometry. Matrix product state (MPS) and tree tensor network (TTN) circuits with constant seam bond dimension can be cut with \(O(1/\varepsilon^2)\) sampling overhead, but remain efficiently classically simulable, ruling out asymptotic quantum advantage within these families. By independently controlling seam and intra-block entanglement, we construct a two-block circuit family that remains cheaply cuttable while requiring a super-polynomial global MPS bond dimension, as supported numerically up to \(n=100\). However, MPS hardness and trainability require incompatible depth regimes, \(d=ω(\log n)\) and \(d=O(\log n)\), respectively. Using magic rather than entanglement as the hardness resource avoids this conflict: shallow Clifford+\(T\) circuits remain cuttable and trainable while their stabiliser-simulation cost grows exponentially with the \(T\)-count.

URL PDF HTML 收藏
2607.17469 2026-07-21 cs.CC cond-mat.stat-mech cs.LG math.CO math.PR 新提交

The Dimension of Nonterminating Resampling Computations

非终止重采样计算的维度

Yunbei Xu

机构 * National University of Singapore(新加坡国立大学)

AI总结 研究非终止重采样计算的维度,通过主定理界定相关量,探讨源幂信息,分析不同修复规则及\(k\)-SAT 等情况,得出有限磁带源下规则的非终止维度差异、\(k\)-SAT 终止条件及公式维度界等结论。

详情
AI中文摘要

随机算法即便存在异常随机磁带使其永远运行,也可能几乎必然终止。本文研究了生存尾部、此类磁带之一的柯尔莫哥洛夫复杂度以及所有磁带的豪斯多夫维度。对于动力修复矩阵可交换的每个\(s>0\),主定理界定了在生存前缀\(w\)上的\(\sum_wP[w]^s\),在确定性非预期选择器上一致。\(s = 1\)的情况控制终止;完整族给出弱源和维度界。源幂包含普通修复核和完整停止时间定律中都没有的信息。在一个常见的有限磁带源下,四顶点路径上的两个重叠分歧修复规则对于每个选择器具有相同的普通核和相同的停止时间定律,但它们的非终止维度可以任意接近零和一。在一个共同的源幂水平上,相同的主导磁带源使一个规则永远运行,但给另一个指数停止尾部。这种分离是由产生相同状态转换的动作标签引起的,因此在幂为一时不可见。对于有界依赖\(k\)-SAT,高于迹增长阈值的条件块最小熵给出指数终止,单个无限运行的有效维度由无限次修复的子句引起的迹增长界定。树公式渐近达到最大度维度和全局源界,而团公式在所述 regime 中达到特定于图的一步阈值。一个精确的反向似然恒等式用每次运行的尾部和编码界补充了这些逐集结果。

英文摘要

A randomized algorithm may terminate almost surely even though exceptional random tapes make it run forever. This paper studies the survival tail, the Kolmogorov complexity of one such tape, and the Hausdorff dimension of all of them. For each $s>0$ at which the powered repair matrices commute, the main theorem bounds $\sum_wP[w]^s$ over surviving prefixes $w$, uniformly over deterministic nonanticipating selectors. The case $s=1$ controls termination; the full family gives weak-source and dimension bounds. The source powers contain information absent even from the ordinary repair kernel and the complete stopping-time law. Under one common finite tape source, two overlapping disagreement-repair rules on a four-vertex path have the same ordinary kernels and the same stopping-time law for every selector, yet their nontermination dimensions can be arbitrarily close to zero and one. At one common source-power level, the same dominated tape source makes one rule run forever but gives the other an exponential stopping tail. The separation is caused by action labels that produce the same state transition and are therefore invisible at power one. For bounded-dependence $k$-SAT, conditional block min-entropy above the trace-growth threshold gives exponential termination, and the effective dimension of an individual infinite run is bounded by the trace growth induced by the clauses repaired infinitely often. Tree formulas asymptotically attain the maximum-degree dimension and global source bounds, while clique formulas attain the graph-specific one-step threshold in the stated regime. An exact backward likelihood identity complements these setwise results with tail and coding bounds for each run.

URL PDF HTML 收藏
2607.04438 2026-07-21 cs.CV cs.AI cs.HC cs.MA cs.MM 版本更新

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

ResearchStudio-Reel:实现从论文到海报、视频和博客的研究最后一公里自动化

Lingao Xiao, Yalun Dai, Yangyu Huang, Qihao Zhao, Wenshan Wu, Hugo He, Ruishuo Chen, Jin Jiang, Qianli Ma, Jiahuan Zhang, Xin Zhang, Ying Xin, Yang Ou, Yan Xia, Scarlett Li, Longbo Huang, Zhipeng Zhang, Yang He, Yap Kim Hui, Yan Lu

机构 * Microsoft Research(微软研究院) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) Tsinghua University(清华大学) Peking University(北京大学) Shanghai Jiao Tong University(上海交通大学) Westlake University(西湖大学) CFAR, A*STAR(计算科学与工程研究所,新加坡科技研究局)

AI总结 研究传播自动化困难,以往方法有局限。该研究提出将最后一公里构建为技能组合,实例化ResearchStudio-Reel,包括共享提取器、可编辑生成器和交互式收敛层,能产出多种可编辑工件,效果优于现有系统。

详情
AI中文摘要

研究传播,即将论文转化为海报、演讲视频和博客文章,仍然是手动的最后一公里。以前的自动化方法孤立地处理每个工件,每个都从头重新提取论文,通常提供单向渲染,作者无法在PowerPoint或Word中重新打开,并且根据软VLM偏好分数来评估质量,而在承载部分仍为空时分数会趋于平稳。我们认为这最后一公里最好构建为技能组合:瘦代理可读契约,共享一个上游提取器,并在测量填充循环中包装确定性原语,其出口是硬通过/失败渲染门。我们将其实例化为ResearchStudio-Reel,五个Claude代码和Codex技能组织成一个共享提取器(Paper2Assets)、三个可编辑生成器(Paper2Poster、Paper2Video、Paper2Blog)和一个交互式收敛层(Paper2Reel)。Paper2Assets将每篇论文提取一次到一个共享包中,供每个下游技能重用;三个生成器生成一个可打印的海报、一个同步的演讲视频和一个双语博客,它们在事实层面上保持一致,并能通过PowerPoint或Word进行往返;Paper2Reel然后将这三个绑定到一个独立的HTML查看器中,其部分级点击会使视频、幻灯片、字幕和博客跳转到匹配的内容。在Paper2Poster基准测试中,我们的海报在美学和信息子标准方面领先于先前的自动化系统和单镜头前沿语言模型,在两名外部VLM评委的评估下,在美学方面超过了作者自己的海报,并在84%至93%的论文中总体获胜;能力审计进一步表明,通过将与叙述对齐的幻灯片亮点与由布局感知DOCX修复控制的双语博客独特配对,ResearchStudio-Reel是唯一能够提供所有三个可编辑工件的管道。项目可在此https URL上获取

英文摘要

Despite growing automation, turning a paper into a coherent poster, talk video, and blog piece often remains a labor-intensive last mile. Recent systems increasingly generate multiple dissemination formats, but a practical workflow must also keep the outputs editable in native tools and bound into one navigable deliverable for revision and reuse. We present ResearchStudio-Reel, a native-editable dissemination workspace that binds its three artifacts into one interactive deliverable at the experience level, implemented as five skills executable in Claude Code and Codex: one shared extractor, three editable artifact generators, and one interactive convergence layer. A shared asset bundle feeds a PowerPoint poster and video deck, plus a bilingual Word blog; rather than re-rendering the paper into a fourth format, Paper2Reel converges these already-produced artifacts at the experience level, binding poster regions, video segments, and blog passages into one interactive viewer. Artifact-specific release checks make this delivery contract testable, and Paper2Poster additionally uses a measured-fill loop. On the Paper2Poster benchmark, our Claude Code configuration achieves the best scores among automated systems on all three aesthetic sub-criteria and the best or tied-best scores on two of three information sub-criteria. Under two VLMjudges, it exceeds the authors' posters in average aesthetics (3.56 vs. 3.03) and wins on overall quality on 74 and 95 of the 100 papers under the two judges. The full pipeline additionally packages the native-editable source artifacts and their aligned viewer. Project is available at https://aka.ms/ResearchStudio

URL PDF HTML 收藏
2607.00573 2026-07-21 cs.CV 版本更新

BrainFIBRE: A Foundation Model via Information Decomposition for Brain Microstructure

BrainFIBRE:基于信息分解的脑微结构基础模型

Zijian Dong, Yi Lin, Fang Ji, Jianxiong Zhou, Kwun Kei Ng, Juan Helen Zhou

机构 * National University of Singapore(新加坡国立大学)

AI总结 提出BrainFIBRE,首个脑微结构基础模型,通过自监督部分信息分解(SPID)和反事实候选构建(CCC)从NODDI图谱中解缠独特、协同和冗余信息,在多种预测任务上达到最优性能。

Comments ECCV 2026. Author name corrected in V2

详情
AI中文摘要

弥散MRI探测脑微结构,对早期脑血管和神经退行性变化特别敏感。神经突方向分散度和密度成像(NODDI)将弥散信号分解为三个生物物理解释图:神经突密度指数(NDI)、方向分散度指数(ODI)和自由水分数(FWF),分别捕捉神经突堆积、纤维相干性和细胞外液。这些3D图为可迁移的微结构表示提供了丰富基底,但整合它们具有挑战性:标准表示学习难以从共享和协同交互中解缠每个图的独特信息。我们提出BrainFIBRE,首个脑微结构基础模型,在来自55,592名UK Biobank参与者的NODDI衍生图上预训练。我们提出自监督部分信息分解(SPID),首次将PID引导的多模态学习扩展到自监督范式。一种新颖的反事实候选构建(CCC)范式通过模态丢弃和交换扰动模态间对齐,为混合专家架构提供对比信号,以解缠独特、协同和冗余信息,无需任何下游标签。在白种人和亚洲人群队列中,BrainFIBRE在预测年龄、性别、脑血管和神经退行性标志物以及认知的多种任务上达到最先进性能,同时产生神经生物学可解释的表示,揭示任务和队列特定的交互模式。BrainFIBRE为微结构水平的神经影像分析建立了多功能基础。

英文摘要

Diffusion MRI probes brain microstructure with particular sensitivity to early cerebrovascular and neurodegenerative changes. Neurite Orientation Dispersion and Density Imaging (NODDI) decomposes the diffusion signal into three biophysically interpretable maps: neurite density index (NDI), orientation dispersion index (ODI), and free water fraction (FWF), capturing neurite packing, fiber coherence, and extracellular fluid. These 3D maps offer a rich substrate for transferable microstructural representations, yet integrating them is challenging: standard representation learning struggles to disentangle the unique information in each map from their shared and synergistic interactions. We present BrainFIBRE, the first foundation model for brain microstructure, pretrained on NODDI-derived maps from 55,592 UK Biobank participants. We propose Self-supervised Partial Information Decomposition (SPID), which extends PID-guided multimodal learning to the self-supervised regime for the first time. A novel Counterfactual Candidate Construction (CCC) paradigm perturbs inter-modality alignment through modality dropping and swapping, providing the contrastive signal for a Mixture-of-Experts architecture to disentangle unique, synergistic, and redundant information without any downstream label. On both Caucasian and Asian cohorts, BrainFIBRE achieves state-of-the-art performance across diverse tasks predicting age, sex, cerebrovascular and neurodegenerative markers, and cognition, while yielding neurobiologically interpretable representations that reveal task- and cohort-specific interaction patterns. BrainFIBRE establishes a versatile foundation for neuroimaging analysis at the microstructural level.

URL PDF HTML 收藏
2606.25325 2026-07-21 cs.AI 版本更新

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning

全模态感知策略优化用于多模态情感推理

Zhiyuan Han, Beier Zhu, Wenwen Tong, Pengyang Shao, Peipei Song, Xinyi Wang, Jiangnan Chen, Lewei Lu, Xun Yang

机构 * University of Science and Technology of China(中国科学技术大学) SenseTime Research(感时间研) National University of Singapore(新加坡国立大学) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)

AI总结 提出OPPO强化学习框架,通过全模态感知奖励和损失优化多模态感知,提升情感推理中模态利用率和忠实度,在多个基准上达到最优。

Comments Accepted at ICML 2026. Project page: https://zjycutieee.github.io/OPPO-page/ Code: https://github.com/ZhiyuanHan-Aaron/OPPO

详情
AI中文摘要

我们发现当前面向情感的全模态多模态大语言模型仍然缺乏可靠的全模态感知:它们(i)在推理轨迹中未充分利用多模态线索,(ii)表现出不忠实的行为,经常从其他模态幻觉出特定模态的陈述。基于这些见解,我们提出了OPPO(全模态感知策略优化),一个明确优化多模态感知的强化学习框架。首先,全模态感知奖励将真实推理分解为细粒度的视觉、声学和情感线索,并奖励语义上恢复这些线索的轨迹。其次,全模态感知损失比较策略在完整输入和单模态掩码输入下的表现,仅对模态特定证据标记施加KL惩罚以抑制跨模态幻觉。我们进一步引入了MEP-Bench,一个诊断基准,用于量化利用率和忠实度。实验表明,OPPO在MER-UniBench和MME-Emotion上达到了最先进的性能,同时在MEP-Bench上大幅提高了利用率和忠实度得分,凸显了充分且忠实全模态感知对多模态情感推理的重要性。

英文摘要

We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning trajectories and (ii) exhibit unfaithful behavior, often hallucinating modality-specific statements from other modalities. Building on these insights, we propose OPPO (Omni-Perception Policy Optimization), a reinforcement learning framework that explicitly optimizes multimodal perception. First, an Omni-Perception Reward decomposes ground-truth reasoning into fine-grained visual, acoustic, and emotion cues and rewards trajectories that semantically recover these cues. Second, an Omni-Perception Loss compares the policy under full and unimodally masked inputs, applying a KL penalty only to modality-specific evidence tokens to suppress cross-modal hallucination. We further introduce MEP-Bench, a diagnostic benchmark that quantifies utilization and faithfulness. Experiments show that OPPO achieves state-of-the-art performance on MER-UniBench and MME-Emotion, while substantially improving utilization and faithfulness scores on MEP-Bench, highlighting the importance of sufficient and faithful omni perception for multimodal emotion reasoning.

URL PDF HTML 收藏
2606.12995 2026-07-21 cs.RO 版本更新

GenHOI: Contact-Aware Humanoid-Object Interaction by Imitating Generated Videos without Task-Specific Training

GenHOI: 通过模仿生成视频实现接触感知的人形机器人-物体交互,无需任务特定训练

Zhihai Bi, Qiang Zhang, Guoyang Zhao, Jiahang Cao, Xueyin Luo, Yushan Zhang, Jinglan Xu, Ruoyu Geng, Yulin Li, Andrew F. Luo, Jun Ma

机构 * The University of Tokyo(东京大学) National University of Singapore(新加坡国立大学) University of California, Los Angeles(加州大学洛杉矶分校) Tsinghua University(清华大学)

AI总结 提出GenHOI框架,通过模仿单个生成视频实现人形机器人零样本执行多种物体交互任务,无需任务特定训练或物理演示数据,利用接触事件和手-物接触区域编码为几何约束优化轨迹。

详情
AI中文摘要

人形机器人-物体交互(HOI)是人形机器人的基本能力,但由于动态平衡与与多样物体稳定交互之间的紧密耦合,它仍然具有挑战性。现有方法通常需要耗时的任务特定策略训练或依赖于刚性轨迹回放,这限制了它们适应新颖交互场景的能力。在这项工作中,我们提出了\textit{GenHOI},一个简单而有效的框架,通过直接模仿单个生成视频,使人类形机器人能够以零样本方式执行多样化的物体交互任务,无需任务特定训练或物理演示数据。GenHOI首先在仿真中重建机器人-物体场景并渲染第一帧图像,该图像与语言命令一起条件化任务导向交互视频的合成。然后分析生成的视频以识别交互相关的接触事件并估计手-物体接触区域,这些被编码为以物体为中心的几何约束,将视觉交互线索转化为物理基础的优化先验。在这些先验的指导下,从视频中恢复的参考运动被细化和平滑,以解决2D视频生成中固有的尺度模糊性,同时将单个参考轨迹适应于未见过的机器人-物体相对姿态。优化后的轨迹最终由闭环跟踪控制器执行。我们在包括箱子抓取、非对称双臂椅子搬运、从下方抬桌子和圆柱物体包裹在内的多样化物体交互任务中,通过大量仿真和真实世界实验验证了所提出的框架。

英文摘要

Humanoid-Object Interaction (HOI) is a fundamental capability for humanoid robots, yet it remains challenging due to the tight coupling between dynamic balance and stable interaction with diverse objects. Existing methods often require time-consuming task-specific policy training or rely on rigid trajectory replay, which limits their ability to accommodate novel interaction scenarios. In this work, we present \textit{GenHOI}, a simple yet effective framework that enables humanoid robots to perform diverse object-interaction tasks in a zero-shot manner by directly imitating a single generated video, without task-specific training or physical demonstration data. GenHOI first reconstructs the robot-object scene in simulation and renders a first-frame image, which, together with the language command, conditions the synthesis of a task-oriented interaction video. The generated video is then analyzed to identify interaction-relevant contact events and estimate hand-object contact regions, which are encoded as object-centric geometric constraints that convert visual interaction cues into physically grounded optimization priors. Guided by these priors, the reference motion recovered from the video is refined and smoothed to resolve the scale ambiguity inherent in 2D video generation, while adapting a single reference trajectory to unseen robot-object relative poses. The optimized trajectory is finally executed by a closed-loop tracking controller. We validate the proposed framework in extensive simulation and real-world experiments across diverse object-interaction tasks, including box grasping, asymmetric bimanual chair carrying, table lifting from below, and cylindrical-object enveloping.

URL PDF HTML 收藏
2602.01167 2026-07-21 cs.AI

Do All Individual Layers Help? An Empirical Study of Task-Interfering Layers in Vision-Language Models

所有个体层都有帮助吗?视觉-语言模型中任务干扰层的实证研究

Zhiming Liu, Yujie Wei, Lei Feng, Xiu Su, Xiaobo Xia, Weili Guan, Zeke Xie, Shuo Yang

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Harbin Institute of Technology(哈尔滨工业大学) Southeast University(东南大学) Central South University(中南大学) National University of Singapore(新加坡国立大学) The Hong Kong University of Science and Technology, Guangzhou(香港科学与技术大学(广州))

AI总结 研究通过层干预发现部分层阻碍下游任务,提出任务自适应层剔除方法提升性能,揭示预训练VLM的意外模块化特性。

Journal ref Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 2026, pp. 9597-9607

详情
AI中文摘要

当前VLM在多种多模态任务中表现出色,但默认启用所有层可能阻碍任务表现。通过干预单层参数发现,某些层反而抑制任务性能。系统研究各层对不同任务的影响,提出任务-层交互向量量化方法,并引入无需训练的测试时适应方法TaLo,动态剔除最干扰的层,提升模型在多个任务和数据集上的性能,包括提升Qwen-VL在ScienceQA地图任务上的准确率。

英文摘要

Current VLMs have demonstrated capabilities across a wide range of multimodal tasks. Typically, in a pretrained VLM, all layers are engaged by default to make predictions on downstream tasks. We find that intervening on a single layer, such as by zeroing its parameters, can improve the performance on certain tasks, indicating that some layers hinder rather than help downstream tasks. We systematically investigate how individual layers influence different tasks via layer intervention. Specifically, we measure the change in performance relative to the base model after intervening on each layer and observe improvements when bypassing specific layers. This improvement can be generalizable across models and datasets, indicating the presence of Task-Interfering Layers that harm downstream tasks' performance. We introduce Task-Layer Interaction Vector, which quantifies the effect of intervening on each layer of a VLM given a task. These task-interfering layers exhibit task-specific sensitivity patterns: tasks requiring similar capabilities show consistent response trends under layer interventions, as evidenced by the high similarity in their task-layer interaction vectors. Inspired by these findings, we propose TaLo (Task-Adaptive Layer Knockout), a training-free, test-time adaptation method that dynamically identifies and bypasses the most interfering layer for a given task. Without parameter updates, TaLo improves performance across various models and datasets, including boosting Qwen-VL's accuracy on the Maps task in ScienceQA by up to 16.6%. Our work reveals an unexpected form of modularity in pretrained VLMs and provides a plug-and-play, training-free mechanism to unlock hidden capabilities at inference time. The source code will be publicly available.

URL PDF HTML 收藏
2504.17356 2026-07-21 cs.AI cs.LG 版本更新

Comprehend, Divide, and Conquer: Feature Subspace Exploration via Multi-Agent Hierarchical Reinforcement Learning

理解、划分与征服:通过多智能体分层强化学习进行特征子空间探索

Weiliang Zhang, Xiaohan Huang, Yi Du, Ziyue Qiao, Qingqing Long, Zhen Meng, Yuanchun Zhou, Meng Xiao

机构 * Computer Network Information Center, Chinese Academy of Sciences(中国科学院计算机网络信息中心) University of Chinese Academy of Sciences(中国科学院大学) Great Bay University(Great Bay大学) Duke-NUS Medical School, National University of Singapore(新加坡国立大学杜克-奈素医学院)

AI总结 研究针对特征选择问题,提出HRLFS方法,先利用基于大语言模型的混合状态提取器捕捉特征特性并聚类,构建分层智能体,通过多智能体分层强化学习进行特征子空间探索,提升了下游机器学习性能并加速运行

Comments 25 pages, keywords: Automated Feature Engineering, Tabular Dataset, Multi-Agent Reinforcement Learning, Feature Selection, Accepted by ACM Transactions on Knowledge Discovery from Data

详情
AI中文摘要

特征选择旨在预处理目标数据集,找到最优且最精简的特征子集以增强下游机器学习任务。基于强化学习的子空间探索策略提供了新视角,但当前强化学习方法在处理复杂数据集时面临挑战。本文提出HRLFS方法,先用基于大语言模型的混合状态提取器捕捉特征特性,进行特征聚类,构建分层智能体。实验证明该方法有效、可扩展且稳健,能提升下游机器学习性能并加速运行。

英文摘要

Feature selection aims to preprocess the target dataset, find an optimal and most streamlined feature subset, and enhance the downstream machine learning task. Among filter, wrapper, and embedded-based approaches, the reinforcement learning (RL)-based subspace exploration strategy provides a novel objective optimization-directed perspective and promising performance. Nevertheless, even with improved performance, current reinforcement learning approaches face challenges similar to conventional methods when dealing with complex datasets. These challenges stem from the inefficient paradigm of using one agent per feature and the inherent complexities present in the datasets. This observation motivates us to investigate and address the above issue and propose a novel approach, namely HRLFS. Our methodology initially employs a Large Language Model (LLM)-based hybrid state extractor to capture each feature's mathematical and semantic characteristics. Based on this information, features are clustered, facilitating the construction of hierarchical agents for each cluster and sub-cluster. Extensive experiments demonstrate the efficiency, scalability, and robustness of our approach. Compared to contemporary or the one-feature-one-agent RL-based approaches, HRLFS improves the downstream ML performance with iterative feature subspace exploration while accelerating total run time by reducing the number of agents involved.

URL PDF HTML 收藏
2506.07691 2026-07-21 cs.CL cs.LG 版本更新

Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models

打破模块限制:保留数据连续性以训练用于指令模型的卓越稀疏自编码器

Jiaming Li, Haoran Ye, Yukun Chen, Xinyue Li, Lei Zhang, Hamid Alinejad-Rokny, Jimmy Chih-Hsien Peng, Min Yang

机构 * Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, China(中国科学院深圳先进技术研究所) University of Chinese Academy of Sciences, China(中国科学院大学) National University of Singapore, Singapore(新加坡国立大学) The University of New South Wales, Australia(新南威尔士大学) Shenzhen University of Advanced Technology, China(深圳先进技术大学)

AI总结 研究针对现有SAEs训练方法在指令模型中因注意力泄漏引入梯度噪声问题,提出FAST顺序训练范式,通过与数据分布和激活模式对齐,提升重建保真度和特征可解释性,实验显示该方法效果显著优于基线。

详情
AI中文摘要

稀疏自编码器(SAEs)是机制可解释性的基石。现有训练方法继承自语言模型预训练的块训练范式,因无关上下文的注意力泄漏在指令模型中引入破坏性梯度噪声。通过GSNR分析,从理论上描述了该问题,并提出了微调对齐顺序训练(FAST),一种专为指令模型设计的顺序训练范式。FAST使SAE训练与指令模型的数据分布和激活模式对齐,显著提高了重建保真度和特征可解释性。实验结果表明,FAST实现了更高的GSNR,与基线的5.1985相比,对数缩放的均方误差显著降低至0.6468,Delta Loss接近零(-0.51%至0.37%)。此外,在Llama-3.2-3B-it上,FAST产生21.1%的高质量特征,大幅优于基线方法的7.0%和10.2%。还发现通过SAEs干预特殊令牌激活可提高生成质量,揭示了细粒度控制的新机会。代码可在指定网址开源获取。

英文摘要

Sparse Autoencoders (SAEs) are a cornerstone of mechanistic interpretability. Existing training methods inherit the Block Training paradigm from LLM pre-training, which introduces destructive gradient noise in instruct models due to attention leakage from unrelated contexts. Using GSNR analysis, we theoretically characterize this issue and propose Finetuning-aligned Sequential Training (FAST), a sequential training paradigm specifically designed for instruct models. FAST aligns SAE training with the data distribution and activation patterns of instruct models, substantially improving both reconstruction fidelity and feature interpretability. Experimental results show that FAST achieves higher GSNR, a significantly lower log-scaled MSE of 0.6468 compared to the baseline's 5.1985, and a near-zero Delta Loss (-0.51\% to 0.37\%). Moreover, on Llama-3.2-3B-it, FAST produces 21.1\% high-quality features, substantially outperforming baseline methods that achieve 7.0\% and 10.2\%. We further find that intervening on special token activations through SAEs can improve generation quality, revealing new opportunities for fine-grained control. Our codes are available as open source at https://github.com/Geaming2002/FAST.

URL PDF HTML 收藏
2607.15808 2026-07-20 cs.CV 新提交

Examining the Associations between Visual and Non-Visual Elements and Cyclists' Route Choices for Various Trip Purposes

研究视觉与非视觉元素与不同出行目的下骑行者路线选择之间的关联

Heyang Hua, Koichi Ito, Filip Biljecki

机构 * Department of Architecture, National University of Singapore(新加坡国立大学建筑系) Department of Real Estate, National University of Singapore(新加坡国立大学房地产系)

AI总结 研究不同出行目的下骑行者路线选择,通过数据驱动案例研究,分析视觉与非视觉因素对其影响,揭示影响积极骑行的时空特征,为街道网络规划和基础设施发展提供参考。

详情
AI中文摘要

了解骑行者对建成环境特征的偏好对于促进可持续城市交通和积极出行很重要。尽管此前有关于骑行者路线选择的研究,但视觉和非视觉因素对不同出行目的选择的影响仍不明确。本文通过加拿大蒙特利尔的数据驱动案例研究填补这一空白。非视觉因素包括社会经济因素和二维环境,视觉因素涉及骑行时的视觉感知并通过街景图像计算得出。研究分两部分,一部分分析时空信息探索骑行起点和终点间的非视觉因素,另一部分研究最短路径与实际路径间这些因素分布的差异。结果揭示了影响积极骑行选择的时空特征,如绿化增加和机动车化程度降低。这些见解可为街道网络规划和基础设施发展提供参考以改善积极交通的使用。

英文摘要

Understanding cyclist preferences for the characteristics of the built environment is important in promoting sustainable urban transportation and active mobility. Despite previous studies on cyclists' route choices, the influence of visual and non-visual factors on these choices for different trip purposes remains unclear; thus, this paper fills this gap through a data-driven case study in Montreal, Canada. Non-visual factors include socioeconomic factors and two-dimensional environments, while visual factors involve visual perception during cycling and are computed using street view images. The study consists of two parts: one part analyzes spatiotemporal information to explore the non-visual factors between the start and end points of cycling trips, and the other part investigates the discrepancies in distributions of these factors between the shortest path and the actual one. The findings reveal the spatiotemporal characteristics that influence active riding choices, such as increased greenery and lower levels of motorization. These insights can inform the planning of street networks and the development of infrastructure to improve the use of active transportation.

URL PDF HTML 收藏
2607.15661 2026-07-20 cs.CV 新提交

Model Merging for Medical LVLMs: A Benchmark and a Winner-Take-All Approach

用于医学视觉语言模型的模型合并:一个基准和一种赢家通吃的方法

Lichao Mou, Shilan Zhang, Chunlei Li, Bingcong Yan, Jingliang Hu, Yilei Shi, Shengwu Xiong, Xiao Xiang Zhu, Lei Li, Yaxiong Chen

机构 * Wuhan University of Technology(武汉理工大学) Technical University of Munich(慕尼黑工业大学) National University of Singapore(新加坡国立大学)

AI总结 研究医学LVLMs的模型合并,提出涵盖多种成像模态和临床任务类型的MergeMedBench基准,评估现有合并方法,提出赢家通吃方法,该方法简单且无超参数,优于现有方法,为LoRA合并提供新视角和实用基线。

Comments Project Page: https://github.com/MedAI-T/MergeMedBench

详情
AI中文摘要

大型视觉语言模型(LVLMs)可通过低秩适应(LoRA)等参数高效微调方法应用于专业医学成像任务,产生了针对特定成像模态和临床场景的专家模型生态系统。然而,在实践中部署多个专家LVLMs会带来大量计算和操作开销。模型合并通过将多个专家模型整合为一个无需重新训练的单一模型提供了一个有前景的解决方案,但在医学领域仍未得到充分探索。在这项工作中,我们首次对医学LVLMs的模型合并进行了系统研究。我们引入了MergeMedBench,这是一个涵盖八种成像模态和多种临床任务类型的综合基准,包括基于两种主流架构构建的16个LoRA微调模型。我们对现有合并方法进行了广泛评估,并进一步提出了赢家通吃方法,这是一种简单且无超参数的方法,只保留专家模型中最具主导性的参数。通过保留控制模型行为的关键参数并丢弃较弱的参数,我们的方法避免了基于平均或对齐策略中固有的信息稀释。尽管简单,赢家通吃方法始终优于现有方法,为LoRA合并提供了新视角,并为未来研究提供了强大的实用基线。

英文摘要

Large vision-language models (LVLMs) can be adapted to specialized medical imaging tasks via parameter-efficient fine-tuning approaches such as low-rank adaptation (LoRA), leading to a growing ecosystem of expert models tailored to specific imaging modalities and clinical scenarios. However, deploying multiple expert LVLMs in practice incurs substantial computational and operational overhead. Model merging provides a promising solution by consolidating multiple experts into a single model without retraining, yet it remains largely unexplored in the medical domain. In this work, we present the first systematic study of model merging for medical LVLMs. We introduce MergeMedBench, a comprehensive benchmark spanning eight imaging modalities and diverse clinical task types, comprising 16 LoRA fine-tuned models built upon two mainstream architectures. We conduct an extensive evaluation of existing merging methods and further propose winner-take-all, a simple and hyperparameter-free approach that retains only the most dominant parameters across expert models. By preserving the critical parameters that govern model behavior and discarding weaker ones, our method avoids the information dilution inherent in averaging- or alignment-based strategies. Despite its simplicity, winner-take-all consistently outperforms existing approaches, offering both a new perspective on LoRA merging and a strong practical baseline for future research.

URL PDF HTML 收藏
2601.13020 2026-07-20 cs.LG cs.AI 版本更新

PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual Learning

PASs-MoE:通过路径激活子空间减轻路由器与专家之间的错位协同漂移以进行持续学习

Zhiyan Hou, Haiyun Guo, Haokai Ma, Yandu Sun, Yonghui Yang, Jinqiao Wang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) National University of Singapore(新加坡国立大学) Southeast University, Nanjing, China(南京东南大学) Wuhan AI Research, Wuhan, China(武汉人工智能研究院)

AI总结 研究持续指令调整中多模态大语言模型的问题,提出基于路径激活子空间的固定容量PASs - MoE - LoRA方法,含PAS引导的重新加权和PAS感知的秩稳定,实验表明该方法在准确性和抗遗忘性上优于基线和变体且不增参数。

Comments Published in the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), Volume 1: Long Papers. 14 pages. Code is available at https://github.com/yueluoshuangtian/PASs-MoE

Journal ref Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 31959--31972, San Diego, California, United States, July 2026. Association for Computational Linguistics

详情
AI中文摘要

持续指令调整(CIT)要求多模态大语言模型(MLLMs)适应一系列任务且不遗忘先前能力。常见策略是通过将输入路由到不同的LoRA专家来隔离更新。然而,现有的基于LoRA的专家混合(MoE)方法通常以不加区分的方式联合更新路由器和专家,导致路由器偏好与专家适应路径协同漂移,偏离早期输入 - 专家专业化,即错位协同漂移,这模糊了专家职责并加剧遗忘。为解决此问题,我们引入路径激活子空间(PASs),它反映输入在每个专家中激活的低秩路径方向,为路由和保存提供能力对齐的坐标系。基于PASs,我们提出基于固定容量PASs的MoE - LoRA方法,包括PAS引导的重新加权和PAS感知的秩稳定。在CIT基准测试中的实验表明,我们的方法在准确性和抗遗忘性方面均优于一系列传统持续学习基线和MoE - LoRA变体,且不增加模型参数。代码可公开获取。

英文摘要

Continual instruction tuning (CIT) requires multimodal large language models (MLLMs) to adapt to a stream of tasks without forgetting prior capabilities. A common strategy is to isolate updates by routing inputs to different LoRA experts. However, existing LoRA-based Mixture-of-Experts (MoE) methods often jointly update the router and experts in an indiscriminate way, causing the router's preferences to co-drift with experts' adaptation pathways and gradually deviate from early-stage input--expert specialization. We term this as Misaligned Co-drift, which blurs expert responsibilities and exacerbates forgetting. To address this, we introduce the pathway activation subspace (PASs), a LoRA-induced subspace that reflects which low-rank pathway directions an input activates in each expert, providing a capability-aligned coordinate system for routing and preservation. Based on PASs, we propose a fixed-capacity PASs-based MoE--LoRA method with two components: PAS-guided Reweighting, which calibrates routing using each expert's pathway activation signals, and PAS-aware Rank Stabilization, which selectively stabilizes rank directions important to previous tasks. Experiments on a CIT benchmark show that our approach consistently outperforms a range of conventional continual learning baselines and MoE--LoRA variants in both accuracy and resistance to forgetting, without increasing model parameters. Our code is publicly available at https://github.com/yueluoshuangtian/PASs-MoE.

URL PDF HTML 收藏
2607.15268 2026-07-17 cs.CV 新提交

Motion-Conditioned Multi-View Fusion for Myocardial Infarction Localization from Echocardiography

用于超声心动图心肌梗死定位的运动条件多视图融合

Guang Yang, Wentian Xu, Siyu Wang, Betty Raman, Lei Li, Vicente Grau

机构 * University of Oxford(牛津大学) National University of Singapore(新加坡国立大学)

AI总结 针对超声心动图心肌梗死定位问题,提出MCF-Net框架,融合心肌运动线索与基础模型表示,通过极稀疏监督建模心脏运动,利用运动衍生软掩码提供先验,跨视图整合运动和视觉优化预测,在节段级定位上性能优于现有方法。

详情
AI中文摘要

心肌梗死是全球主要死因。超声心动图是评估心肌梗死的常用方法,局部室壁运动异常是关键指标。以往基于学习的心肌运动分析方法存在局限性。基础模型改进了基于视觉的超声心动图分析,但多数方法基于单视图,在视图依赖模糊性下,尤其是心尖视图,节段级定位不可靠。为此提出MCF-Net,一种运动引导的多视图融合框架,融合心肌运动线索与基础模型表示来定位梗死。使用EchoPrime提取视觉特征,通过极稀疏监督建模心脏运动,运动衍生的段感知软掩码提供空间先验,运动条件融合机制跨视图整合运动和视觉来优化预测。在节段级心肌梗死定位上,MCF-Net取得了72.4%的F1值和84.9%的准确率,优于现有方法。

英文摘要

Myocardial infarction (MI) remains a leading cause of mortality worldwide. Echocardiography (Echo) is a widely available modality for MI assessment, where regional wall motion abnormality is a key indicator. Prior learning based methods for myocardial motion analysis often use handcrafted descriptors or densely supervised estimation, but the need for extensive annotation limits applicability. Foundation models have recently improved vision-based Echo analysis; however, most methods operate on single views and segment-level localization remains unreliable under view-dependent ambiguity, especially in apical views. To address this, we propose MCF-Net, a novel motion-guided multi-view fusion framework that fuses myocardial motion cues with foundation model representations to localize infarction. Visual features are extracted using EchoPrime, a pretrained Echo foundation model shared across dual views. Cardiac motion is modeled with extremely sparse supervision: a single annotated template frame is transferred across videos to initialize point tracking, avoiding dense labels. Motion-derived segment-aware soft masks provide coarse spatial priors that selectively enhance features for challenging myocardial segments. A motion-conditioned fusion mechanism then integrates motion and vision across views, refining predictions without overriding strong appearance cues. On segment-level MI localization, MCF-Net achieves 72.4\% F1 and 84.9\% accuracy, outperforming state-of-the-art motion-only, vision-only, and fusion baselines.

URL PDF HTML 收藏
2607.15255 2026-07-17 cs.CV 新提交

HoloGeo: Mitigating Landmark Bias in Geo-localization via Evidence-Driven Reasoning

HoloGeo:通过证据驱动推理减轻地理定位中的地标偏差

Pengcheng Zhou, Xuanyu Liu, Yanchen Yin, Bobo Li, Shengqiong Wu, Mong-Li Lee, Wynne Hsu

机构 * National University of Singapore(新加坡国立大学) Shandong University of Science and Technology(山东科技大学) University of Oxford(牛津大学)

AI总结 研究针对视觉语言模型地理定位易受地标偏差影响的问题,提出证据驱动推理框架HoloGeo,借助高质量数据集,通过多维度奖励实现平衡关注与联合推理,经实验验证其在多个数据集上能有效减轻地标偏差,提升地理定位性能。

详情
AI中文摘要

视觉语言模型(VLMs)的进展改善了图像地理定位,但现有模型仍易受地标偏差影响,导致忽视地理线索或形成虚假关联,造成定位不准确。为此设计了偏差强度(BI)和偏差危害(BH)两个量化指标,建立了LandmarkBias - 3K基准。提出证据驱动推理框架HoloGeo,借助高质量BF - 30k数据集,通过纳入多维度奖励鼓励对多样视觉线索的平衡关注,实现证据驱动联合推理。实验表明HoloGeo在多个数据集上表现出色,验证了其对稳健地理空间推理的有效性。

英文摘要

Recent advances in Vision-Language Models (VLMs) have significantly improved image geo-localization, yet existing models remain susceptible to landmark bias, causing them to overlook geographical cues or form spurious correlations, ultimately resulting in inaccurate localization. To systematically investigate this issue, we first design two quantitative metrics, Bias Intensity (BI) and Bias Harmfulness (BH), to characterize the impact of landmarks exerted on model reasoning, and establish a comprehensive benchmark, LandmarkBias-3K. To mitigate landmark bias, we further propose an evidence-driven reasoning framework, HoloGeo, to improve the reliability of geo-localization. HoloGeo is supported by a high-quality dataset, BF-30k, annotated with structured multi-evidence bias-free reasoning chains. By incorporating multi-dimensional rewards, HoloGeo explicitly encourages balanced attention over diverse visual cues and achieves evidence-driven joint reasoning. Extensive experiments demonstrate that HoloGeo not only maintains excellent performance on IM2GPS3K and YFCC4k but also significantly outperforms existing open-source VLMs on LandmarkBias-3K, validating its effectiveness for robust geospatial reasoning.

URL PDF HTML 收藏
2607.15207 2026-07-17 cs.LG cs.RO 新提交

BadWAM: When World-Action Models Dream Right but Act Wrong

BadWAM:当世界-动作模型想得对但做得错时

Qi Li, Xingyi Yang, Xinchao Wang

机构 * National University of Singapore(新加坡国立大学) The Hong Kong Polytechnic University(香港理工大学)

AI总结 研究针对世界-动作模型(WAMs)提出BadWAM框架,用于建模和评估世界-动作漂移攻击,包括仅动作攻击和保持想象攻击,通过不同标准刻画攻击面,评估结果显示能大幅降低任务成功率,揭示WAM漏洞。

详情
AI中文摘要

世界-动作模型(WAMs)正成为具身控制的一个有前景的基础:它们学习将动作生成与未来世界预测相结合的表示,而非仅预测动作。这种耦合常被视为鲁棒性、可解释性和安全性的来源。本文表明该假设是脆弱的。我们引入BadWAM,一个用于建模和评估世界-动作漂移攻击的统一框架,这是一类新的针对WAM的对抗攻击,利用小的视觉扰动打破WAM想象与执行之间的对齐。BadWAM根据攻击强度和隐蔽性这两个自然标准来刻画这种攻击面。当对手优先考虑破坏时,BadWAM实例化仅动作对抗攻击,直接驱使模型走向导致任务失败的动作。当对手还优先考虑隐蔽性时,BadWAM实例化保持想象的对抗攻击,试图在使模型预测的未来接近其纯净想象的同时引发有害的动作转变。我们在不同的WAM变体上评估BadWAM。结果表明,我们的攻击在闭环执行下大幅降低了任务成功率。例如,我们的仅动作攻击将模型性能从96.5%的成功率降至43.1%。我们的保持想象攻击结果进一步揭示了WAM特有的漏洞:适度的未来保持正则化可以保持强大的攻击性能,同时减少未来想象漂移。

英文摘要

World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction. This coupling is often viewed as a source of robustness, interpretability, and safety, as a robot's action can in principle be checked against its imagined future. In this paper, we show that this assumption is fragile. We introduce BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes. BadWAM characterizes this attack surface along two natural criteria: attack strength and stealthiness. When the adversary prioritizes disruption, BadWAM instantiates an action-only adversarial attack, which directly drives the model toward task-failing actions. When the adversary additionally prioritizes stealth, BadWAM instantiates an imagination-preserving adversarial attack, which seeks to induce harmful action shifts while keeping the model's predicted future close to its clean imagination. Together, these two attacks capture a spectrum of WAM-specific failures: from overt action hijacking to stealthier cases where the model appears to imagine a plausible future but executes a desynchronized action. We evaluate BadWAM across different variants of WAMs. Results show that our attacks substantially reduce task success rates under closed-loop execution. For example, our action-only attack reduces the model performance from 96.5% to 43.1% success. The results of our imagination-preserving attack further exposes a WAM-specific vulnerability: moderate future-preserving regularization can maintain strong attack performance while reducing future imagination drift.

URL PDF HTML 收藏
2607.15092 2026-07-17 cs.CL 新提交

Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence

试验中的评分标准:通过合成成对证据从单个查询中演化评分标准

Haocheng Yang, Licheng Pan, Xiaoxi Li, Zhichao Chen, Zhiheng Zhang, Yuan Lu, Haoxuan Li, Hao Wang

机构 * School of Computing, National University of Singapore(新加坡国立大学计算学院) Xiaohongshu Inc.(小红书公司) School of Cyber Science and Technology, Zhejiang University(浙江大学网络空间安全学院) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) School of Statistics and Data Science, Shanghai University of Finance and Economics(上海财经大学统计与数据科学学院) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院)

AI总结 研究针对构建可靠特定查询评分标准难的问题,提出“试验中的评分标准”框架,仅从查询演化评分标准集,靠合成响应对获取监督并验证,实验证明该框架有效,平均准确率最佳且在多数评估集领先。

详情
AI中文摘要

评分标准为训练和评估大语言模型提供结构化、细粒度的信号。然而,可靠的特定查询评分标准很难构建。现有方法通常从人工编写的评分标准、偏好数据或采样响应中获取监督。直接的查询到评分标准生成避免了这些资源,但没有明确检查合理的评分标准是否有用。这样的评分标准可能无法区分答案质量、奖励可选风格或惩罚有效的替代策略。我们引入了“试验中的评分标准”,这是一个仅基于查询的框架,它从空集演化出一组评分标准,无需外部注释或模型训练。它仅从合成的评分标准条件响应对中获取监督,并在添加每个提议的评分标准之前进行验证,筛选出无区分性、过于特定和仅风格的候选评分标准。在五个偏好基准套件上的实验证明了“试验中的评分标准”的有效性,它实现了最佳平均准确率,并在七个评估集中的六个上领先。

英文摘要

Rubrics provide structured, fine-grained signals for training and evaluating large language models (LLMs). Yet reliable query-specific rubrics are difficult to construct. Existing approaches often derive supervision from human-written rubrics, preference data, or sampled responses. Direct query-to-rubric generation avoids these resources, but provides no explicit check that a plausible rubric is useful. Such a rubric may fail to distinguish answer quality, reward an optional style, or penalize a valid alternative strategy. We introduce Rubrics on Trial, a query-only framework that evolves a rubric set from an empty set without external annotations or model training. It derives supervision solely from synthetic rubric-conditioned response pairs and validates each proposed rubric before adding it, screening out non-discriminative, over-specific, and style-only candidate rubrics. Experiments across five preference benchmark suites demonstrate the effectiveness of Rubrics on Trial, which achieves the best average accuracy and leads on six of seven evaluation sets.

URL PDF HTML 收藏
2607.14673 2026-07-17 cs.AI cs.HC 新提交

Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications

万花筒计划:面向现实世界人工智能应用的情境化、符合人类需求的评估

Leanne Tan, Rohan Jaggi, Shaun Khoo, Roy Ka-Wei Lee

机构 * GovTech(新加坡政府科技局) National University of Singapore(新加坡国立大学) University of British Columbia(英属哥伦比亚大学)

AI总结 该项目旨在解决现实世界AI应用评估瓶颈,提出万花筒计划,通过集成基于角色的测试生成、情境化评分标准和人工审查,实现可靠性门控自动评分,经试点和实验验证了其在端到端可靠自动评分方面的有用特征。

详情
AI中文摘要

评估是现实世界人工智能应用的部署瓶颈:公共基准很少能匹配团队的用户、情境或政策,人工审查往往难以扩展。受我们在公共部门人工智能应用工作的启发,该项目解决了应用必须满足当地政策和治理要求时反复出现的评估挑战。我们提出了万花筒计划,这是一种用于情境功能评估的集成工作流程,它将基于角色的测试生成、情境化评分标准和人工审查联系起来,以实现可靠性门控自动评分。生成的测试用例根据特定于应用的评分标准进行评分;人工注释提供可审查的标签;只有当大语言模型判断与这些标签的一致性达到配置的阈值时,才会自动进行评分。因此,万花筒计划是产品团队实用、可检查、迭代的工作流程。我们报告了在四个组织用例上进行的为期三周的试点以及针对跨越四个领域和14个评估维度的108个带注释问答对进行的自定义评分标准判断实验的早期证据。结果突出了端到端可靠自动评分的有用特征。

英文摘要

Evaluations (Evals) are a deployment bottleneck for real-world AI applications: public benchmarks rarely match a team's users, context, or policies, and human review is often tedious to scale. Motivated by our work with AI applications in the public sector, this project addresses recurring evaluation challenges encountered when applications must satisfy local policy and governance requirements. We present Kaleidoscope, an integrated workflow for contextual functional evaluation that links persona-based test generation, contextualized rubrics, and human review for reliability-gated automated scoring. Generated test cases are scored against application-specific rubrics; human annotations provide reviewable labels; and LLM judges automate scoring only when their agreement with those labels meets a configured threshold. Kaleidoscope is therefore a practical, inspectable, iterative workflow for product teams. We report early evidence from a three-week pilot across four organizational use cases and custom-rubric judge experiments on 108 annotated Q\&A pairs spanning four domains and 14 evaluation dimensions. The results highlight useful features for end-to-end reliable, automated scoring.

URL PDF HTML 收藏
2607.14114 2026-07-17 cs.CL cs.AI 新提交

CoEvoT: Co-Evolving Chain-of-Thought Prompting for Graph-LLM Reasoning

CoEvoT:用于图语言模型推理的协同进化思维链提示

Haohua Niu, Xingtong Yu, Yang Liu, Junfeng Fang, Xuanting Xie, Jie Tan, Zhongjian Zhang, Hong Cheng, Yuan Fang

机构 * Sun Yat-Sen University(中山大学) The Chinese University of Hong Kong(香港中文大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) National University of Singapore(新加坡国立大学) University of Electronic Science and Technology of China(电子科技大学) Beijing University of Posts and Telecommunieations(北京邮电大学) Singapore Management University(新加坡管理大学)

AI总结 研究分布转移下的图学习问题,提出CoEvoT框架,通过文本到图令牌重写与图到文本推理指导的闭环协同进化,实现逐步的、状态感知的证据细化,在八个数据集实验中性能优于现有基准模型。

Comments Under review

详情
AI中文摘要

分布转移下的图学习面临持续挑战,模型需在有限或无监督下适应新图。近期图语言模型方法通过将图线性化为提示并使用大语言模型作为预测器来实现高效标签预测,还可采用思维链提示利用大语言模型的多步推理能力。然而,现有基于思维链的图语言模型方法在固定图令牌条件下生成中间思维,限制了结构线索的逐步细化。本文提出CoEvoT,一种简单而有效的用于图语言模型推理的协同进化思维链提示框架。CoEvoT在闭环中结合文本到图令牌重写和图到文本推理指导:每个中间文本思维通过轻量级条件网络更新图令牌证据状态,更新后的令牌反馈到下一步指令以指导后续大语言模型推理。这实现了逐步的、状态感知的证据细化,而非基于固定图快照进行推理。在八个数据集上的大量实验表明,CoEvoT始终优于现有基准模型。

英文摘要

Graph learning under distribution shift presents a persistent challenge, where models adapt to new graphs with limited or even no supervision. Recent graph--LLM approaches move toward label-efficient prediction by linearizing graphs into prompts and using large language models (LLMs) as predictors, and can adopt Chain-of-Thought (CoT) prompting to exploit LLM's multi-step reasoning capability. However, existing CoT-based graph--LLM methods generate intermediate thoughts while conditioning on fixed graph tokens, limiting step-wise refinement of structural cues. In this paper, we propose CoEvoT, a simple yet effective co-evolving CoT prompting framework for graph--LLM reasoning. CoEvoT couples text-to-graph token rewriting and graph-to-text reasoning guidance in a closed loop: each intermediate textual thought is used to update the graph token evidence state via a lightweight condition network, and the updated tokens are fed back into the next-step instruction to guide subsequent LLM reasoning. This enables step-wise, state-aware evidence refinement, rather than reasoning over a fixed graph snapshot. Extensive experiments on eight datasets demonstrate that CoEvoT consistently outperforms state-of-the-art baselines.

URL PDF HTML 收藏
2604.22433 2026-07-17 cs.LG 版本更新

From physical surfaces to human-centric heat stress: LST and UTCI heat mapping reveals nonlinear effects of urban morphology

超越地表温度:可解释的空间机器学习揭示城市形态对以人类为中心的热压力的影响

Yuan Wang, Shengao Yi, Xiaojiang Li, Pengyuan Liu, Zhiwei Yang, Ronita Bardhan, Rudi Stouffs

机构 * Department of Architecture, National University of Singapore, Singapore 117566, Singapore Cambridge Centre for Advanced Research Sustainable Design Group, Department of Architecture, University of Cambridge, Cambridge, United Kingdom Department of City Regional Planning, University of Pennsylvania, Philadelphia, PA 19104, USA Urban Analytics Subject Group, Urban Studies \& Social Policy Division, University of Glasgow Laboratory for Earth Surface Processes, Ministry of Education, College of Urban Environmental Sciences, Peking University, Beijing 100871, China

AI总结 本文通过比较地表温度与通用热气候指数,揭示城市形态对人类热压力的影响,采用可解释的机器学习方法分析两者在空间分布和机制上的差异。

Comments Accepted manuscript. The final published version is available at https://doi.org/10.1016/j.scs.2026.107659

详情
AI中文摘要

热量暴露连接了建成环境与公共卫生,直接影响城市区域的宜居性和可持续性。理解热量暴露的空间异质性及其驱动因素对气候适应性城市规划至关重要。然而,大多数规划导向研究依赖于地表温度(LST),而LST是否足以代表人类热量暴露以及其与生理相关热压力的差异仍缺乏充分研究。本文采用Landsat获取的30米LST和新加坡的GPU加速1米通用热气候指数(UTCI),建立了一个综合的“建模-比较-评估”框架,系统评估两种指标的空间和机制差异。进一步,通过采用新的地理加权XGBoost(GW-XGBoost)和广义加性模型(GAM)工作流程,研究了两种指标与城市因素之间显著的非平稳和阈值型定量关系。研究结果表明,LST和UTCI的空间模式存在显著差异,以及2D和3D城市因素对这两种热指标影响的空间异质性,通过可解释的GW-XGBoost模型(LST的全局袋外R2为0.855,UTCI为0.905)得到揭示。关键的是,空间明确的SHAP解释表明,天空视因子在解释UTCI变化中起核心作用,但对LST的独立贡献相对较小,表明LST无法充分捕捉由遮荫和辐射过程决定的实际人类热压力。值得注意的是,SHAP-GAM分析表明,较高的反照率与增加的UTCI相关。这些新发现为整合生理相关的热指数以指导有针对性的热风险管理和气候适应性城市规划提供了证据。

英文摘要

Heat exposure connects the built environment and public health, directly shaping the livability and sustainability of urban areas. Understanding the spatial heterogeneity of heat exposure and its drivers is vital for climate-adaptive urban planning. However, most planning-oriented studies rely on land surface temperature (LST), and whether LST adequately represents human heat exposure and how it differs from physiologically relevant heat stress remains insufficiently examined. Here, using Landsat-retrieved 30-m LST and GPU-accelerated 1-m universal thermal climate index (UTCI) in Singapore, this study establishes a comprehensive "Modeling-Comparing-Assessing" framework to systematically evaluate the spatial and mechanistic differences between these two metrics. We further investigate their pronounced non-stationary and threshold-based relationships with urban factors using a novel geographically weighted XGBoost (GW-XGBoost) and generalized additive model (GAM) workflow. Our results reveal substantial differences in the spatial patterns of LST and UTCI, along with marked spatial heterogeneity in how 2D and 3D urban factors impact these thermal metrics, as demonstrated by explainable GW-XGBoost models (test R2 = 0.855 for LST and 0.905 for UTCI). Crucially, spatially explicit SHAP shows that sky view factor plays a central role in explaining UTCI variability but exhibits a comparatively marginal independent contribution to LST, indicating that LST inadequately captures shading-driven and radiative processes governing actual human heat stress. Moreover, SHAP-GAM analysis indicates that higher albedo is associated with increased UTCI. These findings provide model-informed planning implications for integrating physiologically relevant thermal indices to support targeted heat risk management and human-centric urban planning.

URL PDF HTML 收藏
2511.22098 2026-07-17 cs.CV 版本更新

WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation

WorldWander:在视频生成中连接自我中心和外中心世界

Quanjian Song, Yiren Song, Kelly Peng, Yuan Gao, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室)

AI总结 研究聚焦视频生成中自我中心与外中心世界转换,提出WorldWander框架,基于视频扩散变换器,整合上下文视角对齐与协作位置编码,精心策划数据集,实验证明该框架在视角同步、角色一致性和泛化能力上表现卓越,设定了新基准。

Comments Accepted by ECCV 2026; Code: https://github.com/showlab/WorldWander

详情
AI中文摘要

视频世界模型的最新进展实现了具有自由导航的交互式环境,使得第一人称(自我中心)和第三人称(外中心)视角之间的转换变得越发重要。然而,现有研究集中于单向的从外中心到自我中心的转换,忽视了参考引导的外中心视角合成。这种能力对游戏和具身人工智能应用至关重要。为此,我们提出了WorldWander,一个为视频生成中自我中心和外中心世界之间的转换量身定制的上下文学习框架。基于先进的视频扩散变换器,WorldWander整合了上下文视角对齐和协作位置编码,以对跨视角同步和角色一致性进行建模。为支持我们的任务,我们精心策划了EgoExo-8K,这是一个包含来自合成和现实世界场景的同步自我中心-外中心三元组的动态且场景丰富的数据集。实验表明,WorldWander实现了卓越的视角同步、角色一致性和泛化能力,为自我中心-外中心视频转换设定了新的基准。

英文摘要

Recent advances in video world models enable interactive environments with free navigation, making translation between first-person (egocentric) and third-person (exocentric) perspectives increasingly important. However, existing studies focus on unidirectional exocentric-to-egocentric translation, overlooking reference-guided exocentric perspective synthesis. This capability is crucial for gaming and embodied AI applications. Motivated by this, we present WorldWander, an in-context learning framework tailored for translating between egocentric and exocentric worlds in video generation. Building upon advanced video diffusion transformers, WorldWander integrates (i) In-Context Perspective Alignment and (ii) Collaborative Position Encoding to model cross-view synchronization and character consistency. To support our task, we curate EgoExo-8K, a dynamic and scene-rich dataset containing synchronized egocentric-exocentric triplets from both synthetic and real-world scenarios. Experiments demonstrate that WorldWander achieves superior perspective synchronization, character consistency, and generalization, setting a new benchmark for egocentric-exocentric video translation.

URL PDF HTML 收藏
2607.13841 2026-07-16 cs.LG stat.ML 新提交

Heavy-Tailed Flow Matching via Random Clocks

通过随机时钟进行重尾流匹配

Zhouhao Yang, Yezhen Wang, Kenji Kawaguchi, Vladimir Braverman, Haoyang Cao

机构 * Johns Hopkins University(约翰斯·霍普金斯大学) National University of Singapore(新加坡国立大学)

AI总结 研究重尾数据匹配问题,提出HTFM框架,将重尾源视为时钟条件高斯源混合,用截断对数签名特征编码时钟。实验显示其在多领域优于高斯流匹配等基线,保留低NFE采样优势,还提供尾部控制接口。

详情
AI中文摘要

重尾数据出现在许多领域,如不平衡图像数据集、金融回报和极端天气等,其中罕见事件具有不成比例的重要性。标准扩散和流匹配模型通常从高斯噪声或高斯源分布开始,对重尾数据的归纳匹配较差。我们提出了通过随机时钟进行重尾流匹配(HTFM)框架,将重尾源描绘为时钟条件高斯源的混合。给定时钟路径时,源分布和流是高斯的;对时钟求边缘分布得到覆盖高斯、α稳定和学生t族的高斯尺度混合。为使时钟条件向量场实用,我们使用截断对数签名特征编码路径值时钟,使速度场能以可忽略的开销适应已实现的条件空间。实验表明,在二维不平衡α稳定混合、CIFAR10-LT和HRRR天气场中,HTFM在模式覆盖、样本质量和尾部统计恢复方面优于高斯流匹配和有竞争力的重尾基线,同时保留了流匹配的低NFE采样优势。此外,随机时钟公式还提供了一个实用的尾部控制接口。

英文摘要

Heavy-tailed data arise in many domains where rare events carry disproportionate importance, such as imbalanced image datasets, financial returns, and weather extremes. Standard diffusion and flow-matching models typically begin from Gaussian noise or Gaussian source distributions, which yield tractable training targets but provide a poor inductive match for heavy-tailed data. We propose Heavy-Tailed Flow Matching via Random Clocks (HTFM), a framework that portrays heavy-tailed sources as mixtures of clock-conditioned Gaussian sources. Conditioning on a given clock path, the source distribution and flow are Gaussian; marginalizing over the clock gives a Gaussian scale mixture covering Gaussian, $α$-stable, and Student-t families. To make the clock-conditioned vector field practical, we encode the path-valued clock using truncated logsignature features, allowing the velocity field to adapt to the realized conditional space with negligible overhead. Empirically, on 2D imbalanced $α$-stable mixtures, CIFAR10-LT, and HRRR weather fields, HTFM improves mode coverage, sample quality, and tail-statistic recovery over Gaussian flow matching and competitive heavy-tailed baselines, while retaining the low-NFE sampling advantage of flow matching. Moreover, the random-clock formulation further provides a practical tail-control interface: by varying only the clock law or tail parameter, the same architecture can calibrate the ``heaviness'' of generated tails across different distribution families.

URL PDF HTML 收藏
2607.13837 2026-07-16 cs.LG cs.AI 新提交

NodeImport: Imbalanced Node Classification with Node Importance Assessment

NodeImport:基于节点重要性评估的不平衡节点分类

Nan Chen, Zemin Liu, Bryan Hooi, Bingsheng He, Jun Hu, Jia Chen

机构 * Johns Hopkins University(约翰斯·霍普金斯大学) Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学) Grabtaxi Holdings Pte Ltd(Grab出租车控股私人有限公司)

AI总结 针对图节点分类的类别不平衡问题,提出利用平衡元集评估节点重要性的方法,推导公式并开发新框架,能过滤有价值节点,构建高质量元集,实验证明该框架在缓解不平衡上具优越性。

Journal ref KDD '25: Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, 2025, Pages 94 - 105

详情
AI中文摘要

在实际应用中,图上的节点分类常面临类别不平衡挑战,多数类主导训练导致模型性能有偏差。传统GNN在此场景中表现不佳。现有解决方案存在不足。本文提出利用平衡元集进行重要性度量的方法,识别可抵消类别不平衡的重要节点用于模型训练,理论推导直接评估节点重要性的公式,开发新框架过滤有价值节点,还介绍构建高质量元集的策略。通过实验证明该框架在缓解类别不平衡方面的优越性。

英文摘要

In real-world applications, node classification on graphs often faces the challenge of class imbalance, where majority classes dominate training, resulting in biased model performance. Traditional GNNs often struggle in such scenarios, as they tend to overfit to majority classes while underrepresenting minority classes. Existing solutions, which either prioritize nodes based on class size or synthesize new nodes for minority classes, often fall short of effectively addressing this imbalance issue. This paper introduces an approach to class-imbalanced node classification by utilizing a balanced meta-set for importance measurement, where a training node is considered significant if it enhances model performance under an unbiased setting. Our method identifies important nodes that can counteract class imbalance and utilizes them for model training, allowing for fine-grained and dynamic node selection throughout the training process. We theoretically derive a formula to directly assess node importance, reducing computational overhead and providing an intuitive threshold for node selection. Guided by this metric, we develop a novel framework that filters valuable labeled, unlabeled, and synthetic nodes that enhance model performance in an unbiased context. A key advantage of this framework is its separation of the synthetic node generation process from the filtering process, ensuring compatibility with various node generation methods. Furthermore, we introduce a strategy to construct a high-quality meta-set that closely approximates the overall feature distribution, ensuring robust representation of each class. We evaluate our framework, NodeImport, across multiple datasets using popular GNN architectures, demonstrating its superiority over existing baselines. Our results highlight the flexibility and effectiveness of the framework in mitigating class imbalance, leading to improved outcomes.

URL PDF HTML 收藏