arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Cambridge(剑桥大学)

至 收录 1276
2607.17117 2026-07-21 cs.LG cs.CL 新提交

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models

持久稀疏自动编码器:在语言模型中学习特征时间尺度

Haoyan Luo, Mateo Espinosa Zarlenga, Mateja Jamnik

机构 * University of Cambridge(剑桥大学) University of Oxford(牛津大学)

AI总结 研究在语言模型中学习特征时间尺度的问题,提出持久稀疏自动编码器,通过为特征学习持久性系数扩展标准SAEs,实验表明其能保持竞争力的重建质量,为解释和监测语言模型带来新机会。

详情
AI中文摘要

稀疏自动编码器(SAEs)将语言模型激活分解为稀疏特征,但标准SAEs独立编码每个令牌,不暴露跨序列持续存在的信息。我们引入了持久稀疏自动编码器(Persistent SAEs),它通过为每个特征学习一个持久性系数来扩展标准SAEs,使模型能够学习哪些特征应该持续以及持续多长时间。我们的实验表明,它们在学习一系列特征时间尺度时保持了有竞争力的重建质量:快速特征表现为局部可解释的检测器,而慢速特征在持久状态下集中主题级信息。此外,如在提示注入监测案例研究中所示,慢速特征保留检测信号并在长上下文中保持因果有效性。这些结果表明,持久稀疏自动编码器通过持久语义表示为解释和监测语言模型开辟了新机会。

英文摘要

Sparse autoencoders (SAEs) decompose language model activations into sparse features, but standard SAEs encode each token independently and do not expose information that persists across a sequence. We introduce Persistent Sparse Autoencoders (Persistent SAEs), which extend standard SAEs by learning a persistence coefficient for each feature, allowing the model to learn which features should persist and for how long. Our experiments show that they retain competitive reconstruction quality while learning a spectrum of feature timescales: fast features behave as locally interpretable detectors, whereas slow features concentrate topic-level information in a persistent state. Moreover, as shown in a prompt-injection monitoring case study, slow features preserve detection signals and remain causally effective over long contexts. These results suggest that Persistent SAEs open up new opportunities for interpreting and monitoring language models through persistent semantic representations.

URL PDF HTML 收藏
2606.26454 2026-07-21 cs.AI 版本更新

Data-driven Machine Learning Cannot Reach Symbolic-level Logical Reasoning -- The Limit of the Scaling Law

数据驱动的机器学习无法达到符号级逻辑推理——缩放定律的极限

Tiansi Dong, Mateja Jamnik, Pietro Liò

机构 * The Alan Turing Institute(艾伦·图灵研究所) University of Cambridge(剑桥大学)

AI总结 本文通过理论分析和实验证明,数据驱动的监督深度学习无法通过增加训练数据和时间达到符号级三段论推理的严谨性,并揭示了训练数据和方法上的根本限制。

详情
AI中文摘要

球面神经网络无需训练数据即可实现符号级三段论推理,这引发了一个问题:逻辑推理的缩放定律极限在哪里?即数据驱动的机器学习系统能否通过增加训练数据和训练时间达到同样的水平。我们展示了两个方法论上的限制,阻碍了监督深度学习达到符号级三段论推理:(1)训练数据无法区分所有24种有效的三段论推理类型;(2)从前提到结论的端到端映射在模式识别和逻辑推理的神经组件之间引入了矛盾的训练目标。除了理论分析,我们通过实验说明欧拉网络无法实现严谨的三段论推理。我们进一步挑战最新的ChatGPTs(GPT-5-nano和GPT-5),以确定三段论语句在四种表面形式(模式)下的可满足性:单词、双词、简单符号和长随机符号,结果表明表面形式影响推理性能,且ChatGPT GPT-5可能达到100%准确率但仍提供不正确的解释。由于经验训练过程在达到100%准确率后停止,我们得出结论:监督机器学习系统无法达到符号逻辑推理的严谨性。

英文摘要

By promoting vectors to spheres and enabling explicit model construction, neural networks can perform symbolic-level syllogistic reasoning without training data. We identify two fundamental limitations that prevent conventional data-driven machine learning systems from achieving this capability: training data generated by the combination table cannot distinguish all 24 valid syllogism types, and end-to-end premise-to-conclusion mapping creates contradictory targets within neural components. Experiments with two representative conventional systems, GPT-5 using linguistic inputs and Euler Net using visual inputs, support this analysis. ChatGPT GPT-5 may reach 100% accuracy in syllogistic reasoning, but with hallucinations. Because the learning process terminates upon reaching 100% accuracy, the system cannot progress beyond empirical accuracy to symbolic level reasoning. Random test data reduced Euler Net's accuracy to 56%. Repeatedly expanding the training set increased its accuracy to 97%, with perfect performance on 8 syllogism types. However, because unintended inputs cannot be exhaustively covered, even 100% test accuracy does not imply symbolic-level reasoning. Since syllogistic reasoning underpins logical reasoning and human rationality, these results suggest that increasing data and training time alone cannot ensure symbolic level logical reasoning.

URL PDF HTML 收藏
2607.15992 2026-07-20 cs.AI 新提交

Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI

缩小人工智能信任差距:可信人工智能独立认证的案例

Trisevgeni Papakonstantinou, Cansu Canca, Farah Nanji, Waheedullah Pardess, Jen Weedon, Jasmijn Remmers, Eliza Krigman, Matthew Ball, Yalda Daryani, Kiran Iqbal, Francielle Vargas, María Llorente Sánchez, Joe Humphreys, Fendi Tsim, Kelly Fitzpatrick, Jeff Dunn, Catherine Feldman

机构 * University College London(伦敦大学学院) AI Ethics Lab(人工智能伦理实验室) University of Cambridge(剑桥大学) Columbia University(哥伦比亚大学) MATS Tech with Intention(技术与意图) University of Southern California(南加州大学) Kyushu University(九州大学) University of Chile(智利大学) BehSci Meets AI(行为科学与人工智能) Digital Trust Council(数字信任委员会)

AI总结 研究指出负责任AI虽有实践但未形成奖励可信度的市场,存在信任差距。原因包括关注重点偏差及三个复合失败。通过审查治理工具发现不足,进而提出以结果为导向的独立认证来缩小信任差距,补充监管与内部治理。

详情
AI中文摘要

在过去十年中,负责任的人工智能(RAI)已产生大量实践,用于识别和减轻人工智能在高风险环境中带来的风险。然而,这项工作尚未产生一个奖励可信度的市场。认真投资于安全、公平和监督的公司无法向消费者、监管机构和股东持续证明其系统超越了最低合规标准。社会缺少一种认可或比较差异的方式,导致了信任差距。我们认为,这种差距部分是由于关注负责任的人工智能(内部流程问题)而非可信的人工智能(可独立验证的现实世界结果问题),并且由于三个复合失败而持续存在:(1)市场无法区分可信系统与其模仿品;(2)评估针对模型和输出而非部署的社会技术系统及其结果;(3)测量生态系统旨在避免伤害而非证明益处。通过审查现有人工智能治理工具并将其与医疗保健、可持续性和安全领域的认证制度进行比较,我们表明没有一个在单一框架中整合治理基线、独立验证的积极结果证据和市场信号。我们提出以结果为导向的独立认证作为可以缩小信任差距的连接层,通过使可信度可衡量、可比较和获得商业回报来补充监管和内部治理。

英文摘要

Over the past decade, responsible AI (RAI) has produced a substantial body of practice for identifying and mitigating the risks AI poses in high-stakes settings. Yet this work has not produced a market that rewards trustworthiness. Firms that invest seriously in safety, fairness, and oversight cannot consistently prove to consumers, regulators, and shareholders that their systems go beyond the bare minimum of compliance. What is missing is a way for society to recognize or compare the difference. The result is a trust gap: a structural condition in which responsible development efforts happen inside organizations but produce no external, independently recognized and verifiable signal of trustworthy outcomes. We argue this gap is sustained in part because of a focus on responsible AI (a matter of internal process) as opposed to trustworthy AI (a matter of independently verifiable real-world outcomes), and that it persists because of three compounding failures: (1) the market cannot distinguish trustworthy systems from their imitations; (2) evaluation targets models and outputs rather than deployed sociotechnical systems and their outcomes; (3) the measurement ecosystem is oriented toward avoiding harm rather than demonstrating benefit. Reviewing existing AI governance instruments and comparing them to certification regimes in healthcare, sustainability, and security, we show that none integrate a governance baseline, independently verified positive-outcome evidence, and market signaling in a single framework. We propose independent, outcome-oriented certification as the connective layer that can close the trust gap, complementing regulation and internal governance by making trustworthiness measurable, comparable, and commercially rewarded.

URL PDF HTML 收藏
2607.15472 2026-07-20 math.DS cs.LG physics.bio-ph 新提交

Ptolemy's Equant Equates to a Universal Dynamical Clock via Machine Learning

通过机器学习,托勒密等距点等同于通用动态时钟

Jingdong Zhang, Luan Yang, Murilo S. Baptista, Zefeng Zhang, Qunxi Zhu, Wei Lin, Celso Grebogi

机构 * School of Mathematical Sciences, Fudan University, Shanghai 200433, China(复旦大学数学科学学院) Research Institute of Intelligent Complex Systems, Fudan University, Shanghai 200433, China(复旦大学智能复杂系统研究所) Department of Mathematics, Imperial College London, London, SW7 2AZ, United Kingdom(伦敦帝国学院数学系) Institute for Complex Systems and Mathematical Biology, University of Aberdeen, Aberdeen AB24 3UE, United Kingdom(阿伯丁大学复杂系统与数学生物学研究所) Department of Psychiatry, University of Cambridge, Cambridge CB2 1TN, United Kingdom(剑桥大学精神病学系)

AI总结 研究非线性高维振荡中相位和相位动力学识别问题,基于托勒密等距点建立通用动态时钟原理,用机器学习框架证明其存在并构建相关动力学,通过四个发现展示价值,为振荡系统研究提供新途径。

Comments 56 pages, 12 figures, 3 tables

详情
AI中文摘要

振荡动力学在非线性系统中普遍存在,但在非线性高维振荡中识别具有物理可解释性的相位和相位动力学仍是一个未解决的核心问题。本文建立了通用动态时钟原理,受托勒密等距点启发,通过面积均匀性原理形式化,将任意维度和几何形状的振荡等效表示为通过等距点诱导的非线性观察坐标的匀速旋转。利用机器学习框架,证明了一类广泛振荡动力学中等距点的存在,并构建了相关动态时钟和相位动力学。通过四个发现展示了其在揭示新物理规则和现象方面的价值,包括大肠杆菌群体中的集体振荡遵循超线性缩放定律、工程遗传电路对基因表达和环境条件变化的响应机制、贝里几何相位的经典力学对应物自然出现以及最优等距点非均匀性为临界转变提供几何预警信号并预测临界参数。动态时钟提供了可直接从数据构建的操作和系统无关的相位动力学,实现了对振荡系统的分类、比较和控制,为理解特定动态机制如何支持网络系统中不同功能行为提供了新途径。

英文摘要

Oscillatory dynamics arise ubiquitously in nonlinear systems, yet identifying a physically interpretable phase and phase dynamics in nonlinear, high-dimensional oscillations remains a central unresolved problem. Here we establish the principle of a universal dynamical clock, a physical perspective in which oscillations of arbitrary dimensionality and geometry are equivalently represented as uniform rotation through an equant-induced nonlinear viewing coordinate, inspired by Ptolemy's equant and formalised through an areal-uniformity principle reminiscent of Kepler's second law. Using a machine-learning framework, we demonstrate the existence of such an equant for a broad class of oscillatory dynamics and construct the associated dynamical clock and phase dynamics under additive forces, including noise, periodic perturbations, and coupling. Its value in uncovering new physical rules and phenomena is demonstrated by four findings: (i) collective oscillations in Escherichia coli populations obey a previously unexplained superlinear scaling law, resolving a long-standing open problem posed in 2004; (ii) the response mechanisms of engineered genetic circuits to changes in gene expression and environmental conditions; (iii) a classical-mechanics counterpart of the Berry geometric phase emerges naturally from the phase of the dynamical clock; and (iv) optimal equant non-uniformity provides a geometric early-warning signal for critical transitions and enables prediction of critical parameters. By providing operational and system-agnostic phase dynamics that can be constructed directly from data, the dynamical clock enables principled classification, comparison, and control of oscillatory systems, and offers a new route to understanding how specific dynamical regimes support distinct functional behaviours in networked systems.

URL PDF HTML 收藏
2606.16900 2026-07-20 cs.LG 版本更新

Factorized Neural Operators Decompose Dynamic and Persistent Responses

因子化神经算子分解动态与持久响应

Hao Tang, Yuechen Duan, Jiongyu Zhu, Zimeng Feng, Hao Li, Chao Li

机构 * School of Medicine, University of Dundee(邓迪大学医学院) School of Data Science, Fudan University(复旦大学数据科学学院) School of Mathematical Sciences, Fudan University(复旦大学数学科学学院) Institute of Science and Technology for Brain-inspired Intelligence, Fudan University(复旦大学类脑智能科学与技术研究院) School of Science and Engineering, University of Dundee(邓迪大学科学与工程学院) Department of Applied Mathematics and Theoretical Physics, University of Cambridge(剑桥大学应用数学与理论物理系)

AI总结 提出因子化神经算子(FaNO),通过分解谱表示为等变动态响应和不变持久响应,提升多尺度物理系统的预测精度、参数效率和跨尺度泛化能力。

详情
AI中文摘要

物理系统通常表现出异质性机制,其中快速演变的动力学与持久结构共存。现有的神经算子通常依赖单一主导归纳偏置,因此将不同的物理响应耦合到共享表示中,难以捕捉这种多尺度物理行为。我们引入了跨域的统一格林函数框架,并提出了因子化神经算子(FaNO),它将谱表示分解为等变动态响应和不变持久响应,从而提高了可解释性和泛化能力。从机制上讲,我们展示了两个算子分支自发地特化为不同的物理角色,这些角色在尺度和域上保持一致:等变分支捕捉快速变化的瞬态动力学,而不变分支提取连贯的持久结构。FaNO的这种因子化机制提高了跨物理系统和域的预测精度、参数效率和跨尺度泛化能力。特别是,它在长时程自回归滚动、跨分辨率外推和物理状态转移下保持一致的预测。这些发现表明,可扩展的物理建模可能受益于从单一归纳偏置公式转向更好地反映物理系统异质性组织的因子化算子表示,从而加速机器学习在科学计算和发现中的可靠部署。

英文摘要

Physical systems often exhibit heterogeneous mechanisms, where rapidly evolving dynamics coexist with persistent structures. Capturing such multiscale physical behavior remains challenging for existing neural operators, which typically rely on single dominant inductive bias and therefore couple distinct physical responses into a shared representation. We introduce the Unified Green's Function Framework across domains and propose the Factorized Neural Operators (FaNO), which decompose spectral representations into equivariant-inspired dynamic responses and invariant-inspired persistent responses, leading to better interpretability and generalization. Mechanistically, we show that the two operator branches spontaneously specialize into distinct physical roles that remain consistent across scales and domains: the equivariant-inspired branch captures rapidly varying transient dynamics, whereas the invariant-inspired branch extracts coherent persistent structures. This factorized mechanism of FaNO consistently improves prediction accuracy, parameter efficiency and cross-scale generalization across physical systems and domains. In particular, it maintains consistent predictions under long-horizon autoregressive rollout, cross-resolution extrapolation and physical-regime shifts. These findings suggest that scalable physical modeling may benefit from moving beyond single-inductive-bias formulations toward factorized operator representations that better reflect the heterogeneous organization of physical systems, accelerating the reliable deployment of machine learning for scientific computing and discovery.

URL PDF HTML 收藏
2604.23786 2026-07-20 cs.AI cs.LG 版本更新

FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment

FAIR_XAI: 通过可解释性提升多模态基础模型公平性以用于幸福感评估

Sophie Chiang, Tom Brennan, Fethiye Irmak Dogan, Jiaee Cheong, Hatice Gunes

机构 * Department of Computer Science & Technology, University of Cambridge(计算机科学与技术系,剑桥大学) Harvard University(哈佛大学)

AI总结 本文研究了多模态基础模型在幸福感评估中的公平性问题,通过可解释性干预框架改善诊断可靠性与公平性,发现不同模型在不同数据集上表现差异显著,且存在性别和种族偏见。

Comments 11 pages, 4 figures

详情
AI中文摘要

近年来,多模态机器学习在幸福感评估中的整合为心理健康监测提供了变革性潜力。然而,随着视觉-语言模型(VLMs)的快速发展,其在临床应用中的部署引发了透明度不足和潜在偏见的担忧。尽管先前研究探讨了公平性与可解释人工智能(XAI)的交集,但将其应用于VLMs进行幸福感评估和抑郁症预测仍显不足。本文研究了VLMs在实验室(AFAR-BSFT)和自然(E-DAIC)数据集上的表现,聚焦诊断可靠性与人口公平性。性能在不同环境和架构间差异显著;Phi3.5-Vision在E-DAIC上达到80.4%的准确率,而Qwen2-VL在同数据集上仅33.9%。此外,两种模型在AFAR-BSFT上均表现出对抑郁症的过度预测倾向。尽管两种架构均存在偏见,但Qwen2-VL显示出更高的性别差异,而Phi-3.5-Vision则表现出更多的种族偏见。我们的XAI干预框架产生了混合结果;公平性提示在Qwen2-VL上实现了完美的相等机会,但以严重的准确率代价。在AFAR-BSFT上,基于可解释性的干预提高了程序一致性,但未保证结果公平性,有时加剧了种族偏见。这些结果突显了程序透明度与公平结果之间的持续差距。我们分析了这些发现并提出了具体的解决建议,强调未来公平性干预必须共同优化预测准确性、人口平等性和跨领域泛化能力。

英文摘要

In recent years, the integration of multimodal machine learning in wellbeing assessment has offered transformative potential for monitoring mental health. However, with the rapid advancement of Vision-Language Models (VLMs), their deployment in clinical settings has raised concerns due to their lack of transparency and potential for bias. While previous research has explored the intersection of fairness and Explainable AI (XAI), its application to VLMs for wellbeing assessment and depression prediction remains under-explored. This work investigates VLM performance across laboratory (AFAR-BSFT) and naturalistic (E-DAIC) datasets, focusing on diagnostic reliability and demographic fairness. Performance varied substantially across environments and architectures; Phi3.5-Vision achieved 80.4% accuracy on E-DAIC, while Qwen2-VL struggled at 33.9%. Additionally, both models demonstrated a tendency to over-predict depression on AFAR-BSFT. Although bias existed across both architectures, Qwen2-VL showed higher gender disparities, while Phi-3.5-Vision exhibited more racial bias. Our XAI intervention framework yielded mixed results; fairness prompting achieved perfect equal opportunity for Qwen2-VL at a severe accuracy cost on E-DAIC. On AFAR-BSFT, explainability-based interventions improved procedural consistency but did not guarantee outcome fairness, sometimes amplifying racial bias. These results highlight a persistent gap between procedural transparency and equitable outcomes. We analyse these findings and consolidate concrete recommendations for addressing them, emphasising that future fairness interventions must jointly optimise predictive accuracy, demographic parity, and cross-domain generalisation.

URL PDF HTML 收藏
2509.18127 2026-07-20 cs.LG cs.AI cs.CL

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework

Safe-SAIL: 通过稀疏自编码解释框架构建大语言模型的细粒度安全景观

Jiaqi Weng, Han Zheng, Hanyu Zhang, Ej Zhou, Qinqin He, Jialing Tao, Hui Xue, Zhixuan Chu, Xiting Wang

机构 * Alibaba Group(阿里巴巴集团) The State Key Laboratory of Blockchain and Data Security, Zhejiang University(浙江大学区块链与数据安全国家重点实验室) Language Technology Lab, University of Cambridge(剑桥大学语言技术实验室) Renmin University of China(中国人民大学)

AI总结 本文提出Safe-SAIL框架,通过稀疏自编码解释方法高效识别安全领域特征,减少解释成本55%,并系统评估1758个安全相关特征,揭示风险特征识别和安全关键实体编码机制。

Journal ref Findings of the Association for Computational Linguistics: ACL 2026, pages 18916-18935 (2026)

详情
AI中文摘要

稀疏自编码(SAEs)通过将纠缠的模型激活分解为单语义特征,推动可解释性研究。然而,SAEs在何种情况下能为低频概念领域生成最细粒度的潜在特征仍不清楚。本文提出Safe-SAIL框架,旨在安全关键领域解释SAE特征,以提升大语言模型的机理理解。Safe-SAIL引入预解释评估指标,高效识别具有强安全领域可解释性的SAEs,并通过段级模拟策略将解释成本降低55%。基于Safe-SAIL,我们训练了涵盖四个领域(色情、政治、暴力和恐怖)的1758个安全相关特征的综合SAE集合,提供可读解释和系统评估。利用此资源,我们进行了实证分析,探讨了Safe-SAIL在风险特征识别中的有效性,以及安全关键实体和概念在模型层间的编码机制。所有模型、解释和工具均在开源工具包和配套产品中公开发布。

英文摘要

Sparse autoencoders (SAEs) enable interpretability research by decomposing entangled model activations into monosemantic features. However, under what circumstances SAEs derive most fine-grained latent features for safety, a low-frequency concept domain, remains unexplored. Two key challenges exist: identifying SAEs with the greatest potential for generating safety domain-specific features, and the prohibitively high cost of detailed feature explanation. In this paper, we propose Safe-SAIL, a unified framework for interpreting SAE features in safety-critical domains to advance mechanistic understanding of large language models. Safe-SAIL introduces a pre-explanation evaluation metric to efficiently identify SAEs with strong safety domain-specific interpretability, and reduces interpretation cost by 55% through a segment-level simulation strategy. Building on Safe-SAIL, we train a comprehensive suite of SAEs with human-readable explanations and systematic evaluations for 1,758 safety-related features spanning four domains: pornography, politics, violence, and terror. Using this resource, we conduct empirical analyses and provide insights on the effectiveness of Safe-SAIL for risk feature identification and how safety-critical entities and concepts are encoded across model layers. All models, explanations, and tools are publicly released in our open-source toolkit and companion product.

URL PDF HTML 收藏
2607.14591 2026-07-17 cs.CL 新提交

How Well Does AI-Generated Feedback Work? Intrinsic and Extrinsic Evaluation across more than 20,000 EFL Essay Drafts

人工智能生成的反馈效果如何?对20000多篇外语作文草稿的内在和外在评估

Steven Coyne, Diana Galvan-Sosa, Ryan Spring, Machi Shimmei, Michael Zock, Keisuke Sakaguchi, Kentaro Inui

机构 * Tohoku University(东北大学) RIKEN(理化学研究所) ALTA Institute, Computer Laboratory, University of Cambridge(剑桥大学ALTA研究所,计算机实验室) CNRS, LIS, Aix-Marseille University(法国国家科学研究中心,艾克斯-马赛大学语言信息处理实验室) MBZUAI(Mohamed bin Zayed大学人工智能学院)

AI总结 研究外语写作中人工智能生成的书面纠正性反馈效果,通过大学外语班级近2000名学生的超20000篇草稿,从教师内在评估和学生外在反馈两角度评估,发现传统专家评估与学生反馈一致性低,强调以学习者为中心评估框架的重要性。

Comments Pre-review version of DOI https://doi.org/10.1007/978-3-032-29788-4_35, presented at AIED 2026 Late Breaking Results. Readers are encouraged to refer to the published version

Journal ref AIED CCIS 3031 (2026) 247-253

详情
AI中文摘要

本研究考察外语写作环境中的反馈,聚焦书面纠正性反馈(WCF)。大语言模型能大规模提供WCF,但使其符合教学最佳实践仍是一项持续挑战。符合事实性或相关性等标准的WCF可能仍不适用于学习情境,凸显基于学习者视角进行外在评估的必要性。我们在一个有近2000名学生的大学外语班级中部署WCF系统,收集了20000多篇草稿。从两个角度评估生成的WCF:经验丰富的英语教师使用评分标准进行内在评估,以及通过学生反馈和参与度指标进行外在评估。结果显示教师专家评分与学生反馈之间的一致性较低。这些发现表明,仅靠传统专家评估可能无法从学习者角度充分捕捉WCF的可用性或帮助性,凸显了以学习者为中心的评估框架对语言教育中基于人工智能的应用的重要性。

英文摘要

This study examines feedback in English as a Foreign Language (EFL) writing contexts, focusing on written corrective feedback (WCF). Large language models (LLMs) can provide WCF at scale, but aligning them with pedagogical best practices remains an ongoing challenge. WCF meeting criteria like factuality or relevance may still be unsuitable for learning contexts, highlighting the need for extrinsic evaluation based on the learner's perspective. We deployed WCF systems in a university-level EFL class with nearly 2,000 students, collecting over 20,000 drafts. We evaluated the generated WCF from two perspectives: intrinsic evaluation by experienced English teachers using a rubric, and extrinsic evaluation via student feedback and engagement metrics. Results revealed low alignment between teacher expert ratings and student feedback. These findings suggest that traditional expert evaluation alone may not fully capture WCF's usability or helpfulness from the learner's perspective, highlighting the importance of learner-centered evaluation frameworks for AI-based applications in language education.

URL PDF HTML 收藏
2607.14480 2026-07-17 cs.CL 新提交

LLM Evaluators are Biased across Languages

语言模型评估器在不同语言间存在偏差

Ej Zhou, Lucas Resck, Zheng Hui, Anna Korhonen

机构 * University of Cambridge(剑桥大学) Language Technology Lab(语言技术实验室)

AI总结 研究发现语言模型评估器在多语言环境中存在偏差,不同语言评分差异显著,与语言资源水平相关,成对准确率无法检测到这些偏差,还探究了资源少的语言得分高的原因,揭示了语言层面的结构错位。

详情
AI中文摘要

语言模型评估器(经过训练的奖励模型和作为评判的提示语言模型)通常通过成对准确率进行验证。在多语言环境中,这基于高成对准确率意味着可靠、语言中立评分的前提。但研究表明该假设不成立。通过对23种语言的语义相同的指令 - 响应对进行实验,发现多语言评估器对不同评估语言给出显著不同分数。偏差具有统计学意义且在不同架构和训练范式的八个开放权重评估器中一致,在前沿评判中也存在,且与语言资源水平强烈相关,资源少的语言评分更宽松。这些偏差成对准确率检测不到,评估器成对准确率超90%,但在全局决策阈值下不同语言接受率差异达43%。研究还探讨了资源少的语言得分高的原因,发现模型不确定性与之有关,且偏差是结构上的语言层面错位,不能仅由内容难度解释。

英文摘要

LLM evaluators (trained reward models and prompted LLM-as-a-Judge) are routinely validated via pairwise accuracy. In a multilingual setting, this operates under the premise that high pairwise accuracy implies reliable, language-neutral scoring. We show that this assumption does not hold. We conduct experiments with semantically identical instruction-response pairs across 23 languages, and find that multilingual evaluators assign significantly different scores to different evaluation languages. The bias is statistically significant and consistent across eight open-weight evaluators of different architectures and training paradigms, persists in frontier judges, and is strongly correlated with language resource level: lower-resource languages are scored more generously. Meanwhile, these biases are invisible to pairwise accuracy: evaluators achieve above 90% pairwise accuracy, yet have up to 43% difference in acceptance rate across languages under a global decision threshold, meaning, for instance, that harmful content in lower-resource languages is more likely to pass safety filters. Per-language thresholds would require language identification, which can be defeated by code-switched prompts. We then investigate why lower-resource languages receive higher rather than lower scores, and we find that model uncertainty is linked with the effect: models tend to give higher scores when less confident, both under negative log-likelihood and under token-free uncertainty measures; however, language identity remains a significant predictor after controlling for uncertainty, and the bias cannot be explained away by content difficulty alone, but is a structural, language-level misalignment.

URL PDF HTML 收藏
2607.13891 2026-07-17 cs.LG cs.CV eess.SP stat.ML 交叉投稿

PiVoT: A Variational Solution for Real-time Large-scale Multi-object Detection and Tracking under Heavy Clutter

PiVoT:一种用于在严重杂波下实时大规模多目标检测与跟踪的变分解决方案

Runze Gan, Qing Li, Simon J. Godsill, Mike E. Davies, James R. Hopgood

机构 * Institute for Imaging, Data and Communications (IDCOM), University of Edinburgh(爱丁堡大学成像、数据与通信研究所(IDCOM)) Department of Engineering, University of Cambridge(剑桥大学工程系) School of Mathematics, University of Edinburgh(爱丁堡大学数学学院)

AI总结 针对数据稀缺雷达应用中多目标检测跟踪难题,PiVoT通过联合推断目标多方面信息,无需外部聚类或检测器,利用变分推断创新实现快速抗杂波跟踪,实验证明其在多方面性能出色,优于现有贝叶斯跟踪器。

详情
AI中文摘要

在许多数据稀缺的雷达应用中,从噪声点云进行多目标检测和跟踪仍然具有挑战性。当前基于泊松测量模型的贝叶斯跟踪器提供了一种无需训练的解决方案,但在严重杂波、大量目标和全分辨率多普勒点云情况下,难以实现准确性和效率。我们使用PiVoT来解决这个问题,它是一种用于位置和多普勒测量的快速、抗杂波多目标跟踪器。PiVoT通过联合推断目标状态、形状、存在概率、数据关联和测量率,对大量且随时间变化的目标进行端到端检测和跟踪,无需外部聚类或检测器。其效率得益于多种变分推断创新,如理论上合理的出生剪枝算法、精确更新的二次到线性复杂度降低以及计算高效的多普勒泊松模型。实验表明,PiVoT在具有挑战性的场景中大大优于现有的贝叶斯跟踪器,同时还展示了对一千个目标的出色可扩展性、对与目标视觉上无法分离的杂波的鲁棒性,以及在全尺寸现代汽车雷达数据集上的实时操作能力,在无需训练的联合检测器和跟踪器方面,其性能可与深度学习检测基准相媲美。

英文摘要

Multi-object detection and tracking from noisy point clouds remain challenging in many data-scarce radar applications. Current Bayesian trackers based on Poisson measurement models offer a training-free solution but struggle to achieve accuracy and efficiency under severe clutter, large object populations, and full-resolution Doppler point clouds. We address this with PiVoT, a fast, clutter-resilient multi-object tracker for both positional and Doppler measurements. PiVoT performs end-to-end detection and tracking of a large and time-varying number of objects without external clustering or detectors, through joint inference of object states, shapes, existence probabilities, data association, and measurement rates. Its efficiency is driven by several variational inference innovations, such as theoretically justified birth pruning, quadratic-to-linear complexity reductions for exact updates, and a computationally efficient Doppler Poisson model. Experiments show that PiVoT substantially outperforms existing Bayesian trackers in challenging scenes, while also demonstrating exceptional scalability to a thousand objects, robustness to clutter visually inseparable from objects, and real-time operation on full-scale modern automotive radar datasets, where it attains performance comparable to a deep-learning detection benchmark as a training-free joint detector and tracker.

URL PDF HTML 收藏
2604.22433 2026-07-17 cs.LG 版本更新

From physical surfaces to human-centric heat stress: LST and UTCI heat mapping reveals nonlinear effects of urban morphology

超越地表温度:可解释的空间机器学习揭示城市形态对以人类为中心的热压力的影响

Yuan Wang, Shengao Yi, Xiaojiang Li, Pengyuan Liu, Zhiwei Yang, Ronita Bardhan, Rudi Stouffs

机构 * Department of Architecture, National University of Singapore, Singapore 117566, Singapore Cambridge Centre for Advanced Research Sustainable Design Group, Department of Architecture, University of Cambridge, Cambridge, United Kingdom Department of City Regional Planning, University of Pennsylvania, Philadelphia, PA 19104, USA Urban Analytics Subject Group, Urban Studies \& Social Policy Division, University of Glasgow Laboratory for Earth Surface Processes, Ministry of Education, College of Urban Environmental Sciences, Peking University, Beijing 100871, China

AI总结 本文通过比较地表温度与通用热气候指数,揭示城市形态对人类热压力的影响,采用可解释的机器学习方法分析两者在空间分布和机制上的差异。

Comments Accepted manuscript. The final published version is available at https://doi.org/10.1016/j.scs.2026.107659

详情
AI中文摘要

热量暴露连接了建成环境与公共卫生,直接影响城市区域的宜居性和可持续性。理解热量暴露的空间异质性及其驱动因素对气候适应性城市规划至关重要。然而,大多数规划导向研究依赖于地表温度(LST),而LST是否足以代表人类热量暴露以及其与生理相关热压力的差异仍缺乏充分研究。本文采用Landsat获取的30米LST和新加坡的GPU加速1米通用热气候指数(UTCI),建立了一个综合的“建模-比较-评估”框架,系统评估两种指标的空间和机制差异。进一步,通过采用新的地理加权XGBoost(GW-XGBoost)和广义加性模型(GAM)工作流程,研究了两种指标与城市因素之间显著的非平稳和阈值型定量关系。研究结果表明,LST和UTCI的空间模式存在显著差异,以及2D和3D城市因素对这两种热指标影响的空间异质性,通过可解释的GW-XGBoost模型(LST的全局袋外R2为0.855,UTCI为0.905)得到揭示。关键的是,空间明确的SHAP解释表明,天空视因子在解释UTCI变化中起核心作用,但对LST的独立贡献相对较小,表明LST无法充分捕捉由遮荫和辐射过程决定的实际人类热压力。值得注意的是,SHAP-GAM分析表明,较高的反照率与增加的UTCI相关。这些新发现为整合生理相关的热指数以指导有针对性的热风险管理和气候适应性城市规划提供了证据。

英文摘要

Heat exposure connects the built environment and public health, directly shaping the livability and sustainability of urban areas. Understanding the spatial heterogeneity of heat exposure and its drivers is vital for climate-adaptive urban planning. However, most planning-oriented studies rely on land surface temperature (LST), and whether LST adequately represents human heat exposure and how it differs from physiologically relevant heat stress remains insufficiently examined. Here, using Landsat-retrieved 30-m LST and GPU-accelerated 1-m universal thermal climate index (UTCI) in Singapore, this study establishes a comprehensive "Modeling-Comparing-Assessing" framework to systematically evaluate the spatial and mechanistic differences between these two metrics. We further investigate their pronounced non-stationary and threshold-based relationships with urban factors using a novel geographically weighted XGBoost (GW-XGBoost) and generalized additive model (GAM) workflow. Our results reveal substantial differences in the spatial patterns of LST and UTCI, along with marked spatial heterogeneity in how 2D and 3D urban factors impact these thermal metrics, as demonstrated by explainable GW-XGBoost models (test R2 = 0.855 for LST and 0.905 for UTCI). Crucially, spatially explicit SHAP shows that sky view factor plays a central role in explaining UTCI variability but exhibits a comparatively marginal independent contribution to LST, indicating that LST inadequately captures shading-driven and radiative processes governing actual human heat stress. Moreover, SHAP-GAM analysis indicates that higher albedo is associated with increased UTCI. These findings provide model-informed planning implications for integrating physiologically relevant thermal indices to support targeted heat risk management and human-centric urban planning.

URL PDF HTML 收藏
2507.03209 2026-07-17 q-bio.QM cs.CE cs.LG q-bio.MN 版本更新

A Machine Learning Benchmarking Framework for Lipid Nanoparticle Transfection Efficiency Prediction

用于脂质纳米颗粒转染效率预测的机器学习基准框架

Asal Mehradfar, Mohammad Shahab Sepehri, Jose Miguel Hernandez-Lobato, Glen S. Kwon, Mahdi Soltanolkotabi, Salman Avestimehr, Morteza Rasoulianboroujeni

机构 * Department of Electrical and Computer Engineering, University of Southern California(电气与计算机工程系,南加州大学) University of Cambridge(剑桥大学) School of Pharmacy, University of Wisconsin-Madison(威斯康星大学麦迪逊分校药学院) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) East Tennessee State University(东田纳西州立大学)

AI总结 研究针对脂质纳米颗粒转染效率预测,提出机器学习基准框架,系统测试不同分子表示与机器学习架构,用特定数据集评估模型,显示利用显式分子子结构编码的模型准确性最高,为相关预测模型发展建立基线。

Comments Published in Communications AI & Computing (Nature Portfolio), 2026

Journal ref Commun. AI Comput. 1, 2 (2026)

详情
AI中文摘要

新的可电离脂质的发现是RNA疗法发展的主要瓶颈。机器学习模型可直接根据脂质结构预测转染效率,但缺乏严格的基准测试。本文提出了一个强大的机器学习基准框架,用于评估基于可电离脂质结构的转染预测模型。该框架系统地对不同分子表示与广泛的机器学习架构进行基准测试,支持模型泛化评估和超越标准回归指标的预测可靠性评估。通过一个精心策划的数据集表明,利用显式分子子结构编码的模型具有最高预测准确性,而一些当前基于图的模型准确性较低。该框架为脂质基RNA递送预测模型的未来发展建立了强大的基线。

英文摘要

The discovery of new ionizable lipids for efficient lipid nanoparticle (LNP)-mediated RNA delivery remains a major bottleneck in RNA therapeutics development. Recent advances demonstrate the potential of machine learning (ML) models to predict transfection efficiency directly from lipid structure, enabling high-throughput virtual screening and accelerating lead identification. However, as new models for LNP transfection prediction continue to emerge, the lack of rigorous and standardized benchmarking poses a significant risk and may undermine confidence in their reliability for discovery. Here, we present a robust ML benchmarking framework for evaluating transfection prediction models based on ionizable lipid structures. This framework systematically benchmarks diverse molecular representations paired with a broad range of ML architectures spanning traditional models, feedforward neural networks, and state-of-the-art graph-based methods. In addition, the presented framework supports assessment of model generalization and evaluates prediction reliability beyond standard regression metrics. Using a curated dataset of 1,100 unique ionizable lipid structures derived from the HeLa transfection dataset originally reported by Xu et al., we show that within this framework, models leveraging explicit molecular substructure encoding consistently achieve the highest predictive accuracy and should serve as essential baselines for the development of new, more sophisticated models. In contrast, some current graph-based models, including AGILE, Chemprop, and KPGT, tend to show comparatively lower accuracy. The presented framework provides a standardized, transparent, and comprehensive benchmarking resource that enables meaningful comparison of emerging architectures and establishes strong baselines for future development of predictive models in lipid-based RNA delivery.

URL PDF HTML 收藏
2607.13431 2026-07-16 cs.LG cs.AI cs.CL 新提交

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation

离散扩散模型:从词元化到生成的统一框架

Ye Yuan, Weien Li, Rui Song, Zeyu Li, Haochen Liu, Xiangyu Kong, Zixuan Dong, Linfeng Du, Zipeng Sun, Weixu Zhang, Jiaxin Huang, Changjiang Han, Yonghan Yang, Zichen Zhao, Xiuyuan Hu, Haolun Wu, Yankai Chen, Fengran Mo, Jikun Kang, Bowei He, Philip S. Yu, Xue Liu

机构 * McGill University(麦吉尔大学) Mila - Quebec AI Institute(米拉-魁北克人工智能研究所) University of Cambridge(剑桥大学) University of Toronto(多伦多大学) MBZUAI - Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Tsinghua University(清华大学) Rochester Institute of Technology(罗彻斯特理工学院) Salesforce(Salesforce公司) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

AI总结 研究离散扩散模型,引入统一框架从离散状态空间构建审视该模型,让现有公式成为共同设计空间实例,揭示训练、推理等方面权衡,为未来研究提供方向。

详情
AI中文摘要

离散去噪扩散模型(DDMs)最近成为离散数据自回归建模的有力替代方案,具有并行生成和迭代全局细化能力。与连续扩散不同,离散扩散模型的状态空间由离散状态空间的构建方式决定。本文引入统一概念框架,通过构建底层离散状态空间来审视离散扩散模型。在此框架下,现有公式成为共同设计空间的不同实例,还揭示了训练目标、推理算法等方面的常见权衡,为未来研究指明方向。

英文摘要

Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete data, offering parallel generation and iterative global refinement capabilities. Unlike continuous diffusion, where the state space is fixed, DDMs are fundamentally shaped by how the discrete state space is constructed: the tokenization scheme, the vocabulary topology, and domain-specific structural alphabets. This work introduces a unified conceptual framework that views discrete diffusion models through the construction of the underlying discrete state space. Within this framework, existing formulations, including transition-matrix, masking/absorbing-state, and score/ratio-based approaches, emerge as different instantiations of a common design space. The framework further exposes common design trade-offs across training objectives, inference algorithms, scaling behavior, systems optimization, and evaluation protocols, suggesting several promising directions for future research.

URL PDF HTML 收藏
2607.13107 2026-07-16 cond-mat.mtrl-sci cond-mat.str-el cs.LG 新提交

DeepCormack: Fermi surface tomography using model-based data-driven algorithms

深度科马克:使用基于模型的数据驱动算法进行费米面断层扫描

Georg F. B. Lovric, Bryn Drury, Carola-Bibiane Schönlieb, Stephen B. Dugdale, Ander Biguri

机构 * Department of Applied Mathematics and Theoretical Physics, University of Cambridge(应用数学与理论物理系,剑桥大学) H.H. Wills Physics Laboratory, University of Bristol(布里斯托大学HH.Wills物理实验室)

AI总结 研究利用电子-正电子湮灭辐射角关联重建三维双光子动量密度来研究材料费米面,提出基于数据驱动模型的深度科马克算法,通过集成深度学习模型增强传统方法,还给出合成数据方法,提高重建质量并加快采集时间。

Comments 28 pages, 17 figures

详情
AI中文摘要

通过电子-正电子湮灭辐射的角关联(ACAR)对三维双光子动量密度(TPMD)进行实验重建,是研究材料费米面的一种特别有用的方法。它不依赖低温、超高真空条件或强磁场,能研究材料的自旋分辨电子结构,但仍是一个具有挑战性的逆问题。通常要测量\(10^8\)次正电子湮灭事件以获取不同角度的TPMD的3至6个投影。标准重建方法是科马克方法(MCM)的ACAR改编版,利用晶体结构的固有对称性。但信噪比差,为费米面研究收集足够质量的数据每个样本可能需要数月。我们提出了深度科马克,这是一族基于数据驱动模型的重建算法,通过在不同阶段集成监督深度学习模型(CNN、MLP和UNet)来增强MCM。为克服缺乏大型实验训练集的问题,我们提出一种利用奇异值分解和动态模式分解生成逼真合成TPMD体积的方法,仅需通过密度泛函理论计算的单个参考动量密度。在测试数据上,深度科马克在200M计数时比MCM的重建质量提高约8.5dB PSNR,在计数减少时仍保持稳定,能显著加快采集时间。对实验数据的泛化很大程度上取决于参考动量密度的训练分布与样本的匹配程度。因此,我们建议将深度科马克与目标材料的DFT计算配对以创建特定样本的训练数据。我们提出的方法能提供更高质量的重建,或能显著加快重建速度,达到几周的量级。

英文摘要

The experimental reconstruction of the 3D two-photon momentum density (TPMD) via angular correlation of electron-positron annihilation radiation (ACAR) is a particularly useful method for studying material Fermi surfaces. It does not rely on low temperatures, UHV conditions, or strong magnetic fields, and enables the study of the spin-resolved electronic structure of materials. Yet, it remains a challenging inverse problem. Typically, 10^8 positron annihilation events are measured for 3--6 projections of the TPMD at different angles. The standard reconstruction approach is an ACAR adaptation of Cormack's method (the MCM) that leverages the inherent symmetry in the crystal's structure. However, the poor signal-to-noise ratio means collecting data of sufficient quality for Fermi surface studies can take months per sample. We present DeepCormack, a family of data-driven model-based reconstruction algorithms that augments the MCM by integrating supervised deep-learning models (CNN, MLP, and UNet) at various stages. To overcome the lack of large experimental training sets, we propose a method which leverages singular value decomposition with dynamic mode decomposition to generate realistic synthetic TPMD volumes, requiring only a single reference momentum density computed via density functional theory. On test data, DeepCormack improves reconstruction quality over MCM by about 8.5 dB PSNR at 200M counts and remains stable at reduced counts, enabling significantly faster acquisition times. Generalisation to experimental data depends strongly on how well the training distribution from the reference momentum density matches the sample. We therefore recommend pairing DeepCormack with a DFT calculation of the target material to create sample-specific training data. Our proposed method offers either much higher quality reconstructions, or enables significantly faster ones, on the order of weeks.

URL PDF HTML 收藏
2606.23531 2026-07-16 cs.RO 版本更新

BiliVLA: Scene-Aware Vision-Language-Action Model with Reinforcement Learning for Autonomous Biliary Endoscopic Navigation

BiliVLA: 基于强化学习的场景感知视觉-语言-动作模型用于自主胆道内镜导航

Jinsong Lin, Chi Kit Ng, Zhiyong Xiong, Zikang Pan, Yihan Hu, Tabassum Tamima, Ziyi Hao, Eddie Cheung, Jiewen Lai, Huxin Gao, Hongliang Ren

机构 * The Chinese University of Hong Kong(香港中文大学) The Third Affiliated Hospital of Sun Yat-sen University(中山大学附属第三医院) University of Cambridge(剑桥大学) University of California, Davis(加州大学戴维斯分校)

AI总结 提出BiliVLA框架,将胆道内镜导航建模为指令条件化的视觉运动学习问题,结合场景感知监督和两阶段训练(SFT+GRPO),在ERCP子任务中实现91.96%的动作精度和84.85%的成功率。

详情
AI中文摘要

内镜逆行胰胆管造影术(ERCP)需要在具有镜面反射、部分遮挡和频繁组织接触的狭窄单目视野中进行精确的内镜导航和稳定的胆管插管。尽管近期的机器人系统和基于视觉的辅助技术改善了操作者的人机工程学并提供了感知线索,但它们在显著的解剖变异和安全关键性视觉伪影下性能下降,阻碍了插管级手术中可靠的自主性。在此,我们提出BiliVLA,一个场景感知的视觉-语言-动作(VLA)框架,将胆道内镜导航建模为指令条件化的视觉运动学习问题。给定内镜观察和阶段特定的语言指令,BiliVLA联合预测目标类别、有根据的边界框以及连续内镜的离散三自由度(DoF)运动命令。该框架结合场景感知监督以增强语义目标一致性,以及安全感知恢复监督以在管腔壁接触下诱导保守后退行为。BiliVLA的一个关键组成部分是两阶段训练范式,将基于接地增强的监督微调(SFT)与群体相对策略优化(GRPO)相结合,显著提高了闭环导航过程中的动作可靠性和决策一致性。在三个ERCP子任务中,BiliVLA在真实世界体模实验中实现了平均动作精度91.96%和总体成功率(SR)84.85%。这些结果表明,整合语义接地、场景感知学习和奖励引导优化改善了感知-动作对齐,并实现了鲁棒的自主内镜导航。

英文摘要

Endoscopic retrograde cholangiopancreatography (ERCP) demands precise endoscopic navigation and stable biliary cannulation within a narrow monocular field characterized by specular reflections, partial occlusions, and frequent tissue contact. Although recent robotic systems and vision-based assistance techniques improve operator ergonomics and provide perceptual cues, their performance degrades under pronounced anatomical variability and safety-critical visual artifacts, which hinders reliable autonomy in cannulation-grade procedures. Here, we present BiliVLA, a scene-aware Vision-Language-Action (VLA) framework that formulates biliary endoscopic navigation as an instruction-conditioned visuomotor learning problem. Given an endoscopic observation and a stage-specific language instruction, BiliVLA jointly predicts the target category, a grounded bounding box, and a discrete three-degree-of-freedom (3-DoF) motor command for a continuum endoscope. The proposed framework incorporates scene-aware supervision to improve semantic target consistency and safety-aware recovery supervision to induce conservative retreat behaviors under luminal wall contact. A key component of BiliVLA is a two-stage training paradigm that combines grounding-enhanced supervised fine-tuning (SFT) with Group Relative Policy Optimization (GRPO), thereby improving action reliability and decision consistency during closed-loop navigation. Across three ERCP subtasks, BiliVLA achieves the best overall performance in physical phantom experiments, with a total mIoU of 0.9625, an overall action precision of 91.96\%, and an overall success rate (SR) of 84.85\%. These results indicate that integrating semantic grounding, scene-aware learning, and reward-guided optimization strengthens perception--action alignment and enables more robust autonomous biliary endoscopic navigation.

URL PDF HTML 收藏
2606.06223 2026-07-16 cs.AI 版本更新

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents

从奖励黑客激活到智能体风险状态:LLM智能体中的上下文校准机制监控

Patrick Wilhelm, Odej Kao

机构 * University of Cambridge(剑桥大学)

AI总结 本研究通过分析ReAct风格智能体在Gameable ALFWorld和WebShop环境中的奖励黑客行为,提出结合激活状态、熵和决策上下文的上下文校准监控方法,以更准确评估智能体风险。

详情
AI中文摘要

语言模型智能体通过观察、推理和动作选择的重复循环运行,使得安全监控依赖于内部模型状态和环境上下文。我们研究了在Gameable ALFWorld和WebShop环境中运行的ReAct风格智能体中的奖励黑客监控。智能体配备了基于激活的奖励黑客分数、token级熵和决策上下文特征。我们发现,在《奖励黑客学校》数据集上微调的适配器可以将奖励黑客倾向转移到智能体动作选择中,尤其是当环境暴露代理奖励可供性时。然而,缓解此类行为不能仅依赖激活动态。高奖励黑客激活识别出潜在策略状态,但并不一定意味着立即的利用动作。在下一步预测任务中,熵和上下文校准的内部特征比单独的奖励黑客激活提高了风险估计。激活方向引导进一步减少了选定混合适配器设置中的代理利用行为。总体而言,我们的结果支持智能体的上下文校准内部监控:奖励黑客激活识别潜在策略状态,而熵和决策上下文有助于确定该状态何时变为风险动作。

英文摘要

Language-model agents act through repeated cycles of observation, reasoning, and action selection, making safety monitoring depend on both internal model state and environment context. We study reward-hacking monitors in ReAct-style agents acting in Gameable ALFWorld and WebShop. Agents are instrumented with activation-based reward-hack scores, token-level entropy, and decision-context features. We find that adapters fine-tuned on \textit{School-of-Reward-Hacks} dataset can transfer reward-hack tendencies into agentic action selection, especially when the environment exposes proxy-reward affordances. However, mitigating such behavior cannot rely on activation dynamics alone. High reward-hack activation identifies a latent policy state, but does not necessarily imply an immediate exploit action. Across next-step prediction tasks, entropy and context-calibrated internal features improve risk estimation over reward-hack activation alone. Activation-direction steering further reduces proxy-exploit behavior in selected mixed-adapter regimes. Overall, our results support context-calibrated internal monitoring for agents: reward-hack activation identifies a latent policy state, while entropy and decision context help determine when that state becomes risky action.

URL PDF HTML 收藏
2607.12145 2026-07-15 stat.ML cs.LG 新提交

Falsifying Causal Graphs With Outlier Events

用异常事件证伪因果图

William Roy Orchard, Philipp M. Faller, Dominik Janzing

机构 * University of Cambridge(剑桥大学) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院) Amazon Research(亚马逊研究)

AI总结 研究如何在无真实因果关系时评估因果图,提出基于异常事件传播证伪候选因果图的方法,利用弱异常很少导致强异常原则,给出相关统计检验,有控制误报等功效,可单样本运行。

Comments Accepted at the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026)

详情
AI中文摘要

真实的因果关系很少为人所知,从数据中推断因果图很困难。一个基本挑战是在没有地面真值的情况下,如何评估给定的因果图是否良好。我们提出基于候选因果图能否解释异常事件的传播来对其进行证伪。我们的方法利用一个关键原则:弱异常很少导致强异常。虽然该原则此前已用于根本原因分析以识别根本原因,但我们将其反过来用于证伪其隐含的异常传播与数据不一致的候选因果图。为此,我们针对候选图是真实因果图这一假设提出了首个统计检验,并表明它们具有误报控制、对错误因果图的功效保证,且可在单个异常样本下运行。

英文摘要

True causal relationships are rarely known, and inferring causal graphs from data is hard. A fundamental challenge is how to assess whether a given causal graph is good in the absence of a ground truth. We propose falsifying candidate causal graphs based on whether they can explain the propagation of an outlier event. Our approach leverages a key principle: weak outliers rarely cause strong ones. While this principle has previously been used in root cause analysis to identify root causes without prior knowledge of the graph, we turn it on its head and use it to falsify candidate causal graphs whose implied outlier propagation is inconsistent with the data. To this end, we present the first statistical tests for the hypothesis that a candidate graph is the true causal graph, and show they have false positive control, power guarantees against incorrect causal graphs, and can operate with a single outlier sample.

URL PDF HTML 收藏
2209.14125 2026-07-15 stat.ML cs.LG 版本更新

Spectral Diffusion Processes

谱扩散过程

Angus Phillips, Thomas Seror, Michael Hutchinson, Valentin De Bortoli, Arnaud Doucet, Emile Mathieu

机构 * University of Oxford(牛津大学) ENS, CNRS, PSL University Paris(巴黎高等师范学院、国家科学研究中心、巴黎大学) University of Cambridge(剑桥大学)

AI总结 研究将扩散模型应用于函数空间上的随机过程,通过谱表示分离随机部分与时空结构,用有限维扩散模型建模谱系数,经截断确保模型有效性,投影回原空间对应相关噪声扩散模型,还展示了方法在多模态数据建模及条件采样上的有效性。

Comments This version (v3) extends the previous workshop version (v2) with conditional sampling and theoretical results. Work carried out in 2022/23. V2 appeared in Score-based Methods Workshop at the 36th Conference on Neural Information Processing Systems (NeurIPS 2022)

详情
AI中文摘要

扩散模型已被证明是在有限维空间上对概率分布进行建模的灵活有效框架。然而,许多物理建模问题如时间序列自然是在函数空间上描述的。本文将扩散模型应用于此类随机过程。为此考虑通过核获得的数据的谱表示,将过程的随机部分与其时空结构分离。过程的随机性完全编码在谱系数中,对其截断并用标准有限维扩散模型建模。通过在谱域截断表示确保所得模型定义有效随机过程,自然满足一致性和可交换性标准。将谱扩散模型投影回原始输入空间表明,对于任何给定边际,该方法对应具有相关噪声的扩散模型,其协方差矩阵由核明确给出。通过相对于上下文集摊销模型,证明了该方法对各种多模态数据集建模以及条件采样的有效性。

英文摘要

Diffusion models have proven to be a flexible and effective framework for modelling probability distributions on finite-dimensional spaces. However, many physical modelling problems such as time series are naturally described over function spaces. In this work we apply diffusion models to such stochastic processes. To do so we consider a spectral representation of the data, obtained using a kernel, thereby dissociating the stochastic part of the processes from their space-time structure. As a result, the stochasticity of the processes is entirely encoded in the spectral coefficients, which we truncate and model using standard finite-dimensional diffusion models. By truncating the representation in the spectral domain we ensure our resulting model defines valid stochastic processes, thereby naturally satisfying consistency and exchangeability criteria. Projecting our spectral diffusion models back to the original input space, we show that for any given marginals our approach corresponds to a diffusion model with correlated noise, with explicit covariance matrix given by the kernel. We demonstrate our method's effectiveness for modelling various multimodal datasets as well as conditional sampling by amortising our models with respect to a context set.

URL PDF HTML 收藏
2607.11570 2026-07-14 cs.RO cs.HC 新提交

ERR@HRI 3.0 Challenge: Multimodal Detection of Errors and Anticipation in Human-Robot Interactions

ERR@HRI 3.0 挑战赛:人机交互中错误与预期的多模态检测

Maria Teresa Parreira, Micol Spitale, Maia Stiber, Shiye Cao, Amama Mahmood, Chien-Ming Huang, Hatice Gunes, Wendy Ju

机构 * Cornell University(康奈尔大学) Microsoft Research(微软研究院) Johns Hopkins University(约翰霍普金斯大学) University of Cambridge(剑桥大学)

AI总结 ERR@HRI 3.0 挑战赛提供自然场景视频数据集,供研究者开发多模态机器学习模型检测人机交互错误与预期,三个团队提交的有效模型超卷积神经网络基线,为构建相关检测系统提供了数据、任务、基线及结果等参考。

详情
AI中文摘要

随着机器人越来越融入人类环境,其检测和应对错误的能力对于维持用户信任和交互质量至关重要。尽管机器学习的最新进展提高了错误检测能力,但大多数方法仅限于特定上下文、受控设置或预提取特征,限制了其在现实世界条件下的通用性和适用性。为应对这一挑战,ERR@HRI 3.0 挑战赛为研究人员提供了两个互补数据集,以实现人机交互中错误检测和预防方法的端到端创新。挑战赛提供了来自自然场景的原始、非匿名视频数据:(1)旁观者影响检测(BAD)数据集,包含 45 名参与者对机器人和人类失败场景的自发反应的网络摄像头记录;(2)坏主意数据集,包含 29 名参与者在预测失败发生前的行动结果时的预期面部反应。两个数据集均通过众包收集,捕捉了现实世界条件的固有变异性。这种自然变异性虽然具有挑战性,但为开发强大的错误检测系统提供了一个真实的测试平台。参与者开发了用于旁观者反应检测(赛道 1)和预期结果预测(赛道 2)的多模态机器学习模型,以及一个可选的跨数据集泛化赛道(赛道 3)。三个团队提交了有效模型,所有模型均超过了我们的卷积神经网络基线。本文描述了 ERR@HRI 3.0 的数据集、任务、基线和结果,并讨论了对构建用于人机交互的通用、上下文感知和预期错误检测系统的意义。

英文摘要

As robots become increasingly integrated into human environments, their ability to detect and respond to errors remains critical for maintaining user trust and interaction quality. While recent advances in machine learning have improved error detection capabilities, most approaches are limited to specific contexts, controlled settings, or pre-extracted features, limiting their generalizability and applicability to real-world conditions. To address this challenge, the third edition of the ERR@HRI Challenge (ERR@HRI 3.0) provided researchers with two complementary datasets that enable end-to-end innovation in methods for both detecting and preventing errors in human-robot interaction. The challenge offered raw, non-anonymized video data from naturalistic settings: (1) the Bystander Affect Detection (BAD) dataset, containing webcam recordings of 45 participants' spontaneous reactions to robot and human failure scenarios; and (2) the Bad Idea dataset, featuring 29 participants' anticipatory facial responses while predicting action outcomes before failures occur. Both datasets were collected via crowdsourcing, capturing the inherent variability of real-world conditions. This naturalistic variability, while challenging, provides an authentic testbed for developing robust error detection systems. Participants developed multimodal machine learning models for bystander reaction detection (Track 1) and anticipatory outcome prediction (Track 2), with an optional cross-dataset generalization track (Track 3). Three teams submitted valid models, all of which surpassed our convolutional neural network baselines. This paper describes the datasets, tasks, baselines, and results of ERR@HRI 3.0, and discusses implications for building generalizable, context-aware, and anticipatory error detection systems for human-robot interaction.

URL PDF HTML 收藏
2607.11312 2026-07-14 cs.CV 新提交

SLVMBench: Skill Learning from Video Memory

SLVMBench:从视频记忆中学习技能

Yudong Yang, Guangzhi Sun, Yixuan Li, Chao Zhang

机构 * Tsinghua University(清华大学) University of Cambridge(剑桥大学)

AI总结 SLVMBench是首个用于评估视频大语言模型从长视频记忆学习技能并应用于实时任务能力的基准测试,通过特定视频流和人工标注进行测试,发现现有视频LLMs在此方面有局限。

详情
AI中文摘要

我们介绍了从视频记忆中学习技能(SLVMBench),这是首个联合评估视频大语言模型(video-LLMs)能否从长视频记忆中学习技能并将其应用于实时任务的基准测试。SLVMBench为模型提供2至3小时的视频流,其中包含嵌入在任意无关视频流中的教程视频,类似现实世界中的人类学习实践。要求Video-LLMs应用所学技能回答关于正在播放视频的实时问题。与强调被动理解的长视频理解基准测试和依赖简短即时演示的技能学习基准测试不同,SLVMBench测试记忆和提取程序知识以及将其转移到实时任务的完整流程。此外,严格的人工标注具有亚秒级时间校准、消除常识猜测的人工设计问题以及整理的教程,以确保所需技能的覆盖范围。对现有先进的专有和开源视频LLMs的评估表明,视频LLMs在从视频中学习和应用技能知识方面存在很大困难。而且,当技能知识置于长视频记忆中时,性能会显著下降。这些结果揭示了现有视频LLMs的一个关键局限性,并将SLVMBench定位为研究从长上下文视频记忆中进行实时技能获取和应用的首个基准测试。

英文摘要

We introduce Skill Learning from Video Memory (SLVMBench), the first benchmark that jointly evaluates whether video large language models (video-LLMs) can learn skills from long video memory and apply them to real-time tasks. SLVMBench presents models with 2-3 hour video streams that contain a tutorial video embedded in a stream of arbitrary irrelevant videos, resembling real-world human learning practices. Video-LLMs are asked to apply the acquired skill to answer real-time questions about an ongoing video. Unlike long-video understanding benchmarks that emphasize passive comprehension and skill-learning benchmarks that rely on short, immediate demonstrations, SLVMBench tests the full pipeline of memorizing and extracting procedural knowledge, as well as transferring it to real-time tasks. Moreover, rigorous human annotations feature sub-second-level temporal calibration, manually engineered questions eliminating common-sense guessing, and collated tutorials to ensure coverage of the required skills. Evaluations on state-of-the-art proprietary and open-source video LLMs show that video-LLMs struggle substantially with learning and applying skill knowledge from videos. Moreover, performance degrades markedly when the skill knowledge is placed within a long video memory. These results reveal a key limitation of existing video LLMs and position SLVMBench as the first benchmark for studying real-time skill acquisition and application from long-context video memory.

URL PDF HTML 收藏
2607.10898 2026-07-14 cs.CV cs.NA math.NA 新提交

Design Choices in Splitting-Based Self-Supervised Sparse-View CT Reconstruction

基于分割的自监督稀疏视图CT重建中的设计选择

Nadja Gruber, Lukas Neumann, Ander Biguri, Gyeongha Hwang, Markus Haltmeier, Johannes Schwab

机构 * University of Innsbruck(因斯布鲁克大学) Institute of Basic Sciences in Engineering Science, University of Innsbruck(因斯布鲁克大学工程科学基础科学研究所) University of Cambridge(剑桥大学) Yeungnam University(岭南大学) University of Applied Sciences Kufstein(库夫施泰因应用科学大学)

AI总结 研究基于分割的自监督稀疏视图CT重建中关键设计选择的影响,引入统一框架分解重建为三组件,通过实验表明最优划分策略依赖噪声结构,多划分分割表现优,结果为方法设计提供指南并凸显独立性假设局限。

Comments 12 pages (main manuscript), 9 pages (supplementary material and figures)

详情
AI中文摘要

自监督数据分割已成为稀疏视图CT重建的一种有前景的范式,可从不完整测量中训练,无需完全采样的真实数据。然而,关键设计选择(包括划分策略、预处理和推理)的影响仍未得到充分理解。本文引入统一框架,将基于分割的重建分解为这三个组件,实现对现有方法及两个增量扩展(多划分分割和替代推理策略)的可控比较。在模拟LoDoPaB-CT数据上的实验以及在真实2DeteCT数据集上的验证表明,最优划分策略强烈依赖于测量噪声结构。基于格点的分割在独立噪声下表现良好,而角度掩蔽在相关噪声和真实测量数据下更稳健。多划分分割在多种设置下始终优于纯投影方式分割。互补的感知和结构指标揭示了掩蔽策略之间的差异,这些差异仅从PSNR和SSIM中不太明显。这些结果为设计自监督稀疏视图CT重建方法提供了实用指南,并突出了现实成像环境中常见独立性假设的局限性。

英文摘要

Self-supervised data splitting has emerged as a promising paradigm for sparse-view CT reconstruction, enabling training from incomplete measurements without fully sampled ground truth. However, the influence of key design choices, including partitioning strategy, preprocessing, and inference, remains insufficiently understood. In this work, we introduce a unified framework that decomposes splitting-based reconstruction into these three components, enabling controlled comparison of existing methods and two incremental extensions: multi-partition splitting and an alternative inference strategy. Experiments on simulated LoDoPaB-CT data under independent and correlated noise, together with validation on the real-world 2DeteCT dataset, show that the optimal partitioning strategy strongly depends on the measurement noise structure. Lattice-based splitting performs favorably under independent noise, whereas angular masking is more robust under correlated noise and real measured data. Multi-partition splitting consistently improves over pure projection-wise splitting in several settings. Complementary perceptual and structural metrics, including LPIPS and HaarPSI, reveal differences between masking strategies that are less apparent from PSNR and SSIM alone. These results provide practical guidelines for designing self-supervised sparse-view CT reconstruction methods and highlight the limitations of common independence assumptions in realistic imaging environments.

URL PDF HTML 收藏
2606.01172 2026-07-14 cs.LG stat.ME stat.ML 版本更新

Revisiting Neural Processes via Fourier Transform and Volterra Series

通过傅里叶变换和Volterra级数重新审视神经过程

Peiman Mohseni, Nick Duffield, Raymond K. W. Wong

机构 * University of Cambridge(剑桥大学)

AI总结 本文利用Volterra展开和集合傅里叶卷积,提出了两种新的条件神经过程模型,解决了现有平移等变神经过程在可解释性和计算效率上的局限性。

详情
AI中文摘要

从有限的、不规则采样的测量中建模未知的潜在函数是科学和工程中的一个反复出现的挑战。神经过程(NPs)是一类概率函数模型,是有前景的解决方案——尤其是当赋予领域特定的对称性(如平移等变性)时,这提高了样本效率和泛化能力。然而,现有的平移等变NPs面临两个局限性:(i)它们堆叠带有非线性的通用组件,模糊了诱导的函数类并限制了可解释性;(ii)卷积设计依赖于具有局部感受野的核,并需要密集的均匀输入网格,而基于注意力的方法避免了这些问题,但随观测数量呈二次方缩放。我们通过两个贡献解决了这两个问题。首先,利用Volterra展开,我们将连续平移等变算子表征为高阶卷积的和,实现了分析透明性,同时允许通过一阶卷积进行高效近似。其次,我们引入了集合傅里叶卷积(SFConvs),这是一种频域参数化方法,直接在不规则采样点上操作,实现近似全局感受野,并在观测数量上线性缩放。基于这些思想,我们提出了两种条件神经过程(CNPs):SFConvCNPs,它堆叠带有非线性的SFConv块,以及SFVConvCNPs,它整合了Volterra公式。在合成和真实世界数据集上的实验证明了我们的方法相对于最先进基线的有效性。

英文摘要

Modeling unknown latent functions from finite, irregularly sampled measurements is a recurring challenge across science and engineering. Neural processes (NPs), a family of probabilistic functional models, are promising solutions -- especially when endowed with domain-specific symmetries like translation equivariance, which improve sample efficiency and generalization. Yet existing translation-equivariant NPs face two limitations: (i) they stack generic components with non-linearities, obscuring the induced function class and limiting interpretability; and (ii) convolutional designs are limited by local receptive fields and the need to embed inputs onto a dense uniform grid, while attention-based alternatives lift these restrictions at quadratic cost in the number of observations. We address both with two contributions. First, using the Volterra expansion, we approximate continuous translation-equivariant operators by sums of higher-order convolutions, yielding analytical transparency while admitting efficient evaluation via first-order convolutions. Second, we introduce set Fourier convolutions (SFConvs), a frequency-domain parameterization that operates directly on irregularly sampled points, achieves approximately global receptive fields, and scales linearly in the number of observations. Building on these ideas, we propose two conditional NPs (CNPs): SFConvCNPs, which stack SFConv blocks with non-linearities, and SFVConvCNPs, which integrate the Volterra formulation. Experiments on synthetic and real-world datasets demonstrate our methods' efficacy against state-of-the-art baselines.

URL PDF HTML 收藏
2604.13213 2026-07-14 stat.ML cs.LG math.OC physics.chem-ph 版本更新

Rare Event Analysis via Stochastic Optimal Control

基于随机最优控制的稀有事件分析

Yuanqi Du, Jiajun He, Dinghuai Zhang, Eric Vanden-Eijnden, Carles Domingo-Enrich

机构 * Microsoft Research New England(微软研究院新英格兰分部) Cornell University(康奈尔大学) University of Cambridge(剑桥大学) Courant Institute of Mathematical Sciences, NYU(纽约大学Courant数学科学研究所)

AI总结 提出将稀有事件分析中的committor函数估计转化为随机最优控制问题,通过反馈控制引导轨迹采样,并开发两种损失函数及处理亚稳态的方法,在基准系统上获得更准确的结果。

详情
AI中文摘要

稀有事件,如生物分子的构象变化、相变和化学反应,是许多物理系统行为的关键,但由于无偏模拟很少产生这些事件,因此计算研究极其困难。过渡路径理论(TPT)为分析此类事件提供了严格的统计框架:它表征了两个指定亚稳态(反应物和产物)之间的反应轨迹集合,其核心对象——committor函数(给出系统下一步到达产物而非反应物的概率)——编码了所有基本的动力学和热力学信息。我们引入了一个框架,将committor估计转化为随机最优控制(SOC)问题。在此公式中,committor定义了一个反馈控制(与其对数梯度成正比),该控制主动引导轨迹朝向反应区域,从而实现对反应路径的高效采样。为了解决由此产生的命中时间控制问题,我们开发了两个互补的目标:直接反向传播损失和基于原理的离策略值匹配损失,并为其建立了一阶最优性保证。我们进一步通过引入一种替代采样过程来解决亚稳态问题(该问题可能使受控轨迹陷入中间势阱),该过程在降低有效能垒的同时保持反应电流。在基准系统上,该框架比现有方法产生了显著更准确的committor估计、反应速率和平衡常数。

英文摘要

Rare events such as conformational changes in biomolecules, phase transitions, and chemical reactions are central to the behavior of many physical systems, yet they are extremely difficult to study computationally because unbiased simulations seldom produce them. Transition Path Theory (TPT) provides a rigorous statistical framework for analyzing such events: it characterizes the ensemble of reactive trajectories between two designated metastable states (reactant and product), and its central object--the committor function, which gives the probability that the system will next reach the product rather than the reactant--encodes all essential kinetic and thermodynamic information. We introduce a framework that casts committor estimation as a stochastic optimal control (SOC) problem. In this formulation the committor defines a feedback control--proportional to the gradient of its logarithm--that actively steers trajectories toward the reactive region, thereby enabling efficient sampling of reactive paths. To solve the resulting hitting-time control problem we develop two complementary objectives: a direct backpropagation loss and a principled off-policy Value Matching loss, for which we establish first-order optimality guarantees. We further address metastability, which can trap controlled trajectories in intermediate basins, by introducing an alternative sampling process that preserves the reactive current while lowering effective energy barriers. On benchmark systems, the framework yields markedly more accurate committor estimates, reaction rates, and equilibrium constants than existing methods.

URL PDF HTML 收藏
2510.11503 2026-07-14 q-bio.NC cs.AI cs.GT 版本更新

People use fast and flat simulation to reason about new games

人们使用快速且扁平的模拟来对新游戏进行推理

Katherine M. Collins, Cedegao E. Zhang, Lionel Wong, Mauricio Barba da Costa, Graham Todd, Adrian Weller, Samuel J. Cheyette, Thomas L. Griffiths, Joshua B. Tenenbaum

机构 * Massachusetts Institute of Technology(麻省理工学院) Princeton University(普林斯顿大学) University of Cambridge(剑桥大学) Stanford University(斯坦福大学) New York University(纽约大学) The Alan Turing Institute(艾伦·图灵研究所)

AI总结 研究人们对新游戏的推理,通过超千名参与者和121种新棋盘游戏的研究,发现人们玩新游戏或评估时具系统性和适应性理性,用“直观玩家”模型解释,为新问题应对及类人AI系统设计提供见解。

详情
AI中文摘要

长期以来,游戏一直是研究自然和人工智能中规划与推理的缩影,常聚焦于专家级甚至超人玩法。但现实生活促使人类智能面临新挑战,即灵活应对从未想过的决策问题。本文通过对超1000名参与者和121种两人策略棋盘游戏(几乎全是参与者陌生的)进行一系列大规模行为研究,发现人们首次玩游戏或在玩之前评估游戏(如公平性或趣味性)时是系统且适应性理性的。我们用“直观玩家”这一计算认知模型解释这些能力,该模型基于快速且扁平(深度受限)的目标导向概率模拟机制。我们的工作为人们遇到新问题时如何快速评估、行动和提建议提供了新见解,还可为设计更灵活、更像人类的人工智能系统提供参考,这类系统不仅能决定如何解决新任务,还能判断任务是否值得思考。

英文摘要

Games have long been a microcosm for studying planning and reasoning in both natural and artificial intelligence (AI), often focusing on expert-level or even super-human play. But real life also pushes human intelligence along a different frontier, requiring people to flexibly navigate decision-making problems that they have never thought about before. Here, we use novice gameplay to study how people reason about new problem settings. Through a series of large-scale behavioral studies with over 1000 participants and 121 two-player strategic board games (almost all novel to our participants), we show that people are systematic and adaptively rational in how they play a game for the first time, or evaluate a game (e.g., how fair or how fun it is likely to be) before they have played it even once. We explain these capacities via a computational cognitive model that we call the 'Intuitive Gamer', a model based on mechanisms of fast and flat (depth-limited) goal-directed probabilistic simulation. Our work offers new insights into how people rapidly evaluate, act, and make suggestions when encountering novel problems, and could inform the design of more flexible and human-like AI systems that can determine not just how to solve new tasks, but whether a task is worth thinking about at all.

URL PDF HTML 收藏
2607.04261 2026-07-13 cs.AI 新提交

Shortcut Learning in Legal Judgment Prediction: Empirical Evidence from the UK Employment Tribunal

法律判决预测中的捷径学习:来自英国就业法庭的实证证据

Joe Watson, Joana Ribeiro de Faria, Marcus Tomalin, Måns Magnusson, Huiyuan Xie, Hao Tian Yeung, Christine Carter, Jonathan Rutherford, Felix Steffek

机构 * Faculty of Law, University of Cambridge(剑桥大学法学院) The Psychometrics Centre, Cambridge Judge Business School, University of Cambridge(剑桥大学心理学测量中心,剑桥Judge商学院) Faculty of English, University of Cambridge(剑桥大学英语学院) Department of Statistics, Uppsala University(乌普萨拉大学统计系) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Department of Engineering, University of Cambridge(剑桥大学工程系)

AI总结 研究英国就业法庭判决中索赔结果预测的捷径学习,用33158条索赔语料库预测结果,发现基于事后司法文本训练的法律判决预测系统性能或被夸大,去除泄漏特征后模型仍能提取有用信号。

详情
AI中文摘要

当前法律判决预测依赖事后司法材料,易进行回顾性分类而非真正预测。本文通过研究英国就业法庭索赔级结果预测实证调查捷径学习。虽预测性能指标看似良好,但基于事后司法文本训练的系统性能可能受材料性质影响。分层测试数据发现泄漏线索影响性能,去除泄漏特征训练的模型表现良好。这表明性能可能被语言假象夸大,但并非致命,去除假象后模型仍能提取有用信号。

英文摘要

Current Legal Judgment Prediction (LJP) is constrained by its reliance on post-hoc judicial materials, increasing the likelihood that models perform retrospective classification rather than true forecasting. This paper empirically investigates shortcut learning in this context by studying claim-level outcome prediction in UK Employment Tribunal (UKET) decisions. Using a corpus of 33,158 individual claims, we predict outcomes from claim texts and LLM-extracted case summaries, evaluating models ranging from interpretable TF-IDF-based classifiers to black-box LLMs. While headline predictive performance figures appear strong, we demonstrate that such performance in LJP systems trained on post-hoc judicial text can be driven by the retrospective nature of the source material. Stratifying the test data by human judgments of leakage reveals that performance increases where outcome-revealing cues are embedded in the narrative. Moreover, a model trained on just the 4% of features identified as leakage achieves high performance, outperforming human experts. These findings substantiate concerns that LJP performance may be exaggerated by linguistic artefacts. Yet this vulnerability is not fatal to the research agenda. Instead, post-hoc judgments might be treated as potentially contaminated texts, requiring active auditing. Retraining models after masking leakage features results in only a negligible reduction in Macro-F1. Hence, while models will opportunistically exploit shortcuts when available, they remain capable of extracting useful predictive signals when these artefacts are removed.

URL PDF HTML 收藏
2606.24477 2026-07-10 cs.CV cs.AI cs.SD 新提交

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

video-SALMONN-R$^3$: 学习重看、重问和重答以实现高效视频理解

Yixuan Li, Guangzhi Sun, Yudong Yang, Chao Zhang

机构 * Tsinghua University(清华大学) ByteDance(字节跳动) University of Cambridge(剑桥大学)

AI总结 提出video-SALMONN-R$^3$,首个通过强化学习实现重看机制的视频大语言模型,无需链式思维冷启动,采用重答和重问策略提升视频问答效率与准确性。

详情
AI中文摘要

视频大语言模型通常受限于计算和内存预算,导致使用降低的帧率和空间分辨率,可能错过问答所需的关键信息。一种实用且高效的解决方案是两阶段范式:首先进行粗粒度视频理解以定位相关片段,然后以更高的时间或空间保真度重看这些片段。本文提出video-SALMONN-R$^3$,这是首个通过强化学习实现重看机制且不依赖链式思维冷启动的端到端视频大语言模型。该设计消除了昂贵的链式思维数据标注需求,并避免了基于链式思维的有监督微调,后者可能损害预训练的视频理解能力。为解决重看引发的推理优先行为与预训练视频大语言模型回答优先倾向之间的不匹配,我们提出重答策略:模型在首次观看时直接给出答案,重看后对其进行修正。最后,为提升重看过程中的问题遵循度,我们提出重问机制,在重新访问定位片段时重新注入查询。实验结果表明,video-SALMONN-R$^3$在显著降低计算成本的同时,持续优于基础模型和问答有监督微调基线,并超越先前基于重看的方法。代码、模型和数据将在论文被接收后公开。

英文摘要

Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to miss critical information for question answering (QA). A practical and efficient solution is a two-stage paradigm: first perform coarse video understanding to localize relevant segments, and then re-watch these segments at higher temporal or spatial fidelity. In this paper, we present video-SALMONN-R$^3$, the first end-to-end video-LLM that enables re-watch through reinforcement learning without relying on chain-of-thought (CoT) cold-start. This design removes the need for costly CoT data annotations and avoids CoT-based supervised fine-tuning (SFT), which can otherwise degrade the pretrained video understanding abilities. To address the mismatch between the reasoning-first behavior induced by re-watch and the answer-first tendency of pretrained video-LLMs, we propose a re-answer strategy, in which the model first produces a direct answer in the first watch and then refines it after re-watching. Finally, to improve question adherence during re-watching, we propose a re-ask mechanism that re-injects the query when revisiting localized segments. Experimental results show that video-SALMONN-R$^3$ consistently outperforms both the base model and the QA-SFT baseline, while surpassing prior re-watch-based approaches with significantly lower computational cost. Code, models, and data will be publicly released upon acceptance.

URL PDF HTML 收藏
2602.10155 2026-07-10 eess.IV cs.CV 版本更新

Data-Driven Registration and Modeling of Brain Deformation for Image-Guided Neurosurgery: A Systematic Review

数据驱动的图像配准与变形建模在图像引导神经外科中的应用:系统综述

Tiago Assis, Colin P. Galvin, Joshua P. Castillo, Nazim Haouchine, Marta Kersten-Oertel, Zeyu Gao, Mireia Crispin-Ortuzar, Stephen J. Price, Thomas Santarius, Yangming Ou, Sarah Frisken, Nuno C. Garcia, Alexandra J. Golby, Reuben Dorent, Ines P. Machado

机构 * LASIGE, Faculty of Sciences, University of Lisbon(里斯本大学科学学院LASIGE) Department of Neurosurgery and Department of Radiology, Brigham and Women's Hospital, Harvard Medical School(哈佛医学院布里洛妇女医院神经外科与放射科) Gina Cody School of Engineering and Computer Science, Concordia University(康科迪亚大学工程与计算机科学学院) Cancer Research UK Cambridge Centre, University of Cambridge(剑桥大学癌症研究英国中心) Department of Oncology, University of Cambridge(剑桥大学肿瘤科) Department of Clinical Neurosciences, University of Cambridge(剑桥大学临床神经科学系) Computational Health Informatics Program (CHIP) and Department of Radiology, Boston Children's Hospital, Harvard Medical School(哈佛医学院波士顿儿童医院计算健康信息学计划与放射科) Sorbonne Université, Institut du Cerveau - Paris Brain Institute - ICM(索邦大学巴黎脑研究所-ICM)

AI总结 系统综述2020-2025年间基于学习的脑变形补偿方法,包括深度学习配准、变形场回归、多模态对齐、切除感知架构及混合模型,指出当前方法在鲁棒性、标准化基准、可解释性和临床部署方面的局限,并展望未来研究方向。

Comments 41 pages, 7 figures, 9 tables. Accepted at Medical Image Analysis

详情
AI中文摘要

精确补偿脑变形对于可靠的图像引导神经外科手术至关重要。手术操作和肿瘤切除会引起组织运动,导致术前规划图像与术中解剖结构错位。本综述考察了2020年至2025年间开发的用于建模和校正脑变形的方法,特别关注基于学习的方法。在PubMed、IEEE Xplore、Scopus和Web of Science中进行了全面的文献检索,并采用预定义的纳入和排除标准,聚焦于应用于神经外科成像中脑变形补偿的计算方法,最终有46项研究符合标准。我们提供了方法策略的统一分析,包括基于深度学习的图像配准、直接变形场回归、合成驱动的多模态对齐、处理缺失对应关系的切除感知架构,以及整合生物力学先验的混合模型。我们还考察了数据集利用情况、报告的评估指标、验证协议,以及各研究如何评估不确定性和泛化能力。尽管基于学习的变形模型展现出有前景的性能和计算效率,但当前方法在分布外鲁棒性、标准化基准测试、可解释性和临床部署准备方面存在局限性。我们的综述强调了这些差距,并概述了未来研究的机会,旨在为神经外科引导实现更鲁棒、更可泛化且临床可转化的变形补偿解决方案。通过组织近期进展并批判性评估实践,本工作为从事数据驱动脑变形建模与校正的研究人员和临床医生提供了全面参考。

英文摘要

Accurate compensation of brain deformation is critical for reliable image-guided neurosurgery. Surgical manipulation and tumor resection induce tissue motion, causing preoperative planning images to become misaligned with the intraoperative anatomy. In this systematic review, we examine data-driven methods developed between 2020 and 2025 for brain deformation registration and modeling, with a particular focus on learning-based approaches. A comprehensive literature search was conducted in PubMed, IEEE Xplore, Scopus, and Web of Science using predefined inclusion and exclusion criteria for computational methods addressing brain deformation in neurosurgical imaging, resulting in 46 eligible studies. We provide a unified analysis of methodological strategies, including deep learning-based image registration, direct deformation field regression, synthesis-driven multimodal alignment, resection-aware architectures for handling missing correspondences, and hybrid models integrating biomechanical priors. We also examine dataset utilization, evaluation metrics, validation protocols, and the assessment of uncertainty and generalization across studies. While learning-based methods demonstrate promising accuracy and computational efficiency, current approaches remain limited by out-of-distribution robustness, standardized benchmarking, interpretability, and readiness for clinical deployment. Our review highlights these gaps and outlines future directions toward more robust, generalizable, and clinically translatable solutions for neurosurgical guidance. By organizing recent advances and critically assessing evaluation practices, this work provides a comprehensive reference for researchers and clinicians working on data-driven registration and modeling of brain deformation.

URL PDF HTML 收藏
2607.06583 2026-07-09 q-bio.QM cs.LG 新提交

Trajectory Inference of Human Aging from Cross-Sectional DNA Methylation Data

从横断面DNA甲基化数据推断人类衰老轨迹

Chandan Gupta, Syed Haider, Pietro Liò

机构 * Independent Researcher(独立研究者) The Institute of Cancer Research(癌症研究所在) University of Cambridge(剑桥大学)

AI总结 该研究从横断面DNA甲基化数据推断人类衰老轨迹,通过年龄正则化变分自编码器和正则化不平衡最优传输构建两阶段计算流程,在大规模数据集上验证,能模拟分子衰老、揭示增长激增并重建衰老原型。

Comments 13 pages, 6 figures

详情
AI中文摘要

DNA甲基化是生物衰老最可靠的分子生物标志物之一。传统表观遗传时钟将衰老视为静态回归任务,只能输出单一分数。为重建连续动态变化,本文将人类表观遗传衰老构建为跨离散年龄快照的轨迹推断问题。引入两阶段计算流程:首先用年龄正则化变分自编码器将高维CpG图谱映射到按时间顺序排列的潜在流形上;其次通过正则化不平衡最优传输模拟潜在空间中的连续运动。在大规模80年泛组织数据集上评估,模型展示出强大的分布插值能力,揭示了晚年显著的增长激增,还重建并验证了不同的生物衰老原型,为模拟人类分子衰老提供了严格的生成范式。

英文摘要

DNA methylation (DNAm) serves as one of the most robust molecular biomarkers of biological aging. While conventional epigenetic clocks accurately predict chronological age from high-dimensional CpG profiles, they treat aging as a static regression task, meaning they can only output a single score rather than simulating how an entire profile continuously changes over time. To reconstruct these continuous dynamics, we frame lifelong human epigenetic aging as a trajectory inference problem across discrete age snapshots derived from widely available cross-sectional data. We introduce a two-stage computational pipeline: first, an age-regularized Variational Autoencoder (VAE) maps high-dimensional CpG profiles onto a chronologically ordered latent manifold while preserving a generative decoder bridge back to the original methylation space. Second, we model the continuous movement across this latent space via Regularized Unbalanced Optimal Transport (RUOT) that unifies deterministic drift, random diffusion, and non-conservative mass changes. By resolving this RUOT formulation using the DeepRUOT framework, our model fluidly accommodates population-level density shifts like survivorship bias and cellular attrition without requiring rigid biological priors. Evaluated on a large-scale, 80-year pan-tissue dataset, our model demonstrates robust distribution interpolation and uncovers a prominent late-life surge in the learned growth field that mathematically captures the variance expansion driven by stochastic epigenetic drift. Finally, by decoding continuous latent paths back to individual CpG sites, we reconstruct and empirically verify distinct biological aging archetypes, offering a rigorous, generative paradigm for simulating human molecular aging.

URL PDF HTML 收藏
2607.02289 2026-07-08 quant-ph cs.AI hep-ex 新提交

Neural-Network Inverse Design of SRF Cavities and Transmons for Bosonic Quantum Computation

用于玻色子量子计算的超导射频腔和Transmon的神经网络逆向设计

Joseph Yaker, Jovan Markovic, Alessandro Reineri, Doga Murat Kurkcuoglu, Silvia Zorzetti

机构 * Superconducting and Quantum Materials System Center (SQMS)(超导与量子材料系统中心) Fermi National Accelerator Laboratory(费米国家加速器实验室) Applied Physics Program, Northwestern University(西北大学应用物理计划) Department of Physics, University of Cambridge(剑桥大学物理系) Illinois Institute of Technology(伊利诺伊理工学院) Department of Physics and Astronomy, Northwestern University(西北大学物理与天文学系)

AI总结 提出两种深度神经网络方法,分别逆向设计超导射频腔几何和Transmon量子比特参数,实现从目标性能到候选设计的快速映射,误差约5%和2%。

详情
AI中文摘要

三维超导射频(SRF)腔提供异常长寿命的电磁模式,当与非线性元件(如transmon量子比特)耦合时,成为有前景的玻色子量子信息处理架构。这类系统的逆向设计,即恢复产生指定电磁和耦合目标的器件几何形状,通常是一个一对多问题。量子比特-腔耦合强度敏感地依赖于transmon几何及其在腔电磁场中的位置。随着这些系统规模扩大和设计参数空间增长,传统迭代模拟的成本变得过高。我们提出了两种深度神经网络(DNN)方法,在设计栈的互补层次上解决这一逆向设计问题。第一种提出产生目标腔可观测量的SRF腔几何。第二种提出产生目标量子比特-腔参数——耦合率、量子比特频率和非谐性$(g, \nu_q, \alpha)$——的transmon量子比特设计。恢复的候选设计与目标匹配在约5%(腔)和约2%(transmon)以内,通过端到端重模拟确认。两种方法都将期望的器件行为直接映射到候选设计,是通常需要的迭代模拟研究的快速替代方案。

英文摘要

Three-dimensional superconducting radio-frequency (SRF) cavities provide exceptionally long-lived electromagnetic modes and, when coupled to nonlinear elements such as transmon qubits, become promising architectures for bosonic quantum information processing. The inverse design of such systems, i.e., recovering device geometries that produce specified electromagnetic and coupling targets, is generally a one-to-many problem. The qubit-cavity coupling strength depends sensitively on both the transmon geometry and its position within the cavity's electromagnetic field. As these systems scale up and their design parameter spaces grow, the cost of conventional iterative simulation becomes prohibitive. We present two deep neural network (DNN) approaches that address this inverse-design problem at complementary levels of the design stack. The first proposes SRF cavity geometries that produce target cavity observables. The second proposes transmon qubit designs that produce target qubit-cavity parameters - the coupling rate, qubit frequency, and anharmonicity $(g, ν_q, α)$. The recovered candidate designs match the targets to within ~5% (cavity) and ~2% (transmon), confirmed by end-to-end re-simulation. Both approaches map desired device behavior directly to candidate designs, a fast alternative to the iterative simulation studies usually required.

URL PDF HTML 收藏
2606.06576 2026-07-08 cs.LG astro-ph.EP astro-ph.IM stat.ML 新提交

Gaussian Process Latent Factor Regression for Low-Data, High-Dimensional Output Problems

高斯过程潜在因子回归用于低数据高维输出问题

Edward T. Stevenson, Eric T. Wolf, Mei Ting Mak, N. J. Mayne, Miles Cranmer

机构 * University of Cambridge(剑桥大学) University of Colorado Boulder(科罗拉多大学博尔德分校) University of Oxford(牛津大学) University of Exeter(埃克塞特大学)

AI总结 提出高斯过程潜在因子回归(GPLFR)模型,通过将输出表示为低维潜在状态的线性高斯解码,联合优化压缩与预测,解决低数据高维输出回归问题,并首次构建岩石系外行星全球气候模型的空间分辨仿真器。

Comments 9 pages content + 23 pages appendix/references. Supporting code at https://github.com/edstevenson/GPLFR

详情
AI中文摘要

在科学领域,回归任务通常需要从少量训练样本预测高维输出。多输出高斯过程在低数据场景中表现出色,但通常难以处理高维输出。PCA-GP(主成分分析加高斯过程回归)等压缩-预测流程处理了高维性,但依赖于为重构而非预测优化的基。为弥补这一差距,我们提出一个模型,将每个输出表示为从高斯过程先验中抽取的低维潜在状态的线性高斯解码。通过解析地边缘化解码器权重,我们将压缩和预测耦合在一个可扩展到高维输出的单一目标中。我们将此模型称为高斯过程潜在因子回归(GPLFR)。我们通过构建首个岩石系外行星全球气候模型的空间分辨仿真器来演示GPLFR。

英文摘要

In the sciences, regression tasks often require predicting high-dimensional outputs from few training examples. Multi-output Gaussian processes excel in low-data regimes but typically struggle with high-dimensional outputs. Compress-then-predict pipelines such as PCA-GP (principal component analysis plus Gaussian process regression) handle high dimensionality, but rely on bases optimized for reconstruction rather than prediction. To address this gap, we propose a model that represents each output as a linear-Gaussian decoding of a low-dimensional latent state drawn from a Gaussian process prior. By analytically marginalizing the decoder weights, we couple compression and prediction in a single objective that scales to high-dimensional outputs. We refer to this model as Gaussian process latent factor regression (GPLFR). We demonstrate GPLFR by building the first spatially resolved emulator of global climate models for rocky exoplanets.

URL PDF HTML 收藏