arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

至 收录 747
2607.17075 2026-07-21 cs.CR cs.CL 新提交

A Systematic Evaluation of Traditional Privacy Policy Analysis Tools Against LLMs

传统隐私政策分析工具针对大语言模型的系统评估

Madhav Aryal, Sudipa Saha, Kaushal Kafle, Anshuman Chhabra, Sunil Manandhar

机构 * University of South Florida(佛罗里达州立大学) IBM T.J. Watson Research Center(IBM 汤普森研究中心)

AI总结 本文系统评估现成大语言模型能否取代专业隐私分析工具,研究六个具三种主要功能及三个中间任务的工具,对比两个先进大语言模型与工具性能,发现大语言模型在隐私政策和法规分析功能上能匹配或超越现有工具。

详情
AI中文摘要

大语言模型(LLMs)的出现显著改变了隐私政策和数据合规性分析研究,使以前需要特定领域工具的任务得以实现。然而,LLMs在多大程度上能真正复制先前工作提供的多样功能、方法和分析尚不清楚。本文首次系统评估现成的LLMs能否取代专业隐私分析工具。研究了涵盖三种主要功能(矛盾检测、法规合规分析、隐私政策总结与聚合)及三个中间任务(使用元组进行结构化数据提取、语义角色标注(SRL)和手动隐私政策标注)的六个代表性工具。通过直接促使模型在10个隐私政策的自定义数据集上执行相应功能和任务,比较了两个最先进的LLMs(不同配置下的GPT-5.2和Gemini-2.5)与工具的性能,以评估现成模型能否在无需进一步工程或特定领域训练的情况下产生工具特定功能。结果表明,LLMs在各项功能上始终能匹配或超越现有工具。在第一方收集实体的手动标注中,LLMs平均精确率达81.8%,召回率达70.9%;在第三方共享实体标注中,与OPP-115数据集相比,平均精确率为91.4%,召回率为70.8%。总体而言,研究结果表明LLMs能有效执行隐私政策和法规分析中以前需要专业工具的广泛功能和任务。

英文摘要

The advent of LLMs has significantly changed the research on privacy policy and data compliance analysis by enabling tasks that previously required specialized, domain-specific tools. However, it remains unclear to what extent LLMs can truly replicate the diverse functionalities, and the wide range of methodologies and analysis offered by prior work. In this paper, we conduct the first systematic evaluation of whether off-the-shelf LLMs can replace specialized privacy analysis tools. We study six representative tools spanning three major functionalities: contradiction detection, regulatory compliance analysis, and privacy policy summarization and aggregation, and across three intermediate tasks: structured data extraction using tuples, Semantic Role Labeling (SRL) and manual privacy policy labeling. We compare the performance of two state-of-the-art LLMs (GPT-5.2 and Gemini-2.5 in various configurations) against the tools by directly prompting the models to perform corresponding functionalities and tasks on a custom dataset of 10 privacy policies, allowing us to assess whether off-the-shelf models can produce tool-specific functionalities without further engineering or domain-specific training, major limitations in prior work. Our results show that LLMs consistently match or exceed the capabilities of existing tools across the functionalities. In manual labeling of first-party collection entities, LLMs achieved an average precision of 81.8% and recall of 70.9%, while for labeling of third-party sharing entities, they achieved an average precision of 91.4% and recall of 70.8% compared to the OPP-115 dataset. Overall, our findings indicate that LLMs can effectively perform a broad range of functionalities and tasks in privacy policy and regulation analysis that previously required specialized tools.

URL PDF HTML 收藏
2607.17872 2026-07-21 quant-ph cs.DC cs.ET cs.LG 新提交

Entanglement geometry separates circuit cutting, classical hardness, and trainability

纠缠几何分离电路切割、经典硬度和可训练性

Maria Gragera Garces, Sabina Drăgoi, Lirandë Pira

机构 * Quantum Software Lab(量子软件实验室) University of Edinburgh, UK(爱丁堡大学) IBM Research(IBM研究院) Centre for Quantum Technologies(量子技术中心) National University of Singapore(新加坡国立大学)

AI总结 研究表明纠缠几何约束电路切割等属性,具恒定接缝键维度的MPS和TTN电路可经典模拟,构建的双块电路家族可廉价切割,MPS硬度和可训练性深度范围不兼容,用魔法作硬度资源可避免冲突,浅Clifford+\(T\)电路有相应特性。

Comments 4 pages, 2 figures

详情
AI中文摘要

电路切割有望扩展量子计算规模,但变分量子优势还需低切割开销、经典硬度和可训练性。我们表明这些属性受纠缠几何强烈约束。具有恒定接缝键维度的矩阵乘积态(MPS)和树张量网络(TTN)电路可在\(O(1/\varepsilon^2)\)采样开销下切割,但仍可高效经典模拟,排除了这些家族内的渐近量子优势。通过独立控制接缝和块内纠缠,我们构建了一个双块电路家族,它可廉价切割且需要超多项式全局MPS键维度,数值上支持到\(n = 100\)。然而,MPS硬度和可训练性需要不兼容的深度范围,分别为\(d=\omega(\log n)\)和\(d=O(\log n)\)。使用魔法而非纠缠作为硬度资源可避免此冲突:浅的Clifford+\(T\)电路可切割且可训练,同时其稳定器模拟成本随\(T\)计数呈指数增长。

英文摘要

Circuit cutting promises to scale quantum computations beyond current hardware, but variational quantum advantage also requires low cutting overhead, classical hardness, and trainability. We show that these properties are strongly constrained by entanglement geometry. Matrix product state (MPS) and tree tensor network (TTN) circuits with constant seam bond dimension can be cut with \(O(1/\varepsilon^2)\) sampling overhead, but remain efficiently classically simulable, ruling out asymptotic quantum advantage within these families. By independently controlling seam and intra-block entanglement, we construct a two-block circuit family that remains cheaply cuttable while requiring a super-polynomial global MPS bond dimension, as supported numerically up to \(n=100\). However, MPS hardness and trainability require incompatible depth regimes, \(d=ω(\log n)\) and \(d=O(\log n)\), respectively. Using magic rather than entanglement as the hardness resource avoids this conflict: shallow Clifford+\(T\) circuits remain cuttable and trainable while their stabiliser-simulation cost grows exponentially with the \(T\)-count.

URL PDF HTML 收藏
2406.18082 2026-07-21 cs.CL cs.HC 版本更新

Octo-planner: On-device Language Model for Planner-Action Agents

Octo-planner:用于规划-行动智能体的设备端语言模型

Wei Chen, Zhiyuan Li, Zhen Guo, Yikang Shen

机构 * Nexa AI & Stanford(Nexa AI 与 斯坦福大学) MIT EECS(麻省理工学院电子工程与计算机科学系) MIT-IBM Watson AI Lab(麻省理工-IBM Watson AI 实验室)

AI总结 研究如何让人工智能智能体有效规划行动,提出设备端规划-行动框架,分离规划与行动执行组件,用模型微调优化性能,多LoRA训练方法应对多域规划挑战,在域内测试成功率达97%,并开源模型权重。

详情
AI中文摘要

人工智能智能体在各领域愈发重要,需有效规划过程。本文提出高效的设备端规划-行动框架,将规划与行动执行分离为两个组件:基于为边缘设备优化的38亿参数语言模型Phi-3 Mini的规划智能体,以及使用章鱼模型执行功能的行动智能体。规划智能体先将任务分解为子步骤响应用户查询,再由行动智能体执行。为在资源受限设备上优化性能,采用模型微调而非上下文学习。利用GPT-4生成规划查询和响应并验证数据质量,在精选数据集上微调Phi-3 Mini模型,在域内测试环境成功率达97%。还开发多LoRA训练方法应对多域规划挑战,开源了模型权重。

英文摘要

AI agents have become increasingly significant in various domains, enabling autonomous decision-making and problem-solving. To function effectively, these agents require a planning process that determines the best course of action and then executes the planned actions. In this paper, we present an efficient on-device Planner-Action framework that separates planning and action execution into two distinct components: a planner agent based on Phi-3 Mini, a 3.8 billion parameter LLM optimized for edge devices, and an action agent using the Octopus model for function execution. The planner agent first responds to user queries by decomposing tasks into a sequence of sub-steps, which are then executed by the action agent. To optimize performance on resource-constrained devices, we employ model fine-tuning instead of in-context learning, reducing computational costs and energy consumption while improving response times. Our approach involves using GPT-4 to generate diverse planning queries and responses based on available functions, with subsequent validations to ensure data quality. We fine-tune the Phi-3 Mini model on this curated dataset, achieving a 97\% success rate in our in-domain test environment. To address multi-domain planning challenges, we developed a multi-LoRA training method that merges weights from LoRAs trained on distinct function subsets. This approach enables flexible handling of complex, multi-domain queries while maintaining computational efficiency on resource-constrained devices. To support further research, we have open-sourced our model weights at https://huggingface.co/NexaAIDev/octopus-planning. For the demo, please refer to https://www.nexa4ai.com/octo-planner.

URL PDF HTML 收藏
2607.15655 2026-07-20 cs.CL cs.LG 新提交

Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models

扩散语言模型的自适应多步前瞻解码

Yingqian Cui, Wei Deng, Lantao Mei, Hang Li, Charu C. Aggarwal, Hui Liu, Yue Xing

机构 * Michigan State University(密歇根州立大学) Morgan Stanley(摩根士丹利) IBM T.J. Watson Research Center(IBM 托马斯·J·沃森研究中心)

AI总结 研究针对扩散语言模型解码,提出自适应多步前瞻框架AdaLook,基于候选分数方差动态决定是否继续展开及进行分支扩展,避免不必要的深度展开,实验证明其在准确性和解码步骤权衡上优于现有一步前瞻解码方法。

详情
AI中文摘要

掩码扩散语言模型(DLMs)通过迭代细化掩码令牌实现并行文本生成,为自回归解码提供了有前景的替代方案。近期基于前瞻的解码方法通过探索未来解码状态改善了准确性与效率的权衡。但现有方法主要依赖浅层次的一步前瞻,对更长解码轨迹并非最优。我们发现深度前瞻的简单扩展也无效。因此,本文提出AdaLook,一个用于DLM解码的自适应前瞻框架。它基于候选分数方差动态决定是否继续展开,并在中间展开状态需要额外探索时进行分支扩展。实验表明AdaLook比现有一步前瞻解码方法在准确性和解码步骤权衡上表现更好。

英文摘要

Masked diffusion language models (DLMs) enable parallel text generation by iteratively refining masked tokens, offering a promising alternative to autoregressive decoding. Recent lookahead-based decoding methods improve the accuracy--efficiency trade-off by exploring future decoding states before committing token updates. However, existing approaches mainly rely on shallow one-step lookahead, which optimizes immediate information gain but can be suboptimal for longer-horizon decoding trajectories. Meanwhile, we find that a naive extension for deeper lookahead is also ineffective, as fixed-depth rollout introduces additional computation and cannot adapt to heterogeneous intermediate decoding states. Thus, in this work, we propose AdaLook, an adaptive lookahead framework for DLM decoding. AdaLook dynamically determines whether to continue rollout based on candidate-score variance and further enables branch expansion when intermediate rollout states require additional exploration. This design avoids unnecessary deep rollout while allowing the decoder to re-trigger lookahead from informative intermediate states. Experiments on various benchmarks and models demonstrate that AdaLook achieves a better accuracy--decoding steps trade-off than existing one-step lookahead decoding methods.

URL PDF HTML 收藏
2607.15313 2026-07-20 cs.LG 新提交

Position: Quantum Program Generation Must Prioritize Validity Over Probabilistic Scaling

立场:量子程序生成必须优先考虑有效性而非概率缩放

Junhao Song, Yu Zhou, William Knottenbelt, Yudong Cao

机构 * IBM(IBM公司) DeepMind(深度思维公司)

AI总结 该论文指出将概率范式用于量子电路合成有误,因量子电路有语法语义差距,未经验证的训练使模型难掌握物理语义,有效子集随量子比特数指数衰减。提出转向以验证器为中心,集成多种元素到生成中,验证意识架构是可行途径,应编码特定规则而非仅靠模仿。

Comments Accepted to ICML 2026: https://openreview.net/forum?id=oX1vWuQ13y

详情
AI中文摘要

缩放假设认为增加模型参数会产生新兴推理能力。本文认为将这种概率范式应用于通用量子电路合成是一个方向性错误。与自然语言不同,量子电路需要严格遵守数学约束,这导致了显著的语法-语义差距。对未经验证的量子程序进行训练意味着模型学习语法但无法捕捉希尔伯特空间的物理语义。由于电路设计的有效子集随量子比特数量呈指数衰减,事后过滤在数学上是难以处理的。我们提出从以人类为中心的副驾驶转向以验证器为中心的代理。我们将分层约束、拓扑掩码和符号代理直接集成到生成过程中。我们的分析表明,仅靠规模无法弥合有效性差距。具有验证意识的架构为模块化量子程序生成提供了一条可行的途径。这些考虑指向了编码量子信息特定任务规则的生成方法,而不是仅仅依赖模仿。

英文摘要

The scaling hypothesis assumes that increasing model parameters yields emergent reasoning capabilities. This position paper argues that applying this probabilistic paradigm to generic quantum circuit synthesis is a directional error. Unlike natural languages, quantum circuits require strict adherence to mathematical constraints that manifest a significant syntax-semantics gap. Training on unverified quantum programs means that models learn syntax but fail to capture the physical semantics of the Hilbert space. Since the valid subset of circuit designs decays exponentially with the number of qubits, post-hoc filtering is mathematically intractable. We propose a pivot from human-centric copilots to verifier-centric agents. We integrate hierarchical constraints, topological masks, and symbolic proxies directly into generation. Our analysis suggests that scale alone cannot bridge the validity gap. Verification-aware architectures offer a viable path for modular quantum program generation. These considerations point toward generation methods that encode task-specific rules of quantum information, rather than relying on imitation alone.

URL PDF HTML 收藏
2607.14318 2026-07-17 cs.LG 新提交

Counterfactual Optimal Action Trees (COAT): Interpretable Prescriptive Policies from Observational Data

反事实最优行动树(COAT):从观测数据中学习可解释的规范性策略

Youssef Drissi, Markus Ettl, Shivaram Subramanian, Wei Sun, Zack Xue

机构 * IBM Research(IBM研究院)

AI总结 研究从观测数据学习可解释规范性策略的问题,核心方法是结合反事实结果估计与大规模混合整数优化的COAT框架,主要贡献是应用于航空公司辅助定价提升收入,推动扩大采用及相关决策举措。

详情
AI中文摘要

我们引入了反事实最优行动树(COAT),这是一个从观测数据中学习可解释规范性策略的框架。COAT将反事实结果估计与大规模混合整数优化相结合,利用列生成将因果预测转化为在业务和监管约束下可行、透明的决策。我们将COAT应用于航空公司辅助定价,该场景具有复杂业务规则和有限实验灵活性。在与一家全球主要航空公司进行的为期17周的实地试点中,COAT使每次预订的追加销售收入提高了6.9%,该航空公司预计在符合条件的国内市场每年将增加5000万至1.5亿美元的高端座位收入。试点的成功促使其扩大采用,并为该组织内更广泛的人工智能驱动决策举措提供了参考。

英文摘要

We introduce COAT (Counterfactual Optimal Action Tree), a framework for learning interpretable prescriptive policies from observational data. COAT combines counterfactual outcome estimation with large-scale mixed-integer optimization, using column generation to translate causal predictions into feasible, transparent decisions under business and regulatory constraints. We apply COAT to airline ancillary pricing, a setting characterized by complex business rules and limited experimental flexibility. In a 17-week field pilot with a major global airline, COAT increased upsell revenue per booking by 6.9%, with the airline projecting \$50-\$150 million in incremental annual premium seat revenue across eligible domestic markets. The success of the pilot led to scaled adoption and informed broader AI-driven decision initiatives within the organization.

URL PDF HTML 收藏
2512.14332 2026-07-17 cs.CL cs.AI 版本更新

Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring

步骤标记:通过步骤监控控制语言推理模型的生成

Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo, John D. Kelleher

机构 * IBM Research Europe(IBM欧洲研究院) Trinity College Dublin(都柏林信任学院) ADAPT Research Centre(ADAPT研究中心)

AI总结 针对语言推理模型效率低、过度生成推理步骤的问题,提出步骤标记框架,引入ReasonType分类法,可在线监控特定步骤计数以产生早期停止标准,经实验在保持准确性时减少令牌数,为控制模型生成及研究其行为提供新途径和工具。

Comments ICML 2026 Workshop on Resource-Adaptive Foundation Model Inference (AdaptFM), Seoul, South Korea

详情
AI中文摘要

在过去几年中,语言推理模型(LRMs)领域非常活跃,训练和推理技术的进步使LRMs能够更准确、更深入地推理。然而,越来越多的研究表明,LRMs仍然效率低下,过度生成验证和反思步骤。为应对这一挑战,我们引入了步骤标记框架,这是一种轻量级句子分类器,能够实时标注LRMs正在生成的推理步骤类型。为监控推理行为,我们引入了ReasonType:一种新颖的推理步骤分类法。在此框架基础上,我们证明了对特定步骤计数的在线监控可以产生有效的可解释的LRM推理早期停止标准。我们在三个开源推理模型上,针对标准基准数据集:MATH500、GSM8K、AIME以及非数学任务(GPQA和MMLU-Pro)评估了步骤标记框架。我们在保持与标准生成相当的准确性的同时,实现了20%到50%的令牌减少,在计算量更大的任务上获得了最大收益。这项工作提供了一种增加对LRMs生成控制的新方法,以及一种研究LRMs行为的新工具。

英文摘要

The field of Language Reasoning Models (LRMs) has been very active over the past few years with advances in training and inference techniques enabling LRMs to reason longer, and more accurately. However, a growing body of studies show that LRMs are still inefficient, over-generating verification and reflection steps. To address this challenge, we introduce the Step-Tagging framework, a lightweight sentence-classifier enabling real-time annotation of the type of reasoning steps that an LRM is generating. To monitor reasoning behaviors, we introduced ReasonType: a novel taxonomy of reasoning steps. Building on this framework, we demonstrated that online monitoring of the count of specific steps can produce effective interpretable early stopping criteria of LRM inferences. We evaluate the Step-tagging framework on three open-source reasoning models across standard benchmark datasets: MATH500, GSM8K, AIME and non-mathematical tasks (GPQA and MMLU-Pro). We achieve 20 to 50% token reduction while maintaining comparable accuracy to standard generation, with largest gains observed on more computation-heavy tasks. This work offers a novel way to increase control over the generation of LRMs, and a new tool to study behaviors of LRMs.

URL PDF HTML 收藏
2607.13897 2026-07-16 cs.LG 新提交

RF Spectrogram Anomaly Detection with Quantum Kitchen Sinks: Architecture, Representation, and Hardware Validation

基于量子随机特征映射的射频频谱图异常检测:架构、表示与硬件验证

Abdallah Aaraba, Alexis Vieloszynski, Remon Polus, Ola Ahmad, Soumaya Cherkaoui

机构 * ibm_quebec(IBM魁北克)

AI总结 研究针对无线射频网络异常检测问题,扩展QKS模板并引入消融协议,通过多深度数据重新上传和环纠缠进行评估。结果表明DCT表示优,适度深度纠缠QKS配置强,QKS优于经典基线,提供了实用可重复的无线网络异常检测框架。

Comments Paper accepted to IEEE quantum week 2026

详情
AI中文摘要

无线信道的广播特性使射频网络易受异常和恶意传输影响,异常检测是安全频谱管理的基本要求。量子随机特征映射(QKS)是适用于近期量子设备的轻量级混合量子特征映射,但其在结构化信号数据上的行为尚不清楚。本文通过多深度数据重新上传和环纠缠扩展了标准QKS模板,并在受控射频频谱图异常检测中评估了所得流程。引入了一个验证锁定的五阶段消融协议,系统地分离了浅层架构、重新上传深度、实验预算、输入表示和经典读出的影响。在完整基准测试中,离散余弦变换(DCT)表示始终优于原始和主成分分析(PCA)输入,适度深度的纠缠QKS配置形成最强操作模式,QKS在所有评估的表示 - 读出对上优于匹配的经典直接读出基线,最佳配置在测试集上达到接收器操作特征曲线下面积(AUROC)为0.8778和测试F1为0.799。该研究在数据方面使用实际测量的低于6GHz蜂窝信号,在计算方面在ibm_quebec量子处理单元(QPU)上进行实际设备验证,AUROC偏差相对于模拟低于0.013。这些结果为在无线网络中部署基于QKS的异常检测提供了一个实用、可重复的框架。

英文摘要

The broadcast nature of wireless channels exposes radio-frequency (RF) networks to anomalous and malicious transmissions, making anomaly detection a fundamental requirement for secure spectrum management. Quantum Kitchen Sinks (QKS) offer a lightweight hybrid quantum feature map suitable for near-term quantum devices, yet their behavior on structured signal data remains poorly understood. In this paper, we extend the standard QKS template with multi-depth data re-uploading and ring entanglement, and evaluate the resulting pipeline on controlled RF spectrogram anomaly detection. We introduce a validation-locked five-stage ablation protocol that systematically separates the effects of shallow architecture, re-uploading depth, episode budget, input representation, and classical readout. Across the completed benchmark, Discrete Cosine Transform (DCT) representations consistently dominate raw and Principal Component Analysis (PCA) inputs, moderate-depth entangled QKS configurations form the strongest operating regime, and QKS improves over matched classical direct-readout baselines across all evaluated representation-readout pairs on the held-out test set, with the best configuration reaching a test Area Under the Receiver Operating Characteristic curve (AUROC) of 0.8778 and a test F1 of 0.7995. The study bridges two levels of realism: real measured sub-6\,GHz cellular signals on the data side and real-device validation on the ibm_quebec Quantum Processing Unit (QPU) on the computing side, with AUROC deviations below 0.013 relative to simulation. These results provide a practical, reproducible framework for deploying QKS-based anomaly detection in wireless networks.

URL PDF HTML 收藏
2607.13416 2026-07-16 cs.LG 新提交

EXPLORE: Exploration with Guided Search for Analog Topology Generation using Language Models

EXPLORE:使用语言模型进行引导搜索以生成模拟拓扑结构

Guanglei Zhou, Chen-Chia Chang, Yikang Shen, Jonathan Ku, Isaac Jacobson, Jingyu Pan, Yiran Chen, Xin Zhang

机构 * Duke University(杜克大学) MIT-IBM Watson AI Lab(麻省理工学院-IBM沃森人工智能实验室) IBM T. J. Watson Research Center(IBM T. J. 沃森研究中心)

AI总结 本文针对自动化模拟电路拓扑设计难题,提出EXPLORE框架,集成模拟器引导蒙特卡罗树搜索与基于变压器的解码,利用语言模型先验优化搜索,在6组件基准测试中显著提升成功率、降低均方误差,推动LLM驱动设计自动化。

Comments MLCAD 26' accepted

详情
AI中文摘要

自动化模拟电路拓扑设计对于减少满足日益多样化和定制化应用需求所需的大量人工工作至关重要。最近的进展是在预训练语言模型上应用序列到序列微调,以单次从用户规范直接生成电路拓扑。然而,由于搜索空间呈指数增长且训练数据集有限,这些一次性生成方法无法生成复杂电路。本文提出了EXPLORE,这是一个搜索增强框架,它将模拟器引导的蒙特卡罗树搜索(MCTS)与基于变压器的解码相结合,以实现模拟拓扑生成的测试时扩展。通过利用语言模型先验并绕过高置信度结构令牌,EXPLORE在搜索过程中将昂贵的模拟器预算主要分配给改变拓扑的决策。在公差为0.01的6组件基准测试中,EXPLORE将一次性生成的成功率从12%和采样与过滤基线的33%提高到65%,并在相同搜索预算下相对于采样与过滤将均方误差降低了20%以上。这些结果使EXPLORE成为第一个将结构化测试时搜索与LM解码集成用于模拟拓扑生成的框架,也是迈向扩展LLM驱动设计自动化的实际一步。

英文摘要

Automating analog circuit topology design is essential to reduce the extensive manual effort required to meet increasingly diverse and customized application demands. Recent advances have applied sequence-to-sequence fine-tuning on pretrained language models to directly generate circuit topologies from user specifications in a single pass. However, these one-shot generation methods failed to generate complex circuits due to their exponentially growing search spaces and limited training datasets. In this paper, we present EXPLORE, a search-enhanced framework that integrates simulator-guided Monte Carlo Tree Search (MCTS) with transformer-based decoding to enable test-time scaling for analog topology generation. By leveraging language-model priors and bypassing high-confidence structural tokens, EXPLORE allocates expensive simulator budget primarily toward topology-altering decisions during search. On a 6-component benchmark at a tight tolerance of 0.01, EXPLORE raises the success rate from 12% for one-shot generation and 33% for a sampling-and-filter baseline to 65%, and lowers MSE by over 20% relative to sampling-and-filter under the same search budget. These results establish EXPLORE as the first framework to integrate structured test-time search with LM decoding for analog topology generation, and a practical step toward scaling LLM-driven design automation.

URL PDF HTML 收藏
2605.15026 2026-07-16 cs.OS cs.AI cs.PF 版本更新

TuxBot: Semantic-Aware Online OS Tuning with Large Language Models

SemaTune: 基于大语言模型的语义感知在线操作系统调优

Georgios Liargkovas, Mihir Nitin Joshi, Hubertus Franke, Kostis Kaffes

机构 * Columbia University(哥伦比亚大学) IBM Research(IBM研究院)

AI总结 SemaTune通过语义感知的在线操作系统调优框架,利用大语言模型进行有限制的指导,提升稳定状态下的性能,通过快速和慢速循环更新配置并验证,实现比传统方法更高的性能提升。

Comments 18 pages, 12 figures

详情
AI中文摘要

在线操作系统调优可以提高长期运行的服务,但现有控制器与实时主机不匹配。它们将调度器、电源、内存和I/O控制视为黑盒变量并优化标量奖励。这种观点忽略了跨控制旋钮的策略结构,当应用指标不可用时会崩溃,并可能导致运行服务进入持续存在的降级区域。我们提出了SemaTune,这是一个主机侧的稳定状态操作系统调优框架,通过有限制的语言模型指导。SemaTune将控制旋钮模式、遥测、当前配置、近期操作-响应历史以及检索到的先前运行转换为紧凑的决策上下文。一个快速循环提出低延迟的更新,一个较慢的循环定期修订搜索策略,且每次提出的更改在通过类型验证后才能到达内核或sysctl接口。这使控制器能够思考操作系统控制的含义和间接性能信号,同时保持模型成本、延迟和权威受限。我们评估了SemaTune在13个活的工作负载上,同时调优多达41个Linux参数。在所有套件中,SemaTune在稳定阶段的性能比默认设置提高了72.5%,比最强的非LLM基线提高了153.3%。一个30窗口会话的模型调用成本约为0.20美元。仅使用主机级别指标,SemaTune在直接应用目标下仍比基线高出93.7个百分点,同时避免了由结构盲探索达到的严重降级区域。

英文摘要

Online OS tuning can improve long-running services, but existing controllers are poorly matched to live hosts. They treat scheduler, power, memory, and I/O controls as black-box variables and optimize a scalar reward. This view ignores cross-knob policy structure, breaks down when application metrics are unavailable, and can send a running service into degraded regions that persist after the bad setting is removed. We present TuxBot, a host-side framework for steady-state OS tuning with bounded language-model guidance. TuxBot turns knob schemas, telemetry, current configuration, recent action--response history, and retrieved prior runs into a compact decision context. A fast loop proposes low-latency updates, a slower loop periodically revises the search strategy, and every proposed change passes through typed validation before reaching kernel or sysctl interfaces. This lets the controller reason about OS-control meaning and indirect performance signals while keeping model cost, latency, and authority constrained. We evaluate TuxBot on 13 live workloads from five benchmark suites while tuning up to 41 Linux parameters. Across the suite, TuxBot improves stable-phase performance by 72.5% over default settings and by 153.3% relative to the strongest non-LLM baseline. A 30-window session costs about $0.20 in model calls. With only host-level metrics, TuxBot still outperforms baselines given direct application objectives by 93.7 percentage points, while avoiding severe degraded regions reached by structure-blind exploration.

URL PDF HTML 收藏
2411.02317 2026-07-16 cs.LG cs.AI cs.CY

Defining and Evaluating Physical Safety for Large Language Models

定义和评估大语言模型的物理安全性

Yung-Chen Tang, Pin-Yu Chen, Tsung-Yi Ho

机构 * The Chinese University of Hong Kong(香港中文大学) IBM Research(IBM研究院)

AI总结 本文提出了一种无人机控制的全面基准,评估大语言模型在物理安全方面的风险与权衡,揭示了模型在安全性和实用性之间的不理想平衡。

Journal ref Communications of the ACM, 2026

详情
AI中文摘要

大语言模型(LLMs)越来越多地用于控制如无人机等机器人系统,但其在现实应用中造成物理威胁和危害的风险仍未经探索。我们的研究通过开发一个全面的无人机控制基准来填补评估LLM物理安全性的关键空白。我们将无人机的物理安全风险分为四类:(1)针对人类的威胁,(2)针对物体的威胁,(3)基础设施攻击,(4)违规行为。我们对主流LLMs的评估揭示了效用与安全之间的不理想权衡,擅长代码生成的模型在关键安全方面表现不佳。此外,虽然结合先进的提示工程技术如上下文学习和思维链可以提高安全性,但这些方法仍难以识别无意攻击。此外,更大模型在安全能力上表现更好,特别是在拒绝危险命令方面。我们的发现和基准可以促进LLM物理安全性的设计和评估。项目页面可在huggingface.co/spaces/TrustSafeAI/LLM-physical-safety上找到。

英文摘要

Large Language Models (LLMs) are increasingly used to control robotic systems such as drones, but their risks of causing physical threats and harm in real-world applications remain unexplored. Our study addresses the critical gap in evaluating LLM physical safety by developing a comprehensive benchmark for drone control. We classify the physical safety risks of drones into four categories: (1) human-targeted threats, (2) object-targeted threats, (3) infrastructure attacks, and (4) regulatory violations. Our evaluation of mainstream LLMs reveals an undesirable trade-off between utility and safety, with models that excel in code generation often performing poorly in crucial safety aspects. Furthermore, while incorporating advanced prompt engineering techniques such as In-Context Learning and Chain-of-Thought can improve safety, these methods still struggle to identify unintentional attacks. In addition, larger models demonstrate better safety capabilities, particularly in refusing dangerous commands. Our findings and benchmark can facilitate the design and evaluation of physical safety for LLMs. The project page is available at huggingface.co/spaces/TrustSafeAI/LLM-physical-safety.

URL PDF HTML 收藏
2605.01965 2026-07-14 cs.LG 版本更新

Retrieval with Multiple Query Vectors through Anomalous Pattern Detection

通过异常模式检测进行多查询向量检索

Allassan Tchangmena A Nken, Baimam Boukar Jean Jacques, Miriam Rateike, Celia Cintas, Skyler Speakman

机构 * IBM Research(IBM研究院)

AI总结 针对复杂任务需多查询向量的问题,提出利用异常模式检测的检索方法,通过识别查询向量中突出的维度子集,扫描数据库检索相关向量集,并在多个数据集上验证,发现多数数据集上大查询集可提升检索性能,1到8个查询向量时提升明显。

详情
AI中文摘要

经典向量检索问题通常将单个查询嵌入向量作为输入,从向量数据库中检索最相似的嵌入向量。然而,复杂的推理和检索任务常常需要多个查询向量。本文提出一种检索方法,它同时考虑多个查询向量,并利用异常模式检测的概念从数据库中检索最相关的向量。具体做法是利用一组查询向量\(Q\)(\(|Q|\geq 1\)),识别出\(Q\)中突出(异常)的向量维度子集,然后扫描向量数据库,检索在这些已识别维度上也异常的向量集并返回。在两个图像数据集、一个文本数据集和一个表格数据集上验证了该方法。总体而言,在大多数数据集上,更大的查询集能提高检索性能,从1增加到8时提升最显著,之后收益变小。

英文摘要

A classical vector retrieval problem typically considers a \emph{single} query embedding vector as input and retrieves the most similar embedding vectors from a vector database. However, complex reasoning and retrieval tasks frequently require \emph{multiple query vectors}, rather than a single one. In this work, we propose a retrieval method that considers multiple query vectors simultaneously and retrieves the most relevant vectors from the database using concepts from anomalous pattern detection. Specifically, our approach leverages a set of query vectors $Q$ (with $|Q|\geq 1$), and identifies the subset of vector dimensions within $Q$ that standout (anomalous) from the rest of dimensions. Next, we scan the vector database to retrieve the set of vectors that are also anomalous across the previously identified vector dimensions and return them as our retrieved set of vectors. We validate our approach on two image datasets, a text dataset, and a tabular dataset. Overall, we observe that, across most datasets, larger query sets lead to improved retrieval performance. The improvement is most pronounced when increasing the query sets from 1 to 8, while the gains become smaller beyond that.

URL PDF HTML 收藏
2607.04819 2026-07-14 cs.LG cs.CR 版本更新

Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers

层并行推理减少了Transformer中加密的非线性深度

Ligong Han, Kai Xu, Hao Wang, Ruijiang Gao, Han Gao, Akash Srivastava

机构 * MBZUAI IFM(穆罕默德·本·扎耶德人工智能大学智能未来研究院) Red Hat AI Innovation(红帽人工智能创新实验室) MIT-IBM Watson AI Lab(麻省理工学院-IBM沃森人工智能实验室) Core AI, IBM(IBM核心人工智能部门) University of Texas at Dallas(德克萨斯大学达拉斯分校)

AI总结 研究结构化牛顿层并行性(SNLP)能否使Transformer层间组合更适合全同态加密(FHE),通过基于切比雪夫多项式近似的模拟框架测量误差积累,结果表明SNLP可减少推理步骤并降低误差放大。

Comments Code is available at https://github.com/phymhan/nanochat-snlp/tree/snlp-fhe

详情
AI中文摘要

全同态加密(FHE)实现加密数据计算,实用的加密Transformer推理受非线性块顺序组合瓶颈限制。研究SNLP能否使层间组合更FHE友好,通过模拟框架测量8个模型和4个架构系列顺序与SNLP推理下的误差积累,结果显示SNLP有优势,且softmax近似主导误差预算。

英文摘要

Fully homomorphic encryption (FHE) enables computation on encrypted data, but practical encrypted Transformer inference is bottlenecked by the sequential composition of many nonlinear blocks. We study whether Structured Newton Layer Parallelism (SNLP) can make this inter-layer composition more FHE-friendly: each Transformer block still requires polynomial approximations for operations such as softmax and RMSNorm, but SNLP reduces the layerwise sequential nonlinear depth from L stages to a small number of solver iterations plus linear structured corrections. Using a simulation framework based on Chebyshev polynomial approximations, we measure error accumulation under sequential versus SNLP inference across 8 models and 4 architecture families. On a 0.5B IDN-trained model, SNLP reduces symbolic bootstraps from 53 to 20 (2.65x) with only +1.2% perplexity degradation, while lowering error amplification (1.36x vs. 1.42x). Across all tested models, SNLP has lower amplification than sequential inference. Ablations show that softmax approximation dominates the error budget and CKKS arithmetic noise is negligible in our setting, suggesting that SNLP is complementary to block-level FHE-friendly operator design rather than a replacement for it.

URL PDF HTML 收藏
2606.27537 2026-07-14 cs.CV 版本更新

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments

MemoBench: 动态变化环境中的世界建模基准测试

Haoyu Chen, Kaichen Zhou, Hang Hua, Kaile Zhang, Jingwen Qian, Wufei Ma, Haonan Chen, Chunjiang Liu, Yizhou Zhao, Xiaoyuan Wang, Weiyue Li, Alan Yuille, Paul Pu Liang, Yilun Du

机构 * Harvard University(哈佛大学) MIT(麻省理工学院) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室) Boston University(波士顿大学) Google(谷歌) JHU(约翰霍普金斯大学) CMU(卡内基梅隆大学) Kempner Institute(肯普纳研究所)

AI总结 提出MemoBench基准,通过目标消失-重现范式评估视频生成模型在动态环境中的记忆一致性,涵盖合成与真实场景,揭示现有模型的关键挑战。

详情
AI中文摘要

视频生成模型旨在模拟动态环境,已有多个基准测试评估帧间的记忆一致性。然而,大多数基准仅在目标保持在视野内时评估一致性,少数迫使目标离开视野的基准则评估遮挡期间无变化的静态场景。为弥补这一差距,我们引入了MemoBench,这是一个围绕动态变化环境中消失-重现范式构建的诊断基准:目标对象经历物理过程,从视野中消失,并必须在重新出现时以更新后的状态正确恢复。我们整理了涵盖合成和真实场景的360个真实剪辑,并设计了一个评估套件,结合自动指标和基于VQA的评估,涵盖四个诊断支柱。对八个最先进模型的评估揭示了在消失-重现范式下关于记忆一致性的关键见解和开放挑战。

英文摘要

Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consistency across frames. However, most assess consistency only while the target remains in view, and the few that force objects out of view evaluate static scenes where nothing changes during occlusion. To bridge this gap, we introduce MemoBench, a diagnostic benchmark built around the disappear-and-reappear paradigm in dynamically changing environments: a target object undergoes a physical process, disappears from view, and must be correctly recovered in its updated state upon reappearance. We curate 360 ground-truth clips spanning synthetic and real-world scenes, and design an evaluation suite combining automated metrics with VQA-based assessment across four diagnostic pillars. Evaluation of eight state-of-the-art models reveals key insights and open challenges regarding memory consistency under the disappear-and-reappear paradigm.

URL PDF HTML 收藏
2602.17665 2026-07-14 cs.CV 版本更新

OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents

OpenEarthAgent:一个用于工具增强地理空间代理的统一框架

Akashah Shabbir, Muhammad Umer Sheikh, Muhammad Akhtar Munir, Hiyam Debary, Mustansar Fiaz, Muhammad Zaigham Zaheer, Paolo Fraccaro, Fahad Shahbaz Khan, Muhammad Haris Khan, Xiao Xiang Zhu, Salman Khan

机构 * Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学) IBM Research(IBM研究院) Linköping University(林霍姆斯大学) Technical University Munich(慕尼黑技术大学) Australian National University(澳大利亚国立大学)

AI总结 本文提出OpenEarthAgent框架,通过整合卫星影像、自然语言查询和结构化推理轨迹,实现多模态地理空间推理,提升遥感任务的执行能力与空间逻辑一致性。

Comments Accepted at the European Conference on Computer Vision (ECCV 2026)

详情
AI中文摘要

近期多模态推理的进展使能够解释图像、连接语言并执行结构化分析任务的代理成为可能。将这些能力扩展到遥感仍然具有挑战性,因为模型必须在空间尺度、地理结构和多光谱指数上进行推理,同时保持连贯的多步逻辑。为了解决这一差距,我们引入OpenEarthAgent,一个统一的框架,用于工具增强的地理空间推理,训练于卫星影像、自然语言查询和结构化推理轨迹。除了作为基准之外,OpenEarthAgent建立了一个围绕统一可执行工具注册表和轨迹基政策学习的统一代理架构。该框架在一致的可调用模式下标准化了异构视觉、光谱、GIS和地理参考栅格操作,使模块化编排和确定性执行成为可能。通过在结构化推理轨迹上进行监督微调,并通过确定性回放验证确保可执行性和空间正确性进行训练。配套的数据集包含14,538个训练实例和1,169个评估实例,超过107,000个推理步骤,涵盖城市、环境、灾害和基础设施领域,并结合GIS操作以及如NDVI、NBR和NDBI等指数分析。基于显式的推理轨迹,学习的代理在多样化的遥感场景中表现出结构化的推理、稳定的空问理解以及可解释的工具驱动行为。我们报告了相对于强基线的一致改进,并在最近的开源和闭源模型中表现出竞争力。我们的代码和训练模型将公开发布。

英文摘要

Recent progress in multimodal reasoning has enabled agents that interpret imagery, connect it with language, and execute structured analytical tasks. Extending these capabilities to remote sensing remains challenging, as models must reason over spatial scale, geographic structures, and multispectral indices while maintaining coherent multi-step logic. To address this gap, we introduce \textit{OpenEarthAgent}, a unified framework for tool-augmented geospatial reasoning trained on satellite imagery, natural-language queries, and structured reasoning traces. Beyond serving as a benchmark, OpenEarthAgent establishes a cohesive agentic architecture built around a unified executable tool registry and trajectory-based policy learning. The framework standardizes heterogeneous visual, spectral, GIS, and georeferenced raster operations under a consistent callable schema, enabling modular orchestration and deterministic execution. Training is performed via supervised fine-tuning on structured reasoning trajectories with deterministic replay validation to ensure executability and spatial correctness. The accompanying corpus comprises 14,538 training and 1,169 evaluation instances with over 107K reasoning steps, spanning urban, environmental, disaster, and infrastructure domains and incorporating GIS operations alongside index analyses such as NDVI, NBR, and NDBI. Grounded in explicit reasoning traces, the learned agent demonstrates structured reasoning, stable spatial understanding, and interpretable tool-driven behaviour across diverse EO scenarios. We report consistent improvements over a strong baseline and competitive performance against recent open and closed-source models. Our code, data and trained models are publicly available: https://github.com/mbzuai-oryx/OpenEarthAgent

URL PDF HTML 收藏
2405.00392 2026-07-14 cs.CR cs.AI

Certified Adversarial Robustness of Machine Learning-based Malware Detectors via (De)Randomized Smoothing

通过(去)随机化平滑实现基于机器学习的恶意软件检测器的认证对抗鲁棒性

Daniel Gibert, Luca Demetrio, Giulio Zizzo, Quan Le, Jordi Planes, Battista Biggio

机构 * CeADAR, University College Dublin(CeADAR,都柏林大学学院) University of Genova(热那亚大学) IBM Research Europe(IBM欧洲研究) University of Lleida(莱里达大学) University of Cagliari(卡利亚里大学)

AI总结 本文提出了一种基于(去)随机化平滑的认证对抗鲁棒性防御方法,通过块划分和多数投票机制提升恶意软件检测器的鲁棒性。

详情
AI中文摘要

基于深度学习的恶意软件检测系统容易受到对抗性示例的攻击——精心设计的恶意程序通过最小扰动逃避检测。因此,社区致力于开发防御对抗性示例的机制。然而,当前基于随机化平滑的防御仍然容易受到注入块状对抗性内容的攻击。在本文中,我们提出了一种可证明的防御方法,保证对于给定的可执行文件和对抗性补丁大小,不存在对抗性示例。我们的方法受到(去)随机化平滑的启发,后者提供了确定性鲁棒性证书。在训练过程中,基础分类器使用连续字节的子集进行训练。在推理阶段,我们的防御将可执行文件分成非重叠的块,独立分类每个块,并通过多数投票计算最终预测,以最小化注入内容的影响。此外,我们引入了一个预处理步骤,将节和头部的大小固定为块大小的整数倍。因此,注入的内容被限制在整数个块内,而不会篡改包含输入示例真实字节的其他块,使我们能够将我们的认证鲁棒性保证扩展到内容插入攻击。我们进行了广泛的消融研究,通过将我们的防御与基于随机化平滑的防御对比,针对各种内容篡改攻击和神经网络架构进行比较。结果表明,我们的方法在强内容插入攻击面前表现出前所未有的鲁棒性,优于文献中的基于随机化平滑的防御。

英文摘要

Deep learning-based malware detection systems are vulnerable to adversarial EXEmples - carefully-crafted malicious programs that evade detection with minimal perturbation. As such, the community is dedicating effort to develop mechanisms to defend against adversarial EXEmples. However, current randomized smoothing-based defenses are still vulnerable to attacks that inject blocks of adversarial content. In this paper, we introduce a certifiable defense against patch attacks that guarantees, for a given executable and an adversarial patch size, no adversarial EXEmple exist. Our method is inspired by (de)randomized smoothing which provides deterministic robustness certificates. During training, a base classifier is trained using subsets of continguous bytes. At inference time, our defense splits the executable into non-overlapping chunks, classifies each chunk independently, and computes the final prediction through majority voting to minimize the influence of injected content. Furthermore, we introduce a preprocessing step that fixes the size of the sections and headers to a multiple of the chunk size. As a consequence, the injected content is confined to an integer number of chunks without tampering the other chunks containing the real bytes of the input examples, allowing us to extend our certified robustness guarantees to content insertion attacks. We perform an extensive ablation study, by comparing our defense with randomized smoothing-based defenses against a plethora of content manipulation attacks and neural network architectures. Results show that our method exhibits unmatched robustness against strong content-insertion attacks, outperforming randomized smoothing-based defenses in the literature.

URL PDF HTML 收藏
1708.04326 2026-07-14 cs.IR cs.CL 交叉投稿

Improved Answer Selection with Pre-Trained Word Embeddings

使用预训练词嵌入改进答案选择

Rishav Chakravarti, Jiri Navratil, Cicero Nogueira dos Santos

机构 * IBM Watson AI Foundations, IBM Research(IBM研究院人工智能基础)

AI总结 研究基于预训练词嵌入的答案选择方法,通过在公开数据集上实验,对比传统方法有显著提升,将词嵌入特征与传统排序学习技术结合,性能可媲美最先进神经网络。

详情
AI中文摘要

本文评估了基于预训练词嵌入的现有和新提出的答案选择方法。词嵌入在各种自然语言处理任务中非常有效,将其集成到传统信息检索(IR)系统中可以捕捉问题和答案之间的语义相关性。在三个公开数据集上的实证结果表明,在有监督和无监督设置下,相对于传统基于词频的方法都有显著提升。我们表明,将这些词嵌入特征与传统排序学习技术相结合,可以达到与为答案选择任务训练的最先进神经网络相似的性能。

英文摘要

This paper evaluates existing and newly proposed answer selection methods based on pre-trained word embeddings. Word embeddings are highly effective in various natural language processing tasks and their integration into traditional information retrieval (IR) systems allows for the capture of semantic relatedness between questions and answers. Empirical results on three publicly available data sets show significant gains over traditional term frequency based approaches in both supervised and unsupervised settings. We show that combining these word embedding features with traditional learning-to-rank techniques can achieve similar performance to state-of-the-art neural networks trained for the answer selection task.

URL PDF HTML 收藏
2607.09094 2026-07-13 cs.CL cs.AI 新提交

PRecG: Legal Precedent Retrieval with Graph Neural Networks and Rhetorical Role Segmentation

PRecG:基于图神经网络和修辞角色分割的法律先例检索

Devanshu Verma, Vasudha Bhatnagar, Vikas Kumar, Balaji Ganesan

机构 * University of Delhi(德里大学) IBM Research(IBM 研究院)

AI总结 研究法律先例检索问题,提出PRecG管道,通过基于句子修辞角色分解文档、构建知识图、学习聚合实体上下文表示等步骤,分层学习法律判决对的表示来计算相似度,经实验验证其有效性。

Comments 23 Pages

详情
AI中文摘要

法律先例检索是法律案件准备、规划、诉讼策略和法律研究中的一项基本任务。当前自动先例检索方法将法律文件映射到低维语义空间并基于表示的接近度计算相似度,忽略了法律技术细节的修辞组织,从而忽略细微法律含义,无法区分法律实体和概念基于文档中修辞角色的上下文意义。为解决这一不足,我们提出PRecG管道,通过分层学习法律判决对的表示来计算相似度。首先基于句子修辞角色将文档分解为不同语义单元,为每个修辞段构建知识图以捕获其中法律实体及其关系,学习并聚合实体的上下文表示以获得段级嵌入,进一步整合这些嵌入以生成统一的文档级表示,最后计算文档对之间的语义相似度。我们在印度法律基准数据集上进行广泛实验验证了该方法的性能,并与现有基线进行比较以证明其有效性。

英文摘要

Legal precedent retrieval is a fundamental task in legal case preparation, planning, litigation strategy, and legal research. Current approaches for automatic precedent retrieval map legal documents to a low-dimensional semantic space and compute similarity based on the proximity of their representations. These approaches treat legal documents as monolithic texts, ignoring the rhetorical organization of the legal technicalities. Ergo, they overlook nuanced legal meanings and fail to distinguish the contextual significance of legal entities and concepts that vary based on their rhetorical roles within the document. To address this insufficiency, we propose the PRecG pipeline that computes the similarity between pairs of legal judgments by hierarchically learning their representations. The process begins by decomposing each document into distinct semantic units (segments) based on the rhetorical roles of sentences. For each rhetorical segment, a knowledge graph is constructed to capture the legal entities and their relationships within the segment. Contextual representations of the entities are then learned and aggregated to derive segment-level embeddings. These embeddings are further integrated to produce a unified document-level representation, and finally, the semantic similarity between a pair of documents is computed. We validate the performance of the proposed approach through extensive experiments on a benchmark Indian legal dataset, comparing it against state-of-the-art baselines to demonstrate its effectiveness.

URL PDF HTML 收藏
2607.09042 2026-07-13 cs.LG 新提交

Learning More from Less: Reinforcement Learning from Hindsight

从更少中学习更多:事后诸葛亮式强化学习

Iris Xu, Sunshine Jiang, John Marangola, Nitish Dashora, Richard Li, Thomas Liu, Zexue He, Yuheng Zhi, Alex Pentland, Pulkit Agrawal, Zhang-Wei Hong

机构 * Massachusetts Institute of Technology(麻省理工学院) MIT-IBM Computing Research Lab(麻省理工学院-IBM计算研究实验室) Stanford University(斯坦福大学) University of California, San Diego(加利福尼亚大学圣地亚哥分校)

AI总结 研究针对强化学习用于视觉-语言-动作模型后期训练时样本效率低的问题,提出事后诸葛亮式学习方法,通过对失败轨迹重标记,让策略联合原始与重标记轨迹训练,在分布外任务上显著提升样本效率并优于基线。

详情
AI中文摘要

强化学习(RL)越来越多地用于视觉-语言-动作(VLA)模型的后期训练,但每次更新都需要消耗机器人的轨迹,而这些轨迹收集起来缓慢且成本高昂,因此样本效率成为核心问题。操纵任务通常只提供稀疏奖励,导致弱策略在训练早期几乎每次轨迹都失败且几乎无学习价值。我们提出了事后诸葛亮式学习(LfH),通过根据失败轨迹实际达成的任务对其进行评分,将事后重标记引入VLA的RL后期训练。单个视觉-语言模型对指令和奖励进行重标记,为一组失败轨迹提出事后诸葛亮式指令并评分,策略在重标记和原始轨迹上联合训练。在分布外的LIBERO-PRO任务上,标准RL进展缓慢,而LfH实现了样本效率5倍的提升,并优于密集进展奖励基线。这些增益在不同VLA主干和实体Franka机器人上均成立。

英文摘要

Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, but every update consumes robot rollouts that are slow and costly to collect, making sample efficiency a central concern. Manipulation tasks typically provide only sparse rewards, so a weak policy fails almost every rollout early in training and has little to learn from, even when those failures execute coherent behavior. Such a failure, however, is a success at a different task. We present Learning from Hindsight (LfH), which brings hindsight relabeling to RL post-training of VLAs by scoring failed rollouts against the tasks they actually achieved. A single vision-language model relabels both the instruction and the reward, proposing a hindsight instruction for a group of failed rollouts and scoring how well each satisfies it, and the policy trains on the relabeled and original rollouts jointly. Because VLAs generalize across language, relabeling in language lets the policy learn more from the same trajectories. On out-of-distribution LIBERO-PRO tasks, where standard RL improves only slowly, LfH achieves $5\times$ improvement in sample efficiency, and outperforms a dense progress-reward baseline. The gains hold across VLA backbones and on a physical Franka robot.

URL PDF HTML 收藏
2607.08837 2026-07-13 cs.LG cs.AI 新提交

Prompt-Driven Exploration

提示驱动的探索

Sunshine Jiang, John Marangola, David Zhang, Raghuram Kowdeed, Ruiyang Luo, Nitish Dashora, Richard Li, Pulkit Agrawal, Zhang-Wei Hong

机构 * Massachusetts Institute of Technology(麻省理工学院) MIT-IBM Computing Research Lab(麻省理工学院-IBM计算研究实验室) Improbable AI Lab(英普罗巴布尔人工智能实验室)

AI总结 研究针对强化学习中探索难的问题,提出提示驱动的探索策略(PDE),利用视觉-语言模型对轨迹视频推理,从弱策略轨迹中优化提示,实现后验采样,能让强化学习从零奖励起步学习成功策略并提升样本效率。

详情
AI中文摘要

探索对于强化学习至关重要,因为策略无法通过重复采样其偏好行为来改进。标准方法在动作空间中注入随机性,但这种抖动只能产生接近原始的轨迹。摆脱弱策略通常需要全局扰动,而动作噪声无法产生。大语言模型和视觉-语言-动作模型提供了一条途径:它们根据自然语言提示来调整策略,修改提示会引发全局变化。挑战在于找到能引发有用全局变化的提示。对于很少成功的弱策略,奖励过于稀疏难以选择。我们的想法是从轨迹本身优化提示:视觉-语言模型对轨迹视频进行推理,诊断策略的响应并重写提示以引出更好的行为。此过程在提示层面实现了后验采样,这是一个经典的强化学习探索框架。我们将此策略称为提示驱动的探索(PDE)。在操纵和推理任务中,PDE使强化学习即使从零奖励开始也能学习到成功的策略,并更广泛地提高样本效率。

英文摘要

Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inject stochasticity in the action space, but such jitter only yields rollouts close to the original. Escaping a weak policy often requires global perturbations that action noise cannot produce. Large language models (LLMs) and vision-language-action (VLA) models offer a pathway: they condition the policy on a natural language prompt, and since the rollout follows from it, modifying the prompt induces global changes. The challenge is finding prompts that induce useful global changes. With a weak policy that rarely succeeds, reward is too sparse to select on. Our idea is to refine prompts from the rollouts themselves: a vision-language model (VLM) reasons over the rollout video, diagnoses how the policy responded, and rewrites the prompt to elicit better behavior next time. This procedure realizes posterior sampling, a classical RL exploration framework, at the level of prompts: the VLM maintains an implicit distribution over useful prompts and updates it from observed rollouts. We call this strategy Prompt-Driven Exploration (PDE). Across manipulation and reasoning tasks, PDE enables RL to learn successful policies even from zero-reward starts, and improves sample efficiency more broadly. Our website is available at https://xinyunsunshine.github.io/prompt-rl.

URL PDF HTML 收藏
2607.07047 2026-07-13 cs.CL cs.AI 新提交

Riemannian Geometry for Pre-trained Language Model Embeddings

预训练语言模型嵌入的黎曼几何

Szczepan Konior, Alexandre Quemy, Przemysław Klocek, Bartłomiej Sobieski, Grégoire Cattan

机构 * IBM Automation and AI(IBM自动化与人工智能) Hother(霍瑟) University of Warsaw(华沙大学) Centre for Credible AI, Warsaw University of Technology(华沙理工大学可信人工智能中心)

AI总结 研究预训练语言模型嵌入的几何结构,提出黎曼均值池化方法,通过提取回拉度量并聚合,在三个数据集上RMP优于欧几里得均值池化,消融实验定位增益来源,训练好的编码器在特定数据集贡献额外信号。

详情
AI中文摘要

理解预训练语言模型嵌入的几何结构对可解释性和安全性很重要。我们研究句子级分类信号是否存在于上下文词元嵌入的黎曼几何中,并通过从学习到的编码器的解析雅可比矩阵中提取每个词元的回拉度量,并在对称正定(SPD)流形上用弗雷歇均值聚合它们来进行探究,我们将此过程称为黎曼均值池化(RMP)。在三个具有非平凡语言结构的数据集(CoLA、CREAK、RTE)上,RMP优于欧几里得均值池化,而在为消除注释驱动的词汇工件而构建的基准FEVER - Symmetric上,该方法正确地保持在随机水平。消融实验表明,随机初始化的编码器与弗雷歇聚合相结合,在三个有信号的数据集中的两个上已经超过了欧几里得池化,将增益来源定位到几何聚合而不是学习到的流形结构;训练好的编码器在CREAK(三个有信号的数据集中知识量最大的)上特别贡献了额外信号。

英文摘要

Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sentence-level classification signal lives in the Riemannian geometry of contextual token embeddings, and probe it by extracting per-token pullback metrics from a learned encoder's analytical Jacobian and aggregating them with the Fréchet mean on the symmetric positive definite (SPD) manifold; we call this procedure Riemannian Mean Pooling (RMP). Across three datasets with non-trivial linguistic structure (CoLA, CREAK, RTE), RMP outperforms Euclidean mean pooling, while on FEVER-Symmetric, a benchmark constructed to remove annotation-driven lexical artifacts, the method correctly stays at chance. Ablations show that a randomly initialised encoder combined with Fréchet aggregation already beats Euclidean pooling on two of the three signal-bearing datasets, localising the source of the gain to the geometric aggregation rather than to learned manifold structure; the trained encoder contributes additional signal specifically on CREAK, the most knowledge-heavy of the three signal-bearing datasets.

URL PDF HTML 收藏
2607.08522 2026-07-10 cs.LG 新提交

Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data

停止猜测何时停止测试:用足够的数据进行高效模型评估

Ofir Arviv, Kristjan Greenewald, Yotam Perlitz, Hadar Mulian, Michal Shmueli-Scheuer, Leshem Choshen

机构 * IBM Research(IBM研究院)

AI总结 研究指出固定大小基准用于模型评估效率低,提出自适应评估框架,结合序贯测试统计范式与定制停止标准,在Open VLM Leaderboard上展示能自适应管理效率与可靠性权衡,降低计算成本并保持统计显著性。

详情
AI中文摘要

固定大小基准的固有局限性使其成为低效的模型评估工具。包括模型排名、模型选择以及整个开发过程中的测试等不同评估目标,需要不同水平的统计能力。固定样本量与这些不同需求之间的不匹配导致计算成本过高或可靠性受损,这是模型评估的关键问题。为克服这些限制,我们呼吁在该领域采用序贯测试。我们提供了一个自适应评估框架,在模型评估中提供了一种在效率和可靠性之间进行权衡的原则方法。我们的框架将既定的序贯测试统计范式与针对常见评估需求(如收益递减检测和最小可检测效应大小)定制的停止标准相结合。我们在Open VLM Leaderboard上展示了其自适应管理效率 - 可靠性权衡的能力,例如,与固定大小评估相比,计算成本降低了80%(置信区间宽度允许2.5个点),同时保持统计显著性。

英文摘要

The inherent rigidity of fixed-size benchmarks makes them an inefficient tool for model evaluation. Diverse evaluation objectives, including model ranking, model selection and testing throughout development, demand varying levels of statistical power. The mismatch between fixed sample sizes and these diverse needs results in either excessive computational cost or compromised reliability - a critical concern for model evaluation. To overcome these limitations, we call for adoption of sequential testing in our field. We provide an adaptive evaluation framework, that provides a principled way to navigate the trade-off between efficiency and reliability in model evaluation. Our framework combines the established statistical paradigm of sequential testing with stopping criteria tailored to common evaluation needs such as diminishing returns detection, and minimum detectable effect size. We demonstrate its ability to adaptively manage the efficiency-reliability trade-off on the Open VLM Leaderboard, including, for example, a 80% reduction in computational cost compared to fixed-size evaluation (with a 2.5-point CI width allowance) while maintaining statistical significance.

URL PDF HTML 收藏
2508.11847 2026-07-10 stat.ML cs.LG 版本更新

Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings

丢弃少量偏好可以改变大型语言模型的排名

Jenny Y. Huang, Yunyi Shen, Dennis Wei, Tamara Broderick

机构 * Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology(麻省理工学院电子工程与计算机科学系) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室) IBM Research(IBM研究院)

AI总结 研究通过评估LLM排名系统的鲁棒性,发现丢弃少量偏好数据可显著影响模型排名,揭示MT-bench偏好更稳定的原因。

详情
AI中文摘要

我们提出了一种评估广泛使用的LLM排名系统——布拉德利-蒂尔模型变种——对丢弃最坏情况下的极小比例偏好数据的鲁棒性的方法。我们的方法计算快速且易于采用。当我们将其应用于流行的LLM排名平台的对决时,包括Chatbot Arena及其衍生平台,我们发现顶级模型的排名对丢弃一小部分偏好数据极为敏感;例如,丢弃仅0.003%的人类偏好数据即可改变Chatbot Arena上的顶级模型。我们的鲁棒性检查识别出最负责此类排名翻转的具体偏好,允许对这些影响性偏好进行检查。我们观察到,来自MT-bench偏好的排名比Chatbot Arena的排名明显更鲁棒,这可能归因于MT-bench使用专家标注员和精心设计的提示。最后,我们发现基于众包人类评估的排名和基于LLM-as-a-judge偏好的排名在系统敏感性上并无系统性差异。

英文摘要

We propose a method for evaluating the robustness of widely used LLM ranking systems -- variants of a Bradley--Terry model -- to dropping a worst-case very small fraction of preference data. Our approach is computationally fast and easy to adopt. When we apply our method to matchups from popular LLM ranking platforms, including Chatbot Arena and derivatives, we find that the rankings of top-performing models can be remarkably sensitive to the removal of a small fraction of preferences; for instance, dropping just 0.003% of human preferences can change the top-ranked model on Chatbot Arena. Our robustness check identifies the specific preferences most responsible for such ranking flips, allowing for inspection of these influential preferences. We observe that the rankings derived from MT-bench preferences are notably more robust than those from Chatbot Arena, likely due to MT-bench's use of expert annotators and carefully constructed prompts. Finally, we find that neither rankings based on crowdsourced human evaluations nor those based on LLM-as-a-judge preferences are systematically more sensitive than the other.

URL PDF HTML 收藏
2607.06608 2026-07-09 cs.CR cs.AI cs.HC 新提交

Security and Privacy in Agentic AI: Grand Challenges and Future Directions

智能体人工智能中的安全与隐私:重大挑战与未来方向

Adam Jenkins, Agnieszka Kitkowska, Caterina Maidhof, Diego Paracuellos, Francesco Sovrano, Gonzalo Gabriel Mendez, Guillermo Suarez-Tangil, Hana Kopecka, Isabel Wagner, Isabel Barbera, Javier Carnerero-Cano, Jide Edu, Jose Luis Martin-Navarro, Jose Such, Josep Domingo-Ferrer, Juan Carlos Carrillo, Kopo Marvin Ramokapane, Mark Cote, Pablo Vellosillo, Ramon Ruiz-Dolz, Rongjun Ma, Ruba Abu-Salma, Sameer Patil, William Seymour, Xiao Zhan

机构 * King’s College London(伦敦国王学院) Jönköping University(琼斯平大学) Universitat Politècnica de València(巴塞罗那理工大学) Inria(法国国家信息与自动化技术研究所) IBM Research(IBM研究院) Aalto University(阿尔托大学) INGENIO (CSIC–Universitat Politècnica de València)(INGENIO(西班牙国家科研委员会-巴塞罗那理工大学)) IMDEA Networks(IMDEA网络研究所) University of Basel(巴塞尔大学) Dutch Data Protection Authority (AP)(荷兰数据保护局) University of Strathclyde(斯特拉斯克莱德大学) Universitat Rovira i Virgili(罗维拉-维古利大学) University of Bristol(布里斯托大学) University of Dundee(邓迪大学)

AI总结 探讨智能体人工智能安全与隐私关键挑战及未来方向,通过汇聚30位国际专家的视野扫描活动,针对人工智能增长带来的新风险展开讨论与协作。

详情
AI中文摘要

我们基于一项视野扫描活动,提出了智能体人工智能安全与隐私方面的关键挑战和未来研究方向。该活动汇聚了来自学术界、 industry和政府的30位国际顶尖专家,就与人工智能日益增长的智能相关的新出现风险进行了深入讨论和协作。

英文摘要

We present key challenges and future research directions in the security and privacy of agentic AI, based on a horizon-scanning exercise that brought together thirty leading international experts from academia, industry, and government to engage in focused discussions and collaborative exercises on the emerging risks associated with the growing agency of AI.

URL PDF HTML 收藏
2607.06760 2026-07-09 cs.AI quant-ph 新提交

QANTIS: Hardware-Calibrated Sequential POMDP Belief Updates on IBM Heron

QANTIS:IBM Heron上的硬件校准顺序POMDP信念更新

Bayram Yuksel Eker, Suayb S. Arslan, Ozgur Nazli, Mustafa Serhat Demirgil, Furkan Deligoz

机构 * IBM(国际商业机器公司)

AI总结 研究在IBM Heron硬件上,QANTIS能否跨顺序Tiger POMDP重复使用校准信念更新服务。通过案例研究比较不同放大方式,结果表明全步FPAA能保留后验,给出了硬件校准信念更新原语的操作包络。

Comments 10 pages, 6 figures

详情
AI中文摘要

部分可观测性下的自主系统基于信念而非原始传感器事件采取行动。QANTIS将量子处理器视为该循环中的校准信念更新服务:接收先验和观测模型,估计罕见事件证据项,并向经典规划器返回普通后验。本文探讨在当前IBM Heron硬件上,该服务能否在顺序Tiger POMDP范围内重复使用而不影响面向规划器的后验。通过可控硬件案例研究给出答案。研究比较了同一轨迹上的无放大、保护Grover放大和全步定点放大,检查返回的后验是否会改变下游行动。全步FPAA在报告的8步和12步主要运行中保留Tiger后验,20步和32步控制保持在同一操作带内。在每个报告的决策检查中,硬件后验和精确贝叶斯后验选择相同的即时行动。边界感知BIQAE在接近零和接近一时稳定幅度估计,罕见事件扫描绘制百万分之一证据的逻辑样本复杂度包络。结果是硬件校准信念更新原语的操作包络,而非独立的硬件优势声明。

英文摘要

Autonomous systems under partial observability act on beliefs, not raw sensor events. QANTIS treats the quantum processor as a calibrated belief-update service in that loop: it receives a prior and an observation model, estimates the rare-event evidence term, and returns an ordinary posterior to a classical planner. This paper asks whether that service can be reused across a sequential Tiger POMDP horizon on present IBM Heron hardware without corrupting the planner-facing posterior. We answer with a controlled hardware case study rather than an end-to-end autonomy or wall-clock speedup claim. The study compares no amplification, guarded Grover amplification, and all-step fixed-point amplification on the same trajectory, then checks whether the returned posterior would change the downstream action. All-step FPAA preserves the Tiger posterior across the reported 8-step and 12-step primary runs, and the 20-step and 32-step controls remain inside the same operating band. In every reported decision check, the hardware posterior and the exact Bayes posterior select the same immediate action. Boundary-aware BIQAE stabilizes amplitude estimation near zero and near one, while a rare-event sweep maps the logical sample-complexity envelope for one-in-a-million evidence. The result is an operating envelope for a hardware-calibrated belief-update primitive, not a standalone hardware-advantage claim.

URL PDF HTML 收藏
2605.17842 2026-07-09 cs.LG 版本更新

SNLP: Layer-Parallel Inference via Structured Newton Corrections

SNLP:通过结构化牛顿校正的层并行推理

Ligong Han, Kai Xu, Hao Wang, Akash Srivastava

机构 * Core AI, IBM(IBM核心AI)

AI总结 提出结构化牛顿层并行(SNLP)框架,通过将Transformer层间依赖视为非线性残差方程并用结构化牛顿校正并行求解,实现推理加速,在0.5B模型上获得高达2.58倍加速。

Comments Project webpage: https://github.com/phymhan/nanochat-snlp

详情
AI中文摘要

自回归语言模型顺序执行Transformer层,造成传统张量或流水线并行无法消除的延迟瓶颈。我们研究是否可以通过将跨层的隐藏状态轨迹视为非线性残差方程的解,并用并行牛顿风格更新来求解,从而放松这种逐层依赖。虽然这一观点在原理上是合理的,但精确的牛顿校正需要昂贵的雅可比向量积,而朴素的固定点迭代在训练好的Transformer上不稳定。我们引入了结构化牛顿层并行(SNLP),一个训练和推理框架,用廉价的架构诱导替代动力学替换精确的层雅可比。在残差Transformer中,这产生了恒等牛顿(IDN),其中校正简化为前缀和式更新;在mHC风格架构中,HC牛顿(HCN)使用模型的残差混合矩阵。我们还研究了SNLP感知训练,包括预训练正则化和直接SNLP前向SFT。在Nanochat规模的Transformer上的实验表明,SNLP揭示了一个实用的速度-质量边界:在0.5B模型上,它实现了高达2.58倍的时钟加速,而一个较不激进的配置在不增加PPL的情况下实现了1.40倍加速。这种有用的权衡来自于IDN/HCN引入的有偏有限迭代计算,而不是精确恢复顺序轨迹。我们进一步表明,SNLP前向SFT可以保持下游任务准确性,并且SNLP可以作为自推测解码的草稿模型,而顺序验证器保持输出正确性。

英文摘要

Autoregressive language models execute Transformer layers sequentially, creating a latency bottleneck that is not removed by conventional tensor or pipeline parallelism. We study whether this layerwise dependency can be relaxed by treating the hidden-state trace across layers as the solution of a nonlinear residual equation and solving it with parallel Newton-style updates. While this view is principled, exact Newton corrections require expensive Jacobian-vector products and naive fixed-point iterations are unstable on trained Transformers. We introduce Structured Newton Layer Parallelism (SNLP), a training and inference framework that replaces exact layer Jacobians with cheap architecture-induced surrogate dynamics. In residual Transformers, this yields Identity Newton (IDN), where the correction reduces to a prefix-sum-like update; in mHC-style architectures, HC Newton (HCN) uses the model's residual mixing matrix. We also study SNLP-aware training, including pretraining regularization and direct SNLP-forward SFT. Experiments on Nanochat-scale Transformers show that SNLP exposes a practical speed-quality frontier: on 0.5B models, it reaches up to 2.58x wall-clock speedup, and a less aggressive configuration reaches 1.40x speedup without increasing PPL. The useful tradeoff comes from the biased finite-iteration computation induced by IDN/HCN rather than exact recovery of the sequential trace. We further show that SNLP-forward SFT can preserve downstream task accuracy, and that SNLP can serve as a drafter for self-speculative decoding while a sequential verifier preserves output correctness.

URL PDF HTML 收藏
2510.06505 2026-07-08 cs.LG cs.AI math.OC stat.ML 版本更新

Medix: Out-of-Distribution Detection from Unlabeled Wild Data via Robust Gradient Statistics

Medix:通过稳健梯度统计从未标记的野生数据中进行分布外检测

Momin Abbas, Ali Falahati, Hossein Goli, Mohammad Mohammadi Amiri

机构 * IBM(IBM公司) University of Waterloo(多伦多大学) CUHK(香港中文大学) Rensselaer Polytechnic Institute(罗切斯特理工学院)

AI总结 研究利用未标记野生数据进行分布外检测的问题,提出Medix框架,通过基于中位数的稳健梯度统计识别异常值,结合标记InD数据训练分类器,理论推导与实证结果表明该方法全面优于现有方法。

Comments Accepted to TMLR. Camera-ready version

详情
AI中文摘要

分布外(OOD)检测对于确保机器学习系统在实际应用中的鲁棒性至关重要。近期方法探索利用未标记数据增强OOD检测能力,但因分布内(InD)和OOD样本混合,有效利用未标记的野生数据仍具挑战。本文介绍Medix框架,利用基于中位数的稳健梯度统计从未标记数据中识别潜在异常值,结合标记的InD数据训练稳健OOD分类器。理论上推导了误差界,实证结果表明Medix在开放世界设置中全面优于现有方法。

英文摘要

Out-of-distribution (OOD) detection plays a crucial role in ensuring the robustness of machine learning systems deployed in real-world applications. Recent approaches have explored the use of unlabeled data, showing potential for enhancing OOD detection capabilities. However, effectively utilizing unlabeled in-the-wild data remains challenging due to the mixed nature of both in-distribution (InD) and OOD samples. The lack of a distinct set of OOD samples complicates the task of training an optimal OOD classifier. In this work, we introduce Medix, a novel framework designed to identify potential outliers from unlabeled data using the median-based robust gradient statistics. We use the median because it provides a stable estimate of the central tendency, as an OOD detection mechanism, due to its robustness against noise and outliers. Using these identified outliers, along with labeled InD data, we train a robust OOD classifier. From a theoretical perspective, we derive error bounds that demonstrate Medix achieves a low error rate. Empirical results further substantiate our claims, as Medix outperforms existing methods across the board in open-world settings.

URL PDF HTML 收藏
2607.04505 2026-07-07 cs.AI 新提交

Why Pure Reasoning is Not Enough: Nature as the Source of Mathematical Innovation

为何纯推理不足:自然作为数学创新之源

Charanjit S. Jutla, Vimal Sharma

机构 * IBM T. J. Watson Research Center(IBM T. J. 沃森研究中心)

AI总结 研究认为人类数学推理受逻辑片段不可判定性和计算难题限制,依赖自然世界的模式匹配。通过追溯傅里叶变换等数学史及逻辑复杂性,表明物理启发的模式匹配是认知必要,为人工智能中语言模型规模大提供依据。

详情
AI中文摘要

我们提出假说,即受适度逻辑片段的不可判定性和计算难题限制,人类数学推理从根本上依赖纯演绎之外领域的模式匹配。自然世界是此类模式的丰富来源,其物理定律和生物系统已历经数十亿年“预计算”。为证实这一观点,我们追溯了傅里叶变换及相关数学的历史……我们认为这些障碍使受物理启发的模式匹配不仅是历史偶然,更是认知必需。最后,我们得出对人工智能的启示:若纯推理本质上不足,那么任何旨在实现人类水平数学创造力的系统都必须嵌入大量跨领域模式,而非仅依赖演绎。这为当代大型语言模型的巨大规模提供了原则性依据。

英文摘要

We advance the hypothesis that human mathematical reasoning, constrained by both the undecidability and the computational intractability of even modest logical fragments, relies fundamentally on pattern matching from domains external to pure deduction. The most prolific reservoir of such patterns is the natural world, whose physical laws and biological systems have undergone billions of years of ``pre-computation'' and already exhibit surprisingly innovative solutions. To ground this claim, we trace the history of the Fourier transform and relevant mathematics, from the vibrating string controversy to the hear equation and subsequent formalisms prevalent in mathematics. At each critical juncture, a physics problem forced the acceptance or creation of a mathematical tool that pure formal reasoning failed to anticipate or, worse, human reasoning had resisted. We further survey the landscape of logical complexity, from NP-hard propositional satisfiability to the non-elementary decision-procedures for monadic second-order theories, to demonstrate that even when a logic is decidable, the resources required for worst-case deduction are astronomically prohibitive. We argue that these barriers make physics-inspired pattern matching not just a historical accident but a cognitive necessity. Finally, we draw the consequence for artificial intelligence: if pure reasoning is constitutively insufficient, then any system aiming at human-level mathematical creativity must embed a vast store of cross-domain patterns rather than rely on deduction alone. This furnishes a principled justification for the enormous scale of contemporary large language models.

URL PDF HTML 收藏
2607.02854 2026-07-07 cs.SE cs.LG 新提交

EvoOtter: Evolutionary Reproduction Test Generator

EvoOtter:进化式重现测试生成器

Toufique Ahmed, Jatin Ganhotra, Avraham Shinnar, Martin Hirzel

机构 * IBM

AI总结 研究如何生成错误重现测试,此前方法成本高且反馈不可靠。新方法EvoOtter通过连续减半控制测试执行成本,用批交叉等控制大语言模型成本,以新适应度分数生成高质量测试,成本低。

详情
AI中文摘要

在修复问题之前,通过生成错误重现测试(BRT)来重现问题很有用。然而,生成BRT本身具有挑战性,因为问题描述往往不正式。先前工作通过推理扩展来解决此问题,但成本高且反馈不可靠。本文探索用于BRT生成的进化编程,以强化反馈并控制成本,新方法EvoOtter能以低得多的成本生成高质量BRT。

英文摘要

Before fixing an issue, it is useful to first reproduce it by generating a bug reproduction test (BRT). However, generating a BRT is itself a challenging task, because issue descriptions tend to be informal, making it difficult to determine whether a candidate BRT indeed fails for the reason in the issue. Prior work has attempted to tackle this problem via inference scaling, using large language models to generate many BRTs and patches, then using execution feedback to select and improve them. Unfortunately, this is expensive and the feedback is unreliable. This paper explores evolutionary programming for BRT generation to sharpen the feedback, while enhancing evolutionary programming to keep costs in check. Our new approach, EvoOtter, controls test execution costs via successive halving. Furthermore, it controls LLM costs via batched crossover for an entire generation in a single LLM call, as well as via rule-based code mutations, with a new fitness score tailored for BRTs. As a result, EvoOtter generates state-of-the-art quality BRTs at the fraction of the cost of prior inference-scaling approaches to this problem. More broadly, this paper points at how to efficiently and effectively combine evolutionary programming with large language models for software engineering.

URL PDF HTML 收藏
2607.02116 2026-07-07 cs.AI 新提交

ContextNest: Verifiable Context Governance for Autonomous AI Agent

ContextNest:自主AI智能体的可验证上下文治理

Misha Sulpovar, Benn R. Konsynski, Qaish Kanchwala, Gabe Goodhart

机构 * PromptOwl, LLC(PromptOwl公司) Goizueta Business School, Emory University(埃默里大学戈伊苏埃塔商学院) IBM Research(IBM研究院)

AI总结 提出ContextNest,一种在检索前提供上下文治理的开放规范,通过版本链、哈希校验和审计追踪确保知识可验证性,实验表明其在防过时攻击和检索确定性上优于传统方法。

Comments 35 pages, 11 tables, 4 figures

详情
AI中文摘要

自主AI智能体越来越依赖外部知识存储,但大多数检索流程仅提供相关性,而不提供关于来源、版本身份、完整性、可追溯性或时间点重建的持久保证。我们将此形式化为上下文治理,并提出了ContextNest,这是一个用于治理的AI可消费知识库的开放规范和参考实现。ContextNest并不取代检索增强生成(RAG);它在检索之下提供治理层,在检索系统操作之前确定哪些工件是批准的、当前的、可归属的和完整性验证的。该规范结合了带元数据的类型化Markdown文档、确定性集合代数选择器、contextnest:// URI引用、SHA-256哈希链版本历史、图级检查点、通过模型上下文协议(MCP)的实时数据源节点,以及智能体上下文消费的审计追踪。这些机制使组织能够重建哪些知识版本影响了智能体输出,以及这些版本在被消费时是否具有AI资格。我们报告了两个受控实验的首个实证结果。在隔离治理与检索失败模式的过时版本攻击中,治理选择严格帕累托支配BM25稀疏检索,在输入令牌成本约三分之一的情况下,答案质量通过率更高(97%对93-90%)。在包含1,060个文档的语料库上的检索确定性实验中,确定性选择器和BM25在重复相同查询时返回稳定的文档集(Jaccard系数1.0),而密集+HNSW基线在80%的查询上是不确定的(平均Jaccard系数0.611,最差情况0.210)。这些结果表明,上下文治理解决了仅靠检索质量无法设计的故障模式。我们在开放许可下发布了核心引擎、CLI和MCP服务器。

英文摘要

Autonomous AI agents increasingly depend on external knowledge stores, yet most retrieval pipelines provide relevance without durable guarantees of provenance, version identity, integrity, traceability, or point-in-time reconstruction. We formalize this as context governance and present ContextNest, an open specification and reference implementation for governed AI-consumable knowledge vaults. ContextNest does not replace Retrieval-Augmented Generation (RAG); it supplies the governance layer beneath retrieval, determining which artifacts are approved, current, attributable, and integrity-verified before retrieval systems operate over them. The specification combines typed Markdown documents with metadata, deterministic set-algebraic selectors, contextnest:// URI references, SHA-256 hash-chained version histories, graph-level checkpoints, source nodes for live data through the Model Context Protocol (MCP), and audit traces of agent context consumption. These mechanisms let organizations reconstruct which knowledge versions informed an agent output and whether those versions were AI-eligible when consumed. We report first empirical results from two controlled experiments. In a stale-version attack isolating the governance-versus-retrieval failure mode, governed selection strictly Pareto-dominates BM25 sparse retrieval, with higher answer-quality pass rate (97% versus 93-90%) at about one-third the input-token cost. In a retrieval-determinism experiment over a 1,060-document corpus, deterministic selectors and BM25 return stable document sets across repeated identical queries (Jaccard 1.0), while a dense+HNSW baseline is non-deterministic on 80% of queries (mean Jaccard 0.611, worst case 0.210). These results suggest that context governance addresses failure modes retrieval quality alone is not designed to resolve. We release a core engine, CLI, and MCP server under open licenses.

URL PDF HTML 收藏