arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Edinburgh(爱丁堡大学)

共收录 816
2608.17933 2026-08-19 cs.AI cs.CE 新提交

EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection

EvoTS-Agent:一种用于金融时间序列变点检测的自进化大语言模型智能体

Lei Jiang, Ye Wei, Xinyu Xi, Jordan Langham-Lopez, Yifan Bao, Raad Khraishi, Yihao Ang, Anthony K. H. Tung, Lukasz Szpruch, Hao Ni

机构 * Alan Turing Institute(阿兰·图灵研究所) University of Oxford(牛津大学) National University of Singapore(新加坡国立大学) NatWest AI Research(国民西敏寺银行人工智能研究部) University of Edinburgh(爱丁堡大学) University College London(伦敦大学学院)

AI总结 该研究针对金融时间序列变点检测难题,提出自进化 LLM 智能体 EvoTS-Agent,经验证引导进化检测流程,在四个基准数据集上表现优于现有同类智能体且执行成功率达100%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17620 2026-08-19 cs.LG 新提交

OOD Detection for EEG-based Machine Learning in High-Risk Environments

高风险环境下基于脑电图(EEG)的机器学习的分布外(OOD)检测

Philipp Bomatter, Henry Gouk

机构 * School of Informatics, University of Edinburgh(爱丁堡大学信息学院)

AI总结 针对脑电图机器学习在高风险环境中易受分布偏移影响的问题,引入EEG OOD检测基准,评估相关方法并结合互补方法构建稳健安全网。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17528 2026-08-19 cs.AI cs.SE 新提交

Agent Lightning v1.0: Towards Harnessed Agentic RL

Agent Lightning v1.0:走向可控的智能体强化学习

Zhiyuan He, Siwei Zhang, Zhiwen Zhou, Yuqing Yang, Yu Kang, Yuge Zhang, Luna K. Qiu, Tin Yan Tsui, Jiahang Xu, Chong Luo

机构 * Microsoft(微软) Fudan University(复旦大学) Zhejiang University(浙江大学) University of Edinburgh(爱丁堡大学)

AI总结 研究针对可控智能体强化学习的挑战,推出轻量级框架 Agent Lightning v1.0,经评估可将 Qwen3.5-9B 在 SWE-bench Verified 上的性能提升14.6个百分点,发布完整流程脚本以推进相关可复现研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31281 2026-08-19 cs.CL 版本更新

Wind Turbine Maintenance Log Labelling Framework: LLM-Driven Data Correction and Enrichment via Semantic Extraction of Reliability Intelligence

风力涡轮机维护日志标注框架:基于LLM驱动的数据校正与语义提取的可靠性智能增强

Max Malyi, Jonathan Shek, Alasdair McDonald, Andre Biscaya

机构 * Institute for Energy Systems, School of Engineering, The University of Edinburgh(能源系统研究所,工程学院,爱丁堡大学) Nadara, Lisbon, Portugal(纳达拉,里斯本,葡萄牙)

AI总结 提出一种利用大语言模型自动标准化和结构化风力涡轮机维护日志的方法,通过纠正系统代码、提取故障模式与维护动作分类,将非结构化文本转化为定量可靠性指标。

Comments An adjustable template containing the Python script architecture, applied dynamic prompts, and data schemas is hosted in an open-source GitHub repository: this https URL (https://github.com/mvmalyi/llm-driven-wind-turbine-maintenance-log-labelling)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24126 2026-08-19 cs.LG cs.PL 版本更新

Likelihood Hacking in Probabilistic Program Synthesis

概率程序合成中的似然黑客行为

Jacek Karwowski, Younesse Kaddar, Zihuiwen Ye, Esmeralda S. Whitammer, Sam Staton

机构 * University of Oxford, Department of Computer Science(牛津大学计算机科学系) University of Edinburgh, School of Informatics(爱丁堡大学信息学院) CIFAR Fellow, Learning in Machines and Brains(CIFAR Fellow, 机器与大脑学习)

AI总结 研究概率程序合成中语言模型通过强化学习生成程序时可能人为夸大边际似然奖励的问题,提出安全语言片段SafeStan以防止似然黑客行为。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16273 2026-08-18 cs.LG cs.AI 新提交

Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic

Foresight-England:面向COVID-19大流行期间医疗事件预测的全国规模电子健康记录生成式AI模型的开发

Simon Ellershaw, Christopher Tomlinson, Zeljko Kraljevic, Spiros Denaxas, Harry Hemingway, Cathie Sudlow, Angela M. Wood, Anoop D. Shah, Richard Dobson

机构 * University College London(伦敦大学学院) King’s College London(伦敦国王学院) University College London Hospitals National Institute for Health Research Biomedical Research Centre(伦敦大学学院医院国家卫生研究院生物医学研究中心) Interdisciplinary Transformation University(跨学科转型大学) British Heart Foundation Data Science Centre(英国心脏基金会数据科学中心) Health Data Research UK(英国健康数据研究中心) The University of Edinburgh(爱丁堡大学) University of Cambridge(剑桥大学) Victor Phillip Dahdaleh Heart and Lung Research Institute, University of Cambridge(剑桥大学维克多·菲利普·达德赫心肺研究所) British Heart Foundation Centre of Research Excellence, University of Cambridge(剑桥大学英国心脏基金会卓越研究中心)

AI总结 本研究开发了首个全国规模的电子健康记录生成式基础模型Foresight-E,在6100万患者数据上训练,可零样本预测医疗事件,为大流行期间的医疗事件预测提供了方法模板。

Comments Methodology and evaluation framework for Foresight-England. As detailed in the Project Status section, NHS England has paused access to data for the Foresight-E project, meaning quantitative results are not currently available. On behalf of the CVD-COVID-UK/COVID-IMPACT Consortium

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26866 2026-08-18 cs.CL cs.LG 版本更新

MoRFI: Monotonic Sparse Autoencoder Feature Identification

MoRFI: 基于单调性的稀疏自编码器特征识别

Dimitris Dimakopoulos, Shay B. Cohen, Ioannis Konstas

机构 * University of Edinburgh, UK(爱丁堡大学) Heriot-Watt University, UK(赫瑞斯泰德大学)

AI总结 本文通过控制微调实验发现,逐步引入新知识会增加幻觉,提出MoRFI方法利用稀疏自编码器分析残差流激活,识别因果相关的潜在特征。

Comments Accepted to the Conference on Language Modeling (COLM) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14435 2026-08-17 cs.CV cs.LG 新提交

Style or Signature? Artist-Disjoint Evaluation of Style Classification in Frozen Vision Embeddings

风格还是签名?冻结视觉嵌入中风格分类的艺术家不相交评估

Rory Ashton

机构 * University of Edinburgh(爱丁堡大学)

AI总结 该研究提出艺术家不相交评估协议,发现CLIP等模型的冻结图像嵌入风格分类准确率在该协议下下降,超现实主义降幅最大,验证了冻结嵌入中风格理解需此类评估。

Comments 15 pages, 3 figures. Accepted at the VISART VIII workshop, ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13684 2026-08-17 cs.AI 新提交

Learning to Assemble Novel Structures with Unfamiliar Parts under Semantic Constraints

在语义约束下学习使用不熟悉部件组装新型结构

Jonghyuk Park, Alex Lascarides, Subramanian Ramamoorthy

机构 * School of Informatics, University of Edinburgh(爱丁堡大学信息学院)

AI总结 该研究提出神经符号架构,在模拟玩具卡车组装领域,借助具身对话与任务演示,通过自然语言传递语义约束提升智能体在线适应的数据效率,解决部署后遇未知语义约束的组装问题。

Comments Accepted to, and to appear in the 20th Conference on Neurosymbolic Learning and Reasoning (NeSy 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11241 2026-08-17 cs.SD 版本更新

The Affective Bridge: Preserving Speech Representations while Enhancing Deepfake Detection vian emotional Constraints

情感桥梁:通过情感约束在增强深度伪造检测的同时保留语音表征

Yupei Li, Chenyang Lyu, Longyue Wang, Weihua Luo, Kaifu Zhang, Björn W. Schuller

机构 * University of Bristol(布里斯托大学) University of Edinburgh(爱丁堡大学) University of Technology Sydney(新南威尔士大学)

AI总结 提出仅用情感识别微调语音编码器,再训练轻量SVM进行深度伪造检测,既保留下游任务表征能力又提升检测性能,发现情感是独特的桥梁任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22366 2026-08-14 cs.CL

Exploratory Semantic Reliability Analysis of Wind Turbine Maintenance Logs using Large Language Models

Max Malyi, Jonathan Shek, Andre Biscaya

机构 * Institute for Energy Systems, School of Engineering, The University of Edinburgh(能源系统研究所,工程学院,爱丁堡大学) Nadara, Lisbon, Portugal(纳达拉,里斯本,葡萄牙)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18064 2026-08-13 cs.AI 版本更新

Towards Human Motion World Models via Executable Behaviour Representations

通过可执行模型理解人类行为

Rimvydas Rubavicius, Manisha Dubey, N. Siddharth, Subramanian Ramamoorthy

机构 * School of Informatics The University of Edinburgh(信息学院爱丁堡大学)

AI总结 本文提出EXACT语言,通过可执行神经符号模型分析人类动作,提升动作分割和异常检测的效率与直观性。

Comments Accepted in ECCV2026 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03818 2026-08-13 cs.LG cs.AI cs.PL 版本更新

Program Semantic Inequivalence Game with Large Language Models

基于大语言模型的程序语义不等价博弈

Antonio Valerio Miceli-Barone, Vaishak Belle, Ali Payani

机构 * University of Edinburgh(爱丁堡大学) Cisco Systems(思科系统)

AI总结 本研究提出基于语义不等价博弈(SInQ)的方法,通过生成器与评估器智能体半对抗训练合成代码推理数据,在跨语言漏洞检测等基准上显著提升LLMs的代码语义理解能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10743 2026-08-12 cs.CL 新提交

Mitigating Context Interference for Reliable and Efficient Search Agents

缓解可靠高效搜索智能体的上下文干扰

Boyang Xue, Bin Wu, Shuofei Qiao, Sheng Wang, Rui Wang, Yiming Du, Hongru Wang, Jeff Z. Pan, Emine Yilmaz, Kam-Fai Wong, Aldo Lipani

机构 * The Chinese University of Hong Kong(香港中文大学) University College London(伦敦大学学院) Zhejiang University(浙江大学) The University of Hong Kong(香港大学) The University of Edinburgh(爱丁堡大学)

AI总结 本文针对多轮搜索智能体的上下文干扰问题,提出基于蒸馏的上下文优化器,将其纳入RL训练后可显著提升搜索智能体的可靠性与效率,开创了AI智能体“先优化上下文再生成”的新范式。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10330 2026-08-12 cs.AI 新提交

Hierarchical Compositionality for An Assistive AI Agent

面向辅助AI智能体的分层组合性

Tianyi Fu, Mohan Sridharan

机构 * University of Edinburgh(爱丁堡大学)

AI总结 本文针对辅助AI智能体的歧义问题,提出嵌入分层组合性原则的架构,结合语义兼容性等模型推理实现歧义消除,实验表明其性能优于当前最优数据驱动基线,可适配特定用户画像。

Comments 25 pages, 9 figures, 4 tables. Project page: https://tianyi-fu.github.io/HCAA

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08392 2026-08-11 cs.AI cs.CL 新提交

CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception

CAP:用于评估具备复杂动作与感知能力的跨站点浏览器智能体的可扩展基准

Zejun Xu, Taiyi Chen, Jin Li, Yongtong Gu, Qi Cheng, Aixuan Lv, Shuai Zhu, Pengfei Zhu, Kaichen Yang, Boyu Sun, Yixian Yang, Mulong Xie, Xin Liu, Dagang Li, Xiaoteng Ma, Hongru Wang

机构 * Macau University of Science and Technology(澳门科技大学) Tsinghua University(清华大学) Southeast University(东南大学) FellouAI ARGUS Lab(ARGUS实验室) The University of Edinburgh(爱丁堡大学)

AI总结 研究人员推出可扩展基准CAP,通过分解-重组流程构建420项跨站点网页任务,实验发现当前浏览器智能体在感知密集型交互上存在明显瓶颈,与真实需求差距较大。

Comments Accepted to COLM 2026. Project page: https://warriorxu0302.github.io/CAP-Bench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10453 2026-08-11 cs.CL cs.AI 版本更新

Reasoning about Intent for Ambiguous Requests

意图推理以应对歧义请求

Irina Saparina, Mirella Lapata

机构 * School of Informatics, University of Edinburgh(爱丁堡大学信息学院)

AI总结 本文提出生成结构化响应以枚举歧义请求的不同解释,通过强化学习训练模型提升覆盖有效解释的召回率和抑制虚假解释的精确率,实验表明方法在覆盖有效答案方面优于基线方法。

Comments COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23016 2026-08-11 cs.LG cs.AI 版本更新

A Sobering Look at Tabular Data Generation via Probabilistic Circuits

对通过概率电路生成表格数据的冷静审视

Davide Scassola, Dylan Ponsford, Adrián Javaloy, Sebastiano Saccani, Luca Bortolussi, Henry Gouk, Antonio Vergari

机构 * School of Informatics University of Edinburgh(信息学院爱丁堡大学) Aindo SpA AREA Science Park(Aindo SpA 面向科学公园) AI lab University of Trieste(人工智能实验室特里este大学)

AI总结 本文质疑表格数据生成的进展观念,指出当前评估方法的局限性,并展示概率电路在成本效益上优于现有最先进的模型,同时揭示SotA模型进展的饱和可能源于不充分的指标。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02984 2026-08-11 hep-lat cs.LG 版本更新

Variance reduction in lattice QCD observables via normalizing flows

通过归一化流减少晶格QCD可观测量的方差

Ryan Abbott, Denis Boyda, Yang Fu, Daniel C. Hackett, Gurtej Kanwar, Fernando Romero-López, Phiala E. Shanahan, Julian M. Urban

机构 * Center for Theoretical Physics, Massachusetts Institute of Technology, Cambridge, MA 02139, USA The NSF AI Institute for Artificial Intelligence Fermi National Accelerator Laboratory, Batavia, IL 60510, U.S.A. Higgs Centre for Theoretical Physics, School of Physics Astronomy, University of Edinburgh, EH9 3FD Edinburgh, United Kingdom Albert Einstein Center, Institute for Theoretical Physics, University of Bern, 3012 Bern, Switzerland Physics Department, Columbia University, New York, NY 10027, USA

AI总结 本文通过归一化流方法有效降低晶格QCD中胶球相关函数和强子结构相关胶子矩阵元的方差,同时减少计算成本。

Comments 15 pages, 4 figures, 2 tables. v2: update to match published version

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11554 2026-08-11 cs.RO cs.CV cs.LG 版本更新

HyperDet: 3D Object Detection with Hyper 4D Radar Point Clouds

HyperDet: 基于超4D雷达点云的3D目标检测

Yichun Xiao, Runwei Guan, Jin Jin, Fangqiang Ding

机构 * University of Edinburgh(爱丁堡大学) HKUST (GZ)(香港科技大学(广州)) University of Oxford(牛津大学) MIT(麻省理工学院)

AI总结 提出一种与检测器无关的框架HyperDet,通过构建任务感知的超4D雷达点云,利用时空累积、跨传感器验证和多普勒引导的运动补偿以及前景生成增强,显著提升仅用雷达的3D目标检测性能。

Comments 11 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17108 2026-08-11 cs.LG cs.AI eess.SP 版本更新

Hybrid Mamba-Attention Neural Architecture for Channel Estimation

MambaNet:结合注意力机制的Mamba辅助信道估计神经网络

Dianxin Luan, Chengsi Liang, Jie Huang, Zheng Lin, Kaitao Meng, John Thompson, Cheng-Xiang Wang, Ozgur Akan

机构 * Institute for Imaging, Data and Communications, School of Engineering, University of Edinburgh(影像、数据与通信研究所,工程学院,爱丁堡大学) James Watt School of Engineering, University of Glasgow(詹姆斯·瓦特工程学院,格拉斯哥大学) National Mobile Communications Research Laboratory, Southeast University(国家移动通信研究中心,东南大学) Purple Mountain Laboratories, Nanjing(紫金山实验室,南京) Department of Electrical and Electronic Engineering, The University of Hong Kong(电气与电子工程系,香港大学) Department of Electrical and Electronic Engineering, University of Manchester(电气与电子工程系,曼彻斯特大学)

AI总结 MambaNet通过结合自注意力机制和定制化Mamba架构,实现低复杂度的OFDM信道估计,尤其在大规模子载波配置中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03020 2026-08-10 cs.AI 版本更新

LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment

LoCA:基于局部信用分配的一次性校准后仅前向的大语言模型调优

Linhan Xia, Rui Liu, Zhaofeng Zhang, Yihao Wang, Binrui Shen, Shengxin Zhu

机构 * University of Oklahoma(俄克拉荷马大学) Imperial College London(伦敦帝国学院) University of Michigan(密歇根大学) Tencent(腾讯) University of Edinburgh(爱丁堡大学) University of Southern California(南加州大学) Beijing Normal University(北京师范大学) Beijing Normal-Hong Kong Baptist University(北京师范大学-香港浸会大学联合国际学院)

AI总结 本文提出LoCA方法,通过一次性校准替换大语言模型调优的重复反向传播,在多个基准上优于LoRA,降低了GPU峰值内存、CPU稳态内存与前向传递时间。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06283 2026-08-07 cs.LG math.OC math.PR stat.ML 新提交

The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity

超越凸性的驯服次梯度未校正朗之万算法

Iosif Lytras, Nikolaos Makras, Sotirios Sabanis

机构 * University of Edinburgh(爱丁堡大学) National Technical University of Athens(雅典国立技术大学) Athena/Archimedes Research Centre(雅典娜/阿基米德研究中心)

AI总结 本文针对非光滑、超线性梯度增长且非凸的目标分布采样问题,提出SG-TULA算法,推导其非渐近收敛界,验证假设并用于GPT-2系列LLM预训练,效果优于AdamW等。

Comments 53 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06139 2026-08-07 cs.SD physics.comp-ph 新提交

Explicit and Stable Pseudospectral Time-Domain Method for the Föppl-von Kármán Equations

用于冯·卡门(Föppl-von Kármán)方程的显式稳定伪谱时域方法

Victor Zheleznov, Stefan Bilbao

机构 * University of Edinburgh(爱丁堡大学) IRCAM(法国声学/音乐研究与协作学院) CNRS(法国国家科学研究中心) Sorbonne Université(索邦大学)

AI总结 本研究针对冯·卡门方程,提出一种在空间域计算乘积项、模态域计算导数的伪谱方法,结合标量辅助变量技术实现显式稳定时间积分,降低了模态合成的计算成本并保留其频率控制优势。

Comments To be presented at Forum Acusticum 2026, Graz, Austria, September 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14538 2026-08-07 cs.AI cs.LG 版本更新

Symbol Grounding in Neuro-Symbolic AI: A Gentle Introduction to Reasoning Shortcuts

神经符号AI中的符号接地:推理快捷方式的入门介绍

Emanuele Marconato, Samuele Bortolotti, Emile van Krieken, Paolo Morettin, Elena Umili, Antonio Vergari, Efthymia Tsamoura, Andrea Passerini, Stefano Teso

机构 * University of Trento(特伦托大学) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) Sapienza University of Rome(罗马大学) University of Edinburgh(爱丁堡大学) Huawei Labs(华为实验室)

AI总结 本文探讨神经符号AI中推理快捷方式的问题,分析其成因与影响,并提供解决方法与策略,以提升模型的可靠性和可信度。

Comments Published on JAIR (Integration of Logical Constraints in Deep Learning special track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04928 2026-08-06 cs.CL 新提交

Does Out-of-Sight Equal Out-of-Mind in CoT Monitorability?

思维链可监控性中“看不见即想不到”吗?

Pedro Ferreira, Wilker Aziz, Ivan Titov

机构 * University of Amsterdam(阿姆斯特丹大学) University of Edinburgh(爱丁堡大学)

AI总结 本研究对比显式与潜在CoT的可监控性,发现其更多取决于任务属性和模型内部访问程度,而非推理模式。

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04285 2026-08-06 cs.AI cs.LG 新提交

The RAIL Principles for Neurosymbolic AI: Reasoning, Assurances, Interfacing and Learning

神经符号AI的RAIL原则:推理(Reasoning)、保证(Assurances)、接口(Interfacing)与学习(Learning)

Agnese Chiatti, Michael Cochez, Cristina Cornelio, Sebastijan Dumancic, Artur d'Avila Garcez, Luis C. Lamb, Lia Morra, Mathias Niepert, Robert Peharz, Alberto Speranzon, Maarten Stol, Annette Ten Teije, Thiviyan Thanapalasingam, Frank Van Harmelen, Emile Van Krieken, Antonio Vergari, Benjie Wang

机构 * Politecnico di Milano(米兰理工大学) ELLIS Institute Finland(芬兰ELLIS研究所) Åbo Akademi University(奥博 Akademi 大学) Samsung AI(三星人工智能研究院) Delft University of Technology(代尔夫特理工大学) Stony Brook University(石溪大学) Politecnico di Torino(都灵理工大学) University of Stuttgart(斯图加特大学) Graz University of Technology(格拉茨工业大学) Lockheed Martin, Advanced Technology Labs(洛克希德·马丁公司先进技术实验室) BrainCreators(BrainCreators公司) Vrije Universiteit Amsterdam(阿姆斯特丹自由大学) University of Amsterdam(阿姆斯特丹大学) University of Edinburgh(爱丁堡大学) UCLA(加利福尼亚大学洛杉矶分校)

AI总结 该文提出神经符号AI的RAIL四项原则,可统一分析多类AI系统,助力工程师更科学地设计部署生产级AI,指导整合神经符号方法到下一代AI技术。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14294 2026-08-06 cs.CV cs.AI cs.LG cs.RO 版本更新

Seeking Physics in Diffusion Noise

在扩散噪声中寻求物理规律

Chujun Tang, Lei Zhong, Fangqiang Ding

机构 * Brown University(布朗大学) University of Edinburgh(爱丁堡大学) MIT(麻省理工学院)

AI总结 研究通过分析预训练扩散变换器的中间去噪表示,发现物理合理与不合理视频在中层特征空间部分可分离,提出渐进轨迹选择策略提升物理一致性并降低推理成本。

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05547 2026-08-06 cs.CL cs.AI cs.LG 版本更新

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

多任务GRPO:跨任务的可靠大语言模型推理

Shyam Sundhar Ramesh, Xiaotong Ji, Matthieu Zimmer, Sangwoong Yoon, Zhiyong Wang, Haitham Bou Ammar, Aurelien Lucchi, Ilija Bogunovic

机构 * UCL Department of EEE(伦敦大学学院电子工程系) UCL Centre for AI(伦敦大学学院人工智能中心) Huawei Noah’s Ark Lab(华为诺亚实验室) UNIST Graduate School of AI(延世大学人工智能研究生院) University of Edinburgh(爱丁堡大学) University of Basel(巴塞尔大学)

AI总结 本文提出MT-GRPO算法,通过动态调整任务权重和比例保持采样器,提升多任务场景下大语言模型的可靠推理性能。

Comments Accepted at ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02005 2026-08-04 cs.AI 新提交

Evolving in the Agent Jungle via History-Informed Opponent Awareness

在智能体丛林中通过历史感知的对手意识进化

Zhaofeng Zhang, Linhan Xia, Rui Liu, Yihao Wang, Binrui Shen, Shengxin Zhu

机构 * University of Edinburgh(爱丁堡大学) University of Oklahoma(俄克拉荷马大学) Imperial College London(伦敦帝国学院) University of Michigan(密歇根大学) University of Southern California(南加州大学) Tencent(腾讯) Beijing Normal University(北京师范大学) Beijing Normal–Hong Kong Baptist University(北京师范大学-香港浸会大学联合国际学院)

AI总结 针对多智能体环境中对手策略持续进化导致静态技能修改方法失效的问题,提出OASE方法,通过历史快照锚定的配对比较选择有益技能修改,在两类场景中实现更低均衡距离与更少无效策略变更。

详情

展开后加载摘要…

URL PDF HTML 收藏