arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

National University of Singapore(新加坡国立大学)

2026-01-30 至 2026-01-30 共收录 15
2601.21742 2026-01-30 cs.AI cs.CL cs.MA

Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems

知识上下文学习:在基于大语言模型的多智能体系统中以正确方式建立信任

Ruiwen Zhou, Maojia Song, Xiaobao Wu, Sitao Cheng, Xunjian Yin, Yuxi Xie, Zhuoqun Hao, Wenyue Hua, Liangming Pan, Soujanya Poria, Min-Yen Kan

机构 * National University of Singapore(国立新加坡大学) Singapore University of Technology(新加坡科技学院) Nanyang Technological University(南洋理工大学) Duke University(杜克大学) University of Waterloo(滑铁卢大学) Microsoft(微软公司) Peking University(北京大学) University of Pennsylvania(宾夕法尼亚大学)

AI总结 本文提出Epistemic Context Learning(ECL),通过历史交互构建同伴资料以提升多智能体系统中信任建模的准确性,使小型模型在性能上超越大模型,并在多种配置中表现出良好的泛化能力。

Comments Codes and data are available at https://github.com/skyriver-2000/epistemic-context-learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21722 2026-01-30 cs.CL cs.AI

Enhancing Language Models for Robust Greenwashing Detection

增强语言模型以实现稳健的绿色洗脑检测

Neil Heinrich Braun, Keane Ong, Rui Mao, Erik Cambria, Gianmarco Mengaldo

机构 * National University of Singapore(新加坡国立大学) Massachusetts Institute of Technology(麻省理工学院) Nanyang Technological University(南洋理工大学)

AI总结 本文提出了一种参数高效框架,通过结合对比学习和顺序排名目标来增强语言模型在绿色洗脑检测中的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21669 2026-01-30 cs.LG cs.AI

Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability Scaling

预期回报导致强化学习中的结果层面模式崩溃及如何通过逆概率缩放修复它

Abhijeet Sinha, Sundari Elango, Dianbo Liu

机构 * National University of Singapore, Singapore(新加坡国立大学)

AI总结 本文提出逆概率缩放方法,通过修正预期回报目标来解决强化学习中的结果层面模式崩溃问题,有效提升多模式策略优化的可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21617 2026-01-30 cs.CV

PathReasoner-R1: Instilling Structured Reasoning into Pathology Vision-Language Model via Knowledge-Guided Policy Optimization

PathReasoner-R1: 通过知识引导的策略优化在病理视觉-语言模型中引入结构化推理

Songhan Jiang, Fengchun Liu, Ziyue Wang, Linghan Cai, Yongbing Zhang

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Microsoft Research(微软研究院) National University of Singapore(新加坡国立大学) Technical University of Dresden(德累斯顿技术大学)

AI总结 PathReasoner-R1通过知识引导的策略优化,在病理视觉-语言模型中引入结构化推理,提升模型的临床推理能力和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21598 2026-01-30 cs.AI

Beyond Imitation: Reinforcement Learning for Active Latent Planning

超越模仿:用于主动潜在规划的强化学习

Zhi Zheng, Wee Sun Lee

机构 * School of Computing, National University of Singapore, Singapore(计算学院,新加坡国立大学)

AI总结 本文提出ATP-Latent方法,通过主动规划和强化学习优化潜在空间,提升链式推理的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21296 2026-01-30 cs.LG cs.AI

Grounding and Enhancing Informativeness and Utility in Dataset Distillation

基于信息性和效用的蒸馏数据集 grounding

Shaobo Wang, Yantai Yang, Guo Chen, Peiru Li, Kaixin Li, Yufa Zhou, Zhaorun Chen, Linfeng Zhang

机构 * EPIC Lab, SJTU(SJTU实验室) Shanghai Jiao Tong University(上海交通大学) National University of Singapore(新加坡国立大学) Duke University(杜克大学) The University of Chicago(芝加哥大学)

AI总结 本文提出InfoUtil框架,通过博弈论和梯度范数优化,提升数据集蒸馏的信息性和效用,实验显示在ImageNet-1K上性能提升6.1%。

Comments Accepted by ICLR 2026, 20 pages, 9 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21251 2026-01-30 cs.RO

Abstracting Robot Manipulation Skills via Mixture-of-Experts Diffusion Policies

通过混合专家扩散策略抽象机器人操作技能

Ce Hao, Xuanran Zhai, Yaohua Liu, Harold Soh

机构 * School of Computing, National University of Singapore(新加坡国立大学计算机学院) Smart Systems Institute, NUS(NUS智能系统研究所) Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) Beijing Zhongguancun Academy(北京中关村学院)

AI总结 本文提出了一种基于扩散的混合专家策略,通过学习紧凑的技能基础和粘性路由机制,实现高效多任务机器人操作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20882 2026-01-30 cs.SE cs.AI cs.CR

DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle

DevOps-Gym: 在软件DevOps周期中评估AI代理的基准测试

Yuheng Tang, Kaijie Zhu, Bonan Ruan, Chuqi Zhang, Michael Yang, Hongwei Li, Suyue Guo, Tianneng Shi, Zekun Li, Christopher Kruegel, Giovanni Vigna, Dawn Song, William Yang Wang, Lun Wang, Yangruibo Ding, Zhenkai Liang, Wenbo Guo

机构 * UC Santa Barbara(加州大学圣巴巴拉分校) National University of Singapore(新加坡国立大学) UC Berkeley(加州大学伯克利分校) Google(谷歌) UC Los Angeles(加州大学洛杉矶分校)

AI总结 DevOps-Gym通过700+真实任务评估AI代理在完整DevOps周期中的能力,揭示其在问题解决和测试生成方面的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19876 2026-01-30 cs.LG

Real-Time Pulsatile Flow Prediction for Realistic, Diverse Intracranial Aneurysm Morphologies using a Graph Transformer and Steady-Flow Data Augmentation

利用图变换器和稳态流数据增强实现颅内动脉瘤形态的实时脉动流预测

Yiying Sheng, Wenhao Ding, Dylan Roi, Leonard Leong Litt Yeo, Hwa Liang Leo, Choon Hwai Yap

机构 * Department of Biomedical Engineering, National University of Singapore(国立新加坡大学生物医学工程系) Department of Biomedical Engineering, Imperial College London(帝国理工学院伦敦校区生物医学工程系) Imperial College Healthcare NHS Trust(帝国理工医疗信托) Department of Medicine, National University Hospital(国立大学医院医学部)

AI总结 本文提出利用图变换器和稳态流数据增强方法,实现颅内动脉瘤形态的实时脉动流预测,有效提升生物力学标记物的计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18196 2026-01-30 cs.CL

LogicReward: Incentivizing LLM Reasoning via Step-Wise Logical Supervision

LogicReward: 通过逐步逻辑监督激励大语言模型推理

Jundong Xu, Hao Fei, Huichi Zhou, Xin Quan, Qijun Huang, Shengqiong Wu, William Yang Wang, Mong-Li Lee, Wynne Hsu

机构 * National University of Singapore(新加坡国立大学) University College London(伦敦大学学院) University of Manchester(曼彻斯特大学) University of Melbourne(墨尔本大学) University of California, Santa Barbara(加州大学圣巴巴拉分校)

AI总结 LogicReward通过逐步逻辑监督提升大语言模型的推理能力,有效增强推理可信度和泛化能力,实验显示其在自然语言推理和逻辑推理任务中优于GPT-4o和o4-mini。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06411 2026-01-30 cs.AI cs.LG

SofT-GRPO: Surpassing Discrete-Token LLM Reinforcement Learning via Gumbel-Reparameterized Soft-Thinking Policy Optimization

SofT-GRPO:通过Gumbel-重新参数化软思考策略优化超越离散标记LLM强化学习

Zhi Zheng, Yu Gu, Wei Liu, Yee Whye Teh, Wee Sun Lee

机构 * School of Computing, National University of Singapore, Singapore(新加坡国立大学计算机学院) Department of Statistics, University of Oxford, United Kingdom(英国牛津大学统计系) School of Intelligence Science(智能科学学院) Technology, Nanjing University, China(技术学院,南京大学,中国)

AI总结 SofT-GRPO通过引入Gumbel噪声和重新参数化技巧,在软思考模式下提升LLM推理性能,使其在Pass@1和Pass@32任务中分别优于离散标记GRPO。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00527 2026-01-30 cs.SE cs.AI

A Hierarchical Imprecise Probability Approach to Reliability Assessment of Large Language Models

一种分层不精确概率方法用于大型语言模型的可靠性评估

Robab Aghazadeh-Chakherlou, Qing Guo, Siddartha Khastgir, Peter Popov, Xiaoge Zhang, Xingyu Zhao

机构 * WMG, University of Warwick(沃里克大学工业与制造业学院) Center for Frontier AI Research, A*STAR(前沿人工智能研究中心) School of Computing, National University of Singapore(新加坡国立大学计算机学院) Center for Software Reliability, City St George's, University of London(伦敦大学圣乔治学院软件可靠性中心) Department of Industrial and Systems Engineering, The Hong Kong Polytechnic University(香港理工大学工业与系统工程系)

AI总结 本文提出HIP-LLM,一种分层不精确概率框架,用于更准确地评估大型语言模型的可靠性。

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25851 2026-01-30 cs.CV

MuSLR: Multimodal Symbolic Logical Reasoning

MuSLR:多模态符号逻辑推理

Jundong Xu, Hao Fei, Yuhui Zhang, Liangming Pan, Qijun Huang, Qian Liu, Preslav Nakov, Min-Yen Kan, William Yang Wang, Mong-Li Lee, Wynne Hsu

机构 * National University of Singapore(新加坡国立大学) Stanford University(斯坦福大学) Peking University(北京大学) UniMelb(墨尔本大学) University of Auckland(奥克兰大学) MBZUAI(穆斯林人工智能研究所) University of California, Santa Barbara(加州大学圣芭芭拉分校)

AI总结 MuSLR提出了一种多模态符号逻辑推理基准,通过形式逻辑规则提升VLMs的推理能力,显著提升链式推理性能及复杂逻辑处理效果。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11311 2026-01-30 cs.AI cs.CY

Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble

通过紧凑的LLM集合模拟人类偏好:提示到代理

Bingchen Wang, Zi-Yu Khoo, Jingtan Wang

机构 * Independent Researcher, China(中国独立研究者) School of Computing, National University of Singapore(新加坡国立大学计算机学院)

AI总结 通过紧凑的LLM集合模拟人类偏好,P2P方法在无需微调和敏感数据的情况下,有效重建目标人群偏好并实现高质量预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03046 2026-01-30 cs.RO

Real-time Dynamics of Soft Manipulators with Cross-section Inflation: Application to the Octopus Muscular Hydrostat

具有横截面膨胀的软执行器实时动态:应用于章鱼肌肉水压系统

Yuchen Sun, Anup Teejo Mathew, Imran Afgan, Federico Renda, Cecilia Laschi

机构 * National University of Singapore(国立新加坡大学) Khalifa University(卡利法大学) Khalifa University Center for Autonomous Robotics Systems(卡利法大学自主机器人系统中心)

AI总结 本文提出扩展的Cosserat杆理论与降阶数值方法,用于研究具有横截面膨胀的软执行器实时动态,应用于章鱼肌肉水压系统的刚度调节与运动研究。

详情

展开后加载摘要…

URL PDF HTML 收藏