arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Nanyang Technological University(南洋理工大学)

2026-01-13 至 2026-01-13 共收录 15
2601.07821 2026-01-13 cs.RO cs.AI cs.LG

Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation

具有自恢复能力的失败感知强化学习:用于现实世界操控的可靠离线到在线强化学习

Huanyu Li, Kun Lei, Sheng Zang, Kaizhe Hu, Yongyuan Liang, Bo An, Xiaoli Li, Huazhe Xu

机构 * Shanghai Qi Zhi Institute(上海启智研究院) Shanghai Jiao Tong University(上海交通大学) IIIS, Tsinghua University(清华大学人工智能研究院) Nanyang Technological University(南洋理工大学) A*STAR Institute for Infocomm Research(新加坡科技动力研究院) University of Maryland(马里兰大学)

AI总结 本研究提出FARL框架,通过整合安全批评者和恢复策略,有效减少现实世界强化学习中的失败并提升性能。

Comments Project page: https://failure-aware-rl.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07454 2026-01-13 cs.RO

WaveMan: mmWave-Based Room-Scale Human Interaction Perception for Humanoid Robots

WaveMan: 基于毫米波的房间级人形机器人人机交互感知

Yuxuan Hu, Kuangji Zuo, Boyu Ma, Shihao Li, Zhaoyang Xia, Feng Xu, Jianfei Yang

机构 * Key Laboratory for Information Science of Electromagnetic Waves, Ministry of Education, School of Information Science and Technology, Fudan University(电磁波信息科学重点实验室、教育部、信息科学与技术学院,复旦大学) School of Mechanical and Aerospace Engineering, Nanyang Technological University(机械与航空航天工程学院,南洋理工大学)

AI总结 WaveMan通过空间自适应感知系统实现房间级人形机器人人机交互的可靠隐私保护,实验显示其在不同用户位置下的交互感知准确率显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07331 2026-01-13 cs.SD cs.LG

SEE: Signal Embedding Energy for Quantifying Noise Interference in Large Audio Language Models

SEE:信号嵌入能量用于量化大型音频语言模型中的噪声干扰

Yuanhe Zhang, Jiayu Tian, Yibo Zhang, Shilinlu Yan, Liang Lin, Zhenhong Zhou, Li Sun, Sen Su

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) North China Electric Power University(华北电力大学) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Nanyang Technological University(新加坡南洋理工大学) Chongqing University of Posts and Telecommunications(重庆邮电大学)

AI总结 本文提出SEE方法,用于量化LALMs中的噪声干扰,通过结构化激活子空间提升噪声感知,实验显示SEE与性能高度相关,传统去噪方法效果有限,提出基于SEE的改进策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07237 2026-01-13 eess.AS cs.SD

The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge

2026年ICASSP自动歌曲美学评估挑战

Guobin Ma, Yuxuan Xia, Jixun Yao, Huixin Xue, Hexin Liu, Shuai Wang, Hao Liu, Lei Xie

机构 * Shanghai Conservatory of Music, Shanghai, China(上海音乐学院,中国上海) College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院,南洋理工大学,新加坡) Nanjing University, Suzhou, China(南京大学,中国苏州)

AI总结 2026年ICASSP挑战通过预测AI生成歌曲的美学评分,推动了音乐生成系统与人类审美偏好的对齐方法发展。

Comments Official summary paper for the ICASSP 2026 ASAE Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06981 2026-01-13 cs.SD eess.AS eess.SP

Directional Selective Fixed-Filter Active Noise Control Based on a Convolutional Neural Network in Reverberant Environments

基于卷积神经网络的定向选择性固定滤波主动降噪:在混响环境中的应用

Boxiang Wang, Zhengding Luo, Haowen Li, Dongyuan Shi, Junwei Ji, Ziyi Yang, Woon-Seng Gan

机构 * School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(南洋理工大学电子与电气工程学院) Center of Intelligent Acoustics and Immersive Communications, Northwestern Polytechnical University, China(西北工业大学智能声学与沉浸式通信中心)

AI总结 本文提出基于卷积神经网络的定向选择性固定滤波主动降噪方法,用于在混响环境中提高噪声抑制效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06573 2026-01-13 cs.AI cs.MM

QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models

QMAVIS:利用大多模态模型融合实现长视频音频理解

Zixing Lin, Jiale Wang, Gee Wah Ng, Lee Onn Mak, Chan Zhi Yang Jeriel, Jun Yang Lee, Yaohao Li

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University, Singapore(南洋理工大学)

AI总结 QMAVIS通过融合大型多模态模型、大型语言模型和语音识别模型,实现了长视频音频理解,展示了在VideoMME数据集上38.75%的性能提升,并在其他数据集上表现出竞争力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00920 2026-01-13 cs.LG cs.AI

MODE: Efficient Time Series Prediction with Mamba Enhanced by Low-Rank Neural ODEs

MODE:基于低秩神经ODEs增强的高效时间序列预测

Xingsheng Chen, Regina Zhang, Bo Gao, Xingwei He, Xiaofeng Liu, Pietro Lio, Kwok-Yan Lam, Siu-Ming Yiu

机构 * School of Computing and Data Science, The University of Hong Kong(计算与数据科学学院,香港大学) Department of Computing and Data Science, Nanyang Technological University(计算与数据科学系,南洋理工大学) School of Information Engineering, Beijing Institute of Graphic Communication(信息工程学院,北京印刷学院) Department of Computing and Data Science, The University of Hong Kong(计算与数据科学系,香港大学) Yale university(耶鲁大学) University of Cambridge(剑桥大学) Nanyang technological university(南洋理工大学) University of Hong Kong(香港大学)

AI总结 MODE通过整合低秩神经ODEs与增强Mamba架构,实现了高效的时间序列预测,提升了准确性和可扩展性。

Comments 12 pages, 6 figures, and 3 tables. Updated description and explanations, and correct some typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05623 2026-01-13 cs.SE cs.AI cs.CL

Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation

以部署为中心的基础设施即代码生成:通过LLM赋能的DevOps模拟实现失败、学习、细化和成功

Tianyi Zhang, Shidong Pan, Zejun Zhang, Zhenchang Xing, Xiaoyu Sun

机构 * Australian National University Canberra Australia New York University \& Columbia University USA Nanyang Technological University Singapore Australian National University Australia Australian National University New York University \& Columbia University Nanyang Technological University

AI总结 本文提出IaCGen框架,通过迭代反馈机制提升IaC模板的部署性,实验表明其在10次迭代内可使模板部署成功率提升至91.6%,并进一步通过人工反馈将性能提升至90%以上。

Comments Accepted by FSE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09930 2026-01-13 cs.CL

Rethinking Prompt Optimizers: From Prompt Merits to Optimization

重新思考提示优化器:从提示优点到优化

Zixiao Zhu, Hanzhang Zhou, Zijian Feng, Tianjiao Li, Chua Jia Jim Deryl, Mak Lee Onn, Gee Wah Ng, Kezhi Mao

机构 * School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore(南洋理工大学电子与电气工程学院) Home Team Science and Technology Agency, Singapore(新加坡科技局)

AI总结 MePO通过显式和可解释的设计优化提示,提升响应质量,适用于多种任务和模型类型。

Comments 30 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20031 2026-01-13 cs.RO cs.CV

MG-SLAM: Structure Gaussian Splatting SLAM with Manhattan World Hypothesis

MG-SLAM:基于曼哈顿世界假设的结构高斯点撒SLAM

Shuhong Liu, Tianchen Deng, Heng Zhou, Liuzhuozheng Li, Hongyu Wang, Danwei Wang, Mingrui Li

机构 * Department of Information Science and Technology and Department of Complexity Science and Engineering, The University of Tokyo(信息科学与技术系和复杂科学与工程系,东京大学) Institute of Medical Robotics and Department of Automation, Shanghai Jiao Tong University(医疗机器人研究所和自动化系,上海交通大学) Department of Mechanical Engineering, Columbia University(机械工程系,哥伦比亚大学) School of Electrical and Electronic Engineering, Nanyang Technological University(电子与电气工程学院,南洋理工大学) Department of Computer Science, Dalian University of Technology(计算机科学系,大连理工大学)

AI总结 MG-SLAM基于曼哈顿世界假设,通过融合线段和平面假设提升室内场景重建的几何精度和完整性,实现高斯SLAM的先进性能。

Comments IEEE Transactions on Automation Science and Engineering

Journal ref IEEE Transactions on Automation Science and Engineering 22 (2025) 17034-17049

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06566 2026-01-13 cs.CV cs.AI

QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models

QCaption: 通过融合大多模态模型实现视频描述与问答

Jiale Wang, Gee Wah Ng, Lee Onn Mak, Randall Cher, Ng Ding Hei Ryan, Davis Wang

机构 * Department of Computer Science and Engineering, Nanyang Technological University, Singapore(南洋理工大学计算机科学与工程系)

AI总结 QCaption通过融合关键帧提取、多模态模型和语言模型,提升视频描述和问答任务的性能,实现44.2%和48.9%的改进。

Journal ref Proceedings of the 27th International Conference on Information Fusion (FUSION), 2024, pp. 1-8

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06195 2026-01-13 cs.LG cs.AI

EntroLnn: Entropy-Guided Liquid Neural Networks for Operando Refinement of Battery Capacity Fade Trajectories

EntroLnn:基于熵引导的液态神经网络用于在役电池容量衰减轨迹的细化

Wei Li, Wei Zhang, Qingyu Yan

机构 * Singapore Institute of Technology(新加坡理工学院) Nanyang Technological University(南洋理工大学)

AI总结 本研究提出 EntroLnn 框架,利用熵引导的液态神经网络对电池容量衰减轨迹进行在线细化,实现高精度的电池健康预测。

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24629 2026-01-13 eess.AS cs.SD

Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis

词级情感表达控制在零样本文本到语音合成中

Tianrui Wang, Haoyu Wang, Meng Ge, Cheng Gong, Chunyu Qiang, Ziyang Ma, Zikang Huang, Guanrou Yang, Xiaobao Wang, Eng Siong Chng, Xie Chen, Longbiao Wang, Jianwu Dang

机构 * Tianjin Key Laboratory of Cognitive Computing and Application, College of Intelligence and Computing, Tianjin University(天津认知计算与应用重点实验室,智能与计算学院,天津大学) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳)) Nanyang Technological University(南洋理工大学) TeleAI, China Telecom(TeleAI,中国电信) Kuaishou Technology(快手科技) Shanghai Jiao Tong University(上海交通大学) Huiyan Technology (Tianjin)(慧颜科技(天津)) Shenzhen Institute of Advanced Technology(深圳先进技术研究院)

AI总结 本文提出WeSCon框架,通过自我训练实现零样本文本到语音合成中词级情感与语速控制,克服数据稀缺问题,达到最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18992 2026-01-13 cs.CV

VPGS-SLAM: Voxel-based Progressive 3D Gaussian SLAM in Large-Scale Scenes

基于体素的渐进式3D高斯SLAM:用于大规模场景的VPGS-SLAM

Tianchen Deng, Wenhua Wu, Junjie He, Yue Pan, Shenghai Yuan, Danwei Wang, Hesheng Wang

机构 * School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(自动化与智能感知学院,上海交通大学) Key Laboratory of System Control and Information Processing, Ministry of Education(系统控制与信息处理重点实验室,教育部) Thrust of Robotics and Autonomous Systems, The Hong Kong University of Science and Technology (Guangzhou)(机器人与自主系统研究 thrust,香港科技大学(广州)) University of Bonn(波恩大学) School of Electrical and Electronic Engineering, Nanyang Technological University(电子与电气工程学院,南洋理工大学)

AI总结 VPGS-SLAM提出了一种基于体素的渐进式3D高斯SLAM方法,适用于大规模室内外场景,通过多子地图实现紧凑准确的场景表示,并结合2D-3D融合跟踪和回环闭合方法提升鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05748 2026-01-13 cs.AI cs.LG

Low-Dimensional Federated Knowledge Graph Embedding via Knowledge Distillation

低维联邦知识图谱嵌入 via 知识蒸馏

Xiaoxiong Zhang, Zhiwei Zeng, Xin Zhou, Chunyan Miao

机构 * Nanyang Technological University(南洋理工大学)

AI总结 本文提出FedKD,通过知识蒸馏实现低维联邦知识图谱嵌入,缓解教师模型过度自信问题,提升通信效率和训练效果。

详情

展开后加载摘要…

URL PDF HTML 收藏