arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

The Hong Kong University of Science and Technology(香港科技大学)

2025-12-16 至 2025-12-16 共收录 9
2512.13297 2025-12-16 cs.AI cs.LG

MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data

MedInsightBench: 通过多步骤洞察发现评估医疗分析代理在多模态医疗数据中的能力

Zhenghao Zhu, Chuxue Cao, Sirui Han, Yuanfeng Song, Xing Chen, Caleb Chen Cao, Yike Guo

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) ByteDance(字节跳动)

AI总结 MedInsightBench通过多步骤洞察发现评估医疗分析代理在多模态医疗数据中的能力,提出MedInsightAgent框架提升医疗数据洞察发现性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12658 2025-12-16 cs.CV

CogDoc: Towards Unified thinking in Documents

CogDoc: 向文档中的统一思维迈进

Qixin Xu, Haozhe Wang, Che Liu, Fangzhen Lin, Wenhu Chen

机构 * Tsinghua University(清华大学) The Hong Kong University of Science and Technology(香港科技大学) University of Waterloo(滑铁卢大学) Imperial College London(伦敦帝国学院)

AI总结 CogDoc提出一种统一的粗到细思维框架,通过直接强化学习提升文档推理性能,优于现有方法并在视觉丰富任务中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12630 2025-12-16 cs.HC cs.AI

ORIBA: Exploring LLM-Driven Role-Play Chatbot as a Creativity Support Tool for Original Character Artists

ORIBA:探索基于大语言模型的对话机器人作为原创角色艺术家创造力支持工具

Yuqian Sun, Xingyu Li, Shunyu Yao, Noura Howell, Tristan Braud, Chang Hee Lee, Ali Asadipour

机构 * Computer Science Research Centre, Royal College of Art(皇家艺术学院计算机科学研究中心) Digital Media, School of Literature, Media, and Communication, Georgia Institute of Technology(佐治亚理工学院数字媒体系) Princeton University(普林斯顿大学) Digital Media, Georgia Institute of Technology(佐治亚理工学院数字媒体系) Division of Integrative Systems and Design, The Hong Kong University of Science and Technology(香港科学大学整合系统与设计 division) Industrial Design Department, College of Engineering, KAIST(韩国科学技术院工程学院工业设计系)

AI总结 ORIBA通过大语言模型支持原创角色艺术家的创意过程,平衡AI辅助与创意自主权。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12367 2025-12-16 physics.optics cs.CV

JPEG-Inspired Cloud-Edge Holography

受JPEG启发的云-边全息术

Shuyang Xie, Jie Zhou, Jun Wang, Renjing Xu

机构 * HKUST(GZ)(香港科技大学(广州)) Sichuan University(四川大学)

AI总结 本文提出了一种基于JPEG的云-边全息术,通过将神经处理转移至云端,利用轻量级解码实现低延迟、高带宽效率的全息图传输。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17364 2025-12-16 cs.CV cs.AI

Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation

条件编织与专家调节:迈向通用且可控的图像生成

Guoqing Zhang, Xingtong Ge, Lu Shi, Xin Zhang, Muqing Xue, Wanru Xu, Yigang Cen, Yidong Li

机构 * State Key Laboratory of Advanced Rail Autonomous Operation(先进轨道交通自主运行国家重点实验室) School of Computer Science and Technology(计算机科学与技术学院) Visual Intellgence +X International Cooperation Joint Laboratory of MOE(教育部视觉智能+X国际合作联合实验室) Hong Kong University of Science and Technology(香港科技大学) SenseTime Research(商汤科技研究院) Beijing Jiaotong University(北京交通大学) SenseTime Research Institute(时光机器研究院)

AI总结 提出UniGen框架,通过CoMoE模块和WeaveNet机制实现通用且可控的图像生成,提升效率和表达性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01950 2025-12-16 cs.RO cs.CV

DualMap: Online Open-Vocabulary Semantic Mapping for Natural Language Navigation in Dynamic Changing Scenes

DualMap: 用于动态变化场景中自然语言导航的在线开放词汇语义映射

Jiajun Jiang, Yiming Zhu, Zirui Wu, Jie Song

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

AI总结 DualMap通过自然语言查询实现动态环境中的高效语义映射与导航,采用混合分割前端和对象级状态检查提升效率,展现先进性能。

Comments 14 pages, 14 figures. Published in IEEE Robotics and Automation Letters (RA-L), 2025. Code: https://github.com/Eku127/DualMap Project page: https://eku127.github.io/DualMap/

Journal ref IEEE Robotics and Automation Letters, Vol. 10, No. 12, pp. 12612-12619, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07293 2025-12-16 cs.CV cs.AI

Fast Wrong-way Cycling Detection in CCTV Videos: Sparse Sampling is All You Need

CCTV视频中快速错误方向骑行检测:稀疏采样足矣

Jing Xu, Wentao Shi, Sheng Ren, Lijuan Zhang, Weikai Yang, Pan Gao, Jie Qin

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) The College of Electronic and Information Engineering, Nanjing University of Aeronautics and Astronautics(南京航空航天大学电子与信息工程学院) The College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院)

AI总结 本文提出WWC-Predictor,通过稀疏采样和轻量级检测器高效估计错误方向骑行比例,实现低误差率和低计算成本。

Comments Accepted by IEEE Transactions on Intelligent Transportation Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12196 2025-12-16 cs.MM cs.CV cs.SD eess.AS

AutoMV: An Automatic Multi-Agent System for Music Video Generation

AutoMV: 一种自动多智能体系统用于音乐视频生成

Xiaoxuan Tang, Xinping Lei, Chaoran Zhu, Shiyun Chen, Ruibin Yuan, Yizhi Li, Changjae Oh, Ge Zhang, Wenhao Huang, Emmanouil Benetos, Yang Liu, Jiaheng Liu, Yinghao Ma

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Nanjing University(南京大学) Queen Mary University of London(伦敦大学女王学院) Hong Kong University of Science and Technology(香港科技大学) University of Manchester(曼彻斯特大学)

AI总结 AutoMV通过多智能体协作生成完整音乐视频,优于现有方法并接近专业水准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01700 2025-12-16 cs.AI

DeepVIS: Bridging Natural Language and Data Visualization Through Step-wise Reasoning

DeepVIS: 通过分步推理连接自然语言与数据可视化

Zhihao Shuai, Boyan Li, Siyu Yan, Yuyu Luo, Weikai Yang

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Hong Kong University of Science and Technology(香港科技大学)

AI总结 DeepVIS通过整合分步推理提升自然语言到可视化的质量,提供透明的推理过程以增强用户理解与调整能力。

Comments IEEE VIS 2025 full paper

Journal ref IEEE Transactions on Visualization and Computer Graphics, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏