arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

The Hong Kong University of Science and Technology(香港科技大学)

2025-12-23 至 2025-12-23 共收录 10
2512.19686 2025-12-23 cs.CV

Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models

视觉感知的CoT:在统一模型中实现高保真的视觉一致性

Zixuan Ye, Quande Liu, Cong Wei, Yuanxing Zhang, Xintao Wang, Pengfei Wan, Kun Gai, Wenhan Luo

机构 * The Hong Kong University of Science and Technology(香港科技大学) Kling Team, Kuaishou Technology(快手科技 Kling 团队)

AI总结 本文提出视觉感知的CoT方法,通过自适应视觉规划和迭代视觉修正提升统一模型的多模态生成能力,实现更高的视觉一致性。

Comments Project Page: https://zixuan-ye.github.io/VACoT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19486 2025-12-23 cs.CV

Dynamic Stream Network for Combinatorial Explosion Problem in Deformable Medical Image Registration

动态流网络用于变形医学图像配准中组合爆炸问题

Shaochen Bi, Yuting He, Weiming Wang, Hao Chen

机构 * Hong Kong University of Science and Technology(香港科技大学) Case Western Reserve University(凯斯西储大学) Hong Kong Metropolitan University(香港理工大学)

AI总结 本文提出动态流网络DySNet,通过自适应流盆地和动态流注意力机制解决变形医学图像配准中的组合爆炸问题,实现更高效的特征建模与匹配。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19240 2025-12-23 cs.CL cs.AI

ChemATP: A Training-Free Chemical Reasoning Framework for Large Language Models

ChemATP: 一种无需训练的化学推理框架用于大语言模型

Mingxu Zhang, Dazhong Shen, Qi Zhang, Ying Sun

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) Shanghai AI Laboratory(上海人工智能实验室)

AI总结 ChemATP通过构建原子级文本知识库,使冻结的大语言模型能够动态检索和推理化学知识,从而在无需训练的情况下实现高效的化学推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18853 2025-12-23 cs.CV cs.HC

VizDefender: Unmasking Visualization Tampering through Proactive Localization and Intent Inference

VizDefender:通过主动定位和意图推断揭示可视化篡改

Sicheng Song, Yanjie Zhang, Zixin Chen, Huamin Qu, Changbo Wang, Chenhui Li

机构 * East China Normal University(华东师范大学) Hong Kong University of Science and Technology(香港科学与技术大学) School of Computer Science and Technology(计算机科学与技术学院)

AI总结 VizDefender通过半脆弱水印和意图分析模块,主动定位和推断可视化篡改,有效检测和分析数据篡改与视觉编码篡改。

Comments IEEE Transactions on Visualization and Computer Graphics (IEEE PacificVis'26 TVCG Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18745 2025-12-23 cs.CV cs.CL cs.LG

InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search

InSight-o3: 通过通用视觉搜索增强多模态基础模型

Kaican Li, Lewei Yao, Jiannan Wu, Tiezheng Yu, Jierun Chen, Haoli Bai, Lu Hou, Lanqing Hong, Wei Zhang, Nevin L. Zhang

机构 * Hong Kong University of Science and Technology(香港科技大学) Huawei(华为)

AI总结 InSight-o3通过通用视觉搜索任务提升多模态基础模型的推理能力,解决复杂视觉信息整合问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23950 2025-12-23 cs.AI

InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback

InterMT:基于人类反馈的多轮交错偏好对齐

Boyuan Chen, Donghai Hong, Jiaming Ji, Jiacheng Zheng, Bowen Dong, Jiayi Zhou, Kaile Wang, Juntao Dai, Xuyao Wang, Wenqi Chen, Qirui Zheng, Wenxin Li, Sirui Han, Yike Guo, Yaodong Yang

机构 * Institute for AI, Peking University(人工智能研究院,北京大学) State Key Laboratory of General Artificial Intelligence, Peking University(通用人工智能国家重点实验室,北京大学) Hong Kong University of Science and Technology(香港科技大学)

AI总结 InterMT通过多轮多模态交互的偏好数据集探索,旨在提升多模态大模型的交互能力,结合人类反馈和专家注释,揭示多轮扩展规律。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18822 2025-12-23 cs.AI cs.CL

AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting

AdaCtrl: 通过难度感知预算分配实现适应性与可控性推理

Shijue Huang, Hongru Wang, Wanjun Zhong, Zhaochen Su, Jiazhan Feng, Bowen Cao, Yi R. Fung

机构 * Hong Kong University of Science and Technology(香港科技大学) The Chinese University of Hong Kong(香港中文大学) Peking University(北京大学)

AI总结 AdaCtrl通过难度感知预算分配实现自适应推理控制,提升模型在不同任务中的效率与效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.14579 2025-12-23 cs.CV cs.AI

GSRender: Deduplicated Occupancy Prediction via Weakly Supervised 3D Gaussian Splatting

GSRender: 通过弱监督的3D高斯点划法实现去重的占用预测

Qianpu Sun, Changyong Shu, Sifan Zhou, Runxi Cheng, Yongxian Wei, Zichen Yu, Dawei Yang, Sirui Han, Yuan Chun

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Houmo AI Dalian University of Technology(大连理工大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

AI总结 GSRender通过弱监督3D高斯点划法和射线补偿模块,有效减少占用预测中的重复问题,提升户外自动驾驶的感知性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18411 2025-12-23 cs.CV cs.AI

AmPLe: Supporting Vision-Language Models via Adaptive-Debiased Ensemble Multi-Prompt Learning

AmPLe: 通过自适应去偏集成多提示学习支持视觉-语言模型

Fei Song, Yi Li, Jiangmeng Li, Rui Wang, Changwen Zheng, Fanjiang Xu, Hui Xiong

机构 * National Key Laboratory of Space Integrated Information System, Institute of Software, Chinese Academy of Sciences(中国科学院空间信息集成系统国家重点实验室,软件研究所) University of Chinese Academy of Sciences(中国科学院大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学)

AI总结 AmPLe通过自适应去偏集成多提示学习方法,解决模型-提示匹配偏差和样本-提示匹配偏差,提升视觉-语言模型在下游任务中的性能。

Comments Accepted by IJCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02800 2025-12-23 cs.CL

Survey and Experiments on Mental Disorder Detection via Social Media: From Large Language Models and RAG to Agents

面向社交媒体的抑郁症检测综述与实验:从大型语言模型和RAG到代理

Zhuohan Ge, Darian Li, Yubo Wang, Nicole Hu, Xinyi Zhu, Haoyang Li, Xin Zhang, Mingtao Zhang, Shihao Qi, Yuming Xu, Han Shi, Chen Jason Zhang, Qing Li

机构 * The Hong Kong Polytechnic University(香港理工大学) The Hong Kong University of Science and Technology(香港科学与技术大学)

AI总结 本文综述并实验了基于社交媒体的抑郁症检测方法,探讨了LLM、RAG和代理系统在提升检测可靠性与推理能力中的应用。

Comments 20 pages, 10 figures. This is an extension of ICDEW 2025

详情

展开后加载摘要…

URL PDF HTML 收藏