PRISM: Prompt Refinement via Image-grounded Self-rewarding Mechanism for Text-to-Image Generation
PRISM:通过基于图像的自我奖励机制进行文本到图像生成的提示优化
Guo Tang, HongJie Luo, Tianxu Wang, Ying Zhang, Hao Wang
机构
*
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
Sun Yat-sen University(中山大学)
;
South China University of Technology(华南理工大学)
;
Guangdong University of Technology(广东工业大学)
ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
ARMOR++:用于对深度伪造检测器进行可转移攻击的多域原语集的智能编排
Christos Korgialas, Gabriel Lee Jun Rong, Dion Jia Xu Ho, Pai Chet Ng, Xiaoxiao Miao, Konstantinos N. Plataniotis
机构
*
Department of Informatics, Aristotle University of Thessaloniki(塞萨洛尼基亚里士多德大学信息学系)
;
Infocomm Technology Cluster, Singapore Institute of Technology(新加坡科技学院信息通信技术集群)
;
Department of Applied Physics and Applied Mathematics, Columbia University(哥伦比亚大学应用物理与应用数学系)
;
Division of Natural and Applied Sciences, Duke Kunshan University(昆山杜克大学自然科学与应用科学部)
;
Department of Electrical and Computer Engineering, University of Toronto(多伦多大学电气与计算机工程系)
A Good Initialization is All You Need for Faithful Visual Attribution
忠实视觉归因只需一个良好的初始化
Zihan Gu, Jiayu Wang, Hua Zhang, Yue Hu
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
机构
*
University of California, Los Angeles(加州大学洛杉矶分校)
;
University of Southern California(南加州大学)
;
DeerLab LLC(DeerLab有限责任公司)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Clemson University(克莱姆森大学)
;
Google(谷歌)
;
San Jose State University(圣何塞州立大学)
;
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构
*
The University of Hong Kong(香港大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Carnegie Mellon University(卡内基梅隆大学)
;
LIGHTSPEED Shenzhen(LIGHTSPEED深圳)
;
LIGHTSPEED Los Angeles(LIGHTSPEED洛杉矶)
ASSCG: Just-Right Gating over Chattering for Fast-Slow LLM Planning in Autonomous Driving
ASSCG:自动驾驶中快慢LLM规划的恰到好处门控
Sining Ang, Yuan Chen, Liu Haiyan, Xuanyao Mao, Jason Bao, Xuliang, Bingchuan Sun, Yan Wang
机构
*
Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院)
;
Department of Automation, University of Science and Technology of China(中国科学技术大学自动化系)
;
Beijing University of Aeronautics and Astronautics(北京航空航天大学)
;
Lenovo Group Limited(联想集团)
P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling
P-MTP: 通过渐进深度缩放的多令牌预测实现高效文档解析
Le Xiang, Chenxi Zhai, Shu Wei, Jingjing Wu, Qunyi Xie, Xiao Tan, Kunbin Chen, Wei He
机构
*
Department of Computer Vision Technology (VIS) Baidu Inc China(百度计算机视觉技术部(VIS))
;
Tsinghua University Shenzhen International Graduate School China(清华大学深圳国际研究生院)
GTA-Net: Cooperative Game Theory for Vision-Language Alignment in Chest X-Ray Report Generation
GTA-Net:合作博弈论在胸部X光报告生成中的视觉-语言对齐
Saif ur Rehman Khan, Imad Ahmed Waqar, Sebastian Vollmer, Andreas Dengel, Muhammad Nabeel Asim
机构
*
Department of Computer Science, Rhineland-Palatinate Technical University of Kaiserslautern-Landau(莱茵兰-普法尔茨凯泽斯劳滕-兰道工业大学计算机科学系)
;
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
;
IntelligentX GmbH