arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2026-01-23 至 2026-01-23 共收录 2 信号源:cs.CV, cs.GR, cs.MM

1. 可控生成 2 篇

2601.15698 2026-01-23 cs.CV cs.AI 74%

Beyond Visual Safety: Jailbreaking Multimodal Large Language Models for Harmful Image Generation via Semantic-Agnostic Inputs

超越视觉安全:通过语义无关输入对多模态大语言模型进行有害图像生成的劫持

Mingyu Yu, Lana Liu, Zhehao Zhao, Wei Wang, Sujuan Qin

机构 * State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(网络与交换技术国家重点实验室,北京邮电大学) School of Cyberspace Security, Beijing University of Posts and Telecommunications(网络安全学院,北京邮电大学)

专题命中 可控生成 :image generation(title);分类 cs.CV

AI总结 本文提出BVS框架,通过语义无关输入对多模态大语言模型进行有害图像生成的劫持,揭示其视觉安全边界的脆弱性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09822 2026-01-23 cs.RO 50%

Multi-Layered Reasoning from a Single Viewpoint for Learning See-Through Grasping

从单一视角进行多层推理以学习透视抓取

Fang Wan, Chaoyang Song

机构 * Southern University of Science and Technology(南方科技大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 可控生成 :inpainting(abstract)

AI总结 本研究提出了一种基于视觉的透视感知架构,通过单一视觉输入实现多模态感知,无需外部摄像头或力传感器即可学习反应抓取。

Comments 39 pages, 13 figures, 2 tables, for supplementary videos, see https://bionicdl.ancorasir.com/?p=1658, for opensourced codes, see https://github.com/ancorasir/SeeThruFinger

详情

展开后加载摘要…

URL PDF HTML 收藏