arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GeminiPainter的序列式流程由感知、认知、规划和行动阶段组成

GeminiPainter's sequence-formed pipeline comprised of perception, cognition, planning, and action stages

Miguel Altamirano Cabrera, Aleksey Fedoseev, Iana Zhura, Dzmitry Tsetserukou

arXiv 2608.00829首次发表:更新:

发表机构

Skolkovo Institute of Science and Technology(斯科尔科沃科学技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出自主机器人肖像生成系统,整合多技术实现绘画,经实验获用户高分评价,产出可识别且具吸引力的机器人肖像。

AI 中文摘要

我们提出一种自主机器人肖像生成系统,结合实时人脸检测、基于AI的草图生成与机器人绘画。该系统采集视频帧,提取面部区域,通过Gemini视觉API将其转换为极简单线草图,基于图的路径规划优化笔画顺序,并在6自由度协作机械臂上执行平滑轨迹。这种感知-认知-行动流程整合了计算机视觉、神经艺术抽象、运动优化与机器人控制。用户以5分制评分,草图质量4.33分、执行感知4.53分、用户体验4.65分,表明机器人肖像可识别、具吸引力且引人入胜。

英文摘要

We present an autonomous robotic portrait-generation system combining real-time face detection, AI-based sketch generation, and robotic drawing. The system captures video frames, extracts facial regions, converts them into minimalist single-line sketches using the Gemini Vision API, optimizes stroke order through graph-based path planning, and executes smooth trajectories on a 6-DoF collaborative manipulator. This perception-cognition-action pipeline integrates computer vision, neural artistic abstraction, motion optimization, and robot control. User ratings on a 5-point scale were high for sketch quality 4.33, perceived execution 4.53, and user experience 4.65, indicating recognizable, appealing, and engaging robotic portraits.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑