发表机构
Institute of Automation, Chinese Academy of Sciences; Zhejiang Gongshang University; Carnegie Mellon University; University of Toronto(中国科学院自动化研究所; 浙江工商大学; 卡内基梅隆大学; 多伦多大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对微观机器人提出语义到物理的框架,集成轨迹并提升定位精度,减少训练时间,提高目标区域覆盖率。
AI 中文摘要
微观机器人需要准确的任务几何信息,即便语言、部件和焦点发生变化。我们提出一种语义到物理的框架,该框架可将指令映射为受约束的几何算子,复用冻结的开放词汇感知,并通过置信度加权和动态规划集成局部可靠的焦平面轨迹。校准的多视图几何将二维路径与物理执行相连。提示、未见部件和几何重构测试的均方根误差(RMSE)为6.30-6.59像素。与使用20-100个标签的部件特定U-Net训练相比,所提出的零新标签配置耗时15分钟,而前者耗时72-165分钟。在9种部件光照条件下,与图像优先的多焦点融合相比,轨迹空间集成将RMSE从14.41像素降至6.28像素(降低56.4%),P95误差从20.07像素降至8.13像素(降低59.5%)。消融实验分离了置信度和路径选择的作用。在代表性机器人实验中,目标区域覆盖率从83.5%提升至92.9%。分配操作提供可测量的物理轨迹,并非该方法的任务特定限制。
英文摘要
Microscopic robots require accurate task geometry despite changes in language, parts, and focus. We present a semantic-to-physical framework that maps instructions to constrained geometric operators, reuses frozen open-vocabulary perception, and integrates locally reliable focal-plane trajectories by confidence weighting and dynamic programming. Calibrated multi-view geometry connects 2-D paths to physical execution. Prompt, unseen-part, and geometry reconfiguration tests yield 6.30-6.59-pixel RMSE. Relative to part-specific U-Net training with 20-100 labels, the proposed zero-new-label configuration takes 15 rather than 72-165 min. Across nine part-illumination conditions, trajectory-space integration reduces RMSE from 14.41 to 6.28 pixels (56.4%) and P95 error from 20.07 to 8.13 pixels (59.5%) compared with image-first multi-focus fusion. An ablation isolates the roles of confidence and path-wise selection. In representative robot experiments, target-region coverage improves from 83.5% to 92.9%. Dispensing provides a measurable physical trace, not a task-specific limitation of the method.