arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkinAgent AI:面向非诊断性护肤支持的安全 grounded 多模态智能体框架

SkinAgent AI: A Safety-Grounded Multimodal Agentic Framework for Non-Diagnostic Skincare Support

Muhammad Muhtasim Shahriar, Abdullah Mohammad Sayem, Tze Hui Liew, M. F. Mridha, Md. Mahiuddin

arXiv 2609.29341首次发表:更新:

发表机构

International Islamic University Chittagong; Multimedia University; American International University-Bangladesh(吉大港国际伊斯兰大学; 多媒体大学; 美国国际大学-孟加拉)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出并评估了 SkinAgent AI,一种非诊断性多模态智能体框架,通过视觉路由、数据库 grounding 和安全检查实现可审计的护肤支持,实验表明其路由准确率高达 99.84%,但严格任务完成率仅 47.08%,仍需独立验证。

AI 中文摘要

面向消费者的护肤人工智能必须在明确的证据和安全边界内协调视觉证据、产品信息、工具使用和面向用户的行动。本研究评估了 SkinAgent AI,这是一个非诊断性多模态框架,将视觉问题路由与基于 LLM 的 grounded 且可审计的编排相结合。该架构包括针对痤疮、毛孔和皱纹的路由;基于照片的肤质估计;基于计数的有序痤疮严重程度支持;类型化工具;基于数据库的推荐和行动功能;确定性安全、隐私和证据检查;状态改变行动前的批准;以及结构化轨迹和重放机制。视觉模型性能和系统级智能体行为分别进行了评估。在三个随机种子下,皮肤状况路由模型达到了 99.84% ± 0.07% 的准确率。肤质估计达到了 88.85% 的准确率,而基于计数的痤疮严重程度支持达到了 84.59% 的准确率,二次加权 kappa 为 0.9076。在一个锁定但非独立的 240 例系统基准测试中,意图准确率为 80.00%,精确工具集匹配率为 62.92%,严格任务完成率为 47.08%。在有限的安全和隐私测试套件中,未观察到违规或成功的跨用户泄漏事件。然而,工具选择错误、产品属性 grounding 不完整以及不可靠的失败回退仍然存在。这些发现支持了有界、基于数据库且可追溯的智能体编排用于非诊断性护肤辅助的可行性。它们并未确立临床就绪性、外部泛化、正式隐私保证或普遍安全性。独立验证、专家评估、鲁棒性和公平性测试以及真实世界环境中的前瞻性评估仍然是必要的。

英文摘要

Consumer-facing skincare AI must coordinate visual evidence, product information, tool use, and user-facing actions within explicit evidence and safety boundaries. This study evaluates SkinAgent AI, a non-diagnostic multimodal framework that combines visual concern routing with grounded and auditable LLM-based orchestration. The architecture includes routing for Acne, Pores, and Wrinkles; photograph-based skin-type estimation; count-informed ordinal acne-severity support; typed tools; database-grounded recommendation and action functions; deterministic safety, privacy, and evidence checks; approval before state-changing actions; and structured trace and replay mechanisms. Visual-model performance and system-level agent behavior were evaluated separately. Across three seeds, the skin-condition routing model achieved 99.84% +/- 0.07% accuracy. Skin-type estimation achieved 88.85% accuracy, while count-informed acne-severity support achieved 84.59% accuracy with a quadratic weighted kappa of 0.9076. On a locked but non-independent 240-case system benchmark, intent accuracy was 80.00%, exact tool-set match was 62.92%, and strict task completion was 47.08%. No violations or successful cross-user leakage events were observed in the finite safety and privacy test suites. Tool-selection errors, incomplete grounding of product attributes, and unreliable failure fallback nevertheless remained. These findings support the feasibility of bounded, database-grounded, and traceable agent orchestration for non-diagnostic skincare assistance. They do not establish clinical readiness, external generalization, formal privacy guarantees, or universal safety. Independent validation, expert assessment, robustness and fairness testing, and prospective evaluation in real-world settings remain necessary.

CommentsSubmitted to JMIR AI and currently under peer review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑