arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

针对监控与安全机器人的身份敏感型大语言模型输出基准测试

Benchmarking Identity-Sensitive LLM Outputs for Surveillance and Security Robots

Nneka Hyman, Jasmine Khan, Raj Korpan

arXiv 2608.16030首次发表:更新:

发表机构

Hunter College, City University of New York; The Graduate Center, City University of New York(纽约城市大学亨特学院; 纽约城市大学研究生中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究测试身份条件提示对LLM生成的监控与安全机器人设计描述的影响,用236个人口统计学身份标签分析可读性,发现存在显著差异,可读性可为相关基准框架提供可解释基线。

AI 中文摘要

大语言模型(LLM)正越来越多地用于在机器人开发初期生成文本形式的机器人设计规范、交互策略及风险评估,这类输出可能影响监控与安全机器人的概念化、文档编制及最终实施。本文评估身份条件提示是否会在LLM生成的监控与安全机器人设计描述中产生系统性差异。我们使用236个人口统计学身份标签,涵盖单标签及模型增强提示条件,将可读性作为初始基准,分析生成的机器人设计描述的可访问性及身份条件差异。结果显示,不同提示条件、设计维度及人口统计学身份间的可读性存在显著差异。尽管可读性无法确定输出是否公平或符合社会规范,但它在包含词汇、语义、情感、句法及公平性分析的更广泛基准框架内提供了可解释的基线。

英文摘要

Large language models (LLMs) are increasingly used to generate textual robot design specifications, interaction policies, and risk assessments during early-stage robot development. Such outputs may influence how surveillance and security robots are conceptualized, documented, and ultimately implemented. This paper evaluates whether identity-conditioned prompts produce systematic differences in LLM-generated surveillance and security robot design descriptions. Using 236 demographic identity labels across single-label and model-augmented prompt conditions, we analyze readability as an initial benchmark for evaluating accessibility and identity-conditioned variation in generated robot design descriptions. The results show significant differences in readability across prompt conditions, design dimensions, and demographic identities. Although readability cannot determine whether an output is fair or socially appropriate, it provides an interpretable baseline within a broader benchmarking framework that also includes lexical, semantic, sentiment, syntactic, and fairness-focused analyses.

CommentsAccepted to the Foundation Models in the Ro-Man Age (FoRMA) Workshop at IEEE RO-MAN 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑