arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14622cs.AI

以人为中心的基准测试大型语言模型(LLM)育儿建议方法

A Human-Centred Approach to Benchmarking LLMs for Parenting Advice

Yunke Zhao, Isobel Voysey, Alastair van Heerden, Rob Hughes, Jun Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

本文以育儿专家的多维度评分为标准,采用LLM作为评判者,评估15个LLM在100个育儿场景中英、中文的表现,发现聚合分数掩盖弱点、模型隐性影响育儿风格等,为相关应用提供见解。

中文摘要 AI 辅助

人们越来越多地使用大型语言模型(LLM)寻求建议,包括育儿建议。育儿是一个关键且具有社会敏感性的领域,因此评估LLM提供的建议需要超越聚合信息质量基准的指标,以考虑回应的关系和行为要素。本文利用育儿专家创建的多维度评分标准,采用LLM作为评判者的方法,在100个育儿场景中以英语和中文两种语言评估了15个LLM。结果显示,聚合分数可能掩盖评分标准特定项目的弱点,模型会隐性鼓励不同的育儿风格,且语言会影响回应。我们强调评估输出可审计性的重要性,以及在育儿等领域评估LLM生成建议面临的挑战。研究结果为选择用于直接用户互动的LLM,以及开发面向用户的育儿建议应用程序提供了重要见解。

英文摘要

People are increasingly using large language models (LLMs) to seek advice, including for parenting. Parenting is a critical and socially sensitive domain. Thus, evaluating advice provided by LLMs requires indicators beyond aggregated information quality benchmarks to consider relational and behavioural elements of the responses. With a multi-dimensional rubric created by parenting experts, this paper evaluates 15 LLMs across 100 parenting scenarios in 2 languages (English and Chinese), using an LLM-as-a-judge method. Results show that aggregate scores can hide rubric item-specific weaknesses, models implicitly encourage different parenting styles, and language influences responses. We highlight the importance of evaluation output auditability and challenges involved in evaluating LLM-generated advice in domains like parenting. Our findings provide important insights for selecting LLMs for direct user engagement and the development of user-facing parenting advice applications.

补充信息

↑