arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不同大语言模型(LLM)生成回复的语义变异性:对基于对话的评估设计的启示

Semantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment

Jiangang Hao

arXiv 2608.24920首次发表:更新:

发表机构

ETS Research Institute; Princeton(ETS研究院; 普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究探究不同LLM生成回复的语义变异性,发现模型选择与对话上下文会影响回复相似度及与人类回复的对齐度,指出仅靠提示词和上下文无法保证跨LLM的回复一致性,需构建相关基础设施与设计策略。

AI 中文摘要

本研究探究当底层大语言模型(LLM)发生变化时,其生成的回复是否仍保持语义一致性。研究采用真实协作对话中的消息,在有无前置对话历史两种条件下,对比了不同LLM生成回复的语义相似度。结果显示,模型选择与对话上下文均会影响回复相似度及与人类回复的对齐度。这些发现表明,仅靠提示词工程和对话上下文可能不足以维持不同LLM间回复的一致性,凸显了在LLM快速持续迭代的背景下,亟需构建能保持回复稳定且可比的基础设施与设计策略。

英文摘要

This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and without preceding chat history. Results show that model choice and conversational context both affect response similarity and alignment with human replies. These findings indicate that prompting and conversational context alone may not be sufficient to preserve response consistency across LLMs, highlighting the need for infrastructure and design strategies that can maintain stable and comparable responses amid the rapid and continuous evolution of LLMs.

Comments9 pages, 4 figures, two tables. Accepted to the AI in Measurement and Education Conference 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑