STONIC:用于LLM价值画像的分层测量契约
STONIC: A Layered Measurement Contract for LLM Value Profiling
AI总结:
STONIC在4家银行5144种情境及35种LLM配置上,验证了问卷评分等三类观测结果描述相同稳定偏好的假设,发现模型存在行为连续性但无统一价值身份,且隐藏状态编码决策更清晰。
AI中文摘要:
大语言模型(LLM)价值研究常将问卷评分、成对选择及从生成文本中推断出的价值合并为一个画像,这种合并假设三类观测结果描述的是相同的稳定偏好。STONIC在来自4家银行的5144种情境和35种固定模型配置上对该假设进行了测试,它比较了孤立评分的响应、在平衡冲突下做出的选择、自发回答,以及模型自身答案与人工编写的替代答案之间的后续选择。在具备可用行为数据的17种配置中,有10种在各银行间保留了认可-选择关系;17种符合条件的配置均更偏好自身较早的答案(中位数效应为0.790),尽管选项位置会改变所有符合条件配置的选择率。画像形态从评分到冲突选择的传递最强,对自发文本的传递则较弱。对200个L3响应的三重标注为语义审计提供了任务局部检查:FULCRA与人类多数意见最为一致,DeBERTa在校准后保留了有用的排序信息;隐藏状态比单独的提示更清晰地编码了完整决策。因此,模型表现出可复现的行为连续性,但证据不支持跨接口存在一个独立于评分器的价值身份。
英文摘要:
LLM value studies often merge questionnaire ratings, pairwise choices, and values inferred from generated text into one profile. That merge assumes that the three observations describe the same stable preference. STONIC tests this assumption on 5,144 situations from four banks and 35 fixed model configurations. It compares responses rated in isolation, choices made under counterbalanced conflict, spontaneous answers, and later choices between a model's own answer and authored alternatives. 10 of 17 configurations with usable behavioral data preserve the endorsement-choice relation across banks. Every one of the 17 eligible configurations prefers its own earlier answer (median effect 0.790), although option position changes the choice rate in every eligible configuration. Profile shape transfers most strongly from ratings to conflict choices and weakens for spontaneous text. Three-way annotation of 200 L3 responses provides a task-local check of the semantic audit: FULCRA agrees most closely with the human majority, while DeBERTa retains useful rank information after calibration. Hidden states encode the completed decision more clearly than the prompt alone. Thus the models show reproducible behavioral continuity, but the evidence does not support one scorer-independent value identity across interfaces.