arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

选择性引导作为商业影响渠道:一个可复现的合成购物智能体压力测试

Selective Elicitation as a Commercial Influence Channel: A Reproducible Synthetic Shopping-Agent Stress Test

Jiapeng Li

arXiv 2609.36614首次发表:更新:

AI 中文总结

本研究通过合成购物智能体压力测试,证明商业激励可通过选择性提问影响推荐结果,而无需进入排序算法,并量化了其效果与风险。

AI 中文摘要

商业激励不必进入最终排序算法即可影响购物助手的推荐:它反而可能影响助手询问哪个偏好问题。我们在一个刻意缩小的合成环境中使这一区别在实验上可观察。每个任务包含两个产品、三个已验证的数值属性、一个价格上限和一个私有的固定偏好向量。一个诚实的模拟用户回答一个成对问题。一个独立的推荐器接收产品和该答案,但不接收赞助分配。我们对比了中性问题、软性商业指令和明确对抗性指令(要求询问赞助商的优势而忽略竞争对手的优势)。在40个保留的赞助分配案例(20个不同的目录偏好上下文)中,软性指令未改变任何选择。对于一种语言模型推荐器,针对性指令将赞助选择提高了0.30,并将平均合成效用相对于中性提问降低了0.0547(95%上下文自助法区间[-0.0828, -0.0291])。一个固定的贝叶斯推荐器显示出类似效果;第二个模型在所有120个冻结的问题-答案输入上做出相同选择。一个终端答案一致性评判器将所有20个采样的针对性答案评为一致,尽管其中五个的合成遗憾高于0.05;一个单独的问题覆盖维度标记了它们单侧引导的问题。一个稳健的部分偏好证书在规定的合成效用下仍然有效,但仅认证了40个针对性案例中的16个,并且并不比直接询问中性问题更好。这些结果既未确立广告激励下的典型行为,也未确立对实际消费者的影响。

英文摘要

A commercial incentive need not enter the final ranking algorithm to affect a shopping assistant's recommendation: it may instead influence which preference question the assistant asks. We make this distinction experimentally observable in a deliberately small, synthetic setting. Each task has two products, three verified numerical attributes, a price limit, and a private fixed preference vector. An honest simulated user answers one pairwise question. A separate recommender receives the products and this answer but not the sponsorship assignment. We contrast a neutral question, a soft commercial instruction, and an explicitly adversarial instruction to ask about the sponsor's advantage while omitting the rival's advantage. Across 40 held-out sponsorship-assignment cases (20 distinct catalog-preference contexts), the soft instruction changes no selections. The targeted instruction raises sponsored selection by 0.30 and reduces mean synthetic utility by 0.0547 relative to neutral questioning (95% context-bootstrap interval [-0.0828, -0.0291]) for one language-model recommender. A fixed Bayesian recommender shows a similar effect; a second model makes the same choices on all 120 frozen question-answer inputs. A terminal-answer consistency judge rates all 20 sampled targeted answers consistent, although five have synthetic regret above 0.05; a separate question-coverage dimension flags their one-sided elicitation. A robust partial-preference certificate remains valid under the stipulated synthetic utility but certifies only 16 of 40 targeted cases and is not better than asking a neutral question directly. These results establish neither typical behavior under advertising incentives nor effects on actual consumers.

Comments7 pages, 1 table; synthetic data and reproduction code in ancillary files; substantial AI-tool use disclosed in the manuscript

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑