arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

我该留下还是展示?学习有选择性地披露信息

Should I stay or should I show? Learning to selectively disclose information

Carlotta Giacchetta, Alessando Bogani, Cesare Barbera, Giovanni De Toni, Michele Caprio, Andrea Pugnana, Andrea Passerini

arXiv 2609.39818首次发表:更新:

发表机构

University of Trento; University of Pisa; ETH AI Center & ETH; University of Warwick(特伦托大学; 比萨大学; 苏黎世联邦理工学院AI中心及苏黎世联邦理工学院; 华威大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在预算约束下学习何时向人类决策者披露支持信息,提出基于信息价值的阈值策略,实验证明其优于不披露和完全披露,且能提升人机团队表现。

AI 中文摘要

在许多高风险场景中,人类决策者在做出决策前可以获取支持信息。然而,获取信息是有成本的,且披露可能无法改善人类决策,甚至可能损害决策。我们通过研究选择性披露来解决这一问题,即在预算约束下学习何时向人类决策者揭示支持信息。我们首先证明最优策略是信息价值(VoI)上的阈值规则,即披露所引起的人类决策风险的预期降低。由于VoI在实践中是未知的,我们估计特定情境下的人类风险,并限定由此产生的插入式策略相对于不披露的可能性能退化,以及其相对于最优策略的遗憾。在基准数据集上的实验表明,无论支持信息是有益还是有害,选择性披露都优于完全不披露和完全披露。两项用户研究表明,当披露由我们学习的策略而非人类自行选择主导时,人机团队的表现可以得到改善,尽管这一优势因任务而异。一个反事实基准(在披露发生时用机器学习预测替代参与者的预测)表明,这些差异可能取决于当信息自动提供而非自行请求时,参与者对建议的遵循度较低。

英文摘要

In many high-stakes settings, human decision-makers can acquire support information before making a decision. However, acquiring information is costly, and disclosure may fail to improve human decisions or may even impair them. We tackle this problem by studying selective disclosure, i.e., the problem of learning when to reveal support information to a human decision-maker under a budget constraint. We first show that the optimal policy is a threshold rule on the Value of Information (VoI), i.e., the expected reduction in human decision risk induced by disclosure. Since VoI is unknown in practice, we estimate the regime-specific human risks and bound the possible degradation of the resulting plug-in policy relative to lack of disclosure, as well as its regret relative to the optimal policy. Experiments on benchmark datasets show that selective disclosure outperforms both no disclosure and full disclosure, regardless of whether the support information is beneficial or harmful. Two user studies show that human-AI team performance can improve when disclosure is led by our learned policy and not human-selected, although this advantage varies across tasks. A counterfactual benchmark, which replaces participants' predictions with a machine-learning prediction when disclosure occurs, suggests that these differences might depend on lower adherence to advice when the information is automatically provided rather than self-requested.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑