CRiDiT:为AI赋能系统中的信任校准实例化运行时测试平台
CRiDiT: Instantiating a run-time testbed for trust calibration in AI-infused systems
- University of Bordeaux, LaBRI(波尔多大学)
- Univ. Paris 1 Panthéon-Sorbonne(巴黎第一大学先贤祠-索邦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文实例化CRiDiT运行时测试平台,在三个高风险场景中验证信任校准,发现校准标准应基于证据而非标量,并指出修复动作词汇需调整。
AI中文摘要:
将AI集成到更大的技术基础设施中,使得人类信任与系统可信度的一致性(即信任校准)成为关键的工程问题,因为任一方向的错误信任都会导致运营和安全风险。虽然概念框架为理解信任校准提供了坚实基础,但其转化为运行系统仍是一个挑战,因为很少有测试平台能让人类信任输入、机器可信度证据、差距检测与修复在闭环中协同运作。本文实例化了CRiDiT(计算风险敏感双向信任),作为运行时测试平台,使用Dempster-Shafer理论和PCR5重分配实现机器侧信任,使用主观逻辑实现人类侧信任,并使用基于阈值的信任差距进行校准。遵循设计科学研究方法,我们在三个高风险场景(招聘、金融、法律)中演练该工件,在十五次会话中产生了144条记录的交互步骤。分析表明,该工件按预期捕获了信任校准动态,并揭示了实例化策略偏离其设计要求的三个点:机器侧估计从全局基准而非任务相关证据开始;风险敏感阈值未产生风险敏感的触发;校准策略将解释性提示分配给过度信任,其中修正缩小了所有6个观察案例中的差距。由于前两个源于同一设计决策,即使用两个估计标量的差异作为校准标准,它们指向一个共同需求:标准应作用于证据而非从证据派生的标量。第三个涉及检测后的处理,表明从信任修复继承的动作词汇与交互日志显示的有效行为不一致。本工作贡献了该工件、其运行时行为的特征描述,以及该特征描述引发的需求。
英文摘要:
The integration of AI into larger technical infrastructures has made the alignment of human trust with system trustworthiness, known as trust calibration, a critical engineering concern, since misplaced trust in either direction leads to operational and safety risks. While conceptual frameworks provide a strong foundation for understanding trust calibration, their translation into running systems remains a challenge, because there are few testbeds in which human trust inputs, machine trustworthiness evidence, gap detection and remediation operate together within a closed loop. This paper instantiates CRiDiT (Computational Risk-Sensitive biDirectional Trust) as a run-time testbed, operationalising machine-side trust with Dempster-Shafer Theory and PCR5 redistribution, human-side trust with Subjective Logic, and calibration with a threshold-based trust gap. Following the Design Science Research methodology, we exercise the artifact across three high-stakes scenarios (hiring, financial, legal), producing 144 logged interaction steps across fifteen sessions. The analysis shows that the artifact captures trust calibration dynamics as intended, and reveals three points at which the instantiated policy departs from its design requirements: the machine-side estimate begins from a global benchmark rather than task-relevant evidence; risk-sensitive thresholds do not produce risk-sensitive triggering; and the calibration policy assigns explanatory prompts to over-trust, where corrections narrowed the gap in all 6 observed cases. Since the first two arise from the same design decision, to make the difference of two estimated scalars the calibration criterion, they point toward a common requirement: that the criterion should operate on the evidence rather than on scalars derived from it. The third concerns what follows detection, and shows that the action vocabulary inherited from trust repair does not align with what the interaction logs show to be effective. The work contributes the artifact, a characterisation of its run-time behaviour, and the requirements this characterisation elicits.