arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RiskChainBench:一个用于混淆平台消息恢复与基于证据的网页调查的基准

RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation

ZhuoXin Liu, Zhiming Ma, Ying Zhang, Mengzheng Yang, Yifan Wang, Zhengqi Huang, Yanhan Zhou, Zekun Lin, Jun Zhang, Shun Zhang, Yue Chen, Qiao Zhao, Peng Chen

arXiv 2609.16900首次发表:更新:

发表机构

Baidu; SmartFlowAI; Tsinghua University; JD Technology; Northeastern University(百度; 智谱AI; 清华大学; 京东科技; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RiskChainBench通过配对混淆消息恢复与网页调查任务,评估AI代理在风险场景中的端到端能力,发现执行失败是主要瓶颈。

AI 中文摘要

平台滥用活动利用表情符号、谐音字、字符拆分和冗余符号来隐藏重定向指令,然后通过伪装链接将用户引导至与色情、欺诈、赌博或非法交易相关的服务。现有基准分别评估混淆文本和风险网页,掩盖了目标恢复如何影响下游证据获取。我们引入了RiskChainBench,将来自600个源会话的3,600个合成令牌文本恢复输入与600个对应的人工标注本地网页环境配对。模型首先恢复消息、操作意图和目的地;同一底层模型随后充当VLM驱动的网页代理,调查正确关联的网站,并生成冻结的、带证据引用的风险报告,而不依赖消息侧语义或域名声誉线索。我们分别对恢复和正确路由的网页调查进行评分,并通过将冻结的初级入口预测作为门控应用于相同的任务2结果来离线组合它们。人工标注决定任务正确性,而固定的多模态证据评判器评估忠实性、充分性、完整性和一致性。在十个模型中,入口Top-1准确率从35.2%到95.2%,网页决策准确率从26.3%到62.8%;领先系统在入口恢复、完整重建、网站决策和细粒度分类方面表现各异。执行失败占网页运行的31.9%,而决策后类型错误仅占0.9%,表明稳定的探索和风险判断是主要瓶颈。我们发布了该基准、协议和可重置的本地沙箱。

英文摘要

Platform abuse campaigns conceal redirection instructions with emojis, homophones, character decomposition, and redundant symbols, then route users through disguised links to services associated with pornography, fraud, gambling, or illicit transactions. Existing benchmarks evaluate obfuscated text and risky webpages separately, obscuring how target recovery affects downstream evidence acquisition. We introduce RiskChainBench, pairing 3,600 synthetic token-text restoration inputs from 600 source sessions with 600 corresponding human-labeled local web environments. A model first restores the message, operational intent, and destination; the same underlying model then acts as a VLM-driven web agent that investigates the correctly associated website and produces a frozen, evidence-cited risk report without message-side semantics or domain-reputation cues. We score restoration and correct-routing web investigation separately and compose them offline by applying the frozen primary-entry prediction as a gate to the same Task 2 result. Human labels determine task correctness, while a fixed multimodal evidence judge assesses faithfulness, sufficiency, completeness, and consistency. Across ten models, Entry Top-1 ranges from 35.2% to 95.2% and web decision accuracy from 26.3% to 62.8%; the leading systems differ across entry recovery, full reconstruction, website decisions, and fine-grained typing. Execution failures account for 31.9% of web runs, whereas post-decision type errors account for only 0.9%, identifying stable exploration and risk judgment as the principal bottlenecks. We release the benchmark, protocol, and resettable local sandbox.

Comments11 pages, 5 figures; 17-page supplementary material included as an ancillary PDF. v2: updated author contribution and correspondence information; scientific content unchanged

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑