AI 中文总结
本研究以RentAHuman平台为案例,通过审核981个悬赏清单,测量证明负担,发现56.2%的清单证明负担评分达4或5分,且智能体/机器人标记的清单更常要求现实世界行动等证明。
AI 中文摘要
在线悬赏市场让需求方发布付费任务,工人不仅需完成任务,还可能被要求提供证明,而证明可能涉及身份或位置暴露、使用个人账号、公开发帖、在现实世界中行动,或后续检查时的重复证据,这些均未在发布的价格中披露。我们将这些公开要求称为证明负担,并在RentAHuman(一个2026年宣传为AI智能体雇佣人类的市场)上对其进行测量。我们研究的是清单要求的内容,而非工人提交或经历的内容。我们手动审核了2026年5月31日的非随机快照:从RentAHuman和另一个同类市场Human Pages中搜索到的所有清单,共981个,除1个外均来自RentAHuman。两名独立编码员记录了13个特征(11种证据、持续监控、现实世界行动)以及我们的0-5分证明负担评分;一名盲态第三方解决了所有分歧。计划的内容筛选后,剩余779个悬赏/任务清单作为主要研究群体;438个(56.2%)评分达4或5分,涵盖154种不同的特征组合:工人需了解清单内容,清单需提供特征组合检查表而非单一评分。平台元数据将部分需求方账户标记为智能体或机器人。探索性比较显示,标记为智能体/机器人的清单中,75.0%包含现实世界行动、位置证明或持续监控,而标记为人类的清单中该比例为55.3%,但评分4或5的比例无明显差异。这些标记为自我报告或平台分配,智能体/机器人标记的清单仅来自20个显示名称,且比较是在看到数据后事后选择的:这是一个假设,而非已证实的差异。我们提供了13项要求的词汇表、经裁定的手动审核以及本描述性案例研究;评分是次要筛选摘要。本研究目前尚未提供经工人验证的测量方法或自动检测器。
英文摘要
Online bounty markets let requesters advertise paid tasks. Workers may be asked not just to complete a task but to prove it, and proof can mean exposure: revealing identity or location, using a personal account, posting publicly, acting in the physical world, or repeated evidence at later checks, none disclosed by the posted price. We call these advertised requirements proof burden and measure them on RentAHuman, a 2026 market publicized as a place for AI agents to hire humans. We study what listings request, not what workers submit or experience. We manually audited a nonrandom May 31, 2026 snapshot: every listing our searches returned from RentAHuman and Human Pages, another such market (981 listings, all but one from RentAHuman). Two independent coders recorded 13 features (11 kinds of evidence, recurring monitoring, physical-world action) and our 0-5 Proof Burden Score; a blinded third resolved all disagreements. A planned content screen leaves 779 bounty/task listings as the primary population; 438 (56.2%) score 4 or 5, spanning 154 distinct feature combinations: a checklist, not a single score, tells workers what a listing entails. Platform metadata labels some requester accounts as agents or bots. Exploratory comparisons show physical-world action, location proof, or recurring monitoring in 75.0% of agent-or-bot-labeled versus 55.3% of human-labeled listings, though score-4-or-5 shares did not clearly differ. The labels are self-reported or platform-assigned, the agent-or-bot-labeled listings come from only 20 displayed names, and the comparison was chosen post hoc, after seeing the data: a hypothesis, not a confirmed difference. We contribute the 13-requirement vocabulary, the adjudicated manual audit, and this descriptive case study; the score is a secondary screening summary. The study offers no worker-validated measure or automated detector yet.
CommentsAccepted at ACM HCOMP 2026 (2026 ACM Conference on Human-AI Complementarity and Alignment), September 27-30, 2026, Alexandria, VA, USA. 10 pages, 4 figures, 3 tables. Camera-ready version; published version DOI: 10.1145/3834580.3838744