arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03846cs.GTcs.LG

EF1约束下的纳什社会福利与相同加性估值:复杂性、保证与实验

EF1-Constrained Nash Social Welfare with Identical Additive Valuations: Complexity, Guarantees, and Experiments

  • National Taiwan Ocean University(国立台湾海洋大学)
  • National Yang Ming Chiao Tung University(国立阳明交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Zih-Sian Yang, Yi-Hao Chen, Yu-Te Kuan, Cheng-Jui Wu, Chuang-Chieh Lin, Po-An Chen

AI总结:

针对相同加性估值下的商品分配,研究EF1约束下的纳什社会福利,提出深度强化学习框架PriorityNet,通过前瞻性动作掩码保证前缀级EF1,实验显示高福利性能。

AI中文摘要:

我们研究了在具有相同加性估值的智能体之间分配不可分割商品的问题,重点关注最多一个商品的无嫉妒性(EF1)和纳什社会福利(NSW)。由于在加性估值下,每个最大NSW分配都是EF1的,相关的阈值问题继承了在相同加性估值下NSW最大化的已知强NP困难性,并且是强NP完全的。因此,我们关注任意EF1分配所满足的福利保证。尽管已知每个这样的分配都能达到无限制最优NSW的$e^{-1/e}$近似,但我们确定了产生更强保证的条件。在均匀估值下,每个EF1分配都是NSW最优的。在$\varepsilon$-小物品条件下,每个EF1分配都达到一个显式近似比$\rho_n(\varepsilon)$,当$\varepsilon\to 0$且$n$固定时,满足$\rho_n(\varepsilon)=1-O(\varepsilon^2)$。我们进一步考虑更强的顺序要求,即在每次物品分配后保持EF1。为此,我们提出了PriorityNet,一个深度强化学习框架,使用近端策略优化训练,并配备前瞻性EF1动作掩码。该掩码将每个决策限制为保持EF1的分配,从而通过构造保证前缀级EF1,无需后处理修复。在离线场景和随机顺序在线场景的3000个测试实例中($n\in[2,20]$,$m\in[5,100]$),PriorityNet分别达到平均归一化NSW值0.9911和0.9701。相对于离线最长处理时间(LPT)和在线最小价值束基线,它实现了实例级胜减负率+27.10%和+17.87%,同时与离线基线的平均归一化福利匹配到小数点后四位,并将在线平均值从0.9694适度提高到0.9701。

英文摘要:

We study the allocation of indivisible goods among agents with identical additive valuations, focusing on envy-freeness up to one good (EF1) and Nash social welfare (NSW). Since every maximum-NSW allocation is EF1 under additive valuations, the associated threshold problem inherits the known strong NP-hardness of NSW maximization under identical additive valuations and is strongly NP-complete. We therefore focus on welfare guarantees satisfied by arbitrary EF1 allocations. Although every such allocation is known to achieve an $e^{-1/e}$-approximation to the unrestricted optimal NSW, we identify conditions yielding stronger guarantees. Under uniform valuations, every EF1 allocation is NSW-optimal. Under an $\varepsilon$-small-item condition, every EF1 allocation achieves an explicit approximation ratio $ρ_n(\varepsilon)$ satisfying $ρ_n(\varepsilon) = 1-O(\varepsilon^2)$ as $\varepsilon\to 0$ for fixed $n$. We further consider the stronger sequential requirement that $\operatorname{EF1}$ be maintained after every item assignment. For this setting, we introduce \emph{PriorityNet}, a deep reinforcement learning framework trained with Proximal Policy Optimization (PPO) and equipped with prospective $\operatorname{EF1}$ action masking, which guarantees prefix-wise $\operatorname{EF1}$ by construction. Across 3,000 test instances in each of the offline full-information and random-order online regimes ($n\in[2,20]$, $m\in[5,100]$), PriorityNet achieves mean normalized $\operatorname{NSW}$ values of $0.9911$ and $0.9701$, respectively. Relative to the offline Longest Processing Time (LPT) heuristic and the online least-valued-bundle rule, it attains instance-wise win-minus-loss rates of $+27.10\%$ and $+17.87\%$. Its aggregate welfare matches the offline LPT baseline to four decimal places and modestly improves upon the online baseline, from $0.9694$ to $0.9701$.

↑