arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

YUKTI:从自然语言情境到稳健、可验证的决策——一种不确定性类型的命题IR、假设稳健帕累托前沿和遗憾证书

YUKTI: From Natural-Language Situations to Robust, Verifiable Decisions An Uncertainty-Typed Proposition IR, Assumption-Robust Pareto Frontiers, and a Regret Certificate

Suyash Mishra

arXiv 2607.09706首次发表:更新:

AI 中文总结

研究从自然语言情境生成稳健决策,核心方法是用类型命题图表示,经多求解器及分布帕累托交接耦合,并引入ARPF等。主要贡献是证明rho与决策遗憾关系,增加可追溯性,合成数据基础,通过多种验证方式展示方法有效性。

AI 中文摘要

语言模型将文字情境转化为数值计划,而主流管道(NL4Opt、OptiMUS、ORLM、OR-LLM-Agent)致力于单一目标和点值系数,然后求解一次。对于分配实际预算、努力或临床关注的决策,这种信心是失败模式,因为每个客观化数字都是一个假设,只有猜测完全正确时最优的计划是脆弱的。YUKTI改变了自动公式化的目标。其表示是一个类型命题图,关系带有形状先验、系数不确定性和来源。YUKTI将每个阶段路由到精确、非线性或进化求解器;通过分布帕累托交接耦合阶段;引入假设稳健帕累托前沿(ARPF),重新采样假设(包括结构ε污染)以评估每个行动存活的频率(rho)。我们证明了一个界限,使rho成为决策遗憾的精确因子,增加可审计的可追溯性,并在不存在时合成一个忠实于基准的数据基础(SRJANA)。我们通过三种方式验证:在受控错误指定下,稳健折衷将均值和尾部遗憾比朴素点计划减少90%以上;在受监管的商业决策中,我们在合法行动空间内优化并以欧元定价下行风险;在41188个决策的真实公共数据集上,样本外回测比记录的现状高出34%,比朴素点规则高出4%,同时减少了优化器的诅咒。求解器是标准的,我们不声称在基准-SOTA中获胜。一个直接比较表明,给定正确数字的语言模型和单目标优化,两者产生的留存遗憾约为YUKTI的47倍——语言模型是公式化者,而非求解器。在长期因果耦合下,前向交接变得不合理,确定其必须成为反向归纳因果策略的位置。

英文摘要

Language models turn a worded situation into a numeric plan, and the dominant pipelines (NL4Opt, OptiMUS, ORLM, OR-LLM-Agent) commit to a single objective and point-valued coefficients, then solve once. For decisions that allocate real budget, effort, or clinical attention, that confidence is the failure mode: every objectified number is an assumption, and a plan optimal only if the guesses are exactly right is fragile -- mimicry of computation. YUKTI changes the target of autoformulation. Its representation is a typed-proposition graph whose relationships carry shape priors, coefficient uncertainty, and provenance. YUKTI routes each stage to an exact, nonlinear, or evolutionary solver; couples stages by a distributional Pareto hand-off; and introduces Assumption-Robust Pareto Frontiers (ARPF), resampling assumptions (including structural epsilon-contamination) to score how often each action survives (rho). We prove a bound making rho an exact factor of decision regret, add auditable traceability, and synthesize a benchmark-faithful data foundation when none exists (SRJANA). We validate three ways: under controlled misspecification the robust compromise cuts mean and tail regret by over 90% versus a naive point plan; on a regulated commercial decision we optimize inside a lawful action space and price the downside in euros; and on a real public dataset of 41,188 decisions an out-of-sample backtest beats the logged status quo by 34% and a naive point rule by 4% while reducing the optimizer's curse. The solvers are standard; we claim no benchmark-SOTA win. A head-to-head shows an LLM given the correct numbers, and single-objective optimization, both incur about 47x the held-out regret of YUKTI -- an LLM is a formulator, not a solver. Under long-range causal coupling, the forward hand-off becomes unsound, locating where it must become a backward-induction causal policy.

Comments19 Pages , 21 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑