arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14118cs.CLcs.SE

用于执行感知大语言模型研究构思的预算子集优化

Budgeted Subset Refinement for Execution-Aware LLM Research Ideation

Micah Zhang

AI总结:

研究针对LLMs生成研究想法存在的问题,引入预算子集优化策略,在统一共享候选池评估中对比不同策略效果,发现随机k优化是低成本基线,多样性感知的MMR - k优化权衡最佳,表明LLM研究构思系统应作为预算支持分配系统评估。

AI中文摘要:

大语言模型(LLMs)能生成对专家评审来说新颖的研究想法,但这些想法往往缺乏多样性,难以可靠评估,也可能无法转化为强大的执行项目。本文评估了一个用于执行前支架问题的受控代理基准:给定一个由LLM生成的有噪声的研究想法池,系统应如何在固定准则下分配有限的优化努力,为人类研究人员构建一个更强、更多样化、更具执行意识的组合?我们引入了预算子集优化,这是一族仅优化选定候选子集而非统一优化所有候选的策略。在跨越10个随机种子和10个研究构思环境的统一共享候选池评估中,仅原始生成和重新排序在基准准则下无法产生强大的非重复研究想法,而优化对于强大的代理评级组合是必要的。统一优化能产生强大的单个想法,但不是计算资源在组合层面的最佳分配。随机k优化是一个强大的低成本基线,而考虑多样性的MMR - k优化给出了最佳的整体代理权衡:最高的强大非重复产出率、成功方法中最低的重复率以及每个强大非重复想法的最佳成本。对一个平衡的72项样本进行的盲法外部评判稳健性检查支持了跨独立模型家族的广泛优化效果,同时表明不同评判下优化策略的单项排名有所不同。这些结果表明,LLM研究构思系统不仅应作为想法生成器进行评估,还应作为预算支持分配系统进行评估。这些主张限于代理评级组合质量,不能替代专家评审或基于执行的验证。

英文摘要:

Large language models (LLMs) can generate research ideas that appear novel to expert reviewers, but recent work also shows that such ideas often lack diversity, are difficult for LLMs to evaluate reliably, and may fail to translate into strong executed projects. This paper evaluates a controlled proxy benchmark for a pre-execution scaffolding problem: given a noisy pool of LLM-generated research ideas, how should a system allocate limited refinement effort to construct a stronger, more diverse, more execution-aware portfolio for human researchers under a fixed rubric? We introduce Budgeted Subset Refinement, a family of strategies that refine only a selected subset of candidates rather than refining all candidates uniformly. In a unified shared-candidate-pool evaluation across 10 random seeds and 10 research-ideation environments, raw generation and reranking alone produce no research-strong nonduplicate ideas under the benchmark rubric, while refinement is necessary for strong proxy-rated portfolios. Uniform refinement produces strong individual ideas but is not the best portfolio-level allocation of compute. Random-k refinement is a strong low-cost baseline, while diversity-aware MMR-k refinement gives the best overall proxy tradeoff: the highest research-strong nonduplicate yield, the lowest duplicate rate among successful methods, and the best cost per research-strong nonduplicate idea. A blinded external-judge robustness check on a balanced 72-item sample supports the broad refinement effect across independent model families, while showing that per-item rankings among refined strategies vary by judge. These results suggest that LLM research ideation systems should be evaluated not only as idea generators, but as budgeted support-allocation systems. The claims are scoped to proxy-rated portfolio quality and do not substitute for expert review or execution-grounded validation.

补充信息

↑