发表机构
University of Georgia; Michigan State University(佐治亚大学; 密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出MathAgent框架,利用多智能体自动优化分子性质预测中的数学不变量表示,通过数据集感知设计、低成本筛选和渐进评估,实现自适应选择并减少手动配置。
AI 中文摘要
数学不变量提供了分子结构的互补表示,但其预测效用可能因数据集和预测任务的不同而有所差异。对于新数据集,从代数拓扑、微分几何、拓扑谱理论和交换代数中适当构建不变量表示是一项具有挑战性的任务,尤其是在评估多种表示计算成本高昂的情况下。我们开发了一种数学智能体模型(MathAgent)来自动优化分子性质预测中的不变量表示。该框架结合了数据集感知的表示设计、低成本候选筛选和渐进式评估。语言模型推理支持任务解释、策略提出和工作流协调,而科学计算和数值决策由确定性工具执行。我们在蛋白质-配体结合亲和力和定量毒性预测任务上评估了MathAgent。不同的数据集偏好不同的数学不变量组合,这表明需要数据集自适应的表示选择,而非固定的通用表示。目标条件下的执行进一步表明,该框架可以根据用户指定的目标调整其评估策略。因此,MathAgent为优化数学不变量表示提供了一种系统性、证据驱动且可追踪的方法,并减少了对分子性质预测中手动配置和/或实验的依赖。
英文摘要
Mathematical invariants provide complementary representations of molecular structure, but their predictive utility can vary across datasets and prediction tasks. For a new dataset, appropriate construction of invariant representations from algebraic topology, differential geometry, topological spectral theory, and commutative algebra is a challenging task, particularly when evaluating multiple representations is computationally expensive. We develop a Mathematical Agentic Model (MathAgent) to automatically optimize invariant representations in molecular property prediction. The framework combines dataset-aware representation design, low-cost candidate screening, and progressive evaluation. Language-model reasoning supports task interpretation, strategy proposal, and workflow coordination, while scientific computations and numerical decisions are carried out by deterministic tools. We evaluate MathAgent on protein--ligand binding-affinity and quantitative toxicity prediction tasks. Different datasets favor different combinations of mathematical invariants, demonstrating the need for dataset-adaptive representation selection rather than a fixed universal representation. Objective-conditioned executions further show that the framework can adjust its evaluation strategy according to user-specified goals. MathAgent therefore provides a systematic, evidence-driven, and traceable approach for optimizing mathematical invariant representations and reducing reliance on manual configuration and/or experimentation in molecular property prediction.