发表机构
Sharda University(夏尔达大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型在天体动力学领域多步推理与符号操作能力不足的问题,提出基于Qwen3-8b的领域自适应框架Taramandal-GPT,结合检索增强生成与回退机制,在APBench基准上通过双重评估展现出竞争性能,尤其擅长思维密集型任务,为专业AI助手发展迈出一步。
AI 中文摘要
大型语言模型(LLM)在自然语言理解方面已展现出显著进展,然而,由于多步推理、符号操作和领域特定术语等方面的挑战,它们在天文学和天体动力学等专业领域的有效性仍然有限。为解决这一问题,我们提出了Taramandal-GPT(星座-GPT),这是一个基于Qwen3-8b骨干构建的领域自适应框架,并增强了检索增强生成(RAG)流水线和回退机制,以提高上下文精度。我们在天体动力学问题基准(APBench)上对其进行了评估,该数据集包含299个问题,涵盖从基础到高级的太空科学水平。通过采用双重评估方法——基于数值边际的评分和语义相似性评估——Taramandal-GPT在与最先进的开源和闭源模型的竞争中表现出色,尤其在思维密集型任务中展现出显著优势。这些结果凸显了专业LLM在需要准确性和可解释性的领域中的价值,将Taramandal-GPT定位为迈向可靠的人工智能(AI)助手(用于天体物理学、航天器工程和太空探索)的一步。
英文摘要
Large language models (LLMs) have shown remarkable progress in natural language understanding, yet their effectiveness in specialized fields like astronomy and astrodynamics remains limited due to challenges in multi-step reasoning, symbolic manipulation, and domain-specific terminology. To address this, we present Taramandal-GPT (Constellation-GPT), a domain-adapted framework built on the Qwen3-8b backbone, enhanced with a Retrieval-Augmented Generation (RAG) pipeline and a fallback mechanism for improved contextual precision. We evaluate it on the Astrodynamics Problems Benchmark (APBench), a dataset of 299 questions covering foundational to advanced levels of space science. Using a dual evaluation method - numeric margin-based scoring and semantic similarity assessment - Taramandal-GPT achieves competitive performance against state-of-the-art open- and closed-source models, with notable strength in thinking-intensive tasks. These results highlight the value of specialized LLMs for domains demanding accuracy and interpretability, positioning Taramandal-GPT as a step toward reliable Artificial Intelligence (AI) assistants for astrophysics, spacecraft engineering, and space exploration.
CommentsProceedings of All India Hindi Technical Conference, 05-06 February 2026