arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Beihang University(北京航空航天大学)

2025-12-29 至 2025-12-29 共收录 6
2512.22087 2025-12-29 cs.CL

Context as a Tool: Context Management for Long-Horizon SWE-Agents

上下文作为工具:长周期SWE代理的上下文管理

Shukai Liu, Jian Yang, Bo Jiang, Yizhi Li, Jinyang Guo, Xianglong Liu, Bryan Dai

机构 * Beihang University(北航大学) Manchester(曼彻斯特) Ubiquant

AI总结 CAT通过整合上下文管理工具,提升SWE代理的长周期推理能力,实验表明其在SWE-Bench-Verified上达到57.6%的解决率,显著优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21220 2025-12-29 cs.AI cs.CV cs.RO

RoboSafe: Safeguarding Embodied Agents via Executable Safety Logic

RoboSafe: 通过可执行的安全逻辑保障具身智能体

Le Wang, Zonghao Ying, Xiao Yang, Quanchen Zou, Zhenfei Yin, Tianlin Li, Jian Yang, Yaodong Yang, Aishan Liu, Xianglong Liu

机构 * Beihang University(北航大学) AI Security Lab(360AI安全实验室) The University of Sydney(悉尼大学) Nanyang Technological University(南洋理工大学) Peking University(北京大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

AI总结 RoboSafe通过可执行的安全逻辑保障具身智能体,减少危险行为并保持任务性能。

Comments 11 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07148 2025-12-29 cs.CL

TCM-Eval: An Expert-Level Dynamic and Extensible Benchmark for Traditional Chinese Medicine

TCM-Eval: 一个专家级动态且可扩展的中医基准测试

Zihao Cheng, Yuheng Lu, Huaiqian Ye, Zeming Liu, Minqi Wang, Jingjing Liu, Zihan Li, Wei Fan, Yuanfang Guo, Ruiji Fu, Shifeng She, Gang Wang, Yunhong Wang

机构 * School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) Beijing Zhimingtang Technology Co., Ltd.(北京智明堂科技有限公司) Beijing Zhiyan AI Technology Co., Ltd.(北京智言AI科技有限公司) Guangzhou University of Chinese Medicine(广州中医药大学)

AI总结 TCM-Eval通过动态可扩展的基准测试和SI-CoTE方法,开发出ZMT模型,显著超越人类中医从业者水平,推动中医领域LLM研究发展。

Comments Work in Progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.14051 2025-12-29 cs.CL cs.AI

GroupDebate: Enhancing the Efficiency of Multi-Agent Debate Using Group Discussion

GroupDebate: 通过群体讨论提升多智能体辩论效率

Tongxuan Liu, Xingyu Wang, Weizhe Huang, Wenjiang Xu, Yuting Zeng, Lei Jiang, Hailong Yang, Jing Li

机构 * University of Science and Technology of China(中国科学技术大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beihang University(北航)

AI总结 本文提出GroupDebate方法,通过将智能体划分为小组并共享中间结果,减少多智能体辩论的token成本,提升效率和准确性。

Comments Accepted by AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21618 2025-12-29 cs.CV cs.RO

SymDrive: Realistic and Controllable Driving Simulator via Symmetric Auto-regressive Online Restoration

SymDrive: 通过对称自回归在线修复实现逼真且可控的驾驶模拟器

Zhiyuan Liu, Daocheng Fu, Pinlong Cai, Lening Wang, Ying Liu, Yilong Ren, Botian Shi, Jianqiang Wang

机构 * School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) State Key Lab of Intelligent Transportation System, Beihang University(北航智能交通系统国家重点实验室)

AI总结 SymDrive通过对称自回归在线修复技术,实现了逼真可控的驾驶模拟,解决了高保真渲染与交互式交通编辑的难题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21450 2025-12-29 cs.LG

RLLaVA: An RL-central Framework for Language and Vision Assistants

RLLaVA: 一种面向语言和视觉助手的强化学习中心框架

Lei Zhao, Zihao Ma, Boyu Lin, Yuhe Liu, Wenjun Wu, Lei Huang

机构 * SKLCCSE, Institute of Artificial Intelligence, Beihang University, Beijing, China(信息与电子技术学院,人工智能研究院,北京航空航天大学,北京,中国) Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University(未来区块链与隐私计算先进创新中心,北京航空航天大学) Hangzhou International Innovation Institute, Beihang University, Hangzhou, China(杭州国际创新研究院,北京航空航天大学,杭州,中国)

AI总结 RLLaVA 提出了一种强化学习中心框架,通过解耦算法逻辑与模型架构,实现高效训练和多任务扩展,提升视觉-语言模型性能。

Comments The code is available at https://github.com/TinyLoopX/RLLaVA

详情

展开后加载摘要…

URL PDF HTML 收藏