MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
MCPEvol-Bench:跨MCP服务器动态演化对大语言模型智能体性能进行基准测试
Huanxi Liu, Kun Hu, Jiaqi Liao, Qiang Wang, Pengfei Qian, YuanZhao Zhai, Dawei Feng, Bo Ding, Huaimin Wang
机构
*
College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机科学与技术学院)
;
State Key Laboratory of Complex & Critical Software Environment(复杂关键软件环境国家重点实验室)
;
National Key Laboratory of Parallel and Distributed Computing(并行与分布式计算国家重点实验室)
Comments19 pages. Keywords: Reasoning, Automated Planning, Item Responses Theory, LLMs as Planner Research Area: NLP and Symbolic Reasoning Research Area Keywords: neurosymbolic, planning in agents, symbolic reasoning Contribution Types: Model analysis & interpretability
Mean-Field Parallel Decoding for Discrete Diffusion Language Models
离散扩散语言模型的平均场并行解码
Tamim Zoabi, Ameen Ali, Liran Ringel, Lior Wolf
机构
*
School of Electrical & Computer Engineering, Tel Aviv University(特拉维夫大学电气与计算机工程学院)
;
School of Computer Science and AI, Tel Aviv University(特拉维夫大学计算机科学与人工智能学院)
;
Department of Computer Science, Technion, Israel Institute of Technology(以色列理工学院计算机科学系)
ttda704 at SemEval-2026 Task 6: Structured Chain-of-Thought Prompting for Political Evasion Detection
ttda704 at SemEval-2026 Task 6: 用于政治回避检测的结构化思维链提示
Tai Tran Tan, An Dinh Thien
机构
*
University of Information Technology, Ho Chi Minh City, Vietnam(胡志明市信息技术大学)
;
Vietnam National University, Ho Chi Minh City, Vietnam(越南国家大学胡志明市分校)
机构
*
Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)
;
Huawei Noah’s Ark Lab(华为诺亚方舟实验室)
;
University of Science and Technology Beijing(北京科技大学)