Comments9 pages, 2 figures, and 8 tables. Accepted for oral presentation at the ACM SIGKDD KDD 2026 Workshop on Personal Intelligence in the Agentic AI Era (PILA 2026)
IMProofBench: Benchmarking AI on Research-Level Mathematical Proof Generation
IMProofBench:在研究级数学证明生成上对人工智能进行基准测试
Johannes Schmitt, Gergely Bérczi, Jasper Dekoninck, Jeremy Feusi, Tim Gehrunger, Raphael Appenzeller, Pieter Belmans, Alessio Bottini, Jim Bryan, João Camarneiro, Ana Cannas da Silva, Niklas Canova, Ana-Maria Castravet, Timo de Wolff, Claudio Fontanari, Filippo Gaia, Baran Hashemi, Daniel Holmes, David Holmes, Aitor Iribar Lopez, Victor Jaeck, Martina Jørgensen, Steven Kelk, Martijn Kool, Stefan Kuhlmann, Adam Kurpisz, Johannes Lengler, Chiara Meroni, Ingmar Metzler, Martin Möller, Samuel Muñoz-Echániz, David Muñoz-Lahoz, Robert Nowak, Georg Oberdieck, Daniel Platt, Dylan Possamaï, Gabriel Ribeiro, Aluna Rizzoli, Daria Sakhanda, Raúl Sánchez Galán, Zheming Sun, Diaaeldin Taha, Josef Teichmann, Richard P. Thomas, Henk van der Pol, Michel van Garrel, Charles Vial, Ignacio Barros, Benjamin Doerr, Peter Grünwald, Henry Liu, David Martins, Aleksandar Mijatović, Sergej Monavari, Marc Roth, Patrick Schnider, Yannik Schuler, Pim Spelier, Yuuji Tanaka, Ronald van Luijk
机构
*
ETH Zurich(苏黎世联邦理工学院)
;
Aarhus University(奥胡斯大学)
Commentsv2: benchmark expanded from 39 to 77 problems; evaluation extended to 14 models including GPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.6; new analyses (IRT-based score aggregation, inter-rater reliability, tool/token usage, non-agentic ablation); contributor author list updated
Comments41 pages, 7 figures, 7 tables. Preliminary cJSON-only evaluation (N=5 main, N=3 ablation; descriptive statistics, no significance claims). Code and 25-run artifacts at https://github.com/Qiao-Zhiyi/fuzz_agent (tag paper01-arxiv-v1). Venue-version Stage-1 pilot on libxml2, sqlite3, openssl_x509 currently in flight; v2 will report those results
机构
*
Key Laboratory of Intraplate Volcanoes and Earthquakes (China University of Geosciences, Beijing), Ministry of Education, Beijing 100083, China(中国大陆地质大学(北京)构造运动与地震重点实验室,教育部,北京100083,中国)
;
School of Geophysics and Information Technology, China University of Geosciences, Beijing 100083, China(中国大陆地质大学(北京)地球物理与信息技术学院,北京100083,中国)
Resilient Strategies for Stochastic Systems: How Much Does It Take to Break a Winning Strategy?
随机系统中的稳健策略:打破获胜策略需要多大的代价?
Kush Grover, Markel Zubia, Debraj Chakraborty, Muqsit Azeem, Nils Jansen, Jan Kretinsky
机构
*
Ruhr University Bochum Germany
;
Nanyang Technological University, Singapore
;
Technical University of Munich \& University of Konstanz Germany
;
Ruhr University Bochum \& Radboud University Nijmegen Germany
;
Masaryk University Czech Republic
;
Ruhr University Bochum
;
Technical University of Munich \& University of Konstanz
;
Ruhr University Bochum \& Radboud University Nijmegen
;
Masaryk University
CommentsTo appear in Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026), Paphos, Cyprus, May 25-29, 2026
Comments11 pages, 4 figures. Presents a DRL agent that mitigates bufferbloat and achieves near-zero packet loss. Validated via NS-3 simulations under a strict training-testing protocol. Code: https://github.com/aglamazlarefe/DRL-TCP
CommentsRajmund Nagy and Taras Kucherenko contributed equally to this work. To be published in the Proceedings of the 20th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2021), Online, May 3-7, 2021, IFAA-MAS, 3 pages, 1 figure
Deep Q-Networks for Accelerating the Training of Deep Neural Networks
Jie Fu
专题命中
Agent评测
:agent(abstract,comments);分类 cs.LG
CommentsWe choose to withdraw this paper. The DQN itself has too many hyperparameters, which makes it almost impossible to be applied to reasonably large datasets. In the later versions (from v4) with SGDR experiments, it seems that the agent only performs random actions
CommentsThis paper was accepted for presentation at the 7th International Workshop on EXplainable, Trustworthy, and Responsible AI and Multi-Agent Systems (EXTRAAMAS 2025). Workshop website: https://extraamas.ehealth.hevs.ch/index.html
Journal refAutonomous Agents and Multiagent Systems. Best and Visionary Papers. AAMAS 2022. Lecture Notes in Computer Science(), vol 13441, pp. 114-129
Comments13 pages, 37 figures, Paper accepted for presentation A Bayesian model of information cascade in International Workshop on Coordination, Organizations, Institutions, Norms and Ethics for Governance of Multi-Agent Systems (COINE), co-located with AAMAS 2021