Journal refZang, X., Jiang, Z., Cheng, J. et al. Instruction-based image editing: a survey on data, models, evaluation, and applications. Vicinagearth 3, 3 (2026)
机构
*
College of Geodesy and Geomatics, Shandong University of Science and Technology(山东科技大学测绘与地理信息学院)
;
School of Environmental Science and Spatial Informatics, China University of Mining and Technology(中国矿业大学环境与测绘学院)
;
State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(武汉大学测绘遥感信息工程国家重点实验室)
;
Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
School of Automation, Southeast University(东南大学自动化学院)
;
Department of Geography, National University of Singapore(新加坡国立大学地理系)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
E-Bench:在现实世界产品场景中对多步工具使用智能体进行基准测试
Weihuang Zheng, Tianyuan Zou, Eileen Ye, Alphet Liu, Youyong Kong, Ya-Qin Zhang, Duran Zheng, Maxm Pan
机构
*
Hunyuan Team, Tencent(腾讯混元团队)
;
Institute for AI Industry Research, Tsinghua University(清华大学人工智能产业研究院)
;
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI
TLA+-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA Specification Generation
TLA$^{+}$-Bench:用于自然语言到TLA+规范生成的基于执行的基准测试和数据集
Arslan Bisharat, Eric Spencer, Brian Ortiz, Khushboo Bhadauria, Mujtaba Nazari, Beatriz Santos, Anisa Ramos, TaiNing Wang, George K. Thiruvathukal, Konstantin Läufer, Mohammed Abuhamad
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI
Comments17 pages, appendix included. Introduces TLA+-Bench, an execution-grounded benchmark and dataset for natural-language to TLA$^{+}$ specification generation. Dataset, evaluation code, and model outputs available at publication
机构
*
School of Cyberspace Science and Technology, Beijing Institute of Technology(航天信息科学技术学院,北京理工大学)
;
School of Computer Science, The University of Auckland(计算机科学学院,奥克兰大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI
机构
*
International Business Machines (IBM)(国际商业机器公司(IBM))
;
Global Atlantic Financial(全球大西洋金融公司)
;
Docusign(DocuSign公司)
;
Salesforce Inc(Salesforce公司)
;
The Home Depot(家得宝公司)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI