TextMineX: Data, Evaluation Framework and Ontology-guided LLM Pipeline for Humanitarian Mine Action
TextMineX: 用于人道主义排雷行动的数据集、评估框架及基于本体的LLM流水线
Chenyue Zhou, Gürkan Solmaz, Flavio Cirillo, Kiril Gashteovski, Jonathan Fürst
机构
*
NEC Laboratories Europe(NEC欧洲实验室)
;
University of Stuttgart(斯图加特大学)
;
VAGO Solutions(VAGO解决方案)
;
Zurich University of Applied Sciences(苏黎世应用科学大学)
;
CAIR, Ss. Cyril and Methodius University of Skopje, North Macedonia(CAIR,斯科普耶塞尔维亚·梅托迪乌斯大学,北马其顿)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
LLM Driven Design of Continuous Optimization Problems with Controllable High-level Properties
基于大语言模型的连续优化问题可控高阶属性设计
Urban Skvorc, Niki van Stein, Moritz Seiler, Britta Grimme, Thomas Bäck, Heike Trautmann
机构
*
Machine Learning and Optimization, Paderborn University, Germany(机器学习与优化,帕德博恩大学,德国)
;
Leiden Institute of Advanced Computer Science, Leiden University, The Netherlands(莱顿先进计算机科学研究所,莱顿大学,荷兰)
;
Data Management and Biometrics, University of Twente, The Netherlands(数据管理与生物特征,埃因霍温大学,荷兰)
专题命中
评测与基准
:LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
机构
*
North China Electric Power University(华北电力大学)
;
Research Institute, China Unicom(中国联合研究院)
;
Wuhan University(武汉大学)
;
Beihang University(北京航空航天大学)
;
Nanyang Technological University(南洋理工大学)
Automated structural testing of LLM-based agents: methods, framework, and case studies
基于大语言模型的智能体的自动化结构测试:方法、框架与案例研究
Jens Kohl, Otto Kruse, Youssef Mostafa, Andre Luckow, Karsten Schroer, Thomas Riedl, Ryan French, David Katz, Manuel P. Luitz, Tanrajbir Takher, Ken E. Friedl, Céline Laurent-Winter
Comments10 pages, 5 figures. Preprint of an accepted paper at IEEE BigData 2025 (main track). Source code for the introduced methods and framework available at https://github.com/awslabs/generative-ai-toolkit
Exploring Graph Learning Tasks with Pure LLMs: A Comprehensive Benchmark and Investigation
探索纯大语言模型在图学习任务中的应用:全面的基准测试与调查
Yuxiang Wang, Xinnan Dai, Wenqi Fan, Yao Ma
机构
*
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Michigan State University(密歇根州立大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Rensselaer Polytechnic Institute(纽约理工学院)
专题命中
评测与基准
:large language model(abstract);language model(abstract);instruction tuning(abstract);分类 cs.LG
Language Models are Symbolic Learners in Arithmetic
语言模型在算术中是符号学习者
Chunyuan Deng, Zhiqi Li, Roy Xie, Ruidi Chang, Hanjie Chen
机构
*
Department of Computer Science(计算机科学系)
;
Rice University(里士满大学)
;
College of Computing(计算学院)
;
Georgia Institute of Technology(佐治亚理工学院)
;
Duke University(杜克大学)
Meaning Is Not A Metric: Using LLMs to make cultural context legible at scale
意义并非一种度量:利用大语言模型在大规模AI社会技术系统中使文化背景可理解
Cody Kommers, Drew Hemment, Maria Antoniak, Joel Z. Leibo, Hoyt Long, Emily Robinson, Adam Sobey
机构
*
The Alan Turing Institute(艾伦·图灵研究所)
;
University of Edinburgh(爱丁堡大学)
;
University of Colorado Boulder(科罗拉多大学丹佛分校)
;
University of Chicago(芝加哥大学)
;
University of Exeter(埃克塞特大学)
;
University of Southampton(南安普顿大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Veri-Sure: A Contract-Aware Multi-Agent Framework with Temporal Tracing and Formal Verification for Correct RTL Code Generation
Veri-Sure:一种具有时间追踪和形式验证的合同感知多智能体框架,用于正确RTL代码生成
Jiale Liu, Taiyu Zhou, Tianqi Jiang
机构
*
School of Physics and Astronomy, The University of Edinburgh, Edinburgh, UK(物理与天文学院,爱丁堡大学,爱丁堡,英国)
;
State Key Laboratory of Analog and Mixed-Signal VLSI, University of Macau, Macau(模拟与混合信号VLSI国家重点实验室,澳门大学,澳门)
;
School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, Shenzhen, China(科学与工程学院,香港中文大学(深圳))
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI
Whitespaces Don't Lie: Feature-Driven and Embedding-Based Approaches for Detecting Machine-Generated Code
空格并不撒谎:基于特征和嵌入的方法用于检测机器生成的代码
Syed Mehedi Hasan Nirob, Shamim Ehsan, Moqsadur Rahman, Summit Haque
机构
*
Computer Science and Engineering(计算机科学与工程)
;
Shahjalal University of Science and Technology(沙赫拉尔大学科学与技术)
;
University of Texas at El Paso(德克萨斯大学埃尔帕索分校)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.LG
UniPCB: A Unified Vision-Language Benchmark for Open-Ended PCB Quality Inspection
UniPCB: 一个统一的视觉-语言基准用于开放式PCB质量检测
Fuxiang Sun, Xi Jiang, Jiansheng Wu, Haigang Zhang, Feng Zheng, Jinfeng Yang
机构
*
Shenzhen Polytechnic University(深圳职业技术大学)
;
Southern University of Science and Technology(南方科技大学)
;
University of Science and Technology Liaoning(辽宁科技大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.AI