Extreme Entropy Machines: Robust information theoretic classification
专题命中 代码评测 :repository(abstract);分类 cs.LG
AI 大模型
代码生成、软件工程智能体、程序修复、测试生成和开发者工具。
专题命中 代码评测 :repository(abstract);分类 cs.LG
专题命中 代码评测 :repository(abstract);分类 cs.AI
Comments ARCOE-Logic 2014 Workshop Notes, pp. 13-24
专题命中 代码评测 :repository(abstract);分类 cs.LG
专题命中 代码评测 :repository(abstract);分类 cs.LG
Comments 8 pages. arXiv admin note: substantial text overlap with arXiv:1403.1946; and text overlap with arXiv:1106.1813 by other authors
Journal ref International Journal of Computer Applications,Vol 69,No 17,pp 28-35,2013
专题命中 代码评测 :repository(abstract);分类 cs.LG
Comments 7 pages
Journal ref World of Computer Science and Information Technology Journal,Vol 3, No 4,pp 70-76,2013
专题命中 代码评测 :repository(abstract);分类 cs.LG
Comments 47 pages, 9 figures; chapter accepted into book 'Support Vector Machine Applications'
专题命中 代码评测 :repository(abstract);分类 cs.LG
Comments 31 pages, 4 figures, 4 tables, Submitted to Pattern Analysis and Applications
专题命中 代码评测 :repository(abstract);分类 cs.CL
Comments 9 pages. This is close to how it appears on the publisher's website (http://bioinformatics.oupjournals.org/cgi/reprint/19/suppl_1/i331) The article wording is the same. Uses bioinformatics-altered.cls, bioinformaticsbib.sty, bioinformaticstitle.sty
Journal ref Bioinformatics Vol. 19 Suppl. 1 2003, pages i331-i339
专题命中 代码评测 :repository(abstract);分类 cs.SE
Comments 4 pages, 2 figures
专题命中 代码评测 :repository(abstract);分类 cs.AI
专题命中 代码评测 :code generation(abstract);分类 cs.PL
Comments 13 pages
面向AI安全评估的对抗语用学:指令冲突、嵌入命令与策略模糊性基准
机构 * Humber Polytechnic(汉博理工学院) ; University of Toronto(多伦多大学)
专题命中 代码评测 :分类 cs.SE、cs.CL、cs.AI;repository(comments)
AI总结 提出对抗语用学基准和标注协议,通过语言学控制的分类法评估模型在指令冲突、嵌入命令等场景下的行为,为安全评估提供实证和方法论工具。
Comments 32-page main paper plus 13-page supplement; 6 figures and 17 tables total; code and data artifact available at the linked repository
现代智能系统中的自我改进:一项综述
机构 * School of Artificial Intelligence, Jilin University(吉林大学人工智能学院) ; King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学) ; University of Alberta(阿尔伯塔大学) ; The Swiss AI Lab IDSIA/USI/SUPSI(瑞士人工智能实验室IDSIA/USI/SUPSI)
专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)
AI总结 综述现代自我改进智能体从研究走向部署,目标是经验驱动的可控进化。提出系统级框架,将智能体视为基础模型与操作支架的耦合配置,自我改进形式化为更新算子,还组织回顾了相关工作、应用、评估等内容并展望未来。
Comments 97 pages, 12 figures. Project page: https://selfimproving-agent.github.io/ Repository: https://github.com/selfimproving-agent/awesome-Self-Improving-Agents
组件消融用于高效混合语言模型架构:性能、鲁棒性和压缩影响
机构 * Doctoral Program in Computer Science, University of Valencia(瓦伦西亚大学计算机科学博士项目)
专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)
AI总结 本文通过组件消融研究混合语言模型,发现注意力机制与替代序列处理路径对性能有显著影响,揭示了模型鲁棒性与压缩优化的关键因素。
Comments 25 pages, 7 figures, 6 tables; revised title, abstract, figures, and data/code repository URL
涌现性失调中的人格模型崩溃
机构 * TELUS Digital Research Hub(TELUS数字研究中心) ; Center for Artificial Intelligence and Machine Learning(人工智能与机器学习中心) ; Institute of Mathematics, Statistics and Computer Science(数学、统计与计算机科学研究所) ; University of São Paulo(圣保罗大学)
专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)
AI总结 提出人格模型崩溃假说,通过道德易感性(S)和道德稳健性(R)两个指标,证明在有害数据上微调大语言模型会导致模型模拟、区分和维持一致角色的内部能力恶化,从而引发涌现性失调。
Comments 23 pages, 7 figures, 7 tables; NeurIPS 2026 submission; Corrected code repository URL
在没有基准的情况下:在无标签的情况下验证比较LLM安全性评分
机构 * Simula Metropolitan Center for Digital Engineering(Simula 数字工程中心) ; Oslo Metropolitan University(奥斯陆 Metropolitan 大学) ; University of Oslo(奥斯陆大学) ; Simula Research Laboratory(Simula 研究实验室) ; Norwegian Directorate of Health(挪威健康 Directorate)
专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)
AI总结 本文提出在无标签情况下验证LLM安全性的方法,通过构建仪器有效性链来替代真实标签,通过实验验证其有效性,并展示了在不同场景下的应用和结果。
Comments SimpleAudit Repository: https://github.com/kelkalot/simpleaudit
通用端到端工具使用强化学习与合成CodeGym
机构 * Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所)
专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)
AI总结 本文提出CodeGym框架,通过合成多样化的多轮工具使用环境,提升LLM代理在不同任务配置下的泛化能力,实验显示Qwen2.5-32B-Instruct在OOD基准测试中准确率提升8.7个百分点。
Comments 24 pages. Accepted to ICLR 2026. Project repository: https://github.com/StigLidu/CodeGym
大模型能帮你清理数据吗?基于大模型的应用级数据准备综述
机构 * Shanghai Jiao Tong University(上海交通大学) ; Tsinghua University(清华大学) ; Microsoft Research(微软研究院) ; MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) ; Shanghai AI Laboratory(上海人工智能实验室) ; Xiaohongshu Inc.(小红书公司) ; Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Alibaba Group(阿里巴巴集团)
专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)
AI总结 本文综述了基于大语言模型的应用级数据准备方法,探讨了数据清洗、整合与丰富三大任务的技术、优势与局限,并提出了可扩展的大语言模型-数据系统和稳健评估协议的未来研究方向。
Comments Please refer to our repository for more details: https://github.com/weAIDB/awesome-data-llm
机构 * University of Chile(智利大学) ; CENIA ; Faculty of mathematics UC(数学学院 UC) ; IMC UC
专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)
Comments 32 pages, 12 figures, repository available
专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)
Comments To appear in NeurIPS 2025. Welcome your submission to challenge our leaderboard at: https://code4db.github.io/parrot-bench/. Also visit our code repository at: https://github.com/weAIDB/PARROT
机构 * Computer Engineering Department, Sharif University of Technology(谢里夫理工大学计算机工程系) ; Iran University of Science and Technology(伊朗科学技术大学) ; Qatar Computing Research Institute(卡塔尔计算研究院)
专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)
Comments 50 main pages, 30 pages appendix, 21 figures, 8 tables, GitHub Repository: https://github.com/llm-lab-org/Generative-AI-for-Character-Animation-Survey
专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)
Comments Journal of Open Source Software; LangFair repository: https://github.com/cvs-health/langfair
Journal ref Journal of Open Source Software, 10(105), 7570 (2025)
专题命中 代码评测 :分类 cs.CL、cs.AI、cs.LG;repository(comments)
Comments Final manuscript published in Language Development Research under CC BY-NC-SA 4.0. Pre-print redistributed through arXiv with permission. Replaces corrupted PsyArXiv pre-print repository at https://psyarxiv.com/37zna
Journal ref Language Development Research, 1(1), 123-191 (2021)
计算机视觉中的持续测试时间适应:方法、基准和未来方向
机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校) ; LIVIA ETS Montreal, ILLS International Laboratory on Learning Systems (ILLS)(蒙特利尔LIVIA ETS,学习系统国际实验室(ILLS)) ; Tulane University(路易斯安那州立大学) ; Nanyang Technological University(南洋理工大学) ; Hanyang University(翰阳大学)
专题命中 代码评测 :repository(abstract)
AI总结 本文针对计算机视觉中训练与测试数据分布不同的问题,定义CTTA问题,分析持续域转移模式,提出分层分类法将现有方法分为三类,回顾代表性方法并展示实验结果,讨论局限性与新兴方向,为持续测试时间适应研究提供路线图。
Comments TMLR 2026 (July edition)
基于课程学习的生成推荐系统中的标记加权多目标学习
专题命中 代码评测 :repository(abstract)
AI总结 本文提出基于课程学习的生成推荐系统中的标记加权多目标学习方法,通过两种信息增益策略提升推荐性能,实验表明其在不同语义ID构造下表现优异。
Comments 12 pages, 3 figures
几何湍流:用于股权协方差动态的测地回归危机指标——来自非洲市场的证据
专题命中 代码评测 :repository(abstract)
AI总结 该研究提出一种基于对称正定锥测地回归的几何湍流指标,可准确识别市场压力事件,其预测性能优于传统欧氏回归,对应的去风险策略能降低最大回撤。
Comments 22 pages, 8 figures
CNOT与Clifford电路的启发式及最优综合
专题命中 代码评测 :repository(abstract)
AI总结 本研究针对CNOT与Clifford电路,提出最优、A*及贪心三类综合算法,可最小化两量子比特门数量或电路深度,经基准测试性能优于现有方法,相关算法已在GitHub仓库开源。
Comments Accepted in Quantum, 6 Aug 2026
Edit2TikZ:面向基于TikZ的科学图像编辑的全面且具有挑战性的基准
专题命中 代码评测 :code generation(abstract)
AI总结 本文推出科学图像编辑基准Edit2TikZ,评估主流多模态大语言模型发现其性能不足,通过构建混合训练集并采用课程学习,可显著提升紧凑模型的编译成功率。
Comments 9 pages, 6 figures, work in progress
AMELI: 镧系离子的角矩阵元
专题命中 代码评测 :repository(abstract)
AI总结 提出基于Slater行列式基和Racah分类的通用框架,计算f^N组态中任意球张量算符的角矩阵元,并构建开源数据库AMELI,以精确算术消除浮点误差,替代传统半经验计算表格。
Comments v2: Abstract, introduction and conclusion rewritten to clarify the novelty and intention of this work. Added Wybourne crystal field Hamiltonians in Section IV G. Improved phase synchronization in Section V B. New Section IX "Application Examples" v3: Minor errors and typographic flaws corrected
Journal ref J. Chem. Phys. 165, 064308 (2026)
VIGIL:应对图像再上下文化中的幻觉检测
机构 * Wroclaw University of Science and Technology(沃拉茨拉夫大学科学与技术学院)
专题命中 代码评测 :repository(abstract)
AI总结 VIGIL通过细粒度分类填补多模态模型图像再上下文化中幻觉检测的空白,提出多阶段检测流程并公开资源促进透明度。
Comments 19 pages, 8 figures, 7 tables. Code and data are available at: https://github.com/mlubneuskaya/vigil and https://huggingface.co/datasets/joannaww/VIGIL