2026-07论文中文摘要
KernelBench验证:大语言模型生成的内核真的能超越PyTorch吗?
KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch?
arXiv 2607.16241 · 2026-07-21偏好优化的归一化奖励
Normalized Rewards for Preference Optimization
arXiv 2607.16240 · 2026-07-21BACON:用于多人工智能评判器建模与评估的预算人类校准
BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges
arXiv 2607.16239 · 2026-07-21用于液滴演化预测的扩散校正自回归傅里叶神经算子
Diffusion-corrected Autoregressive Fourier Neural Operator for Droplet Evolution Prediction
arXiv 2607.16238 · 2026-07-21量化递归推理模型
Quantizing Recursive Reasoning Models
arXiv 2607.16237 · 2026-07-21基于边际影响的全局时间序列解释归因方法的失效
The Failures of Marginal Influence-Based Attribution Methods for Global Time Series Explanations
arXiv 2607.16236 · 2026-07-21汉塔病毒监测:用于汉塔病毒基因组监测的联邦学习
HantaWatch: Federated Learning for Hantavirus Genomic Surveillance
arXiv 2607.16234 · 2026-07-21用于乳腺癌亚型分类和生存预测的具有对比多任务学习的令牌级跨模态变换器
Token-Level Cross-Modal Transformer with Contrastive Multi-Task Learning for Breast Cancer Subtype Classification and Survival Prediction
arXiv 2607.16233 · 2026-07-21从权重到文字:用自然语言表达和编辑偏好模型推理
From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language
arXiv 2607.16232 · 2026-07-21正交梯度约束塑造噪声标签记忆动态
Orthogonal Gradient Constraints Shape Noisy-Label Memorization Dynamics
arXiv 2607.16231 · 2026-07-21FinBench:用于智能金融预测的时间门控校准与不确定性基准测试
FinBench: Time-Gated Calibration and Uncertainty Benchmarking for Agentic Financial Forecasting
arXiv 2607.16229 · 2026-07-21用于网络防御的大语言模型遗忘:方法、挑战及新出现威胁的综述
LLM Unlearning for Cyber Defense: A Survey on Methods, Challenges, and Emerging Threats
arXiv 2607.16227 · 2026-07-21随机和自然地震噪声对用于在2025年堪察加半岛地震前恢复低震级地震活动的波形互相关性能的影响
Effects of stochastic and natural seismic noise on the performance of waveform cross-correlation used to recover low-magnitude seismicity prior to the July 29, 2025, Kamchatka earthquake
arXiv 2607.16226 · 2026-07-21克服脑机接口校准瓶颈:一种基于黎曼对齐和随机权重平均的临床基础架构
Overcoming the BCI Calibration Bottleneck: A Clinically-Grounded Architecture using Riemannian Alignment and Stochastic Weight Averaging
arXiv 2607.16225 · 2026-07-21限制前沿人工智能的国际协议:目标与退出
International Agreements to Limit Frontier AI: Objectives and Exit
arXiv 2607.16224 · 2026-07-21用于设备端机器学习的全传感器智能眼镜平台
Fully-sensorized smart-eyewear platform for on-device Machine Learning
arXiv 2607.16222 · 2026-07-21使用卷积神经网络比较用于异常心音检测的频谱图前端
Comparing Spectrogram Front-Ends for Abnormal Heart-Sound Detection with a Convolutional Neural Network
arXiv 2607.16220 · 2026-07-21MDAF:一种用于自动外交政策分析的多维注释框架
MDAF: A Multi-Dimensional Annotation Framework for Automated Foreign Policy Analysis
arXiv 2607.16219 · 2026-07-21一种用于社区感知犯罪热点检测的模糊逻辑框架:城市计算平台的原型应用架构与探索性验证
A Fuzzy Logic Framework for Community-Aware Crime Hotspot Detection: Prototype Application Architecture and Exploratory Validation for an Urban Computing Platform
arXiv 2607.16218 · 2026-07-21慢变量带噪声的随机快-慢系统中依赖曲率的路径集中
Curvature-Dependent Path Concentration in Stochastic Fast-Slow Systems with Noise on the Slow Variable
arXiv 2607.16217 · 2026-07-21全息共形场论热力学作为弱引力与弱宇宙审查猜想之间关于爱因斯坦-麦克斯韦-幂-杨-米尔斯-反德西特黑洞的桥梁
Holographic CFT Thermodynamics as a Bridge Between the Weak Gravity and Weak Cosmic Censorship Conjectures for EMPYM--AdS Black Holes
arXiv 2607.16216 · 2026-07-21RAIL Guard:弥合大语言模型智能体负责任人工智能中评估与修复的差距
RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents
arXiv 2607.16215 · 2026-07-21是什么让语言表征成为人类大脑中高级视觉感知的良好模型?
What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?
arXiv 2607.16214 · 2026-07-21SelKV:基于逐令牌合并或丢弃及注意力补偿的选择性键值缓存合并
SelKV: Selective KV Cache Merging with Per-Token Merge-or-Drop and Attention Compensation
arXiv 2607.16213 · 2026-07-21符号增强弥补神经事实核查器中规范等价盲点
Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers
arXiv 2607.16212 · 2026-07-21用于大语言模型智能体的准确高效长期记忆
Accurate and Efficient Long-Term Memory for LLM Agents
arXiv 2607.16211 · 2026-07-21强化学习策略验证的综述
A Survey on the Verification of Reinforcement Learning Policies
arXiv 2607.16210 · 2026-07-21Shapley上下文剪枝:一种用于上下文重排和剪枝的合作博弈视角
Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning
arXiv 2607.16209 · 2026-07-21ColGraphRAG:用于多模态GraphRAG的后期交互证据检索
ColGraphRAG: Late-Interaction Evidence Retrieval for Multimodal GraphRAG
arXiv 2607.16208 · 2026-07-21PPO-HSC:基于广域策略覆盖优化的探索性强化学习框架
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization
arXiv 2607.16206 · 2026-07-21需要8个令牌:通过辅助分支实现从弱到强的离策略强化学习
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
arXiv 2607.16205 · 2026-07-21DocOCR-Eval:一种无需真实标签的基于校正的OCR工具选择框架
DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
arXiv 2607.16203 · 2026-07-21通过小语言模型实现人工智能民主化:面向本地部署的结构化基准测试和参数高效微调
Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Efficient Fine-Tuning for Local Deployment
arXiv 2607.16202 · 2026-07-21生成式本体归纳:使用大语言模型从文档语料库中进行领域无关的模式发现
Generative Ontology Induction: Domain-Agnostic Schema Discovery from Document Corpora Using Large Language Models
arXiv 2607.16201 · 2026-07-21人工智能代理系统的确定性重放
Deterministic Replay for AI Agent Systems
arXiv 2607.16200 · 2026-07-21PlanFlip:通过规划阶段提示注入攻击多智能体大语言模型系统
PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection
arXiv 2607.16199 · 2026-07-21基于图神经网络的链接预测综述:技术、应用与挑战
A Survey on GNN-based Link Prediction: Techniques, Applications, and Challenges
arXiv 2607.16198 · 2026-07-21一些大语言模型表现出一致的风险态度
Some Large Language Models Exhibit Consistent Risk Attitudes
arXiv 2607.16197 · 2026-07-21基于灰色关联系数增强的强化学习引导的NSGA-II在多目标优化中的应用:以纳斯达克投资组合优化为例
Reinforcement Learning-Guided NSGA-II Enhanced with Gray Relational Coefficient for Multi-Objective Optimization: Application to NASDAQ Portfolio Optimization
arXiv 2607.16194 · 2026-07-21通过富集范畴实现参数化量子电路语义
Parameterized Quantum Circuit Semantics Through Enriched Categories
arXiv 2607.16114 · 2026-07-21全局域阿贝尔扩张的等分布
Equidistribution for abelian extensions of global fields
arXiv 2607.16079 · 2026-07-21循环循环者!
Loop the Loopies!
arXiv 2607.16051 · 2026-07-21PIXIE:用于具有装配缺陷的未见物体的零样本纹理不变6D姿态估计框架
PIXIE: A Zero-Shot texture-invariant 6D pose estimation framework for unseen objects with assembly defects
arXiv 2607.16015 · 2026-07-21扩散诱导的不稳定性促进生态进化网络中的合作
Diffusion-induced instabilities promote cooperation in eco-evolutionary networks
arXiv 2607.15989 · 2026-07-21快速升温步骤作为对Tool-Narayanaswamy形式主义的检验
Fast temperature up steps as a test of the Tool-Narayanaswamy formalism
arXiv 2607.15988 · 2026-07-21多智能体再保险链中的均衡分析
Equilibrium analysis in a multi-agent reinsurance chain
arXiv 2607.15962 · 2026-07-21双向归纳:掩码扩散语言模型中上下文学习的机制分析
Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models
arXiv 2607.15893 · 2026-07-21DECODEM:通过增强方法从公司组织文档中提取数据
DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods
arXiv 2607.15879 · 2026-07-21用于SAT和XSAT的双线程覆盖蒙特卡洛树搜索
Two-Thread Coverage MCTS for SAT and XSAT
arXiv 2607.15834 · 2026-07-21全同态加密的密文和多项式级优化
Ciphertext- and Polynomial-Level Optimization for Fully Homomorphic Encryption
arXiv 2607.15750 · 2026-07-21