2026-07论文中文摘要
相同问题,不同答案:超越准确率评估大语言模型的可靠性
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
arXiv 2607.22554 · 2026-07-28评估审稿人指南设计对基于大语言模型的自动同行评审的影响
Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review
arXiv 2607.22553 · 2026-07-28MioFFAn:一种具有大语言模型自动化能力的公式形式化注释软件
MioFFAn: an Annotation Software for Formula Formalization with LLM Automation Capabilities
arXiv 2607.22552 · 2026-07-28具有非光滑不等式约束的类星体凸优化问题的镜像下降方法
Mirror Descent Methods for Quasar Convex Optimization Problems With Non-Smooth Inequality Constraints
arXiv 2607.22551 · 2026-07-28大规模学习优化:用于随机组合优化的Benders分解-Transformer框架
Learning to Optimize at Scale: A Benders Decomposition-TransfORmers Framework for Stochastic Combinatorial Optimization
arXiv 2607.22550 · 2026-07-28QFoldAgent:一种用于蛋白质结构预测的自主量子优化多智能体系统
QFoldAgent: An Autonomous Quantum Optimization Multi-Agent System for Protein Structure Prediction
arXiv 2607.22549 · 2026-07-28SeT-Diff:面向高性能计算遥测和时间序列的语义基础模型
SeT-Diff: Towards Semantic Foundation Models for HPC Telemetry and Time-Series
arXiv 2607.22548 · 2026-07-28使用模糊高斯过程的贝叶斯模糊优化
Bayesian fuzzy optimization using fuzzy Gaussian process
arXiv 2607.22547 · 2026-07-28解释GAND:关于性别模糊自然数据与对比归因的资源
Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution
arXiv 2607.22546 · 2026-07-28Semalith v1.4:一个经过校准的1.84亿参数安全分类器,在参数比Llama - Guard - 3 - 8B少44倍的情况下实现了先进的提示注入检测
Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B
arXiv 2607.22545 · 2026-07-28基于扩散模型的基于概念的视觉反事实解释
Concept-based Visual Counterfactual Explanations with Diffusion Models
arXiv 2607.22544 · 2026-07-28一种用于具有负载相关成本的中国邮递员问题的混合启发式框架
A Hybrid Matheuristic Framework for the Chinese Postman Problem with Load-Dependent Costs
arXiv 2607.22542 · 2026-07-28随机微分方程引导的蒙特卡罗强化学习:一种用于噪声环境中稳健决策的随机最大值原理方法
SDE Guided Monte Carlo Reinforcement Learning: A Stochastic Maximum Principle Approach for Robust Decision Making in Noisy Environments
arXiv 2607.22541 · 2026-07-28去中心化交易网络中的多路径路由:凸分配与改进路径证书
Multi-Path Routing in Decentralized Exchange Networks: Convex Allocation and an Improving-Path Certificate
arXiv 2607.22540 · 2026-07-28比较放射治疗计划的优化模型
Comparing Optimization Models for Radiotherapy Scheduling
arXiv 2607.22539 · 2026-07-28基于模型的无导数优化的低秩KKT更新及并行翻转机制
Low-Rank KKT Updates and a Parallel Flipping Mechanism for Model-Based Derivative-Free Optimization
arXiv 2607.22538 · 2026-07-28具有不可分离哈密顿量的有限状态平均场博弈的唯一性结果
A uniqueness result for finite-state mean field games with non-separable Hamiltonian
arXiv 2607.22537 · 2026-07-28具有切换偏好的最优投资
Optimal Investment with Switching Preferences
arXiv 2607.22536 · 2026-07-28不透明的认知中介:大型语言模型部署配置如何塑造伪科学的验证
Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
arXiv 2607.22513 · 2026-07-28用于状态条件波动率预测的易感性蓄水池架构
Susceptible Reservoir Architectures for Regime-Conditional Volatility Forecasting
arXiv 2607.22491 · 2026-07-28用于控制阀静摩擦检测的最优传输图像表示与深度协方差对齐(CORAL)
Optimal Transport Image Representation and Deep Covariance Alignment (CORAL) for Control Valve Stiction Detection
arXiv 2607.22486 · 2026-07-28来自全局单极子宇宙演化的南布-戈德斯通玻色子发射
Nambu-Goldstone emissions from the cosmological evolution of global monopoles
arXiv 2607.22481 · 2026-07-28关于受控世界模型的可识别性
On the Identifiability of Controlled World Models
arXiv 2607.22430 · 2026-07-28ANDES,用于极大望远镜的高分辨率光谱仪:YJH光谱仪的设计与性能分析
ANDES, the high-resolution spectrograph for the ELT: design and performance analysis of the YJH spectrograph
arXiv 2607.22388 · 2026-07-28海王星:用于复杂多相流基准测试的综合机器学习框架
Neptuna: A Comprehensive Machine Learning Framework for Benchmarking Complex Multiphase Flows
arXiv 2607.22280 · 2026-07-28关于黑塞猜想的一个五变量反例,以及雅可比猜想和黑塞猜想的低维状态
A five-variable counterexample to the Hessian conjecture, and the low-dimensional status of the Jacobian and Hessian conjectures
arXiv 2607.22198 · 2026-07-28GLI-AL:一个具有统一解剖-病变标签的多模态胶质瘤MRI标签资源
GLI-AL: A Multi-Modal Glioma MRI Label Resource with Unified Anatomy-Lesion Labels
arXiv 2607.22135 · 2026-07-28南贝格4.2 - 3B:以紧凑模式解锁智能体能力
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model
arXiv 2607.22083 · 2026-07-28战略退出与单边控制
Strategic Exit and Unilateral Control
arXiv 2607.21898 · 2026-07-28闭环:自回归生成渲染的无训练重访一致性
Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering
arXiv 2607.21848 · 2026-07-28基于泡利本征态的高效不可克隆加密
Efficient Unclonable Encryption from Pauli Eigenstates
arXiv 2607.21811 · 2026-07-28版权法合规性的智能体评估
Agentic Evaluation of Copyright Law Compliance
arXiv 2607.21799 · 2026-07-28DONDO:用于非洲语言的开放w2v-BERT语音识别基础模型
DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages
arXiv 2607.21540 · 2026-07-28一种片上可编程机械量子换能器
An on-chip programmable mechano-quantum transducer
arXiv 2607.21487 · 2026-07-28洛伦兹几何中时间顺序钻石的加倍条件
Doubling for chronological diamonds in Lorentzian geometry
arXiv 2607.21423 · 2026-07-28资本市场大语言模型可靠性评分(CM-LRS):从似是而非到可信赖
Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable
arXiv 2607.21340 · 2026-07-28利用数字孪生支持的多无线电微波回传赋能农村地区实现基于5G IAB的FWA
Empowering Rural Areas with Multi-radio Microwave Backhaul Supported by Digital Twin for 5G IAB-based FWA
arXiv 2607.21310 · 2026-07-28专家行为先验强化学习
Expert Behavior Prior Reinforcement Learning
arXiv 2607.21302 · 2026-07-28V-DEAL:将视频安全去校准诊断为理解-拒绝耦合失败
V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure
arXiv 2607.21151 · 2026-07-28VibeVoice-ASR-BitNet技术报告
VibeVoice-ASR-BitNet Technical Report
arXiv 2607.21075 · 2026-07-28HyperImageNet:一个用于细粒度高光谱土地覆盖理解的大规模基准
HyperImageNet: A Large-Scale High-Spatial Resolution Hyperspectral Imagery Classification Benchmark
arXiv 2607.21050 · 2026-07-28EmoAgent-R1:基于强化学习的动态智能体专业化实现多模态情感理解
EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization
arXiv 2607.21013 · 2026-07-28光子集成电路中的偏振矢量共轭
Polarization vector conjugation in a photonic integrated circuit
arXiv 2607.20980 · 2026-07-28沉默的权重:潜在国际象棋推理中权重优于暂存器的因果案例
The Weight of Silence: A Causal Case for Weights Over the Scratchpad in Latent Chess Reasoning
arXiv 2607.20952 · 2026-07-28MatDiffract:用于X射线粉末衍射的材料信息自动分析平台
MatDiffract: A Material-Informed Automated Analysis Platform for X-ray Powder Diffraction
arXiv 2607.20880 · 2026-07-28固定 D3Q125 速度集上的分层对数高斯松弛
Hierarchical Log-Gaussian Relaxation on a Fixed D3Q125 Velocity Set
arXiv 2607.20846 · 2026-07-28基于轻量级时间卷积网络的高效且可解释的身体情感识别
Efficient and Interpretable Body-Based Emotion Recognition with Lightweight Temporal Convolutional Networks
arXiv 2607.20820 · 2026-07-28基于差分动作单元的可解释图注意力网络用于压力识别(StressGAT)
Explainable graph attention network for stress recognition (StressGAT) via differential action units
arXiv 2607.20819 · 2026-07-28混合专家VLA中的涌现组合技能
Emergent Compositional Skills in Mixture-of-Experts VLAs
arXiv 2607.20771 · 2026-07-28块模型中的组件结构与渗流
Component structure and percolation in block models
arXiv 2607.20719 · 2026-07-28