2026-07论文中文摘要
语言模型的经济评估
Economic Evaluations of Language Models
arXiv 2607.19375 · 2026-07-23Euclean:在Lean中通过统一验证实现几何问题的自动形式化
Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean
arXiv 2607.19374 · 2026-07-23通过表示对齐减轻苏格拉底式导师中的支架坍塌
Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment
arXiv 2607.19371 · 2026-07-23Aircast-Mars:一种用于全球天气预报的火星基础模型,采用HEALPix感知卷积
Aircast-Mars: A Mars Foundation Model for Global Weather Forecasting with HEALPix-Aware Convolutions
arXiv 2607.19370 · 2026-07-23超越跟踪或捷径:扑克自回归模型中组合受限的预测状态
Beyond Tracking or Shortcut: Composition-Bounded Predictive States in Poker Autoregressive Models
arXiv 2607.19369 · 2026-07-23谱局部敏感哈希:通过Krylov投影局部敏感哈希实现次二次方的提示压缩
Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing
arXiv 2607.19368 · 2026-07-23重新思考大语言模型中的不确定性评估
Rethinking Uncertainty Evaluation in Large Language Models
arXiv 2607.19367 · 2026-07-23用于大语言模型安全分类的几何引导约束学习
Geometry-Guided Constraint Learning for LLM Safety Classification
arXiv 2607.19366 · 2026-07-23使用回答集编程和大语言模型进行逻辑引导的数据提取
Logic-Guided Data Extraction with Answer Set Programming and Large Language Models
arXiv 2607.19365 · 2026-07-23GraphContainer:用于比较和调试图RAG方法的统一平台
GraphContainer: A Unified Platform for Comparing and Debugging Graph RAG Methods
arXiv 2607.19362 · 2026-07-23多轮大语言模型系统的有状态防护栏:一个对话风险累积框架
Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework
arXiv 2607.19361 · 2026-07-23用于大语言模型智能体的概要图内存:通过叙事概要进行隐式跨实体遍历
Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles
arXiv 2607.19359 · 2026-07-23LISA:用于高效长上下文推理的线性索引稀疏注意力
LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning
arXiv 2607.19358 · 2026-07-23多目标生成推荐系统的随机原始对偶解码
Stochastic Primal-Dual Decoding for Multiobjective Generative Recommender Systems
arXiv 2607.19357 · 2026-07-23NEXUS:工具使用型大语言模型智能体的结构化运行时安全
NEXUS: Structured Runtime Safety for Tool-Using LLM Agents
arXiv 2607.19356 · 2026-07-23大语言模型中的信息辨别
Information Discernment in Large Language Models
arXiv 2607.19355 · 2026-07-23FormulaSPIN:用于自然语言到电子表格公式生成的自博弈微调
FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation
arXiv 2607.19354 · 2026-07-23在英特尔 TDX 下对 NVIDIA H100 上的机密 GPU 推理进行基准测试
Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX
arXiv 2607.19353 · 2026-07-23验证单项卡哇伊度量
Validating the Single Item Kawaii Measure
arXiv 2607.19352 · 2026-07-23OpenEvoShield:面向开放世界多智能体系统攻击的双重非平稳持续防御
OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks
arXiv 2607.19351 · 2026-07-23用于稳健金融欺诈检测和对抗弹性的混合LSTM-图神经网络框架
Hybrid LSTM-Graph Neural Framework for Robust Financial Fraud Detection and Adversarial Resilience
arXiv 2607.19350 · 2026-07-23FineServe:一个细粒度的数据集以及对全球大语言模型服务工作负载的特征描述
FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads
arXiv 2607.19349 · 2026-07-23卷积神经网络中的联想情感学习
Associative Emotional Learning in Convolutional Neural Networks
arXiv 2607.19327 · 2026-07-23NAPTIME:用于鲁宾警报分类的神经过程框架
NAPTIME: A Neural-Process Framework for Rubin Alert Classification
arXiv 2607.19236 · 2026-07-23基于结构化场景知识和可验证推理-行动一致性的自动驾驶认知双过程规划
Cognitive Dual-Process Planning for Autonomous Driving with Structured Scene Knowledge and Verifiable Reasoning-Action Consistency
arXiv 2607.19194 · 2026-07-23探索等离子体中带电尘埃二聚体的自组织
Exploring Self-Organization of Charged Dust Dimers in Plasma
arXiv 2607.19180 · 2026-07-23控制中的一致性:通过成本统一连接多核映射与路由
Coherence in Control: Bridging Many-Core Mapping and Routing through Cost Unification
arXiv 2607.19158 · 2026-07-23Mage-Flow:用于图像生成和编辑的高效原生分辨率基础模型
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
arXiv 2607.19064 · 2026-07-23现在你看到了仇恨:用于隐藏仇恨幻觉的自适应视图检索
Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions
arXiv 2607.19061 · 2026-07-23股票的可观测矩阵动力学
Observable Matrix Dynamics of Stocks
arXiv 2607.19005 · 2026-07-23非中心对称超导 NbRe 薄膜中退火增强的自旋轨道效应
Annealing-enhanced spin-orbit effects in non-centrosymmetric superconducting NbRe films
arXiv 2607.18994 · 2026-07-23通过退火梯度下降增强神经量子态
Enhanced Neural Quantum State via Annealed Gradient Descent
arXiv 2607.18865 · 2026-07-23TSGR:淘宝搜索生成式检索
TSGR: Taobao Search Generative Retrieval
arXiv 2607.18796 · 2026-07-23RoboInter1.5:用于具身世界建模和机器人操作的整体中间表示套件
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
arXiv 2607.18709 · 2026-07-23为不可见之物设计:支持盲人和低视力音乐家的弦乐学习
Designing for What Cannot Be Seen: Supporting Embodied String Learning for Musicians with Blindness and Low-Vision
arXiv 2607.18598 · 2026-07-23HALO:科学假设生成中的交互式协同溯因推理
HALO: Interactive Co-abductive Reasoning in Scientific Hypothesis Generation
arXiv 2607.18564 · 2026-07-23利用高分辨率图像约束光曲线建模对微引力透镜行星系统OGLE-2014-BLG-0676L进行特征描述
Characterizing Microlensing Planetary System OGLE-2014-BLG-0676L with High-Resolution Image Constrained Light Curve Modeling
arXiv 2607.18408 · 2026-07-23迈向用于四足机器人运动的扭矩驱动强化学习
Towards Torque-Driven Reinforcement Learning for Quadruped Locomotion
arXiv 2607.18365 · 2026-07-23随机奇偶博弈的值迭代
Value Iteration for Stochastic Parity Games
arXiv 2607.18355 · 2026-07-23FSDBN:通过动态脑网络实现前景感知的脑电图-视觉对齐
FSDBN: Foreground-Aware EEG-Visual Alignment via Dynamic Brain Networks
arXiv 2607.18344 · 2026-07-23ChemHyperMag:基于物理信息的磁超图学习改进分子ADMET预测
ChemHyperMag: Physics-informed magnetic hypergraph learning improves molecular ADMET prediction
arXiv 2607.18332 · 2026-07-23重要的不是你说了什么,而是你怎么说:评估大语言模型对信念表达的回应
It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief
arXiv 2607.18232 · 2026-07-23调和分析中的离散类似物:多参数拉东平均值
Discrete analogues in harmonic analysis: Multi-parameter Radon averages
arXiv 2607.18160 · 2026-07-23艾萨克模拟到现实:基于强化学习的四足动物运动
Isaac Sim-to-Real: Reinforcement Learning based Locomotion for Quadrupeds
arXiv 2607.18135 · 2026-07-23具有速度依赖阻尼的欧拉-泊松方程
Euler-Poisson equations with velocity-dependent damping
arXiv 2607.18035 · 2026-07-23AdaHome:一种使用本地小语言模型的自适应智能家居助手
AdaHome: An Adaptive Smart Home Assistant using Local Small Language Models
arXiv 2607.18034 · 2026-07-23用于平面光学中最优光束整形的物理信息神经网络
Physics-Informed Neural Networks for Optimal Beam Shaping in Flat Optics
arXiv 2607.18012 · 2026-07-23用大语言模型智能体对概念擦除进行压力测试
Stress Testing Concept Erasure with Large Language Model Agents
arXiv 2607.17890 · 2026-07-23PRiSM:少样本视觉语言模型的原型正则化
PRiSM: Prototype Regularization for Few-Shot VLMs
arXiv 2607.17820 · 2026-07-23最优模式分选日冕仪:宜居世界天文台单模测量的极限
Optimal mode-sorting coronagraphy: limits of single-moded measurements for the Habitable Worlds Observatory
arXiv 2607.17637 · 2026-07-23