2026-09论文中文摘要
超越对称智能体:小语言模型中的认知多样性与多智能体辩论
Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models
arXiv 2609.35875 · 2026-09-30风险规避的在线POMDP规划:基于即时成本CVaR并具有性能保证
Risk-Averse Online POMDP Planning via CVaR of the Immediate Cost with Performance Guarantees
arXiv 2609.35874 · 2026-09-30更多程序还是更多轮次?在LLM工具框架中区分覆盖率与专门化
More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses
arXiv 2609.35873 · 2026-09-30SameFact:相同的安全事实在不同接口下导致不同响应
SameFact: The Same Safety Facts Lead to Different Responses Across Interfaces
arXiv 2609.35872 · 2026-09-30温度和外加电场对光电倍增管暗计数率影响的建模
Modeling the Effects of Temperature and Electric Field on the Dark Count Rate of Photomultiplier Tubes
arXiv 2609.35871 · 2026-09-30说会阻止,实际仍行动:为何LLM安全判断无法约束LLM智能体的行动
Says Block, Still Acts: Why LLM Safety Judgments Fail to Govern Action in LLM Agents
arXiv 2609.35870 · 2026-09-30Token 边界的代价:压缩证书与预测
The Price of Token Boundaries: Compression Certificates and Prediction
arXiv 2609.35869 · 2026-09-30人类可读文本对于大语言模型的有效微调是否必要?
Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?
arXiv 2609.35868 · 2026-09-30ReLOBGen:可回放的限价订单簿消息生成
ReLOBGen: Replayable Limit Order Book Message Generation
arXiv 2609.35867 · 2026-09-30弱偏好域上聚合的一个拓扑条件
A Topological Condition for Aggregation on Weak Preference Domains
arXiv 2609.35866 · 2026-09-30PACT:面向单标记类型化决策的成对锚定校准调优
PACT: Pairwise-Anchored Calibrated Tuning for Single-Token Typed Decisions
arXiv 2609.35865 · 2026-09-30解决财务数据缺失危机:用于SEC 10-K提取的生成式AI流水线
Resolving the Missing Financial Data Crisis: A Generative AI Pipeline for SEC 10-K Extraction
arXiv 2609.35864 · 2026-09-30超越判别:用于生物声学监测的校准地理先验融合
Beyond Discrimination: Calibrated Geoprior Fusion for Bioacoustic Monitoring
arXiv 2609.35863 · 2026-09-30$f(Q,B)$ 引力标量表示中的五维厚膜
Five-dimensional thick branes in the scalar representation of $f(Q,B)$ gravity
arXiv 2609.35862 · 2026-09-30面向与先进CMOS工艺三维集成的12英寸晶圆LGAD传感器技术开发
Development of LGAD Sensor Technology on 12 inch Wafers for 3D Integration with advanced CMOS processes
arXiv 2609.35861 · 2026-09-30可检测性差距:语言模型幻觉检测中的隐藏异质性
The Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models
arXiv 2609.35860 · 2026-09-30WLCG 迷你能力挑战:主机调优以改进 WAN 数据传输
WLCG Mini-Capability Challenge: Host Tuning to Im- prove WAN Data Transfers
arXiv 2609.35859 · 2026-09-30Cannon-Thurston映射、调和映射与纤维化双曲3-流形中的极小曲面
Cannon-Thurston Maps, Harmonic Maps, and Minimal Surfaces in Fibered Hyperbolic 3-Manifolds
arXiv 2609.35858 · 2026-09-30Hamel悖论之迷思
The Myth of Hamel's Paradox
arXiv 2609.35857 · 2026-09-30超越关键词:利用生成式大语言模型和标签聚合对新闻文章中的经济政策不确定性进行分类
Beyond Keywords: Leveraging Generative LLMs and Label Aggregation to Classify Economic Policy Uncertainty in News Articles
arXiv 2609.35856 · 2026-09-30Mara Chain:将失败重新构想为AI系统自动进化的垫脚石
Mara Chain: Rethinking Failure as a Stepping Stone for AI System Auto-Evolution
arXiv 2609.35855 · 2026-09-30立场声明:若无法强制可复现性,就让我们加强可验证性
Position: Let's Strengthen Verifiability If We Can't Enforce Reproducibility
arXiv 2609.35854 · 2026-09-30希尔伯特平面的一种斜边-角公理化
A Hypotenuse-Angle Axiomatization of the Hilbert Plane
arXiv 2609.35853 · 2026-09-30正实数上的次可加双射
Subadditive bijections on the positive reals
arXiv 2609.35851 · 2026-09-30超导量子比特作为粒子探测器的定量表征
Quantitative characterization of superconducting qubits as particle detectors
arXiv 2609.35850 · 2026-09-30SAR相位历史数据的学习压缩:关于GOTCHA的速率诚实可行性研究
Learned Compression of SAR Phase-History Data: A Rate-Honest Feasibility Study on GOTCHA
arXiv 2609.35848 · 2026-09-30Lean 中有限单群的一个构造性图册
A constructive ATLAS of finite simple groups in Lean
arXiv 2609.35847 · 2026-09-30振荡矩阵在矩阵与逆矩阵稀疏性下的指数轮廓
Exponent Profiles of Oscillatory Matrices under Matrix and Inverse Sparsity
arXiv 2609.35846 · 2026-09-30超球面语义轨迹分析:跨学术预印本、专利信号与算力扩展的技术扩散映射
Hyperspherical Semantic Trajectory Analysis: Mapping Technological Diffusion across Academic Preprints, Patent Signals, and Compute Scaling
arXiv 2609.35845 · 2026-09-30一种无分支的通用对算术重归一化方案及其性能评估
A Branch-Free General Renormalization Scheme for Pair Arithmetic and Its Performance Evaluation
arXiv 2609.35844 · 2026-09-30CMS开放数据可视化与FireworksWeb
CMS Open Data Visualization with FireworksWeb
arXiv 2609.35843 · 2026-09-30Crown 图使二部图的表示数最大化
Crown graphs maximise the representation number of bipartite graphs
arXiv 2609.35842 · 2026-09-30超越基于规则的变异测试:使用大型语言模型的测试感知变异体生成
Beyond Rule-Based Mutation Testing: Test-Aware Mutant Generation Using Large Language Models
arXiv 2609.35841 · 2026-09-30低秩近似的矩阵-向量复杂度
Matrix-Vector Complexity of Low-Rank Approximation
arXiv 2609.35840 · 2026-09-30从拍手声中估计房间冲激响应
Estimation of Room Impulse Responses from Handclaps
arXiv 2609.35839 · 2026-09-30记忆诱导的动态Ising-Kuramoto模型中的最优切换
Memory-induced optimal switching in a dynamic Ising-Kuramoto model
arXiv 2609.35838 · 2026-09-30“命运之门”——一个量子游戏概念
The "Gate of Fate" - a quantum gaming concept
arXiv 2609.35837 · 2026-09-30广义Tuza猜想的一个紧分数版本
A Tight Fractional Version of Generalized Tuza's Conjecture
arXiv 2609.35836 · 2026-09-30量子转译器的一种防泄漏、成本感知的回归测试方法
A Leakage-Safe, Cost-Aware Regression Testing Methodology for the Quantum Transpiler
arXiv 2609.35834 · 2026-09-30神经符号路由:在资源受限边缘设备上实现可靠推理
Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge Devices
arXiv 2609.35833 · 2026-09-30LLM何时应信任自己的修订?内在自我修正的风险感知研究
When Should LLMs Trust Their Own Revisions? A Risk-Aware Study of Intrinsic Self-Correction
arXiv 2609.35832 · 2026-09-30超越上下文窗口:一种基于自适应熵的混合检索与长上下文语言模型路由框架
Beyond the Context Window: An Adaptive Entropy-Based Routing Framework for Hybrid Retrieval and Long-Context Language Models
arXiv 2609.35831 · 2026-09-30翻译LHCb文档:初步经验与维护保障
Translating LHCb's Documentation: First experiences and ensuring maintenance
arXiv 2609.35830 · 2026-09-30MPD实验在NICA能区双电子测量中的性能研究
Performance of the MPD experiment in dielectron measurements at NICA
arXiv 2609.35829 · 2026-09-30时间投票中的完全正当代表:高效计算与验证
Full Justified Representation in Temporal Voting: Efficient Computation and Verification
arXiv 2609.35826 · 2026-09-30用于精密实验的可扩展堆叠电极印刷电路板射频四极杆(PCB-RFQ)
A Scalable Stacked-Electrode Printed Circuit Board Radio-Frequency Quadrupole (PCB-RFQ) for Precision Experiments
arXiv 2609.35825 · 2026-09-30可靠但对设计敏感:LLM标注中的工具不确定性
Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation
arXiv 2609.35824 · 2026-09-30CoVLM-Bench:面向协同驾驶问答与规划的真实世界基准
CoVLM-Bench: A Real-World Benchmark for Cooperative Driving Question Answering and Planning
arXiv 2609.35823 · 2026-09-30语言模型中奉承性同意的机制追踪
Tracing mechanisms of sycophantic agreement in language models
arXiv 2609.35822 · 2026-09-30$τ$-Multilingual:跨语言语音代理基准测试
$τ$-Multilingual: Benchmarking Voice Agents Across Languages
arXiv 2609.35820 · 2026-09-30