📄 arXiv 论文速递

📅 2026-09-06cs.AI + cs.LG 最新提交 | DeepSeek 点评:这篇为什么重要

💡 针对自然语言规则难以实现、远程大模型调用成本高的问题,提出将文本规范训练为本地神经函数,兼顾描述灵活性与推理效率,对实际应用具有重要意义。

Many recurring text functions are easy to describe but difficult to implement with rules, while calling a large remote model for every input introduces repeated cost, latency, and

💡 揭示黑盒大模型作为评测工具在共享端点上的测量不稳定性,直接挑战当前依赖LLM评判的基准与训练流程,对AI评估可靠性提出关键警示。

Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the sa

💡 针对进化式提示优化器提示膨胀、长度剧增的问题,提出结构化诊断、多样化与稳定化策略,显著提升优化效率与提示质量,推动自动提示工程发展。

Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, producing prompts up to 3$\times$ longer yet no more accurate. We tr

💡 质疑思维链推理痕迹的可解释性,通过对比人类判断与真实重要性,揭示表面可读性与深层可解释性的鸿沟,对AI安全与可解释性研究具有重要价值。

Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges

Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We in

We introduce the first Probably Approximately Correct (PAC) learning framework for general-sum concurrent stochastic games (CSGs) with transition uncertainty, while addressing the

Weakly-Supervised Dense Video Captioning aims to localize and describe multiple events in untrimmed videos given only an ordered set of event-level captions per video. Recent work

Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally

💡 为因果概率解释提供计算可行框架,解决结果归因与责任分配的核心难题,兼具哲学严谨性与工程实用性,对可解释AI与决策分析有广泛影响。

Explaining why a specific outcome occurred, and which inputs deserve the blame or credit, is central to philosophical, scientific, and policy analysis. Existing tools split into tw

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leavi

Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate, coordinate, and build on each other's work. Yet this shared infrastructure can als

As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges. Traditional pipelining tech