公开论文雷达

公开 arXiv 研究简报 · 2026-09-29T00:59:21.676926+00:00

先看结论和关键数字,再决定要不要读原文。候选只在首次出现时展示;旧候选若后来通过深读门槛,仍会进入重点。

先读SWE-Proof:测试全绿不等于修对

八张卡里只有SWE-Proof、VSpector、Agentic-IC3、RankGround给了基线对照和数字。先读SWE-Proof,它在500个真实issue上量出约25%测试通过的补丁仍有反例;接着读实验设置那篇,明白同一基准的修复率为什么不能直接横比。其余几张偏框架和草稿,当思路看,不要当结论引。

推荐阅读顺序

  1. 2609.21190:证据最硬,直接改你的合并门禁:25%测试通过补丁有反例,85.0%→95.1%有对照。
  2. 2609.17993:读完上一篇你会想比数字,这篇先告诉你同基准的修复率其实做的是不同任务。
  3. 2609.23517:把规范当判据落到实处:148个真实违规、42个新缺陷,模糊测试24小时一个没找到。
  4. 2609.27162:看模型只提案、求解器只校验的分工怎么写;14个基准解出10个,含4个基线全败。
  5. 2609.27105:同一分工的另一种形态,判定权交给Lean内核;13个域12个通过,样本小,看模式即可。
  6. 2609.28575:提醒你上线门禁要同时卡误报:flat-RAG误报16%-43%,部署系统召回只有42%。
  7. 2609.18690:和前面几篇无关,单看一个省调用的工程结论:1次调用比次优多裁剪高5.5%、快1.4倍。
  8. 2609.28896:放最后。方向和前面几篇一致,但没有实验数字,只能借它的流程顺序。
共性方法
这批卡片使力的地方是同一处:把“跑通了就算对”换成一道独立的机器校验。SWE-Proof用形式规约查补丁,Agentic-IC3和广义计划把最终裁决交给求解器和Lean内核,VSpector拿官方规范文本当判据,TWIST给每个检测配表面相似的安全负例。共同点是判定权不交给生成方自己。
关键分歧
分歧在证据够不够硬。SWE-Proof、VSpector、Agentic-IC3、RankGround都有基线对照和具体数字;广义计划只有单模型13个域;TWIST四条轨道只验证了一条,标的还是v0.3草稿;规范驱动APR基准那篇没有实验,只有概念框架和一个没给细节的示例。
选择准则
先看卡片有没有基线对照和数字:有,按证据用;没有,只取流程结构,别当结论引用。真要加校验关卡,优先挑判定权在求解器或内核的那种,模型自写规约先别信,那篇里只有56%通过审计。

重点深读(8 / 8 篇)

形式化与程序验证(3 篇)

形式化与程序验证 10/30

SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

测试通过不等于补丁正确:四分之一藏反例,形式规约把解决率从85.0%推到95.1%:如果你用测试通过来给大模型补丁放行,这个缺口会直接进生产:SWE-Proof在500个真实issue上发现,约25%测试通过补丁仍有反例。它的做法是给补丁补机器可检验的形式规约与证明;对Claude Opus 4.8,给定正确形式规约后解决率从85.0%升到95.1%,而模型自写规约相对无规约基线没有提升。

两句看懂

团队为SWE-bench Verified的500个真实issue补上机器可检验的形式规约与证明,用来检查测试通过是否足够。结果很明确:四分之一测试通过补丁仍有反例,正确形式规约把Claude Opus 4.8解决率从85.0%提升到95.1%,但模型自写规约完全无效且仅56%通过审计。

核心判断

核心判断是:held-out测试不足以证明真实issue补丁正确;SWE-Proof用500个真实issue给出反例率约25%,并显示正确形式规约可把Claude Opus 4.8解决率从85.0%提高到95.1%,但模型自写规约只有56%达标且对解决率无助益。

关键要点

1) 旧假设失败:held-out测试被当作补丁正确性判据,但约25%测试通过补丁仍被形式规约找出反例。 2) 方法与受控检查:Benchproofer写规约、把既有函数公理化、再过机械验证与对抗式LLM审计,两关都过才收录。 3) 决定性结果与动作:正确形式规约使解决率85.0%→95.1%;自写规约无提升,所以不要省掉规约审计。

证据与结果

评测覆盖SWE-bench Verified全部500个真实issue并构成SWE-Proof,同一流水线扩展到SWE-bench Pro;后端为NAGINI、VELVET(Lean DSL)、纯Lean。模型只看Claude Opus 4.8,设置对比无规约基线、模型自写规约、给定正确形式规约;指标是issue解决率与规约审计通过率。结果:测试通过补丁约25%有反例;正确形式规约85.0%→95.1%;自写规约相对基线无提升且56%通过审计;未解决实例92%规约审计不通过,已解决实例51%。

打开论文原文
它要解决什么
工程上真正的问题是:held-out测试能不能判定真实repo级issue补丁已经修对?如果不能,给补丁配形式规约是否能堵住测试盲区,还是只是把成本搬到规约写作上?
研究路径
机制分五步:1) 对每个issue的新增/改动代码写形式规约;2) 把补丁调用的既有函数写成公理化行为假设,不展开验证以控制规模;3) 机械验证器检查实现是否满足规约;4) 对抗式LLM审计检查规约是否忠实覆盖完整意图;5) 机械验证和审计都通过,样本才进入SWE-Proof。
这对工程意味着什么
第一步:给高风险补丁追加形式规约与机器验证,不因测试全绿就合并。要避开的捷径:不要直接采用模型自写或结构化自然语言规约,它们会漏掉约25%错误补丁,自写规约也只有56%通过审计。
证据定位
对Claude Opus 4.8,测试通过补丁里约25%在形式规约下暴露反例;结构化自然语言规约补不上这个缺口。给定正确形式规约时,解决率从85.0%升至95.1%;让模型自写规约相对无规约基线零提升,且只有56%自写规约通过审计。规约质量与结果强相关:未解决实例中92%规约审计不通过,已解决实例中该比例为51%;主要失败模式是不忠实,即只约束部分要求行为。(筛选维度:形式化验证、可复核评测、软件工程方法)
适用边界
主要证据来自SWE-bench Verified的500个issue,评测模型只有Claude Opus 4.8;规约是否忠实依赖对抗式LLM审计,而审计器本身可靠性未单独验证;扩展到SWE-bench Pro没有给出具体数字。
方法与英文摘要

作者取SWE-bench Verified全部500个真实issue,跑Benchproofer流水线:先给新增/改动代码写形式规约;再把补丁调用到的既有函数概括成公理化假设,避免展开验证整个仓库;随后用机械验证器查规约与实现一致,再用对抗式LLM审计查规约是否忠实覆盖意图。两类关卡都通过才收入SWE-Proof;同一流程扩展到SWE-bench Pro,验证后端含NAGINI、VELVET(Lean DSL)与纯Lean。

Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check correctness with held-out test suites, which are inherently incomplete and increasingly susceptible to memorization. Formal verification avoids both problems, but existing work covers only standalone tasks whose specifications are given as input, not real issues, which touch large repositories and state intent in vague natural language. We present Benchproofer, a pipeline that turns a coding task with a known correct patch into a formally verified one: it writes a specification for the new code, summarizes the existing functions that code calls with axioms, and admits an instance only after mechanical and adversarial gates agree. Applying it to SWE-bench Verified yields SWE-Proof, 500 real issues whose correctness is formally verified rather than tested, and it extends to SWE-bench Pro. Evaluating Claude Opus 4.8, we find that verification catches what tests miss: a quarter of test-passing patches admit counterexamples, which a structured natural-language specification does not fix, while a correct formal one lifts resolution from 85% to 95%. Writing that specification is the hard part: an agent that must write its own gains nothing over an unaided baseline, and only 56% of those specifications pass our audit. The usual failure is faithfulness, a specification that constrains part of the required behavior and leaves the rest free. Specification quality still tracks the outcome, failing on 92% of unresolved instances against 51% of resolved ones, making faithful specification synthesis a concrete open problem.

形式化与程序验证 7/30

Agentic-IC3: Enabling Semantic Proof Search in IC3 Model Checking

语言模型代理读RTL引导IC3证明搜索,14个基准解出10个,含4个三条基线都解不出…:你在做IC3形式化验证时,证明常常卡在泛化启发式上:位级实现用不上RTL高层语义,高层实现又挖不透设计结构。Agentic-IC3的做法是把语言模型代理接到Pono字级IC3后端前面,让代理读RTL和实时证明状态来提引理,求解器只负责校验。结果:14个基准、1小时超时内解出10个,其中4个是rIC3、Pono-IC3Bits、A-IC3三条基线全部未解出的用例。

两句看懂

IC3证明搜索依赖泛化启发式,位级实现难以利用RTL高层语义,高层实现也难以充分挖掘设计结构;Agentic-IC3让语言模型代理接入Pono字级IC3后端,依据RTL与证明状态提出引理,并可回溯修正无效分支。在14个信息流与协议/处理器/功能单元基准上,1小时超时内解出10个,其中4个是rIC3、Pono-IC3Bits、A-IC3三条基线全部未解出的用例。

核心判断

语言模型代理可以在RTL语义层面有效引导IC3证明搜索。证据:14个基准中解出10个,其中4个是rIC3、Pono-IC3Bits、A-IC3三条基线全部未解出的用例。

关键要点

1. 旧假设的局限:位级IC3把RTL降到位级逻辑,难以表达高层语义关系;字级/高层IC3保留多位操作,但通用化启发式仍难充分利用设计结构、难跨设计迁移,依赖专家经验。 2. 方法与受控校验:代理接入Pono字级IC3后端,读RTL、性质、证明状态与求解器反馈,用SAT/UNSAT泛化提引理、可加衍生观测信号、可回溯修正;求解器校验通过后才更新证明状态。 3. 决定性结果与行动:14个基准、1小时超时内解出10个,含4个三条基线(rIC3、Pono-IC3Bits、A-IC3)全部未解出的用例;给IC3类工具接语义代理时,务必保留求解器校验关卡。

证据与结果

评测集为14个基准,覆盖安全信息流验证和通信协议、处理器、功能单元的功能验证。对比基线为rIC3、Pono-IC3Bits、A-IC3。单基准超时1小时,指标为解出数量。结果:Agentic-IC3解出10/14,其中4个是三条基线全部未解出的用例,体现了基线在语义层面的求解难点。

打开论文原文
它要解决什么
语言模型代理能否读取RTL寄存器级语义,替代位级或受限高层启发式,来指导IC3归纳证明搜索?
研究路径
代理依次读取RTL、性质、当前证明状态(已接受引理与证明义务)和求解器反馈,定位使义务不可达的信号关系;然后通过SAT或UNSAT泛化提出候选引理,或引入衍生观测信号来简化关系表达;后端用SAT/SMT求解器校验每个提案;若提案导致无效分支,代理回溯到更早的证明义务去修正;只有校验通过后才更新引理集合与义务队列。
这对工程意味着什么
第一步行动:给你的IC3类归纳证明工具接一个能读取RTL和实时证明状态的代理去提引理,而不是只依赖位级泛化规则。要避免的捷径:不要让LLM直接产出不变式、跳过求解器校验就采信——该校验关卡是可靠性的全部来源。
证据定位
决定性对比:在14个基准、1小时超时内,Agentic-IC3解出10个;其中4个是rIC3、Pono-IC3Bits、A-IC3三条基线全部未解出的用例。这说明代理的语义引导带来了位级和受限高层方法给不出的增量求解能力。(筛选维度:形式化验证、可复核评测)
适用边界
评测仅14个基准,集中在安全信息流与通信协议、处理器、功能单元几类设计;超时统一设为1小时;对比基线仅rIC3、Pono-IC3Bits、A-IC3三种,未见更大规模工业基准的结果。
方法与英文摘要

基于Pono字级IC3后端,给语言模型代理开一个接口,让它读取RTL、性质、当前证明状态和求解器反馈。代理通过SAT/UNSAT泛化提出候选引理,也可以插入衍生观测信号来表达信号关系;如果提案导致无效分支,代理回溯到更早的证明义务去修正。任何提案都必须经SAT/SMT求解器校验通过后,才更新证明状态。评测用14个基准,覆盖安全信息流验证和通信协议、处理器、功能单元的功能验证,单基准超时1小时。

IC3 is a state-of-the-art algorithm for hardware model checking that proves safety properties by incrementally constructing an inductive invariant consisting of a set of lemmas. Its effectiveness depends on generalization heuristics that identify useful lemmas and guide proof search. However, many leading IC3 hardware model checkers operate on lowered, bit-level representations, where high-level design relationships are difficult to exploit for generalization. Those operating at a higher level remain limited in exploiting high-level design structure and semantics. We present Agentic-IC3, built on Pono's word-level model-checking infrastructure, which integrates a language-model agent into IC3 to guide semantic proof search using register-transfer-level (RTL) design information. The framework exposes an agent-oriented interface to a persistent IC3 backend, allowing the agent to interact with an explicit, evolving proof state throughout verification. Across successive proof obligations, the agent relates intermediate proof states and solver feedback to the RTL and proposes high-level lemmas through both SAT and UNSAT generalization. Beyond generalization, the agent can introduce derived observation signals to express design relationships succinctly and obtain more informative feedback, and backtrack to revise proposals that lead to unproductive proof branches. The backend checks proposals before updating the proof state, preserving soundness and providing feedback for further reasoning. On a suite of 14 benchmarks spanning security information-flow verification and functional verification of communication protocols, processors, and functional units, Agentic-IC3 solves 10 cases within a one-hour timeout, including 4 unsolved by all three evaluated baselines: rIC3, Pono-IC3Bits, and A-IC3.

形式化与程序验证 7/30

Provably Complete Generalized Planning with LLMs

LLM生成的广义计划,现在能用机器自动证明对全部实例成立:你在意这件事,是因为LLM生成的规划代码即使在测试集上100%通过,也无法说明它对域内所有实例都有效。这项工作给出了一条出路:让LLM同时生成计划和Lean完备性证明,由Lean内核机器检查,不再靠人工抽查。13个基准域中12个通过验证,1个失败。

两句看懂

以往LLM生成的广义计划只能靠人工检查是否对域内所有实例都成立,这项研究让LLM同时生成计划和形式化完备性证明。在13个基准域上用GPT-5.6-Sol测试,12个域的证明通过了Lean内核检查,1个域未获得有效证明。

核心判断

可以自动证明LLM生成的广义计划完备:在13个基准域中,12个域上LLM生成的Lean完备性证明通过了Lean内核的机器检查,1个域未通过。

关键要点

1. 旧缺口:此前工作(如Stein等2026)用LLM生成Python广义计划,测试数据上100%覆盖,但域内是否所有实例都被解决,只能靠人工逐案例分析,没有自动化完备性证据。 2. 方法与检查:语义保持的PDDL-to-Lean转换编码域约束和状态不变式,GPT-5.6-Sol同时生成广义计划(Lean程序solve)和完备性定理证明,是否成立由Lean内核类型检查判定,不是人工也不是LLM自评。 3. 结果与行动:13个基准域中12个域的证明通过Lean内核验证,1个失败;对可形式化正确性条件的生成代码,应把机器可检查的证明作为验收依据。

证据与结果

评测覆盖13个常用规划基准域,是Stein等2026年所用14个域的子集。每个域的输入是人工给定的PDDL域约束和状态不变式,用同一个模型GPT-5.6-Sol对每个域生成广义计划和完备性证明,再由Lean内核统一做类型检查。结果为12/13个域证明通过,1个域未通过。摘录未给出失败域的名称和具体诊断信息。

打开论文原文
它要解决什么
LLM生成的广义计划在测试集上做到100%覆盖,能否自动证明它对规划域中所有实例都成立,而不是人工逐案例抽查?
研究路径
流程是一条链:人工给定PDDL域约束和状态不变式→语义保持的PDDL-to-Lean转换生成State/Action等Lean结构→GPT-5.6-Sol生成广义计划solve(Lean程序)→同一模型生成完备性定理solveComplete的证明(对任意满足约束的实例,solve的计划有效且达成目标)→Lean内核对证明做类型检查(#check)→通过即该域完备性得证,不通过记为失败。判断权在Lean内核,不在人,也不在LLM自己。
这对工程意味着什么
第一步行动:对LLM生成的算法性代码,先检查能否形式化其正确性条件;若可以,让LLM连带生成机器可检查的证明(如Lean)并据此验收。要避免的捷径:不要把测试集100%覆盖当成全量正确性的证据,那只是抽样通过,不是完备性证明。
证据定位
在13个基准域上,12个域同时拿到了广义计划和通过Lean内核机器检查的完备性证明,1个域未能得到有效证明。对照基线是:此前同类工作(如Stein等2026年用LLM生成Python广义计划,测试数据上100%覆盖)只能靠人工逐案例分析确认完备性,没有自动化的完备性证据。(筛选维度:形式化验证、可复核评测)
适用边界
评测只覆盖13个常用基准规划域,且1个域未获得有效完备性证明,失败原因摘录未说明。只用了单一模型GPT-5.6-Sol。域约束和状态不变式需要人工预先给定,之后才是自动化步骤。
方法与英文摘要

输入是PDDL描述的规划域约束和状态不变式。第一步做语义保持的PDDL-to-Lean转换,把域约束和状态不变式编码为Lean的State/Action结构。第二步用GPT-5.6-Sol同时生成广义计划(Lean程序solve)和一条完备性定理证明,定理陈述为:对任意满足约束的实例,solve给出的计划有效且达成目标。第三步由Lean内核对证明做类型检查(#check),通过即记为该域完备性得证。评测数据是13个常用规划基准域,取自Stein等2026年使用的14个域的子集。

Generalized planning aims to compute a plan that solves all instances of a planning domain. Recent work has used LLMs to automatically generate and debug such generalized plans in the form of Python programs and achieved perfect test data coverage for several domains. However, whether these generalized plans are actually complete, i.e. solve all instances of the domain, could only be determined by manual evaluation. Here, we present an approach for automatically generating generalized plans in Lean together with proofs of their completeness relative to a specification of the domain constraints provided as input. We introduce a semantic-preserving PDDL-to-Lean conversion, and use an LLM to generate both the generalized plan and the formal proof that it solves every instance satisfying the domain constraints. The correctness of the completeness proof is determined by Lean's kernel. We evaluate our approach on 13 commonly used benchmark domains, using GPT-5.6-Sol as the LLM. For 12 of the domains we obtain generalized plans together with valid completeness proofs. This is a major advancement of the state of the art in automatic generalized-plan completeness proofs.

软件工程与仓库智能(2 篇)

软件工程与仓库智能 11/30

Specification-Driven Benchmarking for Automated Program Repair From Static Corpora to Executable Specifications

静态程序修复基准撑不住可信评测,应改成按规范生成:评自动程序修复时,最大的风险是基准本身不可信:Defects4J、BugsInPy、SWE-bench覆盖有限,随复用增加更易被大模型训练污染,难度和故障类型也继承自历史缺陷库。论文给出的方向是先用可执行规范声明要什么基准,再由流水线生成实例。

两句看懂

Defects4J、BugsInPy、SWE-bench这类静态APR基准覆盖有限,且越复用越可能被大模型训练数据污染,特性也无法独立控制。论文提出用可执行规范驱动生成流水线替代固定语料,但摘录只支持到概念框架和示例说明,没有量化结果。

核心判断

静态APR基准的结构性问题是覆盖有限、易污染、特性不可控;论文主张用可执行规范生成基准,但支撑证据只到概念论证和一个端到端示例,未见量化实验。

关键要点

1. 旧假设失效:Defects4J、BugsInPy、SWE-bench把基准特性继承自历史缺陷语料,覆盖有限且更易被训练污染。 2. 方法与受控点:用五维规范声明需求,生成、验证、语料管理三模块分离,验证模块独立于生成模块。 3. 结果与行动:摘录无量化结果;可采信的是结构原则,落地时先写规范并独立验证。

证据与结果

摘录没有数据集规模、划分、指标或成功率。它只把Defects4J、BugsInPy、SWE-bench列为静态基准局限的例子,并提到完整论文含一个端到端示例;示例细节不在摘录中,因此不能把该框架当作已验证有效。

打开论文原文
它要解决什么
静态APR基准是否已难以支撑可信评测?能否不写死数据集,改用可执行规范按需生成基准?
研究路径
机制分五步:研究者写规范,声明程序上下文、故障分类、难度、验证策略、语料约束;生成模块按规范产出候选修复任务;独立验证模块检查每个实例是否满足规范;语料管理模块汇总和管理实例;同一规范换随机种子重跑,得到分布等价但内容不同的新语料,而不是反复使用固定数据集。
这对工程意味着什么
第一步先写机器可读规范,再分别实现生成器和独立验证器,并用换种子重生成语料。要避免的捷径是只扩大静态数据集规模;论文指出这不能消除污染,也不能让难度和故障类型变得可控。
证据定位
证据边界很窄:原文摘录没有实证评测、数据集规模、修复系统性能数字,也没有污染检测结果。只提到Defects4J、BugsInPy、SWE-bench作为静态范式局限的例子,并称完整论文有一个端到端示例,说明规范可传导为可独立验证的基准实例;该示例的语言、任务数、验证结果未在摘录中给出。(筛选维度:形式化验证、可复核评测、软件工程方法)
适用边界
这是单一作者的概念性框架论文。原文摘录未给出实证数据集规模、生成成功率、污染检测结果或修复系统实测;唯一提到的端到端示例细节未包含在摘录中,可行性仍待实证验证。
方法与英文摘要

论文不跑实证实验,而是搭概念框架。它把基准规范拆成五个维度:程序上下文、故障分类、难度、验证策略、语料约束。生成流水线分成三个独立模块:生成、验证、语料管理。规范先声明目标,生成模块产出候选修复任务,独立验证模块检查是否符合规范,语料管理模块汇总实例。

Automated Program Repair (APR) benchmarks have traditionally been constructed as static datasets whose characteristics are inherited from the defects they contain. While this paradigm has enabled decades of progress, finite corpora provide limited experimental control, become increasingly susceptible to contamination as they are reused, and cannot be systematically regenerated or adapted as evaluation requirements evolve. We propose specification-driven benchmarking, a paradigm in which benchmarks are defined by executable specifications and realized through benchmark generation. The specification explicitly declares the intended properties of the benchmark (including program context, fault taxonomy, difficulty, validation strategy, and corpus constraints) while a generation pipeline realizes those requirements through independent generation, validation, and corpus management components. We develop the conceptual foundations of this approach by introducing a taxonomy of benchmark specification dimensions, establishing how each specification dimension maps to deterministic architectural responsibilities, and arguing that independent validation is a structural requirement for trustworthy benchmark generation. An end-to-end example illustrates how specification choices propagate through the pipeline to produce benchmark instances whose properties are independently verifiable. By treating the benchmark as an executable specification rather than a static dataset, the proposed paradigm shifts benchmark construction from artifact curation to declarative experimental design.

软件工程与仓库智能 9/30

Experimental Settings in LLM-Based Program Repair: A Study of Inputs, Tool Access, Feedback, and Validation

同一基准的APR修复率不能直接横向比较:如果你要引用或复现LLM程序修复系统的评测结果,这份研究先给你提个醒:只看基准名和修复数量会得出错误结论。作者用一个七维度框架拆解了10个已发表实验设置,发现同一基准下的修复任务本身可能完全不同。

两句看懂

APR评测通常只报告基准、修复数量和最终验证测试,但同一基准下系统拿到的初始输入、工具权限和失败反馈可能完全不同,等于做的是不同的修复任务。作者对Defects4J和SWE-bench上7项研究的10个实验设置逐一编码,确认基准加分数不足以还原任务范围,部分细节甚至无法从论文中恢复。

核心判断

同一基准下的APR修复率不能直接比较:对Defects4J、SWE-bench上7项研究10个实验设置的编码显示,基准相同但输入、工具权限、反馈与验证方式不同,任务本身就不同。

关键要点

1. 旧假设失效:基准名+修复数量+最终验证测试曾被视为足以定义修复任务,但同一基准可实例化出从局部patch生成到仓库级迭代修复的完全不同任务。 2. 方法与受控检查:作者提出七维度编码框架(任务单元/故障定位假设/初始输入/工具访问/修复时反馈/最终验证/资源预算)及机器可读schema,对Defects4J与SWE-bench上7项研究的10个实验设置逐条编码比较。 3. 决定性结果与行动:同为Defects4J,既有oracle给出故障方法的设置,也有要求系统自行定位故障的设置,部分细节无法从论文还原——报告和引用APR结果时必须附带实验设置记录。

证据与结果

数据来自Defects4J与SWE-bench两个基准,语料为7项研究共10个实验设置,每项研究聚焦一个焦点APR系统。编码维度覆盖任务粒度(方法/缺陷/仓库级)、故障定位设置、初始输入、修复时工具访问、修复时反馈、最终验证方式和资源预算。核对结果显示,部分设置的关键细节(如反馈内容、工具权限范围)无法从论文中还原。

打开论文原文
它要解决什么
在Defects4J、SWE-bench这类同一基准上,不同LLM修复系统报告的修复率能不能直接拿来比较?
研究路径
具体做法是:先收集Defects4J与SWE-bench上7项研究、每项聚焦一个焦点APR系统的10个实验设置;再按七个维度——任务单元、故障定位假设、初始输入、修复时工具访问、修复时反馈、最终验证、资源预算——逐条编码;最后比较同一基准下不同设置的差异,并标记那些无法从已发表报告中还原的细节。
这对工程意味着什么
第一步行动:引用或复现任何APR修复率之前,先核对论文是否写清故障定位假设、工具访问权限和反馈机制。要避免的捷径:仅凭'同基准+同修复率'就横向比较不同系统,这很容易得出错误结论。
证据定位
同样是Defects4J:有的设置由oracle直接给出故障方法,系统只需生成patch;有的要求系统从项目检出和失败信息出发自行定位故障。测试的使用也不同:有时只用于最终验证,有时每次失败后把反馈送回系统指导下一轮修复。此外,部分实验的关键细节无法从论文文本中还原。(筛选维度:可复核评测、软件工程方法)
适用边界
分析只覆盖7项研究、10个实验设置,且限于Defects4J与SWE-bench两个基准;结论仅适用于已发表且能获取实验细节的APR系统,未涵盖全部LLM修复方法。
方法与英文摘要

作者收集了Defects4J与SWE-bench上7项已发表研究、共10个实验设置(每项研究聚焦一个焦点修复系统),按任务单元、故障定位假设、初始输入、修复时工具访问、修复时反馈、最终验证、资源预算七个维度逐条编码,并给出一个机器可读的记录schema。

Evaluations of automated program repair (APR) systems commonly report the benchmark, the number of repaired defects, and the tests used for final patch validation, but these items no longer fully specify the repair task presented to a system. Recent LLM-based systems differ in the information supplied before repair, the repository and testing operations permitted during repair, and the feedback returned after unsuccessful attempts, allowing the same benchmark to instantiate substantially different repair tasks ranging from localized patch generation to repository-level diagnosis and iterative repair. We present a framework for explicitly specifying the experimental settings associated with reported APR results. We analyze reported experimental settings from systems evaluated on Defects4J and SWE-bench and characterize each result by its task unit, fault-localization assumptions, initial input, tool access, repair-time feedback, final validation, and resource budget. Our analysis shows that benchmark identity alone is insufficient to reconstruct the evaluated task or determine the appropriate scope of comparison across reported repair rates. We therefore introduce a machine-readable schema for specifying each experimental setting to improve reproducibility and make the scope of cross-system comparisons explicit.

代码质量与优化(1 篇)

代码质量与优化 6/30

VSpector: Specification-Driven Bug Detection for RISC-V CPUs

用官方规范文本揪出RISC-V芯片漏洞:VSpector把RISC-V官方规范转成规则,让LLM直接比对CVA6、XiangShan的RTL代码,发现73个缺陷中42个此前未知,精度68.2%,而SOTA模糊测试工具24小时未发现这些新缺陷。

两句看懂

传统RISC-V缺陷检测依赖参考模型、形式化属性或人工缺陷模式,这些人工产物本身可能遗漏缺陷,VSpector改用LLM直接比对官方规范文本与RTL代码。在CVA6和XiangShan上,217个候选中148个被人工确认为真实违规(精度68.2%),对应42个未知缺陷,而SOTA模糊测试工具DiveFuzz跑24小时一个都没找到。

核心判断

可以:官方RISC-V规范文本可直接作为LLM检测RTL缺陷的信息源,无需参考模型或人工缺陷模式;证据是148个真实违规(68.2%精度)、42个未知缺陷,而SOTA模糊测试24小时未发现。

关键要点

1. 传统CPU缺陷检测(模糊测试、形式化验证、静态分析)依赖参考模型、形式化属性或人工总结的缺陷模式,这些人工产物本身可能不可靠——例如常用ISA模拟器NEMU和Spike因实现简化而漏检某些缺陷,导致依赖它们的检测方法同样漏检。 2. VSpector面向CVA6和XiangShan两个工业级RISC-V CPU,分四阶段处理:规则抽取(从官方规范文档提取自然语言规则)、实现定位(找到对应RTL代码)、候选识别(必要时用ReAct流程按需检索补充代码上下文)、违规审计(逐条先查规范侧再查代码侧,降低误报),用于平衡上下文覆盖广度与LLM推理准确性之间的矛盾。 3. 217个候选违规经人工核查确认148个为真(精度68.2%),对应73个缺陷、42个此前未知;SOTA模糊测试工具DiveFuzz对相同代码跑24小时未发现这42个缺陷中任何一个;42个新缺陷已上报上游,19个已修复、11个已确认,共30个得到官方响应。

证据与结果

评测对象:CVA6、XiangShan两个工业级RISC-V CPU的RTL实现。VSpector共报告217个候选违规,人工逐一核查确认148个为真实违规,精度68.2%,对应73个不同缺陷,其中42个此前未知。对照:SOTA模糊测试工具DiveFuzz对相同代码提交各跑24小时,未发现这42个新缺陷中任何一个。跟进:42个新缺陷均已上报上游仓库,开发者已修复19个、确认11个,共30个获官方响应。

打开论文原文
它要解决什么
能否绕开参考模型、形式化属性和缺陷模式,直接用RISC-V官方规范文本检测RTL缺陷?
研究路径
四阶段流程:①规则抽取——从官方RISC-V规范文档解析出自然语言规则;②实现定位——将规则映射到对应RTL代码位置;③候选识别——检查定位代码是否违反规则,上下文不足时触发ReAct按需检索补充RTL代码;④违规审计——对每个候选依次检索补充规范条款与代码上下文,先做规范侧复核再做代码侧复核,降低误报。
这对工程意味着什么
可将规范驱动审计作为模糊测试/形式化验证之外的补充检测层,因为它能发现参考模型本身缺陷导致的漏检;但不能仅看68.2%精度就跳过人工复核,32%的候选仍是误报。
证据定位
217个候选中人工确认148个真实违规,精度68.2%,对应73个缺陷,其中42个此前未知;DiveFuzz对相同代码24小时模糊测试未发现这42个缺陷中任何一个;42个新缺陷已上报上游,19个已修复、11个已确认。(筛选维度:形式化验证、可复核评测)
适用边界
仅在CVA6、XiangShan两个RTL实现上验证,未覆盖更多CPU设计;217个候选中仍有约32%为误报,依赖人工复核确认;方法效果受官方规范文档完整性和LLM规则抽取准确性影响。
方法与英文摘要

数据:CVA6、XiangShan两个工业级RISC-V CPU的RTL实现与官方RISC-V规范文档。流程分四阶段:规则抽取→实现定位→候选识别(必要时ReAct检索补充RTL代码)→逐条违规审计(先查规范侧再查代码侧,降低误报)。对照实验:SOTA模糊测试工具DiveFuzz对同一代码提交跑24小时。

Detecting RTL design bugs in open-source RISC-V CPU implementations is critical for ensuring system reliability. Traditional detection approaches inherently rely on predefined artifacts. In this paper, we leverage the official,natural-language RISC-V specifications as an effective information source for bug detection. We present VSpector, a specification-driven bug detection pipeline that directly checks whether CPU register-transfer level (RTL) implementations adhere to official specification rules, without requiring specialized construction of reference models, formal properties, or custom bug patterns. To resolve the key technical trade-off between broad context scope and model reasoning accuracy when using Large Language Models (LLMs), VSpector employs a stepwise context refinement scheme across a four-stage pipeline: rule extraction, implementation localization, candidate identification, and sequential violation auditing. We evaluate VSpector on two industrial-strength RISC-V CPUs, CVA6 and XiangShan. Out of 217 reported candidates, manual inspection confirmed 148 true violations, representing a 68.2% precision. These violations correspond to 73 distinct bugs, including 42 previously unknown bugs. In our comparative experiments, DiveFuzz, a state-of-the-art CPU fuzzer, detected none of these new bugs during 24-hour runs per CPU. All 42 new bugs have been reported upstream, with developers already fixing 19 and confirming an additional 11 (30 in total), demonstrating that specification-driven auditing is a practical and complementary strategy for CPU bug detection.

UI 与 GUI Agent(1 篇)

UI 与 GUI Agent 6/30

RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection

轻量重排序器选一张裁剪图,一次VLM调用就超过多裁剪方法:如果你在做高分辨率GUI定位,整图推理会漏小控件,多次裁剪又慢。RankGround用一个轻量重排序器GroundRanker先选出最优裁剪图,只做一次VLM调用,平均精度反而比次优多裁剪方法高5.5%,推理快1.4倍(ScreenSpot-Pro)。

两句看懂

整图推理常漏检小控件,多裁剪多次VLM调用虽准但慢;RankGround用轻量重排序器GroundRanker选优裁剪图,仅一次VLM调用完成定位。ScreenSpot-Pro多骨干多尺度测试显示,RankGround平均精度超次优方法5.5%,推理快1.4倍。

核心判断

单次VLM调用可以达到甚至超过多裁剪方法的精度:GroundRanker选出的裁剪图让RankGround以1次调用超过次优多裁剪方法5.5%精度,且推理快1.4倍(ScreenSpot-Pro)。

关键要点

1. 旧假设失效:直接整图推理在高分辨率密集布局下漏检小/相似控件,ScreenSpot-Pro上Direct与Oracle上限差距超18%。2. 方法与对照:无现成排序数据,作者用严格包含准则加边界感知正样本增强构造监督,GroundRanker先pointwise学包含、再listwise区分相似裁剪图;对照ZoomIn(2次调用)和MVP(3次以上)。3. 结果与行动:RankGround仅1次调用,平均精度超次优方法5.5%、快1.4倍;高分辨率GUI定位应改为先重排序选图再单次调用。

证据与结果

评测基准为ScreenSpot-Pro(8B规模骨干),对比对象包括Direct整图推理、Oracle上限、多裁剪方法ZoomIn(2次VLM调用)和MVP(3次以上调用)。Figure1显示Direct与Oracle差距超18%。RankGround在全部测试骨干和屏幕尺度上,平均定位精度比次优方法高5.5%,推理速度快1.4倍。

打开论文原文
它要解决什么
能否只用一次VLM调用,就在高分辨率GUI截图上达到多次裁剪调用的定位精度?
研究路径
步骤分四步:1)从已有grounding数据集用严格包含准则筛选候选裁剪图,并用边界感知正样本增强构造排序监督;2)GroundRanker先用pointwise目标学习裁剪图是否包含目标;3)再用listwise目标在候选集内排序,区分视觉相似裁剪图的细微语义/空间差异;4)推理时对密集候选裁剪图打分,选最优一张送入VLM,单次调用输出坐标。
这对工程意味着什么
第一步行动:在高分辨率GUI定位任务里,先加一个轻量重排序器筛选候选裁剪图,再单次调用VLM,省去多次裁剪调用。要避开的捷径:不要默认裁剪调用次数越多精度越高,实际会增加延迟,还可能产生冲突预测。
证据定位
ScreenSpot-Pro(8B规模)上,Direct整图推理与Oracle上限差距超18%;多裁剪方法ZoomIn需2次、MVP需3次以上VLM调用才接近该上限。RankGround只用1次调用,在全部骨干和屏幕尺度上平均精度超次优方法5.5%,推理快1.4倍。(筛选维度:可复核评测、GUI Agent 方法)
适用边界
排序监督数据由已有grounding数据集构造,不是原生排序标注,效果依赖严格包含准则的判定质量。评测限于ScreenSpot-Pro等公开基准及文中列出的骨干规模,未覆盖更广场景。
方法与英文摘要

作者先从已有GUI grounding数据集构造裁剪图排序监督:用严格包含准则判定裁剪图是否覆盖目标,并做边界感知正样本增强,改善密集布局下的空间覆盖。GroundRanker分两段训练:先用pointwise目标学习粗粒度包含关系,再用listwise目标区分视觉相似的裁剪图。推理时它对密集候选裁剪图打分,选出最优一张送入VLM,单次调用输出坐标。

Graphical User Interface (GUI) grounding is a fundamental perception task for multimodal agents, enabling them to interpret natural language instructions and interact with digital interfaces. Existing methods face a fundamental trade-off between accuracy and efficiency: direct full-image inference often fails to capture small or visually similar UI elements, while multi-crop strategies improve localization at the cost of multiple expensive Vision-Language Model (VLM) calls per query. To address this challenge, we propose RankGround, a two-stage framework that achieves accurate GUI grounding with a single VLM call per query. Central to our approach is GroundRanker, a lightweight multimodal reranker that identifies the most promising crop from a dense candidate set. Because no off-the-shelf ranking dataset is available, we construct ranking supervision data from existing grounding datasets. A strict containment criterion and boundary-aware positive augmentation improve alignment and spatial coverage in cluttered layouts. GroundRanker is then trained with a two-stage curriculum: a pointwise objective first learns coarse containment, and a listwise objective refines subtle semantic and spatial distinctions among visually similar crops. Experimental results show that RankGround consistently outperforms strong baselines while reducing computational cost. It achieves 1.4 times faster inference and improves localization accuracy by 5.5% on average over the second-best method across all backbones and screen scales, establishing a new state of the art in both efficiency and precision for GUI grounding.

个人知识与本体(1 篇)

个人知识与本体 6/30

TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment

记忆系统还不能既敢纠错又不乱插话:上线前只看召回,会把“乱拦截”或“漏拦截”的记忆系统放进生产;TWIST用表面相似的安全草稿硬负例把误报代价量出来:161条人工核验草稿校验数据(κ=0.85)上,flat-RAG误报16%-43%安全草稿,已部署连贯性系统特异性0.98-1.00却只捕获42%真实矛盾。

两句看懂

长对话记忆评测常只问能否召回,不问系统该不该在信念变化时主动介入或拦截,TWIST给每个检测指标配等量安全草稿对照来定价误报。161条人工核验草稿校验数据(κ=0.85)上,flat-RAG误报16%-43%安全草稿,已部署连贯性系统特异性0.98-1.00但仅捕获42%真实矛盾。

核心判断

当前记忆系统做不到“该出手时出手、不该出手时不出手”:161条人工核验数据显示检测率与误报控制此消彼长;13种基线对比确认差距主要来自检索覆盖不足,而非任务本身不可解。

关键要点

1. 旧基准把记忆当召回,不测信念变化点是否该纠错、拦截或拒答,介入失误在标准召回分上不可见。 2. TWIST在LoCoMo上扩四轨道,并给每个检测指标配表面相似安全草稿硬负例;Track B完成人工验证,161条,双人盲标+仲裁,κ=0.85。 3. 13种基线无三指标兼优:flat-RAG召回0.76-0.97但误报16%-43%,部署系统特异性0.98-1.00但召回42%;金证据召回1.000说明差距指向检索覆盖,上线要同时卡检测率和误报率。

证据与结果

数据源为LoCoMo语料及评测框架扩展。已公开人工验证版是Track B草稿校验v1.0,161条,双人盲标+仲裁后κ=0.85。指标为矛盾检测召回、硬负例特异性、证据归因准确率。结果:flat-RAG召回0.76-0.97、误报率16%-43%;部署中连贯性系统特异性0.98-1.00、召回仅42%;金证据基线召回1.000;全文本模型近乎解题;仅草稿基线暴露模型风格先验差异。

打开论文原文
它要解决什么
部署中的对话记忆系统,能否在信念变化点正确介入:检出矛盾、拦截错误草稿;同时不把安全内容误报成问题?
研究路径
1)基于LoCoMo构造四轨道任务,每个正例配表面相似安全草稿负例;2)独立标注者盲标后双人核验加仲裁定金标,Track B为161条,κ=0.85;3)用LLM裁判诱饼样本校准评分可信度;4)做可分性审计确认任务非纯检索可解;5)跑13种基线(flat-RAG、金证据、全文本、仅草稿等),比较召回、特异性、归因三项。
这对工程意味着什么
第一步:上线前用表面相似的安全样本做硬负例压测,报告检测率同时报告误报率。要避开的捷径:只报单一召回分,它会掩盖16%-43%误报或42%召回这类真实代价。
证据定位
Track B(161条,κ=0.85)显示:flat-RAG召回0.76-0.97,但误报16%-43%安全草稿;已部署连贯性系统特异性0.98-1.00,但召回仅42%;金证据基线召回1.000,全文本模型近乎解题,13种配置无一同时兼顾召回、特异性、归因。(筛选维度:置信度与不确定性、可复核评测)
适用边界
仅Track B(草稿校验)完成人工验证并公开161条金标;Track A、C、D只有任务模板和指标设计,未做人工验证,计划v2实现;论文标注版本号为v0.3草稿。
方法与英文摘要

用LoCoMo语料及评测框架扩出四条轨道:矛盾检测、草稿校验、信念更迭、敏感召回管控;已人工验证的是Track B草稿校验,161条金标,独立盲标双人标注加仲裁,κ=0.85。每个检测指标都配表面相似的安全草稿硬负例;对照13种基线配置,包括flat-RAG、金证据、全文本、仅草稿。

Long-conversation memory benchmarks increasingly test recall and prompted knowledge updates, and recent work studies evolving user beliefs and memory state. TWIST is a proposed benchmark suite for a complementary, unmeasured property: intervention quality -- whether a deployed memory system, exercised through its own ingest/recall/vet surface, acts correctly at belief change points. Four tracks cover unprompted tension detection, vetting outgoing drafts against the record, answering with current beliefs while preserving supersession history, and governing sensitive recall. The suite extends LoCoMo's corpora and harness, pairing every detect/block metric with a matched do-not-over-detect control: surface-matched hard negatives price false intervention, so no track can be gamed by flagging everything. The benchmark itself is validated first: independent, gold-blind double annotation with adjudication, judge decoy calibration, and a separability audit. On the human-validated Track B v1.0 key (161 items, post-adjudication kappa = 0.85), no tested configuration simultaneously achieves high contradiction recall, high hard-negative specificity, and high attribution: flat-RAG baselines detect 0.76-0.97 of true contradictions but falsely flag 16-43% of surface-matched safe drafts depending on backend, while a deployed coherence-oriented system almost never over-flags (0.98-1.00 specificity) yet catches 42% of true contradictions -- a trade-off no recall-only score can see. A 13-configuration baseline ladder localizes causes: every gold contradiction is detectable from its evidence alone (recall 1.000), calibrated models nearly solve the track given the full transcript -- consistent with substantial retrieval-coverage gaps -- and draft-only floors reveal model-dependent style priors. A system's TWIST profile, beside its recall score, measures whether memory knows when to intervene and when not to.

人机协同与对齐(0 篇)

本轮没有通过深读证据门的重点论文。

本轮分类概览

同一论文只归入一个最先命中的赛道,避免重复计数;“新增候选”只统计首次展示的论文。

赛道新增候选重点
形式化与程序验证143
软件工程与仓库智能152
代码质量与优化71
UI 与 GUI Agent21
个人知识与本体171
人机协同与对齐00

近一个季度监测日历

北京时间。绿色表示有可阅读的新候选,灰蓝表示已监测但无新增,橙色表示部分降级;“无记录”不等于失败。

2026 年 8 月

一二三四五六日

2026 年 9 月

一二三四五六日

近 14 次监测窗口

仅展示公开源的聚合运行状态,不含提示词、全文或个人数据。

本轮新增候选(55 篇)

按赛道、评分和日期展开;中文标签用于导航,英文摘要用于核验。已展示过的旧候选不会每日重复。

形式化与程序验证(14 篇)

形式化与程序验证 · 7/30 · 2026-09-22Agentic-IC3:语义证明搜索在 IC3 中引入语言模型代理,用 RTL 设计信息引导泛化Agentic-IC3: Enabling Semantic Proof Search in IC3 Model Checking

IC3 is a state-of-the-art algorithm for hardware model checking that proves safety properties by incrementally constructing an inductive invariant consisting of a set of lemmas. Its effectiveness depends on generalization heuristics that identify useful lemmas and guide proof search. However, many leading IC3 hardware model checkers operate on lowered, bit-level representations, where high-level design relationships are difficult to exploit for generalization. Those operating at a higher level remain limited in exploiting high-level design structure and semantics. We present Agentic-IC3, built on Pono's word-level model-checking infrastructure, which integrates a language-model agent into IC3 to guide semantic proof search using register-transfer-level (RTL) design information. The framework exposes an agent-oriented interface to a persistent IC3 backend, allowing the agent to interact with an explicit, evolving proof state throughout verification. Across successive proof obligations, the agent relates intermediate proof states and solver feedback to the RTL and proposes high-level lemmas through both SAT and UNSAT generalization. Beyond generalization, the agent can introduce derived observation signals to express design relationships succinctly and obtain more informative feedback, and backtrack to revise proposals that lead to unproductive proof branches. The backend checks proposals before updating the proof state, preserving soundness and providing feedback for further reasoning. On a suite of 14 benchmarks spanning security information-flow verification and functional verification of communication protocols, processors, and functional units, Agentic-IC3 solves 10 cases within a one-hour timeout, including 4 unsolved by all three evaluated baselines: rIC3, Pono-IC3Bits, and A-IC3.

阅读 arXiv 原文
形式化与程序验证 · 7/30 · 2026-09-22LLM 生成可证完备的广义规划将 PDDL 语义保持地转成 Lean,由内核检查规划完备性证明Provably Complete Generalized Planning with LLMs

Generalized planning aims to compute a plan that solves all instances of a planning domain. Recent work has used LLMs to automatically generate and debug such generalized plans in the form of Python programs and achieved perfect test data coverage for several domains. However, whether these generalized plans are actually complete, i.e. solve all instances of the domain, could only be determined by manual evaluation. Here, we present an approach for automatically generating generalized plans in Lean together with proofs of their completeness relative to a specification of the domain constraints provided as input. We introduce a semantic-preserving PDDL-to-Lean conversion, and use an LLM to generate both the generalized plan and the formal proof that it solves every instance satisfying the domain constraints. The correctness of the completeness proof is determined by Lean's kernel. We evaluate our approach on 13 commonly used benchmark domains, using GPT-5.6-Sol as the LLM. For 12 of the domains we obtain generalized plans together with valid completeness proofs. This is a major advancement of the state of the art in automatic generalized-plan completeness proofs.

阅读 arXiv 原文
形式化与程序验证 · 6/30 · 2026-09-22SLED-IFV:硬件信息流验证分解以求解器校验的 LLM 引导流程,做功能简化与关系强化分解SLED-IFV: Solver-Validated LLM-Guided Decomposition for Scalable Hardware Information-Flow Verification

Formal hardware information-flow verification (IFV) provides strong guarantees against secret-dependent timing and control behavior, but often scales poorly on realistic RTL. We identify two recurring proof barriers in self-composed IFV: implementation complexity, where proof-hard datapath logic dominates even though the property needs only a compact boundary relation, and relational inductive complexity, where the proof depends on cross-copy public-control facts that the backend prover does not infer efficiently. To address them, we introduce two semantic proof decomposition forms: functional simplification, which replaces a proof-hard RTL region with a validated over-approximate summary, and relational strengthening, which exposes and proves the cross-copy relations needed for induction. We further present SLED-IFV, a solver-validated LLM-guided flow that automates the selection of these forms and their concrete targets. Given a self-composed miter and an oracle-free decision sheet, the LLM proposes a decomposition, then materializes it into proof artifacts under controller checks. The controller compiles the checked artifacts into proof obligations, and the formal verification backend remains the sole authority for acceptance. Across nine nontrivial benchmarks constructed from real RTL, SLED-IFV achieves up to 603x solver-only speedup and converts two 12-hour timeouts into completed proofs. The closed-loop flow produces verifier-accepted decompositions for all cases.

阅读 arXiv 原文
形式化与程序验证 · 6/30 · 2026-09-22定理证明搜索中生成器的直接优化把计算对齐训练扩展到树搜索,并另提搜索无关的预算分配损失Direct Optimization of Generators for Search in Automated Theorem Proving

Fine-tuned Large Language Models (LLMs) significantly advance Automated Theorem Proving (ATP), but are often deployed as guiding policies within tree search rather than for single-attempt generation. Recent work shows cross entropy is suboptimal for an LLM used in flat search strategies such as aggregation or filtering and that work has developed new loss functions to correct this misalignment. Extending this alignment to tree search is more challenging: proof discovery depends on exploration and recovery through off-trace states that supervised demonstrations do not reveal. We extend Compute-Aligned Training (CAT) to this setting through an abstraction of policy-guided search, deriving tractable, trace-supported losses. Alongside these search-aware losses, we introduce a search-agnostic uniform-allocation (UA) loss that accounts for the budget without specifying the specific search. Both induce scalar weights on per-tactic cross-entropy gradients. We characterize how off-trace behavior affects the search-aware weights, including conditions for vanishing approximation error at large budgets. On a Lean benchmark, both approaches achieve higher observed proof-success rates than cross-entropy across six search strategies, with strong results from a single shared UA adapter. Budget sweeps show larger gains over cross-entropy at 16 than at 256 expansions, implying CAT scales with test time compute.

阅读 arXiv 原文
形式化与程序验证 · 6/30 · 2026-09-21分形决策图:边缘实时分诊摘要自称零权重张量运行并给出若干提升,但证据仅为自述Universal Fractal Natural Language Decision Map: Real-Time Edge Triage Across Heterogeneous Domains

Deploying Large Language Models for runtime operational triage incurs prohibitive latency (>100-500 ms), high VRAM requirements (>4-8 GB), and excessive energy dissipation. Extending Mandelbrot Fractal Neural Synthesis (Dagli et al., 2026), this paper presents the Universal Fractal Natural Language Decision Map, realized via the werr machine-native edge reflex runtime and the production answerr platform (https://answerr.me). Operating entirely without stored weight tensors (0 Bytes VRAM), the engine synthesizes deterministic decisions---noul (Boolean), choice (categorical), and score (ordinal)---by dynamically modulating 24-byte coordinate seeds along the chaotic boundary of the Mandelbrot set and evaluating multi-scale escape dynamics. Drawing inspiration from biological System-One reflex arcs, the engine introduces: (i) an Auto-Seed Router with domain projector Phi_D yielding a +28.8% accuracy gain over linear baselines; (ii) an Information-Theoretic Semantic Token Damping Filter (T_desc = 0.045) insulating against prompt injections (0.0% empirical bypass; 95% Wilson CI: [0.0%, 27.8%]) while pruning iterations by 45.8% (accelerating throughput 2.5x to 3.31 ms latency); (iii) a Multi-Scale Harmonic Tripod Fusion; (iv) a Coupled Margin Expansion Operator (Pitchfork Bifurcation Offset); and (v) a Cyclic Z/9Z Modular Resonant Grid Discretization based on the closed sub-ideal {0,3,6} (Lean 4 Mathlib ZMod 9), reducing FLOPs by 68.4%. Evaluated on JevBench (N=231), werr achieves 100.00% TypeSafe compliance and 81.65% calibrated accuracy with 7.08 ms median latency. We provide an OpenAI-compatible API and demonstrate deployment on 32-byte EVM smart contracts via the open-source werracle on-chain oracle (21,438 gas).

阅读 arXiv 原文
形式化与程序验证 · 4/30 · 2026-09-21Lara:面向自主科学的机检语言把研究论证编码为可执行工件,支持自动校验与跨论文桥接Beyond Natural Language: An Agent-Native Language for Autonomous Science

As autonomous AI agents take on every stage of scientific inquiry, research output is expanding far beyond human review capacity. Yet scientific communication still relies on natural-language prose: an informal medium prone to ambiguity, hidden assumptions, and untracked limitations that machines cannot reliably audit. We introduce Lara, a machine-checkable language and protocol for checking and revising support for research claims. By turning research arguments into executable artifacts, Lara provides an epistemic kernel for autonomous science: it enables automated validation pipelines for research agents, lets declared bridges connect arguments across papers into an auditable network, and allows both humans and machines to recheck the standing of an encoded claim in milliseconds. In a Lara program, authors explicitly declare their claims, supporting evidence and assumptions, and known objections or limitations. A lightweight, deterministic checker adjudicates these interactions, assigning each claim a reproducible status: "justified", "defeated", "contested", or "gap", which marks a claim whose support is incomplete and locates the unanswered question. Case studies cover empirical review, a philosophical debate without measurements, and the loss of support when an assumed axiom is withdrawn. We establish the metatheory of claim checking and cross-context argument transport, and mechanize the semantic guarantees in Lean 4 (roughly 117,000 lines), leaving three arguments on paper. The audited public metatheory is "sorry"-free and uses only Lean's three standard axioms; some executable examples additionally trust native evaluation.

阅读 arXiv 原文
形式化与程序验证 · 1/30 · 2026-09-21Lean Pool:AI 维护的形式化数学库摘要极短,仅称由 AI 代理扩增与优化,细节证据有限Lean Pool: An AI-Maintained Archive of Formalized Mathematics

Lean Pool is a repository of formalized mathematics. It is grown, maintained and optimized by AI agents.

阅读 arXiv 原文
形式化与程序验证 · 7/30 · 2026-09-21规范驱动的 AI 辅助开发生命周期结合 AGENTS.md 与分阶段技能文件,在课程中以学生问卷评估A Lean and Spec-Driven AI-Assisted Software Development Lifecycle for Applied AI Education: The AI-SDLC Approach

AI coding agents increasingly support software development beyond code completion, including planning, implementation, testing, and repository-level task execution. Their practical use, however, often remains only weakly connected to established software engineering practices. The aim of this work is to develop and evaluate a lightweight, spec-driven lifecycle for governed agentic software engineering. The lifecycle combines established software engineering practices with repository-local guidance through specifications, AGENTS.md, and phase-specific agent skill files. The approach was developed in the context of the FHNW course AI-assisted Software Development and applied by students to business-oriented software use cases. Its educational and practical applicability is explored through a student survey combining closed rating items with open-ended questions. The contribution of this work is a process-oriented framework that enables AI coding agents to operate with bounded autonomy within an explicit, reviewable, and test-oriented software development lifecycle.

阅读 arXiv 原文
形式化与程序验证 · 10/30 · 2026-09-18SWE-Proof:机检证明的议题修复为已知正确补丁写规范并转为形式验证任务,减少对测试集依赖SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check correctness with held-out test suites, which are inherently incomplete and increasingly susceptible to memorization. Formal verification avoids both problems, but existing work covers only standalone tasks whose specifications are given as input, not real issues, which touch large repositories and state intent in vague natural language. We present Benchproofer, a pipeline that turns a coding task with a known correct patch into a formally verified one: it writes a specification for the new code, summarizes the existing functions that code calls with axioms, and admits an instance only after mechanical and adversarial gates agree. Applying it to SWE-bench Verified yields SWE-Proof, 500 real issues whose correctness is formally verified rather than tested, and it extends to SWE-bench Pro. Evaluating Claude Opus 4.8, we find that verification catches what tests miss: a quarter of test-passing patches admit counterexamples, which a structured natural-language specification does not fix, while a correct formal one lifts resolution from 85% to 95%. Writing that specification is the hard part: an agent that must write its own gains nothing over an unaided baseline, and only 56% of those specifications pass our audit. The usual failure is faithfulness, a specification that constrains part of the required behavior and leaves the rest free. Specification quality still tracks the outcome, failing on 92% of unresolved instances against 51% of resolved ones, making faithful specification synthesis a concrete open problem.

阅读 arXiv 原文
形式化与程序验证 · 4/30 · 2026-09-17Gardam 格无唯一乘积的机检证明用 Lean 4 给出显式有限子集反例,证明经内核检查A Machine-Checked Proof that Gardam's $\tilde{A}_2$ Lattice Does Not Have Unique Products

Kaplansky's zero-divisor conjecture asserts that the group ring of a torsion-free group over a field has no zero divisors. It holds for every group with the unique-product property, so a counterexample can only come from a torsion-free group without unique products. In lectures in 2021, Gardam announced that the torsion-free $\tilde{A}_2$ lattice $Γ= \langle a, b \mid a b a^2 b^{-1} a^2 b^{-2}, a b^3 a b^4 a^{-1} b \rangle$ does not have unique products and presented it as a new candidate group: it has property (T), and the known methods for proving the conjecture do not apply to it. To our knowledge, no proof of the announcement has been published. We give a proof checked by the Lean 4 kernel and stated against Mathlib's UniqueProds class. The witness is an explicit pair of finite subsets with $|A| = 32$ and $|B| = 28$ in which each of the 896 products coincides with another product. For 658 products the certificate is an identity in the free group; the other 238 certificates are explicit products of conjugated relators, 970 conjugates in all, checked by free reduction. A homomorphism onto $\mathbb{Z}/42$ shows that each pair $(u,v)$ differs from its partner $(u',v')$ as a pair of group elements, which is all the theorem requires. Together with a homomorphism onto the alternating group $A_4$ it also shows that the listed words are pairwise distinct, so the sets have exactly 32 and 28 elements. The witness and certificates come from an untrusted search program and are re-checked by Lean. The development uses only the axioms propext, Classical.choice and Quot.sound, with no sorry and no native_decide. The mathematical statement is Gardam's. To our knowledge this is the first verification in a proof assistant of a unique-product failure in a torsion-free group; torsion-freeness of $Γ$ is taken from Gardam and is not formalized here.

阅读 arXiv 原文
形式化与程序验证 · 0/30 · 2026-09-17大模型回答中的语言地缘差异对同一议题以多语言提问,回应倾向随提问语言不同而不同Geopolitical Divisions Across Languages in Large Language Models

People increasingly turn to AI chatbots for news and explanations of world events. But do they receive the same political answers when they ask in different languages? Here we show that the language of a question can change how the same AI systems assess the war in Ukraine. We ask GPT, Claude and Gemini to evaluate twenty statements about the war in 112 languages, collecting 67,200 responses. The balance between Russia-leaning and Ukraine-leaning responses differs across languages. When we group responses by countries' official languages, they follow a pattern resembling worldwide political divisions: relatively more Russia-leaning answers correspond to more favourable public views of Russia, less support for Ukraine in United Nations votes, and less aid to Ukraine. The broad pattern recurs across all three models and remains when individual statement pairs are removed. Our findings suggest a possible route through which information warfare may shape the text used to train AI models, which may in turn spread geopolitical biases.

阅读 arXiv 原文
形式化与程序验证 · 6/30 · 2026-09-17长周期自动形式化核心定理以共享蓝图协调 AI 证明代理,完成 Lean 4 机检证明Long-horizon autoformalization of a core theorem underlying MIP* = RE

Landmark mathematical formalizations have taken specialist teams years to complete. We present FormalFlow, a system that coordinates AI proving agents under human supervision to address statement drift and proof composition in long-horizon formalization. Drawing on software engineering principles and practices, it uses a shared blueprint to guide nested planning, proving and review loops. Agents strengthen verification and review throughout formalization. We completed a machine-checked Lean 4 proof of the quantum soundness of the classical low individual-degree test, a core theorem underlying MIP* = RE. Developing the proof took 63 days; greater parallelism could further reduce this time. The final library contains 126,367 lines of Lean code, all generated by agents. The formalization corrects side conditions and intermediate errors while preserving the published final error bound under corrected assumptions. This work provides a verified foundation for quantum complexity and demonstrates a route to affordable verification of major research proofs by small teams.

阅读 arXiv 原文
形式化与程序验证 · 4/30 · 2026-09-17LLM 与 Lean 辅助的 LLVM 翻译验证生成结构化证明骨架,自动处理可由确定性方法解决的义务LLVM Translation Validation Automated with Large Language Models and Lean

LLVM is the cornerstone of modern compilers, but its subtle intermediate representation (IR) semantics make transformations error-prone and necessitate formal verification. Alive2, a state-of-the-art translation validator based on satisfiability modulo theories, has achieved substantial success in automating the validation of LLVM transformations. However, it still faces scalability limitations, does not support symbolic bitwidths, and offers only bounded guarantees for loops. In contrast, interactive theorem provers such as Lean can address these cases but require substantial proof engineering. In this paper, we present Trivet, a framework combining large language models (LLMs) and Lean for automated translation validation of LLVM transformations. Trivet generates structured proof scaffolds based on source and target functions, automatically discharges obligations amenable to deterministic reasoning, and delegates transformationspecific obligations to LLMs. It produces refinement proofs or counterexample-based refutations, with every successful verdict checked by the Lean kernel. On 148 LLVM transformations, Trivet verifies or refutes 147, leaving one invalid case unresolved. Successful cases include 60 loop-free transformations with symbolic bitwidths, 27 cases from a restricted class of loop-containing transformations, and 10 complex valid fixed-bitwidth cases on which Alive2 times out. Compared with an unscaffolded baseline, scaffolding enables 26 additional proofs. On cases solved by both configurations, it reduces mean proof time by 75.9% and mean monetary cost by 88%.

阅读 arXiv 原文
形式化与程序验证 · 7/30 · 2026-09-16MAGS:多代理自动形式化保安全以 Dafny 作验证中间表示,并用验证器反馈修复违规MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approaches, including fuzz testing, static analysis, and LLM-as-a-Verifier, can detect many failures but struggle to cover all possible edge cases. Formal verification addresses this by providing machine-checkable guarantees over specified properties, but traditionally demands substantial manual specification and proof engineering. We introduce a unified multi-agent framework, MAGS, that generates executable programs with formal safety guarantees, using Dafny as a verification-aware intermediate representation where safety properties can be mechanically checked. MAGS formalizes and freezes human-audited APIs and safety requirements, translates generated code into Dafny, repairs violations using verifier feedback, and compiles verified programs back into executable code. We evaluate MAGS on 100 CUDA kernels, 100 terminal scripts, and 20 robotic-arm tasks. Across all 220 examples, it achieves a 100% success rate in producing programs with non-trivial safety guarantees against frozen specifications. Independent safety and functional evaluations further show strong performance across all three domains, while revealing failures when the auto-formalized semantics do not fully capture the target behavior.

阅读 arXiv 原文

软件工程与仓库智能(15 篇)

软件工程与仓库智能 · 4/30 · 2026-09-24软件工程自我效能感量表开发87 题五维量表在 527 名本科生试测,仅报告初步效度与信度证据Design, development, and preliminary validity and reliability evidence of the Software Engineering Self-Efficacy Scale (SESES)

The purpose of this research is to design, develop, implement, and provide preliminary validity and reliability evidence of the Software Engineering Self-Efficacy Scale (SESES). Framed by a conceptual framework using guidance in software engineering curriculum and concepts along with the notion of self-efficacy, we generated an initial item pool of n = 87 items to operationalize and measure software engineering self-efficacy among undergraduate computing students. The conceptual framework traces five dimensions: 1) Requirements Engineering, 2) Teamwork and Collaboration, 3) Software Quality Management, 4) Software Design and Architecture, and 5) Software Agile Methodologies. We pilot tested the SESES with n = 527 undergraduate computing students who had completed a software engineering course in the current semester or a previous academic semester. We employed Exploratory Factor Analysis (EFA) with the Principal Axis Factoring method and an oblique (Promax) rotation to examine the underlying structure of the SESES, resulting in the same five internally consistent latent constructs in the conceptual framework with minimal cross-loading and a simple structure in the pattern matrix, explaining approximately 57% of the variability in these data. Our findings suggest that software engineering self-efficacy is a multidimensional construct of five theorized and correlated, yet distinct latent factors. We unpack the limitations and delimitations of the research while exploring undergraduate computing students' software engineering self-efficacy using necessary domain-specific measurements.

阅读 arXiv 原文
软件工程与仓库智能 · 11/30 · 2026-09-24规范驱动的自动程序修复基准主张以可执行规范生成基准,替代静态语料并降低污染风险Specification-Driven Benchmarking for Automated Program Repair From Static Corpora to Executable Specifications

Automated Program Repair (APR) benchmarks have traditionally been constructed as static datasets whose characteristics are inherited from the defects they contain. While this paradigm has enabled decades of progress, finite corpora provide limited experimental control, become increasingly susceptible to contamination as they are reused, and cannot be systematically regenerated or adapted as evaluation requirements evolve. We propose specification-driven benchmarking, a paradigm in which benchmarks are defined by executable specifications and realized through benchmark generation. The specification explicitly declares the intended properties of the benchmark (including program context, fault taxonomy, difficulty, validation strategy, and corpus constraints) while a generation pipeline realizes those requirements through independent generation, validation, and corpus management components. We develop the conceptual foundations of this approach by introducing a taxonomy of benchmark specification dimensions, establishing how each specification dimension maps to deterministic architectural responsibilities, and arguing that independent validation is a structural requirement for trustworthy benchmark generation. An end-to-end example illustrates how specification choices propagate through the pipeline to produce benchmark instances whose properties are independently verifiable. By treating the benchmark as an executable specification rather than a static dataset, the proposed paradigm shifts benchmark construction from artifact curation to declarative experimental design.

阅读 arXiv 原文
软件工程与仓库智能 · 0/30 · 2026-09-23需求演化差异可视化的 LLM 流程以语义图快照做并排比较,摘要简短,可核证据有限LLM-Assisted Workflow for Structural Difference Visualization in Evolving Software Requirements

This paper presents an LLM-assisted workflow for visualizing structural differences in evolving software require- ments. Implemented in the OntologyWeb environment, the work- flow represents baseline and current requirements as triple-based semantic graphs and supports side-by-side comparison of curated graph snapshots. The comparison view aligns matched entities and uses visual encoding to highlight structural changes.

阅读 arXiv 原文
软件工程与仓库智能 · 4/30 · 2026-09-22依赖更新 PR 的修复代理路由仅用创建时标题与元数据排序是否需修复代理,摘要给出一项 F1When Should Dependency Updates Invoke Repair Agents? A Lightweight Routing Study

Dependency-update pull requests are frequent and mostly routine, but a small subset requires non-trivial compatibility repair. Recent repository-level coding agents make such repair increasingly plausible, yet invoking them on every dependency update wastes model calls, CI time, repository context, and review attention. We frame this as a pre-agent routing problem: deciding which dependency-update pull requests should be escalated before downstream diagnosis or repair attempts. We introduce DepFixRouter, a lightweight router that ranks dependency updates by historical compatibility-repair likelihood using creation-time textual and metadata signals. On 497 labeled GitHub dependency-update candidates, only 72 require substantive repair. A creation-time-safe LinearSVC using only PR titles and bot/dependency flags reaches 0.488 repair F1 and captures 51.4% of repairs within the top 20% routed pull requests, improving calls per captured repair from 6.90 under route-all or random policies to 2.68. Retrospective full-history signals improve top-20% recall to 65.3%, revealing substantial hindsight leakage in pull-request histories rather than deployment-time routing utility. In a 60-case diagnosis-agent pilot, router-gated diagnosis reduces actual LLM calls by 66.7% and tokens by 66.1%, suggesting budgetaware escalation while measuring diagnosis rather than patch generation. DepFixRouter can serve as a lightweight escalation layer between routine dependency-update automation and expensive repository-level agents, enabling budget-aware maintenance without relying on retrospective repair evidence for deployment-time routing.

阅读 arXiv 原文
软件工程与仓库智能 · 7/30 · 2026-09-22AI 代理 PR 合并后的修复归属追踪已合并代理 PR 的后续修复,并与同期人工 PR 基线对比Who Finishes the Job? A Study of Follow-Up Fixes and Commit Authorship on AI Coding Agent Pull Requests

AI coding agents now author a large share of pull requests (PRs) merged into popular open-source projects. A merged agent PR is usually considered finished work; yet, prior studies have reported issues in agent code after the merge (e.g., code smells and static-analysis issues). However, little is known about how often a merged agent PR is fixed afterward, and who actually authors the fixing. In this paper, we follow 6,774 merged agent PRs across five AI coding agents (OpenAI Codex, GitHub Copilot, Devin, Cursor, and Claude Code) from the AIDev-pop dataset (open-source repositories with at least 500 stars) into their follow-up fixes, against a baseline of 5,044 contemporaneous human PRs from the same repositories. We link each merge to its candidate fixes, verify every candidate with human annotators and an LLM judge that matches human-level agreement (binary Cohen's Kappa=0.78 against a human-human K=0.77, Direct-fix precision 90%), and attribute the fixing work at the PR and the commit level. Our findings show that (1) merged agent PRs attract verified fixes at 1.62 times the odds of merged human PRs in the same repositories over the same period of time; (2) 69.6% of verified fixes in agent merges come from the same agent; and (3) 76.4% of the verified fix PRs are agent-authored throughout all commits. These results show that agents currently largely finish their own job, but their merges still require fixing more often than human merges.

阅读 arXiv 原文
软件工程与仓库智能 · 4/30 · 2026-09-22Galaxy 生态维护与支持的跨空间研究用议题、PR 与论坛讨论刻画维护关切及其跨空间关联Understanding Maintenance and Support in a Community-Driven Scientific Workflow Ecosystem: A Cross-Space Study of Galaxy

Galaxy is a widely used, community-driven scientific workflow system whose sustainability depends on continuous maintenance across its software, tools, workflows, infrastructure, documentation, and user-support ecosystem. However, maintenance knowledge in Galaxy is distributed across development and community-support spaces, making it difficult to understand what is maintained, how maintenance artifacts are resolved, and how user-facing concerns connect to repository-level development. We conduct a large-scale empirical study of Galaxy using 11,762 GitHub issues, 52,203 pull requests, and 6,235 Community Forum discussions. We characterize maintenance and support concerns, examine factors associated with resolution outcomes and resolution time, and investigate explicit and candidate connections among maintenance artifacts across these spaces. Using BERTopic modeling, we identify nine issue topics, 14 pull-request topics, and 14 forum topics, revealing a maintenance landscape spanning workflow execution, data management, tools and dependencies, infrastructure, testing, scientific resources, documentation, and user support. Resolution analyses show that coordination, diagnostic, contributor, automation, and engagement characteristics exhibit different associations with whether artifacts are resolved and how quickly resolution occurs. We further find limited explicit traceability between development and support spaces: 97.77\% of 16,426 resolved explicit relationships occur within GitHub, while only 294 connect GitHub artifacts with Community Forum discussions, despite additional semantic and technical relatedness across these spaces. Together, these findings characterize Galaxy maintenance as a distributed ecosystem-level process and identify opportunities to improve diagnostic reporting, lifecycle-aware triage, cross-space traceability, and the reuse of community-support knowledge.

阅读 arXiv 原文
软件工程与仓库智能 · 7/30 · 2026-09-22Python 跨操作系统可移植性问题重跑测试加议题分析,提出七类失败分类与诊断特征An Empirical Analysis of Cross-OS Portability Issues in Python Projects

While Python is designed as a cross-platform language, real-world applications encounter portability failures when deployed across different operating systems. We present the first large-scale empirical study of cross-OS portability issues in Python, analyzing 2,042 open-source repositories using two complementary approaches: systematic cross-OS test reexecution and manual analysis of GitHub issues. Our cross-platform testing of 500 projects reveals that 11.2% exhibit OS-dependent test failures. Through systematic analysis of 240 GitHub issues, we confirm 102 genuine portability problems spanning 95 additional projects. We develop a comprehensive taxonomy identifying 7 primary failure categories - with file/directory operations, process management, and library dependencies being most prevalent - along with 24 distinct sub-categories, 15 diagnostic signatures, and 4 systematic repair patterns. Our evaluation reveals that existing static analysis tools provide minimal support for portability detection, while large language models achieve 40-79% accuracy in identifying issues and 50-77% success in generating fixes when provided with structured guidance. Through 33 contributed pull requests, we demonstrate practical applicability and developer acceptance (17 merged, zero rejected) of our findings. This work establishes the first comprehensive baseline for understanding and addressing cross-OS portability issues in Python, providing actionable insights for developers, tool designers, and the broader research community.

阅读 arXiv 原文
软件工程与仓库智能 · 7/30 · 2026-09-21并行编码代理的语义协调基准同一测试分别跑单补丁与合并补丁,构造任务中干扰较常见Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development

Parallel coding agents can produce patches that work alone but fail when merged. This happens when one agent changes an interface or rule that another agent still relies on. We study these failures with stale, a benchmark for semantic coordination. Our evaluation runs the same tests on each patch alone and on their combination, counting only failures introduced by combining the patches. We use three tiers: synthetic tasks with controlled interface changes, pairs of merged pull requests, and constructed tasks that use real Django helpers. Among 834 runs on 417 mined Django pairs, only one showed interference after correcting the grading procedure. On constructed tasks using 12 Django helpers, interference occurred in 97% of runs. A message describing the completed concurrent change recovered 82% of runs. Reviewed pull requests may contain few unresolved parallel changes, even when agents fail on controlled tasks using real code. The constructed failure rates do not estimate how often these problems occur in practice.

阅读 arXiv 原文
软件工程与仓库智能 · 6/30 · 2026-09-19ChatGPT 增强修复的跨基准泛化在三套基准上比较增强方法,发现增益方向随模型与基准变化Towards the Generalizability of Leveraging ChatGPT in APR via Self-enhancing: An Empirical Study

Automated Program Repair (APR) increasingly relies on Large Language Models (LLMs). ChatGPT-enhanced APR uses techniques such as self-correction and autonomous agents to improve repair without modifying model parameters. Although these approaches report strong results on Defects4J and SWE-bench, the stability of enhancement gains across benchmarks remains under-explored. We evaluate three ChatGPT-enhanced APR methods on three representative, long-standing benchmarks. With GPT-3.5-Turbo, SRepair achieves a larger absolute gain on HumanEval-Java than on Defects4J, while SRepair and FixAgent without extrinsic information yield negative gains on BugsInPy. With GPT-5.4-mini, the evaluated methods achieve larger absolute gains on Defects4J than on HumanEval-Java, while gains on BugsInPy are non-negative but limited. We investigate benchmark-related factors through code transformations and benchmark-specific fine-tuning. Code transformations reduce enhancement gains on Defects4J, while benchmark-specific fine-tuning increases gains on BugsInPy. Directly supplying GPT-3.5-Turbo with error messages and triggering tests yields more correct repairs than the evaluated ChatGPT-enhanced APR methods on BugsInPy. These findings highlight the need to evaluate generalizability across benchmarks and models using multiple metrics, and suggest that directly providing repair-specific extrinsic information may be more effective than enhancement methods when their gains are limited.

阅读 arXiv 原文
软件工程与仓库智能 · 6/30 · 2026-09-18拉取请求为何沉默:完成障碍分析大量停滞 PR 与评审评论,归类贡献类型与停滞原因Why Do Pull Requests Go Silent? Uncovering the Barriers to Contribution Completion in Open-Source Code Review

Pull requests (PRs) underpin pull-based software development by enabling distributed code review and collaborative contribution in open-source projects. Yet many become inactive before integration and are eventually abandoned or closed, wasting contributor and maintainer effort. Although prior work has examined PR abandonment and review delays, less is known about the contribution types, discussion-level barriers, and post-stalling collaboration patterns associated with inactivity. We investigate which PR types most often stall, why inactivity occurs from authors' and reviewers' perspectives, and how stalling relates to later contributor and reviewer engagement. We analyzed 14,234 stalled PRs and 164,562 review comments from 19 popular GitHub repositories using stale-bot workflows. An LLM-based voting classifier categorized PRs by contribution type, while quantitative analysis was combined with qualitative coding of general and inline review discussions. Feature-enhancement and issue-fixing PRs formed the largest share, together exceeding 77% of classified stalled PRs. General comments linked stalling mainly to communication and coordination breakdowns, including missing interaction, delayed feedback, and unclear follow-up. Inline comments showed that inactivity does not always reflect disengagement: many PRs were blocked by technical or dependency issues, including failing checks, configuration problems, compatibility concerns, and environment mismatches. Only 39.56% of contributors later submitted another PR, and reviewer re-engagement with the same contributors was approximately 21%. PR inactivity is a socio-technical coordination problem involving communication, technical readiness, review ownership, and automation practices. We recommend type-aware triage, clearer review feedback, explicit ownership of next actions, CI blocker management, and cause-aware stale-bot interventions.

阅读 arXiv 原文
软件工程与仓库智能 · 3/30 · 2026-09-18文件排序对代码评审效果的影响挖掘大量多文件 PR,考察文件位置与后续修复变更的关联Does Order Matter? An Empirical Investigation into the Impact of File Ordering on Code Review Effectiveness

Modern code review is central to software quality, but its effectiveness depends on reviewer expertise, change characteristics, and how tools present changes. Most platforms display modified files alphabetically by default, although our prior work shows that developers find this ordering cognitively misaligned with how they understand multi-file pull requests. Whether these ordering-related attention patterns affect outcomes at scale remains unclear. We present a large-scale study of file ordering and review effectiveness, mining 330,343 multi-file pull requests and 756,814 file instances from 182 GitHub projects in five programming languages. We examine whether file position is associated with later bug-fixing changes, whether pull request size moderates this relationship, and whether reviewer attention, proxied by comments, aligns with latent bug outcomes. Results show statistically significant but modest associations among file position, pull request size, review activity, and latent bug likelihood. Latent bug rates rise from 56.7% at position 1 to 61.5% at position 30. Pull request size has a non-linear relationship with latent bug likelihood: pull requests of about ten files have the lowest risk, while very small and very large ones have elevated rates. Hurdle models show that attention is diluted as pull request size grows: each additional modified file reduces the odds of receiving any review comment by about 8.7%. These findings reveal an attention-effectiveness gap: visible review activity does not necessarily prevent defects. Alphabetical ordering is therefore not a neutral interface default, but a structural feature shaping attention allocation, review coverage, and confidence in review outcomes. We propose context-aware file ordering, dependency-aware grouping, risk-aware prioritization, and per-file coverage indicators to make review attention more visible and actionable.

阅读 arXiv 原文
软件工程与仓库智能 · 7/30 · 2026-09-17AdaRepair-Mem:自适应经验编排指出仓库级记忆检索的三点局限,提出覆盖感知的检索回退AdaRepair-Mem: Adaptive Experience Orchestration for Repository-Level Program Repair

Recent memory-augmented repository-level program repair methods reuse historical repair experiences to improve LLM-based issue resolution. However, our analysis reveals three limitations in existing repository-level memory retrieval. First, episodic memory is highly imbalanced across repositories, leaving low-resource repositories with little effective support. Second, more memory does not monotonically lead to higher repair success, suggesting that relevance, quality, and redundancy matter more than raw memory volume. Third, memory accumulation is phase-misaligned: repositories may contain many reproduction experiences but few patch or refinement experiences. To address these problems, we propose an adaptive experience retrieval framework for repository-level program repair. Our framework introduces coverage-aware retrieval, which falls back to cross-repository or repair-type-based memories when same-repository memory is insufficient; quality-aware selection, which ranks memories by relevance, historical utility, specificity, and redundancy; and stage-aware routing, which separates and retrieves memories for reproduction, localization, patch generation, patch refinement, and validation. Evaluated on SWE-Bench-Lite and SWE-Bench-Verified, the proposed framework improves repair performance on under-covered repositories, reduces noisy memory retrieval, and better supports failed-to-fixed patch refinement. Our results show that the key to memory-augmented repair is not simply accumulating more experiences, but retrieving the right experiences for the right repair context.

阅读 arXiv 原文
软件工程与仓库智能 · 9/30 · 2026-09-16LLM 程序修复实验设置的规范指出基准与修复数不足以规定任务,提出显式描述框架Experimental Settings in LLM-Based Program Repair: A Study of Inputs, Tool Access, Feedback, and Validation

Evaluations of automated program repair (APR) systems commonly report the benchmark, the number of repaired defects, and the tests used for final patch validation, but these items no longer fully specify the repair task presented to a system. Recent LLM-based systems differ in the information supplied before repair, the repository and testing operations permitted during repair, and the feedback returned after unsuccessful attempts, allowing the same benchmark to instantiate substantially different repair tasks ranging from localized patch generation to repository-level diagnosis and iterative repair. We present a framework for explicitly specifying the experimental settings associated with reported APR results. We analyze reported experimental settings from systems evaluated on Defects4J and SWE-bench and characterize each result by its task unit, fault-localization assumptions, initial input, tool access, repair-time feedback, final validation, and resource budget. Our analysis shows that benchmark identity alone is insufficient to reconstruct the evaluated task or determine the appropriate scope of comparison across reported repair rates. We therefore introduce a machine-readable schema for specifying each experimental setting to improve reproducibility and make the scope of cross-system comparisons explicit.

阅读 arXiv 原文
软件工程与仓库智能 · 4/30 · 2026-09-15OpenClaw 技能生态的增长与治理基于仓库历史与三次注册表快照,测度热潮后的留存与扫描After the Party: Growth, Governance, and Security Scanning in the OpenClaw Agent Skill Ecosystem

AI agents increasingly act through agent skills, i.e., natural-language instructions, that direct a host agent toward shell, network, credential, file, and process actions, and public registries distribute them at scale. In the first half of 2026, the OpenClaw AI agent went viral, and its public skill registry boomed: the observable stock nearly doubled in 91 days, and a majority of the listings visible in June were created in just two months. By the end of our study window, the wave had crested, and monthly listing creation and core-repository activity were falling from their spring peaks. This paper measures what the boom left behind, drawing on the OpenClaw Git history, its GitHub issues and pull requests, and three ClawHub registry snapshots. Attention is concentrated: the top 10% of skills received 46.93% of all downloads. No simple skill features (like size or download counts) remained a stable predictor of continued listing once creation cohort and skill age were controlled. Human scrutiny did not stay: 77.86% have zero stars and zero comments, while 85.06% of the readable skills carry privilege evidence. And automated cleanup is not ready: the three security scanners disagreed on 23,702 of the 61,990 skills they all cover. After human adjudication, weighted scanner sensitivity against the reference standard ranged from 21.67% to 61.06%. Governing fast-growing agent-skill registries cannot rely on simple metadata or single scanner scores; it requires robust, transparent measurement and independent validation.

阅读 arXiv 原文
软件工程与仓库智能 · 7/30 · 2026-09-15水下机器人视觉语言模型蜕变测试用多目标搜索找最少的图像变换以诱发错误预测Search-Based Metamorphic Testing of Vision-Language Models in Autonomous Underwater Robotic Software

Our industry partner focuses on quality assurance for industrial systems across multiple domains, including maritime systems, such as overwater vessels and autonomous underwater robots (AURs). Despite the strong performance of vision-language models (VLMs) in scene understanding, image captioning, and object recognition, their use in AUR software operating in underwater environments is underexplored. Therefore, in this context, it is important to evaluate the quality of VLMs for integration into AUR software and, so, automated software testing tools are needed to assess their suitability and improve their dependability. To this end, we propose a search-based metamorphic testing approach (MetaVLM) that identifies a minimal set of transformations on underwater images to induce incorrect model predictions, thereby revealing VLM failures. We employ NSGA-II as a multi-objective search algorithm and evaluate it over open-source VLMs, BLIP and CLIP, against a random search baseline. Results demonstrate the strengths and limitations of each VLM in the context of AUR software systems. Based on the results, we derive lessons for software engineering practitioners and researchers working on quality assurance of VLM-based software systems.

阅读 arXiv 原文

代码质量与优化(7 篇)

代码质量与优化 · 3/30 · 2026-09-24AI 辅助代码修复的社会技术瓶颈15 天工业 C++ 仓库个案,考察持续集成、评审与团队协作约束Orchestrating AI-Assisted Code Remediation: Socio-Technical Bottlenecks in a Large Industrial Repository

Background: Code degradation in large, long-lived codebases is costly to remediate through manual refactoring and opportunistic clean-ups. LLM-based coding assistants can perform mechanical remediation at scale, but their impact on industrial workflows is underexplored. Objective: We investigate how massive AI-assisted code remediation affects build-on-commit continuous integration (CI), code review, and team coordination in a large industrial repository, and which socio-technical bottlenecks constrain such remediation when source editing becomes cheap through AI assistance. Method: We report on a 15-day exploratory single-case field study in which an experienced developer used a command-line AI coding buddy to remediate widespread issues in a closed-source industrial C++ repository. We triangulate Gerrit metadata with a developer diary and team chat, analyzed through descriptive statistics and qualitative coding. Results: AI-assisted remediation rapidly generated hundreds of commits touching thousands of lines, saturating CI and reviewer attention. Naïve per-file commits overloaded build-on-commit CI; Switching to directory-based batching and capping the number of files per change restored throughput, but still required explicit review solicitation, negotiation of acceptable commit granularity, and iterative follow-up to resolve build and static-analysis failures. Conclusion: When mechanical editing is cheap, CI capacity, review effort, and change orchestration become primary bottlenecks. Sustainable AI-assisted remediation in very large repositories requires deliberate control of commit, review, and CI batch granularity and treating semantic change sets, such as ``fix all instances of warning X'', as first-class units of work that can be sliced differently for developers, reviewers, and CI.

阅读 arXiv 原文
代码质量与优化 · 3/30 · 2026-09-23AST 自动消除 Java 跳转语句引入辅助布尔变量重构控制流以保持语义,为抽取方法铺路AST-Based Automated Elimination of break and continue Statements in Java Code

This work presents the development of an automatic refactoring tool for Java code built on top of the Eclipse JDT API. The proposed approach transforms control structures containing break and continue statements within different types of loops into semantically equivalent constructs that avoid their explicit use. To achieve this, auxiliary boolean variables are introduced to restructure the control flow while preserving the original program behavior. The main objective of this transformation is to improve code structure and enable the application of subsequent automated refactorings, particularly those based on the Extract Method operation, which are typically restricted by the presence of jump statements. The implementation relies on the analysis and rewriting of the Abstract Syntax Tree (AST), ensuring semantic equivalence in all addressed scenarios. The tool was validated through 54 manually designed test cases and 151 units tests, all of which produced satisfactory results. In addition, it was applied to 139 methods from seven open-source projects, generating code without compilation errors and preserving the original behavior as verified by the projects' test suites. The results demonstrate that the proposed approach safely automates the restructuring of code containing break and continue statements, facilitating further evolution and structural analysis.

阅读 arXiv 原文
代码质量与优化 · 6/30 · 2026-09-20VSpector:RISC-V 规范驱动查错用自然语言规范直接检查 RTL 是否符合规则,无需参考模型VSpector: Specification-Driven Bug Detection for RISC-V CPUs

Detecting RTL design bugs in open-source RISC-V CPU implementations is critical for ensuring system reliability. Traditional detection approaches inherently rely on predefined artifacts. In this paper, we leverage the official,natural-language RISC-V specifications as an effective information source for bug detection. We present VSpector, a specification-driven bug detection pipeline that directly checks whether CPU register-transfer level (RTL) implementations adhere to official specification rules, without requiring specialized construction of reference models, formal properties, or custom bug patterns. To resolve the key technical trade-off between broad context scope and model reasoning accuracy when using Large Language Models (LLMs), VSpector employs a stepwise context refinement scheme across a four-stage pipeline: rule extraction, implementation localization, candidate identification, and sequential violation auditing. We evaluate VSpector on two industrial-strength RISC-V CPUs, CVA6 and XiangShan. Out of 217 reported candidates, manual inspection confirmed 148 true violations, representing a 68.2% precision. These violations correspond to 73 distinct bugs, including 42 previously unknown bugs. In our comparative experiments, DiveFuzz, a state-of-the-art CPU fuzzer, detected none of these new bugs during 24-hour runs per CPU. All 42 new bugs have been reported upstream, with developers already fixing 19 and confirming an additional 11 (30 in total), demonstrating that specification-driven auditing is a practical and complementary strategy for CPU bug detection.

阅读 arXiv 原文
代码质量与优化 · 0/30 · 2026-09-17闭环机器人软件的习得与迁移把验证选出的闭环实现作为可复用经验存档,用于新任务生成Learning and Transferring Closed-Loop Robot Software

Closed-loop robot policies require observation processing, state management, and situation-dependent branching, making them costly to design and tune manually. Although coding agents increasingly support control-code generation and optimization, it remains unclear whether implementations improved on source tasks also support policy acquisition for new tasks. We study this question by treating complete closed-loop implementations as reusable execution experience. For each source task, a coding agent generates policy code from a few successful demonstrations and iteratively improves it using simulation feedback. The validation-selected implementations are retained in a software archive. For new tasks, the agent generates and improves policies using archived implementations, target demonstrations, and execution feedback. The resulting policy is then frozen and executes without further model calls. Across four source tasks in RoboCasa, iterative optimization increases mean success from 28.3% to 64.2%. Across nine target tasks and three independent runs, mean success is 45.2% without references, 41.5% with initial source code, and 57.0% with optimized source code. Optimized references outperform initial references in all three runs on the nine-task average, with a mean gain of 15.6 percentage points. These results demonstrate the value of execution-improved software as a resource for acquiring new policies in this setting, although initial references remain better on two target tasks when averaged across runs.

阅读 arXiv 原文
代码质量与优化 · 0/30 · 2026-09-17AutoData:预训练数据选择的智能搜索把数据选择视为可执行算法搜索,用代理模型反馈迭代改进AutoData: Agentic Search for Pre-training Data Selection

LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop. We frame pre-training data selection as heuristic engineering over per-document features, i.e., lexical statistics, categorical labels, and perplexity. We introduce AutoData, an agent that searches directly over executable selection algorithms. Unlike prior data mixture methods that optimise weights over a fixed set of domains, AutoData searches a richer program space of scoring, stratification, and stochastic selection rules, discovering feature interactions automatically by iteratively refining algorithms with validation feedback from a proxy model. Within an overnight search, AutoData discovers a selection algorithm that outperforms existing human-designed curation pipelines. Despite being searched only on this small proxy, the discovered recipe transfers to larger scales and improves the downstream metric CORE. These results suggest that data engineering can be treated as an agentic machine learning problem, extending autonomous research from model and training-code optimization to the data.

阅读 arXiv 原文
代码质量与优化 · 3/30 · 2026-09-15PrimeScientist:研究投入的策略分配把方向选择与资源投入建为序贯决策,用可执行计划树保留备选PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research

Autonomous research agents aim to automate scientific workflows, from proposing ideas to conducting experiments and analyzing results. Yet current AI and research agents can propose more directions than available resources allow them to pursue. Moreover, each attempt could consume substantial resources, requiring agents to reconsider how to invest in subsequent research. Thus, deciding how to invest research effort strategically should be a defining capability of autonomous research agents. Accordingly, we introduce PrimeScientist, which jointly determines research direction and resource investment across successive research attempts. Specifically, we formulate this challenge of strategic research effort allocation as a sequential decision problem where remaining resources should explicitly guide the research policy. We first introduce an executable plan tree that preserves competing plans and their outcomes across attempts. Building on this representation, we propose an adaptive MCTS-based allocation policy that balances exploration and exploitation using experimental feedback and remaining resources. Comprehensive evaluations across AI research, systems and code optimization, and machine learning engineering show that strategic allocation improves research quality and sample efficiency together. Across 12 AI research tasks, PrimeScientist improves average reward by 10.3% with 50.6% fewer research attempts than AutoResearch under the same resource budget. We believe making research effort allocation an explicit optimization target establishes effective resource use as a core research capability for autonomous agents to drive scientific breakthroughs at scale.

阅读 arXiv 原文
代码质量与优化 · 3/30 · 2026-09-15非对比表示学习的四型克隆检测基于 VICReg 加跨层一致性与深度加权,规避负采样偏差Type-IV Code Clone Detection via Layer-Wise Non-Contrastive Representation Learning

Software clones are fragments of code that are similar or functionally equivalent to each other. They pose significant challenges for maintenance, refactoring, and bug detection. Detecting Type-IV clones, which are semantically equivalent but may differ syntactically, is particularly difficult for traditional token- or syntax-based methods. Recent machine learning approaches rely on contrastive learning, which requires careful negative sampling and can introduce bias. In this paper, we propose LWVIC4Code, a non-contrastive representation learning approach specifically designed for Type-IV clone detection. Building on the Variance-Invariance-Covariance Regularization (VICReg) framework and prior layer-wise VICReg training, LWVIC4Code introduces cross-layer consistency regularization and depth-dependent layer weighting to progressively refine semantic information across transformer layers, producing robust and discriminative code representations. We conduct an empirical study comparing LWVIC4Code against a contrastive learning baseline and zero-shot large language models on Python (Kamino) and multi-language (GPTCloneBench) datasets. Results show that LWVIC4Code achieves competitive or superior performance without negative samples, benefits from layer-wise supervision, and generalizes effectively from Python to other languages, particularly Java and C#. These results demonstrate that non-contrastive, layer-wise representation learning is a promising direction for robust semantic code clone detection.

阅读 arXiv 原文

UI 与 GUI Agent(2 篇)

UI 与 GUI Agent · 5/30 · 2026-09-24Jev-Mobile:移动 GUI 轻量执行低频 VLM 规划配高频类型化决策模型,摘要称 AndroidWorld 达 79% 成功率Jev-Mobile: Jev as an Executor for Mobile GUI Agents

Vision-language models (VLMs) have become a common foundation for autonomous mobile GUI agents, but most existing systems rely on the VLM for both planning and action grounding at nearly every interaction step, leading to substantial latency and model-serving cost. We introduce Jev-Mobile, which shifts this paradigm to low-frequency VLM planning and high-frequency lightweight execution: the VLM specifies local goals, the accessibility tree defines a structured executable action space, and Jev, a fast typed decision model, repeatedly selects actions within this space. This design allows multiple GUI actions to be executed under a single VLM decision, reducing expensive VLM inference while preserving adaptive interaction. On the full AndroidWorld task suite, Jev-Mobile achieves 79% task success, compared with 78% for SeeAct-V and 84% for a Step-wise VLM baseline. Among successful trajectories, it reduces mean end-to-end execution time by 32.7% and mean model API cost by 73.4% relative to Step-wise VLM. These results show that decoupling high-level VLM reasoning from low-level action execution can substantially improve mobile GUI agent efficiency while maintaining competitive task performance.

阅读 arXiv 原文
UI 与 GUI Agent · 6/30 · 2026-09-16RankGround:轻量重排引导裁剪两阶段框架以单次 VLM 调用完成高分辨率 GUI 定位RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection

Graphical User Interface (GUI) grounding is a fundamental perception task for multimodal agents, enabling them to interpret natural language instructions and interact with digital interfaces. Existing methods face a fundamental trade-off between accuracy and efficiency: direct full-image inference often fails to capture small or visually similar UI elements, while multi-crop strategies improve localization at the cost of multiple expensive Vision-Language Model (VLM) calls per query. To address this challenge, we propose RankGround, a two-stage framework that achieves accurate GUI grounding with a single VLM call per query. Central to our approach is GroundRanker, a lightweight multimodal reranker that identifies the most promising crop from a dense candidate set. Because no off-the-shelf ranking dataset is available, we construct ranking supervision data from existing grounding datasets. A strict containment criterion and boundary-aware positive augmentation improve alignment and spatial coverage in cluttered layouts. GroundRanker is then trained with a two-stage curriculum: a pointwise objective first learns coarse containment, and a listwise objective refines subtle semantic and spatial distinctions among visually similar crops. Experimental results show that RankGround consistently outperforms strong baselines while reducing computational cost. It achieves 1.4 times faster inference and improves localization accuracy by 5.5% on average over the second-best method across all backbones and screen scales, establishing a new state of the art in both efficiency and precision for GUI grounding.

阅读 arXiv 原文

个人知识与本体(17 篇)

个人知识与本体 · -3/30 · 2026-09-25PIA:健康对话转为结构化记录主张健康代理记忆需类型化记录与时间规则,摘要式检索难支持趋势问答PIA: A Personal Intelligence Agent Turning Health Conversations into Records and Records into Understanding

General-purpose agent memory summarizes conversations: it extracts salient snippets, embeds them, and retrieves the top-k into the prompt. A health agent cannot run on summaries: a dose becomes a sentence, "since last week" is resolved at the model's discretion, and a three-month glucose trend cannot be answered by text similarity. We present PIA, a personal intelligence agent deployed alongside a consumer health agent. PIA receives the agent's natural-language requests, decides for itself whether and how to write or read, and turns conversations into typed clinical records and records into a synthesized understanding of the user. Its memory harness consists of four controls -- extraction, memory, retrieval, and understanding -- each a domain-agnostic mechanism with a pluggable health module: schema, medical alias dictionary, knowledge graph, and temporal rules. We show how the same query receives a different answer as the memory injected into the response context deepens from one-dimensional recall, to a two-dimensional health snapshot, to a three-dimensional trajectory with causality, and report lessons from operation: self-reported health data are missing not at random, question phrasing governs the quality of synthesized understanding, and nearly a third of candidate causal links are structural noise that rules alone remove.

阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-25共享代理记忆的信念准入基准CPB 评估主张是否应写入共享记忆,关注来源去重与答案覆盖的权衡A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory

Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent agents to it. To study this problem, we introduce the Correlated Promotion Benchmark (CPB), which evaluates whether candidate claims should be admitted to shared memory.CPB-Static constructs a frozen test split from publicly annotated sources with fixed gold actions. CPB-Live runs multi-agent teams over a shared store, records all writes and retrievals, and tracks source lineage defined by each scenario. A separate consumer answers from the store alone. We evaluate eight admission policies across four agent families. Our results show that policies which deduplicate sources reject many true claims alongside false ones, whereas policies preserving answer coverage admit nearly as many false claims as unrestricted sharing. Gating on declared source type reduces false adoption to 0.06--0.09, compared with 0.22--0.47 for other answering policies. Once an uncontested false belief enters memory, the consumer asserts it in 0.97--0.99 of probes across all families. No non-oracle policy consistently rejects false claims across verbatim copies, paraphrases, and paraphrases declared authoritative. These findings reveal the limitations of admission policies without access to source lineage.

阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-24MemProbe:代理记忆稳定性探查借认知实验范式诊断记忆更新与保持权衡,并分解为行为剖面Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms

Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cognitive-science-inspired framework for diagnosing stability-plasticity tradeoffs in agent memory. The framework is motivated by a core insight from cognitive memory research: memory is reconstructive and shaped by interference, source reliability, reinforcement, and reactivation. MemProbe turns this insight into four reusable experimental paradigms (interference, misinformation, consolidation strength, and reconsolidation window) that manipulate when a memory should be updated, preserved, or treated as uncertain. It further decomposes correctness into behavioral profiles that reveal how systems update, preserve, attribute, and temporally organize information. We instantiate these paradigms in a 56-episode diagnostic suite and evaluate six incremental memory systems under a unified protocol. Results show that systems with similar aggregate scores exhibit distinct behavioral profiles. MemProbe provides such a diagnostic lens, turning aggregate performance into interpretable profiles of memory maintenance over time. Code is available at https://github.com/jq-ding/MemProbe.

阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-24生产规模 AutoResearch 的失败模式十二周运行报告五种失败模式,含记忆衰减与搜索方向停滞AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework

Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language model that iteratively edits a training script and retains modifications that improve a held-out scalar metric -- to automate this exploration. We report on twelve weeks of running this paradigm at production scale, where iterations consume hours of multi-GPU compute, evaluation involves competing criteria, and campaigns span weeks across many training jobs. Across two independently developed representation-learning systems for a book recommendation pipeline, we ran 220+ experiments and observed five recurring failure modes absent from the original setting: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation. We contribute a three-principle scaffolding design -- prevent, persist, redirect -- that maps each failure mode to a structural remedy and whose instantiation scales with iteration cost. The framework produced a 1.82x Recall@6 lift and a 2.1x coherence lift over hand-tuned baselines, and the agent autonomously designed a text-only fallback that expanded catalog coverage by 5.8x. The two systems span nearly three orders of magnitude in per-iteration cost yet exhibit the same failure modes, suggesting these are structural properties of production-scale autonomous research rather than artifacts of either application.

阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-24持久记忆的作用域匹配与干扰按任务族限定技能检索范围,摘要称可减少有害部署次数Scope Before You Persist: Preventing Cross-Family Interference in Agent Memory

Persistent memory lets language-model agents improve prompts and skills without updating model weights. We show that matching retrieval scope to certification scope enables these edits to support reliable repeated adaptation across recurring task families. We study frozen-model agents on ProcStream-RSI, a 12-round code-repair stream, using Orthogonal Regression Control (ORC), an execution-grounded gate for persistent skill edits. In an intervention that holds proposals and gate decisions fixed, retrieving each accepted skill only for its originating family raises mean hidden trajectory utility from 0.713 under global memory to 0.816 and changes harmful deployments from six of eight to none. In 27 paired randomized-order streams, Scoped-ORC improves mean trajectory utility by 0.063 [0.037, 0.094] over Global-ORC, accepts 63 rather than 12 updates, and produces multiple accepted updates in 19/27 streams, with 0/63 harmful acceptances. The global control reaches 0.713, below the static agent's 0.775, because locally valid edits can interfere with unrelated families. These results establish scope matching as a complementary control for persistent agent memory: certification determines whether an edit is supported, while retrieval scope determines where that evidence authorizes its use.

阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-23META:情景记忆增强交易代理多指标代理加记忆模块,构建无状态分析之外的决策框架Agent Memory with Episodic Retrieval for Financial Decision-Making

Large language models (LLMs) have demonstrated strong capabilities in financial analysis and reasoning, inspiring recent advances in agent-based trading frameworks. While these systems show promise, prior approaches either emphasize long-horizon forecasting or operate as stateless analyzers, limiting their applicability to the demands of trading in complicated settings. To address these gaps, we introduce META (Memory Enhanced Trading Agent), the first RAG-like episodic-memory-augmented multi-agent framework for financial decision making. META integrates a family of specialized indicator agents (e.g., Trend, MACD, Stochastic, RSI, SMA, AVWAP, Heikin-Ashi) with a Decision Agent that fuses their reports, and a Memory module that retrieves and updates past trading episodes encoded as market state embeddings with outcomes and reflections. By recalling relevant experiences and adaptively reweighting signals under similar market regimes, META achieves improved directional accuracy and robustness under short-horizon evaluation. Our results demonstrate that episodic memory provides a powerful mechanism for regime-aware, interpretable, and low-latency decision-making in trading and decision making. The code of this project is released on GitHub.

阅读 arXiv 原文
个人知识与本体 · 6/30 · 2026-09-23TWIST:对话记忆干预质量基准四条赛道考察信念变更时是否恰当干预,并配对假阳性控制TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment

Long-conversation memory benchmarks increasingly test recall and prompted knowledge updates, and recent work studies evolving user beliefs and memory state. TWIST is a proposed benchmark suite for a complementary, unmeasured property: intervention quality -- whether a deployed memory system, exercised through its own ingest/recall/vet surface, acts correctly at belief change points. Four tracks cover unprompted tension detection, vetting outgoing drafts against the record, answering with current beliefs while preserving supersession history, and governing sensitive recall. The suite extends LoCoMo's corpora and harness, pairing every detect/block metric with a matched do-not-over-detect control: surface-matched hard negatives price false intervention, so no track can be gamed by flagging everything. The benchmark itself is validated first: independent, gold-blind double annotation with adjudication, judge decoy calibration, and a separability audit. On the human-validated Track B v1.0 key (161 items, post-adjudication kappa = 0.85), no tested configuration simultaneously achieves high contradiction recall, high hard-negative specificity, and high attribution: flat-RAG baselines detect 0.76-0.97 of true contradictions but falsely flag 16-43% of surface-matched safe drafts depending on backend, while a deployed coherence-oriented system almost never over-flags (0.98-1.00 specificity) yet catches 42% of true contradictions -- a trade-off no recall-only score can see. A 13-configuration baseline ladder localizes causes: every gold contradiction is detectable from its evidence alone (recall 1.000), calibrated models nearly solve the track given the full transcript -- consistent with substantial retrieval-coverage gaps -- and draft-only floors reveal model-dependent style priors. A system's TWIST profile, beside its recall score, measures whether memory knows when to intervene and when not to.

阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-23即时记忆:读时策展的代理记忆主张保留原始轨迹、推迟到读时策展,避免写入时不可逆丢弃Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents

Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. This forces the system to decide what is worth remembering before the future query is known, irreversibly discarding information and producing a query-independent summary that must serve many possible downstream tasks. Learning such a write-time curator is also difficult because the value of a storage decision may only become apparent when a relevant query arrives, potentially many tasks later, creating a long-horizon credit-assignment problem. We instead retain raw trajectories and defer curation until read time, when the current task is known. Given the retrieved traces and the new task, a memory curator synthesizes a compact, task-adaptive payload tailored to the immediate need. Because this payload is consumed on the same task, the curator can be trained directly from immediate task success, avoiding delayed utility signals and the need to artificially group related tasks. Across ALFWorld, WebShop, and $τ^2$-bench, our Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively. Notably, even an untrained curator is already competitive with or surpasses these baselines, showing that task-adaptive read-time curation itself is a major source of the gain; training the curator further compounds the improvement.

阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-23EnSIMem:实体结构化长期记忆离线建实体—属性索引条目并保留出处与时序,在线按证据需求检索EnSIMem: Entity-Structured Indexing for Long-Term Agent Memory

An agent that interacts with users over long periods must recall facts, preferences, events, and changes from a continuously growing interaction history. Existing memory systems often compress interactions into generic summaries or retrieve anonymous text chunks, making it difficult for an agent to identify the correct entity, property, and supporting evidence. We present EnSIMem, an entity-structured long-term memory architecture for an agent. During offline construction, the system organizes interactions into theme-coherent episodes and builds dialogue-grounded index entries of the form [entity][entity type][property:value]. Each entry preserves its source turns, temporal information, and available multimodal fields. During online interaction, the agent's request is decomposed into evidence requirements whose properties are aligned with the memory index. Entity-property lookup and adaptive retrieval then collect the evidence needed for point, temporal, compositional, and aggregation reasoning. The agent generates its response from the preserved source evidence rather than from lossy memory summaries. On long-term agent-memory benchmarks, EnSIMem achieves high answer accuracy while maintaining compact contexts and favorable online efficiency. These results show that entity-structured indexing and episode-level provenance provide a reliable foundation for long-term memory in agents. The code of our model is available at https://github.com/RamonMeng/EnSIMem.

阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-22AkasicMEM:受治理的企业记忆强调来源到记忆的授权连续性,防派生复用导致信息泄露AkasicMEM: Governed Enterprise Memory for Agents

Agent memory enables enterprise agents to retain knowledge acquired during work and reuse it across tasks and agents, turning execution experience into persistent organizational knowledge. Realizing this potential requires both source--memory integration, through which enterprise sources and accumulated memory can be utilized together, and memory governance, through which shared memory remains subject to organizational policies throughout its lifecycle. These requirements interact when information from enterprise sources persists in memory. As this information is repeatedly derived and reused under changing principals and policies, source restrictions may be bypassed, resulting in information leakage. Preventing such leakage requires authorization continuity, under which source restrictions remain effective throughout source-to-memory and memory-to-memory derivation and reuse. Existing approaches address these concerns individually, but do not treat source--memory integration, memory governance, and authorization continuity as combined core design targets across the memory lifecycle. We define Governed Enterprise Memory as agent memory designed around this combined scope and present AkasicMEM as its realization. AkasicMEM realizes authorization continuity through transitive lineage, policy composition during memory formation, and policy re-evaluation during retrieval. It is built on GraphAI's AkasicDB, a unified vector--graph--relational database whose storage and execution substrate enables the underlying operations of these mechanisms to be jointly optimized and executed.

阅读 arXiv 原文
个人知识与本体 · 5/30 · 2026-09-21DolphinBench:代理记忆帕累托前沿以任务完成度评记忆,用有历史与无历史对照验证任务有效性DolphinBench: Mapping the Pareto Frontier of Agent Memory

Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question itself signals that some fact must be retrieved, and often which one. Moreover, benchmarks rarely require anything beyond accuracy from submissions, allowing memory systems to make unreasonable cost/time tradeoffs to achieve higher scores. We present DolphinBench, a benchmark that evaluates memory directly through an agent's task completion. DolphinBench includes three knowledge-work personas with roughly 500k tokens of user messages per persona and evaluates agents on tasks that depend on information from that history. We verify all 200 tasks per persona by running an agent with and without the relevant history, requiring success with it and failure without it. Finally, we require all evaluations to report total cost and latency alongside accuracy, which enables us to evaluate agent memory systems holistically. No existing memory benchmark combines all three. The dataset and evaluation code are available at https://dolphinbench.ai.

阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-21代理 AI 安全事件报告要素依专家输入梳理报告内容,如记忆访问与工具使用,问题仍开放Beyond Predictable Paths: Redefining AI Security Incident Reporting for Agents

AI agents are being deployed rapidly, accompanied by a growing number of AI-specific attacks and corresponding incidents. As incident reporting becomes increasingly important for legal compliance, governance, accountability, and security; current frameworks must be adapted to the unique characteristics of AI agents. In this paper, two editorial authors compare AI systems and AI agents and, drawing on input from 23 experts in academia and industry, identify the information required for reporting incidents where the security of AI agents is harmed. %involving AI agents. Potential reporting elements include, for example, agent memory and memory accesses, actual and potential levels of autonomy, and tool usage. Based on these findings, we identify several open research questions, including how to efficiently record incidents and how to determine whether vulnerabilities and incidents generalize. Expert feedback also highlighted potential reporting weaknesses, such as risks of data leakage and attacks targeting the reporting infrastructure itself, creating additional research needs. Lastly, we summarize privacy requirements and outline research directions for the secure and trustworthy deployment of AI agents.

阅读 arXiv 原文
个人知识与本体 · 5/30 · 2026-09-21MemCalib:代理记忆使用校准评估模型是否恰当使用记忆影响,发现常过度或不足使用MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents

The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over its response. Yet this capability has remained largely overlooked. To assess this capability, we introduce MemCalib, a benchmark grounded in realistic memory-system scenarios for evaluating memory use and advancing optimization algorithms. Results on the MemCalib test set reveal that frontier open- and closed-source models struggle to use memory appropriately. They frequently over-use or under-use memory rather than matching each proposition's actual use to its target level, leading to biased, low-quality responses. Experiments with common post-training algorithms, including group relative policy optimization and on-policy self-distillation, further reveal a clear directional skew: trained models improve in one direction while deteriorating in the other. We therefore propose MemCalib-RL, an ordered bidirectional counterfactual credit-assignment algorithm that separates over- and under-use signals and localizes their credit to response tokens through exact atom ablation. Results across model families and scales (Qwen3-8B, Ministral-3-8B-Instruct, and Qwen3.5-35B-A3B) show that MemCalib-RL achieves the best overall performance while better balancing over-use and under-use, with gains generalizing beyond MemCalib in external benchmark evaluation. Further experiments support its design choices and robustness and provide insight into its training dynamics.

阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-21Jev-Mem:系统一直觉控制记忆用快速控制面管记忆组织与检索,减少记忆路径上的生成开销Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

Agentic memory is becoming essential for long-horizon AI agents, yet many existing systems rely on autoregressive LLMs to control how memories are organized, retrieved, and used, placing expensive generation on the critical path of memory operations. We introduce \textbf{\method}, a new agentic memory architecture inspired by System-One/System-Two cognition. System One captures fast, lightweight decision-making, whereas System Two performs slower, deliberative reasoning. Jev-Mem brings this division of labor to agentic memory through a dedicated System-One control plane, a structured multi-relational memory plane, and a System-Two reasoning plane. The System-One controller governs memory typing and relational organization during construction, and dynamically performs query routing, retrieval-budget allocation, graph traversal, candidate scoring, and adaptive stopping during retrieval. System Two is invoked only for complex reasoning and answer synthesis. This design improves both memory effectiveness and system efficiency: on LoCoMo Jev-Mem achieves an overall LLM-as-a-Judge score of 0.777, an 11.0\% relative improvement over the strongest baseline, while reducing memory construction time to 158\,s, a 6.6$\times$ speedup over the fastest competing memory system, and lowering average query latency to 0.93\,s, a 36.7\% reduction.

阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-17自演化检索索引让索引自主诊断检索短板并修改索引键,减少人工介入Self-Evolving Search Index

Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index keys, retrieval quality depends heavily on how effectively these keys expose the knowledge contained in each document. However, effective index representations vary across retrieval environments, making it difficult for any fixed optimization strategy to perform consistently. Yet evolving an index to its retrieval environment remains largely human-driven, requiring humans to diagnose retrieval failures, refine the optimization strategy, and reprocess the index accordingly. We propose SELF-INDEX, a framework that enables an index to self-evolve without human intervention. Its Optimizer autonomously diagnoses retrieval shortfalls, selectively revises the responsible index keys, and validates each revision before updating the index. Beyond reacting to observed retrieval demands, SELF-INDEX proactively explores additional demands through a Query Simulator, allowing the index to evolve beyond the queries already available for optimization. Across diverse corpora and retrievers, SELF-INDEX consistently improves retrieval performance while outperforming existing index optimization methods. We further show that these benefits extend to downstream applications, improving the effectiveness and efficiency of search agents and helping agent memory systems retrieve useful past interactions.

阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-16双过程语言代理的记忆与反思扩展在 ScienceWorld 上做特征开关消融,摘要称全系统平均分最高Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments

Language agents remain brittle in interactive environments, where success requires long-horizon state tracking, valid action execution, and recovery from failed steps. We extend SwiftSage, a dual-process agent that combines a fast action proposer with a slower planner, using two modular cognitive extensions: an Adaptive Memory Module (AMM) for salience-gated episodic storage and trigger-driven retrieval, and a Self-Reflection Module (SRM) for bounded execution-time validation and corrective intervention. Both modules are implemented as feature-flagged extensions over the same execution substrate, enabling controlled ablations on ScienceWorld. Across four configurations---baseline, baseline+AMM, baseline+SRM, and the full system---the full system achieves the best mean final score (64.62), success rate (43.17%), and successful-step efficiency (19.33 steps), while SRM is the strongest standalone contributor. The results suggest that execution-time control is the dominant bottleneck in this setting, while episodic memory becomes most useful once the runtime loop is stabilized.

阅读 arXiv 原文
个人知识与本体 · 4/30 · 2026-09-16WFM:复杂代理推理的知识表示主张稀疏图与稠密文档结合的代理原生知识表示,细节有限WFM: Wiki Foundation Model for Complex Agentic Reasoning

Real-world agents fundamentally require persistent non-parametric knowledge for dynamic reasoning, i.e., long-term memory and retrieval-augmented generation. While graphs have shown reliable advantages in providing structured evidence, the sparse graph representations naturally restrict machine readability and semantic density required for complex agentic workflows. Driven by this limitation, the entire industry is witnessing a paradigm shift from traditional sparse graphs to LLM Wiki, an agent-native knowledge representation that couples dense document contexts with markdown files containing multi-layered topological linkages. However, parameterizing such rich semantics is challenging to encode dense textual contexts using traditional sparse graph embeddings. Moreover, learning LLM Wiki with existing graph encoders could overwhelm distributed system overheads that hinder deployment in large-scale commercial scenarios. To this end, we propose a novel paradigm Wiki Foundation Model, i.e., WFM, tailored for scalable, agent-native representation and retrieval. Specifically, (i) we formalize a Wiki Graph schema that seamlessly bridges fine-grained structures with dense contexts, maintaining explicit topologies alongside continuous semantics; (ii) A query-conditioned attentive aggregation is tailored for rich wiki message passing and explicit attention variance regularization; (iii) We engineer an infrastructural NCCL boundary exchange protocol that hoists static partition indices and leverages fixed-shape GPU-to-GPU collectives, bypassing CPU serialization and memory copy overheads. Extensive evaluations across five long-term agent memory and multi-hop reasoning benchmarks demonstrate the remarkable performance of WFM, while achieving a 10.5 times training acceleration on distributed clusters.

阅读 arXiv 原文

人机协同与对齐(0 篇)

本轮该赛道没有候选论文。