公开论文雷达

公开 arXiv 研究简报 · 2026-09-30T00:57:09.575514+00:00

先看结论和关键数字,再决定要不要读原文。候选只在首次出现时展示;旧候选若后来通过深读门槛,仍会进入重点。

先读哪几篇:把"写了"和"真生效"分开

八篇里真正互相呼应的是四篇:方法论套件在仓库里落地率只有21.7%、Relic把规则绑到运行时后新成员正确率41.2%、WideSWE最佳配置只做完42.50%、Path2Spec指出至少34.6%通过验证的规约没覆盖真实分支。先按这个顺序读,剩下四篇按手头的活挑着看。

推荐阅读顺序

  1. 2609.32014:先看现状:5435个仓库里只有21.7%出现过至少一种机制,写了规则文件不等于整套方法论生效。
  2. 2609.32965:看完现状接着看对策:规则绑运行时,换人接手正确率41.2%,纯文本只有34.6%。
  3. 2609.33382:交付前先知道边界:120个跨仓库任务最佳配置只解决42.50%,69个失败里48个是半成品。
  4. 2609.33438:同一个坑换到规约层:通过验证的规约里至少34.6%没覆盖行为分支,按路径拆着生成。
  5. 2609.33796:打算自己训策略生成模型再读:把验证器的失败检查和反例喂给RL,别只给通过/失败。
  6. 2609.32049:要做长期对话的记忆层才需要:用量驱动图加替代边,矛盾幻觉从70%降到0%。
  7. 2609.30186:被调用成本和延迟卡住时读:成功率79%对基线84%,成本降73.4%,自己权衡这笔账。
  8. 2609.35281:放最后:只有手头真有UML生成的遗留代码才读,而且它没给数值、自称preliminary。
共性方法
这批卡片反复戳同一个点:写了规则、验证通过、整体测试全绿,都不等于事情真做到了。它们给的办法也一致——把检查绑到能执行的东西上,运行时协议、逐条路径的规约、按仓库粒度的fail-to-pass、验证器给出的具体失败原因,而不是留一份文档或一个总通过率。
关键分歧
证据分量差得很远。5435个仓库的分层抽样、360次受控运行、120个跨仓库任务,是能拿来下结论的;另一头,EngramRAG的矛盾测试是50个合成情节,RAISE的1408个机构场景出自单一合成机构,ARTHUR自建数据集只有几个项目、论文自称preliminary且连具体数值都没给。
选择准则
只有一小时:读32014和32965,再扫一遍33382。要动手改东西,再加读对应那篇——写规约看33438,训模型看33796,做记忆层看32049,压GUI成本看30186。35281没数值,只当思路参考,别当结论引用。

重点深读(8 / 8 篇)

形式化与程序验证(2 篇)

形式化与程序验证 7/30

RAISE: Reinforcing Access Control Policy Synthesis in LLMs via Symbolic Evaluation

符号验证诊断改进访问策略生成:CedarInstruct数据集(5800场景/44领域)+RAISE两阶段训练:SFT保证语法合法,RL用验证器失败诊断而非纯通过/失败信号,9B模型语义成功率超越前沿模型zero-shot效果。

两句看懂

前沿模型生成的Cedar策略合法率超97%但完整通过验证仅约三分之一,直接用pass/fail做RL时约三分之一样本全部采样失败导致无梯度信号。RAISE-OC改用验证器的失败检查项和反例引导探索,训练出的9B模型语义成功率反超GPT-6 Astra、Claude Opus 5共13.33和16.26个百分点。

核心判断

能否训练小模型生成语义正确的访问控制策略?能,但关键不在验证器信号量而在用法——RAISE-OC用失败检查与反例引导RL探索,9B模型语义成功率超越两个前沿模型13.33/16.26个百分点。

关键要点

1.旧假设与评测缺口:合法策略被默认可信,但GPT-6 Astra与Claude Opus 5合法率超97%,完整通过验证计划的仅约三分之一,说明语法正确不等于语义正确。 2.构建与协议:CedarInstruct含5800场景/44领域(其中1408个单一机构场景),每条配需求、schema、验证目标策略与可执行验证计划;训练分两阶段,先用验证过的策略做SFT,再用约5.4K场景+LoRA做RL,对比六种消费不同粒度验证器信号(二元通过/失败到失败检查描述)的RL方案。 3.决定性结果与失败诊断:纯pass/fail奖励下RL阶段约三分之一输入的所有采样策略均失败,梯度信号缺失;RAISE-OC把失败检查与符号反例转成引导式探索并用off-context GRPO学习,是唯一显著优于SFT的方案,9B模型语义成功率超越GPT-6 Astra 13.33个百分点、Claude Opus 5 16.26个百分点,并迁移至CedarBench。

证据与结果

CedarInstruct共5800个场景,44个领域,其中1408个属于单一合成机构;每条含自然语言需求、Cedar schema、验证目标策略及可执行验证计划。在held-out场景比较语义成功率:GPT-6 Astra、Claude Opus 5合法率超97%但仅约1/3通过完整验证计划;RAISE-OC训练的Qwen3.5-9B(约5.4K验证场景+LoRA)分别超出13.33、16.26个百分点,并迁移到独立构建的CedarBench;诊断发现RL阶段约1/3输入的全部采样策略验证失败,纯二元奖励无信号。

打开论文原文
它要解决什么
能否用符号验证信号训练小模型直接生成语义正确的访问控制策略,替代部署时反复调用验证器修复?
研究路径
先用CedarInstruct中验证过的目标策略对模型做SFT,使输出从多数不合法变为几乎全部合法;再进入RL阶段,对每个失败案例用验证器给出的违反属性与具体反例请求构造引导上下文,采样探索后用off-context GRPO更新策略,而非直接用通过/失败作奖励。
这对工程意味着什么
训练策略生成模型应把验证器的失败原因与反例喂给RL而非只给通过/失败标签,否则约三分之一样本会全失败而无梯度;只提高合法率(SFT)不能替代语义正确性检验。
证据定位
GPT-6 Astra、Claude Opus 5合法策略率超97%但仅约1/3通过完整验证计划;RAISE-OC训练的Qwen3.5-9B在held-out场景语义成功率超过两者13.33和16.26个百分点,并迁移到CedarBench。(筛选维度:形式化验证、可复核评测)
适用边界
1408个机构场景来自单一合成机构,44个领域是否代表真实企业策略分布未说明;六种RL方案对比基于同一验证器与同一LoRA/9B模型规模,能否泛化到其他模型规模未验证。
方法与英文摘要

构建CedarInstruct:5800个场景/44领域,含1408个单一合成机构场景,每条配验证目标策略与可执行验证计划。两阶段训练:先用验证过的目标策略做SFT,再用约5.4K验证场景+LoRA做RL,比较六种消费不同粒度验证器信号的RL方案。

Translating natural-language access-control requirements into policies requires careful reasoning about permissions, constraints, and exceptions, and even frontier LLMs often produce policies that violate the intended authorization semantics. We construct CedarInstruct, to our knowledge the first dataset that supports both training and semantic evaluation for formally verifiable Cedar policy synthesis. It contains 5,800 scenarios across 44 domains and 1,408 representing a single synthetic organization, each with a verified target policy and an executable verification plan. On this data we introduce RAISE, which trains policy synthesizers from formal verification in two stages, verified supervised fine-tuning (SFT) followed by a reinforcement learning (RL) stage that learns from verifier signal. We find that SFT succeeds largely by letting models express authorization logic they already have, since untrained models rarely write valid Cedar but often reason correctly when they do. After SFT, how the verifier's information is used matters more than how much of it is used. Of six RL instantiations that consume progressively richer verifier signal, only RAISE-OC improves meaningfully on SFT; it turns failed checks and symbolic counterexamples into guided exploration and learns from the result with off-context GRPO. With about 5.4K verified scenarios and LoRA fine-tuning, RAISE-OC trains Qwen3.5-9B to surpass zero-shot GPT-6 Astra and Claude Opus 5 by 13.33 and 16.26 percentage points in semantic success on held-out scenarios, and training transfers to the independently constructed CedarBench.

形式化与程序验证 7/30

Path2Spec: Path-Aware Specification Generation via Large Language Models

验证通过不代表规约合格:按路径拆分生成,通过率从44.2%提到83.0%:如果你用LLM生成形式化规约,只看"验证通过"会漏掉大问题:至少34.6%通过验证的规约其实没覆盖程序的真实行为分支。Path2Spec的做法是把程序按执行路径拆开,逐条路径生成规约再合并,在SG-Bench和SV-COMP上验证通过率达87.5%和83.0%,明显高于基线的66.7%和44.2%。

两句看懂

现有LLM规约生成把程序当整体处理,导致约束过粗,至少34.6%通过验证的规约没覆盖真实行为分支。Path2Spec按路径生成规约再合并,在SG-Bench和SV-COMP上通过率达87.5%和83.0%,远超基线的66.7%和44.2%。

核心判断

验证通过不等于规约质量达标:至少34.6%通过验证的规约未覆盖程序真实行为分支,而按路径拆分生成的Path2Spec把通过率从66.7%/44.2%提升到87.5%/83.0%。

关键要点

1.旧假设失效:验证器只查前置/后置条件是否满足,不查规约是否覆盖全部执行路径,SpecGen通过验证的规约中至少34.6%未捕捉程序真实行为。 2.方法与受控检验:Path2Spec逐路径生成规约再合并,失败则按逻辑分支递归拆解重试(decompose-then-retry),在SG-Bench(120个程序)和SV-COMP(265个程序)上与SpecGen对比并加人工评审。 3.决定性结果:通过率87.5%/83.0%对66.7%/44.2%,人工评审确认语义对齐更精确;实际使用时应按路径拆分生成,而不是只信验证通过率。

证据与结果

基准是SG-Bench(120个程序)和SV-COMP(265个程序),主指标为验证通过率,辅以人工评审规约与代码的语义对齐度。对比基线SpecGen:SG-Bench上87.5% vs 66.7%,SV-COMP上83.0% vs 44.2%。另有分析指出,SpecGen式整体生成中至少34.6%通过验证的规约未覆盖程序真实行为分支,标准验证器无法检测,要靠Logical Strength类指标或人工评审发现。

打开论文原文
它要解决什么
LLM生成的形式化规约,验证通过是否等于它真正覆盖了程序的所有执行行为?
研究路径
第一步,LLM从输入程序抽取全部执行路径。第二步,对每条路径单独生成前置/后置条件规约。第三步,把各路径规约合并成整体规约。第四步,若路径生成失败,按逻辑分支把程序递归拆成子程序,分别生成规约后再合并回原程序(decompose-then-retry)。
这对工程意味着什么
第一步行动:生成程序规约时先枚举执行路径,逐路径生成再合并,不要把程序当整体一次性生成。要避免的捷径:只看验证通过率,它检测不出至少34.6%规约未覆盖行为分支的问题。
证据定位
SG-Bench上验证通过率87.5%,基线SpecGen为66.7%;SV-COMP上83.0%,基线为44.2%。人工评审确认Path2Spec的规约与代码语义对齐更精确。另一项分析发现,SpecGen式整体生成方法中至少34.6%通过验证的规约未覆盖程序真实行为分支,标准验证器检测不出这类问题,需要Logical Strength类指标或人工评审才能发现。(筛选维度:形式化验证、可复核评测)
适用边界
评测只覆盖SG-Bench(120个程序)和SV-COMP(265个程序)两个公开基准;34.6%的覆盖不足比例来自对所分析基线方法样本的统计,具体抽样细节摘录中未给出。
方法与英文摘要

Path2Spec的流程是:LLM先抽取程序的全部执行路径,为每条路径单独生成前置/后置条件规约,再合并成整体规约。对路径生成失败的复杂程序,按逻辑分支递归拆解成更小的子程序,分别生成规约后合并回原程序,即decompose-then-retry策略。评测用SG-Bench(120个程序)和SV-COMP(265个程序)两个基准,对比基线SpecGen。

Formal specifications are critical for program verification, comprehension, and maintenance. However, manually writing them is costly and difficult to scale. Recent studies have shown that Large Language Models (LLMs) are promising for automated specification generation, but existing methods suffer from quality issues. We analyze a state-of-the-art approach and find that at least 34.6% of successfully verified specifications actually fail to meaningfully capture the program's distinct behavior, which is a quality issue not captured by metrics that only measure verification success. We further found that a major factor contributing to such hidden quality issues stems from the design of existing methods: these methods treat a program as a single unit, resulting in overly general, coarse-grained constraints. To this end, we introduce Path2Spec, a divide-and-conquer framework that addresses these limitations through systematic path-based reasoning. Path2Spec leverages LLMs to extract all execution paths from an input program, generates path-specific specifications for each, and merges them into a comprehensive overall specification. For complex programs where path-based generation struggles, Path2Spec employs a decompose-then-retry strategy that recursively breaks a program into smaller subprograms based on logical branches, generates specifications for each, and merges them back. We evaluate Path2Spec on two public benchmarks: SG-Bench (120 programs) and SV-COMP (265 programs). Results show that Path2Spec can outperform the state-of-the-art baseline SpecGen: 87.5% versus 66.7% on SG-Bench, and 83.0% versus 44.2% on SV-COMP. Human evaluation further validates that Path2Spec generates higher-quality specifications with precise semantic alignment to the code.

软件工程与仓库智能(3 篇)

软件工程与仓库智能 10/30

Relic: From Multi-Agent Collaboration to Persistent Organizational Capability

换人后协作约定不失效:把规则绑定运行时,新成员正确率41.2%对纯文本34.6%:多智能体团队靠对话达成的协作规则,一换成员就失效,交付质量随之下降。Relic的做法是把反复出现的协作摩擦固化成绑定运行时的可执行协议:360次受控运行中完整契约交付率从14.06%升至19.76%,新成员在可执行协议下正确率41.2%,高于纯文本规则的34.6%和无协议的25.4%。

两句看懂

多智能体团队靠对话达成的协作约定,换掉参与者后往往失效;Relic把反复出现的协作摩擦固化为绑定运行时的可执行协议。360次受控运行显示完整契约交付率从14.06%升至19.76%,新成员在可执行协议下正确率41.2%,高于纯文本规则的34.6%。

核心判断

换人后协议能否继续起作用?能——把规则绑定到运行时(触发/角色/证据/后果)比纯文本更有效:新成员正确率41.2%对34.6%,完整契约交付率提升5.71个百分点(14.06%→19.76%)。

关键要点

1. 旧假设失效:对话/群聊式协调只在当前成员间建立约定,成员更替后约定不自动生效;共享技能库保留了代码,但不指派独立评审人、不要求集成前提供最新证据,接口变更类故障反复出现。 2. 方法与对照:Relic把摩擦转成经验证批准、绑定运行时的协议(触发条件+责任角色+证据+后果),可修订可废止;360次受控运行覆盖10个任务、3种模型,新成员接手设无协议/纯文本/可执行绑定三档对照。 3. 结果与行动:交付率14.06%→19.76%(+5.71pp),新成员正确率41.2%对纯文本34.6%、无协议25.4%,CooperBench固定子集29/48胜Solo 26/48——应把协作教训写成运行时协议,而不是留文档给新人读。

证据与结果

10个软件协作工作负载、3种模型、360次受控运行,对比有/无协议生命周期的结构化团队。新成员接手三档对照:无协议25.4%、纯文本协议34.6%、可执行协议绑定41.2%。CooperBench评测:剔除失效题对后全量477对中达成367对(76.9%);固定48对同模型子集Relic 29/48对Solo 26/48,反转了官方对等基线中出现的协作损失。

打开论文原文
它要解决什么
对话达成的协作共识,能否在成员更替后继续约束团队,而不是靠新人重新学一遍?
研究路径
成员观察共享工作中的摩擦,提出规则草案;团队验证并批准后,协议绑定到运行时,固定四要素:触发条件、责任角色、所需证据(如批准评审+最新测试)、执行后果。协议保持可修订、可废止。新成员加入时按同一协议接手职责,规则约束的是运行时行为,不依赖新人读过历史对话或继承作者私有记忆。
这对工程意味着什么
第一步行动:挑一个团队里反复出现的协作故障(如接口变更导致客户端过时),把它写成一条带触发条件、责任人、所需证据和后果的协议,绑定到运行时执行。要避开的捷径:不要把规则只写成文本文档留给新人阅读——本研究中纯文本只有34.6%,比可执行绑定的41.2%低6.5个百分点,替代不了运行时绑定。
证据定位
完整契约交付率14.06%→19.76%(+5.71个百分点),在每个模型层级都改善全部四个已验证生产端点。新成员接手:可执行协议正确率41.2%,比纯文本34.6%高6.5个百分点,远高于无协议25.4%。CooperBench剔除失效题对后全量477对中达367对(76.9%);固定48对同模型子集Relic 29/48胜过Solo 26/48,反转官方对等基线出现的协作损失。(筛选维度:形式化验证、可复核评测、软件工程方法)
适用边界
CooperBench评测排除了失效题对后再统计;固定同模型子集仅48对;核心追踪案例(接口评审规则)是单一示例,用于说明机制,不作为统计证据。
方法与英文摘要

Relic让智能体基于可见的工作摩擦提出规则草案,团队验证批准后把协议绑定到运行时,固定触发条件、责任角色、所需证据与执行后果,协议可修订可废止。评测覆盖10个软件协作任务、3种模型、360次受控运行,对比有/无协议生命周期的结构化团队;另设新成员接手三档对照:无协议、纯文本协议、可执行协议绑定;并在CooperBench剔除失效题对后的477对全量及固定48对同模型子集上评测。

Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the lesson continue to govern the team? We introduce Relic, which turns recurring collaboration failures into organization-owned, executable protocols. Members reflect on visible work, propose rules, and govern their adoption. Adopted protocols bind triggers, responsibilities, required evidence, and execution consequences to the runtime, while remaining open to revision and retirement. In one traced case, repeated integration friction produces an interface-review rule that governs later pull requests and is revised as work continues. Across 360 controlled runs over ten software workloads and three models, Relic raises complete-contract delivery from 14.06% to 19.76% (+5.71 percentage points) over a matched structured team without the protocol lifecycle, improving all four verified production endpoints in every model stratum. Under fresh-member transfer, behavioral correctness is 25.4% with no inherited protocol, 34.6% with the same rules provided as readable text, and 41.2% with executable bindings, a +6.5-point advantage over text alone. On the full CooperBench benchmark, after excluding broken benchmark pairs, Relic achieves 367/477 (76.9%), establishing the best reported result among peer-structured systems. On the fixed 48-pair same-model subset, Relic also exceeds Solo (29/48 vs. 26/48), reversing the coordination loss exhibited by the official peer baseline. Together, these results show how collaboration experience can become persistent organizational state that remains useful beyond the members who created it.

软件工程与仓库智能 10/30

Methodological Harness in Agentic Software Engineering: An Empirical Study on Mining Software Repositories

智能体方法论套件落地率仅21.7%,多数是零散单点机制:如果你团队在用AI智能体写代码,这项研究说明:别指望写一份规则文件就等于用上了整套方法论——对5435个GitHub仓库的检测显示,仅21.7%的仓库出现至少一种机制,高星仓库为65.3%,但多机制协同配置罕见。研究用分层抽样加自动产物检测和人工编码验证了这一点。

两句看懂

此前'方法论套件'八机制框架仅基于灰色文献构建,未在真实仓库中验证是否落地。研究对5435个GitHub仓库分层抽样检测机制产物,并对150个规则文件定性编码,发现总体21.7%仓库出现至少一种机制、高星仓库达65.3%,但多机制协同配置罕见。

核心判断

方法论套件在真实仓库中部分落地但很零散:21.7%的仓库出现至少一种机制(高星仓库65.3%),多机制整合配置罕见,规则文件内容多止步于指令与规范。

关键要点

1. 旧框架的问题:'方法论套件'八机制(情境工程、持久共享知识、可执行规格等)只基于灰色文献,未经同行评审的仓库级检验,是否被采用完全未知。 2. 方法与核查:从AIDev的116211个仓库按星标分层抽样5435个,自动扫描规则文件、规格文档、ADR等七类机制产物并记录引入时间,再对150个规则文件人工编码核对归类。 3. 结果与做法:总体21.7%仓库出现至少一种机制,高星仓库65.3%,但多机制组合罕见;规则文件几乎只含指令与规范,因此采用智能体开发时应逐项核查八机制,而非默认已成系统。

证据与结果

数据来自AIDev数据集的116211个GitHub仓库,分层抽样5435个做RQ1产物检测,另抽150个规则文件做RQ2定性编码。指标包括机制普及率、机制广度、共现模式和时间演变。结果:总体普及率21.7%,高可见度仓库65.3%;多机制组合罕见;规则文件几乎都含'引导智能体'与'声明规范',而持久共享知识、可执行规格、结构化会商、分级自主权仅在少数文件中出现。

打开论文原文
它要解决什么
此前提出的'智能体软件工程方法论套件'八机制框架只是概念提议:真实GitHub仓库到底有没有采用这些机制?以什么方式落地?
研究路径
第一步,从AIDev的116211个仓库按星标可见度分层抽样5435个。第二步,自动扫描仓库中的规则/上下文文件、规格文档、ADR等产物,记录各机制首次出现日期,计算普及率、每仓库机制数、机制共现和随时间变化趋势。第三步,另抽150个仓库的规则文件,由人工定性编码,判断内容对应八机制中的哪几项。
这对工程意味着什么
第一个行动:引入智能体协作规范时,逐项核查八机制——尤其持久共享知识、可执行规格、结构化会商、分级自主权——是否真正落地。要避开的捷径:不要以为写一份规则文件就等于整套方法论生效,多数仓库的规则文件只做到指令与规范两项。
证据定位
总体仓库中至少出现一种机制的比例为21.7%,高可见度仓库升至65.3%;但多机制组合配置在两类仓库中都很罕见。已有规则文件的仓库几乎只覆盖'引导智能体'和'声明规范'两类内容;持久共享知识、可执行规格、结构化会商、分级自主权四项机制仅在少数规则文件中出现。(筛选维度:形式化验证、可复核评测、软件工程方法)
适用边界
样本仅取自AIDev数据集中标记有智能体活动的GitHub仓库,按星标分层抽样;规则文件定性编码只覆盖150个仓库。研究只衡量套件是否被采用,不评估它是否提升团队实际效能。
方法与英文摘要

研究从AIDev数据集(116211个GitHub仓库)按星标可见度分层抽样5435个仓库。RQ1用自动化脚本扫描七类可观测机制的产物——规则/上下文文件、规格说明、ADR架构决策记录——并记录首次引入时间,统计普及率、机制广度、共现模式和时间演变。RQ2另抽150个仓库的规则文件,人工定性编码,核对内容对应八机制中的哪几项。

Context. Agentic software engineering requires mechanisms to coordinate and govern agents' work. The framework motivating this study proposes a methodological harness with eight mechanisms: context engineering, persistent shared knowledge, executable specifications, N-version mindset and parallel agents, normative specifications, structured consultation, evidence-based acceptance, and graduated autonomy. These are realized through artifacts such as rule or context files, specifications, and architectural decision records, but the framework had not been empirically validated at repository scale. Objective. We analyze to what extent and how this harness is observable in repositories with agentic activity. RQ1 characterizes adoption through artifact prevalence, breadth, co-occurrence, and temporal evolution; RQ2 examines rule files - a key observable artifact and persistent source of agent instructions - to assess how their content reflects the proposed mechanisms. Method. From the AIDev dataset of 116,211 GitHub repositories, we analyzed 5,435 using a design stratified by visibility, measured through stars as an indicator of popularity and/or reputation. For RQ1, we detected and quantified artifacts and introduction dates for the seven observable mechanisms. For RQ2, we qualitatively coded rule files from 150 repositories. Results. Population prevalence of at least one mechanism is 21.7%, versus 65.3% among the most visible repositories. Multi-mechanism configurations are rare. Where rule files exist, they almost always guide the agent and state norms, while persistent shared knowledge, executable specifications, structured consultation, and graduated autonomy appear only in a minority. Conclusions. The harness is empirically observable, but mainly through isolated mechanisms rather than the integrated system proposed by the framework.

软件工程与仓库智能 8/30

WideSWE: Can Coding Agents Coordinate Changes Across Repositories?

编码智能体做跨仓库协同修改:最佳配置只完成42.50%:如果你打算让编码智能体改跨多个仓库的需求,这份测试提醒你先别信任它:在120个真实跨仓库任务(60缺陷修复+60功能)上,最好的配置也只解决42.50%,而且近七成失败案例是部分仓库改完、其余没改的半成品。

两句看懂

单仓库评测无法暴露编码智能体在跨仓库协同任务中的缺陷,WideSWE用120个真实跨仓库任务(60缺陷+60功能)重新检验了这一问题。测试7种配置后,最佳的Codex CLI+GPT-5.6-sol仅解决42.50%的任务,69个失败案例中48个存在部分仓库通过F2P测试而其余未通过的情况。

核心判断

编码智能体尚不能可靠地把同一需求贯穿到所有相关仓库:最佳配置在120个跨仓库任务上仅解决42.50%,近七成失败案例处于部分仓库通过、部分未通过的半成品状态。

关键要点

1. 旧评测假设的问题:现有基准(SWE-bench、DeepSWE、ProgramBench等)主要在单仓库内评估,但生态系统中大量修改需跨多仓库协同(172.9万条PR中10.9万条明确引用同生态其他仓库)。 2. 方法与受控检查:从103个生态系统构建120个任务(60缺陷+60功能),改写隐藏测试兼容多种实现并保留回归检查,对比联合执行与逐仓库独立执行两种协议。 3. 决定性结果与行动:最佳配置仅解决42.50%(区间10.83%-42.50%),69个失败案例中48个是半成品;工程上应按仓库粒度校验测试并对比两种执行方式。

证据与结果

以隐藏fail-to-pass测试判定任务是否完全解决。7种配置成功率10.83%-42.50%,最佳的Codex CLI+GPT-5.6-sol为42.50%;69个失败案例中48个是部分仓库通过F2P测试、部分未通过。89个同提示任务上,联合执行解出36个、独立执行解出32个,20个结果翻转:独立执行更擅长补齐遗漏改动,但更难纠正已尝试未成功的实现;联合执行能借助关联仓库信息指导实现与验证。

打开论文原文
它要解决什么
编码智能体能否识别并完成跨越多个代码仓库的协同修改请求,而不是只在单个仓库里把问题修掉?
研究路径
构造流程:1)从200个高star GitHub组织中选103个生态系统;2)挖掘关联PR并人工复核,筛出需多仓库实质修改的任务;3)合并issue与PR需求生成提示词;4)改写隐藏测试以兼容多种正确实现并保留回归检查;5)运行7种智能体配置,记录F2P测试通过情况;6)对89个任务对比联合执行与独立执行,统计结果翻转数。
这对工程意味着什么
第一步行动:评估跨仓库改动任务时,逐仓库检查测试通过情况,而不是只看整体是否成功。要避免的捷径:仅凭单仓库测试全部通过就判定任务完成——这会漏掉半成品修改。
证据定位
Codex CLI+GPT-5.6-sol表现最佳,但也只解决42.50%的任务(7种配置整体在10.83%-42.50%之间);69个失败案例中,48个是部分仓库通过F2P测试、其余未通过;在89个同提示对照任务上,联合执行解出36个、独立执行解出32个,20个任务的结果发生翻转。(筛选维度:可复核评测、软件工程方法)
适用边界
任务来源限定于star数最高的200个GitHub组织中的103个生态系统,样本聚焦公开、活跃且高关注度的项目,对小众或私有生态系统的代表性文中未说明。
方法与英文摘要

从200个高star GitHub组织中筛出103个活跃生态系统,挖掘172.9万条PR,经人工复核加可执行验证得到120个任务(60缺陷+60功能);提示词融合原始issue与PR需求,隐藏测试被改写以兼容多种正确实现并保留回归检查;测试7种智能体配置,并对比联合执行与逐仓库独立执行两种方式。

Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multiple repositories. We introduce WideSWE to evaluate coding agents on such cross-repository tasks. Mining and reviewing changes across 103 software ecosystems yields 120 real-world tasks, balanced between 60 bug fixes and 60 features. We derive prompts from related issues and pull requests. We systematically review and adapt hidden tests to support diverse correct implementations while preserving required behavior and regression checks. Across seven agent configurations, full task success ranges from 10.83% to 42.50%, with the configuration pairing Codex CLI with GPT-5.6-sol achieving the highest rate. Trajectories show agents failing to identify necessary changes, recognizing changes but leaving them unfinished, or modifying the required repositories without fully satisfying the request. To examine whether working on one repository at a time can alleviate these difficulties, we compare it with joint execution under identical prompts. Independent execution mainly recovers omitted work and is less effective at correcting previously attempted but unsuccessful implementations. Joint execution can use information from related repositories to guide implementation and verification. Code is available at https://github.com/ZJU-ACES-ISE/WideSWE.

代码质量与优化(1 篇)

代码质量与优化 6/30

Is there a future for models in the LLM era?

LLM时代MDE没过时:先重构再校验,优于让LLM从零生成:如果你的系统里有传统MDE工具从UML生成的遗留Java代码,先别急着扔掉模型规格。ARTHUR框架用多智能体LLM重构这些代码以支持Spring Boot等现代栈,再用Model-Based Testing校验它仍符合原模型,初步对比显示这条路优于直接生成。

两句看懂

传统MDE生成的遗留Java代码难以支持Spring Boot等现代框架,ARTHUR用多智能体LLM对其重构,并保留原有模型约束。作者在自建多项目数据集上以compile@k、pass@k、pass^k及通过率、时间成本对比重构版与直接生成版,初步结果显示MDE路线仍具优势。

核心判断

LLM智能体能够改进MDE生成的遗留代码:ARTHUR重构版在自建数据集上,按compile@k/pass@k/pass^k及通过率、时间成本评测,初步优于直接从模型生成的版本,但具体数值未披露。

关键要点

1. 旧假设:业界认为MDE复杂度高、已被脚本式开发取代,此前少有工作把LLM智能体与UML/MDE代码生成严格结合并做测试验证。2. 方法与校验:ARTHUR输入需求、UML图和MDE遗留代码,多智能体重构以支持Spring Boot/React,并用Model-Based Testing校验仍符合原模型;对照组为直接从概念模型生成。3. 结果与行动:初步结论是重构版优于直接生成('MDE远未被淘汰'),但无具体数值且为preliminary;工程上先用LLM重构+模型驱动测试,而非让LLM从零生成。

证据与结果

数据为作者自建,含数个项目(具体数量未披露),遗留代码由传统MDE工具(如Merlin Prototyper)从UML图生成;无公开tier/split。评测两种模式:MDE生成+ARTHUR重构 vs. 直接从概念模型生成。指标为时间与成本、测试通过率、compile@k、pass@k、pass^k。摘录仅给出定性结论'MDE远未被淘汰',未列出具体数值,论文自称结果为preliminary。

打开论文原文
它要解决什么
LLM智能体能否改善由UML图与文本需求经MDE工具生成的Java遗留代码?
研究路径
1.收集文本需求、UML图及传统MDE工具生成的Java遗留代码;2.ARTHUR多智能体据此重构代码,加入Spring Boot、React等现代框架支持;3.用Model-Based Testing核对重构代码是否仍满足原始UML模型规格;4.分别测量重构版与直接生成版的时间成本、通过率、compile@k、pass@k、pass^k。
这对工程意味着什么
第一步行动:改造遗留MDE代码时,先用LLM多智能体重构,再用Model-Based Testing校验是否仍符合原始UML模型,而不是直接丢弃模型。要避开的捷径:让LLM从零生成——它看似快,但可能丢失原始约束;在本文未给出可验证具体数值前,也不要直接采信'重构更优'的结论。
证据定位
论文用compile@k、pass@k、pass^k及通过率、时间成本对比ARTHUR重构版与直接生成版,给出定性结论'MDE远未被淘汰',即重构路线更优;但摘录未提供任何具体数值,且论文自称结果为preliminary。(筛选维度:形式化验证、可复核评测)
适用边界
数据集为作者自建、项目数量少('several projects'),未公开具体规模;论文自称preliminary,摘录未给出compile@k/pass@k具体数值,结论可推广性有限。
方法与英文摘要

作者自建数据集,含数个项目的、由传统MDE工具(如Merlin Prototyper)从UML图生成的Java遗留代码。输入文本需求、UML图和该遗留代码,ARTHUR多智能体重构代码以支持Spring Boot/React,并用Model-Based Testing校验重构后代码是否仍满足原始模型规格。对照组为直接从概念模型生成代码、不经重构的版本。

In an era where software development is deeply tied with Large Language Models, does Model Driven Engineering (MDE) still make sense? This raises the question of the extent to which MDE can be successfully combined with an LLM approach to address the downsides of each approach separately. In this paper, we try to answer that question by investigating a central Research Question: Can Agents powered by Large Language Models improve code generated from UML diagrams and text specifications by Model Driven Engineering (MDE) tools? To investigate this problem, we developed a novel LLM-powered Multi-Agentic approach called ARTHUR (Architecture Refactoring Through Hybrid UML Reasoning), a framework that aims at combining the reliability of MDE and the ease of use of LLMs. ARTHUR is designed to refactor legacy Java code produced by traditional, rule-based, MDE code generators from UML diagrams. To asses the answer to our question, we have refactored the legacy code of several projects from our custom dataset crafted for this purpose. \name{} which made it possible to add support for modern frameworks like Spring Boot, while ensuring compliance with Model-Based Testing techniques to verify that the code still corresponds to the initial model's specifications. We then measured the results obtained in terms of time and cost, passing test rate, and \texttt{compile@k}, \texttt{pass@k} and \texttt{pass$^k$} metrics. We also observed the effect of generating code directly from the conceptual model without the refactoring. Our preliminary test results show that MDE could not be more far from retirement, after all.

UI 与 GUI Agent(1 篇)

UI 与 GUI Agent 5/30

Jev-Mobile: Jev as an Executor for Mobile GUI Agents

移动GUI智能体:VLM低频规划+轻量执行降本增效:移动GUI智能体常需VLM逐步规划并定位动作,延迟高、成本贵。Jev-Mobile由VLM设定局部目标,轻量模型Jev在无障碍树动作空间中连续执行。AndroidWorld测试成功率79%,成功轨迹执行时间降32.7%,API成本降73.4%。

两句看懂

多数移动GUI智能体依赖VLM在每一步同时做规划和动作定位,延迟高、调用成本大,Jev-Mobile改为VLM低频设定局部目标、轻量模型Jev连续执行无障碍树候选动作。AndroidWorld全量任务测试显示,其成功率79%接近逐步式VLM基线的84%,且成功轨迹执行时间降32.7%、模型API成本降73.4%。

核心判断

轻量决策模型可在无障碍树动作空间内接管多步执行,使VLM降为低频规划者;AndroidWorld全量测试显示成功率79%接近84%基线,而成功轨迹执行时间与API成本分别降32.7%和73.4%。

关键要点

1. 多数移动GUI智能体让VLM在几乎每一步同时做规划和动作定位,导致每步都要一次昂贵VLM推理,带来高延迟与高模型调用成本。 2. Jev-Mobile在AndroidWorld全量任务集上评测:VLM读取任务、截图/无障碍树与历史动作,只输出局部目标与精确文本;程序从原始无障碍树节点(可见、启用、边界有效)枚举候选动作(点击/长按/滚动/聚焦/输入/返回/主页/回车/开应用);Jev据候选反复选择动作直至DONE或BLOCKED;评测分别统计候选覆盖率、选择正确性、交还控制时机,而不仅看最终成功率;对照组为SeeAct-V(UI-TARS定位)与逐步式VLM基线。 3. 全量任务成功率:Jev-Mobile 79%、SeeAct-V 78%、逐步式VLM基线84%;成功轨迹中平均端到端执行时间降32.7%,平均模型API成本降73.4%;失败诊断:截图出现无障碍树未暴露的图标或所需控件缺少标志位时,Jev只能返回BLOCKED,无法从图像直接合成坐标动作,这是当前确定性候选构造方式的边界。

证据与结果

评测基于AndroidWorld全量任务套件,由独立评测器在任务结束后判定终端设备状态是否成功。对照组:SeeAct-V(UI-TARS视觉定位)、逐步式VLM基线,三者共用同一VLM但执行机制不同。指标包括任务成功率、成功轨迹的平均端到端执行时间、平均模型API成本,并单独统计候选覆盖率、选择正确性与控制交还时机。结果:成功率Jev-Mobile 79%、SeeAct-V 78%、逐步式VLM基线84%;相对逐步式VLM,成功轨迹执行时间降32.7%,API成本降73.4%。失败诊断:截图含无障碍树未暴露图标或必要控件缺少标志位时,Jev返回BLOCKED,无法从图像合成坐标动作。

打开论文原文
它要解决什么
能否让轻量决策模型接管大部分GUI动作执行,减少VLM逐步调用同时保持任务成功率?
研究路径
VLM读取任务指令、当前截图/无障碍树及分组历史动作,输出局部目标与精确文本,不生成委托后摘要;程序据当前无障碍树枚举候选(可见、启用、边界有效的节点,支持点击/长按/滚动/聚焦/输入/返回/主页/回车/开应用);Jev仅凭候选描述与局部目标反复选择动作ID或返回DONE/BLOCKED;每次动作后重新观察并重建候选集,直至局部目标完成或阻塞交还VLM。
这对工程意味着什么
可将高频动作选择交给基于无障碍树候选的轻量模型执行,由VLM只负责低频目标规划,从而降低推理延迟与API成本;但不可想当然省去失败兜底——截图信息缺失于无障碍树时仍需VLM介入。
证据定位
AndroidWorld全量任务:Jev-Mobile成功率79%,SeeAct-V 78%,逐步式VLM基线84%;在成功轨迹中,相对逐步式VLM,平均端到端执行时间降32.7%,平均模型API成本降73.4%。(筛选维度:GUI Agent 方法)
适用边界
候选动作仅由确定性程序从原始无障碍树节点生成,不推断缺失的视觉坐标、不修复界面缺陷;节点缺失标志位则保持未知,不可见/禁用/边界无效节点不产生候选;当前输入工具拒绝非ASCII文本;对照系统各自执行机制不同,结果反映整体系统而非Jev的独立效应。
方法与英文摘要

基于AndroidWorld全量任务集;VLM读取任务、当前截图/无障碍树及历史动作,输出局部目标与精确文本;程序从原始无障碍树节点(可见、启用、边界有效)枚举候选动作(点击/长按/滚动/聚焦/输入/返回/主页/回车/开应用);Jev据候选描述连续选择动作直至DONE或BLOCKED;对照组为SeeAct-V(UI-TARS定位)与逐步式VLM基线。

Vision-language models (VLMs) have become a common foundation for autonomous mobile GUI agents, but most existing systems rely on the VLM for both planning and action grounding at nearly every interaction step, leading to substantial latency and model-serving cost. We introduce Jev-Mobile, which shifts this paradigm to low-frequency VLM planning and high-frequency lightweight execution: the VLM specifies local goals, the accessibility tree defines a structured executable action space, and Jev, a fast typed decision model, repeatedly selects actions within this space. This design allows multiple GUI actions to be executed under a single VLM decision, reducing expensive VLM inference while preserving adaptive interaction. On the full AndroidWorld task suite, Jev-Mobile achieves 79% task success, compared with 78% for SeeAct-V and 84% for a Step-wise VLM baseline. Among successful trajectories, it reduces mean end-to-end execution time by 32.7% and mean model API cost by 73.4% relative to Step-wise VLM. These results show that decoupling high-level VLM reasoning from low-level action execution can substantially improve mobile GUI agent efficiency while maintaining competitive task performance.

个人知识与本体(1 篇)

个人知识与本体 6/30

EngramRAG: Dynamic Usage-Weighted Topology and Synaptic Consolidation for Multi-Hop Agentic Memory

智能体记忆新架构:用量驱动图谱防幻觉:多会话智能体记忆易多跳检索失败、身份被衰减抹除、图谱不随使用更新。EngramRAG用用量驱动拓扑与巩固衰减机制,在LoCoMo上使Recall@5从38.29%提至53.21%,矛盾抑制从70%降至0%。

两句看懂

传统向量或静态图记忆难做多跳关联、易被时间衰减抹除人设、拓扑不随使用调整,EngramRAG改用可塑性拓扑与分层巩固衰减。经LoCoMo共1982条QA验证,Recall@5从38.29%升至53.21%,矛盾抑制从70.0%降到0.0%,90天模拟人设保留率达100.0%。

核心判断

用量驱动的可塑性拓扑与分层巩固衰减能替代静态图与时钟衰减做长期智能体记忆:LoCoMo上Recall@5达53.21%对比静态图8.50%,知识矛盾从70%降到0%。

关键要点

1. 稠密向量RAG只按语义相似度检索,跨会话多跳关联(Associative Blindness)会丢失;静态知识图谱把结构当作固定不变,不随实际检索频率调整边权重(Static Topology Stagnation),传统时钟衰减还会抹除长期人设(Scaffolding Amnesia)。 2. 在LoCoMo基准全部10段长期对话、1982条QA上评测:U-PPR用共激活的Hebbian塑性把持久实体推成高中心度的认知枢纽,CATD按拓扑承重而非时间设定半衰期并设N_grace≥4的冷启动宽限,SUPERSEDES有向图过滤被替代事实,三路检索(稠密向量+BM25+U-PPR)经动态RRF融合;另设50个合成情节的知识矛盾测试和90天纵向模拟。 3. Recall@5从稠密向量的38.29%、BM25的48.66%、静态图谱的8.50%提升到53.21%(p<0.001),MRR提升43.1%;矛盾测试中split-brain幻觉从70.0%降到0.0%;90天模拟里人设保留率达100.0%(基线衰减到60.0%),交互检索延迟仍为26.21ms。

证据与结果

数据:LoCoMo基准全部10段长期对话、1982条QA。指标Recall@5/MRR。EngramRAG Recall@5为53.21%,对比稠密向量38.29%(+38.9%,配对t=13.57,p<0.001)、BM25 48.66%(配对t=5.90)、静态图8.50%(超6倍);MRR 0.4203对0.2937(+43.1%)。时序推理子集Recall@5为62.33%(高于稠密向量16.67点、BM25 6.23点)。本地7B阅读模型下游生成:token F1持平(17.29%对18.34%),时序生成精度更优(18.18%对11.77%)。50个合成情节矛盾测试中,SUPERSEDES把split-brain幻觉从70.0%降到0.0%;90天模拟中CATD保留率100.0%对基线60.0%,交互延迟26.21ms。

打开论文原文
它要解决什么
多跳时序记忆能否用用量驱动拓扑替代静态图与时间衰减,同时不引入矛盾幻觉?
研究路径
Waking State实时接收对话,写入稠密向量、BM25索引与动态属性图;检索时按查询语义与U-PPR(共激活频率驱动的Hebbian边权)做个性化PageRank扩散,定位认知枢纽;三路结果经动态RRF按意图路由融合排序;后台Dreaming State异步执行CATD,按节点拓扑承重而非最近访问时间重算半衰期,新节点享N_grace≥4次访问的冷启动保护;事实更新时SUPERSEDES有向边标记旧节点失效并过滤其被检索。
这对工程意味着什么
给智能体记忆加图检索时,应按拓扑承重设定遗忘半衰期而非单纯按最近访问时间衰减,并用显式替代边隔离过期事实;仅靠纯向量Top-K或不区分新旧事实的静态图,容易在长期部署中丢人设或产生矛盾回答。
证据定位
LoCoMo全部1982条QA上,EngramRAG的Recall@5为53.21%,对比稠密向量RAG的38.29%(+38.9%,p<0.001)、BM25的48.66%、静态图RAG的8.50%;矛盾抑制测试中split-brain幻觉从70.0%降到0.0%。(筛选维度:形式化验证、可复核评测)
适用边界
评测限于LoCoMo基准10段对话共1982条QA,知识矛盾抑制测试为50个合成情节,下游生成用本地7B阅读模型,结果能否推广到更大规模真实部署或更大模型未在此说明。
方法与英文摘要

数据:LoCoMo基准全部10段长期多会话对话共1982条QA。架构分两态:低延迟Waking State实时检索,异步Dreaming State巩固。检索融合稠密向量、BM25与U-PPR(用量调制个性化PageRank,Hebbian塑性生成认知枢纽)三路,经动态RRF排序;CATD按拓扑承重而非时钟时间设定保留半衰期(冷启动宽限N_grace≥4);SUPERSEDES有向图过滤已被替代的旧事实。另设50个合成情节的知识矛盾测试与90天纵向模拟,并用本地7B阅读模型做下游生成评测。

As autonomous LLM agents are deployed across multi-session environments, conventional memory architectures suffer from Associative Blindness (inability to traverse multi-hop relational dependencies), Scaffolding Amnesia (temporal decay evicting core persona invariants), and Static Topology Stagnation (immutable graphs ignoring usage dynamics). Grounded in Complementary Learning Systems (CLS) principles, we propose EngramRAG, an adaptive memory architecture coupling a low-latency Waking State reflex with an asynchronous background Dreaming State consolidation cycle. EngramRAG introduces: (1) Usage-Modulated Personalized PageRank (U-PPR), where transition probabilities adapt via Hebbian plasticity to promote persistent entities into high-centrality Epistemic Macro-Hubs; (2) Consolidation-Activated Topology Decay (CATD), which scales retention half-life by topological load-bearing weight rather than wall-clock recency, protected by a cold-start grace period (N_grace >= 4); (3) Directed SUPERSEDES DAG filtering to suppress obsolete state during fact mutations; and (4) Triple-source hybrid retrieval fusing dense vectors, BM25, and U-PPR via dynamic Reciprocal Rank Fusion (RRF). Evaluating on all 1,982 QA pairs across 10 long-term conversations in the LoCoMo benchmark, EngramRAG achieves +38.9% relative improvement in Recall@5 (53.21% vs. 38.29%, p < 0.001) and +43.1% in MRR (0.4203 vs. 0.2937) over dense vector RAG, significantly outperforming Okapi BM25 (48.66%) and isolated static graph retrieval (8.50%). On temporal reasoning, EngramRAG reaches 62.33% Recall@5 (+16.67 points over dense vectors). In controlled mutation tests, SUPERSEDES suppresses split-brain hallucinations from 70.0% to 0.0%, while 90-day simulations show 100.0% scaffolding retention under a 26.21ms interactive retrieval reflex.

人机协同与对齐(0 篇)

本轮没有通过深读证据门的重点论文。

本轮分类概览

同一论文只归入一个最先命中的赛道,避免重复计数;“新增候选”只统计首次展示的论文。

赛道新增候选重点
形式化与程序验证82
软件工程与仓库智能43
代码质量与优化31
UI 与 GUI Agent01
个人知识与本体141
人机协同与对齐00

近一个季度监测日历

北京时间。绿色表示有可阅读的新候选,灰蓝表示已监测但无新增,橙色表示部分降级;“无记录”不等于失败。

2026 年 8 月

一二三四五六日

2026 年 9 月

一二三四五六日

近 14 次监测窗口

仅展示公开源的聚合运行状态,不含提示词、全文或个人数据。

本轮新增候选(29 篇)

按赛道、评分和日期展开;中文标签用于导航,英文摘要用于核验。已展示过的旧候选不会每日重复。

形式化与程序验证(8 篇)

形式化与程序验证 · 4/30 · 2026-09-28证明义务驱动的 Lean 理论构建ProofLoom自动构建Lean模型与理论,由证明义务驱动,含独立Judge审核;未给出具体通过率ProofLoom: Proof-Obligation-Driven Theory Construction for Autoformalizing Research-Level Stochastic Optimization

Formalizing research-level stochastic optimization in Lean requires both an algorithm model and domain theory connecting foundational libraries to convergence proofs. Revising a model to restore provability can change the mathematical claim. We introduce ProofLoom, a fully automated LLM-agent system for Proof-Obligation-Driven Theory Construction. Given a published algorithm, target theorem, and source proof, ProofLoom autonomously constructs the Lean model and supporting theory. Open proof obligations drive the development of definitions, interfaces, lemmas, and proof plans. Signature contracts record evidence and obligations for model revisions; an independent Judge rejects unsupported assumptions and weakened conclusions. Planner expands the published argument into intermediate claims, and Audit checks whether the Lean proof follows it. Across tasks, SOptLib accumulates verified mathematics and construction experience: reusable results are extracted, generalized, and verified, while modeling decisions and failed proof routes are recorded. Later tasks retrieve these results and records and contribute new developments, forming a cycle of construction, accumulation, and reuse. On fifteen textbook and research-paper tasks, ProofLoom obtains mean human ratings of 6.3/7 and 6.4/7, compared with 4.9/7 and 5.0/7 for the strongest of six baselines. Across 33 developments, it produces 490,693 lines of algorithm-local Lean code with no sorry. The formalizations also expose 28 incorrect formulas, proof gaps, and algorithm-analysis mismatches in published sources across 22 developments, each with checked evidence.

阅读 arXiv 原文
形式化与程序验证 · 4/30 · 2026-09-28SPIN 验证分布式检测降级链用Promela建模端点判定流水线与降级回退链,验证六项安全与活性性质;基于生产架构抽象Verifying Graceful Degradation in a Distributed Malware-Detection System with SPIN

Modern endpoint malware detection is distributed: a lightweight agent on each endpoint collects features from a scanned file or process, sends them to a remote server for analysis, and then enforces the returned verdict locally by blocking, quarantining, or disinfecting. Because the endpoint acts on the verdict, the distributed machinery surrounding detection must never turn a transient server failure into a wrong action. We present a formal model, in Promela, of the endpoint decision pipeline of such a system, abstracted from a production architecture at Bitdefender. The model captures the system's graceful-degradation fallback chain: when the primary analysis server times out, the endpoint falls back to an older legacy-protocol server, and failing that to a reduced-signature local scan, before enforcing a verdict. Assuming detection signatures are sound, we specify six safety and liveness properties in linear temporal logic (LTL) and verify them exhaustively with the SPIN model checker. We prove that the fallback machinery never causes a false positive (an enforcement action against a benign file), commits to exactly one verdict per scan even when timed-out responses arrive late, weakens detection strength only in an explicit and ordered way, and always terminates in an enforcement decision, so the pipeline is deadlock-free. Each property is checked to hold non-vacuously, and we report how the state space grows with concurrent scans and endpoints. The work shows how model checking can give strong correctness guarantees for the failure-handling logic of a production security system, a layer that has received little direct formal attention.

阅读 arXiv 原文
形式化与程序验证 · 6/30 · 2026-09-28结构化知识何时助定理证明构建MathKG语义知识图谱并在miniF2F做消融,比较五模型与四种增强模式;未见完整结果When Does Structured Knowledge Help Neural Theorem Proving?

Does structured mathematical knowledge help LLMs prove theorems in Lean 4? If so, for which models, and does the answer vary by problem? Formal libraries such as Mathlib encode 285,000+ verified theorems with syntactic dependencies, but the semantic layer mathematicians rely on for discovery (analogies, generalizations, cross-domain bridges) remains implicit. We introduce MathAgent, which builds this layer as a knowledge graph, MathKG, and uses it to augment LLM theorem provers. MathKG connects 364 Mathlib theorems and definitions by 9,434 typed semantic edges inferred via LLM-based relation extraction anchored to verified Mathlib declarations. We run a controlled ablation across four augmentation modes (no context, knowledge-graph context, Mathlib retrieval, both) and five models: Qwen3-8B/32B, their Lean-specialized derivatives Goedel-Prover-V2-8B/32B, and Claude Sonnet 4.6, on miniF2F, plus PutnamBench and MathOlympiadBench for Sonnet. Three findings emerge. (i) Specialization dominates augmentation: Lean fine-tuning adds 33-38 percentage points of solve rate in every mode, and a specialized 8B model beats a $4\times$ larger general one by 29-35 points, while no augmentation mode improves solve rate by more than 3 points. (ii) Augmentation is capability-conditioned: knowledge-graph context helps small models but hurts large ones, with the specialized model gaining more relative to its general base at every scale. (iii) Yet the augmentation modes solve different problems: an oracle selecting the best mode per problem solves 6% to 58% more than the unaugmented prover, a complementarity effect that strengthens on harder problems (32% more on PutnamBench). These results motivate adaptive strategies that select augmentation by model capability and problem. Code, data, and artifacts are available at https://github.com/sarehnabi/mathagent

阅读 arXiv 原文
形式化与程序验证 · 7/30 · 2026-09-27符号验证强化访问控制策略合成构建CedarInstruct数据集并以验证器信号做两阶段训练,提升可验证Cedar策略合成RAISE: Reinforcing Access Control Policy Synthesis in LLMs via Symbolic Evaluation

Translating natural-language access-control requirements into policies requires careful reasoning about permissions, constraints, and exceptions, and even frontier LLMs often produce policies that violate the intended authorization semantics. We construct CedarInstruct, to our knowledge the first dataset that supports both training and semantic evaluation for formally verifiable Cedar policy synthesis. It contains 5,800 scenarios across 44 domains and 1,408 representing a single synthetic organization, each with a verified target policy and an executable verification plan. On this data we introduce RAISE, which trains policy synthesizers from formal verification in two stages, verified supervised fine-tuning (SFT) followed by a reinforcement learning (RL) stage that learns from verifier signal. We find that SFT succeeds largely by letting models express authorization logic they already have, since untrained models rarely write valid Cedar but often reason correctly when they do. After SFT, how the verifier's information is used matters more than how much of it is used. Of six RL instantiations that consume progressively richer verifier signal, only RAISE-OC improves meaningfully on SFT; it turns failed checks and symbolic counterexamples into guided exploration and learns from the result with off-context GRPO. With about 5.4K verified scenarios and LoRA fine-tuning, RAISE-OC trains Qwen3.5-9B to surpass zero-shot GPT-6 Astra and Claude Opus 5 by 13.33 and 16.26 percentage points in semantic success on held-out scenarios, and training transfers to the independently constructed CedarBench.

阅读 arXiv 原文
形式化与程序验证 · 7/30 · 2026-09-27路径感知的程序规约生成指出约三成验证通过的规约未刻画程序特异行为,Path2Spec以分治按路径生成约束Path2Spec: Path-Aware Specification Generation via Large Language Models

Formal specifications are critical for program verification, comprehension, and maintenance. However, manually writing them is costly and difficult to scale. Recent studies have shown that Large Language Models (LLMs) are promising for automated specification generation, but existing methods suffer from quality issues. We analyze a state-of-the-art approach and find that at least 34.6% of successfully verified specifications actually fail to meaningfully capture the program's distinct behavior, which is a quality issue not captured by metrics that only measure verification success. We further found that a major factor contributing to such hidden quality issues stems from the design of existing methods: these methods treat a program as a single unit, resulting in overly general, coarse-grained constraints. To this end, we introduce Path2Spec, a divide-and-conquer framework that addresses these limitations through systematic path-based reasoning. Path2Spec leverages LLMs to extract all execution paths from an input program, generates path-specific specifications for each, and merges them into a comprehensive overall specification. For complex programs where path-based generation struggles, Path2Spec employs a decompose-then-retry strategy that recursively breaks a program into smaller subprograms based on logical branches, generates specifications for each, and merges them back. We evaluate Path2Spec on two public benchmarks: SG-Bench (120 programs) and SV-COMP (265 programs). Results show that Path2Spec can outperform the state-of-the-art baseline SpecGen: 87.5% versus 66.7% on SG-Bench, and 83.0% versus 44.2% on SV-COMP. Human evaluation further validates that Path2Spec generates higher-quality specifications with precise semantic alignment to the code.

阅读 arXiv 原文
形式化与程序验证 · 4/30 · 2026-09-26有限轨迹 LTL 的预测式监控比较progression与自动机式预测监控,主张progression方法缓解自动机方案的复杂度问题Progression- vs Automata-based Anticipatory Monitoring of LTL over Finite Traces (Extended Version)

When safety-critical systems are developed from a known internal specification, their correctness can be established by model checking. In the frequent case where such a specification is unknown or inaccessible, runtime verification presents an attractive alternative, e.g., to ascertain that autonomous and agentic systems as well as business processes satisfy desirable properties and/or comply with safety requirements. In this paper we study anticipatory monitoring, an advanced form of runtime verification, where the monitoring state is determined by both the trace prefix seen so far, and all its possible finite-length, future continuations. We focus on monitoring linear-time properties that may involve arithmetic constraints. Automata-based approaches, the de-facto standard in this setting, are notorious for their computational complexity. We propose an alternative approach based on progression and LTLf satisfiability checking, for both propositional and arithmetic settings. We experimentally compare the automata- and progression-based approaches, and a third method that combines the two. Our experiments suggest that the progression-based approach often succeeds in producing a verdict when the automata constructions do not terminate, especially for the arithmetic setting. For the propositional setting, the combined technique provides a good tradeoff.

阅读 arXiv 原文
形式化与程序验证 · 7/30 · 2026-09-26用词汇蕴含支撑形式化推理考察LLM能否提供逻辑式NLI所需的词汇知识,并衡量其对证明搜索的贡献Explaining Textual Entailment with Lexical Entailments: Using LLMs to Supply Lexical Relations for Formal Proofs

Large Language Models (LLMs) are highly capable of natural language reasoning and appear to store a great deal of lexical knowledge, but it is still unclear how much of this knowledge they actually use when reasoning, and whether they use it in the right way. On the other hand, logic-based Natural Language Inference (NLI) systems provide transparent and formally grounded reasoning, but they need to be supplied with rich lexical knowledge to prove inferences beyond purely logical ones. In this paper, we evaluate whether LLMs can identify all lexical knowledge needed to solve NLI problems and how much this knowledge contributes to proof search in a logic-based NLI system. Our research focuses exclusively on structured lexical entailments (e.g., chinchilla$\sqsubseteq$small animal) as a proxy for structured explanations for NLI problems with an entailment label. First, we curate a dataset for a new task of explaining sentential entailments with a set of lexical entailments. The dataset is used to intrinsically evaluate LLMs on generating structured lexical explanations. Then, we use NLI as an extrinsic evaluation in a simple neuro-symbolic setting, assessing whether LLMs can supply sufficient lexical relations to LangPro, a natural-logic theorem prover for natural language. The results show that the proposed task remains challenging even for hosted proprietary LLMs, and that their contribution to theorem proving is moderate: generated relations are often only partially sound and may be tailored to the specific NLI problem rather than representing generally valid lexical knowledge.

阅读 arXiv 原文
形式化与程序验证 · 7/30 · 2026-09-25分布式多智能体自动形式化协议Choir把形式化项目拆为可由独立贡献者完成的任务,经确定性门禁合并,支持多种证明助手Choir: An Open Protocol for Distributed Multi-Agent Autoformalization

AI agents can now formalize entire textbooks and major theorems in proof assistants such as Lean, but current efforts are typically centralized: a single team runs all agents and bears the full computational cost. We introduce Choir, an open protocol for distributed formalization. Choir decomposes a project into tasks that can be completed by independent contributors, each running their own agent with their own LLM subscription, while coordinating entirely through the project's GitHub repository. To support open participation, every contribution is checked by a deterministic gate before merge. Choir supports Lean 4, Isabelle, and Rocq, and is open source and modular, allowing projects to replace individual components or extend the protocol.

阅读 arXiv 原文

软件工程与仓库智能(4 篇)

软件工程与仓库智能 · 8/30 · 2026-09-27编码智能体能否跨仓库协作WideSWE从103个生态挖掘120个跨仓库任务,七种配置完整成功率约11%至43%WideSWE: Can Coding Agents Coordinate Changes Across Repositories?

Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multiple repositories. We introduce WideSWE to evaluate coding agents on such cross-repository tasks. Mining and reviewing changes across 103 software ecosystems yields 120 real-world tasks, balanced between 60 bug fixes and 60 features. We derive prompts from related issues and pull requests. We systematically review and adapt hidden tests to support diverse correct implementations while preserving required behavior and regression checks. Across seven agent configurations, full task success ranges from 10.83% to 42.50%, with the configuration pairing Codex CLI with GPT-5.6-sol achieving the highest rate. Trajectories show agents failing to identify necessary changes, recognizing changes but leaving them unfinished, or modifying the required repositories without fully satisfying the request. To examine whether working on one repository at a time can alleviate these difficulties, we compare it with joint execution under identical prompts. Independent execution mainly recovers omitted work and is less effective at correcting previously attempted but unsuccessful implementations. Joint execution can use information from related repositories to guide implementation and verification. Code is available at https://github.com/ZJU-ACES-ISE/WideSWE.

阅读 arXiv 原文
软件工程与仓库智能 · 10/30 · 2026-09-26把协作失败沉淀为组织协议Relic将反复出现的协作失败转为可执行协议,绑定触发、责任与证据,并称有受控运行评估Relic: From Multi-Agent Collaboration to Persistent Organizational Capability

Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the lesson continue to govern the team? We introduce Relic, which turns recurring collaboration failures into organization-owned, executable protocols. Members reflect on visible work, propose rules, and govern their adoption. Adopted protocols bind triggers, responsibilities, required evidence, and execution consequences to the runtime, while remaining open to revision and retirement. In one traced case, repeated integration friction produces an interface-review rule that governs later pull requests and is revised as work continues. Across 360 controlled runs over ten software workloads and three models, Relic raises complete-contract delivery from 14.06% to 19.76% (+5.71 percentage points) over a matched structured team without the protocol lifecycle, improving all four verified production endpoints in every model stratum. Under fresh-member transfer, behavioral correctness is 25.4% with no inherited protocol, 34.6% with the same rules provided as readable text, and 41.2% with executable bindings, a +6.5-point advantage over text alone. On the full CooperBench benchmark, after excluding broken benchmark pairs, Relic achieves 367/477 (76.9%), establishing the best reported result among peer-structured systems. On the fixed 48-pair same-model subset, Relic also exceeds Solo (29/48 vs. 26/48), reversing the coordination loss exhibited by the official peer baseline. Together, these results show how collaboration experience can become persistent organizational state that remains useful beyond the members who created it.

阅读 arXiv 原文
软件工程与仓库智能 · 3/30 · 2026-09-25以调用图重组代码审查变更基于PR实证发现约四成PR的改动函数由调用图相连,CG-Diff按子图重组审查界面CG-Diff: Organizing Code Changes Around Call Graphs

Tools for code review present code changes file-by-file. We argue that, oftentimes, changes can be better organized around a call graph of the changed code. From an empirical study of GitHub pull requests (PRs) in-the-wild, we find that (1) in around 40% of the PRs, more than half of the changed functions (and methods) are connected by call graphs, and (2) necessary callees are often located in different files from their callers. Based on these findings, we develop the notion of CG-Diffs, subgraphs derived by decomposing a call graph of the changed code into smaller and more navigable directed graphs. We then implement a web-based interface for viewing PRs that restructures the PR around these CG-Diffs. Through a within-subject study comparing our interface against GitHub's PR view, we find that CG-Diffs help orient participants and provide them with more meaningful structures to navigate the PR. We also found several limitations: function nodes repeated across multiple CG-Diffs can be disorienting, and changes not contained in a function (e.g. globals, imports) are not as immediately apparent. Our study shows promise in using call graphs to help contextualize and navigate unfamiliar codebases, which may benefit new contributors to open source, and reviewers of unfamiliar LLM-generated PRs.

阅读 arXiv 原文
软件工程与仓库智能 · 10/30 · 2026-09-25仓库级方法学护栏的实证观察在含智能体活动的仓库中考察八类护栏机制与规则文件的采用广度、共现与时间演变Methodological Harness in Agentic Software Engineering: An Empirical Study on Mining Software Repositories

Context. Agentic software engineering requires mechanisms to coordinate and govern agents' work. The framework motivating this study proposes a methodological harness with eight mechanisms: context engineering, persistent shared knowledge, executable specifications, N-version mindset and parallel agents, normative specifications, structured consultation, evidence-based acceptance, and graduated autonomy. These are realized through artifacts such as rule or context files, specifications, and architectural decision records, but the framework had not been empirically validated at repository scale. Objective. We analyze to what extent and how this harness is observable in repositories with agentic activity. RQ1 characterizes adoption through artifact prevalence, breadth, co-occurrence, and temporal evolution; RQ2 examines rule files - a key observable artifact and persistent source of agent instructions - to assess how their content reflects the proposed mechanisms. Method. From the AIDev dataset of 116,211 GitHub repositories, we analyzed 5,435 using a design stratified by visibility, measured through stars as an indicator of popularity and/or reputation. For RQ1, we detected and quantified artifacts and introduction dates for the seven observable mechanisms. For RQ2, we qualitatively coded rule files from 150 repositories. Results. Population prevalence of at least one mechanism is 21.7%, versus 65.3% among the most visible repositories. Multi-mechanism configurations are rare. Where rule files exist, they almost always guide the agent and state norms, while persistent shared knowledge, executable specifications, structured consultation, and graduated autonomy appear only in a minority. Conclusions. The harness is empirically observable, but mainly through isolated mechanisms rather than the integrated system proposed by the framework.

阅读 arXiv 原文

代码质量与优化(3 篇)

代码质量与优化 · 6/30 · 2026-09-28LLM 时代 MDE 还需保留吗提出ARTHUR多智能体框架,用LLM重构MDE生成的遗留Java代码;摘要未见量化结果Is there a future for models in the LLM era?

In an era where software development is deeply tied with Large Language Models, does Model Driven Engineering (MDE) still make sense? This raises the question of the extent to which MDE can be successfully combined with an LLM approach to address the downsides of each approach separately. In this paper, we try to answer that question by investigating a central Research Question: Can Agents powered by Large Language Models improve code generated from UML diagrams and text specifications by Model Driven Engineering (MDE) tools? To investigate this problem, we developed a novel LLM-powered Multi-Agentic approach called ARTHUR (Architecture Refactoring Through Hybrid UML Reasoning), a framework that aims at combining the reliability of MDE and the ease of use of LLMs. ARTHUR is designed to refactor legacy Java code produced by traditional, rule-based, MDE code generators from UML diagrams. To asses the answer to our question, we have refactored the legacy code of several projects from our custom dataset crafted for this purpose. \name{} which made it possible to add support for modern frameworks like Spring Boot, while ensuring compliance with Model-Based Testing techniques to verify that the code still corresponds to the initial model's specifications. We then measured the results obtained in terms of time and cost, passing test rate, and \texttt{compile@k}, \texttt{pass@k} and \texttt{pass$^k$} metrics. We also observed the effect of generating code directly from the conceptual model without the refactoring. Our preliminary test results show that MDE could not be more far from retirement, after all.

阅读 arXiv 原文
代码质量与优化 · 3/30 · 2026-09-27LLM 协同演化通用博弈算法多智能体LLM元学习在C++中共同演化搜索机制与启发式,在400余环境中评测Modular Discovery of General Game-Playing Algorithms with Large Language Models

General Game Playing across arbitrary games from rules alone remains challenging due to differing algorithmic requirements across game classes and strict decision-time constraints. Rather than hand-designing search heuristics for specific domains, can we leverage Large Language Models (LLMs) to discover general game-playing algorithms? Because language models can propose and refactor structured code, they provide an expressive proposal engine for exploring the space of algorithmic designs. We introduce a multi-agent LLM meta-learning system to co-evolve game-agnostic procedural search mechanisms in C++ alongside domain heuristics synthesized directly from game rules. Controlling the compute budget, we benchmark the discovered mechanisms across more than 400 diverse environments, including OpenSpiel training and held-out games, procedural simulation engines, and games with deep neural policy-value representations trained via PPO. Evaluated via AlphaRank stationary distributions and Soft Condorcet Optimization (SCO) against 15 established MCTS baselines, the discovered search mechanisms consistently achieve top-tier ratings and pairwise ballot majorities over most baselines across independent evolutionary runs, generalizing to unseen human-designed and procedurally synthesized games and remaining competitive with baselines on frozen neural network representations.

阅读 arXiv 原文
代码质量与优化 · -3/30 · 2026-09-24机器人操作中的经验复用与迁移研究报告在仿真与实体电梯按钮任务中,身体知识与经验记录缩短完成时间;样本规模有限Robot Manipulation with GPT-6-Astra: Body Knowledge, Experience Reuse, Emergent Skills, and Sim2Real Transfer

General-purpose multimodal agents can write robot-control programs, but repeated exploration and model-mediated action selection can make execution slow. We study how external body knowledge, successful experience, and executable skills improve an XLeRobot controlled by GPT-6-Astra in a simulated and a physical elevator-button task. In 30 fixed-start simulation trials, complete robot geometry and camera information reduce mean completion time by 57.4% relative to a baseline with only the common control interface and no prior experience; images with synchronized action and state records reduce it by 68.6% without additional body assets. In nine paired comparisons (18 trials) at starts displaced by 10-100 cm, experience recorded at the original start reduces mean time by 58-63% relative to no experience, demonstrating generalization to the tested new starting positions. During experience experiments, GPT-6-Astra spontaneously generates a short visual-feedback program. Researcher-refactored versions reduce mean local-task time by 29-31% in 27 simulation trials. Finally, 12 real-robot trials using operator-confirmed button contact demonstrate sim2real reuse: at a shared nominal start, simulation XML assets and simulation experience reduce mean time by 53.0% and 49.9%, respectively; real experience also transfers to two new starts. These results suggest a practical way to build general-purpose manipulation experiments around GPT-6-Astra: supply machine-readable body descriptions and synchronized demonstrations, and turn useful agent-generated feedback routines into reusable skills, while the agent adapts actions from current images. We release all task prompts, trial-level experimental data, and acquired skill implementations at https://github.com/hesd10/astra-robot-sim2real.

阅读 arXiv 原文

UI 与 GUI Agent(0 篇)

本轮该赛道没有候选论文。

个人知识与本体(14 篇)

个人知识与本体 · 3/30 · 2026-09-28面向社交关系的弹性隐私记忆EP-Mem以用户自定策略区分人物与事件,跨社交角色控制记忆披露;仅摘要所述设计EP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM Agents

Large language model (LLM) agents face critical privacy risks when acting as delegates in human-agent-human communication. To prevent such breaches, agents must understand users' social relationships and adhere to context-dependent social information disclosure boundaries. Current studies on agent memory privacy focus on instantaneous interactions, leaving the long-term relational disclosure problem unexplored. In this paper, we propose EP-Mem, an Elastic Privacy Memory architecture that reframes privacy as user-owned boundary control across social roles. EP-Mem introduces (1) token-level memory driven by user-configurable a privacy policy that stratifies persons and events, combining domain-level default circulation rules with fact-level whitelist/blacklist exceptions; and (2) a pluggable sidecar with a privacy engine that aligns disclosure controls with memory across summary, detail, and boundary granularities, enforced throughout generation, storage, and retrieval. We construct EP-Bench, to our knowledge the first long-term multi-party benchmark with cross-session correlated events for policy-conditioned relational disclosure. Experiments show that EP-Mem achieves 94.0% privacy classification accuracy, improves disclosure-permission judgment from 22% to 68%, and reduces privacy leakage by 75.6%, while maintaining retrieval performance and cross-benchmark generalization.

阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-28编码智能体记忆的强化学习训练提出CAMG长程智能体RL环境,含shell与持久工作区,让记忆行为从预训练文件操作中习得Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning

Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environments of relatively short horizons. This setup ties learned memory behavior to environment-specific interfaces that lie outside the base model's pre-training and must be learned from scratch, so even after post-training, agents struggle to use memory in long-horizon tasks. To address these limitations, we introduce Coding Agent Memory Gym (CAMG), a suite of long-horizon agentic-RL environments spanning Shop, Coding, DeepResearch, and AutoResearch. Alongside each environment's native task interface, CAMG provides executable shell access and an episode-persistent workspace, enabling agents to create, revise, search, and reuse files as memory throughout an episode. We also introduce CAMG-RL, which trains a single policy jointly across all four environments with fully asynchronous PPO, learning this file-based memory behavior directly from downstream task reward, and we train CAMG-RL-4B and CAMG-RL-9B from Qwen3.5 models of matching size. On SWE-bench Verified and MLE-bench Lite, CAMG-RL-4B and CAMG-RL-9B are competitive with Qwen3.5-35B-A3B and Qwen3.5-122B-A10B, respectively.

阅读 arXiv 原文
个人知识与本体 · 4/30 · 2026-09-28运行时按需构建的智能体记忆JAM以分层页存储保留原始历史,由Researcher逐次检索整合,并用Memory-Gym数据做监督微调Just-In-Time Agent Memory with Runtime Agentic Research

Memory is critical for AI agents. Many existing agent-memory systems follow an Ahead-of-Time (AOT) design, constructing memory before a specific request arrives. While this reduces online serving cost, such request-agnostic memory construction can discard fine-grained information that later becomes important. To address this limitation, we propose Just-In-Time Agent Memory (JAM), a trainable framework for query-conditioned context construction at runtime. A Memorizer preserves complete raw histories in a hierarchical page-store with compact navigational summaries, while a Researcher iteratively retrieves, inspects, and integrates evidence for each request. To train these memory-use behaviors, we introduce Memory-Gym, an evidence-grounded data synthesis pipeline covering nine task types across six domains, and optimize the Researcher through verified-trajectory supervised fine-tuning followed by Hint-guided Group Relative Policy Optimization. We demonstrate the effectiveness of JAM across a variety of benchmarks on agent memory and long-context processing, where it achieves stronger task performance than AOT-style memory systems while remaining substantially more efficient than prior trained agentic memory approaches. To support reproducibility and future research, we release our anonymized source code at https://github.com/VectorSpaceLab/general-agentic-memory.

阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-28把人格融入双通路智能体记忆PersMem将人格映射为记忆处理参数,控制情感评估与记忆保留检索;摘要概述为四个步骤PersMem: Internalizing Personality into Dual-Pathway Memory for LLM Agents

The profile of a role-playing agent usually depends on the pre-defined personality in a system prompt, whereas its memory processing pipeline, including prioritisation of stored memories and subsequent retrieval, remains independent of this personality. This separation causes the agent's memory processing to be inconsistent with the pre-defined personality, and makes it difficult to validate whether agent behaviours follow this personality. In this paper, we propose Personality-Integrated Memory (PersMem), which integrates personality into the agent's memory processing pipeline, making it consistently personality-dependent. PersMem processes memory using four steps, where the personality is mapped to operation-specific parameters controlling: (i) affective appraisal annotating emotion states of the user input; (ii) retention of previously stored memories along with the current input; (iii) passive affect-driven memory retrieval exploring memories similar to user input in semantics and personality-guided emotions; and (iv) active goal-driven memory retrieval that refines and selects passively retrieved memories for the reply. Consequently, consistency with the pre-defined personality can be examined by inspecting memory-processing traces during human-agent interactions. We evaluate these personality-dependent differences in attachment and Big Five settings. PersMem exceeds the chance baseline for four-way attachment classification by 23.1 percentage points. In Big Five dialogue comparisons, PersMem achieves 67.5% accuracy, 6.7 percentage points above a baseline using uniformly sampled memories. On CoSER, PersMem achieves an average score of 66.13, with scores of 69.33 for Character Fidelity and 84.33 for Storyline Quality. Together, these results show that PersMem produces distinguishable personality-related memory-processing patterns.

阅读 arXiv 原文
个人知识与本体 · 4/30 · 2026-09-28会话智能体的说话人索引记忆Stashbird以来源溯源组织情节、语义与社区摘要,报告大幅降低摄入提示词tokenStashbird: Efficient Speaker-Indexed Memory for Conversational Agents

AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episodes to derived memory state through explicit provenance. Stashbird organizes memory into episodic records, semantic relations, community summaries, and persisted graph state, with lifecycle operations for incremental updates and episode-level deletion. We evaluate question-answering accuracy and model-facing workload across four long-term memory benchmarks. On LoCoMo, Stashbird uses 76.4x fewer ingestion prompt tokens than Graphiti. Compared with reproduced Hindsight on the same benchmark, it uses 8.1x fewer retrieval prompt tokens, with accuracy 1.6 percentage points lower. It achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.

阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-28选择原始轮次能替代抽取吗预注册研究在LoCoMo与LongMemEval上比较原始轮次选择与LLM抽取,称结果非劣于抽取When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model

Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-extraction memory (one-sided 95% bound -3.0 points against a -5-point margin). Blind human grading narrows the margin but does not change the result. Raw turns cost 3,061 times less to write, and the result holds with a second answer model. Within this study, reranking's gain shrinks as the budget grows. It adds 17.4 points on LoCoMo and 9.1 on LongMemEval when three of 30 candidates are kept. At generous budgets it adds 1.5 and 1.1, and extraction systems are more accurate. This suggests why published results disagree. At matched context, Jev selects as accurately as an LLM reranker (non-inferiority bound -2.0) at a third of the latency, and more accurately than a multi-call graph traversal. Reranking lowers correct abstention. Plans, code and graded answers are released.

阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-28追踪记忆构建顺序的影响RoutePrism用同源两次不同顺序构建记忆并对比,恢复单条记录可挽回大幅精度损失RoutePrism: Tracing Construction Order Effects in Agent Memory

Processing the same records in a different order can discard different evidence, yet endpoint accuracy alone cannot reveal what changed or whether it mattered. We introduce RoutePrism, a diagnostic protocol that builds memory twice from the same source pool in two processing orders, then traces which sources, compiled contexts, and answers differ. Because record content, timestamps, policy, and the answer model all stay fixed, any observed difference is localized to the memory construction step. A matched four-condition intervention tests whether a record displaced by reordering actually carried task-relevant evidence: restoring that single record recovers over 60 percentage points of lost accuracy, while substituting a non-supporting record of equal length does not. We evaluate the protocol on PersonaMem-32K (63 primary queries, 29 users) and 470 LongMemEval-S questions with histories spanning 38 to 62 sessions, replicating the core intervention across five answer models. Survivor selection, defined as the choice of which record a cluster retains, drives most source-level changes, while different memory policies (compaction, bounded recency, MemoChat-style summarization, A-MEM) produce distinct failure signatures at the source, context, and metadata layers.

阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-28记忆投毒攻击的严重度度量提出反事实记忆后悔指标与MemHarm,通过离线配对损失筛选稀疏语义编辑并给出类内认证From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents

As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior. Persistent-memory attacks are typically evaluated by whether they succeed, yet successful attacks can leave persistent states with substantially different downstream consequences. We study this severity as a distinct attack-design objective and formalize it with counterfactual memory regret (CMR), the paired increase in expected downstream loss relative to clean memory. We introduce MemHarm, which predeclares a finite class of sparse, grounded semantic edits, evaluates candidates through the normal agent memory interface using offline paired-loss feedback, and certifies resolved selections within that class. Compared with attack-success optimization, CMR-guided selection produces substantially larger downstream loss while retaining most of the success-rate gain. Across two agent benchmarks and diverse memory designs, MemHarm attains the highest CMR point estimates among the evaluated general attacks on identical support. Factor-removal interventions link this harm to the selected semantic factor, and native-agent deployments verify the write-to-fresh-process attack path.

阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-26智能体记忆固化的可治理性四阶段研究关注记忆固化决策质量与可信度,报告未达阈值的负结果与防泄漏协议The Epistemics of Agent Memory: Measuring, and Governing, the Consolidation Decision in Long-Horizon LLM Agents

Long-horizon LLM agents must convert accumulated experience into durable memory, deciding what to keep, compress, abstract into reusable skills and rules, or forget. We report a four-phase research program on this consolidation problem whose central finding is a shift in what is measured: from how much an agent remembers, to whether its consolidation decisions are any good, to whether those decisions can be trusted. Phase 1 learns episodic boundaries from agent traces by downstream utility; an honest near-miss (oracle correlation 0.691 vs a 0.70 bar) whose lasting output is a three-gate anti-leakage protocol. Phase 2 learns when to promote experience and to which abstraction level under a token budget, achieving a verified +22.7% task-success improvement with 7x compression, but exposing a degenerate-forgetting failure and a distribution-shift failure mode we name lambda-prevalence coupling. Phase 3 introduces ConsolidationBench, an oracle-by-construction benchmark that scores consolidation decisions against a known optimum on three non-circular axes; production retrieval systems retain information yet score zero on cross-level transfer. Phase 4 introduces governed consolidation: the decision wrapped in poison-resistance, reversibility, and auditability guarantees with a quality gate. Governance is statistically distinct from the quality score ($r^2 = 0.43$; partial $r = 0.27$; identical-quality policies differ threefold in governance), so the contribution survives independently of the metric's external validity. On that question we report a resolved negative: after a graded-reuse redesign removed a structural ceiling, a two-benchmark study with 2,532 real answer cells finds the quality score does not predict real transfer accuracy (pooled Spearman $ρ= -0.24$, n = 12, CI spanning zero). An adversarial self-critique pass cleared the final claim set with zero surviving overclaims.

阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-26异构记忆提供者的路由管理评估13种记忆方法后,MemAgent将记忆视为路由问题,按内容选择检索与存储来源MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents

Current agents remain largely stateless across tasks, limiting their ability to continually improve from prior interactions and making memory essential for long-horizon agentic behavior. Existing memory methods seek to reuse past experience, but most rely on a single memory representation (e.g., trajectories, reflections, skills, structured knowledge) whose effectiveness varies across task distributions. Rethinking this design space, we evaluate 13 memory methods and find that no single method generalizes across benchmarks, revealing the potential of managing heterogeneous memory providers. We formulate agent memory as a routing problem in which a memory agent decides which memory provider to retrieve from, whether to inject short-term memory, and which providers should store the resulting experience. Based on this perspective, we propose MemAgent, featuring a content-aware routing architecture and a training-data synthesis pipeline. The routing architecture combines content-aware probing before retrieval, short-term memory gating during execution, and selective multi-provider storage, while the training pipeline synthesizes phase-specific supervision for routing decisions. Across GAIA, WebWalkerQA, and xBench-DS, MemAgent improves average accuracy by 10.0% and outperforms every individual memory method across all three benchmarks. These gains come with less than 0.3% routing overhead and a 12% reduction in average task steps.

阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-26带检索分数界的有损记忆压缩LAM用确定性去重给出分数扰动界,报告在600条轨迹上删除约22%观测tokenLAM: Efficient Lossy Agent Memory Framework With A Retrieval-Score Error Bound

Agent memory grows as agents read inputs, reason, and call tools. Longer histories increase inference cost and eventually exceed the context window. LLM-based summarization reduces this history but adds latency and provides no explicit bound on information loss. We propose LAM, a Lossy Agent Memory system with three components: a deterministic deduplication rule with a substitution bound on retrieval scores - a bound on score perturbation, not a certificate of unchanged ranking; a memory manager that preserves the cached prefix and overlaps compaction with inference; and a performance model that estimates compaction costs before deployment. On 600 agent trajectories, LAM removes 22.47% of observation tokens while retaining 99.984% of the measured gold-patch evidence. At a fixed deletion set, the performance model predicts a 71.4x-91.6x end-to-end speedup from removing records before prefill instead of deleting them from a prefilled context. That benefit comes from the schedule rather than the rule and applies to any prefix-preserving test.

阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-25把智能体记忆做成一等中间件立场论文主张记忆应成为可插拔中间件,提出双向可插拔、多租户隔离等六项挑战Memory as Middleware for Self-Improving AI Agents

AI agents are stateless across sessions by default and therefore operationally amnesic: each session begins with little durable knowledge of prior failures, repairs, preferences, or successful strategies. As a result, agents repeat the same mistakes and discard hard-won experience. The dominant fix is \emph{bespoke memory}---retrieval, persistence, and learning logic hand-wired into one agent and bound to one storage engine. This creates a fragmented landscape where memory cannot be swapped, shared, isolated, or reasoned about independently of the agent that owns it. We argue that this is a middleware problem: agent memory deserves a first-class, pluggable layer, just as data access, messaging, and persistence each became middleware concerns. We develop this vision through six systems challenges: two-sided pluggability, host-native interposition, multi-tenant isolation, write-path consistency, federated sharing with provenance, and lifecycle governance. We present ALTK-Evolve, a reference implementation of memory middleware for self-improving agents, and use it to motivate a broader research agenda for future memory middleware.

阅读 arXiv 原文
个人知识与本体 · 6/30 · 2026-09-25多跳智能体记忆的拓扑巩固EngramRAG结合唤醒与梦境两态,用使用加权PageRank与拓扑衰减缓解多跳检索问题EngramRAG: Dynamic Usage-Weighted Topology and Synaptic Consolidation for Multi-Hop Agentic Memory

As autonomous LLM agents are deployed across multi-session environments, conventional memory architectures suffer from Associative Blindness (inability to traverse multi-hop relational dependencies), Scaffolding Amnesia (temporal decay evicting core persona invariants), and Static Topology Stagnation (immutable graphs ignoring usage dynamics). Grounded in Complementary Learning Systems (CLS) principles, we propose EngramRAG, an adaptive memory architecture coupling a low-latency Waking State reflex with an asynchronous background Dreaming State consolidation cycle. EngramRAG introduces: (1) Usage-Modulated Personalized PageRank (U-PPR), where transition probabilities adapt via Hebbian plasticity to promote persistent entities into high-centrality Epistemic Macro-Hubs; (2) Consolidation-Activated Topology Decay (CATD), which scales retention half-life by topological load-bearing weight rather than wall-clock recency, protected by a cold-start grace period (N_grace >= 4); (3) Directed SUPERSEDES DAG filtering to suppress obsolete state during fact mutations; and (4) Triple-source hybrid retrieval fusing dense vectors, BM25, and U-PPR via dynamic Reciprocal Rank Fusion (RRF). Evaluating on all 1,982 QA pairs across 10 long-term conversations in the LoCoMo benchmark, EngramRAG achieves +38.9% relative improvement in Recall@5 (53.21% vs. 38.29%, p < 0.001) and +43.1% in MRR (0.4203 vs. 0.2937) over dense vector RAG, significantly outperforming Okapi BM25 (48.66%) and isolated static graph retrieval (8.50%). On temporal reasoning, EngramRAG reaches 62.33% Recall@5 (+16.67 points over dense vectors). In controlled mutation tests, SUPERSEDES suppresses split-brain hallucinations from 70.0% to 0.0%, while 90-day simulations show 100.0% scaffolding retention under a 26.21ms interactive retrieval reflex.

阅读 arXiv 原文
个人知识与本体 · 4/30 · 2026-09-25世界模型验证的反事实记忆COUNTERMEM在失败动作后借助可执行世界模型评估替代动作,构建经验证的反事实记忆COUNTERMEM: World-Model Verified Counter-Factual Memory for Language Agents

Existing agent memory frameworks mainly create memory through an agent's interaction with the factual world, e.g., remembering feedback from actions taken to improve performance on future tasks. However, these frameworks seldom ask the "what if" question during memory construction: what if a different action had been taken, would the feedback have changed, and how could this feedback become useful memory? Obtaining such feedback directly in an active environment can be expensive and can alter the state needed for comparison. In this work, we introduce COUNTERMEM, a reinforcement-learning framework for constructing and using verified counterfactual memory across tasks. After a failed action, COUNTERMEM evaluates local alternatives from a copy or reset of the original state using executable world models, such as tests, proof checkers, and solvers. It stores improvements with the original and corrected actions, checked outcomes, and conditions for reuse. A learned memory-use policy selects a retrieved record or skips memory to balance task success and interaction cost, while the base LLM remains fixed. Both memory and policy are frozen during held-out evaluation. We evaluate COUNTERMEM on 12 benchmark settings across six domains. With gpt-oss-120b, COUNTERMEM improves both ReAct and Reflexion on all 12 benchmarks across six domains, averaging a gain of 12.6 percentage points over their unaugmented versions. In the four-domain comparison across two backbones, task-run tokens decrease by 7.7-42.0%, excluding offline selector-training costs. Further analyses show that removing verification or persistent storage weakens the gains, while applying verified corrections to unsuitable decisions can reverse them. Code will be released upon acceptance.

阅读 arXiv 原文

人机协同与对齐(0 篇)

本轮该赛道没有候选论文。