形式化与程序验证(8 篇)
形式化与程序验证 · 4/30 · 2026-09-28证明义务驱动的 Lean 理论构建ProofLoom自动构建Lean模型与理论,由证明义务驱动,含独立Judge审核;未给出具体通过率ProofLoom: Proof-Obligation-Driven Theory Construction for Autoformalizing Research-Level Stochastic Optimization
Formalizing research-level stochastic optimization in Lean requires both an algorithm model and domain theory connecting foundational libraries to convergence proofs. Revising a model to restore provability can change the mathematical claim. We introduce ProofLoom, a fully automated LLM-agent system for Proof-Obligation-Driven Theory Construction. Given a published algorithm, target theorem, and source proof, ProofLoom autonomously constructs the Lean model and supporting theory. Open proof obligations drive the development of definitions, interfaces, lemmas, and proof plans. Signature contracts record evidence and obligations for model revisions; an independent Judge rejects unsupported assumptions and weakened conclusions. Planner expands the published argument into intermediate claims, and Audit checks whether the Lean proof follows it. Across tasks, SOptLib accumulates verified mathematics and construction experience: reusable results are extracted, generalized, and verified, while modeling decisions and failed proof routes are recorded. Later tasks retrieve these results and records and contribute new developments, forming a cycle of construction, accumulation, and reuse. On fifteen textbook and research-paper tasks, ProofLoom obtains mean human ratings of 6.3/7 and 6.4/7, compared with 4.9/7 and 5.0/7 for the strongest of six baselines. Across 33 developments, it produces 490,693 lines of algorithm-local Lean code with no sorry. The formalizations also expose 28 incorrect formulas, proof gaps, and algorithm-analysis mismatches in published sources across 22 developments, each with checked evidence.
阅读 arXiv 原文
形式化与程序验证 · 4/30 · 2026-09-28SPIN 验证分布式检测降级链用Promela建模端点判定流水线与降级回退链,验证六项安全与活性性质;基于生产架构抽象Verifying Graceful Degradation in a Distributed Malware-Detection System with SPIN
Modern endpoint malware detection is distributed: a lightweight agent on each endpoint collects features from a scanned file or process, sends them to a remote server for analysis, and then enforces the returned verdict locally by blocking, quarantining, or disinfecting. Because the endpoint acts on the verdict, the distributed machinery surrounding detection must never turn a transient server failure into a wrong action. We present a formal model, in Promela, of the endpoint decision pipeline of such a system, abstracted from a production architecture at Bitdefender. The model captures the system's graceful-degradation fallback chain: when the primary analysis server times out, the endpoint falls back to an older legacy-protocol server, and failing that to a reduced-signature local scan, before enforcing a verdict. Assuming detection signatures are sound, we specify six safety and liveness properties in linear temporal logic (LTL) and verify them exhaustively with the SPIN model checker. We prove that the fallback machinery never causes a false positive (an enforcement action against a benign file), commits to exactly one verdict per scan even when timed-out responses arrive late, weakens detection strength only in an explicit and ordered way, and always terminates in an enforcement decision, so the pipeline is deadlock-free. Each property is checked to hold non-vacuously, and we report how the state space grows with concurrent scans and endpoints. The work shows how model checking can give strong correctness guarantees for the failure-handling logic of a production security system, a layer that has received little direct formal attention.
阅读 arXiv 原文
形式化与程序验证 · 6/30 · 2026-09-28结构化知识何时助定理证明构建MathKG语义知识图谱并在miniF2F做消融,比较五模型与四种增强模式;未见完整结果When Does Structured Knowledge Help Neural Theorem Proving?
Does structured mathematical knowledge help LLMs prove theorems in Lean 4? If so, for which models, and does the answer vary by problem? Formal libraries such as Mathlib encode 285,000+ verified theorems with syntactic dependencies, but the semantic layer mathematicians rely on for discovery (analogies, generalizations, cross-domain bridges) remains implicit. We introduce MathAgent, which builds this layer as a knowledge graph, MathKG, and uses it to augment LLM theorem provers. MathKG connects 364 Mathlib theorems and definitions by 9,434 typed semantic edges inferred via LLM-based relation extraction anchored to verified Mathlib declarations. We run a controlled ablation across four augmentation modes (no context, knowledge-graph context, Mathlib retrieval, both) and five models: Qwen3-8B/32B, their Lean-specialized derivatives Goedel-Prover-V2-8B/32B, and Claude Sonnet 4.6, on miniF2F, plus PutnamBench and MathOlympiadBench for Sonnet. Three findings emerge. (i) Specialization dominates augmentation: Lean fine-tuning adds 33-38 percentage points of solve rate in every mode, and a specialized 8B model beats a $4\times$ larger general one by 29-35 points, while no augmentation mode improves solve rate by more than 3 points. (ii) Augmentation is capability-conditioned: knowledge-graph context helps small models but hurts large ones, with the specialized model gaining more relative to its general base at every scale. (iii) Yet the augmentation modes solve different problems: an oracle selecting the best mode per problem solves 6% to 58% more than the unaugmented prover, a complementarity effect that strengthens on harder problems (32% more on PutnamBench). These results motivate adaptive strategies that select augmentation by model capability and problem. Code, data, and artifacts are available at https://github.com/sarehnabi/mathagent
阅读 arXiv 原文
形式化与程序验证 · 7/30 · 2026-09-27符号验证强化访问控制策略合成构建CedarInstruct数据集并以验证器信号做两阶段训练,提升可验证Cedar策略合成RAISE: Reinforcing Access Control Policy Synthesis in LLMs via Symbolic Evaluation
Translating natural-language access-control requirements into policies requires careful reasoning about permissions, constraints, and exceptions, and even frontier LLMs often produce policies that violate the intended authorization semantics. We construct CedarInstruct, to our knowledge the first dataset that supports both training and semantic evaluation for formally verifiable Cedar policy synthesis. It contains 5,800 scenarios across 44 domains and 1,408 representing a single synthetic organization, each with a verified target policy and an executable verification plan. On this data we introduce RAISE, which trains policy synthesizers from formal verification in two stages, verified supervised fine-tuning (SFT) followed by a reinforcement learning (RL) stage that learns from verifier signal. We find that SFT succeeds largely by letting models express authorization logic they already have, since untrained models rarely write valid Cedar but often reason correctly when they do. After SFT, how the verifier's information is used matters more than how much of it is used. Of six RL instantiations that consume progressively richer verifier signal, only RAISE-OC improves meaningfully on SFT; it turns failed checks and symbolic counterexamples into guided exploration and learns from the result with off-context GRPO. With about 5.4K verified scenarios and LoRA fine-tuning, RAISE-OC trains Qwen3.5-9B to surpass zero-shot GPT-6 Astra and Claude Opus 5 by 13.33 and 16.26 percentage points in semantic success on held-out scenarios, and training transfers to the independently constructed CedarBench.
阅读 arXiv 原文
形式化与程序验证 · 7/30 · 2026-09-27路径感知的程序规约生成指出约三成验证通过的规约未刻画程序特异行为,Path2Spec以分治按路径生成约束Path2Spec: Path-Aware Specification Generation via Large Language Models
Formal specifications are critical for program verification, comprehension, and maintenance. However, manually writing them is costly and difficult to scale. Recent studies have shown that Large Language Models (LLMs) are promising for automated specification generation, but existing methods suffer from quality issues. We analyze a state-of-the-art approach and find that at least 34.6% of successfully verified specifications actually fail to meaningfully capture the program's distinct behavior, which is a quality issue not captured by metrics that only measure verification success. We further found that a major factor contributing to such hidden quality issues stems from the design of existing methods: these methods treat a program as a single unit, resulting in overly general, coarse-grained constraints. To this end, we introduce Path2Spec, a divide-and-conquer framework that addresses these limitations through systematic path-based reasoning. Path2Spec leverages LLMs to extract all execution paths from an input program, generates path-specific specifications for each, and merges them into a comprehensive overall specification. For complex programs where path-based generation struggles, Path2Spec employs a decompose-then-retry strategy that recursively breaks a program into smaller subprograms based on logical branches, generates specifications for each, and merges them back. We evaluate Path2Spec on two public benchmarks: SG-Bench (120 programs) and SV-COMP (265 programs). Results show that Path2Spec can outperform the state-of-the-art baseline SpecGen: 87.5% versus 66.7% on SG-Bench, and 83.0% versus 44.2% on SV-COMP. Human evaluation further validates that Path2Spec generates higher-quality specifications with precise semantic alignment to the code.
阅读 arXiv 原文
形式化与程序验证 · 4/30 · 2026-09-26有限轨迹 LTL 的预测式监控比较progression与自动机式预测监控,主张progression方法缓解自动机方案的复杂度问题Progression- vs Automata-based Anticipatory Monitoring of LTL over Finite Traces (Extended Version)
When safety-critical systems are developed from a known internal specification, their correctness can be established by model checking. In the frequent case where such a specification is unknown or inaccessible, runtime verification presents an attractive alternative, e.g., to ascertain that autonomous and agentic systems as well as business processes satisfy desirable properties and/or comply with safety requirements. In this paper we study anticipatory monitoring, an advanced form of runtime verification, where the monitoring state is determined by both the trace prefix seen so far, and all its possible finite-length, future continuations. We focus on monitoring linear-time properties that may involve arithmetic constraints. Automata-based approaches, the de-facto standard in this setting, are notorious for their computational complexity. We propose an alternative approach based on progression and LTLf satisfiability checking, for both propositional and arithmetic settings. We experimentally compare the automata- and progression-based approaches, and a third method that combines the two. Our experiments suggest that the progression-based approach often succeeds in producing a verdict when the automata constructions do not terminate, especially for the arithmetic setting. For the propositional setting, the combined technique provides a good tradeoff.
阅读 arXiv 原文
形式化与程序验证 · 7/30 · 2026-09-26用词汇蕴含支撑形式化推理考察LLM能否提供逻辑式NLI所需的词汇知识,并衡量其对证明搜索的贡献Explaining Textual Entailment with Lexical Entailments: Using LLMs to Supply Lexical Relations for Formal Proofs
Large Language Models (LLMs) are highly capable of natural language reasoning and appear to store a great deal of lexical knowledge, but it is still unclear how much of this knowledge they actually use when reasoning, and whether they use it in the right way. On the other hand, logic-based Natural Language Inference (NLI) systems provide transparent and formally grounded reasoning, but they need to be supplied with rich lexical knowledge to prove inferences beyond purely logical ones. In this paper, we evaluate whether LLMs can identify all lexical knowledge needed to solve NLI problems and how much this knowledge contributes to proof search in a logic-based NLI system. Our research focuses exclusively on structured lexical entailments (e.g., chinchilla$\sqsubseteq$small animal) as a proxy for structured explanations for NLI problems with an entailment label. First, we curate a dataset for a new task of explaining sentential entailments with a set of lexical entailments. The dataset is used to intrinsically evaluate LLMs on generating structured lexical explanations. Then, we use NLI as an extrinsic evaluation in a simple neuro-symbolic setting, assessing whether LLMs can supply sufficient lexical relations to LangPro, a natural-logic theorem prover for natural language. The results show that the proposed task remains challenging even for hosted proprietary LLMs, and that their contribution to theorem proving is moderate: generated relations are often only partially sound and may be tailored to the specific NLI problem rather than representing generally valid lexical knowledge.
阅读 arXiv 原文
形式化与程序验证 · 7/30 · 2026-09-25分布式多智能体自动形式化协议Choir把形式化项目拆为可由独立贡献者完成的任务,经确定性门禁合并,支持多种证明助手Choir: An Open Protocol for Distributed Multi-Agent Autoformalization
AI agents can now formalize entire textbooks and major theorems in proof assistants such as Lean, but current efforts are typically centralized: a single team runs all agents and bears the full computational cost. We introduce Choir, an open protocol for distributed formalization. Choir decomposes a project into tasks that can be completed by independent contributors, each running their own agent with their own LLM subscription, while coordinating entirely through the project's GitHub repository. To support open participation, every contribution is checked by a deterministic gate before merge. Choir supports Lean 4, Isabelle, and Rocq, and is open source and modular, allowing projects to replace individual components or extend the protocol.
阅读 arXiv 原文
软件工程与仓库智能(4 篇)
软件工程与仓库智能 · 8/30 · 2026-09-27编码智能体能否跨仓库协作WideSWE从103个生态挖掘120个跨仓库任务,七种配置完整成功率约11%至43%WideSWE: Can Coding Agents Coordinate Changes Across Repositories?
Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multiple repositories. We introduce WideSWE to evaluate coding agents on such cross-repository tasks. Mining and reviewing changes across 103 software ecosystems yields 120 real-world tasks, balanced between 60 bug fixes and 60 features. We derive prompts from related issues and pull requests. We systematically review and adapt hidden tests to support diverse correct implementations while preserving required behavior and regression checks. Across seven agent configurations, full task success ranges from 10.83% to 42.50%, with the configuration pairing Codex CLI with GPT-5.6-sol achieving the highest rate. Trajectories show agents failing to identify necessary changes, recognizing changes but leaving them unfinished, or modifying the required repositories without fully satisfying the request. To examine whether working on one repository at a time can alleviate these difficulties, we compare it with joint execution under identical prompts. Independent execution mainly recovers omitted work and is less effective at correcting previously attempted but unsuccessful implementations. Joint execution can use information from related repositories to guide implementation and verification. Code is available at https://github.com/ZJU-ACES-ISE/WideSWE.
阅读 arXiv 原文
软件工程与仓库智能 · 10/30 · 2026-09-26把协作失败沉淀为组织协议Relic将反复出现的协作失败转为可执行协议,绑定触发、责任与证据,并称有受控运行评估Relic: From Multi-Agent Collaboration to Persistent Organizational Capability
Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the lesson continue to govern the team? We introduce Relic, which turns recurring collaboration failures into organization-owned, executable protocols. Members reflect on visible work, propose rules, and govern their adoption. Adopted protocols bind triggers, responsibilities, required evidence, and execution consequences to the runtime, while remaining open to revision and retirement. In one traced case, repeated integration friction produces an interface-review rule that governs later pull requests and is revised as work continues. Across 360 controlled runs over ten software workloads and three models, Relic raises complete-contract delivery from 14.06% to 19.76% (+5.71 percentage points) over a matched structured team without the protocol lifecycle, improving all four verified production endpoints in every model stratum. Under fresh-member transfer, behavioral correctness is 25.4% with no inherited protocol, 34.6% with the same rules provided as readable text, and 41.2% with executable bindings, a +6.5-point advantage over text alone. On the full CooperBench benchmark, after excluding broken benchmark pairs, Relic achieves 367/477 (76.9%), establishing the best reported result among peer-structured systems. On the fixed 48-pair same-model subset, Relic also exceeds Solo (29/48 vs. 26/48), reversing the coordination loss exhibited by the official peer baseline. Together, these results show how collaboration experience can become persistent organizational state that remains useful beyond the members who created it.
阅读 arXiv 原文
软件工程与仓库智能 · 3/30 · 2026-09-25以调用图重组代码审查变更基于PR实证发现约四成PR的改动函数由调用图相连,CG-Diff按子图重组审查界面CG-Diff: Organizing Code Changes Around Call Graphs
Tools for code review present code changes file-by-file. We argue that, oftentimes, changes can be better organized around a call graph of the changed code. From an empirical study of GitHub pull requests (PRs) in-the-wild, we find that (1) in around 40% of the PRs, more than half of the changed functions (and methods) are connected by call graphs, and (2) necessary callees are often located in different files from their callers. Based on these findings, we develop the notion of CG-Diffs, subgraphs derived by decomposing a call graph of the changed code into smaller and more navigable directed graphs. We then implement a web-based interface for viewing PRs that restructures the PR around these CG-Diffs. Through a within-subject study comparing our interface against GitHub's PR view, we find that CG-Diffs help orient participants and provide them with more meaningful structures to navigate the PR. We also found several limitations: function nodes repeated across multiple CG-Diffs can be disorienting, and changes not contained in a function (e.g. globals, imports) are not as immediately apparent. Our study shows promise in using call graphs to help contextualize and navigate unfamiliar codebases, which may benefit new contributors to open source, and reviewers of unfamiliar LLM-generated PRs.
阅读 arXiv 原文
软件工程与仓库智能 · 10/30 · 2026-09-25仓库级方法学护栏的实证观察在含智能体活动的仓库中考察八类护栏机制与规则文件的采用广度、共现与时间演变Methodological Harness in Agentic Software Engineering: An Empirical Study on Mining Software Repositories
Context. Agentic software engineering requires mechanisms to coordinate and govern agents' work. The framework motivating this study proposes a methodological harness with eight mechanisms: context engineering, persistent shared knowledge, executable specifications, N-version mindset and parallel agents, normative specifications, structured consultation, evidence-based acceptance, and graduated autonomy. These are realized through artifacts such as rule or context files, specifications, and architectural decision records, but the framework had not been empirically validated at repository scale. Objective. We analyze to what extent and how this harness is observable in repositories with agentic activity. RQ1 characterizes adoption through artifact prevalence, breadth, co-occurrence, and temporal evolution; RQ2 examines rule files - a key observable artifact and persistent source of agent instructions - to assess how their content reflects the proposed mechanisms. Method. From the AIDev dataset of 116,211 GitHub repositories, we analyzed 5,435 using a design stratified by visibility, measured through stars as an indicator of popularity and/or reputation. For RQ1, we detected and quantified artifacts and introduction dates for the seven observable mechanisms. For RQ2, we qualitatively coded rule files from 150 repositories. Results. Population prevalence of at least one mechanism is 21.7%, versus 65.3% among the most visible repositories. Multi-mechanism configurations are rare. Where rule files exist, they almost always guide the agent and state norms, while persistent shared knowledge, executable specifications, structured consultation, and graduated autonomy appear only in a minority. Conclusions. The harness is empirically observable, but mainly through isolated mechanisms rather than the integrated system proposed by the framework.
阅读 arXiv 原文
个人知识与本体(14 篇)
个人知识与本体 · 3/30 · 2026-09-28面向社交关系的弹性隐私记忆EP-Mem以用户自定策略区分人物与事件,跨社交角色控制记忆披露;仅摘要所述设计EP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM Agents
Large language model (LLM) agents face critical privacy risks when acting as delegates in human-agent-human communication. To prevent such breaches, agents must understand users' social relationships and adhere to context-dependent social information disclosure boundaries. Current studies on agent memory privacy focus on instantaneous interactions, leaving the long-term relational disclosure problem unexplored. In this paper, we propose EP-Mem, an Elastic Privacy Memory architecture that reframes privacy as user-owned boundary control across social roles. EP-Mem introduces (1) token-level memory driven by user-configurable a privacy policy that stratifies persons and events, combining domain-level default circulation rules with fact-level whitelist/blacklist exceptions; and (2) a pluggable sidecar with a privacy engine that aligns disclosure controls with memory across summary, detail, and boundary granularities, enforced throughout generation, storage, and retrieval. We construct EP-Bench, to our knowledge the first long-term multi-party benchmark with cross-session correlated events for policy-conditioned relational disclosure. Experiments show that EP-Mem achieves 94.0% privacy classification accuracy, improves disclosure-permission judgment from 22% to 68%, and reduces privacy leakage by 75.6%, while maintaining retrieval performance and cross-benchmark generalization.
阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-28编码智能体记忆的强化学习训练提出CAMG长程智能体RL环境,含shell与持久工作区,让记忆行为从预训练文件操作中习得Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning
Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environments of relatively short horizons. This setup ties learned memory behavior to environment-specific interfaces that lie outside the base model's pre-training and must be learned from scratch, so even after post-training, agents struggle to use memory in long-horizon tasks. To address these limitations, we introduce Coding Agent Memory Gym (CAMG), a suite of long-horizon agentic-RL environments spanning Shop, Coding, DeepResearch, and AutoResearch. Alongside each environment's native task interface, CAMG provides executable shell access and an episode-persistent workspace, enabling agents to create, revise, search, and reuse files as memory throughout an episode. We also introduce CAMG-RL, which trains a single policy jointly across all four environments with fully asynchronous PPO, learning this file-based memory behavior directly from downstream task reward, and we train CAMG-RL-4B and CAMG-RL-9B from Qwen3.5 models of matching size. On SWE-bench Verified and MLE-bench Lite, CAMG-RL-4B and CAMG-RL-9B are competitive with Qwen3.5-35B-A3B and Qwen3.5-122B-A10B, respectively.
阅读 arXiv 原文
个人知识与本体 · 4/30 · 2026-09-28运行时按需构建的智能体记忆JAM以分层页存储保留原始历史,由Researcher逐次检索整合,并用Memory-Gym数据做监督微调Just-In-Time Agent Memory with Runtime Agentic Research
Memory is critical for AI agents. Many existing agent-memory systems follow an Ahead-of-Time (AOT) design, constructing memory before a specific request arrives. While this reduces online serving cost, such request-agnostic memory construction can discard fine-grained information that later becomes important. To address this limitation, we propose Just-In-Time Agent Memory (JAM), a trainable framework for query-conditioned context construction at runtime. A Memorizer preserves complete raw histories in a hierarchical page-store with compact navigational summaries, while a Researcher iteratively retrieves, inspects, and integrates evidence for each request. To train these memory-use behaviors, we introduce Memory-Gym, an evidence-grounded data synthesis pipeline covering nine task types across six domains, and optimize the Researcher through verified-trajectory supervised fine-tuning followed by Hint-guided Group Relative Policy Optimization. We demonstrate the effectiveness of JAM across a variety of benchmarks on agent memory and long-context processing, where it achieves stronger task performance than AOT-style memory systems while remaining substantially more efficient than prior trained agentic memory approaches. To support reproducibility and future research, we release our anonymized source code at https://github.com/VectorSpaceLab/general-agentic-memory.
阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-28把人格融入双通路智能体记忆PersMem将人格映射为记忆处理参数,控制情感评估与记忆保留检索;摘要概述为四个步骤PersMem: Internalizing Personality into Dual-Pathway Memory for LLM Agents
The profile of a role-playing agent usually depends on the pre-defined personality in a system prompt, whereas its memory processing pipeline, including prioritisation of stored memories and subsequent retrieval, remains independent of this personality. This separation causes the agent's memory processing to be inconsistent with the pre-defined personality, and makes it difficult to validate whether agent behaviours follow this personality. In this paper, we propose Personality-Integrated Memory (PersMem), which integrates personality into the agent's memory processing pipeline, making it consistently personality-dependent. PersMem processes memory using four steps, where the personality is mapped to operation-specific parameters controlling: (i) affective appraisal annotating emotion states of the user input; (ii) retention of previously stored memories along with the current input; (iii) passive affect-driven memory retrieval exploring memories similar to user input in semantics and personality-guided emotions; and (iv) active goal-driven memory retrieval that refines and selects passively retrieved memories for the reply. Consequently, consistency with the pre-defined personality can be examined by inspecting memory-processing traces during human-agent interactions. We evaluate these personality-dependent differences in attachment and Big Five settings. PersMem exceeds the chance baseline for four-way attachment classification by 23.1 percentage points. In Big Five dialogue comparisons, PersMem achieves 67.5% accuracy, 6.7 percentage points above a baseline using uniformly sampled memories. On CoSER, PersMem achieves an average score of 66.13, with scores of 69.33 for Character Fidelity and 84.33 for Storyline Quality. Together, these results show that PersMem produces distinguishable personality-related memory-processing patterns.
阅读 arXiv 原文
个人知识与本体 · 4/30 · 2026-09-28会话智能体的说话人索引记忆Stashbird以来源溯源组织情节、语义与社区摘要,报告大幅降低摄入提示词tokenStashbird: Efficient Speaker-Indexed Memory for Conversational Agents
AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episodes to derived memory state through explicit provenance. Stashbird organizes memory into episodic records, semantic relations, community summaries, and persisted graph state, with lifecycle operations for incremental updates and episode-level deletion. We evaluate question-answering accuracy and model-facing workload across four long-term memory benchmarks. On LoCoMo, Stashbird uses 76.4x fewer ingestion prompt tokens than Graphiti. Compared with reproduced Hindsight on the same benchmark, it uses 8.1x fewer retrieval prompt tokens, with accuracy 1.6 percentage points lower. It achieves higher accuracy than Hindsight on LongMemEval-S and GroupMemBench and comparable accuracy on EverMemBench.
阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-28选择原始轮次能替代抽取吗预注册研究在LoCoMo与LongMemEval上比较原始轮次选择与LLM抽取,称结果非劣于抽取When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model
Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-extraction memory (one-sided 95% bound -3.0 points against a -5-point margin). Blind human grading narrows the margin but does not change the result. Raw turns cost 3,061 times less to write, and the result holds with a second answer model. Within this study, reranking's gain shrinks as the budget grows. It adds 17.4 points on LoCoMo and 9.1 on LongMemEval when three of 30 candidates are kept. At generous budgets it adds 1.5 and 1.1, and extraction systems are more accurate. This suggests why published results disagree. At matched context, Jev selects as accurately as an LLM reranker (non-inferiority bound -2.0) at a third of the latency, and more accurately than a multi-call graph traversal. Reranking lowers correct abstention. Plans, code and graded answers are released.
阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-28追踪记忆构建顺序的影响RoutePrism用同源两次不同顺序构建记忆并对比,恢复单条记录可挽回大幅精度损失RoutePrism: Tracing Construction Order Effects in Agent Memory
Processing the same records in a different order can discard different evidence, yet endpoint accuracy alone cannot reveal what changed or whether it mattered. We introduce RoutePrism, a diagnostic protocol that builds memory twice from the same source pool in two processing orders, then traces which sources, compiled contexts, and answers differ. Because record content, timestamps, policy, and the answer model all stay fixed, any observed difference is localized to the memory construction step. A matched four-condition intervention tests whether a record displaced by reordering actually carried task-relevant evidence: restoring that single record recovers over 60 percentage points of lost accuracy, while substituting a non-supporting record of equal length does not. We evaluate the protocol on PersonaMem-32K (63 primary queries, 29 users) and 470 LongMemEval-S questions with histories spanning 38 to 62 sessions, replicating the core intervention across five answer models. Survivor selection, defined as the choice of which record a cluster retains, drives most source-level changes, while different memory policies (compaction, bounded recency, MemoChat-style summarization, A-MEM) produce distinct failure signatures at the source, context, and metadata layers.
阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-28记忆投毒攻击的严重度度量提出反事实记忆后悔指标与MemHarm,通过离线配对损失筛选稀疏语义编辑并给出类内认证From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents
As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior. Persistent-memory attacks are typically evaluated by whether they succeed, yet successful attacks can leave persistent states with substantially different downstream consequences. We study this severity as a distinct attack-design objective and formalize it with counterfactual memory regret (CMR), the paired increase in expected downstream loss relative to clean memory. We introduce MemHarm, which predeclares a finite class of sparse, grounded semantic edits, evaluates candidates through the normal agent memory interface using offline paired-loss feedback, and certifies resolved selections within that class. Compared with attack-success optimization, CMR-guided selection produces substantially larger downstream loss while retaining most of the success-rate gain. Across two agent benchmarks and diverse memory designs, MemHarm attains the highest CMR point estimates among the evaluated general attacks on identical support. Factor-removal interventions link this harm to the selected semantic factor, and native-agent deployments verify the write-to-fresh-process attack path.
阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-26智能体记忆固化的可治理性四阶段研究关注记忆固化决策质量与可信度,报告未达阈值的负结果与防泄漏协议The Epistemics of Agent Memory: Measuring, and Governing, the Consolidation Decision in Long-Horizon LLM Agents
Long-horizon LLM agents must convert accumulated experience into durable memory, deciding what to keep, compress, abstract into reusable skills and rules, or forget. We report a four-phase research program on this consolidation problem whose central finding is a shift in what is measured: from how much an agent remembers, to whether its consolidation decisions are any good, to whether those decisions can be trusted. Phase 1 learns episodic boundaries from agent traces by downstream utility; an honest near-miss (oracle correlation 0.691 vs a 0.70 bar) whose lasting output is a three-gate anti-leakage protocol. Phase 2 learns when to promote experience and to which abstraction level under a token budget, achieving a verified +22.7% task-success improvement with 7x compression, but exposing a degenerate-forgetting failure and a distribution-shift failure mode we name lambda-prevalence coupling. Phase 3 introduces ConsolidationBench, an oracle-by-construction benchmark that scores consolidation decisions against a known optimum on three non-circular axes; production retrieval systems retain information yet score zero on cross-level transfer. Phase 4 introduces governed consolidation: the decision wrapped in poison-resistance, reversibility, and auditability guarantees with a quality gate. Governance is statistically distinct from the quality score ($r^2 = 0.43$; partial $r = 0.27$; identical-quality policies differ threefold in governance), so the contribution survives independently of the metric's external validity. On that question we report a resolved negative: after a graded-reuse redesign removed a structural ceiling, a two-benchmark study with 2,532 real answer cells finds the quality score does not predict real transfer accuracy (pooled Spearman $ρ= -0.24$, n = 12, CI spanning zero). An adversarial self-critique pass cleared the final claim set with zero surviving overclaims.
阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-26异构记忆提供者的路由管理评估13种记忆方法后,MemAgent将记忆视为路由问题,按内容选择检索与存储来源MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents
Current agents remain largely stateless across tasks, limiting their ability to continually improve from prior interactions and making memory essential for long-horizon agentic behavior. Existing memory methods seek to reuse past experience, but most rely on a single memory representation (e.g., trajectories, reflections, skills, structured knowledge) whose effectiveness varies across task distributions. Rethinking this design space, we evaluate 13 memory methods and find that no single method generalizes across benchmarks, revealing the potential of managing heterogeneous memory providers. We formulate agent memory as a routing problem in which a memory agent decides which memory provider to retrieve from, whether to inject short-term memory, and which providers should store the resulting experience. Based on this perspective, we propose MemAgent, featuring a content-aware routing architecture and a training-data synthesis pipeline. The routing architecture combines content-aware probing before retrieval, short-term memory gating during execution, and selective multi-provider storage, while the training pipeline synthesizes phase-specific supervision for routing decisions. Across GAIA, WebWalkerQA, and xBench-DS, MemAgent improves average accuracy by 10.0% and outperforms every individual memory method across all three benchmarks. These gains come with less than 0.3% routing overhead and a 12% reduction in average task steps.
阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-26带检索分数界的有损记忆压缩LAM用确定性去重给出分数扰动界,报告在600条轨迹上删除约22%观测tokenLAM: Efficient Lossy Agent Memory Framework With A Retrieval-Score Error Bound
Agent memory grows as agents read inputs, reason, and call tools. Longer histories increase inference cost and eventually exceed the context window. LLM-based summarization reduces this history but adds latency and provides no explicit bound on information loss. We propose LAM, a Lossy Agent Memory system with three components: a deterministic deduplication rule with a substitution bound on retrieval scores - a bound on score perturbation, not a certificate of unchanged ranking; a memory manager that preserves the cached prefix and overlaps compaction with inference; and a performance model that estimates compaction costs before deployment. On 600 agent trajectories, LAM removes 22.47% of observation tokens while retaining 99.984% of the measured gold-patch evidence. At a fixed deletion set, the performance model predicts a 71.4x-91.6x end-to-end speedup from removing records before prefill instead of deleting them from a prefilled context. That benefit comes from the schedule rather than the rule and applies to any prefix-preserving test.
阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-25把智能体记忆做成一等中间件立场论文主张记忆应成为可插拔中间件,提出双向可插拔、多租户隔离等六项挑战Memory as Middleware for Self-Improving AI Agents
AI agents are stateless across sessions by default and therefore operationally amnesic: each session begins with little durable knowledge of prior failures, repairs, preferences, or successful strategies. As a result, agents repeat the same mistakes and discard hard-won experience. The dominant fix is \emph{bespoke memory}---retrieval, persistence, and learning logic hand-wired into one agent and bound to one storage engine. This creates a fragmented landscape where memory cannot be swapped, shared, isolated, or reasoned about independently of the agent that owns it. We argue that this is a middleware problem: agent memory deserves a first-class, pluggable layer, just as data access, messaging, and persistence each became middleware concerns. We develop this vision through six systems challenges: two-sided pluggability, host-native interposition, multi-tenant isolation, write-path consistency, federated sharing with provenance, and lifecycle governance. We present ALTK-Evolve, a reference implementation of memory middleware for self-improving agents, and use it to motivate a broader research agenda for future memory middleware.
阅读 arXiv 原文
个人知识与本体 · 6/30 · 2026-09-25多跳智能体记忆的拓扑巩固EngramRAG结合唤醒与梦境两态,用使用加权PageRank与拓扑衰减缓解多跳检索问题EngramRAG: Dynamic Usage-Weighted Topology and Synaptic Consolidation for Multi-Hop Agentic Memory
As autonomous LLM agents are deployed across multi-session environments, conventional memory architectures suffer from Associative Blindness (inability to traverse multi-hop relational dependencies), Scaffolding Amnesia (temporal decay evicting core persona invariants), and Static Topology Stagnation (immutable graphs ignoring usage dynamics). Grounded in Complementary Learning Systems (CLS) principles, we propose EngramRAG, an adaptive memory architecture coupling a low-latency Waking State reflex with an asynchronous background Dreaming State consolidation cycle. EngramRAG introduces: (1) Usage-Modulated Personalized PageRank (U-PPR), where transition probabilities adapt via Hebbian plasticity to promote persistent entities into high-centrality Epistemic Macro-Hubs; (2) Consolidation-Activated Topology Decay (CATD), which scales retention half-life by topological load-bearing weight rather than wall-clock recency, protected by a cold-start grace period (N_grace >= 4); (3) Directed SUPERSEDES DAG filtering to suppress obsolete state during fact mutations; and (4) Triple-source hybrid retrieval fusing dense vectors, BM25, and U-PPR via dynamic Reciprocal Rank Fusion (RRF). Evaluating on all 1,982 QA pairs across 10 long-term conversations in the LoCoMo benchmark, EngramRAG achieves +38.9% relative improvement in Recall@5 (53.21% vs. 38.29%, p < 0.001) and +43.1% in MRR (0.4203 vs. 0.2937) over dense vector RAG, significantly outperforming Okapi BM25 (48.66%) and isolated static graph retrieval (8.50%). On temporal reasoning, EngramRAG reaches 62.33% Recall@5 (+16.67 points over dense vectors). In controlled mutation tests, SUPERSEDES suppresses split-brain hallucinations from 70.0% to 0.0%, while 90-day simulations show 100.0% scaffolding retention under a 26.21ms interactive retrieval reflex.
阅读 arXiv 原文
个人知识与本体 · 4/30 · 2026-09-25世界模型验证的反事实记忆COUNTERMEM在失败动作后借助可执行世界模型评估替代动作,构建经验证的反事实记忆COUNTERMEM: World-Model Verified Counter-Factual Memory for Language Agents
Existing agent memory frameworks mainly create memory through an agent's interaction with the factual world, e.g., remembering feedback from actions taken to improve performance on future tasks. However, these frameworks seldom ask the "what if" question during memory construction: what if a different action had been taken, would the feedback have changed, and how could this feedback become useful memory? Obtaining such feedback directly in an active environment can be expensive and can alter the state needed for comparison. In this work, we introduce COUNTERMEM, a reinforcement-learning framework for constructing and using verified counterfactual memory across tasks. After a failed action, COUNTERMEM evaluates local alternatives from a copy or reset of the original state using executable world models, such as tests, proof checkers, and solvers. It stores improvements with the original and corrected actions, checked outcomes, and conditions for reuse. A learned memory-use policy selects a retrieved record or skips memory to balance task success and interaction cost, while the base LLM remains fixed. Both memory and policy are frozen during held-out evaluation. We evaluate COUNTERMEM on 12 benchmark settings across six domains. With gpt-oss-120b, COUNTERMEM improves both ReAct and Reflexion on all 12 benchmarks across six domains, averaging a gain of 12.6 percentage points over their unaugmented versions. In the four-domain comparison across two backbones, task-run tokens decrease by 7.7-42.0%, excluding offline selector-training costs. Further analyses show that removing verification or persistent storage weakens the gains, while applying verified corrections to unsuitable decisions can reverse them. Code will be released upon acceptance.
阅读 arXiv 原文