形式化与程序验证(14 篇)
形式化与程序验证 · 7/30 · 2026-09-22Agentic-IC3:语义证明搜索在 IC3 中引入语言模型代理,用 RTL 设计信息引导泛化Agentic-IC3: Enabling Semantic Proof Search in IC3 Model Checking
IC3 is a state-of-the-art algorithm for hardware model checking that proves safety properties by incrementally constructing an inductive invariant consisting of a set of lemmas. Its effectiveness depends on generalization heuristics that identify useful lemmas and guide proof search. However, many leading IC3 hardware model checkers operate on lowered, bit-level representations, where high-level design relationships are difficult to exploit for generalization. Those operating at a higher level remain limited in exploiting high-level design structure and semantics. We present Agentic-IC3, built on Pono's word-level model-checking infrastructure, which integrates a language-model agent into IC3 to guide semantic proof search using register-transfer-level (RTL) design information. The framework exposes an agent-oriented interface to a persistent IC3 backend, allowing the agent to interact with an explicit, evolving proof state throughout verification. Across successive proof obligations, the agent relates intermediate proof states and solver feedback to the RTL and proposes high-level lemmas through both SAT and UNSAT generalization. Beyond generalization, the agent can introduce derived observation signals to express design relationships succinctly and obtain more informative feedback, and backtrack to revise proposals that lead to unproductive proof branches. The backend checks proposals before updating the proof state, preserving soundness and providing feedback for further reasoning. On a suite of 14 benchmarks spanning security information-flow verification and functional verification of communication protocols, processors, and functional units, Agentic-IC3 solves 10 cases within a one-hour timeout, including 4 unsolved by all three evaluated baselines: rIC3, Pono-IC3Bits, and A-IC3.
阅读 arXiv 原文
形式化与程序验证 · 7/30 · 2026-09-22LLM 生成可证完备的广义规划将 PDDL 语义保持地转成 Lean,由内核检查规划完备性证明Provably Complete Generalized Planning with LLMs
Generalized planning aims to compute a plan that solves all instances of a planning domain. Recent work has used LLMs to automatically generate and debug such generalized plans in the form of Python programs and achieved perfect test data coverage for several domains. However, whether these generalized plans are actually complete, i.e. solve all instances of the domain, could only be determined by manual evaluation. Here, we present an approach for automatically generating generalized plans in Lean together with proofs of their completeness relative to a specification of the domain constraints provided as input. We introduce a semantic-preserving PDDL-to-Lean conversion, and use an LLM to generate both the generalized plan and the formal proof that it solves every instance satisfying the domain constraints. The correctness of the completeness proof is determined by Lean's kernel. We evaluate our approach on 13 commonly used benchmark domains, using GPT-5.6-Sol as the LLM. For 12 of the domains we obtain generalized plans together with valid completeness proofs. This is a major advancement of the state of the art in automatic generalized-plan completeness proofs.
阅读 arXiv 原文
形式化与程序验证 · 6/30 · 2026-09-22SLED-IFV:硬件信息流验证分解以求解器校验的 LLM 引导流程,做功能简化与关系强化分解SLED-IFV: Solver-Validated LLM-Guided Decomposition for Scalable Hardware Information-Flow Verification
Formal hardware information-flow verification (IFV) provides strong guarantees against secret-dependent timing and control behavior, but often scales poorly on realistic RTL. We identify two recurring proof barriers in self-composed IFV: implementation complexity, where proof-hard datapath logic dominates even though the property needs only a compact boundary relation, and relational inductive complexity, where the proof depends on cross-copy public-control facts that the backend prover does not infer efficiently. To address them, we introduce two semantic proof decomposition forms: functional simplification, which replaces a proof-hard RTL region with a validated over-approximate summary, and relational strengthening, which exposes and proves the cross-copy relations needed for induction. We further present SLED-IFV, a solver-validated LLM-guided flow that automates the selection of these forms and their concrete targets. Given a self-composed miter and an oracle-free decision sheet, the LLM proposes a decomposition, then materializes it into proof artifacts under controller checks. The controller compiles the checked artifacts into proof obligations, and the formal verification backend remains the sole authority for acceptance. Across nine nontrivial benchmarks constructed from real RTL, SLED-IFV achieves up to 603x solver-only speedup and converts two 12-hour timeouts into completed proofs. The closed-loop flow produces verifier-accepted decompositions for all cases.
阅读 arXiv 原文
形式化与程序验证 · 6/30 · 2026-09-22定理证明搜索中生成器的直接优化把计算对齐训练扩展到树搜索,并另提搜索无关的预算分配损失Direct Optimization of Generators for Search in Automated Theorem Proving
Fine-tuned Large Language Models (LLMs) significantly advance Automated Theorem Proving (ATP), but are often deployed as guiding policies within tree search rather than for single-attempt generation. Recent work shows cross entropy is suboptimal for an LLM used in flat search strategies such as aggregation or filtering and that work has developed new loss functions to correct this misalignment. Extending this alignment to tree search is more challenging: proof discovery depends on exploration and recovery through off-trace states that supervised demonstrations do not reveal. We extend Compute-Aligned Training (CAT) to this setting through an abstraction of policy-guided search, deriving tractable, trace-supported losses. Alongside these search-aware losses, we introduce a search-agnostic uniform-allocation (UA) loss that accounts for the budget without specifying the specific search. Both induce scalar weights on per-tactic cross-entropy gradients. We characterize how off-trace behavior affects the search-aware weights, including conditions for vanishing approximation error at large budgets. On a Lean benchmark, both approaches achieve higher observed proof-success rates than cross-entropy across six search strategies, with strong results from a single shared UA adapter. Budget sweeps show larger gains over cross-entropy at 16 than at 256 expansions, implying CAT scales with test time compute.
阅读 arXiv 原文
形式化与程序验证 · 6/30 · 2026-09-21分形决策图:边缘实时分诊摘要自称零权重张量运行并给出若干提升,但证据仅为自述Universal Fractal Natural Language Decision Map: Real-Time Edge Triage Across Heterogeneous Domains
Deploying Large Language Models for runtime operational triage incurs prohibitive latency (>100-500 ms), high VRAM requirements (>4-8 GB), and excessive energy dissipation. Extending Mandelbrot Fractal Neural Synthesis (Dagli et al., 2026), this paper presents the Universal Fractal Natural Language Decision Map, realized via the werr machine-native edge reflex runtime and the production answerr platform (https://answerr.me). Operating entirely without stored weight tensors (0 Bytes VRAM), the engine synthesizes deterministic decisions---noul (Boolean), choice (categorical), and score (ordinal)---by dynamically modulating 24-byte coordinate seeds along the chaotic boundary of the Mandelbrot set and evaluating multi-scale escape dynamics. Drawing inspiration from biological System-One reflex arcs, the engine introduces: (i) an Auto-Seed Router with domain projector Phi_D yielding a +28.8% accuracy gain over linear baselines; (ii) an Information-Theoretic Semantic Token Damping Filter (T_desc = 0.045) insulating against prompt injections (0.0% empirical bypass; 95% Wilson CI: [0.0%, 27.8%]) while pruning iterations by 45.8% (accelerating throughput 2.5x to 3.31 ms latency); (iii) a Multi-Scale Harmonic Tripod Fusion; (iv) a Coupled Margin Expansion Operator (Pitchfork Bifurcation Offset); and (v) a Cyclic Z/9Z Modular Resonant Grid Discretization based on the closed sub-ideal {0,3,6} (Lean 4 Mathlib ZMod 9), reducing FLOPs by 68.4%. Evaluated on JevBench (N=231), werr achieves 100.00% TypeSafe compliance and 81.65% calibrated accuracy with 7.08 ms median latency. We provide an OpenAI-compatible API and demonstrate deployment on 32-byte EVM smart contracts via the open-source werracle on-chain oracle (21,438 gas).
阅读 arXiv 原文
形式化与程序验证 · 4/30 · 2026-09-21Lara:面向自主科学的机检语言把研究论证编码为可执行工件,支持自动校验与跨论文桥接Beyond Natural Language: An Agent-Native Language for Autonomous Science
As autonomous AI agents take on every stage of scientific inquiry, research output is expanding far beyond human review capacity. Yet scientific communication still relies on natural-language prose: an informal medium prone to ambiguity, hidden assumptions, and untracked limitations that machines cannot reliably audit. We introduce Lara, a machine-checkable language and protocol for checking and revising support for research claims. By turning research arguments into executable artifacts, Lara provides an epistemic kernel for autonomous science: it enables automated validation pipelines for research agents, lets declared bridges connect arguments across papers into an auditable network, and allows both humans and machines to recheck the standing of an encoded claim in milliseconds. In a Lara program, authors explicitly declare their claims, supporting evidence and assumptions, and known objections or limitations. A lightweight, deterministic checker adjudicates these interactions, assigning each claim a reproducible status: "justified", "defeated", "contested", or "gap", which marks a claim whose support is incomplete and locates the unanswered question. Case studies cover empirical review, a philosophical debate without measurements, and the loss of support when an assumed axiom is withdrawn. We establish the metatheory of claim checking and cross-context argument transport, and mechanize the semantic guarantees in Lean 4 (roughly 117,000 lines), leaving three arguments on paper. The audited public metatheory is "sorry"-free and uses only Lean's three standard axioms; some executable examples additionally trust native evaluation.
阅读 arXiv 原文
形式化与程序验证 · 1/30 · 2026-09-21Lean Pool:AI 维护的形式化数学库摘要极短,仅称由 AI 代理扩增与优化,细节证据有限Lean Pool: An AI-Maintained Archive of Formalized Mathematics
Lean Pool is a repository of formalized mathematics. It is grown, maintained and optimized by AI agents.
阅读 arXiv 原文
形式化与程序验证 · 7/30 · 2026-09-21规范驱动的 AI 辅助开发生命周期结合 AGENTS.md 与分阶段技能文件,在课程中以学生问卷评估A Lean and Spec-Driven AI-Assisted Software Development Lifecycle for Applied AI Education: The AI-SDLC Approach
AI coding agents increasingly support software development beyond code completion, including planning, implementation, testing, and repository-level task execution. Their practical use, however, often remains only weakly connected to established software engineering practices. The aim of this work is to develop and evaluate a lightweight, spec-driven lifecycle for governed agentic software engineering. The lifecycle combines established software engineering practices with repository-local guidance through specifications, AGENTS.md, and phase-specific agent skill files. The approach was developed in the context of the FHNW course AI-assisted Software Development and applied by students to business-oriented software use cases. Its educational and practical applicability is explored through a student survey combining closed rating items with open-ended questions. The contribution of this work is a process-oriented framework that enables AI coding agents to operate with bounded autonomy within an explicit, reviewable, and test-oriented software development lifecycle.
阅读 arXiv 原文
形式化与程序验证 · 10/30 · 2026-09-18SWE-Proof:机检证明的议题修复为已知正确补丁写规范并转为形式验证任务,减少对测试集依赖SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?
Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check correctness with held-out test suites, which are inherently incomplete and increasingly susceptible to memorization. Formal verification avoids both problems, but existing work covers only standalone tasks whose specifications are given as input, not real issues, which touch large repositories and state intent in vague natural language. We present Benchproofer, a pipeline that turns a coding task with a known correct patch into a formally verified one: it writes a specification for the new code, summarizes the existing functions that code calls with axioms, and admits an instance only after mechanical and adversarial gates agree. Applying it to SWE-bench Verified yields SWE-Proof, 500 real issues whose correctness is formally verified rather than tested, and it extends to SWE-bench Pro. Evaluating Claude Opus 4.8, we find that verification catches what tests miss: a quarter of test-passing patches admit counterexamples, which a structured natural-language specification does not fix, while a correct formal one lifts resolution from 85% to 95%. Writing that specification is the hard part: an agent that must write its own gains nothing over an unaided baseline, and only 56% of those specifications pass our audit. The usual failure is faithfulness, a specification that constrains part of the required behavior and leaves the rest free. Specification quality still tracks the outcome, failing on 92% of unresolved instances against 51% of resolved ones, making faithful specification synthesis a concrete open problem.
阅读 arXiv 原文
形式化与程序验证 · 4/30 · 2026-09-17Gardam 格无唯一乘积的机检证明用 Lean 4 给出显式有限子集反例,证明经内核检查A Machine-Checked Proof that Gardam's $\tilde{A}_2$ Lattice Does Not Have Unique Products
Kaplansky's zero-divisor conjecture asserts that the group ring of a torsion-free group over a field has no zero divisors. It holds for every group with the unique-product property, so a counterexample can only come from a torsion-free group without unique products. In lectures in 2021, Gardam announced that the torsion-free $\tilde{A}_2$ lattice $Γ= \langle a, b \mid a b a^2 b^{-1} a^2 b^{-2}, a b^3 a b^4 a^{-1} b \rangle$ does not have unique products and presented it as a new candidate group: it has property (T), and the known methods for proving the conjecture do not apply to it. To our knowledge, no proof of the announcement has been published. We give a proof checked by the Lean 4 kernel and stated against Mathlib's UniqueProds class. The witness is an explicit pair of finite subsets with $|A| = 32$ and $|B| = 28$ in which each of the 896 products coincides with another product. For 658 products the certificate is an identity in the free group; the other 238 certificates are explicit products of conjugated relators, 970 conjugates in all, checked by free reduction. A homomorphism onto $\mathbb{Z}/42$ shows that each pair $(u,v)$ differs from its partner $(u',v')$ as a pair of group elements, which is all the theorem requires. Together with a homomorphism onto the alternating group $A_4$ it also shows that the listed words are pairwise distinct, so the sets have exactly 32 and 28 elements. The witness and certificates come from an untrusted search program and are re-checked by Lean. The development uses only the axioms propext, Classical.choice and Quot.sound, with no sorry and no native_decide. The mathematical statement is Gardam's. To our knowledge this is the first verification in a proof assistant of a unique-product failure in a torsion-free group; torsion-freeness of $Γ$ is taken from Gardam and is not formalized here.
阅读 arXiv 原文
形式化与程序验证 · 0/30 · 2026-09-17大模型回答中的语言地缘差异对同一议题以多语言提问,回应倾向随提问语言不同而不同Geopolitical Divisions Across Languages in Large Language Models
People increasingly turn to AI chatbots for news and explanations of world events. But do they receive the same political answers when they ask in different languages? Here we show that the language of a question can change how the same AI systems assess the war in Ukraine. We ask GPT, Claude and Gemini to evaluate twenty statements about the war in 112 languages, collecting 67,200 responses. The balance between Russia-leaning and Ukraine-leaning responses differs across languages. When we group responses by countries' official languages, they follow a pattern resembling worldwide political divisions: relatively more Russia-leaning answers correspond to more favourable public views of Russia, less support for Ukraine in United Nations votes, and less aid to Ukraine. The broad pattern recurs across all three models and remains when individual statement pairs are removed. Our findings suggest a possible route through which information warfare may shape the text used to train AI models, which may in turn spread geopolitical biases.
阅读 arXiv 原文
形式化与程序验证 · 6/30 · 2026-09-17长周期自动形式化核心定理以共享蓝图协调 AI 证明代理,完成 Lean 4 机检证明Long-horizon autoformalization of a core theorem underlying MIP* = RE
Landmark mathematical formalizations have taken specialist teams years to complete. We present FormalFlow, a system that coordinates AI proving agents under human supervision to address statement drift and proof composition in long-horizon formalization. Drawing on software engineering principles and practices, it uses a shared blueprint to guide nested planning, proving and review loops. Agents strengthen verification and review throughout formalization. We completed a machine-checked Lean 4 proof of the quantum soundness of the classical low individual-degree test, a core theorem underlying MIP* = RE. Developing the proof took 63 days; greater parallelism could further reduce this time. The final library contains 126,367 lines of Lean code, all generated by agents. The formalization corrects side conditions and intermediate errors while preserving the published final error bound under corrected assumptions. This work provides a verified foundation for quantum complexity and demonstrates a route to affordable verification of major research proofs by small teams.
阅读 arXiv 原文
形式化与程序验证 · 4/30 · 2026-09-17LLM 与 Lean 辅助的 LLVM 翻译验证生成结构化证明骨架,自动处理可由确定性方法解决的义务LLVM Translation Validation Automated with Large Language Models and Lean
LLVM is the cornerstone of modern compilers, but its subtle intermediate representation (IR) semantics make transformations error-prone and necessitate formal verification. Alive2, a state-of-the-art translation validator based on satisfiability modulo theories, has achieved substantial success in automating the validation of LLVM transformations. However, it still faces scalability limitations, does not support symbolic bitwidths, and offers only bounded guarantees for loops. In contrast, interactive theorem provers such as Lean can address these cases but require substantial proof engineering. In this paper, we present Trivet, a framework combining large language models (LLMs) and Lean for automated translation validation of LLVM transformations. Trivet generates structured proof scaffolds based on source and target functions, automatically discharges obligations amenable to deterministic reasoning, and delegates transformationspecific obligations to LLMs. It produces refinement proofs or counterexample-based refutations, with every successful verdict checked by the Lean kernel. On 148 LLVM transformations, Trivet verifies or refutes 147, leaving one invalid case unresolved. Successful cases include 60 loop-free transformations with symbolic bitwidths, 27 cases from a restricted class of loop-containing transformations, and 10 complex valid fixed-bitwidth cases on which Alive2 times out. Compared with an unscaffolded baseline, scaffolding enables 26 additional proofs. On cases solved by both configurations, it reduces mean proof time by 75.9% and mean monetary cost by 88%.
阅读 arXiv 原文
形式化与程序验证 · 7/30 · 2026-09-16MAGS:多代理自动形式化保安全以 Dafny 作验证中间表示,并用验证器反馈修复违规MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs
LLM coding agents now generate complex programs at a scale that makes thorough human review increasingly difficult, raising the risk of safety and security failures. Common approaches, including fuzz testing, static analysis, and LLM-as-a-Verifier, can detect many failures but struggle to cover all possible edge cases. Formal verification addresses this by providing machine-checkable guarantees over specified properties, but traditionally demands substantial manual specification and proof engineering. We introduce a unified multi-agent framework, MAGS, that generates executable programs with formal safety guarantees, using Dafny as a verification-aware intermediate representation where safety properties can be mechanically checked. MAGS formalizes and freezes human-audited APIs and safety requirements, translates generated code into Dafny, repairs violations using verifier feedback, and compiles verified programs back into executable code. We evaluate MAGS on 100 CUDA kernels, 100 terminal scripts, and 20 robotic-arm tasks. Across all 220 examples, it achieves a 100% success rate in producing programs with non-trivial safety guarantees against frozen specifications. Independent safety and functional evaluations further show strong performance across all three domains, while revealing failures when the auto-formalized semantics do not fully capture the target behavior.
阅读 arXiv 原文
软件工程与仓库智能(15 篇)
软件工程与仓库智能 · 4/30 · 2026-09-24软件工程自我效能感量表开发87 题五维量表在 527 名本科生试测,仅报告初步效度与信度证据Design, development, and preliminary validity and reliability evidence of the Software Engineering Self-Efficacy Scale (SESES)
The purpose of this research is to design, develop, implement, and provide preliminary validity and reliability evidence of the Software Engineering Self-Efficacy Scale (SESES). Framed by a conceptual framework using guidance in software engineering curriculum and concepts along with the notion of self-efficacy, we generated an initial item pool of n = 87 items to operationalize and measure software engineering self-efficacy among undergraduate computing students. The conceptual framework traces five dimensions: 1) Requirements Engineering, 2) Teamwork and Collaboration, 3) Software Quality Management, 4) Software Design and Architecture, and 5) Software Agile Methodologies. We pilot tested the SESES with n = 527 undergraduate computing students who had completed a software engineering course in the current semester or a previous academic semester. We employed Exploratory Factor Analysis (EFA) with the Principal Axis Factoring method and an oblique (Promax) rotation to examine the underlying structure of the SESES, resulting in the same five internally consistent latent constructs in the conceptual framework with minimal cross-loading and a simple structure in the pattern matrix, explaining approximately 57% of the variability in these data. Our findings suggest that software engineering self-efficacy is a multidimensional construct of five theorized and correlated, yet distinct latent factors. We unpack the limitations and delimitations of the research while exploring undergraduate computing students' software engineering self-efficacy using necessary domain-specific measurements.
阅读 arXiv 原文
软件工程与仓库智能 · 11/30 · 2026-09-24规范驱动的自动程序修复基准主张以可执行规范生成基准,替代静态语料并降低污染风险Specification-Driven Benchmarking for Automated Program Repair From Static Corpora to Executable Specifications
Automated Program Repair (APR) benchmarks have traditionally been constructed as static datasets whose characteristics are inherited from the defects they contain. While this paradigm has enabled decades of progress, finite corpora provide limited experimental control, become increasingly susceptible to contamination as they are reused, and cannot be systematically regenerated or adapted as evaluation requirements evolve. We propose specification-driven benchmarking, a paradigm in which benchmarks are defined by executable specifications and realized through benchmark generation. The specification explicitly declares the intended properties of the benchmark (including program context, fault taxonomy, difficulty, validation strategy, and corpus constraints) while a generation pipeline realizes those requirements through independent generation, validation, and corpus management components. We develop the conceptual foundations of this approach by introducing a taxonomy of benchmark specification dimensions, establishing how each specification dimension maps to deterministic architectural responsibilities, and arguing that independent validation is a structural requirement for trustworthy benchmark generation. An end-to-end example illustrates how specification choices propagate through the pipeline to produce benchmark instances whose properties are independently verifiable. By treating the benchmark as an executable specification rather than a static dataset, the proposed paradigm shifts benchmark construction from artifact curation to declarative experimental design.
阅读 arXiv 原文
软件工程与仓库智能 · 0/30 · 2026-09-23需求演化差异可视化的 LLM 流程以语义图快照做并排比较,摘要简短,可核证据有限LLM-Assisted Workflow for Structural Difference Visualization in Evolving Software Requirements
This paper presents an LLM-assisted workflow for visualizing structural differences in evolving software require- ments. Implemented in the OntologyWeb environment, the work- flow represents baseline and current requirements as triple-based semantic graphs and supports side-by-side comparison of curated graph snapshots. The comparison view aligns matched entities and uses visual encoding to highlight structural changes.
阅读 arXiv 原文
软件工程与仓库智能 · 4/30 · 2026-09-22依赖更新 PR 的修复代理路由仅用创建时标题与元数据排序是否需修复代理,摘要给出一项 F1When Should Dependency Updates Invoke Repair Agents? A Lightweight Routing Study
Dependency-update pull requests are frequent and mostly routine, but a small subset requires non-trivial compatibility repair. Recent repository-level coding agents make such repair increasingly plausible, yet invoking them on every dependency update wastes model calls, CI time, repository context, and review attention. We frame this as a pre-agent routing problem: deciding which dependency-update pull requests should be escalated before downstream diagnosis or repair attempts. We introduce DepFixRouter, a lightweight router that ranks dependency updates by historical compatibility-repair likelihood using creation-time textual and metadata signals. On 497 labeled GitHub dependency-update candidates, only 72 require substantive repair. A creation-time-safe LinearSVC using only PR titles and bot/dependency flags reaches 0.488 repair F1 and captures 51.4% of repairs within the top 20% routed pull requests, improving calls per captured repair from 6.90 under route-all or random policies to 2.68. Retrospective full-history signals improve top-20% recall to 65.3%, revealing substantial hindsight leakage in pull-request histories rather than deployment-time routing utility. In a 60-case diagnosis-agent pilot, router-gated diagnosis reduces actual LLM calls by 66.7% and tokens by 66.1%, suggesting budgetaware escalation while measuring diagnosis rather than patch generation. DepFixRouter can serve as a lightweight escalation layer between routine dependency-update automation and expensive repository-level agents, enabling budget-aware maintenance without relying on retrospective repair evidence for deployment-time routing.
阅读 arXiv 原文
软件工程与仓库智能 · 7/30 · 2026-09-22AI 代理 PR 合并后的修复归属追踪已合并代理 PR 的后续修复,并与同期人工 PR 基线对比Who Finishes the Job? A Study of Follow-Up Fixes and Commit Authorship on AI Coding Agent Pull Requests
AI coding agents now author a large share of pull requests (PRs) merged into popular open-source projects. A merged agent PR is usually considered finished work; yet, prior studies have reported issues in agent code after the merge (e.g., code smells and static-analysis issues). However, little is known about how often a merged agent PR is fixed afterward, and who actually authors the fixing. In this paper, we follow 6,774 merged agent PRs across five AI coding agents (OpenAI Codex, GitHub Copilot, Devin, Cursor, and Claude Code) from the AIDev-pop dataset (open-source repositories with at least 500 stars) into their follow-up fixes, against a baseline of 5,044 contemporaneous human PRs from the same repositories. We link each merge to its candidate fixes, verify every candidate with human annotators and an LLM judge that matches human-level agreement (binary Cohen's Kappa=0.78 against a human-human K=0.77, Direct-fix precision 90%), and attribute the fixing work at the PR and the commit level. Our findings show that (1) merged agent PRs attract verified fixes at 1.62 times the odds of merged human PRs in the same repositories over the same period of time; (2) 69.6% of verified fixes in agent merges come from the same agent; and (3) 76.4% of the verified fix PRs are agent-authored throughout all commits. These results show that agents currently largely finish their own job, but their merges still require fixing more often than human merges.
阅读 arXiv 原文
软件工程与仓库智能 · 4/30 · 2026-09-22Galaxy 生态维护与支持的跨空间研究用议题、PR 与论坛讨论刻画维护关切及其跨空间关联Understanding Maintenance and Support in a Community-Driven Scientific Workflow Ecosystem: A Cross-Space Study of Galaxy
Galaxy is a widely used, community-driven scientific workflow system whose sustainability depends on continuous maintenance across its software, tools, workflows, infrastructure, documentation, and user-support ecosystem. However, maintenance knowledge in Galaxy is distributed across development and community-support spaces, making it difficult to understand what is maintained, how maintenance artifacts are resolved, and how user-facing concerns connect to repository-level development. We conduct a large-scale empirical study of Galaxy using 11,762 GitHub issues, 52,203 pull requests, and 6,235 Community Forum discussions. We characterize maintenance and support concerns, examine factors associated with resolution outcomes and resolution time, and investigate explicit and candidate connections among maintenance artifacts across these spaces. Using BERTopic modeling, we identify nine issue topics, 14 pull-request topics, and 14 forum topics, revealing a maintenance landscape spanning workflow execution, data management, tools and dependencies, infrastructure, testing, scientific resources, documentation, and user support. Resolution analyses show that coordination, diagnostic, contributor, automation, and engagement characteristics exhibit different associations with whether artifacts are resolved and how quickly resolution occurs. We further find limited explicit traceability between development and support spaces: 97.77\% of 16,426 resolved explicit relationships occur within GitHub, while only 294 connect GitHub artifacts with Community Forum discussions, despite additional semantic and technical relatedness across these spaces. Together, these findings characterize Galaxy maintenance as a distributed ecosystem-level process and identify opportunities to improve diagnostic reporting, lifecycle-aware triage, cross-space traceability, and the reuse of community-support knowledge.
阅读 arXiv 原文
软件工程与仓库智能 · 7/30 · 2026-09-22Python 跨操作系统可移植性问题重跑测试加议题分析,提出七类失败分类与诊断特征An Empirical Analysis of Cross-OS Portability Issues in Python Projects
While Python is designed as a cross-platform language, real-world applications encounter portability failures when deployed across different operating systems. We present the first large-scale empirical study of cross-OS portability issues in Python, analyzing 2,042 open-source repositories using two complementary approaches: systematic cross-OS test reexecution and manual analysis of GitHub issues. Our cross-platform testing of 500 projects reveals that 11.2% exhibit OS-dependent test failures. Through systematic analysis of 240 GitHub issues, we confirm 102 genuine portability problems spanning 95 additional projects. We develop a comprehensive taxonomy identifying 7 primary failure categories - with file/directory operations, process management, and library dependencies being most prevalent - along with 24 distinct sub-categories, 15 diagnostic signatures, and 4 systematic repair patterns. Our evaluation reveals that existing static analysis tools provide minimal support for portability detection, while large language models achieve 40-79% accuracy in identifying issues and 50-77% success in generating fixes when provided with structured guidance. Through 33 contributed pull requests, we demonstrate practical applicability and developer acceptance (17 merged, zero rejected) of our findings. This work establishes the first comprehensive baseline for understanding and addressing cross-OS portability issues in Python, providing actionable insights for developers, tool designers, and the broader research community.
阅读 arXiv 原文
软件工程与仓库智能 · 7/30 · 2026-09-21并行编码代理的语义协调基准同一测试分别跑单补丁与合并补丁,构造任务中干扰较常见Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development
Parallel coding agents can produce patches that work alone but fail when merged. This happens when one agent changes an interface or rule that another agent still relies on. We study these failures with stale, a benchmark for semantic coordination. Our evaluation runs the same tests on each patch alone and on their combination, counting only failures introduced by combining the patches. We use three tiers: synthetic tasks with controlled interface changes, pairs of merged pull requests, and constructed tasks that use real Django helpers. Among 834 runs on 417 mined Django pairs, only one showed interference after correcting the grading procedure. On constructed tasks using 12 Django helpers, interference occurred in 97% of runs. A message describing the completed concurrent change recovered 82% of runs. Reviewed pull requests may contain few unresolved parallel changes, even when agents fail on controlled tasks using real code. The constructed failure rates do not estimate how often these problems occur in practice.
阅读 arXiv 原文
软件工程与仓库智能 · 6/30 · 2026-09-19ChatGPT 增强修复的跨基准泛化在三套基准上比较增强方法,发现增益方向随模型与基准变化Towards the Generalizability of Leveraging ChatGPT in APR via Self-enhancing: An Empirical Study
Automated Program Repair (APR) increasingly relies on Large Language Models (LLMs). ChatGPT-enhanced APR uses techniques such as self-correction and autonomous agents to improve repair without modifying model parameters. Although these approaches report strong results on Defects4J and SWE-bench, the stability of enhancement gains across benchmarks remains under-explored. We evaluate three ChatGPT-enhanced APR methods on three representative, long-standing benchmarks. With GPT-3.5-Turbo, SRepair achieves a larger absolute gain on HumanEval-Java than on Defects4J, while SRepair and FixAgent without extrinsic information yield negative gains on BugsInPy. With GPT-5.4-mini, the evaluated methods achieve larger absolute gains on Defects4J than on HumanEval-Java, while gains on BugsInPy are non-negative but limited. We investigate benchmark-related factors through code transformations and benchmark-specific fine-tuning. Code transformations reduce enhancement gains on Defects4J, while benchmark-specific fine-tuning increases gains on BugsInPy. Directly supplying GPT-3.5-Turbo with error messages and triggering tests yields more correct repairs than the evaluated ChatGPT-enhanced APR methods on BugsInPy. These findings highlight the need to evaluate generalizability across benchmarks and models using multiple metrics, and suggest that directly providing repair-specific extrinsic information may be more effective than enhancement methods when their gains are limited.
阅读 arXiv 原文
软件工程与仓库智能 · 6/30 · 2026-09-18拉取请求为何沉默:完成障碍分析大量停滞 PR 与评审评论,归类贡献类型与停滞原因Why Do Pull Requests Go Silent? Uncovering the Barriers to Contribution Completion in Open-Source Code Review
Pull requests (PRs) underpin pull-based software development by enabling distributed code review and collaborative contribution in open-source projects. Yet many become inactive before integration and are eventually abandoned or closed, wasting contributor and maintainer effort. Although prior work has examined PR abandonment and review delays, less is known about the contribution types, discussion-level barriers, and post-stalling collaboration patterns associated with inactivity. We investigate which PR types most often stall, why inactivity occurs from authors' and reviewers' perspectives, and how stalling relates to later contributor and reviewer engagement. We analyzed 14,234 stalled PRs and 164,562 review comments from 19 popular GitHub repositories using stale-bot workflows. An LLM-based voting classifier categorized PRs by contribution type, while quantitative analysis was combined with qualitative coding of general and inline review discussions. Feature-enhancement and issue-fixing PRs formed the largest share, together exceeding 77% of classified stalled PRs. General comments linked stalling mainly to communication and coordination breakdowns, including missing interaction, delayed feedback, and unclear follow-up. Inline comments showed that inactivity does not always reflect disengagement: many PRs were blocked by technical or dependency issues, including failing checks, configuration problems, compatibility concerns, and environment mismatches. Only 39.56% of contributors later submitted another PR, and reviewer re-engagement with the same contributors was approximately 21%. PR inactivity is a socio-technical coordination problem involving communication, technical readiness, review ownership, and automation practices. We recommend type-aware triage, clearer review feedback, explicit ownership of next actions, CI blocker management, and cause-aware stale-bot interventions.
阅读 arXiv 原文
软件工程与仓库智能 · 3/30 · 2026-09-18文件排序对代码评审效果的影响挖掘大量多文件 PR,考察文件位置与后续修复变更的关联Does Order Matter? An Empirical Investigation into the Impact of File Ordering on Code Review Effectiveness
Modern code review is central to software quality, but its effectiveness depends on reviewer expertise, change characteristics, and how tools present changes. Most platforms display modified files alphabetically by default, although our prior work shows that developers find this ordering cognitively misaligned with how they understand multi-file pull requests. Whether these ordering-related attention patterns affect outcomes at scale remains unclear. We present a large-scale study of file ordering and review effectiveness, mining 330,343 multi-file pull requests and 756,814 file instances from 182 GitHub projects in five programming languages. We examine whether file position is associated with later bug-fixing changes, whether pull request size moderates this relationship, and whether reviewer attention, proxied by comments, aligns with latent bug outcomes. Results show statistically significant but modest associations among file position, pull request size, review activity, and latent bug likelihood. Latent bug rates rise from 56.7% at position 1 to 61.5% at position 30. Pull request size has a non-linear relationship with latent bug likelihood: pull requests of about ten files have the lowest risk, while very small and very large ones have elevated rates. Hurdle models show that attention is diluted as pull request size grows: each additional modified file reduces the odds of receiving any review comment by about 8.7%. These findings reveal an attention-effectiveness gap: visible review activity does not necessarily prevent defects. Alphabetical ordering is therefore not a neutral interface default, but a structural feature shaping attention allocation, review coverage, and confidence in review outcomes. We propose context-aware file ordering, dependency-aware grouping, risk-aware prioritization, and per-file coverage indicators to make review attention more visible and actionable.
阅读 arXiv 原文
软件工程与仓库智能 · 7/30 · 2026-09-17AdaRepair-Mem:自适应经验编排指出仓库级记忆检索的三点局限,提出覆盖感知的检索回退AdaRepair-Mem: Adaptive Experience Orchestration for Repository-Level Program Repair
Recent memory-augmented repository-level program repair methods reuse historical repair experiences to improve LLM-based issue resolution. However, our analysis reveals three limitations in existing repository-level memory retrieval. First, episodic memory is highly imbalanced across repositories, leaving low-resource repositories with little effective support. Second, more memory does not monotonically lead to higher repair success, suggesting that relevance, quality, and redundancy matter more than raw memory volume. Third, memory accumulation is phase-misaligned: repositories may contain many reproduction experiences but few patch or refinement experiences. To address these problems, we propose an adaptive experience retrieval framework for repository-level program repair. Our framework introduces coverage-aware retrieval, which falls back to cross-repository or repair-type-based memories when same-repository memory is insufficient; quality-aware selection, which ranks memories by relevance, historical utility, specificity, and redundancy; and stage-aware routing, which separates and retrieves memories for reproduction, localization, patch generation, patch refinement, and validation. Evaluated on SWE-Bench-Lite and SWE-Bench-Verified, the proposed framework improves repair performance on under-covered repositories, reduces noisy memory retrieval, and better supports failed-to-fixed patch refinement. Our results show that the key to memory-augmented repair is not simply accumulating more experiences, but retrieving the right experiences for the right repair context.
阅读 arXiv 原文
软件工程与仓库智能 · 9/30 · 2026-09-16LLM 程序修复实验设置的规范指出基准与修复数不足以规定任务,提出显式描述框架Experimental Settings in LLM-Based Program Repair: A Study of Inputs, Tool Access, Feedback, and Validation
Evaluations of automated program repair (APR) systems commonly report the benchmark, the number of repaired defects, and the tests used for final patch validation, but these items no longer fully specify the repair task presented to a system. Recent LLM-based systems differ in the information supplied before repair, the repository and testing operations permitted during repair, and the feedback returned after unsuccessful attempts, allowing the same benchmark to instantiate substantially different repair tasks ranging from localized patch generation to repository-level diagnosis and iterative repair. We present a framework for explicitly specifying the experimental settings associated with reported APR results. We analyze reported experimental settings from systems evaluated on Defects4J and SWE-bench and characterize each result by its task unit, fault-localization assumptions, initial input, tool access, repair-time feedback, final validation, and resource budget. Our analysis shows that benchmark identity alone is insufficient to reconstruct the evaluated task or determine the appropriate scope of comparison across reported repair rates. We therefore introduce a machine-readable schema for specifying each experimental setting to improve reproducibility and make the scope of cross-system comparisons explicit.
阅读 arXiv 原文
软件工程与仓库智能 · 4/30 · 2026-09-15OpenClaw 技能生态的增长与治理基于仓库历史与三次注册表快照,测度热潮后的留存与扫描After the Party: Growth, Governance, and Security Scanning in the OpenClaw Agent Skill Ecosystem
AI agents increasingly act through agent skills, i.e., natural-language instructions, that direct a host agent toward shell, network, credential, file, and process actions, and public registries distribute them at scale. In the first half of 2026, the OpenClaw AI agent went viral, and its public skill registry boomed: the observable stock nearly doubled in 91 days, and a majority of the listings visible in June were created in just two months. By the end of our study window, the wave had crested, and monthly listing creation and core-repository activity were falling from their spring peaks. This paper measures what the boom left behind, drawing on the OpenClaw Git history, its GitHub issues and pull requests, and three ClawHub registry snapshots. Attention is concentrated: the top 10% of skills received 46.93% of all downloads. No simple skill features (like size or download counts) remained a stable predictor of continued listing once creation cohort and skill age were controlled. Human scrutiny did not stay: 77.86% have zero stars and zero comments, while 85.06% of the readable skills carry privilege evidence. And automated cleanup is not ready: the three security scanners disagreed on 23,702 of the 61,990 skills they all cover. After human adjudication, weighted scanner sensitivity against the reference standard ranged from 21.67% to 61.06%. Governing fast-growing agent-skill registries cannot rely on simple metadata or single scanner scores; it requires robust, transparent measurement and independent validation.
阅读 arXiv 原文
软件工程与仓库智能 · 7/30 · 2026-09-15水下机器人视觉语言模型蜕变测试用多目标搜索找最少的图像变换以诱发错误预测Search-Based Metamorphic Testing of Vision-Language Models in Autonomous Underwater Robotic Software
Our industry partner focuses on quality assurance for industrial systems across multiple domains, including maritime systems, such as overwater vessels and autonomous underwater robots (AURs). Despite the strong performance of vision-language models (VLMs) in scene understanding, image captioning, and object recognition, their use in AUR software operating in underwater environments is underexplored. Therefore, in this context, it is important to evaluate the quality of VLMs for integration into AUR software and, so, automated software testing tools are needed to assess their suitability and improve their dependability. To this end, we propose a search-based metamorphic testing approach (MetaVLM) that identifies a minimal set of transformations on underwater images to induce incorrect model predictions, thereby revealing VLM failures. We employ NSGA-II as a multi-objective search algorithm and evaluate it over open-source VLMs, BLIP and CLIP, against a random search baseline. Results demonstrate the strengths and limitations of each VLM in the context of AUR software systems. Based on the results, we derive lessons for software engineering practitioners and researchers working on quality assurance of VLM-based software systems.
阅读 arXiv 原文
代码质量与优化(7 篇)
代码质量与优化 · 3/30 · 2026-09-24AI 辅助代码修复的社会技术瓶颈15 天工业 C++ 仓库个案,考察持续集成、评审与团队协作约束Orchestrating AI-Assisted Code Remediation: Socio-Technical Bottlenecks in a Large Industrial Repository
Background: Code degradation in large, long-lived codebases is costly to remediate through manual refactoring and opportunistic clean-ups. LLM-based coding assistants can perform mechanical remediation at scale, but their impact on industrial workflows is underexplored. Objective: We investigate how massive AI-assisted code remediation affects build-on-commit continuous integration (CI), code review, and team coordination in a large industrial repository, and which socio-technical bottlenecks constrain such remediation when source editing becomes cheap through AI assistance. Method: We report on a 15-day exploratory single-case field study in which an experienced developer used a command-line AI coding buddy to remediate widespread issues in a closed-source industrial C++ repository. We triangulate Gerrit metadata with a developer diary and team chat, analyzed through descriptive statistics and qualitative coding. Results: AI-assisted remediation rapidly generated hundreds of commits touching thousands of lines, saturating CI and reviewer attention. Naïve per-file commits overloaded build-on-commit CI; Switching to directory-based batching and capping the number of files per change restored throughput, but still required explicit review solicitation, negotiation of acceptable commit granularity, and iterative follow-up to resolve build and static-analysis failures. Conclusion: When mechanical editing is cheap, CI capacity, review effort, and change orchestration become primary bottlenecks. Sustainable AI-assisted remediation in very large repositories requires deliberate control of commit, review, and CI batch granularity and treating semantic change sets, such as ``fix all instances of warning X'', as first-class units of work that can be sliced differently for developers, reviewers, and CI.
阅读 arXiv 原文
代码质量与优化 · 3/30 · 2026-09-23AST 自动消除 Java 跳转语句引入辅助布尔变量重构控制流以保持语义,为抽取方法铺路AST-Based Automated Elimination of break and continue Statements in Java Code
This work presents the development of an automatic refactoring tool for Java code built on top of the Eclipse JDT API. The proposed approach transforms control structures containing break and continue statements within different types of loops into semantically equivalent constructs that avoid their explicit use. To achieve this, auxiliary boolean variables are introduced to restructure the control flow while preserving the original program behavior. The main objective of this transformation is to improve code structure and enable the application of subsequent automated refactorings, particularly those based on the Extract Method operation, which are typically restricted by the presence of jump statements. The implementation relies on the analysis and rewriting of the Abstract Syntax Tree (AST), ensuring semantic equivalence in all addressed scenarios. The tool was validated through 54 manually designed test cases and 151 units tests, all of which produced satisfactory results. In addition, it was applied to 139 methods from seven open-source projects, generating code without compilation errors and preserving the original behavior as verified by the projects' test suites. The results demonstrate that the proposed approach safely automates the restructuring of code containing break and continue statements, facilitating further evolution and structural analysis.
阅读 arXiv 原文
代码质量与优化 · 6/30 · 2026-09-20VSpector:RISC-V 规范驱动查错用自然语言规范直接检查 RTL 是否符合规则,无需参考模型VSpector: Specification-Driven Bug Detection for RISC-V CPUs
Detecting RTL design bugs in open-source RISC-V CPU implementations is critical for ensuring system reliability. Traditional detection approaches inherently rely on predefined artifacts. In this paper, we leverage the official,natural-language RISC-V specifications as an effective information source for bug detection. We present VSpector, a specification-driven bug detection pipeline that directly checks whether CPU register-transfer level (RTL) implementations adhere to official specification rules, without requiring specialized construction of reference models, formal properties, or custom bug patterns. To resolve the key technical trade-off between broad context scope and model reasoning accuracy when using Large Language Models (LLMs), VSpector employs a stepwise context refinement scheme across a four-stage pipeline: rule extraction, implementation localization, candidate identification, and sequential violation auditing. We evaluate VSpector on two industrial-strength RISC-V CPUs, CVA6 and XiangShan. Out of 217 reported candidates, manual inspection confirmed 148 true violations, representing a 68.2% precision. These violations correspond to 73 distinct bugs, including 42 previously unknown bugs. In our comparative experiments, DiveFuzz, a state-of-the-art CPU fuzzer, detected none of these new bugs during 24-hour runs per CPU. All 42 new bugs have been reported upstream, with developers already fixing 19 and confirming an additional 11 (30 in total), demonstrating that specification-driven auditing is a practical and complementary strategy for CPU bug detection.
阅读 arXiv 原文
代码质量与优化 · 0/30 · 2026-09-17闭环机器人软件的习得与迁移把验证选出的闭环实现作为可复用经验存档,用于新任务生成Learning and Transferring Closed-Loop Robot Software
Closed-loop robot policies require observation processing, state management, and situation-dependent branching, making them costly to design and tune manually. Although coding agents increasingly support control-code generation and optimization, it remains unclear whether implementations improved on source tasks also support policy acquisition for new tasks. We study this question by treating complete closed-loop implementations as reusable execution experience. For each source task, a coding agent generates policy code from a few successful demonstrations and iteratively improves it using simulation feedback. The validation-selected implementations are retained in a software archive. For new tasks, the agent generates and improves policies using archived implementations, target demonstrations, and execution feedback. The resulting policy is then frozen and executes without further model calls. Across four source tasks in RoboCasa, iterative optimization increases mean success from 28.3% to 64.2%. Across nine target tasks and three independent runs, mean success is 45.2% without references, 41.5% with initial source code, and 57.0% with optimized source code. Optimized references outperform initial references in all three runs on the nine-task average, with a mean gain of 15.6 percentage points. These results demonstrate the value of execution-improved software as a resource for acquiring new policies in this setting, although initial references remain better on two target tasks when averaged across runs.
阅读 arXiv 原文
代码质量与优化 · 0/30 · 2026-09-17AutoData:预训练数据选择的智能搜索把数据选择视为可执行算法搜索,用代理模型反馈迭代改进AutoData: Agentic Search for Pre-training Data Selection
LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop. We frame pre-training data selection as heuristic engineering over per-document features, i.e., lexical statistics, categorical labels, and perplexity. We introduce AutoData, an agent that searches directly over executable selection algorithms. Unlike prior data mixture methods that optimise weights over a fixed set of domains, AutoData searches a richer program space of scoring, stratification, and stochastic selection rules, discovering feature interactions automatically by iteratively refining algorithms with validation feedback from a proxy model. Within an overnight search, AutoData discovers a selection algorithm that outperforms existing human-designed curation pipelines. Despite being searched only on this small proxy, the discovered recipe transfers to larger scales and improves the downstream metric CORE. These results suggest that data engineering can be treated as an agentic machine learning problem, extending autonomous research from model and training-code optimization to the data.
阅读 arXiv 原文
代码质量与优化 · 3/30 · 2026-09-15PrimeScientist:研究投入的策略分配把方向选择与资源投入建为序贯决策,用可执行计划树保留备选PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research
Autonomous research agents aim to automate scientific workflows, from proposing ideas to conducting experiments and analyzing results. Yet current AI and research agents can propose more directions than available resources allow them to pursue. Moreover, each attempt could consume substantial resources, requiring agents to reconsider how to invest in subsequent research. Thus, deciding how to invest research effort strategically should be a defining capability of autonomous research agents. Accordingly, we introduce PrimeScientist, which jointly determines research direction and resource investment across successive research attempts. Specifically, we formulate this challenge of strategic research effort allocation as a sequential decision problem where remaining resources should explicitly guide the research policy. We first introduce an executable plan tree that preserves competing plans and their outcomes across attempts. Building on this representation, we propose an adaptive MCTS-based allocation policy that balances exploration and exploitation using experimental feedback and remaining resources. Comprehensive evaluations across AI research, systems and code optimization, and machine learning engineering show that strategic allocation improves research quality and sample efficiency together. Across 12 AI research tasks, PrimeScientist improves average reward by 10.3% with 50.6% fewer research attempts than AutoResearch under the same resource budget. We believe making research effort allocation an explicit optimization target establishes effective resource use as a core research capability for autonomous agents to drive scientific breakthroughs at scale.
阅读 arXiv 原文
代码质量与优化 · 3/30 · 2026-09-15非对比表示学习的四型克隆检测基于 VICReg 加跨层一致性与深度加权,规避负采样偏差Type-IV Code Clone Detection via Layer-Wise Non-Contrastive Representation Learning
Software clones are fragments of code that are similar or functionally equivalent to each other. They pose significant challenges for maintenance, refactoring, and bug detection. Detecting Type-IV clones, which are semantically equivalent but may differ syntactically, is particularly difficult for traditional token- or syntax-based methods. Recent machine learning approaches rely on contrastive learning, which requires careful negative sampling and can introduce bias. In this paper, we propose LWVIC4Code, a non-contrastive representation learning approach specifically designed for Type-IV clone detection. Building on the Variance-Invariance-Covariance Regularization (VICReg) framework and prior layer-wise VICReg training, LWVIC4Code introduces cross-layer consistency regularization and depth-dependent layer weighting to progressively refine semantic information across transformer layers, producing robust and discriminative code representations. We conduct an empirical study comparing LWVIC4Code against a contrastive learning baseline and zero-shot large language models on Python (Kamino) and multi-language (GPTCloneBench) datasets. Results show that LWVIC4Code achieves competitive or superior performance without negative samples, benefits from layer-wise supervision, and generalizes effectively from Python to other languages, particularly Java and C#. These results demonstrate that non-contrastive, layer-wise representation learning is a promising direction for robust semantic code clone detection.
阅读 arXiv 原文
个人知识与本体(17 篇)
个人知识与本体 · -3/30 · 2026-09-25PIA:健康对话转为结构化记录主张健康代理记忆需类型化记录与时间规则,摘要式检索难支持趋势问答PIA: A Personal Intelligence Agent Turning Health Conversations into Records and Records into Understanding
General-purpose agent memory summarizes conversations: it extracts salient snippets, embeds them, and retrieves the top-k into the prompt. A health agent cannot run on summaries: a dose becomes a sentence, "since last week" is resolved at the model's discretion, and a three-month glucose trend cannot be answered by text similarity. We present PIA, a personal intelligence agent deployed alongside a consumer health agent. PIA receives the agent's natural-language requests, decides for itself whether and how to write or read, and turns conversations into typed clinical records and records into a synthesized understanding of the user. Its memory harness consists of four controls -- extraction, memory, retrieval, and understanding -- each a domain-agnostic mechanism with a pluggable health module: schema, medical alias dictionary, knowledge graph, and temporal rules. We show how the same query receives a different answer as the memory injected into the response context deepens from one-dimensional recall, to a two-dimensional health snapshot, to a three-dimensional trajectory with causality, and report lessons from operation: self-reported health data are missing not at random, question phrasing governs the quality of synthesized understanding, and nearly a third of candidate causal links are structural noise that rules alone remove.
阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-25共享代理记忆的信念准入基准CPB 评估主张是否应写入共享记忆,关注来源去重与答案覆盖的权衡A Benchmark and Diagnostic Study of Epistemic Admission in Shared Agent Memory
Evaluating claim admission in shared agent memory is challenging because repeated claims may be mistaken for independent evidence. An agent may copy or paraphrase a retrieved belief, while admitting a false claim exposes subsequent agents to it. To study this problem, we introduce the Correlated Promotion Benchmark (CPB), which evaluates whether candidate claims should be admitted to shared memory.CPB-Static constructs a frozen test split from publicly annotated sources with fixed gold actions. CPB-Live runs multi-agent teams over a shared store, records all writes and retrievals, and tracks source lineage defined by each scenario. A separate consumer answers from the store alone. We evaluate eight admission policies across four agent families. Our results show that policies which deduplicate sources reject many true claims alongside false ones, whereas policies preserving answer coverage admit nearly as many false claims as unrestricted sharing. Gating on declared source type reduces false adoption to 0.06--0.09, compared with 0.22--0.47 for other answering policies. Once an uncontested false belief enters memory, the consumer asserts it in 0.97--0.99 of probes across all families. No non-oracle policy consistently rejects false claims across verbatim copies, paraphrases, and paraphrases declared authoritative. These findings reveal the limitations of admission policies without access to source lineage.
阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-24MemProbe:代理记忆稳定性探查借认知实验范式诊断记忆更新与保持权衡,并分解为行为剖面Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms
Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cognitive-science-inspired framework for diagnosing stability-plasticity tradeoffs in agent memory. The framework is motivated by a core insight from cognitive memory research: memory is reconstructive and shaped by interference, source reliability, reinforcement, and reactivation. MemProbe turns this insight into four reusable experimental paradigms (interference, misinformation, consolidation strength, and reconsolidation window) that manipulate when a memory should be updated, preserved, or treated as uncertain. It further decomposes correctness into behavioral profiles that reveal how systems update, preserve, attribute, and temporally organize information. We instantiate these paradigms in a 56-episode diagnostic suite and evaluate six incremental memory systems under a unified protocol. Results show that systems with similar aggregate scores exhibit distinct behavioral profiles. MemProbe provides such a diagnostic lens, turning aggregate performance into interpretable profiles of memory maintenance over time. Code is available at https://github.com/jq-ding/MemProbe.
阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-24生产规模 AutoResearch 的失败模式十二周运行报告五种失败模式,含记忆衰减与搜索方向停滞AutoResearch at Production Scale: Failure Modes and a Multi-Agent Framework
Optimizing embedding systems for production recommendation pipelines demands systematic exploration that consumes disproportionate engineering effort at scale. We apply Andrej Karpathy's AutoResearch paradigm -- a large language model that iteratively edits a training script and retains modifications that improve a held-out scalar metric -- to automate this exploration. We report on twelve weeks of running this paradigm at production scale, where iterations consume hours of multi-GPU compute, evaluation involves competing criteria, and campaigns span weeks across many training jobs. Across two independently developed representation-learning systems for a book recommendation pipeline, we ran 220+ experiments and observed five recurring failure modes absent from the original setting: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation. We contribute a three-principle scaffolding design -- prevent, persist, redirect -- that maps each failure mode to a structural remedy and whose instantiation scales with iteration cost. The framework produced a 1.82x Recall@6 lift and a 2.1x coherence lift over hand-tuned baselines, and the agent autonomously designed a text-only fallback that expanded catalog coverage by 5.8x. The two systems span nearly three orders of magnitude in per-iteration cost yet exhibit the same failure modes, suggesting these are structural properties of production-scale autonomous research rather than artifacts of either application.
阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-24持久记忆的作用域匹配与干扰按任务族限定技能检索范围,摘要称可减少有害部署次数Scope Before You Persist: Preventing Cross-Family Interference in Agent Memory
Persistent memory lets language-model agents improve prompts and skills without updating model weights. We show that matching retrieval scope to certification scope enables these edits to support reliable repeated adaptation across recurring task families. We study frozen-model agents on ProcStream-RSI, a 12-round code-repair stream, using Orthogonal Regression Control (ORC), an execution-grounded gate for persistent skill edits. In an intervention that holds proposals and gate decisions fixed, retrieving each accepted skill only for its originating family raises mean hidden trajectory utility from 0.713 under global memory to 0.816 and changes harmful deployments from six of eight to none. In 27 paired randomized-order streams, Scoped-ORC improves mean trajectory utility by 0.063 [0.037, 0.094] over Global-ORC, accepts 63 rather than 12 updates, and produces multiple accepted updates in 19/27 streams, with 0/63 harmful acceptances. The global control reaches 0.713, below the static agent's 0.775, because locally valid edits can interfere with unrelated families. These results establish scope matching as a complementary control for persistent agent memory: certification determines whether an edit is supported, while retrieval scope determines where that evidence authorizes its use.
阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-23META:情景记忆增强交易代理多指标代理加记忆模块,构建无状态分析之外的决策框架Agent Memory with Episodic Retrieval for Financial Decision-Making
Large language models (LLMs) have demonstrated strong capabilities in financial analysis and reasoning, inspiring recent advances in agent-based trading frameworks. While these systems show promise, prior approaches either emphasize long-horizon forecasting or operate as stateless analyzers, limiting their applicability to the demands of trading in complicated settings. To address these gaps, we introduce META (Memory Enhanced Trading Agent), the first RAG-like episodic-memory-augmented multi-agent framework for financial decision making. META integrates a family of specialized indicator agents (e.g., Trend, MACD, Stochastic, RSI, SMA, AVWAP, Heikin-Ashi) with a Decision Agent that fuses their reports, and a Memory module that retrieves and updates past trading episodes encoded as market state embeddings with outcomes and reflections. By recalling relevant experiences and adaptively reweighting signals under similar market regimes, META achieves improved directional accuracy and robustness under short-horizon evaluation. Our results demonstrate that episodic memory provides a powerful mechanism for regime-aware, interpretable, and low-latency decision-making in trading and decision making. The code of this project is released on GitHub.
阅读 arXiv 原文
个人知识与本体 · 6/30 · 2026-09-23TWIST:对话记忆干预质量基准四条赛道考察信念变更时是否恰当干预,并配对假阳性控制TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment
Long-conversation memory benchmarks increasingly test recall and prompted knowledge updates, and recent work studies evolving user beliefs and memory state. TWIST is a proposed benchmark suite for a complementary, unmeasured property: intervention quality -- whether a deployed memory system, exercised through its own ingest/recall/vet surface, acts correctly at belief change points. Four tracks cover unprompted tension detection, vetting outgoing drafts against the record, answering with current beliefs while preserving supersession history, and governing sensitive recall. The suite extends LoCoMo's corpora and harness, pairing every detect/block metric with a matched do-not-over-detect control: surface-matched hard negatives price false intervention, so no track can be gamed by flagging everything. The benchmark itself is validated first: independent, gold-blind double annotation with adjudication, judge decoy calibration, and a separability audit. On the human-validated Track B v1.0 key (161 items, post-adjudication kappa = 0.85), no tested configuration simultaneously achieves high contradiction recall, high hard-negative specificity, and high attribution: flat-RAG baselines detect 0.76-0.97 of true contradictions but falsely flag 16-43% of surface-matched safe drafts depending on backend, while a deployed coherence-oriented system almost never over-flags (0.98-1.00 specificity) yet catches 42% of true contradictions -- a trade-off no recall-only score can see. A 13-configuration baseline ladder localizes causes: every gold contradiction is detectable from its evidence alone (recall 1.000), calibrated models nearly solve the track given the full transcript -- consistent with substantial retrieval-coverage gaps -- and draft-only floors reveal model-dependent style priors. A system's TWIST profile, beside its recall score, measures whether memory knows when to intervene and when not to.
阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-23即时记忆:读时策展的代理记忆主张保留原始轨迹、推迟到读时策展,避免写入时不可逆丢弃Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents
Agentic memory systems reuse past experience to improve future performance, yet most existing designs curate memory at write time: once a task is completed, its trajectory is distilled into a fixed artifact, such as a reflection, workflow, skill, or reasoning strategy, that is later retrieved by similarity. This forces the system to decide what is worth remembering before the future query is known, irreversibly discarding information and producing a query-independent summary that must serve many possible downstream tasks. Learning such a write-time curator is also difficult because the value of a storage decision may only become apparent when a relevant query arrives, potentially many tasks later, creating a long-horizon credit-assignment problem. We instead retain raw trajectories and defer curation until read time, when the current task is known. Given the retrieved traces and the new task, a memory curator synthesizes a compact, task-adaptive payload tailored to the immediate need. Because this payload is consumed on the same task, the curator can be trained directly from immediate task success, avoiding delayed utility signals and the need to artificially group related tasks. Across ALFWorld, WebShop, and $τ^2$-bench, our Just-in-Time Memory (JitMem) consistently outperforms no-memory agents as well as heuristic and learned write-time memory methods, improving over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively. Notably, even an untrained curator is already competitive with or surpasses these baselines, showing that task-adaptive read-time curation itself is a major source of the gain; training the curator further compounds the improvement.
阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-23EnSIMem:实体结构化长期记忆离线建实体—属性索引条目并保留出处与时序,在线按证据需求检索EnSIMem: Entity-Structured Indexing for Long-Term Agent Memory
An agent that interacts with users over long periods must recall facts, preferences, events, and changes from a continuously growing interaction history. Existing memory systems often compress interactions into generic summaries or retrieve anonymous text chunks, making it difficult for an agent to identify the correct entity, property, and supporting evidence. We present EnSIMem, an entity-structured long-term memory architecture for an agent. During offline construction, the system organizes interactions into theme-coherent episodes and builds dialogue-grounded index entries of the form [entity][entity type][property:value]. Each entry preserves its source turns, temporal information, and available multimodal fields. During online interaction, the agent's request is decomposed into evidence requirements whose properties are aligned with the memory index. Entity-property lookup and adaptive retrieval then collect the evidence needed for point, temporal, compositional, and aggregation reasoning. The agent generates its response from the preserved source evidence rather than from lossy memory summaries. On long-term agent-memory benchmarks, EnSIMem achieves high answer accuracy while maintaining compact contexts and favorable online efficiency. These results show that entity-structured indexing and episode-level provenance provide a reliable foundation for long-term memory in agents. The code of our model is available at https://github.com/RamonMeng/EnSIMem.
阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-22AkasicMEM:受治理的企业记忆强调来源到记忆的授权连续性,防派生复用导致信息泄露AkasicMEM: Governed Enterprise Memory for Agents
Agent memory enables enterprise agents to retain knowledge acquired during work and reuse it across tasks and agents, turning execution experience into persistent organizational knowledge. Realizing this potential requires both source--memory integration, through which enterprise sources and accumulated memory can be utilized together, and memory governance, through which shared memory remains subject to organizational policies throughout its lifecycle. These requirements interact when information from enterprise sources persists in memory. As this information is repeatedly derived and reused under changing principals and policies, source restrictions may be bypassed, resulting in information leakage. Preventing such leakage requires authorization continuity, under which source restrictions remain effective throughout source-to-memory and memory-to-memory derivation and reuse. Existing approaches address these concerns individually, but do not treat source--memory integration, memory governance, and authorization continuity as combined core design targets across the memory lifecycle. We define Governed Enterprise Memory as agent memory designed around this combined scope and present AkasicMEM as its realization. AkasicMEM realizes authorization continuity through transitive lineage, policy composition during memory formation, and policy re-evaluation during retrieval. It is built on GraphAI's AkasicDB, a unified vector--graph--relational database whose storage and execution substrate enables the underlying operations of these mechanisms to be jointly optimized and executed.
阅读 arXiv 原文
个人知识与本体 · 5/30 · 2026-09-21DolphinBench:代理记忆帕累托前沿以任务完成度评记忆,用有历史与无历史对照验证任务有效性DolphinBench: Mapping the Pareto Frontier of Agent Memory
Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are built for a conversational question-answer format, where the question itself signals that some fact must be retrieved, and often which one. Moreover, benchmarks rarely require anything beyond accuracy from submissions, allowing memory systems to make unreasonable cost/time tradeoffs to achieve higher scores. We present DolphinBench, a benchmark that evaluates memory directly through an agent's task completion. DolphinBench includes three knowledge-work personas with roughly 500k tokens of user messages per persona and evaluates agents on tasks that depend on information from that history. We verify all 200 tasks per persona by running an agent with and without the relevant history, requiring success with it and failure without it. Finally, we require all evaluations to report total cost and latency alongside accuracy, which enables us to evaluate agent memory systems holistically. No existing memory benchmark combines all three. The dataset and evaluation code are available at https://dolphinbench.ai.
阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-21代理 AI 安全事件报告要素依专家输入梳理报告内容,如记忆访问与工具使用,问题仍开放Beyond Predictable Paths: Redefining AI Security Incident Reporting for Agents
AI agents are being deployed rapidly, accompanied by a growing number of AI-specific attacks and corresponding incidents. As incident reporting becomes increasingly important for legal compliance, governance, accountability, and security; current frameworks must be adapted to the unique characteristics of AI agents. In this paper, two editorial authors compare AI systems and AI agents and, drawing on input from 23 experts in academia and industry, identify the information required for reporting incidents where the security of AI agents is harmed. %involving AI agents. Potential reporting elements include, for example, agent memory and memory accesses, actual and potential levels of autonomy, and tool usage. Based on these findings, we identify several open research questions, including how to efficiently record incidents and how to determine whether vulnerabilities and incidents generalize. Expert feedback also highlighted potential reporting weaknesses, such as risks of data leakage and attacks targeting the reporting infrastructure itself, creating additional research needs. Lastly, we summarize privacy requirements and outline research directions for the secure and trustworthy deployment of AI agents.
阅读 arXiv 原文
个人知识与本体 · 5/30 · 2026-09-21MemCalib:代理记忆使用校准评估模型是否恰当使用记忆影响,发现常过度或不足使用MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents
The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over its response. Yet this capability has remained largely overlooked. To assess this capability, we introduce MemCalib, a benchmark grounded in realistic memory-system scenarios for evaluating memory use and advancing optimization algorithms. Results on the MemCalib test set reveal that frontier open- and closed-source models struggle to use memory appropriately. They frequently over-use or under-use memory rather than matching each proposition's actual use to its target level, leading to biased, low-quality responses. Experiments with common post-training algorithms, including group relative policy optimization and on-policy self-distillation, further reveal a clear directional skew: trained models improve in one direction while deteriorating in the other. We therefore propose MemCalib-RL, an ordered bidirectional counterfactual credit-assignment algorithm that separates over- and under-use signals and localizes their credit to response tokens through exact atom ablation. Results across model families and scales (Qwen3-8B, Ministral-3-8B-Instruct, and Qwen3.5-35B-A3B) show that MemCalib-RL achieves the best overall performance while better balancing over-use and under-use, with gains generalizing beyond MemCalib in external benchmark evaluation. Further experiments support its design choices and robustness and provide insight into its training dynamics.
阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-21Jev-Mem:系统一直觉控制记忆用快速控制面管记忆组织与检索,减少记忆路径上的生成开销Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents
Agentic memory is becoming essential for long-horizon AI agents, yet many existing systems rely on autoregressive LLMs to control how memories are organized, retrieved, and used, placing expensive generation on the critical path of memory operations. We introduce \textbf{\method}, a new agentic memory architecture inspired by System-One/System-Two cognition. System One captures fast, lightweight decision-making, whereas System Two performs slower, deliberative reasoning. Jev-Mem brings this division of labor to agentic memory through a dedicated System-One control plane, a structured multi-relational memory plane, and a System-Two reasoning plane. The System-One controller governs memory typing and relational organization during construction, and dynamically performs query routing, retrieval-budget allocation, graph traversal, candidate scoring, and adaptive stopping during retrieval. System Two is invoked only for complex reasoning and answer synthesis. This design improves both memory effectiveness and system efficiency: on LoCoMo Jev-Mem achieves an overall LLM-as-a-Judge score of 0.777, an 11.0\% relative improvement over the strongest baseline, while reducing memory construction time to 158\,s, a 6.6$\times$ speedup over the fastest competing memory system, and lowering average query latency to 0.93\,s, a 36.7\% reduction.
阅读 arXiv 原文
个人知识与本体 · 0/30 · 2026-09-17自演化检索索引让索引自主诊断检索短板并修改索引键,减少人工介入Self-Evolving Search Index
Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index keys, retrieval quality depends heavily on how effectively these keys expose the knowledge contained in each document. However, effective index representations vary across retrieval environments, making it difficult for any fixed optimization strategy to perform consistently. Yet evolving an index to its retrieval environment remains largely human-driven, requiring humans to diagnose retrieval failures, refine the optimization strategy, and reprocess the index accordingly. We propose SELF-INDEX, a framework that enables an index to self-evolve without human intervention. Its Optimizer autonomously diagnoses retrieval shortfalls, selectively revises the responsible index keys, and validates each revision before updating the index. Beyond reacting to observed retrieval demands, SELF-INDEX proactively explores additional demands through a Query Simulator, allowing the index to evolve beyond the queries already available for optimization. Across diverse corpora and retrievers, SELF-INDEX consistently improves retrieval performance while outperforming existing index optimization methods. We further show that these benefits extend to downstream applications, improving the effectiveness and efficiency of search agents and helping agent memory systems retrieve useful past interactions.
阅读 arXiv 原文
个人知识与本体 · 3/30 · 2026-09-16双过程语言代理的记忆与反思扩展在 ScienceWorld 上做特征开关消融,摘要称全系统平均分最高Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments
Language agents remain brittle in interactive environments, where success requires long-horizon state tracking, valid action execution, and recovery from failed steps. We extend SwiftSage, a dual-process agent that combines a fast action proposer with a slower planner, using two modular cognitive extensions: an Adaptive Memory Module (AMM) for salience-gated episodic storage and trigger-driven retrieval, and a Self-Reflection Module (SRM) for bounded execution-time validation and corrective intervention. Both modules are implemented as feature-flagged extensions over the same execution substrate, enabling controlled ablations on ScienceWorld. Across four configurations---baseline, baseline+AMM, baseline+SRM, and the full system---the full system achieves the best mean final score (64.62), success rate (43.17%), and successful-step efficiency (19.33 steps), while SRM is the strongest standalone contributor. The results suggest that execution-time control is the dominant bottleneck in this setting, while episodic memory becomes most useful once the runtime loop is stabilized.
阅读 arXiv 原文
个人知识与本体 · 4/30 · 2026-09-16WFM:复杂代理推理的知识表示主张稀疏图与稠密文档结合的代理原生知识表示,细节有限WFM: Wiki Foundation Model for Complex Agentic Reasoning
Real-world agents fundamentally require persistent non-parametric knowledge for dynamic reasoning, i.e., long-term memory and retrieval-augmented generation. While graphs have shown reliable advantages in providing structured evidence, the sparse graph representations naturally restrict machine readability and semantic density required for complex agentic workflows. Driven by this limitation, the entire industry is witnessing a paradigm shift from traditional sparse graphs to LLM Wiki, an agent-native knowledge representation that couples dense document contexts with markdown files containing multi-layered topological linkages. However, parameterizing such rich semantics is challenging to encode dense textual contexts using traditional sparse graph embeddings. Moreover, learning LLM Wiki with existing graph encoders could overwhelm distributed system overheads that hinder deployment in large-scale commercial scenarios. To this end, we propose a novel paradigm Wiki Foundation Model, i.e., WFM, tailored for scalable, agent-native representation and retrieval. Specifically, (i) we formalize a Wiki Graph schema that seamlessly bridges fine-grained structures with dense contexts, maintaining explicit topologies alongside continuous semantics; (ii) A query-conditioned attentive aggregation is tailored for rich wiki message passing and explicit attention variance regularization; (iii) We engineer an infrastructural NCCL boundary exchange protocol that hoists static partition indices and leverages fixed-shape GPU-to-GPU collectives, bypassing CPU serialization and memory copy overheads. Extensive evaluations across five long-term agent memory and multi-hop reasoning benchmarks demonstrate the remarkable performance of WFM, while achieving a 10.5 times training acceleration on distributed clusters.
阅读 arXiv 原文