New papers on Software & programming
135 new papers on software & programming in the last 7 days, within Computing. These are the 50 Pipette rates most worth reading, with the main result in the authors' own words.
The best of the week
CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding
We applied our pipeline to the 167 million files of the Stack-Edu corpus, creating the first known large-scale open-taxonomy knowledge graph for source code.
Preprint with a published versionClaims a big stepBeyond Natural Language: An Agent-Native Language for Autonomous Science
We introduce Lara, a machine-checkable language and protocol for checking and revising support for research claims.
PreprintClaims a big stepHuman-guided physics-constrained AI agents construct an auditable model of soil-plug evolution
We introduce a human-in-the-loop, physics-constrained multi-agent workflow where human experts define admissible physics and modeling boundaries, while agents retrieve evidence, derive equations, implement solvers, and audit the theory-to-code chain.
PreprintReal-world useUnity Insight: A Production Code--Asset Index for LLM Coding Agents in Unity Projects
We present Unity Insight, to our knowledge the first persistent, LLM-facing, agent-integrated cross-file code--asset index for Unity projects, shipping in production with Tuanjie Codely, the agent CLI of Tuanjie Engine, since its public launch on 2026-07-28.
PreprintBold claims, read criticallyClaims a big stepReal-world useXtrace: High-Fidelity GPU Intra-Kernel Tracing via Binary-Level Instruction Splicing
Xtrace is the first GPU kernel tracing system with near-zero compile-time interference and minimized runtime overhead.
PreprintBold claims, read criticallyClaims a big stepReal-world useThe Refutation Gap: Certifying Both Halves of an Optimality Claim
We close the gap with a pipeline that synthesizes minimal linear straight-line programs over GF(2), where every decisive UNSAT answer emits a DRAT proof checked by an independent third-party checker.
PreprintEVAGE: Autonomous MEV Generation and Adaptation via Multi-Agent Harness
We present EVAGE, the first fully autonomous multi-agent framework for end-to-end MEV strategy generation and adaptation.
PreprintBold claims, read criticallyClaims a big stepReal-world useA Governance-Aware Large Language Model Orchestrated Agentic Digital Twin for Transmission System Operator Control Room Decision Support
This paper presents a governance-aware agentic digital twin for transmission grid control rooms.
PreprintReal-world useFrom Code to Requirements: Agentic Reverse Engineering of Business Rules at Enterprise Scale
This paper presents an agentic framework that autonomously generates Business Requirements Documents (bRDs) through reverse engineering of undocumented enterprise software.
PreprintBold claims, read criticallyReal-world useDjinnlang: Higher-Level Programming by Unambiguous Specification with an LLM in the Compiler
To demonstrate that our LLM-in-the-compiler paradigm is feasible when supported by our unambiguity constraint, we present Djinnlang, a high-level specification language built for this future.
PreprintBold claims, read criticallyClaims a big stepFormal Verification of Proofs from Automated Theorem Provers for Higher-Order Logic
The resulting prototype reconstructs about 80% of generated proof steps automatically, making Leo-III the first higher-order automated theorem prover to support independently checkable proof reconstruction and providing a basis for cross-system reuse.
Preprint with a published versionClaims a big stepVisual Graph Reasoning via Knowledge Compilation
To address this limitation, we propose VGCompiler, a compilation-centric paradigm for visual graph reasoning via knowledge compilation.
PreprintReal-world useQuality over Quantity: Diversity-Aware Data Selection for Efficient Verilog Code Generation
To bridge this gap, we propose VeriSelector, the first data selection framework for Verilog code generation that jointly optimizes quality and diversity.
PreprintInvestorNerd: An Investment and Financial Insights System Based on User Profiles
This paper details the design, implementation, and evaluation of InvestorNerd, demonstrating how generative AI and open financial data can be integrated to create scalable, insight-rich tools for financial literacy.
PreprintReal-world useEasy to readPresage: Prefetch Search via Agent-Guided Experiments
Through this method, Presage is able to insert prefetches that improve performance by a geomean of 10% across 81 workloads by exploring tens of different prefetching alternatives per workload, including a 2.7% runtime reduction on SPEC CPU 2026 where prior SOTA fails.
PreprintReal-world useSpecification Before Generation: A Pre-Registered, Five-Model Paired Evaluation of a Specification Frame for LLM-Generated Code in Money, Time, Idempotency, and Access Tasks
The frame reduced defects in all five models (mean reduction 0.16 to 0.70 findings per task, every Holm-adjusted sign test significant, every bootstrap confidence interval excluding zero).
PreprintReal-world useSensorWF: A FAIR Generalizable Workflow Framework for Scientific Time-Series Analysis
This work introduces SensorWF, a FAIR-annotated workflow framework for generalizable scientific time-series analysis.
PreprintBetween the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase
We present: (i) a new dataset consisting of the full development history of a 21,000-line Python tool built entirely by Claude AI, with no human-authored code or tests, (ii) two code-provenance tracing tools, (iii) three taxonomies for instruction intent, commit provenance, and response reliability, (iv) application of these to analyse the dataset.
PreprintTraceVIC: Causal Reasoning over Code Evolution for Identifying Vulnerability-Inducing Commits
We present TraceVIC, a temporal graph-based approach for identifying and ranking VICs by reasoning over code evolution.
PreprintReal-world useSpecification-Driven Benchmarking for Automated Program Repair From Static Corpora to Executable Specifications
We propose specification-driven benchmarking, a paradigm in which benchmarks are defined by executable specifications and realized through benchmark generation.
PreprintTiga: Compiling Graph Message Passing at Scale
We present Tiga, a just-in-time compiler that separates the definition of a message-passing program from how its interactions are traversed, computed, and stored.
PreprintVSpector: Specification-Driven Bug Detection for RISC-V CPUs
We present VSpector, a specification-driven bug detection pipeline that directly checks whether CPU register-transfer level (RTL) implementations adhere to official specification rules, without requiring specialized construction of reference models, formal properties, or custom bug patterns.
PreprintReal-world useMetrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen
We argue that compile rate is a scientifically unreliable metric for single-function vulnerability repair, and we support this with five controlled experiments over 203 vulnerable functions from Big-Vul, three open-source code LLMs (350M to 6.7B parameters), and three prompting strategies.
PreprintCode availableForte: A sensitivity type system for imperative Rust
We introduce Forte, a sensitivity type system for Rust whose soundness rests on ownership.
PreprintCONCURDEP: Event-Guided Analysis of Dependency Invalidation in CPython Concurrency
We present CONCURDEP, a source-level static analysis of dependency invalidation.
PreprintReal-world useHow Much of a Real Workload Can LLM-Generated GPU Kernels Actually Reach?
Profiling seven workloads across three domains, we find the addressable fraction ranges from 8.9% to 58.2%.
PreprintCode availableSchedules Are Solvable Symbols: Tuning-Free Compilation of Tile Programs on Dataflow Architectures
We present Loom, a tuning-free symbolic compiler framework for tile-based SPMD programs on spatial dataflow architectures.
PreprintBold claims, read criticallyReal-world useSWE-PolyVision: Benchmarking Cross-Image Abductive Reasoning for Repository-Level Software Engineering
We present SWE-PolyVision, an executable benchmark of 92 real tasks from 36 open-source organizations, with 48 public tasks and 44 private holdouts.
PreprintPretraining and adapting a language model on a dependency-free stack: GPT-2 124M from random weights, reproduced against llm.c, and a clinical adapter for Qwen3-0.6B
Using numbat, a machine-learning stack written in Zig with no third-party runtime dependencies, we pretrain a 124.4 M-parameter GPT-2 from random initialisation over 9.91 B tokens of web text, then adapt a separate small model to clinical question answering.
PreprintCode availableThe Canonical Parallel Form as a Substrate for Parallelizing Compilers and Agentic Optimizers
We introduce the Canonical Parallel Form (CPF), a device-neutral program state from which every ordering constraint our analyses prove unnecessary has been removed.
PreprintBold claims, read criticallyReal-world useFoundations of Algebraic Architecture Theory: A Rising Sea of Geometry, Transport, Comparison, and Reconstruction
The main reconstruction theorem identifies the category of full geometries and all their structure-preserving morphisms with an independently defined category of local models, up to equivalence.
PreprintSkelOT: Reusing AOT Compilation Across EVM Contract Families
We present \textsc{SkelOT}, an AOT framework that lifts the unit of compilation reuse from code hash to family skeleton.
PreprintReal-world useBridging the Vendor Gap: Enabling AMD GPU Support for Awkward Array via ROCm/HIP for the HL-LHC Era
We show that a small, reusable set of optimization patterns---loop flattening, -bit vectorized loads, splitting fused kernels, and profile-guided launch configuration---recovers CUDA-class performance without changing the public API.
PreprintReal-world useSLED-IFV: Solver-Validated LLM-Guided Decomposition for Scalable Hardware Information-Flow Verification
Across nine nontrivial benchmarks constructed from real RTL, SLED-IFV achieves up to 603x solver-only speedup and converts two 12-hour timeouts into completed proofs.
PreprintConstraint-Driven Context Engineering: Designing Domain Interfaces for AI Systems
This paper proposes Constraint-Driven Context Engineering (CDCE), a design approach for engineering domain interfaces for AI systems.
PreprintReal-world usePackaged, But Not Portable: Why Conforming to the Agent Plugin Standard Is Rare, and Why Conforming Would Not Be Enough
This paper argues the community standardised a packaging format when composition needs a model, names the four concepts such a model must add - qualified capability identity, a declared capability surface, a precedence rule, and inter-plugin relations - and shows they fit an additive v1.1 profile of the same specification rather than a competing standard.
PreprintReal-world useCode availableSatisfaction Is Not Explanation: Auditing Vacuity and Training Influence in Temporal-Logic-Guided Reinforcement Learning
This paper introduces an audit layer that can.
PreprintQuantitative coverability for probabilistic well-structured transition systems
For an effective subclass, we solve the approximate quantitative coverability problem over bounded horizons, and over infinite horizons under decisiveness, requiring no probabilistic information beyond individual transition probabilities.
PreprintKernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization
We present KernelOPT, a multi-agent system that treats compiled models as structured artifacts.
PreprintReal-world useFeatLens: Feature-Guided Dynamic Code Graph Construction and Retrieval for Repository-Level Code Generation
This paper presents FeatLens, a feature-guided dynamic code graph construction and retrieval approach for repository-level code generation.
PreprintReal-world useDoes Order Matter? An Empirical Investigation into the Impact of File Ordering on Code Review Effectiveness
Results show statistically significant but modest associations among file position, pull request size, review activity, and latent bug likelihood.
PreprintReal-world useFrom Agent Output to Authorized Transition
The paper contributes a precise vocabulary, compositional architecture, domain profiles, mapping to open-source implementations, and an adversarial evaluation agenda.
PreprintReal-world useVidTutorAssistant: Automating Responses to Programming Tutorial Questions
We present VidTutorAssistant, a web platform that automates responses to viewer questions on programming video tutorials.
Preprint with a published versionReal-world useCharacterizing Feedback Statements in Machine Learning Jupyter Notebooks
We contribute a taxonomy of feedback statements in ML notebooks, organized along the functional intent of the statement and the ML pipeline stage in which it appears.
PreprintEntangle: Uncovering Collaboration in the GitHub Quantum Software Ecosystem
Starting from 71 domain keywords, Entangle identifies more than 1,500 quantum repositories, 27,000 contributors and 400 organizations, revealing an ecosystem strongly organized around four leading industrial vendors, but also supported by 2,387 contributors who connect projects, organizations and domains.
Preprint with a published versionConstrained Program Generation for 3D Reaction Animation with a 0.8B Model
We present ChemXRG, a domain-specific language (DSL) framework that addresses this gap by representing a reaction animation as an executable program.
PreprintJudgment-Centred Software Engineering Education: A Post-Hype Review and Framework for AI-Augmented Learning
We then refine the AI-Augmented Software Engineering Education (AASEE) framework into five non-linear integration levels and four cross-cutting evidence obligations: explain, verify, modify, and account.
PreprintAgentic-IC3: Enabling Semantic Proof Search in IC3 Model Checking
We present Agentic-IC3, built on Pono's word-level model-checking infrastructure, which integrates a language-model agent into IC3 to guide semantic proof search using register-transfer-level (RTL) design information.
PreprintTeach-to-Crash: A Closed-Loop Student-Teacher LLM Framework for Collision-Inducing Test Scenario Generation
We introduce Teach-to-Crash, a closed-loop testing framework that combines a constrained ego-centric scenario representation, stagnation-aware search control, and a dual-LLM architecture for adaptive failure discovery.
Preprint with a published versionReal-world useDueList: A Theory of Lists with Combinators for SMT Solvers
In this work, we provide first-class support for reasoning about lists within SMT solvers.
Preprint