Papers nuevos sobre Software y programación
135 papers nuevos sobre software y programación en los últimos 7 días, dentro de Computación. Acá están los 50 que Pipette considera más valiosos, con el resultado principal en palabras de sus autores.
Lo mejor de la semana
CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding
We applied our pipeline to the 167 million files of the Stack-Edu corpus, creating the first known large-scale open-taxonomy knowledge graph for source code.
Preprint con versión publicadaDice ser un gran avanceBeyond Natural Language: An Agent-Native Language for Autonomous Science
We introduce Lara, a machine-checkable language and protocol for checking and revising support for research claims.
PreprintDice ser un gran avanceHuman-guided physics-constrained AI agents construct an auditable model of soil-plug evolution
We introduce a human-in-the-loop, physics-constrained multi-agent workflow where human experts define admissible physics and modeling boundaries, while agents retrieve evidence, derive equations, implement solvers, and audit the theory-to-code chain.
PreprintUso en el mundo realUnity Insight: A Production Code--Asset Index for LLM Coding Agents in Unity Projects
We present Unity Insight, to our knowledge the first persistent, LLM-facing, agent-integrated cross-file code--asset index for Unity projects, shipping in production with Tuanjie Codely, the agent CLI of Tuanjie Engine, since its public launch on 2026-07-28.
PreprintAfirmaciones fuertes, leer con cuidadoDice ser un gran avanceUso en el mundo realXtrace: High-Fidelity GPU Intra-Kernel Tracing via Binary-Level Instruction Splicing
Xtrace is the first GPU kernel tracing system with near-zero compile-time interference and minimized runtime overhead.
PreprintAfirmaciones fuertes, leer con cuidadoDice ser un gran avanceUso en el mundo realThe Refutation Gap: Certifying Both Halves of an Optimality Claim
We close the gap with a pipeline that synthesizes minimal linear straight-line programs over GF(2), where every decisive UNSAT answer emits a DRAT proof checked by an independent third-party checker.
PreprintEVAGE: Autonomous MEV Generation and Adaptation via Multi-Agent Harness
We present EVAGE, the first fully autonomous multi-agent framework for end-to-end MEV strategy generation and adaptation.
PreprintAfirmaciones fuertes, leer con cuidadoDice ser un gran avanceUso en el mundo realA Governance-Aware Large Language Model Orchestrated Agentic Digital Twin for Transmission System Operator Control Room Decision Support
This paper presents a governance-aware agentic digital twin for transmission grid control rooms.
PreprintUso en el mundo realFrom Code to Requirements: Agentic Reverse Engineering of Business Rules at Enterprise Scale
This paper presents an agentic framework that autonomously generates Business Requirements Documents (bRDs) through reverse engineering of undocumented enterprise software.
PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo realDjinnlang: Higher-Level Programming by Unambiguous Specification with an LLM in the Compiler
To demonstrate that our LLM-in-the-compiler paradigm is feasible when supported by our unambiguity constraint, we present Djinnlang, a high-level specification language built for this future.
PreprintAfirmaciones fuertes, leer con cuidadoDice ser un gran avanceFormal Verification of Proofs from Automated Theorem Provers for Higher-Order Logic
The resulting prototype reconstructs about 80% of generated proof steps automatically, making Leo-III the first higher-order automated theorem prover to support independently checkable proof reconstruction and providing a basis for cross-system reuse.
Preprint con versión publicadaDice ser un gran avanceVisual Graph Reasoning via Knowledge Compilation
To address this limitation, we propose VGCompiler, a compilation-centric paradigm for visual graph reasoning via knowledge compilation.
PreprintUso en el mundo realQuality over Quantity: Diversity-Aware Data Selection for Efficient Verilog Code Generation
To bridge this gap, we propose VeriSelector, the first data selection framework for Verilog code generation that jointly optimizes quality and diversity.
PreprintInvestorNerd: An Investment and Financial Insights System Based on User Profiles
This paper details the design, implementation, and evaluation of InvestorNerd, demonstrating how generative AI and open financial data can be integrated to create scalable, insight-rich tools for financial literacy.
PreprintUso en el mundo realFácil de leerPresage: Prefetch Search via Agent-Guided Experiments
Through this method, Presage is able to insert prefetches that improve performance by a geomean of 10% across 81 workloads by exploring tens of different prefetching alternatives per workload, including a 2.7% runtime reduction on SPEC CPU 2026 where prior SOTA fails.
PreprintUso en el mundo realSpecification Before Generation: A Pre-Registered, Five-Model Paired Evaluation of a Specification Frame for LLM-Generated Code in Money, Time, Idempotency, and Access Tasks
The frame reduced defects in all five models (mean reduction 0.16 to 0.70 findings per task, every Holm-adjusted sign test significant, every bootstrap confidence interval excluding zero).
PreprintUso en el mundo realSensorWF: A FAIR Generalizable Workflow Framework for Scientific Time-Series Analysis
This work introduces SensorWF, a FAIR-annotated workflow framework for generalizable scientific time-series analysis.
PreprintBetween the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase
We present: (i) a new dataset consisting of the full development history of a 21,000-line Python tool built entirely by Claude AI, with no human-authored code or tests, (ii) two code-provenance tracing tools, (iii) three taxonomies for instruction intent, commit provenance, and response reliability, (iv) application of these to analyse the dataset.
PreprintTraceVIC: Causal Reasoning over Code Evolution for Identifying Vulnerability-Inducing Commits
We present TraceVIC, a temporal graph-based approach for identifying and ranking VICs by reasoning over code evolution.
PreprintUso en el mundo realSpecification-Driven Benchmarking for Automated Program Repair From Static Corpora to Executable Specifications
We propose specification-driven benchmarking, a paradigm in which benchmarks are defined by executable specifications and realized through benchmark generation.
PreprintTiga: Compiling Graph Message Passing at Scale
We present Tiga, a just-in-time compiler that separates the definition of a message-passing program from how its interactions are traversed, computed, and stored.
PreprintVSpector: Specification-Driven Bug Detection for RISC-V CPUs
We present VSpector, a specification-driven bug detection pipeline that directly checks whether CPU register-transfer level (RTL) implementations adhere to official specification rules, without requiring specialized construction of reference models, formal properties, or custom bug patterns.
PreprintUso en el mundo realMetrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen
We argue that compile rate is a scientifically unreliable metric for single-function vulnerability repair, and we support this with five controlled experiments over 203 vulnerable functions from Big-Vul, three open-source code LLMs (350M to 6.7B parameters), and three prompting strategies.
PreprintCódigo disponibleForte: A sensitivity type system for imperative Rust
We introduce Forte, a sensitivity type system for Rust whose soundness rests on ownership.
PreprintCONCURDEP: Event-Guided Analysis of Dependency Invalidation in CPython Concurrency
We present CONCURDEP, a source-level static analysis of dependency invalidation.
PreprintUso en el mundo realHow Much of a Real Workload Can LLM-Generated GPU Kernels Actually Reach?
Profiling seven workloads across three domains, we find the addressable fraction ranges from 8.9% to 58.2%.
PreprintCódigo disponibleSchedules Are Solvable Symbols: Tuning-Free Compilation of Tile Programs on Dataflow Architectures
We present Loom, a tuning-free symbolic compiler framework for tile-based SPMD programs on spatial dataflow architectures.
PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo realSWE-PolyVision: Benchmarking Cross-Image Abductive Reasoning for Repository-Level Software Engineering
We present SWE-PolyVision, an executable benchmark of 92 real tasks from 36 open-source organizations, with 48 public tasks and 44 private holdouts.
PreprintPretraining and adapting a language model on a dependency-free stack: GPT-2 124M from random weights, reproduced against llm.c, and a clinical adapter for Qwen3-0.6B
Using numbat, a machine-learning stack written in Zig with no third-party runtime dependencies, we pretrain a 124.4 M-parameter GPT-2 from random initialisation over 9.91 B tokens of web text, then adapt a separate small model to clinical question answering.
PreprintCódigo disponibleThe Canonical Parallel Form as a Substrate for Parallelizing Compilers and Agentic Optimizers
We introduce the Canonical Parallel Form (CPF), a device-neutral program state from which every ordering constraint our analyses prove unnecessary has been removed.
PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo realFoundations of Algebraic Architecture Theory: A Rising Sea of Geometry, Transport, Comparison, and Reconstruction
The main reconstruction theorem identifies the category of full geometries and all their structure-preserving morphisms with an independently defined category of local models, up to equivalence.
PreprintSkelOT: Reusing AOT Compilation Across EVM Contract Families
We present \textsc{SkelOT}, an AOT framework that lifts the unit of compilation reuse from code hash to family skeleton.
PreprintUso en el mundo realBridging the Vendor Gap: Enabling AMD GPU Support for Awkward Array via ROCm/HIP for the HL-LHC Era
We show that a small, reusable set of optimization patterns---loop flattening, -bit vectorized loads, splitting fused kernels, and profile-guided launch configuration---recovers CUDA-class performance without changing the public API.
PreprintUso en el mundo realSLED-IFV: Solver-Validated LLM-Guided Decomposition for Scalable Hardware Information-Flow Verification
Across nine nontrivial benchmarks constructed from real RTL, SLED-IFV achieves up to 603x solver-only speedup and converts two 12-hour timeouts into completed proofs.
PreprintConstraint-Driven Context Engineering: Designing Domain Interfaces for AI Systems
This paper proposes Constraint-Driven Context Engineering (CDCE), a design approach for engineering domain interfaces for AI systems.
PreprintUso en el mundo realPackaged, But Not Portable: Why Conforming to the Agent Plugin Standard Is Rare, and Why Conforming Would Not Be Enough
This paper argues the community standardised a packaging format when composition needs a model, names the four concepts such a model must add - qualified capability identity, a declared capability surface, a precedence rule, and inter-plugin relations - and shows they fit an additive v1.1 profile of the same specification rather than a competing standard.
PreprintUso en el mundo realCódigo disponibleSatisfaction Is Not Explanation: Auditing Vacuity and Training Influence in Temporal-Logic-Guided Reinforcement Learning
This paper introduces an audit layer that can.
PreprintQuantitative coverability for probabilistic well-structured transition systems
For an effective subclass, we solve the approximate quantitative coverability problem over bounded horizons, and over infinite horizons under decisiveness, requiring no probabilistic information beyond individual transition probabilities.
PreprintKernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization
We present KernelOPT, a multi-agent system that treats compiled models as structured artifacts.
PreprintUso en el mundo realFeatLens: Feature-Guided Dynamic Code Graph Construction and Retrieval for Repository-Level Code Generation
This paper presents FeatLens, a feature-guided dynamic code graph construction and retrieval approach for repository-level code generation.
PreprintUso en el mundo realDoes Order Matter? An Empirical Investigation into the Impact of File Ordering on Code Review Effectiveness
Results show statistically significant but modest associations among file position, pull request size, review activity, and latent bug likelihood.
PreprintUso en el mundo realFrom Agent Output to Authorized Transition
The paper contributes a precise vocabulary, compositional architecture, domain profiles, mapping to open-source implementations, and an adversarial evaluation agenda.
PreprintUso en el mundo realVidTutorAssistant: Automating Responses to Programming Tutorial Questions
We present VidTutorAssistant, a web platform that automates responses to viewer questions on programming video tutorials.
Preprint con versión publicadaUso en el mundo realCharacterizing Feedback Statements in Machine Learning Jupyter Notebooks
We contribute a taxonomy of feedback statements in ML notebooks, organized along the functional intent of the statement and the ML pipeline stage in which it appears.
PreprintEntangle: Uncovering Collaboration in the GitHub Quantum Software Ecosystem
Starting from 71 domain keywords, Entangle identifies more than 1,500 quantum repositories, 27,000 contributors and 400 organizations, revealing an ecosystem strongly organized around four leading industrial vendors, but also supported by 2,387 contributors who connect projects, organizations and domains.
Preprint con versión publicadaConstrained Program Generation for 3D Reaction Animation with a 0.8B Model
We present ChemXRG, a domain-specific language (DSL) framework that addresses this gap by representing a reaction animation as an executable program.
PreprintJudgment-Centred Software Engineering Education: A Post-Hype Review and Framework for AI-Augmented Learning
We then refine the AI-Augmented Software Engineering Education (AASEE) framework into five non-linear integration levels and four cross-cutting evidence obligations: explain, verify, modify, and account.
PreprintAgentic-IC3: Enabling Semantic Proof Search in IC3 Model Checking
We present Agentic-IC3, built on Pono's word-level model-checking infrastructure, which integrates a language-model agent into IC3 to guide semantic proof search using register-transfer-level (RTL) design information.
PreprintTeach-to-Crash: A Closed-Loop Student-Teacher LLM Framework for Collision-Inducing Test Scenario Generation
We introduce Teach-to-Crash, a closed-loop testing framework that combines a constrained ego-centric scenario representation, stagnation-aware search control, and a dual-LLM architecture for adaptive failure discovery.
Preprint con versión publicadaUso en el mundo realDueList: A Theory of Lists with Combinators for SMT Solvers
In this work, we provide first-class support for reasoning about lists within SMT solvers.
Preprint