Papers nuevos sobre Sistemas y redes
123 papers nuevos sobre sistemas y redes en los últimos 7 días, dentro de Computación. Acá están los 50 que Pipette considera más valiosos, con el resultado principal en palabras de sus autores.
Lo mejor de la semana
CerebroSim: Scalable Whole-Brain Simulator at 100-Trillion-Synapse Scale on the LineShine Supercomputer
Using a model derived from magnetic resonance imaging and diffusion-weighted imaging, CerebroSim simulates 86 billion neurons and 100 trillion synapses on 18,432 nodes across 11.2 million cores of the LineShine Supercomputer, sustaining 24.44 PFlop/s, 91% weak-scaling efficiency, and 94% strong-scaling efficiency.
PreprintDice ser un gran avanceUso en el mundo realReading the Sky to Forecast the Ground: Physics-Informed Link-State Forecasting for LEO Networks at Any Location
In this paper, we introduce Gnomon, a physics-informed system that forecasts user-perceived low-Earth-orbit (LEO) downlink throughput, uplink throughput, and round-trip time (RTT) under different levels of trace availability.
PreprintUso en el mundo realVQ-LIC: Shared Vector-Quantized Learned Image Compression on a Resource-Constrained FPGA
We present VQ-LIC, an asymmetric edge-cloud codec in which a compact INT8 depthwise (DW)-pointwise (PW) analysis transform and multi-codebook vector quantization (VQ) run at the edge on a reusable DW/PW engine pair, while reconstruction is handled by a larger cloud decoder.
PreprintDice ser un gran avanceUso en el mundo realHydrozoan: Latency-Adaptive DAG Consensus under Mixed Byzantine and Crash Faults
This paper introduces Hydrozoan, the first DAG protocol with a dual commit path under a hybrid fault model of f Byzantine and c crashed validators, on n = 3f+c+2p+1 validators.
PreprintFlux: Optimal Scheduling of Optical Circuit Switches for LLM Training
We show that Flux reduces training iteration time by up to and peak NIC buffer requirements by more than three orders of magnitude compared to traditional periodic schedulers.
PreprintUso en el mundo realFast Recovery for LLM Serving via Decoupled Device Memory Lifetime in Dynamo
We present fast recovery for Dynamo based on this principle.
PreprintUso en el mundo realCódigo disponibleScalable Packet Tracking on FPGAs for Erasure-Coded RDMA over Lossy WANs
We present COmpact Multi-path Erasure-coded Tracking (COMET), the first fully hardware-offloaded packet-arrival tracking design implemented on an FPGA-based network interface card (NIC) for multi-path RDMA over lossy WANs.
PreprintAfirmaciones fuertes, leer con cuidadoDice ser un gran avanceUso en el mundo realSPLASH: Co-Designing Sparse Attention with High-Bandwidth Flash for Efficient Long-Context Inference
Across models and context lengths, SPLASH improves decode throughput per GPU by 3.5x-11.4x over the evaluated baselines under a 100 ms per-token latency objective, while keeping accuracy within 4% of dense attention across long-context suites.
PreprintUso en el mundo realDLB: Distributed Load Balancing at Scale for Generative AI Inference
This paper introduces DLB, the Distributed Load Balancer, a novel system designed to minimize end-to-end user latency for large-scale, heterogeneous workloads.
PreprintUso en el mundo realA Carbon-Aware Quantum Computing Framework for LCA-Driven Sustainability in Quantum Cloud Services
Conclusion: Superconducting quantum computers are structurally embodied-carbon-dominated, inverting classical sustainability intuition and motivating direct power measurement and cross-architecture validation as quantum infrastructure scales.
PreprintUso en el mundo realToki: Profiling HBM Performance on FPGA Systems with RISC-V Soft Cores and PCIe Host DMA Traffic
Toki, released as open source, is the first hardware-software framework that enables profiling the performance of HBM on FPGA accelerator cards by jointly considering (i) the execution of workloads on RISC-V soft cores instantiated on the FPGA and (ii) the injection of memory traffic from the host system via DMA over PCIe, providing insights that cannot be obtained with synthetic traffic generators alone.
Preprint con versión publicadaPacket-Level In-Network Semantic Adaptation for Unstable Mobile Emergency Networks
This paper presents DINA, a packet-level in-network semantic adaptation method.
PreprintUso en el mundo realSemantics Delivery Network: Rethinking Web Retrieval Infrastructure for LLM Agents
We propose Semantics Delivery Network (SemDN): an origin-authorized, hierarchical edge substrate that indexes, searches, and smart-caches web content at chunk granularity.
PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo realRoutingBench: Can Agentic Routing Analysis Scale to Production Datacenter Networks?
Our results show that agentic analysis is promising: agent skills curated with a principle termed "explore more; digest less" enables routing-path analysis on hyperscale networks of 50K routers with an accuracy of 99.5%, significantly outreaching the scalability of traditional symbolic analysis.
PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo realFrom Idle to Urgent: A Resource-Harvested HPC Workflow for High-Fidelity Seismic Estimation
By developing two specialized HPC kernels, the proposed method reduces energy-to-solution by 76% and improves throughput by 3.7-fold during normal operations to efficiently construct training datasets, while during emergencies, it couples NN-based inverse analysis with physics-based simulations and dynamic refinement to reduce conventional computational costs by over 97.9%, enabling the generation of highly reliable spatial time-history ground motion distributions within 30 minutes post-earthquake.
PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo realAccelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures
We present a solution employing a set of fixed-configuration fused-upcast GEMM kernels that load 16-bit weights from memory, upcast them to FP32 in registers, and accumulate with IEEE-754 arithmetic in a reduction order that is a pure function of the problem shape and is therefore independent of the device, its SM count, or kernel scheduling.
PreprintUso en el mundo realConduit: An Experience Data Plane for Distributed Reinforcement Learning
We present Conduit, a framework-agnostic runtime that exposes RL experience management as an explicit systems optimization problem.
PreprintHot-Cold Tiering of HBM and High Bandwidth Flash for Agentic LLM Serving
On agentic workloads with Qwen3-Coder-30B-A3B, our design delivers 14 ms time-between-tokens (TBT) and adds only 0.1 ms of resume latency on top of prefill, while hosting more concurrent sessions per GPU.
Preprint con versión publicadaUso en el mundo realNostrAgent: A Decentralized Identity and Delegation Architecture for Sovereign Agentic Systems
We present NostrAgent, a decentralized architecture that unifies all five over Nostr relays using three custom event kinds: Kind 38100 identity declarations authenticated by BIP340 Schnorr signatures with pre-rotation commitments, Kind 38101 scoped delegation chains whose every hop verifiably narrows granted capabilities, and Kind 38102 peer attestations forming a Sybil-deterrent trust graph, with Lightning HTTP 402 (L402) binding payment to agent identity.
PreprintUso en el mundo realBrain API: An Intent-Aware Control Plane for Policy-Governed Agentic Systems
Its central contribution is the decision artifact: a durable, versioned, auditable record of how an intent became an executable plan, capturing which policies applied, which capabilities were evaluated, which alternatives were rejected, and why.
PreprintUso en el mundo realBackstitch: Restoring Request Causality Across a Production Microservice Fleet
Repairs restore the causal chain without disturbing the work it describes: breaks at 240 of the repaired calls fell from 90.46% to 4.69%, and over 112 days the fleet's break rate more than halved.
PreprintUso en el mundo realDHSched: Stateless Control for Stateful Real-Time Avatar Serving
We present DHSched, a control plane that manages long-lived stateful sessions through stateless peer Dispatch replicas.
PreprintUso en el mundo realDeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale
This report presents DeepSeek Elastic Compute (DSec), a production sandbox platform that exposes FnCall, container, microVM, and full-VM sandbox backends through a unified SDK.
PreprintUso en el mundo realAdaptive and Cost-Efficient Joint Scheduling of UAV Routes and Analytics with Transit-Borne Fog
Across 36 workload configurations in a rural region, derived from real cellular and transit data, DA achieves up to 20% higher utility than the strongest heuristic and up to 41% higher utility than the strongest adapted-prior scheduler, while incurring the lowest aggregate cost.
PreprintUso en el mundo realReusing Spare Vehicle Computing Capacity: Is It Viable, Profitable and Sustainable?
Vehicles absorb most of the offloaded traffic within tens of milliseconds; participation yields up to 157 km of monthly driving range per vehicle; and life-cycle emissions drop by over 99% when using VCC compared to that edge infrastructure.
PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo realBridging LLM Serving and CXL-SSDs with Chunk-Aware KV Cache Management
Across five LLM models, LM-CXD reduces average TTFT over a stock CXL-SSD by up to 2.6 with compute asynchronous prefetching and 4.03 with layerwise prefetching, achieving TTFT within 1.5 of local DRAM on average.
PreprintUso en el mundo realNebulaSD: Many-for-Many Speculative Decoding
We present NebulaSD, a many-for-many, or M-for-N, speculative decoding system that organizes draft and target workers into independently schedulable resource pools and dynamically reconstructs stage-specific batches from shared request pools.
PreprintUso en el mundo realCross-Model Autoscaling for Shared LLM Serving
Across seven LLM serving traces, TRE reduces P95 end-to-end latency by 11.9--79.0% and P99 latency by 12.5--72.6% compared with a state-of-the-art KV-cache-based reactive autoscaler running on the same hot-switch runtime.
PreprintUso en el mundo realCódigo disponibleDecoupling Logical Masks from GPU Execution for Dynamic Block-Sparse Attention
We present Tessera, a specialized runtime for dynamic BSA that decouples logical masks from GPU execution while preserving specified attention interactions.
PreprintUso en el mundo realNetInspector: Measuring and Improving LLM Capabilities for Reliable Intent-Based Networking Policy Generation
Motivated by this finding, we introduce NetInspector, a three-layer agentic framework that enforces a verify-then-act protocol, decoupling information retrieval from reasoning so that the LLM focuses on symbolic reasoning while every policy decision is grounded in verifiable network facts retrieved from a live Environment Layer before approval.
PreprintUso en el mundo realFrom WPT to Encrypted Telemetry: A Battery-Free Backscattering-based Polarimetric Wireless Sensor
The proposed platform targets secure, energyefficient active sensing and overcomes key limitations of many prior battery-free approaches, which commonly provide neither on-node computation nor cryptographic protection.
Preprint con versión publicadaUso en el mundo realSoK: From Finding to Deployment: Systematizing the OS Kernel Bug Lifecycle
This SoK systematizes the Linux kernel bug lifecycle from discovery to deployment.
PreprintUso en el mundo realEMA: Elastic and Performance Transparent Memory Across GPUs
We present EMA, a memory sharing system that allows GPUs within a server to borrow and reclaim memory from each other, forming an elastic pool of capacity.
PreprintUso en el mundo realASTRA: Toward Agentic AI for Intelligent Device-Network-Cloud Synergy in Next-Generation Mobile Communication
Validated through system-level simulations in two representative scenarios, ASTRA achieves a 13.1% average throughput gain in dense-crowd cell selection by redistributing UEs from congested cells via semantic load exchange, and an 18.2% passive handover reduction in high-speed mobility through predictive trajectory-aware coordination, providing initial evidence that the proposed agentic framework accesses solution regions structurally inaccessible under protocol-constrained architectures.
PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo realFrom functioning to evolving: A complex systems perspective on future self-organised federated energy communities
Drawing lessons from two highly successful large-scale engineered systems, the Internet and agile software engineering, we highlight how prioritising design for evolution over traditional design for functionality enables energy systems to adapt to net-zero dynamics and handle unforeseen uncertainties.
PreprintUso en el mundo realCo-Fabric: Breaking Host-Domain Boundaries for Unified xPU Interconnection
This paper presents Co-Fabric, a bus-based interconnect that, unlike conventional bus designs, breaks host-domain boundaries to deliver unified xPU interconnection for scale-up superpods.
PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo realMVP: A Motion-Predictive Speculative Vision Pipeline with Non-Blocking Drift Correction
We present MVP, a motion-predictive speculative vision pipeline that operates entirely in the motion domain.
PreprintUso en el mundo realSKYE: Write-Optimized Key-Value Store with Fine-Grained Control over Persistent Memory Accesses
We present SKYE, a write-optimized PM key-value store that achieves high throughput and scalability.
PreprintUso en el mundo realDon't let your Memory defy you: Fragmentation-Aware Serverless Allocation with Elastic Memory Locality
In this paper, we present Memoryless, a fragmentation-aware resource manager that exposes memory locality as a serverless control-plane primitive.
PreprintUso en el mundo realZOCheck: CPU-Shadow Checkpointing for Zeroth-Order LLM Fine-Tuning
We present ZOCheck, a fault-tolerant ZO training system that exploits this replayable structure through a CPU shadow process that continuously replays logged updates, materializes consistent recovery images off the GPU critical path, and persists them asynchronously.
PreprintSemord: Learned Semantic-Preserving Placement and Low-Fanout Routing for Distributed Vector Search
We present Semord, a decentralized vector search overlay system that achieves high recall by routing each ANN query to a small set of relevant peers, without relying on a centralized coordinator.
PreprintUso en el mundo realTetris: Circuit Scheduling for Rearrangeably Non-Blocking Photonic Interconnects
We present Tetris, a circuit scheduling algorithm for RNB photonic interconnects.
PreprintUso en el mundo realAccurate Simulation of Distributed Training Jobs with Network Contention Modeling
This paper introduces MoSim, a GPU-cluster simulator that models DT job execution under dynamic network contention.
PreprintCódigo disponibleKREX: Concurrent Kernel Benchmarking on Shared GPUs via Region-Granular Exclusivity
We present KREX, a runtime for concurrent kernel agent benchmarking with region-granular exclusivity.
PreprintWhat Stops a Small Language Model From Driving a Database Agent
Five server changes, touching no model, prompt or sampling setting, moved six models by 6 to 21 cells out of 30.
PreprintDecomposing Predictive Kubernetes Autoscaling for Large Language Model Serving Under Long Startup Delays
Our main finding is that a simple exponentially weighted moving average (EWMA) predictor with delay-aware lookahead and an upper confidence bound (UCB) margin captures most of the benefit, reducing time-to-first-token (TTFT) service-level-objective (SLO) violations from 53% (reactive, queries-per-second based) to 0.5% across five random seeds; lookahead alone is the single largest factor, a 14 reduction.
PreprintUso en el mundo realA principled approach for energy-efficient training via phase-aware GPU frequency tuning
We present PAFT, a phase-aware, dynamically adaptable GPU frequency tuning system that reduces energy consumption of training workloads with minimal performance overhead.
PreprintUso en el mundo realxTier: Intelligent Tiering for CXL-Enabled Memory
We present xTier, a kernel-resident learned memory-tiering system. xTier attaches eBPF programs to PEBS events and uses a compact quantized MLP to score sampled pages inside the kernel at microsecond-scale latency.
PreprintUso en el mundo realLayerCheck: Adaptive Layer-wise Checkpointing for Large Language Model Post-training
Inspired by this observation, we present LayerCheck, a layer-wise adaptive checkpointing framework that selectively persists layers whose updates exceed a threshold.
PreprintUso en el mundo realDissecting How Die Scaling Breaks GPU Fine-grained Scheduling
Across full-GPU kernel execution, intra-application multiplexing, and inter-application co-location, asymmetry-aware scheduling improves mainstream kernels by up to 1.22x, multiplexed LLM inference by up to 14.3%, and avoids up to 1.33x performance variation.
PreprintUso en el mundo real