New papers on Data & databases
31 new papers on data & databases in the last 7 days, within Computing. These are the 31 Pipette rates most worth reading, with the main result in the authors' own words.
The best of the week
Residential Electricity Consumption Dataset for Sri Lanka (RECON-SL)
As the first dataset of its kind in Sri Lanka, and among the few worldwide to link smart meter records with longitudinal survey data at this scale, it provides a unique resource for energy, machine learning and policy research, with known gaps in smart meter coverage documented for users.
PreprintReal-world useEasy to readDecoding the Legalese: A Scalable and Quantitative Framework for Analyzing Corporate Privacy Policies
Building on our structured representations, we introduce the first standardized and repeatable quantitative metrics for evaluating privacy policies along four dimensions: completeness, transparency, commitment to user protection, and emphasis on business-driven data practices.
PreprintPierce: GPU Ray Tracing for Spatial Joins over Complex 3D Data
In this paper, we introduce Pierce, an approach that reformulates three-dimensional spatial joins over complex polyhedral meshes as ray-tracing operations and leverages the hardware ray-tracing units (RT cores) of modern GPUs to accelerate query execution.
Preprint with a published versionReal-world useSpend Classification Without Leakage: An Evaluation Harness and What It Changed in a Deployed System
We build a harness with four protocols over 1.26 million labelled purchase lines from two US state governments.
PreprintReal-world useCode availablePaperAtlas: an automatically constructed atlas of computational methods and software from 6.4 million open-access articles
We present PaperAtlas, an automatically constructed atlas derived from the PubMed Central open-access corpus.
PreprintDeciphering the Babel of Play: A Human-AI Collaborative Approach for Large-Scale Cross-Language Analysis of Game Reviews
This work offers empirical insights into cross-cultural game evaluation and a scalable methodological approach to multilingual content analysis that preserves human interpretation.
PreprintThe Stochastic Shift: A New Evaluation Paradigm for Text-to-SQL with AI Operators
By validating the standard relational logic and the AI operations separately, our framework achieves state-of-the-art overall accuracy across both platforms (up to 97.2%), proposing a reliable standard for benchmarking AI-powered SQL generators.
PreprintSame Pattern, Different Answer: A Reference Semantics and Divergence Map for GQL and SQL/PGQ Path Patterns
Nine of seventeen constructs draw more than one answer across the engines that accepted them.
PreprintCode availableData Agents: Agentic Data Systems
To address these challenges, we propose a new paradigm called the Data Agent, designed to manage, process, and analyze data with minimal human intervention.
Preprint with a published versionBold claims, read criticallyIt's the Geometry, Not the Model: Effective Rank and Subspace Alignment in Functional Connectivity Classification
These results identify subspace orientation as a key factor in transfer degradation under controlled perturbations.
PreprintThe Tethys Dataset: Seven Years of Hourly Smart Water Metering and a Pipeline for Making It Usable
We present Tethys: 91 months of hourly water consumption data from 24 buildings of a municipal water network, published with a quantitative account of its quality, rather than in place of one.
PreprintReal-world useThe Geometry of Alliances: Vote Transfer Modelling in French Two-Round Elections
We show that ideological proximity alone explains the large majority of vote transfers, that a simple left-right axis is insufficient to capture the relevant distances, and that the structural configuration of the 2024 electorate was far less favorable to the far right than pre-election forecasts suggested.
PreprintBold claims, read criticallyA framework for linking literature-based knowledge integration and infrastructure-supported knowledge integration: Opportunities and challenges from a case study
These findings show that literature-based synthesis can be complemented by infrastructure-supported integration where outputs are accessible and usable, and that advancing knowledge integration depends not only on infrastructures but on making data, code, and workflows accessible, executable, and reusable.
PreprintHistoRAG: A Citation-Grounded Question Answering Assistant for Teaching with Scanned Local History and Heritage Archives
This paper presents HistoRAG, a question answering assistant that answers from one regional collection and cites a volume and a page for every fact.
PreprintReal-world useEfficient Dense Vector Search within Knowledge Graph Content Embeddings
We present QLever-Unified Indexed Vector Embedding Retrieval (QUIVER), an extension to QLever that adds native support for dense vector retrieval within RDF knowledge graphs.
PreprintExploiting Residual Reachability for Cross-Model Migration of Graph-Based Indexes in Approximate Nearest Neighbor Search
Across eight text and image migrations, our DGM methods can achieve up to 17.43 times speedup on constructing the new graph index than the fastest degree-matched reconstruction method while keeping competitive recalls.
PreprintReal-world useKathDB-FAO: Synthesized Query Plans in a Multimodal DBMS
On SemBench, KathDB-FAO cuts execution cost by 58.8% on average across scenarios compared with the next best system, at comparable or better quality.
PreprintReading the Data Back: Enriching Variable-Level Metadata for Model-Data Consistency Checks
Conclusion: The profile makes analysis-variable relationships queryable while preserving evidence and review history; screening yields review candidates with context.
PreprintCode availableStructured Spatio-Temporal Evidence Graphs for Open-Vocabulary Object Retrieval in Videos
To address it, we propose STEG-OVR, a structured spatio-temporal evidence graph for open-vocabulary object retrieval.
PreprintSynthetic Human Mobility Data Generation: A Structured Review of Representations, Methods, and Practical Capabilities
Building on this synthesis, we introduce a Meaning-Population-Autonomy framework that characterises these methods along three dimensions: behavioural meaning, population grounding and scale, and generation autonomy.
PreprintFrom Code Archival to Knowledge Graph: Bridging Software Heritage, COAR Notify and Wikidata
Our core contribution is a pair of application profiles that turn Wikidata into a connector between the scholarly record and archived source code; we openly release all code, application profiles, and harvested datasets.
PreprintCode availableBenchmarking Automated Knowledge Graph Construction from Semi-Structured Data
This work contributes a realistic task definition, extends current state of the art evaluation frameworks, allows evaluation of systems that predict mappings and RDF data alike, and combines evaluation of both KG construction and downstream usage.
PreprintDatabases with Missing Values that are Governed by Missingness Mechanisms
The combination of the MG and the observed DB allows us to build a block-independent probabilistic DB.
PreprintTAILOR: Template-Preserving Augmentation for Long-Tailed Log Parsing
To address this challenge, we propose TAILOR, a log parsing framework that improves template inference for rare log groups through template-preserving augmentation.
PreprintGitScholar: A Dataset for Predicting AI Research Impact from GitHub Engagement
In this work, we propose using GitHub engagement as an additional source and demonstrate that it provides both a timely and accurate signal.
PreprintCode availableGraph Memory for LLM Agents: At What Cost? A Comparative Evaluation of Query, Ingest, and Update Performance Across Graph Database Engines
Our central finding is that no system in this sample is categorically fastest: a native graph engine (Ladybug) outperforms Corvic AI on narrow, bounded-neighborhood shapes, while Corvic AI is faster on shapes that scan or join a large fraction of the graph, and a system implementing graph query syntax via SQL/PGQ (DuckPGQ) is measurably slower purely due to query-plan choice.
Preprintkgsteward: a tool for building, reproducing and maintaining distributed knowledge graphs
To tackle this challenge, we present kgsteward, a Python command-line tool that builds and maintains knowledge graphs inside RDF stores from a single, version-controlled configuration file. kgsteward supports multiple triplestores, keeps the local graph up-to-date with its external sources possibly already in RDF, or transformed into it on the fly, and uses SPARQL 1.1 UPDATE commands to amend further imported RDF on the fly.
PreprintReal-world useAdapting a Large Language Model Crash-Severity Pipeline to Tennessee: Performance Across Sampling Strategies
The results highlight the importance of interpreting LLM crashseverity performance together with sampling design, class balance, class-level metrics, and evaluation-population composition.
PreprintReal-world useSsgCaps: A controlled dataset for the evaluation of sound scene generation algorithms
We introduce SsgCaps, a publicly available dataset of human-engineered sound scenes wherein each scene matches a precisely structured prompt that guides the sampling process.
PreprintIdentifying Suspected Mislabeled Apps in Google Play Application Removal Prediction: An Empirical Comparison of Label Noise Detection Methods
The value of the detectors lies in characterizing this label noise.
PreprintFrom OA colors to observable access states: A commentary on the 2026 OpenAlex proposal
The analysis therefore recommends treating the location/version observation as the atomic unit of OA metadata and deriving work-level labels, legacy colors and user-facing summaries from those observations.
Preprint