New papers on Search & recommendation
88 new papers on search & recommendation in the last 7 days, within AI & machine learning. These are the 50 Pipette rates most worth reading, with the main result in the authors' own words.
The best of the week
Machine-Interpretable Information: Compiling Documents into Searchable and Readable Protocol States
We introduce Machine-Interpretable Information (MII), the first agent-to-agent (A2A) document-to-state protocol.
PreprintBold claims, read criticallyClaims a big stepUniK: Universal Knowledge Perception for Digital and Physical AI
Across five digital AI domains (medical literature, open-domain QA, chemistry, legal video proceedings, and government open data) UniK combined with an open-source 70-billion-parameter model consistently matches or outperforms frontier proprietary LLMs that are orders of magnitude larger: 76% RAG accuracy on government data versus 47% for GPT-5; 77.9% on medical QA without fine-tuning; topping all open-source chemistry pipelines.
PreprintBold claims, read criticallyReal-world usePRISM-RAG: Multimodal Hypergraph Retrieval-Augmented Generation for Tobacco Product and Legislative Policy Reasoning
PRISM-RAG retrieves passages from the correct jurisdiction in 93.9% of policy compliance queries, a 48.6 percentage point advantage over standard RAG (p<0.001), using zero LLM calls at index time and one at query time, and is competitive with or outperforms SOTA RAG frameworks across keyword, semantic, jurisdiction-, and compliance-accuracy metrics.
PreprintReal-world useMeet, Compare, or Abstain: LatWeave for Deterministic Multi-Hop Question Answering on Knowledge Lattices
We present LatWeave, which organizes knowledge into a multidimensional knowledge lattice and compiles multi-hop QA into three deterministic operators -- meet (constraint intersection), compare (lattice-order comparison), and abstain (structural abstention); LLMs appear only on the construction side (one-shot extraction) and the query-planning side, while the answer-generation path is zero-LLM, zero-task-training, and auditable end to end -- so that question answering over Web-published knowledge becomes reproducible item by item.
PreprintWhen More Evidence Hurts: Publication-Bias Drift and Principled Stopping for Biomedical Causal Search
Against full-budget retrieval, DACG-agent reduces evidence drift from 15.7% to 6.4% and improves null-effect accuracy by 21 percentage points (40.0%61.4%) while using 67% fewer retrieval steps; overall accuracy rises from 61.4% to 69.3% (95% CI 61--77).
PreprintReal-world useComputation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
We show that, for frozen off-the-shelf encoders and language models, it is not: identity is computed when both sentences share one forward pass, and is not a property of the embedding geometry those systems ship.
PreprintLarge Knowledge Model: From Papers to a Scientific Reasoning Landscape
We introduce the Large Knowledge Model (LKM), a scientific knowledge infrastructure that transforms the literature into a shared, computationally accessible reasoning resource.
PreprintBold claims, read criticallyBeyond Raw Engagement: A Counterfactual Observability Framework for Recommender Systems at Netflix
In this work, we present a general evaluation framework that enhances observability across multiple recommender systems at Netflix and demonstrate its effectiveness through several production deployments.
PreprintReal-world useHybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large Scale
We present a hybrid GPU-CPU co-serving system that resolves the paradox through orchestration rather than a new model class.
PreprintReal-world useOneTrans-V2: Unifying Retrieval, Pre-rank, and Fine-rank with One Transformer in Industrial Recommender
Building on OneTrans' model-level unification, we present OneTrans-V2, one Transformer that unifies the entire cascade.
PreprintReal-world useSeek: Self-Evaluative Exploration for Knowledge Retrieval
We introduce Seek, Self-Evaluative Exploration for Knowledge Retrieval, a training-free framework that addresses this limitation through iterative corpus interaction at test time.
Preprint with a published versionReducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education
The findings suggest that carefully designed course-specific AI systems may reduce barriers to academic support by occupying an intermediary space between independent study and formal support.
PreprintReal-world useAn Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency
Evaluations on mainstream large language models and multiple open-source datasets show that MDL reduces the hallucination rate under conflicting memories by about 56.04% in general scenarios and approaches zero hallucination in high-risk scenarios.
PreprintReal-world useGraded-Relevance Composed Multimodal Retrieval for E-commerce Visual Search at Scale
We propose a methodology for training CIR retrievers on graded relevance, consisting of: (i) a VLM to curate training data, generating both queries (object detection + modifier synthesis) and 4-level relevance labels without manual annotation, (ii) an iterative relevance-feedback loop that expands the training set by mining hard negatives from the in-training retriever, and (iii) a hierarchy-aware angular objective to train the retriever directly on the graded labels rather than collapsing them to a binary split.
PreprintReal-world useOvis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings
In this report, we introduce Ovis-Embedding, a state-of-the-art omni-modal embedding family built on native integration of text, image, video, and audio.
PreprintBold claims, read criticallyReFilter: Bridging Embeddings and LLM Filtering for Similar Mobile App Retrieval
To address this gap, we propose ReFilter, a hybrid framework that first Retrieves semantically related candidate apps using embeddings and then applies LLM-based contextual Filtering to identify true functionally similar apps with higher precision.
PreprintReal-world useCross-Country Code-Mixing for Generative Recommendation
Inspired by code-switching corpora in multilingual natural language processing, we propose CMRec, a cross-country GR framework that injects cross-country supervision at the data level via dual-constrained, context-aware code-mixing.
PreprintReal-world useExplainable Recommendations at Scale: LLM Rationales for YouTube Music Artist Discovery
We demonstrate that combining LLM-backed recommendations with these explanatory rationales significantly reduces the trust barrier for new content, yielding statistically significant improvements in both user exploration and overall engagement on the discovery surfaces.
PreprintReal-world useRAG-NAROK: Retrieval-Aware Knowledge Corpus Poisoning in RAG with Source-specific Refutation
Our results demonstrate that RAG-NAROK significantly outperforms static baselines across diverse domains, revealing a fundamental tension between RAG transparency and AI security.
PreprintBold claims, read criticallyCRISS: A Retrieval-Augmented AI Chatbot for Assisting Cancer Registrars
Domain-specific RAG improved evidence grounding and response quality for cancer registry questions while enabling citation-supported assistance across complexity levels.
PreprintReal-world usePredictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention
Multi-hop retrieval failures are not uniformly distributed across queries: they cluster in structurally predictable subpopulations.
PreprintAutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory
We present AutoViewMem, a data-driven framework that organizes long-term conversational memory into self-configuring, low-overlap semantic views before indexing.
PreprintConnected Content Retriever: Dense Graph Edge Features Powering Pre-Ranking at LinkedIn
At its core is a sorted-search GPU primitive that joins dense graph affinity features (viewer to author) with document level features stored on the GPU at runtime in 5-10 ms. The shift to GPU served scoring enabled a 50x scale up of the ranking model's parameters and delivered a +2.5% lift in content time spent on the LinkedIn Feed in online experiments, significantly higher than the typical gains observed in LinkedIn Feed experiments.
PreprintReal-world useKnowledge-as-Skill: A Structural Design for Autonomous Knowledge-Base Use by LLM Agents
We propose Knowledge-as-Skill, an organization scheme that makes a knowledge base discoverable, navigable, and self-descriptive.
PreprintReal-world useTEMPS: Temporal Sentence Embeddings for Temporal Information Retrieval
On three temporal benchmarks, TEMPS improves MRR for every semantic backbone tested and, on TS- Retriever, lifts R@1 from 19.92 to 25.39 over the prior temporal state of the art.
PreprintReal-world useBeyond a Scalar: Distributional Serving Interfaces for Watch-Time Prediction
To address this limitation, we propose the Distributional Serving Interface (DSI), which has a distribution provider, a compact, low-dimensional summary, and lightweight readouts tailored to each task.
PreprintReal-world useThe Visual Target Matters: Learning across the Visual Hierarchy for Brain-to-Image Retrieval
To this end, we introduce NeuroGlyph, which learns a trial-independent visual target from multiple depths of a frozen visual backbone.
PreprintLightweight Ranking Heads: Accelerating Multi-Task Experimentation in Production Recommender Systems
To address the critical challenge of slow experimentation velocity, we introduce the Lightweight Ranking Heads (Light Heads) framework.
PreprintReal-world useOBLIQ-IR: Training a Dense Retriever for Oblique Queries
We address this with OBLIQ-IR, a single-vector dense retriever whose training mixture combines per-mechanism synthetic queries with a new form of cross-model supervision: kNN-graph distillation from a frozen authorship encoder, which transfers a style-versus-topic inductive bias into the student.
PreprintCode availableX-Rec Technical Report
To address these limitations, we propose X-Rec to directly learn the recommendation distribution in the continuous item embedding space through flow matching and generate embedding triggers for approximate nearest neighbor retrieval.
PreprintReal-world useEfficient Iterative Retrieval with Heterogeneous Batching
To address these, we present Orthrus, a serving system that performs heterogeneous batching within a unified inference loop.
PreprintReal-world useCode availableScoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores
A score computed without a live engine can estimate page quality or query-page fit, while end-to-end visibility additionally depends on engine-specific exposure and selection.
PreprintLearned Cross-Task Relationships in Multi-Task Models
We propose a framework that learns cross-task relationships in multi-task models by approximating the joint distribution of task labels through targeted pairwise relationships.
Preprint with a published versionReal-world useA Systematic Multi-Domain Evaluation of Document Retrievers
To address this gap, we conduct a large-scale empirical evaluation of document retrievers, covering three families (sparse, dense, and expansion-based) and evaluating 33 retrievers across seven IR datasets, analyzing retrieval quality, runtime, and failure points.
PreprintEvidenT: Building Trustworthy Enterprise Assistants through Evidence Groundedness and Traceability
On approximately 500 real enterprise queries, EvidenT improves gold-source hit rate by an average of 29% over prompting baselines, produces no citations to nonretrieved urls, and achieves near-saturated answer-to-source lexical coverage.
PreprintReal-world usePotential for Enhanced Learning in Machine Learning Classes by Using Wiki LLM Indexing
The compiled wiki remained accurate and grounded (9.93; 100% grounded in cited sources), while retrieval scored lower and was markedly less grounded (8.14; 64%).
PreprintReal-world useAdaMerge: Tuning-Free Patch Compression for Multi-Vector Visual Document Retrieval
On the long-document benchmark ViDoRe-V2 (4 datasets, two backbones), AdaMerge significantly outperforms tuned PtM across the operating range (p < 10^-4); on the short-document benchmark ViDoRe-V1 (10 datasets, two backbones), where all merging methods are already near-lossless, AdaMerge matches tuned PtM without any per-dataset tuning.
Preprint with a published versionAsymmetric Dynamic Routing: Balancing Reasoning Depth and Computational Efficiency in Hypergraph RAG
To balance reasoning quality and inference efficiency, we propose Asymmetric Dynamic Routing (ADR), an intent-conditioned retrieval framework operating over hierarchical knowledge graphs.
PreprintAnalyzing Public Discourse on Urbanism: Topic Clustering, Sentiment Analysis and Retrieval-Augmented Generation using YouTube Comments
We present a pipeline and conversational system that combines geographic entity resolution, topic modeling, sentiment analysis, and Retrieval-Augmented Generation (RAG) over 22,788 chunks of YouTube transcripts and comments spanning 309 North American cities.
PreprintUNIQUE: A Unified Retrieval and Ranking System for Large-Scale Feed Recommendation
To address them, we present UNIQUE, a unified retrieval and ranking recommendation framework with single-layer flat quantization.
PreprintReal-world useFrom Prompt to Recommendation: A Fitted Stage Model of Brand Visibility in AI Search
We fit a chronological diagnostic model using prior-run history and contemporaneous retrieval indicators: On the latest 30% holdout, the full model achieves AUC 0.963 on GPT and 0.942 on Gemini, compared with 0.937/0.917 for prior history alone and 0.880/0.840 for live signals alone.
PreprintWhat Should We Ask Next? Retrieval-Aware Question Learning under Partial Evidence
We establish that candidate discriminativeness and perceived usefulness provide weak supervision for this objective, then introduce RAVEL, a retrieval-aware online reinforcement learning framework for interactive person re-identification.
PreprintBoundaryMORPH: Budgeted Reranking via Active Set Selection for Diffuse Retrieval
We demonstrate that BoundaryMORPH achieves state-of-the-art set retrieval quality across multiple models and datasets with open-ended queries ( nCG@100 over the strongest baseline).
PreprintGroundedGEO: Auditing the Evidence Gap in Generative Search Rankings
On the frozen listwise ranker Qwen2.5-7B, unsupported-rich variants show significant normalized rank gain over clean candidates (+0.065 to +0.092 across claim profiles, Holm-corrected), while supported and neutral controls do not; the effect is model-dependent (marginal on MiMo-v2.5, absent on GLM-5.3-Flash).
PreprintMuSeR: Scalable Long-sequence Recommendation with Multi-interest Modeling
Rather than proposing a new modeling primitive, our contribution is a system-level integration that makes long-term, multi-interest, and multimodal modeling jointly deployable in a real-time production pipeline, together with the engineering practices required to sustain it.
PreprintReal-world useCalibrating Reproduced Claims in Recommender Systems
We introduce claim calibration as a way of stating the strongest claim supported by a follow-up study, together with the conditions under which it holds and the parts that remain untested.
PreprintAuditing Source Exposure in Baidu and Google AI Search
The results reveal substantial differences across platform-language settings in overview availability and visible source exposure.
PreprintDecoupled Learning and Selection in Slate Recommendation for Privacy and Stability Under Noisy Scores
The contribution is therefore a privacy-scope contract and a certifiable score-to-slate stability mechanism, not a universal utility claim.
Preprint with a published versionBridging Static and Agentic RAG for Taiwanese Historical Question Answering
We therefore introduce a post-hoc selector that compares the two responses and their cited evidence, significantly outperforming either individual pipeline and recovering 60.34% of the oracle headroom.
PreprintTest-Time Adaptation with Query-Dependent Residuals for Visual Document Retrieval
We introduce Q-REACT, a query-side test-time adaptation method that converts limited reranker feedback into reusable retrieval improvements.
Preprint