Papers nuevos sobre Modelos de lenguaje
273 papers nuevos sobre modelos de lenguaje en los últimos 7 días, dentro de IA y aprendizaje automático. Acá están los 50 que Pipette considera más valiosos, con el resultado principal en palabras de sus autores.
Lo mejor de la semana
You Can Tell Who's Asking: What the Web's Questions Are Made Of, and Where They Come From
We present the first diachronic, occurrence-level measurement of web question provenance, and find the crawlable web's questions have shifted from being asked by humans toward manufactured for machines to read.
PreprintDice ser un gran avanceCódigo disponibleCalibrated Decisions at Scale: Converting Police Crash Narratives into Probabilistic Crash Variables with a System One Model (Jev)
This paper formulates narrative coding as gated, typed decisions answered by Jev, a System One model that returns probabilities over analyst-defined options and generates no text.
PreprintUso en el mundo realRheoSampling: Resolving the One-Hot Dilemma in Stochastic Dynamic-Tree Speculative Decoding
Specifically, we inject a sampled token among the deterministic top-K slots and treat it with different probabilities during construction and verification, making RheoSampling the first dynamic-tree method with both context-aware top-K construction and stochastic sampling while maintaining losslessness.
PreprintDice ser un gran avanceUso en el mundo realWhat, When, and How: Audio Description as Constrained Global Optimization
When evaluated on REFRAMED, a benchmark for realistic AD of movies, our approach makes better decisions than prompted LLMs about what to describe and when to describe it, establishing a new SOTA on narrative QA and temporally grounded metrics.
PreprintUso en el mundo realLingLan: An Advancing Traditional Chinese Medicine Diagnosis LLM with Multimodal Data
Experimental results show that our method significantly enhances diagnostic accuracy, achieving a relative improvement of 103.5% over the baseline (62.72% vs. 30.82%) and reaching an F1-score of up to 82%.
PreprintUso en el mundo realDirecting large language models to follow the letter or spirit of the law
With minimal modifications, our method significantly changed LLM behavior across diverse measures, novel vignettes, real-world scenarios, and influential legal cases.
PreprintClarification Is Not Correction: LLMs Fail to Let Go
We argue this misses a deeper problem: in many conversations the model does not forget, it commits too early.
PreprintUso en el mundo realThe Wisdom of Artificial Deliberative Crowds
Across domains, deliberation reduced collective error beyond passive aggregation of independent responses, and post-deliberation individual judgments retained this collective gain.
PreprintBeyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?
Our results show that prompting strategy plays a strong role in stylistic quality: few-shot prompting most consistently improves saj density, while its effects on rhetoric and coherence vary by model, with the strongest models (GPT-4o and GPT-5.4-mini) benefiting most from rule-based prompting on these dimensions, though zero-shot prompting yields the highest aggregate scores across all five models.
PreprintOnline Automated Algorithm Design with Large Language Models
To address these limitations, we introduce online LLM-based AAD, a novel optimization paradigm that treats the algorithm itself as a state-dependent decision variable.
PreprintReceptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models
Finally, we introduce a simple approach that substantially increases receptiveness without increasing substantive deference, demonstrating that conversational receptiveness and substantive independence can be achieved together.
PreprintOmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
Training Qwen3-Omni-Instruct with OmniVChat-RL on synthesized dialogues improves its performance on both OmniVChat-Bench and the human-recorded OmniVChat-Bench-Human.
PreprintBeetle: A Bilingual Model Suite for Modelling Second-Language Processing
Evaluating models on human bilingual and second language reading-time prediction and grammaticality judgement tasks, we find that staged and temporally structured curricula consistently improve alignment with language learner reading time compared to balanced bilingual training, with the largest gains at smaller data scales and for typologically closer language pairs.
PreprintLearning to Discover Interesting Mathematics
Optimizing for our metric creates a model capable of producing more interesting theorems, while also reducing substantial or full overlap with Mathlib from 91.9% to 30.6%, showcasing the creation of more out-of-distribution math.
PreprintAfirmaciones fuertes, leer con cuidadoSelf-Play Pretraining with Zero Data
We introduce Self-Play Pretraining with Zero Data, an initial proof-of-concept towards realizing this vision.
PreprintSignGPT: Toward LLM-Mediated Sign Language Interaction through Gloss-Free Translation and Generation
We present SignGPT, a unified, pose-based framework for gloss-free SLT and SLG.
PreprintUso en el mundo realOne Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction
We formulate enterprise IE as per-user prompt adaptation under interaction feedback and propose Self-Meta-Evolve, a hierarchical framework that maintains a dedicated prompt for each user and continuously refines it through a dual-loop process: an inner loop that edits structured prompts based on persona-conditioned feedback, and an outer loop that evolves the meta-prompt itself by distilling successful editing patterns.
PreprintUso en el mundo realBeyond Repeated Sampling: Learning Search Policies for LLM Reasoning
On hard mathematical reasoning problems, the trained concept generator substantially improves the answer generator's pass@k over naive repeated sampling at the same answer generation allocation, surpasses concepts drawn from much larger untuned models, and transfers to answer generators it was never trained against, including a model from a different family.
PreprintDocMIDE: Learning Multi-Hop Implicit Derivation in Visually Rich Documents
On a 4,151-pair implicit extraction benchmark, DocMIDE raises accuracy from 70.8% to 95.9% on Qwen3.5-4B from only a small set of annotated examples, and transfers to a second backbone architecture.
PreprintUso en el mundo realGreedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference
We show it is not precision-invariant: the same model, prompt, and decoding algorithm produce different outputs in BF16 versus FP16 on identical hardware.
Preprint con versión publicadaEviStreams: Human-in-the-Loop AI Data Extraction for Systematic Reviews in Medicine
We present EviStreams, a live, open-source, no-code web platform that puts review teams in control of AI-assisted extraction at three key stages: program design (a structured decomposition approved before any code runs), field specification (typed field definitions calibrated from a pilot), and extracted predictions (reviewer-blinded dual review with adjudication).
PreprintUso en el mundo realRAILS: Retrieval-Augmented Incremental LLM Clustering at Scale
We present RAILS, a retrieval-augmented incremental LLM clusterer that turns clustering into a simple loop over a growing label pool and scales through document batching with bounded concurrency.
PreprintUso en el mundo realThe Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance
Point accuracy alone therefore misjudges these simulators, and current post-training trades diversity for consensus.
PreprintEvaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction
These findings establish CharBM25 as an effective, GPU-free retrieval strategy that matches or exceeds dense retrieval at negligible computational cost, and show that combining it with a general-purpose LLM of 12B+ parameters delivers reliable, training-free Devanagari post-OCR correction without task-specific fine-tuning.
PreprintUso en el mundo realCódigo disponibleTalking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue
We observe a consistent dissociation: AI produces the surface features of cooperative communication without the underlying social architecture.
PreprintDomain-Adaptive Pretraining Enhances Water Treatment Semantic Representation for Large-Scale Structured Literature Mining
Here, we address this by developing WaterBERT, a domain-adapted encoder model designed for semantic representation and structured information extraction from water treatment texts.
PreprintUso en el mundo realBeyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency
We propose TRACER, a multi-turn user simulator that explicitly models users' evolving intent and learns to align simulated behavior with real interaction trajectories.
PreprintUso en el mundo realEmpath: Tracing Multi-Level Emotion Dynamics in Crisis Counseling Dialogues
We introduce EMPATH, a framework for understanding affective dynamics in mental health dialogues across three granularities: turn-level labels, transition probabilities, and global conversation archetypes.
PreprintOmniEdu: Open Foundation Models for Learning and Teaching
We present OmniEdu, an open family of foundation models for K-12 learning and teaching.
PreprintUso en el mundo realPropose, Verify, Commit: Evidence-Grounded Memory for Long-Horizon Multi-Actor Conversations
We introduce EGMEMORY, which formulates long-horizon multi-actor memory as a searchable state machine that separates persistent message-level evidence from an explicit active state.
PreprintRepresentation-guided in-context learning for medical image interpretation with multimodal large language models
Across eight datasets spanning histopathology, radiology and retinal fundoscopy, RG-ICL improved classification (mean gain 20 percentage points) and visual question answering (VQA) (mean gain 13 percentage points) over no-context and conventional ICL, approaching or exceeding training-based comparators.
PreprintUso en el mundo realFrom Decorative to Load-Bearing: Task Difficulty Shapes the Causal Role of Chain-of-Thought
Across Gemma-2-9B-IT, Llama-3.1-8B-Instruct, and DeepSeek-R1-Distill-Qwen-7B on GSM8K, MMLU, and BIG-Bench Hard, CoT load-bearingness tracks model-relative task difficulty: on easy tasks models silently bypass their own reasoning; on hard tasks they follow corrupted steps and propagate errors.
Preprint con versión publicadaWhen AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation
Our study shows that introducing synthetic reviews compresses rating distributions and reduces both same-paper and corpus-level semantic diversity.
PreprintUso en el mundo realRealize What Matters: Principled Context Representation for Large-Scale Reasoning
In this work, drawing on the cognitive theory of relevance realization, we propose concrete principles for designing AI systems that construct effective representations of very large contexts.
PreprintAfirmaciones fuertes, leer con cuidadoCódigo disponibleRewired or Gated? How Instruction Tuning Shapes Knowledge-Conflict Circuits in LLMs
Five independent methods, node and edge attribution, superposition role analysis, causal ablation, and path patching, converge on gating, with the same heads, in the same late-layers, are found to be reweighted rather than replaced with a high node overlap (0.60-0.82).
PreprintAfirmaciones fuertes, leer con cuidadoHuman-LLM Deliberation as Interactive Proof: Conditions for Verifiability Without Transparency
We prove anytime-valid soundness against adaptive provers: the probability of ever accepting a false claim is at most a chosen error level, provided the task supplies bounds on false passes and human checking errors that remain valid after every relevant history.
PreprintCOMED: The Missing Middle Between Routing and Collaboration in Multi-LLM Inference
We introduce COMED (Controlled Model Escalation for Multi-LLM Deliberation), a post-anchor controller for selective cross-model collaboration.
PreprintUso en el mundo realCollapse, Not Complexity: Failure-Conditioned Decomposition Repair for End-to-End Document Parsing
Repair gains 1.40 Overall (95% CI [0.68, 2.16]) at 1.13x tokens, replicates across three checkpoints, and, with all parameters frozen, gains 2.41 (CI [1.64, 3.46]) on the remaining 1,175 benchmark pages.
PreprintUso en el mundo realRepurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders
We propose LLMAE, a method for repurposing a pretrained decoder-only language model as a continuous text autoencoder by exposing an intermediate fixed-length latent bottleneck within its internal activations.
PreprintExtracting Arguments, Not Just Classifying Them: Instruction-Tuned LLMs for Generative Component Detection
In this paper, we propose ITFACD, a novel approach based on instruction-tuned Large Language Models (LLMs) using compact instruction-based prompts, and reframe ACD as a language generation task, enabling arguments to be identified directly from plain text without relying on pre-segmented components.
Preprint con versión publicadaDisentangling interaction and bias effects in opinion dynamics of large language models
By exposing stark differences between LLMs and providing quantitative tools for comparing interaction and bias contributions to opinion shifts in LLM agent discussions, our approach highlights both promises and pitfalls of using LLMs as proxies for human behavior.
Revista con revisión por paresPERSONAWEAVER: Controllable Diversity Beyond Conventional Archetypes in Procedural Character Generation
Across ten realistic and fantastical settings and three LLM(s), PersonaWeaver produces broader moral and interactional response distributions than prior work.
PreprintCódigo disponibleAligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation
Together, this work shows that while curating community-driven data improves the alignment of LLM responses, model performance remains disparate across distinct sub-communities and specific mental health needs.
PreprintUso en el mundo realPINNForge: Execution-Grounded Evolutionary Design of Physics-Informed Neural Networks for PDE Solving via Large Language Models
We propose PINNsForge, an LLM-driven evolutionary framework for execution-feedback-based automated PINN design.
PreprintCapable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models
We find that the extracted reasoning matches native reasoning performance and substantially outperforms no-reasoning baselines, across competition mathematics, science, and code generation.
PreprintTelecomGPT-R1: Unified Post-Training for Reasoning Across Heterogeneous Telecom Tasks
To fill this gap, we introduce TelecomGPT-R1, a family of open source unified telecom reasoning models structured around four complementary axes: protocol, knowledge, modeling, and fault.
PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo realBlaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament
Taken together, these patterns suggest that the perceived rise in harsh political language reflects not merely a general rhetorical drift, but an ideologically asymmetric hardening of political discourse.
PreprintCódigo disponibleLexLattice: Multilingual Extractive Summarization via Neural Cellular Automata on Document Hierarchies
We introduce LexLattice, an extractive summarizer that reifies a legal act's hierarchy as a two-dimensional semantic lattice and consolidates over it with a masked 2D neural cellular automata before selection.
PreprintAfirmaciones fuertes, leer con cuidadoLIMIT: Less Is More for Instruction Tuning in Text-to-SQL
On the BIRD and Spider benchmark, LIMIT selects only 796 and 863 samples while achieving 100% table coverage, enabling Qwen3-8B to reach 69.1% and 88.9% execution accuracy.This result surpasses methods trained on 20 times more data and establishes a new state-of-the-art among open-source approaches.
PreprintAfirmaciones fuertes, leer con cuidadoGenerative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention Outcomes
GenAI-MI chatbots can deliver MI-consistent interactions perceived favorably, but evidence for sustained behavioral or functional change is limited.
PreprintUso en el mundo real