New papers on Benchmarks & evaluation
203 new papers on benchmarks & evaluation in the last 7 days, within AI & machine learning. These are the 50 Pipette rates most worth reading, with the main result in the authors' own words.
The best of the week
StudentBench: AI and human tutoring yield equivalent GRE learning gains
We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average.
PreprintReal-world useEasy to readCode availableM3GA-Wild: A Large-Scale Dataset and Benchmark for Multi-Modal Multi-session Ground-to-Aerial Place Recognition in Forests
We present M3GA-Wild, the first benchmark for multi-modal, multi-session ground-to-aerial place recognition in forests.
PreprintClaims a big stepCinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models
To address this, we introduce CinematicVQA, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing our introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative functions.
PreprintSynthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark
We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers.
PreprintFinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability
To address this gap, we introduce FinFIRST (Financial Information Retrieval, Sourcing and Traceability), the first financial benchmark to jointly evaluate answers and supporting evidence through atomic rubrics.
PreprintCode availableSeeing Is Not Perceiving: When Synthetic Consumers Can and Cannot Pretest Visual Marketing
Every resulting configuration passed the manipulation checks; however, none of the configurations reproduced more than two of the six human effects, and the remainder were nonsignificant.
PreprintReal-world useASR ensembling for phoneme intelligibility evaluation of speech anonymizers
Our results show that simple hard-voting ASR metric reaches correlations above 0.9 with human ratings when aggregated by feature, test-type, or condition, provided that multiple ASR models are combined; evaluating stimuli with and without a carrier sentence further improves the correlation at the stimulus level.
PreprintARAFA: An LLM-Generated Arabic Fact-Checking Dataset
In this manuscript, we introduce Arafa, a new large-scale dataset for fact-checking in Modern Standard Arabic, constructed through an automated framework leveraging large language models (LLMs).
PreprintEnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise
We present EnterpriseVal, a use-case-level evaluation system that closes this gap.
PreprintReal-world useThe Illinois Social Attitudes Aggregate Corpus (ISAAC): An Open Tool and Reproducible Pipeline for Analyzing Social Group Discourse at Scale
We introduce the Illinois Social Attitudes Aggregate Corpus (ISAAC), an open, modular, and accessible corpus of 527 million+ English-language Reddit posts selected for relevance to six key social group distinctions based on race, sexuality, age, ability, body weight, and skin tone, covering the 17-year period from 2007 to 2023.
PreprintBold claims, read criticallyA Living Benchmark for Information Retrieval from Electronic Health Records
We present a scalable framework that automatically generates question--answer pairs from longitudinal EHR notes.
PreprintReal-world useVox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models
To address this challenge, we introduce Vox-Infinity, the first benchmark specifically designed to evaluate long-context understanding in spoken language models.
PreprintCogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition
We introduce CogGym, a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials.
PreprintFinInteract: Benchmarking Clarification and Intent Integration in Ambiguous Financial Question Answering
We release FinInteract, a bilingual (English/Chinese) benchmark of 173 instances that pairs each question with a default and an intended interpretation across a five-category ambiguity taxonomy, and grades whether an agent elicits the right clarification and then integrates it.
PreprintObject Detection Benchmarks are Incomplete: The Role of Label Errors and Annotation Uncertainty
While object detection has advanced through improved architectures and open-vocabulary models, we provide strong evidence that benchmark quality is limited by annotation incompleteness.
PreprintVimarsha: Faithful ASR Evaluation for Indian Languages with Demographic Diversity, In-the-Wild Audio and Spelling Variations
We introduce Vimarsha, a 100-hour benchmark spanning all 22 scheduled Indian languages, designed to address both distortions.
PreprintGameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
To address these challenges, we introduce GameHorizon, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families.
PreprintCode availableBehavior is Not Enough: A Mechanism-Based Evaluation of Social Norm Emergence in LLM Societies
By making expectations observable, our framework allows us to attribute each mechanism's contribution separately, offering designers of multi-agent systems a principled basis for selecting the social processes that sustain cooperation.
PreprintA Hierarchy-Aware Video-Language Model Evaluation and Hyperbolic Baseline for Surgery
In this paper we make two contributions to address this problem, (i) we introduce SurgHiBench, the first hierarchy-aware evaluation suite for surgical video understanding, with three tasks measuring recognition, consistency, and severity across granularity levels.
PreprintBeyond Overlap: Estimating the Causal Effect of Benchmark Exposure
We present LeakScale, an interventional framework for estimating this missing quantity.
PreprintInvisible in Space, Visible in Time: Motion Vision CAPTCHA against GUI Agents
Results reveal a substantial human--agent gap: on the full MVCAP-Bench, human accuracy reaches 99.6%, whereas the best GUI agent achieves only 16.8%, close to the six-way chance level.
PreprintCIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords
Our results show that LLMs continue to struggle with the cross-lingual understanding of Chinese internet buzzwords, particularly in fine-grained non-literal interpretation, robust equivalent matching under option perturbations, and calibrated harmfulness detection.
PreprintCode availableAugur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes
Our central finding is methodological and negative: most of the measured gap between frontier cloud models and open-weight models we fine-tune and serve offline is attributable to an under-specified evaluation, not a difference in capability. We show this three ways.
PreprintBabelArena: A Large-Scale Multilingual Benchmark for LLM Agents
We introduce BabelFlow, a benchmark-general agentic workflow that adapts existing agent benchmarks to new languages by analyzing runtime dependencies, coordinating structure-preserving translation, and combining multi-layer verification with human review to preserve task and evaluation semantics.
PreprintBeyond the Leaderboard: Counterfactual Diagnosis of End-to-End and VLA Driving Policies Under Domain Shift
We propose a counterfactual check-up: a few hundred real frames, each edited two ways (pedestrian removed, or re-lit by a night-style perturbation), every edit verified by an independent detector, and the change in the planned trajectory read as a diagnosis rather than a score.
PreprintReal-world useTrains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers
We answer it on a governed delivery plane, where an agent drives ten stages and an oracle scores each stage from platform-recorded facts.
PreprintOmni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction
We establish Omni Demand Understanding (ODU) as a distinct multimodal contextual inference problem: given an interaction stream, a model must detect whether a user demand is present and infer intent from multimodal and conversational context.
PreprintEnigmaForge: The Question Is Hidden in the Story
Intuition reshuffles the leaderboard: a 22x spread where fact recovery spans 1.6x, the second-best fact-recoverer ranks fourteenth, one model is indifferent to being told the question, and another is significantly better without it.
PreprintLow-Cost Assays for Measuring Model Behavior Across Vendors and Releases
To address these constraints, we present a simple, cheap, scalable, and replicable model for studying model behavior.
PreprintCode availableProbabilistic Modelling of Operational Design Domains, A New Approach for Testing AI Systems
To this end, we introduce probabilistically extended ontologies (PEONs): ontologies describing the ODD, augmented with a probability distribution over the partitioning they induce.
PreprintReal-world useFinancial Language Models as Applied Artificial Intelligence Systems for News-Based Trading under Market Frictions
We introduce MFAST, a Market-Friction-Aware Sentiment-to-Trading framework that converts timestamped financial text into auditable, reproducible, and market-feasible trading decisions.
PreprintReal-world useBeneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models
Across 60,008 runs on three video-agent architectures, we find that grounding is the dominant source of downstream error, with roughly four times the causal impact of corrupting visual observations.
PreprintAn open benchmark for machine learning-based polymer property prediction
We introduce Polymer Benchmark 2026 (PolyBench26), an open dataset comprising nearly 250,000 polymer-property datapoints across eight physical properties, including data from experimental measurements, density functional theory, and molecular dynamics.
PreprintCode availableOSWorld-Pro: Process-based Evaluation for Computer Use Agents
We introduce OSWorld-Pro: a set of over 300 tasks containing over 2800 subgoals to enable the procedural evaluation of CUAs grounded in over 67,000 human annotations.
PreprintMuLA-Bench: A Multilingual Long-Form Audio Understanding Benchmark via Multi-Tier Auditing
We introduce MuLA-Bench: 5,038 open-ended questions over 1,769 in-the-wild recordings totaling 1,377.9 hours, covering 16 languages and eight domains.
PreprintTWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment
TWIST is a proposed benchmark suite for a complementary, unmeasured property: intervention quality -- whether a deployed memory system, exercised through its own ingest/recall/vet surface, acts correctly at belief change points.
PreprintWPBench: A Comprehensive Benchmark for Wind Power Forecasting
To address these limitations, we propose WPBench, a comprehensive, fair, and extensible benchmark for wind power forecasting.
PreprintReal-world useFrom Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost
Across two datasets spanning four tasks, we show that: (1) sessions with identical quality ratings can differ by up to 70 times in interaction cost; (2) quality-cost relationships vary by task, with some tasks rewarding extended interaction and others favoring fast convergence; (3) subjective user ratings are not reliable substitutes for productivity; and (4) productive sessions are characterized by agents probing earlier and users spending less effort repairing the interaction.
PreprintNo More Free Lunch: Corpus Task Complexity Matters as Corpora Grow
We find that high-CTC tasks not only grow much more challenging on average at longer contexts for LCLMs, they reverse many modeling conclusions drawn solely from low-CTC evaluations.
PreprintWhat Looks Like a Capability Limit in Vision-Language Models Is a Readout Limit
We find it is not, and that the limits a benchmark reports can belong to the readout rather than to the model.
PreprintAn Evolutionary Agentic Approach for Open-ended Image Quality Perception
Across traditional IQA, structural fidelity, context-aware aesthetics, and newly defined open-ended dimensions, PACE consistently improves its MLLM backbone, achieving competitive performance across diverse IQA settings, and reduces the Holistic Override Rate (HOR) from 44.4% to 8.6%.
PreprintNeither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models
Scoring replaces binary windows with a strictly proper threshold-weighted continuous ranked probability score whose weights are intent-conditioned timing kernels fitted to human floor-transfer-offset distributions, proving boundedness, consistency, and binary reduction.
PreprintPAWS: Policy-driven Agentic World Simulation
We introduce PAWS, a Policy-driven Agentic World Simulation dataset covering 36 verified U.S. financial and economic policy episodes, 12,727 policy-linked news records, and 65,291 source-grounded stakeholder actions.
PreprintBeyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction
Across SEEDA, RM-EVAL achieves strong agreement with human rankings, and RGTG yields consistent gains in reward and external validation, demonstrating a unified framework for both assessing and enhancing GEC systems without relying on gold references.
PreprintPotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks
We introduce PotARCin, a benchmark that extends ARC by assessing understanding of a task's underlying abstract rule across five dimensions: Definition, Classification, Constrained Generation, Editing, and Inversion.
PreprintExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
To this end, we introduce ExplorationBench, which turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds: their rules are executable, so every answer can be checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks.
PreprintIntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law
We introduce IntLawNER, a NER dataset and benchmark for codified sources of international law, covering 2,987 gold-annotated sentences and 8,094 entity spans from International Court of Justice (ICJ) decisions, UN Security Council resolutions, and European Court of Human Rights (ECtHR) judgments, annotated with seven institution-specific entity types.
PreprintPredicting Out-of-Distribution Generalization of Neural Operators via Observable Spectral Error Decomposition
Through four controlled experiments, we show that consistently predicts energy error under in-distribution, in-band spectral shift, out-of-band tail, and compound shifts, whereas global metrics can be systematically misleading.
PreprintOmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation
To address these gaps, we introduce OmniVBench and the Omni-R2V Dataset for evaluating and training omni R2V models.
PreprintAre Coreset Selection Methods Worth Their Cost?
Across eight wall-clock budget anchors on each of CIFAR-10 and Tiny ImageNet, no anchor is won by a sophisticated selector: every winner is class-balanced random sampling, repeated random sampling, or full-data training.
Preprint