LitBench: A Benchmark for Retrieval-Grounded Multi-Paper Evidence Synthesis in Epilepsy
In the authors' words
BackgroundClinicians increasingly use AI systems to search the medical literature, but current benchmarks do not jointly test whether a response identifies the originating paper and its supporting passage. LitBench evaluates four dimensions: stated versus interpretation-requiring facts, the number and composition of competing papers, one-versus two-paper evidence, and refusal when evidence is absent. MethodsLitBench comprises 1,980 open-access papers (1,000 epilepsy papers holding the answers, 980 decoys from unrelated fields) and 9,472 human-reviewed facts, giving 2,188 single-paper questions and 170 requiring a fact from each of two papers. Difficulty rose by burying the answer among up to 2,000 distractors, then again over live PubMed Central. The same questions were then asked with the answering paper removed, so that refusal was the correct response. Four systems were tested: Gemma-4B, Gemma-12B, Sonnet-5, and DeepSeek-V4-Flash (refusal conditions only). Three model judges scored each answer by majority; two-paper questions counted only when both facts were found. ResultsAcross fixed-corpus single-paper conditions, accuracy ranged from 55.9% to 71.9% for Gemma-4B, 61.2% to 75.5% for Gemma-12B, and 90.0% to 91.3% for Sonnet-5. Similar epilepsy papers were harder than mixed candidate sets for both Gemma configurations. Two-paper accuracy was 5.7%, 17.1%, and 37.8%, respectively. With evidence absent, Gemma-12B refusal fell from 88.3% to 39.9% as candidate sets grew; Gemma-4B almost never refused. Sonnet-5 refused on 94.9% of sampled single-paper and all sampled two-paper cells. DeepSeek-V4-Flash also refused frequently, but on 80.3% of matched answer-present controls. ConclusionsNo system was reliable across every axis tested: accuracy fell as competing papers were added, interpretation-requiring facts were harder than stated ones, two-paper synthesis was rarely achieved, and only some systems recognized absent evidence. What counts as "successful" literature search carries considerable nuance; LitBench can evaluate proposed tools at a fine-grained level.
Appeared: Thursday, September 24. medRxiv. Preprint, not yet peer-reviewed.