pipette
ENEnglish

Open-world fungal ITS embeddings improve higher-rank placement and cross-view retrieval, but percent identity is stronger on a matched ITS2 novelty benchmark

A. O'Brien, A. Gardette

Preprint

En palabras de los autores

1. Learned sequence embeddings are increasingly proposed for fungal ITS classification and novelty detection, but are usually evaluated against weaker versions of themselves or default-configured alignment, by AUROC alone, and with queries treated as independent. We asked whether a purpose-trained encoder outperforms percent identity when both score identical queries under a firewalled open-world design, and which evaluation choices decide the answer. 2. From the UNITE release of 19 February 2025 we built ITS-core, ITS1 and ITS2 views and a genus-separated split with sealed calibration and test partitions. A convolutional encoder trained with genus-proxy, cross-view, hierarchical and episodic objectives was frozen and hash-sealed before test data were opened. We compared it with exhaustive, coverage-filtered VSEARCH identity on the same queries and references using genus-clustered paired intervals, and logged every post-opening correction before computing corrected metrics. 3. Three evaluation choices changed the comparison. Default VSEARCH heuristics returned a lower-identity hit than exhaustive search for 48.3% of benchmark queries; without a coverage filter, identity placed only 64.9% of novel ITS-core queries in the correct class, against 98.7% with it, because the conserved 5.8S let partial alignments win; and the encoder's embeddings depended on inference batching. Once corrected, identity exceeded the encoder in known-genus accuracy and novel-family placement at every view (family 80.0% to 84.6% against 53.8% to 64.8%) and detected more novel genera at a lower false-novelty rate. The encoder's novelty AUROC was within 0.023 of identity's, and no genus-clustered interval excluded zero. Identity's own development-to-test gap exceeded the encoder's, so that gap reflects partition composition rather than selection. Of the genera held out by a historical benchmark, 85% had been present in training; on a leakage-safe version, identity's AUROC advantage was 0.065, with a genus-clustered interval excluding zero. 4. Correctly configured alignment matched or exceeded the learned embedding throughout. The decisive results came from the evaluation, not the representation: baseline configuration, batch invariance, genus-level uncertainty, a parameter-free control for selection and a leakage audit each changed a conclusion, and each is inexpensive to apply to any learned barcode method.

Resultado principalLimitación que admiten los autores

Apareció: jueves, 24 de septiembre. bioRxiv. Preprint, todavía sin revisión por pares.

DOI: 10.64898/2026.09.17.752437