pipette
ENEnglish

Identifying Scientists on X

Philipp Meier, Katarina Boland, Laura Kallmeyer, Stefan Dietze

Preprint con versión publicada

En palabras de los autores

With the growing importance of science-related discourse on the Web and the erosion of the classical knowledge order, it is important to identify different user groups, such as scientists, automatically. This work proposes an approach for identifying scientists and non- scientists on X/Twitter based on their user biographies and tweets. We show that we are able to classify accounts as scientists and non- scientists on two different datasets, reaching an F1 score of up to 0.88 using Random Forests with linguistic features and up to 0.96 using a contrastively fine-tuned DeBERTa model in an ensemble setup. Furthermore, we provide two datasets with X users labeled as scientists or non scientists and their respective tweets and user biographies.

Resultado principalEl resumen no menciona limitaciones.

Apareció: lunes, 28 de septiembre. arXiv. Preprint con versión publicada.

DOI: 10.1145/3795513.3810448

Versión publicada: Companion Publication of the 18th ACM Web Science Conference (2026) 155-164

Comentario de los autores: Corrected version of Identifying Scientists on X published at Companion Publication of the 18th ACM Web Science Conference 2026