pipette
ESEspañol

Identifying Scientists on X

Philipp Meier, Katarina Boland, Laura Kallmeyer, Stefan Dietze

Preprint with a published version

In the authors' words

With the growing importance of science-related discourse on the Web and the erosion of the classical knowledge order, it is important to identify different user groups, such as scientists, automatically. This work proposes an approach for identifying scientists and non- scientists on X/Twitter based on their user biographies and tweets. We show that we are able to classify accounts as scientists and non- scientists on two different datasets, reaching an F1 score of up to 0.88 using Random Forests with linguistic features and up to 0.96 using a contrastively fine-tuned DeBERTa model in an ensemble setup. Furthermore, we provide two datasets with X users labeled as scientists or non scientists and their respective tweets and user biographies.

Main resultThe abstract does not state a limitation.

Appeared: Monday, September 28. arXiv. Preprint with a published version.

DOI: 10.1145/3795513.3810448

Published version: Companion Publication of the 18th ACM Web Science Conference (2026) 155-164

Authors' comment: Corrected version of Identifying Scientists on X published at Companion Publication of the 18th ACM Web Science Conference 2026