pipette
ENEnglish

Unmasking digital despair: A language model framework for predicting depression on social media

M. Xie, H. He, Z. Xie

PreprintAfirmaciones fuertes, leer con cuidadoUso en el mundo real

En palabras de los autores

Depression is one of the most prevalent mental health disorders globally. Social media content can reflect emotional states, offering real-time signals of depressive symptoms. This study aimed to develop a deep learning model to predict depression using Twitter/X data, and to characterize potential causes of depression using a large language model (LLM). Twitter/X posts from April 20, 2023, to July 24, 2024, were obtained using the key phrase ("I" or "me") and "diagnosed with depression". After data cleaning and GPT-4o labelling, 2,275 depressive Twitter users, and an additional 1,661 non-depressive users from tweets using the keyword "today", were identified. A deep learning RoBERTa model for predicting depression was built on tweets from both user groups, with an 80/20 training-testing split. GPT-4o models were subsequently employed to analyze depressive users posts to understand potential causes for depression. The RoBERTa model achieved strong performance in predicting depression among Twitter users from their tweets, with an accuracy of 0.822, an F1 score of 0.855, and an AUC of 0.809. Common reasons for depression identified by GPT-4o models included societal pressure, low self-esteem, cultural influences, and identity-related challenges. These findings highlight the potential of deep learning models in early screening of depression using social media data. Insights into potential reasons for depression may inform targeted prevention strategies, public health interventions, and improved mental health support for at-risk populations.

Resultado principalEl resumen no menciona limitaciones.

Apareció: jueves, 24 de septiembre. medRxiv. Preprint, todavía sin revisión por pares.

DOI: 10.64898/2026.09.21.26363552