BDI-batch: Leveraging Standardized Clinical Questionnaires for Contrastive Learning in Psychological Marker Retrieval
Automatic detection of signs of depression has received increasing attention in recent years, evolving from user-level classification (binary) to fine-grained symptom detection.
Despite recent advances, state-of-the-art sentence retrieval models still struggle to discriminate between sentences associated with closely related symptoms.
In this work, we propose \texttt{BDI-batch}, a novel contrastive learning method
that leverages clinically grounded sentence pairs to encode fine-grained distinctions among depression-related symptoms.
The method exploits a standardized clinical questionnaire and does not rely on annotated data.
We demonstrate the effectiveness of our approach on two early risk detection datasets and under multiple retrieval strategies.
Additionally, we explore the use of LLMs to generate synthetic contrastive samples and observe that their performance falls short of questionnaire-derived pairs. Finally, we conduct an error analysis that reveals the most common errors of the sentence retrievers.
Palabras clave:
Publicación: Congreso
1788260369137
1 de septiembre de 2026
/research/publications/bdi-batch-leveraging-standardized-clinical-questionnaires-for-contrastive-learning-in-psychological-marker-retrieval
Automatic detection of signs of depression has received increasing attention in recent years, evolving from user-level classification (binary) to fine-grained symptom detection.
Despite recent advances, state-of-the-art sentence retrieval models still struggle to discriminate between sentences associated with closely related symptoms.
In this work, we propose \texttt{BDI-batch}, a novel contrastive learning method
that leverages clinically grounded sentence pairs to encode fine-grained distinctions among depression-related symptoms.
The method exploits a standardized clinical questionnaire and does not rely on annotated data.
We demonstrate the effectiveness of our approach on two early risk detection datasets and under multiple retrieval strategies.
Additionally, we explore the use of LLMs to generate synthetic contrastive samples and observe that their performance falls short of questionnaire-derived pairs. Finally, we conduct an error analysis that reveals the most common errors of the sentence retrievers. - Marcos Fernández-Pichel, David E. Losada
publications_es