PhD Proposal: Voice as a Scalable Biomarker for Health: Interpretable Modeling, Cross-Lingual Generalization, and Toward Translation into Practice

Talk
Roksana Khanom
Time: 
08.20.2026 14:00 to 15:30

Speech and voice provide a rich, non-invasive window into human physiology and have emerged as promising digital biomarkers for monitoring respiratory, neurological, and mental health conditions. Recent advances in self-supervised learning have substantially improved speech representation learning by reducing reliance on manual annotation and enabling robust transfer across downstream tasks. However, existing self-supervised models are primarily optimized for semantic and linguistic understanding, while overlooking physiological characteristics that reflect subtle changes in vocal production. Consequently, their learned representations often fail to capture disease relevant biomarkers that generalize across speakers, recording conditions, and clinical populations. This thesis proposes a unified self-supervised representation learning framework for learning physiology aware speech representations that preserve clinically meaningful vocal biomarkers while remaining robust to nuisance variations such as speaker identity, recording devices, and acoustic environments. First, it establishes that unscripted, naturalistic speech carries measurable respiratory disease information and develops a principled analysis of modern self-supervised speech representations, using the Language Invariance Score for cross lingual biomarker selection and layer wise geometric probing of frozen foundation models, to characterize which physiological biomarkers they capture, which they neglect, and how these limitations affect downstream clinical inference. Building on these insights, it develops biomarker-aware self-supervised learning objectives that explicitly encourage latent representations to encode vocal attributes associated with respiratory health while reducing dependence on disease-specific annotations. Finally, it introduces efficient transformer architectures that enable scalable modeling of long-duration speech recordings through sparse attention mechanisms, allowing the framework to capture long-range temporal dependencies without prohibitive computational cost. Together, these contributions establish a general framework for learning clinically meaningful speech representations that generalize across diseases, populations, and recording settings. Beyond improving the detection and monitoring of respiratory diseases such as COPD, the proposed research advances the broader field of speech based digital biomarkers by enabling more reliable, interpretable, and data efficient models for mobile health, remote patient monitoring, telemedicine, and personalized longitudinal healthcare