A speech waveform turning into a slow physiological pulse, beside a wrist-worn sensor band

Speech and Wearables as Health Markers

Speech is the cheapest clinical signal there is. It needs no sensor, no clinic and no appointment, and it changes measurably with cognitive decline, with mood and with relapse. The research question is not whether the signal is there — it is whether it survives outside controlled recordings, across languages and across people.

Clinical speech

We work on dementia detection from speech with cross-lingual features, which matters because almost all available clinical speech data is English and almost no patient is. On depression, our line goes back to real-time estimation from mid-term audio features, later fused with facial information.

How something is said

The same paralinguistic signals carry quality rather than pathology. We assess speaking skills from aural and textual information together, released the PuSQ public speaking quality dataset, and work on detecting filled pauses from audio and ASR — one of the few disfluencies that is both measurable and meaningful.

Beyond the microphone

Where speech is not available continuously, wearables are. Our self-supervised approach to detecting non-psychotic relapses from wearable-based digital phenotyping took second place in the 2nd e-Prevention Challenge, a Signal Processing Grand Challenge at ICASSP 2024.

Adjacent to the clinical work, the same speech stack supports accessibility: SmartSubs streams real-time subtitles from theatrical speech to smartglasses for people with hearing loss.

Selected publications

  • Melistas, T., Kapelonis, L., Antoniou, N., Mitseas, P., Sgouropoulos, D., Giannakopoulos, T., Katsamanis, A., Narayanan, S., & Demokritos, N.C.S.R. Cross-lingual features for alzheimer's dementia detection from speech Proc. INTERSPEECH (pp. 3008-3012) (2023) doi
  • Kaliosis, P., Eleftheriou, S., Nikou, C., & Giannakopoulos, T. A self-supervised learning approach for detecting non-psychotic relapses using wearable-based digital phenotyping 2024 IEEE ICASSPW. IEEE doi
  • Eleftheriou, S., Koromilas, P., & Giannakopoulos, T. Automatic Assessment of Speaking Skills Using Aural and Textual Information Proceedings of The Fourth International Conference on Natural Language and Speech Processing (ICNLSP 2021) (pp. 166-1)
  • Chatziagapi, A., Sgouropoulos, D., Karouzos, C., Melistas, T., Giannakopoulos, T., Katsamanis, A., & Narayanan, S. Audio and ASR-based Filled Pause Detection 2022 10th International Conference on Affective Computing and Intelligent Interaction (ACII) (pp. 1-7). IEEE doi
  • Smailis, C., Sarafianos, N., Giannakopoulos, T., & Perantonis, S. Fusing Active Orientation Models and Mid-term Audio Features for Automatic Depression Estimation Proceedings of the 9th ACM International Conference on PErvasive Technologies Related to Assistive Environments (PETRA 2016) doi
  • Giannakopoulos, T., Smailis, C., Perantonis, S. J., & Spyropoulos, C. D. Realtime depression estimation using mid-term audio features AI-AM/NetMed@ECAI (2014)

Open resources: PuSQ and readys, our speech quality assessment tool.