Research Areas

MagCIL works on multimodal machine learning for audio, speech, music, image and video. Four core areas carry most of our scientific output; each of them reaches practice through a small number of application domains, where the problem, the data and the constraints come from a real setting rather than a benchmark.

Core areas

Speech Analytics

We model both what people say and how they say it, with an emphasis on behaviour that survives real-world conditions: noise, channel shift, accent and language.

  • Speech emotion recognition, and its robustness under noise and distribution shift
  • Paralinguistics: speaking style, public-speaking quality, engagement and speaker diarization
  • Clinical speech: cognitive decline, depression and mental-health monitoring
  • Speech recognition for Greek and other low-resource settings
  • Expressive speech synthesis and synthetic-speech detection

In depth:
Speech emotion recognition that survives the real world · Speech and wearables as health markers

Music Information Retrieval and Music Generation

A long-running line of work on musical content, increasingly focused on idioms that mainstream models represent poorly, Greek traditional, folk and laiko music among them.

  • Datasets and annotation frameworks for underrepresented repertoires
  • Recognition of genre, instrument, playing technique, regional style, mood and emotion
  • Audio fingerprinting for song identification and music-use monitoring
  • Text-to-audio and text-to-music generation
  • Symbolic music generation

In depth:
AI for Greek music and cultural heritage · Music identification and monitoring · Generative models: control and verification

Computer Vision and Image Generation

Representation learning and generative modelling for images and video, built on vision foundation models and self-supervised pretraining.

  • Controllable image generation with latent diffusion
  • Self-supervised and few-shot visual representation learning
  • Video understanding: summarization, classification and retrieval
  • Multimodal fusion of sound, image and text
  • Human activity, gesture and intention recognition

In depth:
Movies and video that are heard, not only seen · Generative models: control and verification

Audio Event Recognition and Acoustic Scene Analysis

Non-speech, non-music audio: what a recording tells you about the place it was made in and the events that happened there.

  • Sound event detection and acoustic scene classification
  • Urban soundscape monitoring at scale
  • Bioacoustics and animal vocalization analysis
  • Few-shot learning for acoustic classes with little labelled data
  • Adversarial robustness of audio models

In depth:
Sound events and urban soundscapes · Frugal and robust audio models

Application domains

Culture and Cultural Heritage

AI for Greek musical and linguistic heritage: semantic infrastructure for traditional and laiko music, automatic transcription and indexing of archival speech, emotion in theatrical performance, and the urban soundscape treated as living heritage.

Health and Wellbeing

Speech and physiological signals as low-cost markers of health: cognitive decline and dementia, depression, mental-health relapse, and unobtrusive monitoring in assisted-living settings. The same work underpins accessibility applications such as real-time subtitling for people with hearing loss.

Defence and Security

Frugal and robust AI for settings where data is scarce, compute is constrained and models are adversarially probed: few-shot acoustic recognition, adversarial robustness evaluation, AI-assisted annotation pipelines, and paralinguistic triage at borders.

Environment and Sustainability

Acoustic and remote sensing for environmental monitoring: urban soundscape quality, transport-related air pollution, marine traffic and environmental risk in the Aegean, and the attribution of chemical pollution sources.

Trustworthy and Synthetic Media

As generative audio and video become routine, so does the need to tell them apart from the real thing. We work on synthetic-media detection and on public-facing AI literacy, including a national deepfake awareness platform.

AI Infrastructure and Open Science

Making the above usable by others: open-source libraries for audio analysis and for training and deploying audio models, released datasets and benchmarks, and contributions to national and European AI platforms.

Methodological foundations

Running underneath all four areas is a shared methodological agenda, which is also where our most general results sit: self-supervised and few-shot representation learning, contrastive learning, generative modelling with diffusion and flow matching, robustness under distribution shift, and multimodal fusion. Methods developed for one modality routinely move to another, which is the main reason the group keeps all four areas under one roof.