Generative Models for Audio, Music, Speech and Images
Generating plausible audio, music or images is no longer the hard part. The hard part is control — getting the attribute you asked for and not the three that came with it — and verification: knowing that what came out is what you wanted, and, separately, being able to tell generated material apart from the real thing.
That is where our work sits, across modalities.
Steering without retraining
In symbolic music generation, we study the inner workings of a music transformer and modulate attributes at inference time through activation steering, rather than by retraining the model. The difficulty is that attributes are entangled: pushing on one moves the others. Our EUSIPCO 2026 work addresses that interference directly, with a dual-steering framework based on Gram-Schmidt orthogonalisation.
The same question appears in images. Reglue Your Latents, at ECCV 2026, works on entangled diffusion using global and local semantics.
Generation in service of recognition
Generative models are also a tool rather than a goal. Our GAN-based augmentation for speech emotion recognition, still the most cited result of this line, exists because labelled affective speech is scarce and always will be. Earlier work on conditional generation of lyrics and vocal melody given an accompaniment belongs to the same family.
Telling the difference
The inverse problem matters as much as the forward one. Through the DeepFake Awareness platform, funded by the Greek Ministry of Digital Governance, we also record how well the general public can distinguish AI-generated voice from genuine recordings — a measurement that is hard to obtain any other way.
Selected publications
- Reglue Your Latents with Global and Local Semantics for Entangled Diffusion European Conference on Computer Vision (ECCV), pp. 444-463, Springer (2026) doi
- Latent Space Disentanglement via Activation Steering for Interpretable Attribute Control in Symbolic Music Generation 34th European Signal Processing Conference (EUSIPCO), Bruges, Belgium (2026) doi
- Lyrics and Vocal Melody Generation conditioned on Accompaniment Proceedings of the 2nd Workshop on NLP for Music and Spoken Audio (NLP4MusA) (2021)
- Data Augmentation Using GANs for Speech Emotion Recognition Proc. Interspeech 2019, 171-175 doi
A demo shows GAN-generated video driven by music.


