Selected Publications
2026 · 2025 · 2024 · 2023 · 2022 · 2021 · 2020 · 2019 · 2018 · 2017 · 2016 · 2015 · 2014 · 2012 · 2010
2026
- Reglue Your Latents with Global and Local Semantics for Entangled Diffusion European Conference on Computer Vision (ECCV), pp. 444-463, Springer doi
- Latent Space Disentanglement via Activation Steering for Interpretable Attribute Control in Symbolic Music Generation 34th European Signal Processing Conference (EUSIPCO), Bruges, Belgium doi
- Neural Collapse by Design: Learning Class Prototypes on the Hypersphere 43rd International Conference on Machine Learning (ICML) doi
- From Pretraining to Robustness: Benchmarking SSL Models for Noise-Robust Speech Emotion Recognition ICASSP 2026, IEEE, Barcelona, pp. 19497-19501 doi
- Ultrafine, Solid Particles and Black Carbon Near Real Time Assessment and Critical Properties at a Traffic Hotspot MI-TRAP Monitoring Station in Athens EGU General Assembly, Vienna, EGU26-21010 doi
2025
- Rethinking Objectives for Multi-View and Multi-Modal Contrastive Learning UniReps Workshop, Conference on Neural Information Processing Systems (NeurIPS)
- An AI-Powered Framework for Streamlined Image and Audio Annotations Artificial Intelligence for Security and Defence Applications III, Vol. 13679, pp. 495-513, SPIE doi
- Contrastive and Transfer Learning for Effective Audio Fingerprinting through a Real-World Evaluation Protocol IJMSTA, 7(1), 68-82 doi
- Prototypical Contrastive Learning For Improved Few-Shot Audio IEEE MLSP doi
- On the Robustness of State-of-the-Art Transformers for Sound Event Classification Against Black Box Adversarial Attacks 2025 EUSIPCO. IEEE doi
2024
- Bridging Mini-Batch and Asymptotic Analysis in Contrastive Learning: From InfoNCE to Kernel-Based Losses Forty-first International Conference on Machine Learning doi
- A multimodal dataset for electric guitar playing technique recognition Data in Brief, Vol 52 doi
- A self-supervised learning approach for detecting non-psychotic relapses using wearable-based digital phenotyping 2024 IEEE ICASSPW. IEEE doi
- RobuSER: A robustness Benchmark for Speech Emotion Recognition 2024 12th International Conference on Affective Computing and Intelligent Interaction (ACII) (pp. 1-7). IEEE doi
- MItigating Transport-Related Air Pollution in Europe: The MI-TRAP project The European Aerosol Conference
- Emotion-aware speech popularity prediction: a use-case on TED talks 2024 12th International Conference on Affective Computing and Intelligent Interaction (ACII). IEEE doi
- Greek2MathTex: A Greek Speech-to-Text Framework for LaTeX Equations Generation 13th Helenic Conference on AI, SETN 2024, Piraeus, Greece, September 11-13, 2024. Proceedings 4 doi
2023
- Cross-lingual features for alzheimer's dementia detection from speech Proc. INTERSPEECH (pp. 3008-3012) doi
- Smart Subs Subtitling App for Watching Live Virtual Dome Performances 2023 SMAP. IEEE doi
- Unsupervised Temporal Analysis of Mouse Vocalizations 2023 IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB) (pp. 1-8). IEEE doi
- Jazz Mapping: An Advanced Framework for Solo Analysis and Discovery in Jazz Music Audio Engineering Society Convention 155
- MMATR: A Lightweight Approach for Multimodal Sentiment Analysis Based on Tensor Methods ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 1-5). IEEE doi
- Designing and Evaluating Speech Emotion Recognition Systems: A reality check case study with IEMOCAP ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (pp. 1-5). IEEE doi
2022
- Detecting the sources of chemicals in the Black Sea using non-target screening and deep learning convolutional neural networks Science of The Total Environment, 157554 doi
- Film Shot Type Classification Based on Camera Movement Styles Iberian Conference on Pattern Recognition and Image Analysis (pp. 602-615) doi
- A Dataset for Speech Emotion Recognition in Greek Theatrical Plays 13th Conference on Language Resources and Evaluation (LREC 2022), pages 1040–1046 doi
- Real-time feasibility of a human intention method evaluated through a competitive human-robot reaching game 2022 ACM/IEEE International Conference on Human-Robot Interaction (pp. 1080-1084) doi
- Object Size Prediction from Hand Movement Using a Single RGB Sensor International Conference on Human-Computer Interaction (pp. 369-386) doi
- Video soundtrack evaluation with machine learning: Data availability, feature extraction, and classification Advances in Speech and Music Technology doi
- Cross linguistic speech emotion recognition using CNNs: a use-case in Greek Theatrical Data Proceedings of the 15th International Conference on PErvasive Technologies Related to Assistive Environments (pp. 662-667) doi
- Photography Style Analysis using Convolutional Neural Networks 2022 SITIS. IEEE doi
- Analysis of Mouse Vocal Communication (AMVOC): a deep, unsupervised method for rapid detection, analysis and classification of ultrasonic vocalisations Bioacoustics, 1-31 doi
- Emotion Recognition in Music Using Deep Neural Networks Advances in Speech and Music Technology: Computational Aspects and Applications. Cham: Springer International Publishing, 193-213 doi
- A Dataset for Greek Traditional and Folk Music: Lyra 23rd International Society for Music Information Retrieval Conference (ISMIR 2022) doi
- Audio and ASR-based Filled Pause Detection 2022 10th International Conference on Affective Computing and Intelligent Interaction (ACII) (pp. 1-7). IEEE doi
2021
- Deep multimodal emotion recognition on human speech: A review Applied Sciences, 11(17), 7962 doi
- Instrument Playing Technique Recognition: A Greek Music Use Case In Worldwide Music Conference (pp. 124-136). Springer doi
- Multimodal Summarization of User-Generated Videos Applied Sciences, 11(11), 5260. doi
- Automatic Assessment of Speaking Skills Using Aural and Textual Information Proceedings of The Fourth International Conference on Natural Language and Speech Processing (ICNLSP 2021) (pp. 166-1)
- Lyrics and Vocal Melody Generation conditioned on Accompaniment Proceedings of the 2nd Workshop on NLP for Music and Spoken Audio (NLP4MusA)
2020
- ARCHEO: A Dataset for Sound Event Detection in Areas of Touristic Interest 2020 15th International Workshop on Semantic and Social Media Adaptation and Personalization (SMA (pp. 1-6). IEEE. doi
2019
- Athens Urban Soundscape (ATHUS): A Dataset for Urban Soundscape Quality Recognition. In International Conference on Multimedia Modeling (pp. 338-348). Springer, Cham. doi
- Recognizing the quality of urban sound recordings using hand-crafted and deep audio features Proceedings of the 12th ACM International Conference on PErvasive Technologies Related to Assistive Environments doi
- Real-Time Arm Gesture Recognition Using 3D Skeleton Joint Data Algorithms, 12(5), 108 doi
- Data Augmentation Using GANs for Speech Emotion Recognition Proc. Interspeech 2019, 171-175 doi
- Unsupervised Low-Rank Representations for Speech Emotion Recognition Proc. Interspeech 2019, 939-943 doi
2018
- Speech-music discrimination using deep visual feature extractors Expert Systems with Applications 114 (2018): 334-344. doi
- Curriculum learning of visual attribute clusters for multi-task classification Pattern Recognition, 80, 94-108. doi
- Enhanced movie content similarity based on textual, auditory and visual information. Expert Systems with Applications 96 (2018): 86-102. doi
2017
- Towards predicting task performance from EEG signals 2017 IEEE International Conference on Big Data (Big Data) (pp. 4423-4425) doi
- Curriculum learning for multi-task classification of visual attributes 2017 In Proceedings of the IEEE International Conference on Computer Vision Workshops (pp. 2608-2615) doi
- Daily Activity Recognition based on Meta-classification of Low-level Audio Events Proceedings of ICT4AWE2017, ISBN: 978-989-758-251-6 doi
- Design for a System of Multimodal Interconnected ADL Recognition Services Components and Services for IoT Platforms, pp. 323-333. Springer International Publishing doi
2016
- A ROS Framework for Audio-Based Activity Recognition Proceedings of the 9th ACM International Conference on PErvasive Technologies Related to Assistive Environments (PETRA 2016) doi
- Bayesian network to predict environmental risk of a possible ship accident International Journal of Risk Assessment and Management, 19(3), 228-239 doi
- Short-term Recognition of Human Activities using Convolutional Neural Networks 2016 International Conference on Signal-Image Technology & Internet Based Systems (SITIS). IEEE doi
- Fusing Active Orientation Models and Mid-term Audio Features for Automatic Depression Estimation Proceedings of the 9th ACM International Conference on PErvasive Technologies Related to Assistive Environments (PETRA 2016) doi
- Content Representation and Similarity of Movies based on Topic Extraction from Subtitles Proceedings of the 9th Hellenic Conference on Artificial Intelligence. ACM doi
- Audio-visual speaker diarization using fisher linear semi-discriminant analysis. Multimedia Tools and Applications, 75(1), 115-130 doi
2015
- pyaudioanalysis: An open-source python library for audio signal analysis. PloS one 10.12 (2015): e0144610. doi
- Automatic soundscape quality estimation using audio analysis Proceedings of the 8th ACM International Conference on PErvasive Technologies Related to Assistive Environments (p. 19) doi
- Long-term marine traffic monitoring for environmental safety in the aegean sea International Archives of the Photogrammetry, Remote Sensing Spatial Information Sciences doi
- Counting and tracking people in a smart room: An IoT approach Semantic and Social Media Adaptation and Personalization (SMAP), 2015 10th International Workshop on (pp. 1-5). IEEE doi
- Fusing multiple audio sensors for acoustic event detection Image and Signal Processing and Analysis (ISPA), 2015 9th International Symposium on (pp. 265-269). IEEE doi
2014
- Introduction to audio analysis: a MATLAB® approach Academic Press
- Realtime depression estimation using mid-term audio features AI-AM/NetMed@ECAI
2012
- Fisher linear semi-discriminant analysis for speaker diarization IEEE TASLP 20.7 (2012): 1913-1922 doi
- Detection and clustering of musical audio parts using Fisher linear semi-discriminant analysis 2012 EUSIPCO. IEEE
2010
- Audio-visual fusion for detecting violent scenes in videos 2010 doi


