Keynote Speaker

Keynote 3 - Language Queried Audio Source Separation
Prof. Wenwu Wang / University of Surrey
2026-09-10 | 08:00-09:00 (Europe/Rome)
Large audio-language models (LALMs) are emerging as a powerful approach for addressing a broad range of audio signal processing problems, including generation, captioning, editing, and source separation. By integrating large language models (LLMs) with audio modeling techniques, they enable more flexible and intelligent audio systems. This talk focuses on language-queried audio source separation (LASS), a recently introduced framework that extracts desired sounds from an audio mixture using natural language queries. LASS offers an intuitive and scalable interface for applications such as automated editing, remixing, and audio rendering. We start by outlining the problem and its motivation, making a connection with conventional paradigms for speech and universal source separation. We then present two LASS methods: AudioSep and FlowSep. AudioSep is a text-driven model that uses query and separation networks to estimate time–frequency masks and extract target sounds. FlowSep is a generative approach based on rectified flow matching (RFM), modeling transitions from noise to target features in a variational autoencoder (VAE) latent space and reconstructing audio via a decoder and vocoder. We also present datasets, evaluation metrics, experimental results, and sound demos. Finally, we link source separation to audio editing and generation, highlighting recent advances (e.g. the RFM-Editing, AudioLDM, and AudioLDM2 models), and conclude the talk with future research directions in this evolving field.
Bio
Wenwu Wang is a Professor in Signal Processing and Machine Learning, and the Associate Head of External Engagement, School of Computer Science and Electronic Engineering, University of Surrey, UK. He is also a Core AI Fellow at the Surrey Institute for People Centred Artificial Intelligence. His current research interests include signal processing, machine learning/AI, and machine audition. He has authored and co-authored over 400 papers in these areas. His work has been recognized with more than 15 accolades, including the Meta Distinguished Faculty Award, Audio Engineering Society Best Technical Paper Award, IEEE Signal Processing Society Young Author Best Paper Award, DCASE awards, and LVA/ICA Best Student Paper Award. He has been elected to IEEE Fellow for contributions to audio classification, generation and source separation, since 2026. He has been an invited Keynote or Plenary Speaker on about 30 international conferences and workshops.