Keynote

Ambisonics Signal Processing: From Spherical Array Theory to Machine Learning Challenges

Abstract

Ambisonics is a versatile framework for spatial audio capture, encoding, and reproduction, with growing adoption in industry as the standard format for immersive spatial audio. Ambisonics is based on the spherical harmonics representation of sound fields, with its signal processing having evolved from classical spherical array theory to modern machine learning. This talk reviews Ambisonics encoding approaches as a means to generate Ambisonics signals from microphone recordings, first for spherical arrays, then for wearable arrays and arrays of unconstrained geometry using methods ranging from linear Ambisonics Signal Matching to neural network-based approaches that implicitly leverage information in the audio data. Neural network-based approaches, once incorporated into spatial audio applications, require loss functions that reflect auditory spatial perception, a topic of growing interest that also opens new challenges. These perceptual loss functions are particularly relevant to recent work on Ambisonics upscaling - moving beyond first-order to higher-order representations. Beyond upscaling, Ambisonics further extends to speech enhancement, both for improving spatial audio quality and, more broadly, for enabling array-agnostic processing through its modular, geometry-independent structure. Together, these topics illustrate the growing role of Ambisonics as a unifying representation in spatial audio and point to promising directions for future research.

Keynote Speaker

Prof. Boaz Rafaely

Boaz Rafaely is a Professor at Ben-Gurion University of the Negev in Israel, where he leads the Acoustics Laboratory, conducting research in spatial audio signal processing. He has published over 200 papers in journals and conferences, and is the author of the book Fundamentals of Spherical Array Processing. Prof. Rafaely has previously served as Head of the School of Electrical and Computer Engineering at BGU, Chair of the Israeli Acoustical Association, and Chair of the Technical Committee on Audio Signal Processing of the European Acoustical Association. He has also been an associate editor for IEEE Transactions on Audio, Speech and Language Processing, IEEE Signal Processing Letters, and Acta Acustica. He is an IEEE fellow.

Prof. Boaz Rafaely

Keynote

Organizing the Acoustic World: Computational Principles of Auditory Attention

Abstract

The ability to navigate complex acoustic environments depends on more than recognizing sounds. It requires identifying what is behaviorally relevant, organizing sensory inputs into meaningful objects, and adapting to changing goals and contexts. This talk will examine the computational principles that support these abilities, drawing on studies of auditory attention, predictive processing, and auditory scene analysis. I will present behavioral and neural evidence illustrating how top-down goals, statistical learning, and bottom-up organization interact to shape auditory representations and guide selective listening in natural acoustic scenes. I will conclude by discussing how these same principles are beginning to influence modern audio systems, offering opportunities for closer integration between auditory neuroscience and machine listening.

Keynote Speaker

Prof. Mounya Elhilali

Mounya Elhilali received her Ph.D. degree in Electrical and Computer Engineering from the University of Maryland, College Park in 2004. She is the Charles Renn faculty scholar and professor of Electrical and Computer Engineering at the Johns Hopkins University, with joint appointment in the department of Psychology and Brain Sciences. She directs the Laboratory for Computational Audio Perception and is affiliated with the center for speech and language processing, the center for hearing and balance and the Kavli neuroscience institute. Her research examines sound processing by humans and machines in noisy soundscapes, and investigates reverse engineering intelligent processing of sounds by brain networks with applications to speech and audio technologies and medical diagnostic systems. Dr. Elhilali is the recipient of the National Science Foundation CAREER award and the Office of Naval Research Young Investigator award. She was recognized as outstanding woman innovator by the Johns Hopkins tech ventures in 2020.

Prof. Mounya Elhilali

Keynote

Language Queried Audio Source Separation

Abstract

Large audio-language models (LALMs) are emerging as a powerful approach for addressing a broad range of audio signal processing problems, including generation, captioning, editing, and source separation. By integrating large language models (LLMs) with audio modeling techniques, they enable more flexible and intelligent audio systems. This talk focuses on language-queried audio source separation (LASS), a recently introduced framework that extracts desired sounds from an audio mixture using natural language queries. LASS offers an intuitive and scalable interface for applications such as automated editing, remixing, and audio rendering. We start by outlining the problem and its motivation, making a connection with conventional paradigms for speech and universal source separation. We then present two LASS methods: AudioSep and FlowSep. AudioSep is a text-driven model that uses query and separation networks to estimate time–frequency masks and extract target sounds. FlowSep is a generative approach based on rectified flow matching (RFM), modeling transitions from noise to target features in a variational autoencoder (VAE) latent space and reconstructing audio via a decoder and vocoder. We also present datasets, evaluation metrics, experimental results, and sound demos. Finally, we link source separation to audio editing and generation, highlighting recent advances (e.g. the RFM-Editing, AudioLDM, and AudioLDM2 models), and conclude the talk with future research directions in this evolving field.

Keynote Speaker

Prof. Wenwu Wang

Wenwu Wang is a Professor in Signal Processing and Machine Learning, and the Associate Head of External Engagement, School of Computer Science and Electronic Engineering, University of Surrey, UK. He is also a Core AI Fellow at the Surrey Institute for People Centred Artificial Intelligence. His current research interests include signal processing, machine learning/AI, and machine audition (listening). He has (co)-authored over 400 papers in these areas. His work has been recognized with more than 15 accolades, including the Meta Distinguished Faculty Award (2026), Audio Engineering Society Best Technical Paper Award (2025), IEEE Signal Processing Society Young Author Best Paper Award (2022), DCASE Judge’s Award (2020, 2023, and 2024), DCASE Reproducible System Award (2019 and 2020), and LVA/ICA Best Student Paper Award (2018). He has been elected to IEEE Fellow for contributions to audio classification, generation and source separation, since 2026. He has been an invited Keynote or Plenary Speaker on about 30 international conferences and workshops. More about his work can be found from his personal page: https://personalpages.surrey.ac.uk/w.wang/.

Prof. Mounya Elhilali