P2-01: Comparing Training Objectives for Neural Post-Filtering of Coded Music

Heinmüller, Alexander*, Brendel, Andreas, Delgado, Pablo, Herre, Jürgen

Subjects (starting with primary): Speech and audio coding

Presented in Poster Session 2

Abstract:

Neural post-filters are an effective tool for improving the reconstruction quality of audio codecs, and hence, numerous models have been proposed in the literature. In this work, we compare several training paradigms—discriminative training, generative adversarial networks (GANs), score-based diffusion (SD), and conditional flow matching—for post-filtering speech and music, and we evaluate the resulting quality improvements across different perceptual audio codecs. We show that GANs can achieve music quality similar to that of SD models at a fraction of the computational complexity. Since the characteristics of a music signal are relevant for the performance of a codec, we further investigate how neural post-filters behave across different signal types.

This PDF is password protected. Enter the proceedings password in the PDF viewer when prompted.