EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding

ICASSP 2026

Abstract

Audio codecs power discrete music generative modelling, music streaming, and immersive media by shrinking PCM audio to bandwidth-friendly bit rates. Recent works have gravitated toward processing in the spectral domain; however, spectrogram-domain approaches typically struggle with phase modelling, which is naturally complex-valued. We introduce an end-to-end complex-valued RVQ-VAE audio codec that preserves magnitude-phase coupling across the entire analysis-quantization-synthesis pipeline and removes adversarial discriminators and diffusion post-filters. Without GANs or diffusion, the model matches or surpasses much longer-trained baselines in-domain and reaches state-of-the-art out-of-domain performance, reducing the training budget by an order of magnitude while preserving high perceptual quality.

Publication
ICASSP 2026
Michele Mancusi
Michele Mancusi
Senior Applied Scientist, PhD

Senior Applied Scientist at Music.AI (Moises), developing and optimizing autoencoders and generative models for speech, music, and general audio enhancement