ICASSP 2026
Audio codecs power discrete music generative modelling, music streaming, and immersive media by shrinking PCM audio to bandwidth-friendly bit rates. Recent works have gravitated toward processing in the spectral domain; however, spectrogram-domain approaches typically struggle with phase modelling, which is naturally complex-valued. We introduce an end-to-end complex-valued RVQ-VAE audio codec that preserves magnitude-phase coupling across the entire analysis-quantization-synthesis pipeline and removes adversarial discriminators and diffusion post-filters. Without GANs or diffusion, the model matches or surpasses much longer-trained baselines in-domain and reaches state-of-the-art out-of-domain performance, reducing the training budget by an order of magnitude while preserving high perceptual quality.