Michele Mancusi

Michele Mancusi

Senior Applied Scientist, PhD

Music.AI (Moises)

Hello world! đź‘‹

I’m Michele Mancusi, a Senior Applied Scientist at Music.AI (Moises), where I develop and optimize autoencoders and generative models for speech, music, and general audio enhancement.

Previously, I was a Research Scientist at Sony, where I conducted research on deep-learning-based generative models for speech, audio, and music, including Large Language Model (LLM) and diffusion-based approaches.

Earlier in my career, I interned at Microsoft and Musixmatch. At Microsoft, I worked on deep learning for unsupervised speech separation, while at Musixmatch I focused on singing voice detection.

I earned my Ph.D. from Sapienza University of Rome under the supervision of Prof. Emanuele RodolĂ  as a member of the Gladia research group. My doctoral research centered on music generation, source separation, and Natural Language Processing (NLP), contributing to advancements in the field of generative AI.

Interests
  • Deep Learning
  • Signal Processing
  • Generative AI
  • Music Generation
  • Source Separation
  • NLP
  • Speech Synthesis
Education

Work Experience

 
 
 
 
 
Music.AI (Moises)
Senior Applied Scientist
Apr 2025 – Present
  • Developing and optimizing autoencoders and generative models for speech, music, and general audio enhancement.
 
 
 
 
 
Sony
Research Scientist
Apr 2024 – Mar 2025 Stuttgart, Germany
  • Conducted research on deep-learning-based generative models for speech, audio, and music, including LLM- and diffusion-based approaches.
 
 
 
 
 
Sony
Visiting Research Scientist
Nov 2023 – Jan 2024 Stuttgart, Germany
  • Worked with Dr. Stefan Uhlich in the AI, Speech and Sound Group. Conducted research on deep learning for effects removal and timbre transfer with diffusion models.
 
 
 
 
 
Microsoft
Research Scientist Intern
Jun 2023 – Sep 2023 Redmond, Washington, USA
  • Worked with Dr. Sebastian Braun in the Audio and Acoustics Research Group. Conducted research on deep learning for unsupervised speech separation.
 
 
 
 
 
Musixmatch
Data Scientist Intern
Sep 2022 – Mar 2023 Bologna, Italy
  • Worked with Dr. Loreto Parisi in the AI Team. Conducted research on deep learning for singing voice detection.

Publications

PHALAR: Phasors for Learned Musical Audio Representations
ICML 2026
SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and Personalized Music Editing
AAAI 2026
COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations
ICASSP 2025
High-Resolution Speech Restoration with Latent Diffusion Model
ICASSP 2025

Achievements

  • First-author paper selected among the top 1% submissions for an oral presentation at ICLR 2024
  • Awarded €50,000 in AWS credit for the best generative AI research project
  • Awarded €20,000 for the best machine translation research project
  • Recognized as one of the top doctoral research projects and awarded research funding

Academic Experience

OsnabrĂĽck University
Invited Talk
  • International Symposium on Diffusion Models for Audio and Music Processing
  • OsnabrĂĽck, Germany
  • Lectured at the Deep Learning and Applied AI MSc course and mentored students for their Master’s Thesis
  • Delivered a guest lecture on the Latent Autoregressive Source Separation paper upon the invitation of Prof. Ronald Coifman

Attended

ISMIR 2025: 26th International Society for Music Information Retrieval Conference
AAAI 2023: The 37th AAAI Conference on Artificial Intelligence
4th International Summer School on AI and Games
IRDTA: 4th International School on Deep Learning
ACDL: 3rd Advanced Course on Data Science & Machine Learning

Skills

lightning
Lightning
wandb
Weights & Biases
pytorch
PyTorch
hydra
Hydra
aws
AWS SageMaker
azure
Azure ML
google-cloud
Google Cloud Platform
slurm
Slurm
condor
HTCondor
kubernetes
Kubernetes

Contact