CLAP Sound Retrieval

Contrastive Language-Audio Pretraining for text-prompted sound retrieval and 2D embedding exploration

Project developed as part of the Deep Learning coursework at IMT Atlantique. Code and materials available on GitHub.

CLAP (Contrastive Language-Audio Pretraining) is a multimodal learning strategy that projects audio signals and natural language descriptions into a shared representation space. The model is trained via contrastive objectives to maximize mutual information between corresponding text-audio pairs while repelling non-matching pairs.

With an aligned representation space, the model unlocks:

  • Zero-shot Sound Classification: Classifying audio clips into unseen categories using natural language prompts.
  • Text-Prompted Audio Retrieval: Searching large sound archives using arbitrary descriptive sentences.
  • Interactive Latent Exploration: Visualizing and analyzing clustering behavior of acoustic features.
2D t-SNE projection of the aligned text-audio latent space evaluated on ESC-50.