This release provides open weights for the following models, introduced in SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval:
SonicCLAP-AR, optimized for audio-text retrieval.SonicCLAP-MOS, optimized for alignment with human perceptual judgments.
Both models were trained on the SonicCaps dataset, with different caption sampling strategies.
.zip files are meant to be decompressed in the root repo folder and each will extract to the right checkpoints and samples folders.
All open weights in this page are released under the CC-BY-NC license.