fusion-embedding-2-2b-preview v0.2
Generation 2, second release: fine-tuned on the expanded AudioCaps 2.0 training pool.
v0.2 fine-tunes the v0.1 checkpoint's trained components on the AudioCaps 2.0-expanded in-domain pool (86,394 pairs, 1,600 steps; train split verified disjoint from every evaluation protocol we report). Non-audio outputs remain bit-for-bit the base model's.
Highlights
- AudioCaps audio-to-text R@10 0.743 -> 0.759 (R@1 0.302 -> 0.310), text-to-audio 0.775 -> 0.785.
- Clotho zero-shot t2a 0.482 -> 0.485; VGGSound within noise of v0.1 under paired protocols; emergent audio-to-image up ~1.5 points.
- Public MAEB leaderboard results updated to this revision (mean 0.361 vs 0.354 across the nine-task sound-event tier).
Weights and model card: https://huggingface.co/EximiusLabs/fusion-embedding-2-2b-preview (revision tag v0.2-preview; v0.1-preview remains pinned for reproducibility). All gains verified against v0.1 under identical protocols.