Repository navigation
[Feature] Add EmbeddingGemma 2 (sub-1B, multilingual, Apache 2.0) as a Smart Search model #32173
Replies: 4 comments 4 replies
|
This discussion has automatically been closed as it is likely a duplicate. We get a lot of duplicate threads each day, which is why we ask you in the template to confirm that you searched for duplicates before opening one. If you're sure this is not a duplicate, please leave a comment and we will reopen the thread if necessary. |
|
Checked before opening, and this isn't a duplicate of any existing thread. Searches I ran on this repo's discussions (open and closed) for On the closest-looking thread, #31529 ( If the detector flagged it against the general "add a CLIP model for Smart Search" cluster (#11845, #14958, #17135, #25762, #31529), those are all for text-image CLIP/SigLIP models with a 512-d CLIP-style interface. This one is a genuinely different architecture with a different integration path, and the reason I think it's worth the effort is the multilingual ones on that cluster are 3854–4675 MiB while this is 440M with MRL truncation. Happy to reopen the conversation on any part of this if it's not the shape you'd want.
|
|
just wanted to write the exact same! Kudos for your suggestion! |
|
I've run some local benchmarks to compare this to our existing models. The OP claims that this model is good for "small hardware" don't hold up for me, and overall the gains (where they do exist) seem slim and the complexity of adding this model doesn't seem worth it to me - but final call is @mertalev's. Setup: text→image retrieval recall (mean of R@1/5/10), fp32, CPU (i9-12900H).
Crossmodal-3600 per language:
XTD-10 per language:
|
Uh oh!
There was an error while loading. Please reload this page.
The feature
Add google/embeddinggemma-2 as a selectable Smart Search model in
Administration > Settings > Machine Learning > Smart Search.Why this model is a strong fit
Google announced EmbeddingGemma 2 today. For the part Immich cares about, the text + vision config (
{"audio_config": None}) is 440M parameters and maps images and search queries into one shared 768-d space — which is exactly the visual/textual pair splitsmart_searchalready assumes.immich-appHF org.The edge case — this is the part I think matters most
Immich's smart search quality today has a cost floor. Looking at the
searching.mdtables, the multilingual options start at 3854–4675 MiB of model memory (ViT-SO400M-16-SigLIP2-384__webli,nllb-clip-base-siglip__v1), and the English-only defaultViT-B-32__openaiscores 69.9 recall. That means the good multilingual models are effectively locked to a decent server, and the ones that fit on small hardware are the weak ones.EmbeddingGemma 2 is purpose-built to break that tradeoff:
mxfp4/nvfp4/qatbuilds existPut plainly: this is one of the first strong multilingual retrieval models where "good search" and "runs on consumer hardware" stop being a trade-off. Google already ships LiteRT-LM, MLX, Ollama and LMStudio paths for it — the edge story is established, not theoretical.
I'm not claiming Immich users on small hardware are numerous, but right now the model tables effectively assume a server with headroom. If there were a 440M multilingual option that keeps recall near the 3.8 GB models, that changes who can self-host well.
Note
float16is a trap here — the model card is explicit that fp16 exceeds the activation range and returns NaN rather than raising, so it has to be handled in export, not left to users.Audio is out of scope (Immich has no audio search). Video maps cleanly, since Immich already embeds a sampled frame through the vision encoder.
Why this is more than an allow-list entry
Contrast with fg-clip2-base #31529, which was two lines because it is open_clip-compatible. EmbeddingGemma 2 is not, so I think this needs a new
ModelSourcerather than an addition to_OPENCLIP_MODELS. Concretely:tokenizer.jsonpreprocess_cfg.json, fixed center crop (224/256/384/…)encodeText(query)verbatimtask: search result | query: {query}prefixCLIP_MODEL_INFOentry +smart_searchmigrationmodel.onnxper roleTwo of these are behavioural rather than plumbing:
task: search result | query: ...on the query side does not error — it silently degrades ranking quality. That should be baked into the textual encoder, not exposed as an option.What I can do
immich-app/ml-models: a newembeddinggemmasource +models.yamlentrymachine-learning/immich_ml/models/constants.py: register the new sourceGemmaEmbeddingVisualEncoder/GemmaEmbeddingTextualEncoderpair undermodels/torch.onnx.exportin bf16, validatedcos = 1.0against the HF reference on all outputsimmich-appHF org (the ML service pulls fromimmich-app/<model>)searching.mdbenchmark tablesThree questions before I start:
ModelSourceis ongoing maintenance, and I don't yet have Immich-harness numbers to prove 440M beats the 3854 MiB incumbent. If the answer is "not worth it", I'd rather hear that than spend a weekend on an export nobody installs. Happy to run the benchmark first and come back.searching.mdrow with memory, execution time, and recall on Crossmodal-3600 / XTD-10 / Flickr30k, same harness as the existing numbers.smart_searchmigrates for anyone who switches. MRL truncation to 512 keeps the current column, and Google's own numbers put 512d within ~0.2 MTEB of full 768d on text. I lean 512d for a first pass — and it's also the friendlier option for exactly the small-hardware case above — but that's a product call.If a PR is easier to review than a feature request, I'm happy to open the export scaffolding against
immich-app/ml-modelsfirst and let this thread be the decision on #1.Platform
All reactions