Skip to content

4.3.1

Latest

Choose a tag to compare

@xenova xenova released this 07 Oct 03:19
511bb61

What's new?

This release adds support for EmbeddingGemma 2. EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs鈥攁nd combinations thereof鈥攊nto a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.

Text search

The feature-extraction pipeline returns normalized embeddings, so the dot product of two embeddings is their cosine similarity:

import { pipeline, matmul } from "@huggingface/transformers";

const extractor = await pipeline("feature-extraction", "onnx-community/embeddinggemma-2-ONNX", {
  device: "webgpu", // or "wasm" (browser) / "cpu" (Node.js)
  dtype: "q4", // see "Choosing a dtype" below
});

const query = "task: search result | query: Which planet is known as the Red Planet?";
const documents = [
  "title: none | text: Venus is often called Earth's twin because of its similar size and proximity.",
  "title: none | text: Mars, known for its reddish appearance, is often referred to as the Red Planet.",
  "title: none | text: Jupiter, the largest planet in our solar system, has a prominent red spot.",
  "title: none | text: Saturn, famous for its rings, is sometimes mistaken for the Red Planet.",
];

const embeddings = await extractor([query, ...documents], { pooling: "mean", normalize: true });
const query_embedding = embeddings.slice([0, 1]);
const document_embeddings = embeddings.slice([1, null]);

const scores = (await matmul(query_embedding, document_embeddings.transpose(1, 0))).tolist()[0];
const ranking = scores.map((score, i) => ({ score, document: documents[i] })).sort((a, b) => b.score - a.score);
console.log(ranking);
// [
//   { score: 0.854, document: "title: none | text: Mars, known for its reddish appearance, ..." },
//   { score: 0.783, document: "title: none | text: Saturn, famous for its rings, ..." },
//   { score: 0.752, document: "title: none | text: Jupiter, the largest planet in our solar system, ..." },
//   { score: 0.684, document: "title: none | text: Venus is often called Earth's twin, ..." },
// ]

Images, audio and video

For other modalities, use the processor and the model directly. Every input maps into the same embedding space, so any embedding can be compared with any other: here, text queries against an image, an audio clip and a video.

import { AutoModel, AutoProcessor, load_image, load_audio, load_video, cat, matmul } from "@huggingface/transformers";

const model_id = "onnx-community/embeddinggemma-2-ONNX";
const processor = await AutoProcessor.from_pretrained(model_id);
const model = await AutoModel.from_pretrained(model_id, { device: "webgpu", dtype: "q4" });

// The processor takes (text, images, audio, videos)
const embed = async (...inputs) => (await model(await processor(...inputs))).sentence_embedding;

// Text queries, with a task prefix (see "Task Instruction Prefixes" below)
const queries = [
  "task: search result | query: cats sleeping on a couch",
  "task: search result | query: a president's speech about serving your country",
  "task: search result | query: a turtle swimming in the ocean",
];
const query_embeddings = await embed(queries);

// An image, an audio clip (mono, 16 kHz) and a video (1 frame per second)
const url = "https://huggingface.co/datasets/Xenova/transformers.js-docs/resolve/main";
const image = await load_image(`${url}/cats.jpg`);
const audio = await load_audio(`${url}/jfk.wav`, 16000);
const video = await load_video(`${url}/sea-turtle.mp4`, { fps: 1 });

const media_embeddings = cat([
  await embed(null, image),
  await embed(null, null, audio),
  await embed(null, null, null, video),
]);

// Embeddings are normalized: the dot product is the cosine similarity
const scores = (await matmul(media_embeddings, query_embeddings.transpose(1, 0))).tolist();
["image", "audio", "video"].forEach((name, i) => console.log(name, scores[i].map((x) => x.toFixed(3))));
// Each input scores highest with its own query:
//          cats     speech   turtle
// image ["0.740", "0.462", "0.508"]
// audio ["0.504", "0.763", "0.487"]
// video ["0.505", "0.500", "0.731"]

To embed several items of a modality at once, pass one list per input: processor(null, [[image1], [image2]]) returns two image embeddings, while a flat list of images, processor(null, [image1, image2]), is a single input made of both images.

Full Changelog: 4.3.0...4.3.1