Skip to content

PaulKinlan/image-embedding-lab

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

55 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Image Embedding Lab

A hands-on lab for building intuition about what image embeddings actually capture. Upload an image, transform it (rotate, crop, scale, flip, compress…), and watch how much the embedding moves. Everything runs in your browser via Transformers.js (and there's a Node CLI for batch runs) — no server, no API key.

Live: https://paulkinlan.github.io/image-embedding-lab/ · How it works (explainer) · Reverse search · Vector visualizer · VLM playground · Batch report · Tiling report · Layout report · Findings

The Reverse image search demo (search.html) puts the whole thesis to work: an 800-image corpus embedded offline with six encoders — CLIP, SigLIP, DINOv2, Florence-2, Gemma 4 (see lib/gemma4.mjs for the story of getting it working) and Gemini Embedding 2 (scripts/build-search-index.mjs + build-search-index-gemma4.mjs), packed into one SQLite file that the page loads with sql.js and queries entirely client-side. Pick any query image, crop and rotate it with live re-search, and compare the models' rankings side by side: every hit shows its cosine, hits above each model's same-image threshold (from the layout experiments) are marked, and the query's source image is outlined so you can watch it fall down the rankings as the transform deepens. Search is exact (one dot product per corpus vector, sub-millisecond at this scale); the page explains where approximate indexes like HNSW actually earn their keep.

New to embeddings? Start with the interactive explainer — it builds the whole mental model from scratch (vectors and cosine, patches and tokens, position embeddings, pooling, training objectives, floors), each concept as a widget you can drag or scramble, then puts this repo's experimental results in context — including "what cosine threshold lets you claim two images are the same?", answered per model with an interactive threshold explorer backed by results-same-image.json (scripts/experiment-same-image-threshold.mjs).

The VLM playground (vlm.html) runs Florence-2 in the browser — caption an image, read its text (OCR), or detect objects. It's the counterpart to the embedding lab: where the lab reduces an image to one vector, a VLM turns it into text, so you can watch it read a page an embedding model can only see as "a document".

The Vector visualizer (vector-viz.html) goes the other way: instead of measuring an embedding, it draws one. Embed an image (CLIP/SigLIP/DINOv2, or Florence-2's DaViT encoder — which, being a VLM, also captions each image it embeds, so you see the model's vector and what it says it sees) or a text string (CLIP — text and image share a space) and render the vector as a 2D image, in fourteen modes: consecutive triples of dimensions as RGB pixels, a heatmap grid, Hilbert-curve and square-spiral walks (neighbouring dims stay neighbouring pixels), a barcode, an audio-style waveform, the vector's self-similarity matrix (vᵢ·vⱼ) and recurrence plot (|vᵢ−vⱼ|), a DFT frequency spectrum of the vector-as-signal, dimension pairs as complex phase/magnitude, a seeded random-projection "fingerprint", an interference texture (each dimension a golden-angle plane wave), or a radial flower. Normalization, dimension ordering, and colormap are all swappable lenses. Load two inputs (any mix of image and text) to get the cosine similarity, a diverging diff image, and a slerp morph slider between the two vectors. A rotation sweep re-embeds an image at every angle and stacks the vectors into a rotation×dimension image (jointly normalized; delta view highlights which dimensions move under rotation), with per-angle cosines and a filmstrip. Hover any pixel to see which dimension it came from and its value. The pure math lives in lib/vector-viz.mjs (tested by scripts/test-vector-viz.mjs); embeddings use the same pooling/normalization as the lab via lib/experiment.mjs, so what you see here is exactly the vector the lab measures. None of these images is "what the model sees" — they're lenses on a point in embedding space; the interesting part is comparing structure across inputs and models.

What this is really about

The starting question was simple: are image embeddings invariant to transformations? If you rotate or crop an image, does a vision encoder still produce "the same" vector? Text embeddings famously do this for meaning — "the cat sat on the mat" and "a feline rested on a rug" land in nearly the same place. Do image encoders treat a rotated photo as the same photo?

Answering it turned into a tour of how these models — and the way you measure them — actually work. This README is the running log of the questions and what each one taught me.

The questions I've been chasing

1. Are the embeddings transformation-invariant? Partly. Flips and photometric changes (brightness, grayscale, blur, occlusion) barely move the embedding; rotation moves it the most. But "how much is a lot?" turned out to depend entirely on the next two questions.

2. Was I even measuring the right thing? (the big one) No — at first. image-feature-extraction returns different shapes per model: CLIP gives a pooled [1, 512] vector, but SigLIP returns [1, 196, 768] and DINOv2 [1, 257, 384] — the raw, unpooled grid of patch tokens. Comparing a flattened patch grid means a rotation just shuffles which patch is at which index, so cosine collapses — and smaller patches (more tokens) collapse harder. That single mistake manufactured a confident-but-wrong headline ("patch size is the dominant factor"). The fix: mean-pool to one vector per image before comparing. After that, SigLIP ties CLIP on rotation and the "patch size" story evaporates. Lesson: with embeddings, how you pool is part of the model.

3. How low is "low"? (the baseline floor) A 70% cosine sounds high — but high compared to what? So the lab now measures the different- image floor: the cosine between unrelated images. It varies wildly — CLIP ~58%, DINOv2 ~15%. That reframes everything: CLIP's high scores are partly just a compressed embedding space where even random images sit at 58%. Every bar in the UI now has a floor tick; a result only means "invariant" if it's well to the right of the floor, not just near 100%.

4. Does JPEG compression change the picture? Each transform is embedded at raw pixels and at JPEG q0.9 / q0.5 / q0.2, so compression is its own axis. So far: at q0.9 it's negligible (~1 point); the interesting question is whether invariance holds as quality drops toward q0.2.

5. Why is everything squished to 224×224? How do LLMs read text then? Because these encoders are fixed at 224 — the ViT patch grid and its positional embeddings are baked in at training, so they physically can't take a bigger image. At 224 a patch is 14–32px, far too coarse for text. That's why these are semantic encoders, not OCR. The big VLMs (GPT-4V, Claude, Qwen2-VL, InternVL, Gemma 3/4) read text by using higher native resolution and tiling — splitting a page into many crops, encoding each, and concatenating thousands of vision tokens. OCR comes from throwing far more patches at the image, not a smarter low-res pass.

6. What do the different encoders "see"?

  • CLIP (OpenAI, 2021) — trained on image↔caption pairs, so it's language-aligned.
  • SigLIP (Google) — same idea, sigmoid loss, generally sharper.
  • DINOv2 (Meta) — self-supervised, never told about language; a good contrast.
  • Gemma 4 vision (via ONNX Runtime) — the encoder pulled out of a VLM. Parked as "broken" for days until one session.inputNames inspection revealed its ONNX takes pre-patchified input (flattened 16px patches + position ids), not an image; preprocessing ported in lib/gemma4.mjs and it now works in the browser and the search index. Comparing language-aligned vs self-supervised is a big part of the fun.

7. Semantic invariance vs surface form. Even fixed, no encoder treats a rotated image as identical — a ViT's positional embeddings are absolute, so a turn is a different input. Image encoders lack the paraphrase signal text encoders get ("this rotated image means the same thing" is never in the training data). That framing holds; the earlier mistake was only in the magnitudes.

See FINDINGS.md for the current numbers and report.html for the batch charts.

Using it

Browser — open the live link (or serve locally, below). Pick an encoder (224px, higher-res 384/336, or experimental Gemma 4), choose a tiling level (whole image / 2×2 / 3×3), drop an image, and either drag the sliders + "Test this transform" for a single check, or "Compare all transforms" for the full sweep. A progress bar shows each embedding as it runs; results show the raw-pixel bar with the JPEG q90/50/20 values beneath, plus the different-image floor.

Tiling splits the image into crops, embeds each at full model resolution, and pools them — so the encoder effectively sees 2×/3× the detail. It's the AnyRes trick VLMs use to read text: try a webpage screenshot at 1×1 vs 3×3 and watch the different-image floor drop as the model starts to tell documents apart.

CLI (batch, Node):

npm install
node cli.mjs test-images/*.jpg --model clip                 # human-readable table
node cli.mjs test-images/*  --model siglip384               # higher-res encoder
node cli.mjs test-images/*  --model dinov2 --tiles 3         # 3×3 AnyRes tiling
node cli.mjs test-images/*  --model dinov2 --json > results-dinov2.json
python3 scripts/generate_report.py                          # regenerate report.html

Models: clip, siglip, dinov2 (224px), siglip384, clip-l (higher-res). --tiles N for AnyRes tiling.

Both share lib/experiment.mjs (transforms, pooling, cosine, JPEG variants) so the browser and CLI measure identically.

Run locally

python3 -m http.server 8000   # then open http://localhost:8000

Open questions / where this is going

  • Tiling + higher-resolution encoders — now in: SigLIP-384 / CLIP-L-336 plus a 2×2 / 3×3 AnyRes tiling mode. Next: measure the floor drop on documents systematically.
  • Does the embedding encode what-is-where, or just what? — the layout report runs seven experiments (patch shuffle, orientation probe, token equivariance, layout swaps, caption stability, mirrored text, caption retrieval). Short version: layout is in every embedding; how much the pooled vector moves measures the aggregation (CLS vs mean-pool), not the encoder; DINOv2 is substantially equivariant while Florence entangles position into its token features; and semantic robustness ≠ vector robustness (CLIP's meaning survives shuffles that wreck its cosine, Florence's cosine survives rotations that wreck its captions).
  • Text↔image alignment — experiment 7 of the layout report: own-caption retrieval survives transforms far better than image↔image cosine implies (12/15 under a 4×4 shuffle).
  • Isolate patch size properly — compare within one family (CLIP-B/16 vs B/32) instead of across three models that differ on everything.
  • CLS vs mean-pool within one model — the layout report's biggest confound, and now its most interesting follow-up: extract both from DINOv2/SigLIP and re-run the shuffle test.
  • Heavier compression / other corruptions — noise, resize artifacts, screenshots of screenshots.

Suggestions and questions welcome — this is a learning project as much as a benchmark.

About

https://paulkinlan.github.io/image-embedding-lab/

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

No releases published

Packages

 
 
 

Contributors