I have searched the existing issues, both open and closed, to make sure this is not a duplicate report.
The bug
When you query random assets with /api/search/random, the assets returned are strongly biased towards lower UUIDs. The figure below was generated by sampling 100k assets (100 queries of size 1000) from a library of 3.5k images and plotting the relative sample rate of each asset compared to the lexicographic order of its UUID within the database.
Assets with low UUIDs (e.g. 00012fdc-95e5-412f-b8aa-65807961c489) are being sampled ~16x more frequently than assets with high UUIDs (e.g. ffdbd0d4-891a-49cb-a296-2d2ef7311768). The result is that any "random" sample, like a randomized slideshow, returns some assets much more frequently than others.
The OS that Immich Server is running on
Unraid 7.2.2
Version of Immich Server
v2.5.2
Version of Immich Mobile App
N/A
Platform with the issue
Device make and model
No response
Your docker-compose.yml content
name: immich
services:
immich-server:
container_name: immich_server
image: ghcr.io/immich-app/immich-server:${IMMICH_VERSION:-release}
# extends:
# file: hwaccel.transcoding.yml
# service: cpu # set to one of [nvenc, quicksync, rkmpp, vaapi, vaapi-wsl] for accelerated transcoding
volumes:
# Do not edit the next line. If you want to change the media storage location on your system, edit the value of UPLOAD_LOCATION in the .env file
- ${UPLOAD_LOCATION}:/usr/src/app/upload
- /etc/localtime:/etc/localtime:ro
env_file:
- .env
ports:
- '2283:2283'
depends_on:
- redis
- database
restart: always
healthcheck:
disable: false
immich-machine-learning:
container_name: immich_machine_learning
# For hardware acceleration, add one of -[armnn, cuda, rocm, openvino, rknn] to the image tag.
# Example tag: ${IMMICH_VERSION:-release}-cuda
image: ghcr.io/immich-app/immich-machine-learning:${IMMICH_VERSION:-release}
# extends: # uncomment this section for hardware acceleration - see https://immich.app/docs/features/ml-hardware-acceleration
# file: hwaccel.ml.yml
# service: cpu # set to one of [armnn, cuda, rocm, openvino, openvino-wsl, rknn] for accelerated inference - use the `-wsl` version for WSL2 where applicable
volumes:
- model-cache:/cache
env_file:
- .env
restart: always
healthcheck:
disable: false
redis:
container_name: immich_redis
image: docker.io/valkey/valkey:8-bookworm@sha256:ff21bc0f8194dc9c105b769aeabf9585fea6a8ed649c0781caeac5cb3c247884
healthcheck:
test: redis-cli ping || exit 1
restart: always
database:
container_name: immich_postgres
image: ghcr.io/immich-app/postgres:14-vectorchord0.3.0-pgvectors0.2.0@sha256:fa4f6e0971f454cd95fec5a9aaed2ed93d8f46725cc6bc61e0698e97dba96da1
environment:
POSTGRES_PASSWORD: ${DB_PASSWORD}
POSTGRES_USER: ${DB_USERNAME}
POSTGRES_DB: ${DB_DATABASE_NAME}
POSTGRES_INITDB_ARGS: '--data-checksums'
# Uncomment the DB_STORAGE_TYPE: 'HDD' var if your database isn't stored on SSDs
# DB_STORAGE_TYPE: 'HDD'
volumes:
# Do not edit the next line. If you want to change the database storage location on your system, edit the value of DB_DATA_LOCATION in the .env file
- ${DB_DATA_LOCATION}:/var/lib/postgresql/data
restart: always
volumes:
model-cache:
Your .env content
# You can find documentation for all the supported env variables at https://immich.app/docs/install/environment-variables
# The location where your uploaded files are stored
UPLOAD_LOCATION=/mnt/user/Media/Immich
# The location where your database files are stored. Network shares are not supported for the database
DB_DATA_LOCATION=/mnt/user/appdata/postgresql/data
# To set a timezone, uncomment the next line and change Etc/UTC to a TZ identifier from this list: https://en.wikipedia.org/wiki/List_of_tz_database_time_zones#List
# TZ=Etc/UTC
# The Immich version to use. You can pin this to a specific version like "v1.71.0"
IMMICH_VERSION=release
# Connection secret for postgres. You should change it to a random password
# Please use only the characters `A-Za-z0-9`, without special characters or spaces
DB_PASSWORD=redacted
# The values below this line do not need to be changed
###################################################################################
DB_USERNAME=postgres
DB_DATABASE_NAME=immich
Reproduction steps
- Make multiple POST requests to the
search/random endpoint.
- Compare the frequency that each asset is returned to the lexicographic order of its UUID.
This can be done with the following Python script and a running Immich server:
import requests
import pandas as pd
KEY = "API_KEY_HERE"
BASE_URL = "IMMICH_URL_HERE"
# Query 100k random assets in batches of 1k
assets = []
for _ in range(100):
assets += requests.request(
"POST",
BASE_URL + "/api/search/random",
headers={"x-api-key": KEY},
json={"size": 1000},
).json()
# Compare asset UUIDs to their sampling frequency
df = pd.DataFrame(assets)
sorted_ids = df["id"].sort_values().unique().tolist()
(
df["id"].value_counts()
.reset_index(name="count")
.assign(lex_order=lambda x: x["id"].apply(lambda id: sorted_ids.index(id)))
.sort_values("lex_order")
).plot(x="lex_order", y="count")
Relevant log output
Additional information
The root cause seems to be the sampling algorithm used in SearchRepository.searchRandom, which queries assets below and above a random cutoff UUID, unions them, and returns a subset of the union.
I tested that SQL query in a dummy Postgres database and found that on average, the final subset is sampled much more heavily (~90%) from the left side of the union (lower UUIDs) than the right side of the union (higher UUIDs).
I have searched the existing issues, both open and closed, to make sure this is not a duplicate report.
The bug
When you query random assets with
/api/search/random, the assets returned are strongly biased towards lower UUIDs. The figure below was generated by sampling 100k assets (100 queries of size 1000) from a library of 3.5k images and plotting the relative sample rate of each asset compared to the lexicographic order of its UUID within the database.Assets with low UUIDs (e.g.
00012fdc-95e5-412f-b8aa-65807961c489) are being sampled ~16x more frequently than assets with high UUIDs (e.g.ffdbd0d4-891a-49cb-a296-2d2ef7311768). The result is that any "random" sample, like a randomized slideshow, returns some assets much more frequently than others.The OS that Immich Server is running on
Unraid 7.2.2
Version of Immich Server
v2.5.2
Version of Immich Mobile App
N/A
Platform with the issue
Device make and model
No response
Your docker-compose.yml content
Your .env content
Reproduction steps
search/randomendpoint.This can be done with the following Python script and a running Immich server:
Relevant log output
Additional information
The root cause seems to be the sampling algorithm used in
SearchRepository.searchRandom, which queries assets below and above a random cutoff UUID, unions them, and returns a subset of the union.I tested that SQL query in a dummy Postgres database and found that on average, the final subset is sampled much more heavily (~90%) from the left side of the union (lower UUIDs) than the right side of the union (higher UUIDs).