Skip to content

/api/search/random results are biased by asset UUID #26049

Description

@aazuspan

I have searched the existing issues, both open and closed, to make sure this is not a duplicate report.

  • Yes

The bug

When you query random assets with /api/search/random, the assets returned are strongly biased towards lower UUIDs. The figure below was generated by sampling 100k assets (100 queries of size 1000) from a library of 3.5k images and plotting the relative sample rate of each asset compared to the lexicographic order of its UUID within the database.

Image

Assets with low UUIDs (e.g. 00012fdc-95e5-412f-b8aa-65807961c489) are being sampled ~16x more frequently than assets with high UUIDs (e.g. ffdbd0d4-891a-49cb-a296-2d2ef7311768). The result is that any "random" sample, like a randomized slideshow, returns some assets much more frequently than others.

The OS that Immich Server is running on

Unraid 7.2.2

Version of Immich Server

v2.5.2

Version of Immich Mobile App

N/A

Platform with the issue

  • Server
  • Web
  • Mobile

Device make and model

No response

Your docker-compose.yml content

name: immich

services:
  immich-server:
    container_name: immich_server
    image: ghcr.io/immich-app/immich-server:${IMMICH_VERSION:-release}
    # extends:
    #   file: hwaccel.transcoding.yml
    #   service: cpu # set to one of [nvenc, quicksync, rkmpp, vaapi, vaapi-wsl] for accelerated transcoding
    volumes:
      # Do not edit the next line. If you want to change the media storage location on your system, edit the value of UPLOAD_LOCATION in the .env file
      - ${UPLOAD_LOCATION}:/usr/src/app/upload
      - /etc/localtime:/etc/localtime:ro
    env_file:
      - .env
    ports:
      - '2283:2283'
    depends_on:
      - redis
      - database
    restart: always
    healthcheck:
      disable: false

  immich-machine-learning:
    container_name: immich_machine_learning
    # For hardware acceleration, add one of -[armnn, cuda, rocm, openvino, rknn] to the image tag.
    # Example tag: ${IMMICH_VERSION:-release}-cuda
    image: ghcr.io/immich-app/immich-machine-learning:${IMMICH_VERSION:-release}
    # extends: # uncomment this section for hardware acceleration - see https://immich.app/docs/features/ml-hardware-acceleration
    #   file: hwaccel.ml.yml
    #   service: cpu # set to one of [armnn, cuda, rocm, openvino, openvino-wsl, rknn] for accelerated inference - use the `-wsl` version for WSL2 where applicable
    volumes:
      - model-cache:/cache
    env_file:
      - .env
    restart: always
    healthcheck:
      disable: false

  redis:
    container_name: immich_redis
    image: docker.io/valkey/valkey:8-bookworm@sha256:ff21bc0f8194dc9c105b769aeabf9585fea6a8ed649c0781caeac5cb3c247884
    healthcheck:
      test: redis-cli ping || exit 1
    restart: always

  database:
    container_name: immich_postgres
    image: ghcr.io/immich-app/postgres:14-vectorchord0.3.0-pgvectors0.2.0@sha256:fa4f6e0971f454cd95fec5a9aaed2ed93d8f46725cc6bc61e0698e97dba96da1
    environment:
      POSTGRES_PASSWORD: ${DB_PASSWORD}
      POSTGRES_USER: ${DB_USERNAME}
      POSTGRES_DB: ${DB_DATABASE_NAME}
      POSTGRES_INITDB_ARGS: '--data-checksums'
      # Uncomment the DB_STORAGE_TYPE: 'HDD' var if your database isn't stored on SSDs
      # DB_STORAGE_TYPE: 'HDD'
    volumes:
      # Do not edit the next line. If you want to change the database storage location on your system, edit the value of DB_DATA_LOCATION in the .env file
      - ${DB_DATA_LOCATION}:/var/lib/postgresql/data
    restart: always

volumes:
  model-cache:

Your .env content

# You can find documentation for all the supported env variables at https://immich.app/docs/install/environment-variables

# The location where your uploaded files are stored
UPLOAD_LOCATION=/mnt/user/Media/Immich

# The location where your database files are stored. Network shares are not supported for the database
DB_DATA_LOCATION=/mnt/user/appdata/postgresql/data

# To set a timezone, uncomment the next line and change Etc/UTC to a TZ identifier from this list: https://en.wikipedia.org/wiki/List_of_tz_database_time_zones#List
# TZ=Etc/UTC

# The Immich version to use. You can pin this to a specific version like "v1.71.0"
IMMICH_VERSION=release

# Connection secret for postgres. You should change it to a random password
# Please use only the characters `A-Za-z0-9`, without special characters or spaces
DB_PASSWORD=redacted

# The values below this line do not need to be changed
###################################################################################
DB_USERNAME=postgres
DB_DATABASE_NAME=immich

Reproduction steps

  1. Make multiple POST requests to the search/random endpoint.
  2. Compare the frequency that each asset is returned to the lexicographic order of its UUID.

This can be done with the following Python script and a running Immich server:

import requests
import pandas as pd

KEY = "API_KEY_HERE"
BASE_URL = "IMMICH_URL_HERE"

# Query 100k random assets in batches of 1k
assets = []
for _ in range(100):
    assets += requests.request(
        "POST", 
        BASE_URL + "/api/search/random", 
        headers={"x-api-key": KEY}, 
        json={"size": 1000},
    ).json()

# Compare asset UUIDs to their sampling frequency
df = pd.DataFrame(assets)
sorted_ids = df["id"].sort_values().unique().tolist()
(
    df["id"].value_counts()
    .reset_index(name="count")
    .assign(lex_order=lambda x: x["id"].apply(lambda id: sorted_ids.index(id)))
    .sort_values("lex_order")
).plot(x="lex_order", y="count")

Relevant log output

Additional information

The root cause seems to be the sampling algorithm used in SearchRepository.searchRandom, which queries assets below and above a random cutoff UUID, unions them, and returns a subset of the union.

I tested that SQL query in a dummy Postgres database and found that on average, the final subset is sampled much more heavily (~90%) from the left side of the union (lower UUIDs) than the right side of the union (higher UUIDs).

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    Status
    Done

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions