Skip to content
frgfmPublic

About

PyTorch implementations of recent Computer Vision tricks (ReXNet, RepVGG, Unet3p, YOLOv4, CIoU loss, AdaBelief, PolyLoss, MobileOne). Other additions: AdEMAMix

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

328 stars

Watchers

4 watching

Forks

CI Status ruff ty Test coverage percentage

PyPi Status GitHub release (latest by date) pyversions License

Huggingface Spaces Open in Colab

Documentation Status

Implementations of recent Deep Learning tricks in Computer Vision, easily paired up with your favorite framework and model zoo.

Holocrons were information-storage datacron devices used by both the Jedi Order and the Sith that contained ancient lessons or valuable information in holographic form.

Source: Wookieepedia

Quick Tour

Open In Colab

Holocron provides research implementations with PyTorch interfaces. Implementations and training recipes may differ from a paper author's code. This example uses an explicit checkpoint and its preprocessing and labels:

import torch
from PIL import Image
from torchvision.transforms.v2 import Compose, ConvertImageDtype, Normalize, PILToTensor, Resize
from holocron.models.classification import ResNet18_Checkpoint, resnet18

checkpoint = ResNet18_Checkpoint.IMAGENETTE.value
model = resnet18(checkpoint=checkpoint).eval()

image = Image.open(path_to_an_image).convert("RGB")
preprocessing = checkpoint.pre_processing

transform = Compose([
    Resize(preprocessing.input_shape[1:], interpolation=preprocessing.interpolation),
    PILToTensor(),
    ConvertImageDtype(torch.float32),
    Normalize(preprocessing.mean, preprocessing.std),
])

input_tensor = transform(image).unsqueeze(0)

with torch.inference_mode():
    probabilities = model(input_tensor).squeeze(0).softmax(dim=0)

class_idx = probabilities.argmax().item()
label = checkpoint.meta.categories[class_idx]
confidence = probabilities[class_idx].item()
print(label, confidence)

Pretrained models

An architecture may be available without pretrained weights. See the checkpoint list and loading guide for classification datasets, metrics and weight compatibility, and the support status for other tasks. Most classification checkpoints target Imagenette (10 classes); selected ReXNet variants also provide ImageNet-1K weights (1,000 classes).

Use the checkpoint's preprocessing and category metadata, as shown above. A matching model name does not make torchvision or timm weights compatible. To adapt pretrained weights to your own classes, load them before replacing the classifier; see the transfer-learning guide.

Legacy weight loaders log a warning and keep the affected model's initial parameters when no weights are available. YOLO26 builders instead raise ValueError for pretrained=True. No full detection checkpoints are published; a pretrained classification backbone is not a pretrained detector. Train models without weights with the reference scripts.

Loss functions

PolyLoss expects raw logits and either torch.int64 class indices or soft class probabilities. See the loss input guide and runnable example for tensor shapes, target types and the current ignore_index limitations.

Installation

Prerequisites

Python 3.11 (or higher) and uv/pip are required to install Holocron.

Latest stable release

You can install the last stable release of the package using pypi as follows:

pip install pylocron

Developer mode

Alternatively, if you wish to use the latest features of the project that haven't made their way to a release yet, you can install the package from source (install Git first):

git clone https://github.com/frgfm/Holocron.git
pip install -e Holocron/.

Paper references

PyTorch layers for every need

Models for vision tasks

Vision-related operations

Trying something else than Adam

More goodies

Documentation

The full package documentation is available here for detailed specifications.

Demo app

The project includes a minimal demo app using Gradio

demo_app

You can check the live demo, hosted on 🤗 HuggingFace Spaces 🤗 over here 👇 Hugging Face Spaces

Reference scripts

Reference scripts are provided to train your models using holocron on famous public datasets. Those scripts currently support the following vision tasks:

YOLO training status

YOLOv4's training implementation has been repaired and verified with regression tests and a reproducible CPU fixed-batch learning check. With a frozen pretrained Imagenette CSPDarknet53-Mish backbone, it produces correctly labeled detections with IoUs of 0.81 and 0.79 on two synthetic examples, with finite losses and gradients and a saved checkpoint. See the detection guide for the command, matched control, and CUDA/VOC smoke recipe.

Full CUDA/VOC training on the repaired implementation remains pending; the synthetic result does not establish dataset accuracy or reproduce the paper. YOLOv1/v2 loss behavior is unchanged, with regression coverage for shared inference changes. YOLOv3 is not implemented as a detector; the Darknet-53 classification backbone is available. No pretrained detection checkpoints are published.

Latency benchmark

You crave for SOTA performances, but you don't know whether it fits your needs in terms of latency?

The table below contains historical results from an older timing method. Use the script below for current comparisons. Its GPU timer waits for execution to finish.

Arch GPU mean (std) CPU mean (std)
repvgg_a0* 3.14ms (0.87ms) 23.28ms (1.21ms)
repvgg_a1* 4.13ms (1.00ms) 29.61ms (0.46ms)
repvgg_a2* 7.35ms (1.11ms) 46.87ms (1.27ms)
repvgg_b0* 4.23ms (1.04ms) 33.16ms (0.58ms)
repvgg_b1* 12.48ms (0.96ms) 100.66ms (1.46ms)
repvgg_b2* 20.12ms (0.31ms) 155.90ms (1.59ms)
repvgg_b3* 24.94ms (1.70ms) 224.68ms (14.27ms)
rexnet1_0x 6.01ms (0.26ms) 13.66ms (0.21ms)
rexnet1_3x 6.43ms (0.10ms) 19.13ms (2.05ms)
rexnet1_5x 6.46ms (0.28ms) 21.06ms (0.24ms)
rexnet2_0x 6.75ms (0.21ms) 31.77ms (3.28ms)
rexnet2_2x 6.92ms (0.51ms) 33.61ms (0.60ms)
sknet50 11.40ms (0.38ms) 54.03ms (3.35ms)
sknet101 23.55 ms (1.11ms) 94.89ms (5.61ms)
sknet152 69.81ms (0.60ms) 253.07ms (3.33ms)
tridentnet50 16.62ms (1.21ms) 142.85ms (5.33ms)
res2net50_26w_4s 9.25ms (0.22ms) 41.84ms (0.80ms)
resnet50d 36.97ms (3.58ms) 36.97ms (3.58ms)
pyconv_resnet50 20.03ms (0.28ms) 178.85ms (2.35ms)
pyconvhg_resnet50 38.41ms (0.33ms) 301.03ms (12.39ms)
darknet24 3.94ms (1.08ms) 29.39ms (0.78ms)
darknet19 3.17ms (0.59ms) 26.36ms (2.80ms)
darknet53 7.12ms (1.35ms) 53.20ms (1.17ms)
cspdarknet53 6.41ms (0.21ms) 48.05ms (3.68ms)
cspdarknet53_mish 6.88ms (0.51ms) 67.78ms (2.90ms)

The reported latency for RepVGG models is the one of the reparametrized version

This benchmark was performed over 100 iterations on (224, 224) inputs, on a laptop to better reflect performances that can be expected by common users. The hardware setup includes an Intel(R) Core(TM) i7-10750H for the CPU, and a NVIDIA GeForce RTX 2070 with Max-Q Design for the GPU.

Measure a classification model on your hardware:

uv sync --locked --extra scripts
uv run --no-sync python scripts/eval_latency.py rexnet1_0x --device cpu --output /tmp/latency.json

The command measures PyTorch on CPU by default. Use --device cuda:0 or --device mps to measure a GPU. Use --backend onnx to measure ONNX Runtime on CPU. Its export runs in a separate process and uses temporary files. The ONNX worker checks its output against the PyTorch export reference.

Each command runs five fresh worker processes by default. It reports the first forward call, median and 95th-percentile warmed batch latency, and sustained throughput in images per second. GPU timing waits for work to finish. The summary uses medians across workers for timings and throughput. It shows the range of worker medians. Peak RSS is the highest resident process memory across workers, recorded after inference. It includes imports, model setup, and warm-up. It excludes the parent process, ONNX export, and output checks. RSS is unavailable on Windows. CUDA workers also save peak tensor allocations and allocator reservations in the JSON file. RSS and CUDA memory are separate measurements. First-call timing excludes imports and model setup. Compare RSS only between runs using the same runtime and device.

Use --batch-size, --size, --threads, --it, --warmup, and --repeat to set the workload. Inputs use float32. Inference uses evaluation mode, disables gradient tracking, and converts reparametrizable models to their inference form. CPU operations use one thread by default. The fixed seed improves repeatability; it does not prove identical weights or inputs across PyTorch versions. These measurements cover prepared-tensor inference. They exclude image loading, preprocessing, and transfers. They do not measure training or model accuracy.

To compare main with a dependency PR, install each checkout into a separate environment with uv sync --locked. Use the same Python version, machine, runtime, device, and settings. Run the same version of this script against both installations, using an absolute path to the script. Save a JSON file from each environment. The files record the package revision and dirty state, dependency versions, runtime settings, script hash, and individual trials. Check that the CUDA math settings match before comparing GPU results. Alternate baseline and candidate commands, and first run main twice to check normal measurement variation. Compare latency, throughput, and peak RSS per workload; avoid treating small changes within that variation as gains. CI runs short CPU checks for PyTorch and ONNX and uploads their JSON results. CI timing is advisory.

All arguments are listed by python scripts/eval_latency.py --help.

Docker container

If you wish to deploy containerized environments, you can use the provided Dockerfile to build a docker image:

make build-api

Minimal API template

Looking for a boilerplate to deploy a model from Holocron with a REST API? Thanks to the wonderful FastAPI framework, you can do this easily.

Deploy your API locally

Run your API in a docker container as follows:

make start-api

In order to stop the container, use make stop-api

What you have deployed

Your API is now running on port 8080, with its documentation http://localhost:8080/redoc and requestable routes:

import requests

with open("/path/to/your/img.jpeg", "rb") as f:
    data = f.read()
response = requests.post("http://localhost:8080/classification", files={"file": data}).json()

Citation

If you wish to cite this project, feel free to use this BibTeX reference:

@software{Fernandez_Holocron_2020,
author = {Fernandez, François-Guillaume},
month = {5},
title = {{Holocron}},
url = {https://github.com/frgfm/Holocron},
year = {2020}
}

Contributing

Any sort of contribution is greatly appreciated!

You can find a short guide in CONTRIBUTING to help grow this project!

License

Distributed under the Apache 2.0 License. See LICENSE for more information.

About

PyTorch implementations of recent Computer Vision tricks (ReXNet, RepVGG, Unet3p, YOLOv4, CIoU loss, AdaBelief, PolyLoss, MobileOne). Other additions: AdEMAMix

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

328 stars

Watchers

4 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages