Implementations of recent Deep Learning tricks in Computer Vision, easily paired up with your favorite framework and model zoo.
Holocrons were information-storage datacron devices used by both the Jedi Order and the Sith that contained ancient lessons or valuable information in holographic form.
Source: Wookieepedia
Holocron provides research implementations with PyTorch interfaces. Implementations and training recipes may differ from a paper author's code. This example uses an explicit checkpoint and its preprocessing and labels:
import torch
from PIL import Image
from torchvision.transforms.v2 import Compose, ConvertImageDtype, Normalize, PILToTensor, Resize
from holocron.models.classification import ResNet18_Checkpoint, resnet18
checkpoint = ResNet18_Checkpoint.IMAGENETTE.value
model = resnet18(checkpoint=checkpoint).eval()
image = Image.open(path_to_an_image).convert("RGB")
preprocessing = checkpoint.pre_processing
transform = Compose([
Resize(preprocessing.input_shape[1:], interpolation=preprocessing.interpolation),
PILToTensor(),
ConvertImageDtype(torch.float32),
Normalize(preprocessing.mean, preprocessing.std),
])
input_tensor = transform(image).unsqueeze(0)
with torch.inference_mode():
probabilities = model(input_tensor).squeeze(0).softmax(dim=0)
class_idx = probabilities.argmax().item()
label = checkpoint.meta.categories[class_idx]
confidence = probabilities[class_idx].item()
print(label, confidence)An architecture may be available without pretrained weights. See the checkpoint list and loading guide for classification datasets, metrics and weight compatibility, and the support status for other tasks. Most classification checkpoints target Imagenette (10 classes); selected ReXNet variants also provide ImageNet-1K weights (1,000 classes).
Use the checkpoint's preprocessing and category metadata, as shown above.
A matching model name does not make torchvision or timm weights compatible.
To adapt pretrained weights to your own classes, load them before replacing the
classifier; see the transfer-learning guide.
Legacy weight loaders log a warning and keep the affected model's initial
parameters when no weights are available. YOLO26 builders instead raise
ValueError for pretrained=True. No full detection checkpoints are published; a pretrained
classification backbone is not a pretrained detector. Train models without
weights with the reference scripts.
PolyLoss expects raw logits and either torch.int64 class indices or soft
class probabilities. See the loss input guide and runnable example
for tensor shapes, target types and the current ignore_index limitations.
Python 3.11 (or higher) and uv/pip are required to install Holocron.
You can install the last stable release of the package using pypi as follows:
pip install pylocronAlternatively, if you wish to use the latest features of the project that haven't made their way to a release yet, you can install the package from source (install Git first):
git clone https://github.com/frgfm/Holocron.git
pip install -e Holocron/.- Activation: HardMish, NLReLU, FReLU
- Loss: Focal Loss, MultiLabelCrossEntropy, MixupLoss, ClassBalancedWrapper, ComplementCrossEntropy, MutualChannelLoss, DiceLoss, PolyLoss
- Convolutions: NormConv2d, Add2d, SlimConv2d, PyConv2d, Involution
- Regularization: DropBlock
- Pooling: BlurPool2d, SPP, ZPool
- Attention: SAM, LambdaLayer, TripletAttention
- Image Classification: Res2Net (based on the great implementation from Ross Wightman), Darknet-24, Darknet-19, Darknet-53, CSPDarknet-53, ResNet, ResNeXt, TridentNet, PyConvResNet, ReXNet, SKNet, RepVGG, ConvNeXt, MobileOne.
- Object Detection: YOLOv1, YOLOv2, YOLOv4
- Semantic Segmentation: U-Net, UNet++, UNet3+
- Optimizer: LARS, Lamb, TAdam, AdamP, AdaBelief, Adan, and customized versions (RaLars), AdEMAMix
- Optimizer wrapper: Lookahead, Scout (experimental)
The full package documentation is available here for detailed specifications.
The project includes a minimal demo app using Gradio
You can check the live demo, hosted on 🤗 HuggingFace Spaces 🤗 over here 👇
Reference scripts are provided to train your models using holocron on famous public datasets. Those scripts currently support the following vision tasks:
YOLOv4's training implementation has been repaired and verified with regression tests and a reproducible CPU fixed-batch learning check. With a frozen pretrained Imagenette CSPDarknet53-Mish backbone, it produces correctly labeled detections with IoUs of 0.81 and 0.79 on two synthetic examples, with finite losses and gradients and a saved checkpoint. See the detection guide for the command, matched control, and CUDA/VOC smoke recipe.
Full CUDA/VOC training on the repaired implementation remains pending; the synthetic result does not establish dataset accuracy or reproduce the paper. YOLOv1/v2 loss behavior is unchanged, with regression coverage for shared inference changes. YOLOv3 is not implemented as a detector; the Darknet-53 classification backbone is available. No pretrained detection checkpoints are published.
You crave for SOTA performances, but you don't know whether it fits your needs in terms of latency?
The table below contains historical results from an older timing method. Use the script below for current comparisons. Its GPU timer waits for execution to finish.
| Arch | GPU mean (std) | CPU mean (std) |
|---|---|---|
| repvgg_a0* | 3.14ms (0.87ms) | 23.28ms (1.21ms) |
| repvgg_a1* | 4.13ms (1.00ms) | 29.61ms (0.46ms) |
| repvgg_a2* | 7.35ms (1.11ms) | 46.87ms (1.27ms) |
| repvgg_b0* | 4.23ms (1.04ms) | 33.16ms (0.58ms) |
| repvgg_b1* | 12.48ms (0.96ms) | 100.66ms (1.46ms) |
| repvgg_b2* | 20.12ms (0.31ms) | 155.90ms (1.59ms) |
| repvgg_b3* | 24.94ms (1.70ms) | 224.68ms (14.27ms) |
| rexnet1_0x | 6.01ms (0.26ms) | 13.66ms (0.21ms) |
| rexnet1_3x | 6.43ms (0.10ms) | 19.13ms (2.05ms) |
| rexnet1_5x | 6.46ms (0.28ms) | 21.06ms (0.24ms) |
| rexnet2_0x | 6.75ms (0.21ms) | 31.77ms (3.28ms) |
| rexnet2_2x | 6.92ms (0.51ms) | 33.61ms (0.60ms) |
| sknet50 | 11.40ms (0.38ms) | 54.03ms (3.35ms) |
| sknet101 | 23.55 ms (1.11ms) | 94.89ms (5.61ms) |
| sknet152 | 69.81ms (0.60ms) | 253.07ms (3.33ms) |
| tridentnet50 | 16.62ms (1.21ms) | 142.85ms (5.33ms) |
| res2net50_26w_4s | 9.25ms (0.22ms) | 41.84ms (0.80ms) |
| resnet50d | 36.97ms (3.58ms) | 36.97ms (3.58ms) |
| pyconv_resnet50 | 20.03ms (0.28ms) | 178.85ms (2.35ms) |
| pyconvhg_resnet50 | 38.41ms (0.33ms) | 301.03ms (12.39ms) |
| darknet24 | 3.94ms (1.08ms) | 29.39ms (0.78ms) |
| darknet19 | 3.17ms (0.59ms) | 26.36ms (2.80ms) |
| darknet53 | 7.12ms (1.35ms) | 53.20ms (1.17ms) |
| cspdarknet53 | 6.41ms (0.21ms) | 48.05ms (3.68ms) |
| cspdarknet53_mish | 6.88ms (0.51ms) | 67.78ms (2.90ms) |
The reported latency for RepVGG models is the one of the reparametrized version
This benchmark was performed over 100 iterations on (224, 224) inputs, on a laptop to better reflect performances that can be expected by common users. The hardware setup includes an Intel(R) Core(TM) i7-10750H for the CPU, and a NVIDIA GeForce RTX 2070 with Max-Q Design for the GPU.
Measure a classification model on your hardware:
uv sync --locked --extra scripts
uv run --no-sync python scripts/eval_latency.py rexnet1_0x --device cpu --output /tmp/latency.jsonThe command measures PyTorch on CPU by default. Use --device cuda:0 or --device mps to measure a GPU.
Use --backend onnx to measure ONNX Runtime on CPU. Its export runs in a separate process and uses temporary files.
The ONNX worker checks its output against the PyTorch export reference.
Each command runs five fresh worker processes by default. It reports the first forward call, median and 95th-percentile warmed batch latency, and sustained throughput in images per second. GPU timing waits for work to finish. The summary uses medians across workers for timings and throughput. It shows the range of worker medians. Peak RSS is the highest resident process memory across workers, recorded after inference. It includes imports, model setup, and warm-up. It excludes the parent process, ONNX export, and output checks. RSS is unavailable on Windows. CUDA workers also save peak tensor allocations and allocator reservations in the JSON file. RSS and CUDA memory are separate measurements. First-call timing excludes imports and model setup. Compare RSS only between runs using the same runtime and device.
Use --batch-size, --size, --threads, --it, --warmup, and --repeat to set the workload.
Inputs use float32. Inference uses evaluation mode, disables gradient tracking, and converts reparametrizable
models to their inference form. CPU operations use one thread by default.
The fixed seed improves repeatability; it does not prove identical weights or inputs across PyTorch versions.
These measurements cover prepared-tensor inference. They exclude image loading, preprocessing, and transfers.
They do not measure training or model accuracy.
To compare main with a dependency PR, install each checkout into a separate environment with uv sync --locked.
Use the same Python version, machine, runtime, device, and settings. Run the same version of this script against
both installations, using an absolute path to the script. Save a JSON file from each environment.
The files record the package revision and dirty state, dependency versions, runtime settings, script hash,
and individual trials. Check that the CUDA math settings match before comparing GPU results.
Alternate baseline and candidate commands, and first run main twice to check normal measurement variation.
Compare latency, throughput, and peak RSS per workload; avoid treating small changes within that variation as gains.
CI runs short CPU checks for PyTorch and ONNX and uploads their JSON results. CI timing is advisory.
All arguments are listed by python scripts/eval_latency.py --help.
If you wish to deploy containerized environments, you can use the provided Dockerfile to build a docker image:
make build-apiLooking for a boilerplate to deploy a model from Holocron with a REST API? Thanks to the wonderful FastAPI framework, you can do this easily.
Run your API in a docker container as follows:
make start-apiIn order to stop the container, use make stop-api
Your API is now running on port 8080, with its documentation http://localhost:8080/redoc and requestable routes:
import requests
with open("/path/to/your/img.jpeg", "rb") as f:
data = f.read()
response = requests.post("http://localhost:8080/classification", files={"file": data}).json()If you wish to cite this project, feel free to use this BibTeX reference:
@software{Fernandez_Holocron_2020,
author = {Fernandez, François-Guillaume},
month = {5},
title = {{Holocron}},
url = {https://github.com/frgfm/Holocron},
year = {2020}
}Any sort of contribution is greatly appreciated!
You can find a short guide in CONTRIBUTING to help grow this project!
Distributed under the Apache 2.0 License. See LICENSE for more information.

