Release 0.6.0
Release Note
Add new model training with imagenet & clip.
-
Introduce the ImageNet1K dataset and use the CLIP model for knowledge distillation
Due to the expansion of the dataset size, the original configuration parameters were no longer sufficient to allow the model to converge smoothly. Therefore, we made some adjustments to the model.
- Settings
- Num of classes: 1,281,833
- Num of epochs: 40
- Num of data per epoch: 2,560,000
- Batch Size: 1024
- Optimizer: AdamW
- Setting:
- squeeze: Conv2d -> Flatten -> Linear
Dataset with CLIP Num_Classes TPR@FPR=1e-4 ROC Indoor X 390,144 0.772 0.9958 ImageNet-1K X 1,281,833 0.813 0.9961 ImageNet-1K V 1,281,833 0.859 0.9982 -
By using ImageNet-1K to expand the number of categories to approximately 1.3 million, the model is provided with a richer variety of visual changes, increasing data diversity and improving performance by 4.1%.
-
Building on the foundation of ImageNet-1K, introducing the CLIP model and conducting knowledge distillation during training can further enhance performance by 4.6% in the comparison benchmark of TPR@FPR=1e-4.
-
TPR@FPR=1e-4: 0.859
FPR 1e-05 1e-04 1e-03 1e-02 1e-01 1 TPR 0.764 0.859 0.926 0.972 0.996 1.0 Threshold 0.756 0.735 0.713 0.684 0.637 0.368
- Settings
Full Changelog: 0.5.0...0.6.0