[ICASSP 2023] ViTASD: Robust Vision Transformer Baselines for Autism Spectrum Disorder Facial Diagnosis

Official PyTorch Implementation

Shenzhen Children's Hospital
New York University

Abstract

Autism spectrum disorder (ASD) is a lifelong neurodevelopmental disorder with very high prevalence around the world. Research progress in the field of ASD facial analysis in pediatric patients has been hindered due to a lack of well-established baselines. In this paper, we propose the use of the Vision Transformer (ViT) for the computational analysis of pediatric ASD. The presented model, known as ViTASD, distills knowledge from large facial expression datasets and offers model structure transferability. Specifically, ViTASD employs a vanilla ViT to extract features from patients' face images and adopts a lightweight decoder with a Gaussian Process layer to enhance the robustness for ASD analysis. Extensive experiments conducted on standard ASD facial analysis benchmarks show that our method outperforms all of the representative approaches in ASD facial analysis, while the ViTASD-L achieves a new state-of-the-art.

Dataset

Publicly available datasets were analyzed in this study. The original data page can be found at: Kaggle. The author update the dataset to a new Google Drive

Other useful dataset for computer vision in Autism Spectrum Disorder detection:

DE-ENIGMA Dataset
Saliency4ASD dataset

We will expand the research for these datasets in the future. And we are also trying to build a new benchmark for ASD facial diagnosis using many new datasets in Shenzhen's children. Any news for this benchmark will be updated to this Github repo until we publish the competition. This project will create a completely non-profit platform for ASD early intervention around the world.

Model

Pre-trained in AffectNet Dataset

python train_affectnet.py fit -c ./configs/config_affectnet_base.yaml

python train_affectnet.py fit -c ./configs/config_affectnet_large.yaml

Training

python train.py fit -c ./configs/config_vitasd_small.yaml

python train.py fit -c ./configs/config_vitasd_base.yaml

python train.py fit -c ./configs/config_vitasd_large.yaml

Monitoring the training ('X' is S, B, or L)

tensorboard --logdir=./lightning_logs/ViTASD-'X'

Evaluation

pending

Name		Name	Last commit message	Last commit date
Latest commit History 13 Commits
configs		configs
datasets		datasets
lib		lib
lightning_logs		lightning_logs
models		models
runs		runs
tools		tools
.gitignore		.gitignore
LICENSE		LICENSE
README.md		README.md
eval.py		eval.py
requirement.txt		requirement.txt
train.py		train.py
train_affectnet.py		train_affectnet.py
visualization_attention.py		visualization_attention.py

License

PediaMedAI/ViTASD

Folders and files

Latest commit

History

Repository files navigation

[ICASSP 2023] ViTASD: Robust Vision Transformer Baselines for Autism Spectrum Disorder Facial Diagnosis

Official PyTorch Implementation

Abstract

Dataset

Model

Pre-trained in AffectNet Dataset

Training

Monitoring the training ('X' is S, B, or L)

Evaluation

About

Topics

Resources

License

Stars

Watchers

Forks

Languages